Functional Weave
Code in TypeScript

legal.conflict-name-match Unreviewed

Fuzzy-match a name against existing clients and parties for a conflict check: normalised, token-sorted Jaro-Winkler.

1.0.1 · published 2026-10-03 by charlie · Anterra

Pinned by 18 tests, run in TypeScript, Python and Rust.

Unreviewed. This capability’s implementations agree in every language and pass its published test vectors, which were worked out from the official sources cited. But no qualified solicitor has yet checked those vectors, or confirmed that the capability covers the cases it claims. Treat it as a draft. Do not use it for real people, money or decisions without your own expert review. Once a qualified reviewer signs off, this notice is replaced with their name, qualification and the date. Each new version needs fresh sign-off.

Not professional advice. This capability calculates legal figures from published rules. It is a software component for developers, not legal advice. Rules change and every rate here has an effective date. Check that the dates cover your case. Verify results against the official sources listed in its README, and have a solicitor review how you use it, before anyone relies on the output. Provided “as is” under its licence, without warranty.

What it does

Before a firm takes on a client it checks the client and the other parties against everyone it has acted for or against. Exact matching misses "Smith, John" against "John Smith", "Acme Holdings Ltd" against "ACME HOLDINGS LIMITED" and a transposed "Marhta"; this returns every candidate whose name is similar enough, best first, so a person can review them.

## How a name is compared

For example

  • conflictNameMatch(Martha, marhta, Dwayne, MARTHA LTD, Jones, 90%) → ×2 Winkler's MARTHA/MARHTA: 0.9611; a company form and case do not stop an exact match
  • conflictNameMatch(Dwayne, Duane, 0%) → ×1 Winkler's DWAYNE/DUANE: exactly 0.84
  • conflictNameMatch(Dixon, Dicksonx, 81.33%) → ×1 Winkler's DIXON/DICKSONX: 0.8133, included at a threshold of 8133

The function

The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.

export function conflictNameMatch(name: string, candidates: readonly string[], thresholdBasisPoints: number): readonly NameMatch[]
namestringthe new client, counterparty or other name to check
candidatesstring[]existing names: clients, parties, former clients
thresholdBasisPointsintthe least similarity reported, 0 to 10000; 8500 is a common starting point, 0 reports every candidate
returnsNameMatch[]

The type it declares, generated into your project

/** One candidate at or above the threshold. */
export interface NameMatch {
  /** its position in candidates */
  readonly index: number;
  /** as given */
  readonly candidate: string;
  /** the normalised, token-sorted form that was compared */
  readonly matchKey: string;
  /** Jaro-Winkler similarity of the match keys, 0 to 10000 */
  readonly score: number;
}

Your code names it in one line, in the file that uses it

import { conflictNameMatch } from "#fune/legal.conflict-name-match@^1";
impl/typescript.ts · 100 lines · open · raw

Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.

import { roundDiv } from "./math_round_div.ts";  ← from math.round-div ^1.0.0 · built alongside by fune
import { normaliseName } from "./text_normalise_name.ts";  ← from text.normalise-name ^1.0.0 · built alongside by fune
import { type NameMatch } from "./legal_conflict_name_match_types.ts";

const COMPANY_FORMS = new Set(["ltd", "limited", "plc", "llp", "llc", "inc"]);
const DROPPED = new Set(["'", "’", "."]);
const MAX_KEY = 500;

function lowerOne(ch: string): string {
  // One-to-one case mappings only, as text.normalise-name, so all three languages agree.
  const mapped = ch.toLowerCase();
  return Array.from(mapped).length === 1 ? mapped : ch;
}

function isAsciiPunctuation(cp: number): boolean {
  return (cp >= 0x21 && cp <= 0x2f) || (cp >= 0x3a && cp <= 0x40) || (cp >= 0x5b && cp <= 0x60) || (cp >= 0x7b && cp <= 0x7e);
}

function compareCodePoints(a: string, b: string): number {
  const x = Array.from(a).map((c) => c.codePointAt(0) as number);
  const y = Array.from(b).map((c) => c.codePointAt(0) as number);
  for (let i = 0; i < Math.min(x.length, y.length); i++) {
    if (x[i] !== y[i]) return x[i] - y[i];
  }
  return x.length - y.length;
}

/** The normalised, punctuation-free, company-form-free, token-sorted key. */
function matchKey(value: string): string {
  let text = "";
  for (const ch of normaliseName(value)) {
    const cp = ch.codePointAt(0) as number;
    if (DROPPED.has(ch)) continue;
    if (ch === "&") text += " and ";
    else if (isAsciiPunctuation(cp) || ch === " ") text += " ";
    else text += lowerOne(ch);
  }
  const tokens = text.split(" ").filter((t) => t !== "" && !COMPANY_FORMS.has(t));
  const key = tokens.sort(compareCodePoints).join(" ");
  if (Array.from(key).length > MAX_KEY) {
    throw new RangeError(`names must be at most ${MAX_KEY} characters after normalising`);
  }
  return key;
}

/** Jaro-Winkler similarity in basis points, from exact integer fractions. */
function jaroWinkler(s1: string[], s2: string[]): number {
  const a = s1.length;
  const b = s2.length;
  if (a === 0 || b === 0) return 0;
  const window = Math.max(0, Math.floor(Math.max(a, b) / 2) - 1);
  const used = new Array<boolean>(b).fill(false);
  const order1: string[] = [];
  const matched1 = new Array<boolean>(a).fill(false);
  for (let i = 0; i < a; i++) {
    for (let j = Math.max(0, i - window); j <= Math.min(b - 1, i + window); j++) {
      if (!used[j] && s1[i] === s2[j]) {
        used[j] = true;
        matched1[i] = true;
        break;
      }
    }
  }
  for (let i = 0; i < a; i++) if (matched1[i]) order1.push(s1[i]);
  const m = order1.length;
  if (m === 0) return 0;
  let k = 0;
  let outOfOrder = 0;
  for (let j = 0; j < b; j++) {
    if (!used[j]) continue;
    if (s2[j] !== order1[k]) outOfOrder += 1;
    k += 1;
  }
  // Jaro = N / D exactly, with t = outOfOrder / 2.
  const n = 2 * m * m * (a + b) + a * b * (2 * m - outOfOrder);
  const d = 6 * a * b * m;
  let prefix = 0;
  while (prefix < 4 && prefix < a && prefix < b && s1[prefix] === s2[prefix]) prefix += 1;
  if (10 * n <= 7 * d) return roundDiv(10000 * n, d, "half-up");
  return roundDiv(10000 * (n * (10 - prefix) + prefix * d), 10 * d, "half-up");
}

/** Candidates similar to a name, best first, for a conflict check. */
export function conflictNameMatch(name: string, candidates: readonly string[], thresholdBasisPoints: number): readonly NameMatch[] {
  if (!Number.isInteger(thresholdBasisPoints) || thresholdBasisPoints < 0 || thresholdBasisPoints > 10000) {
    throw new RangeError(`thresholdBasisPoints must be between 0 and 10000, received ${thresholdBasisPoints}`);
  }
  const key = matchKey(name);
  if (key === "") {
    throw new RangeError("name is empty after normalising");
  }
  const s1 = Array.from(key);
  const matches: NameMatch[] = [];
  candidates.forEach((candidate, index) => {
    const other = matchKey(candidate);
    const score = jaroWinkler(s1, Array.from(other));
    if (score >= thresholdBasisPoints) matches.push({ index, candidate, matchKey: other, score });
  });
  return matches.sort((x, y) => y.score - x.score || x.index - y.index);
}

Install

fune build

With that line in your source, in a TypeScript project (language typescript in fune.project), fune build resolves it and its 2 dependencies, pins them in fune.lock, downloads only the TypeScript package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:

fune add legal.conflict-name-match
Download for TypeScript legal.conflict-name-match-1.0.1-typescript.fune · 17,629 bytes sha256 12899b70615105b8198f2eeed229330b6ca895cbc85eede9a454fb56ce2cdd08

The manifest, vectors and README with only the TypeScript implementation. Install it without the registry with fune add ./legal.conflict-name-match-1.0.1-typescript.fune, or fetch it from a terminal with fune pull legal.conflict-name-match@1.0.1:typescript.

The whole function, every language, is one file too: legal.conflict-name-match-1.0.1.fune, 26,132 bytes, sha256 c27233e64cee06c1f59edb65c614a9919cbf8fb9f4683e8bcda553f267fa09b8. It installs into a project of any language.

Customise it in your app

The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.

before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.

// fune: before legal.conflict-name-match

after — your function gets the result and the arguments, and returns the final result.

// fune: after legal.conflict-name-match

replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.

// fune: replace math.round-div in legal.conflict-name-match
// fune: replace text.normalise-name in legal.conflict-name-match

step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show legal.conflict-name-match --steps.

// fune: step legal.conflict-name-match after <n|label>

Tests

A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.

CaseArgumentsExpected
Winkler's MARTHA/MARHTA: 0.9611; a company form and case do not stop an exact match Martha, marhta, Dwayne, MARTHA LTD, Jones, 90% → ×2
Winkler's DWAYNE/DUANE: exactly 0.84 Dwayne, Duane, 0% → ×1
Winkler's DIXON/DICKSONX: 0.8133, included at a threshold of 8133 Dixon, Dicksonx, 81.33% → ×1
and excluded one basis point higher Dixon, Dicksonx, 81.34% →
word order does not matter: 'Smith, John' is 'john smith' John Smith, Smith, John, SMITH John, 100% → ×2
Ltd, Limited and PLC are dropped; equal scores stay in input order Acme Holdings Limited, Acme Holdings PLC, ACME HOLDINGS LTD., Acme Holdings Ltd, 100% → ×3
'&' is 'and', and LLP is dropped Smith & Jones LLP, Smith and Jones, 100% → ×1
apostrophes and full stops vanish: O'Neill is ONeill O'Neill, ONEILL, O’Neill, 100% → ×2
below Jaro 0.7 there is no prefix boost: Amir/Amos is 6667, not 7333 Amir, Amos, 0% → ×1
a dropped letter: John/Jon 0.9333 John, Jon, 90% → ×1
Show the other 8 tests
CaseArgumentsExpected
an empty or form-only candidate scores 0, reported only at threshold 0 John, , Ltd, Jon, 0% → ×3
accented letters are kept, not folded: José/Jose 0.8833 José, Jose, 80% → ×1
best first: higher scores lead regardless of input order Martha, Dwayne, marhta, Martha, 90% → ×2
no candidates, no matches Anyone, , 0% →
a name that normalises to nothing is an error Ltd., Acme, 85% → error: name is empty after normalising
a threshold over 10000 is an error Acme, Acme, 100.01% → error: thresholdBasisPoints must be between 0 and 10000
a negative threshold is an error Acme, Acme, -0.01% → error: thresholdBasisPoints must be between 0 and 10000
a name over 500 characters is an error aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa… → error: names must be at most 500 characters after normalising

More from the author

1. **Normalise** with text.normalise-name (trim, collapse Unicode whitespace), then lower-case each character (one-to-one case mappings only, as text.normalise-name does, so all three languages agree). 2. **Punctuation.** Apostrophes (' and ’) and full stops are removed, so "O'Neill" is "oneill" and "J.P." is "jp"; "&" becomes the word "and"; every other ASCII punctuation character becomes a space. Letters outside ASCII are kept as they are: "José" and "Jose" are close, not equal. 3. **Company forms** are dropped as whole words: ltd, limited, plc, llp, llc, inc. "Acme Ltd" and "Acme Limited" are the same key, "acme". 4. **Token sort.** The words are sorted (by code point) and joined with single spaces, so word order does not matter: "Smith, John" and "John Smith" are both "john smith". This is the "token sort" idea of fuzzy matching.

The result for each candidate carries this `matchKey`, so a reviewer can see why two names matched.

## The score

Jaro-Winkler similarity of the two keys, over Unicode code points, as defined by Winkler (W. E. Winkler, "String Comparator Metrics and Enhanced Decision Rules in the Fellegi-Sunter Model of Record Linkage", Proceedings of the Section on Survey Research Methods, American Statistical Association, 1990, pp. 354-359; the definition as set out at https://en.wikipedia.org/wiki/Jaro%E2%80%93Winkler_distance):

- matching characters are equal characters no further apart than floor(max(|s1|, |s2|) / 2) - 1, each used once, scanning s1 left to right; - t is half the number of matched characters that are out of order; - Jaro = (m/|s1| + m/|s2| + (m - t)/m) / 3, or 0 when m is 0; - Winkler adds l x 0.1 x (1 - Jaro) for a common prefix of l characters (at most 4), **only when Jaro is above 0.7**, Winkler's boost threshold.

It is computed exactly as a fraction of integers and rounded half-up once to basis points (10000 = identical), so the three languages give the same score and a threshold means the same thing everywhere. The standard examples: MARTHA/MARHTA 9611, DWAYNE/DUANE 8400, DIXON/DICKSONX 8133.

Results are the candidates scoring at least `thresholdBasisPoints`, highest score first, and in input order among equal scores, so the output is deterministic. A candidate that normalises to nothing (an empty string, "Ltd") scores 0. A query that normalises to nothing is an error, and so is a key over 500 characters (the exact arithmetic is bounded).

## Limits

A similarity score is a screen, not a decision: it will miss a name that changed on marriage, a trading name, a transliteration ("Mohammed" and "Muhammad" score well; "Ivanov" and "Иванов" do not), and it will flag common surnames. Choosing the threshold is a risk decision for the firm; set it low enough that a person reviews near misses. The company-form list is short and English; "Co", "Company", "Group" and "The" are kept on purpose, because dropping them joins genuinely different names.

## Before you rely on this

**Not professional advice.** This capability calculates legal figures from published rules. It is a software component for developers, not legal advice. Rules change and every rate here has an effective date. Check that the dates cover your case. Verify results against the official sources listed above, and have a solicitor review how you use it, before anyone relies on the output. Provided "as is" under its licence, without warranty.

**Unreviewed.** This capability's implementations agree in every language and pass its published test vectors, which were worked out from the official sources cited. But no qualified solicitor has yet checked those vectors, or confirmed that the capability covers the cases it claims. Treat it as a draft. Do not use it for real people, money or decisions without your own expert review. Once a qualified reviewer signs off, this notice is replaced with their name, qualification and the date. Each new version needs fresh sign-off.

1.0.1 marks it unreviewed. The code and the tests are unchanged.

Files

PathBytes
README.md4,454
impl/python.py3,489
impl/rust.rs4,659
impl/typescript.ts4,011
vectors.json5,569