Functional Weave
Code in TypeScript

legal.conflict-name-match

Fuzzy-match a name against existing clients and parties for a conflict check: normalised, token-sorted Jaro-Winkler.

1.0.0 (not the latest) · published 2026-10-03 by charlie · Anterra

Pinned by 18 tests, run in TypeScript, Python and Rust.

Not professional advice. This capability calculates legal figures from published rules. It is a software component for developers, not legal advice. Rules change and every rate here has an effective date. Check that the dates cover your case. Verify results against the official sources listed in its README, and have a solicitor review how you use it, before anyone relies on the output. Provided “as is” under its licence, without warranty.

What it does

Before a firm takes on a client it checks the client and the other parties against everyone it has acted for or against. Exact matching misses "Smith, John" against "John Smith", "Acme Holdings Ltd" against "ACME HOLDINGS LIMITED" and a transposed "Marhta"; this returns every candidate whose name is similar enough, best first, so a person can review them.

## How a name is compared

For example

  • conflictNameMatch(Martha, marhta, Dwayne, MARTHA LTD, Jones, 90%) → ×2 Winkler's MARTHA/MARHTA: 0.9611; a company form and case do not stop an exact match
  • conflictNameMatch(Dwayne, Duane, 0%) → ×1 Winkler's DWAYNE/DUANE: exactly 0.84
  • conflictNameMatch(Dixon, Dicksonx, 81.33%) → ×1 Winkler's DIXON/DICKSONX: 0.8133, included at a threshold of 8133

The function

The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.

export function conflictNameMatch(name: string, candidates: readonly string[], thresholdBasisPoints: number): readonly NameMatch[]
namestringthe new client, counterparty or other name to check
candidatesstring[]existing names: clients, parties, former clients
thresholdBasisPointsintthe least similarity reported, 0 to 10000; 8500 is a common starting point, 0 reports every candidate
returnsNameMatch[]

The type it declares, generated into your project

/** One candidate at or above the threshold. */
export interface NameMatch {
  /** its position in candidates */
  readonly index: number;
  /** as given */
  readonly candidate: string;
  /** the normalised, token-sorted form that was compared */
  readonly matchKey: string;
  /** Jaro-Winkler similarity of the match keys, 0 to 10000 */
  readonly score: number;
}

Your code names it in one line, in the file that uses it

import { conflictNameMatch } from "#fune/legal.conflict-name-match@^1";
impl/typescript.ts · 100 lines · open · raw

Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.

import { roundDiv } from "./math_round_div.ts";  ← from math.round-div ^1.0.0 · built alongside by fune
import { normaliseName } from "./text_normalise_name.ts";  ← from text.normalise-name ^1.0.0 · built alongside by fune
import { type NameMatch } from "./legal_conflict_name_match_types.ts";

const COMPANY_FORMS = new Set(["ltd", "limited", "plc", "llp", "llc", "inc"]);
const DROPPED = new Set(["'", "’", "."]);
const MAX_KEY = 500;

function lowerOne(ch: string): string {
  // One-to-one case mappings only, as text.normalise-name, so all three languages agree.
  const mapped = ch.toLowerCase();
  return Array.from(mapped).length === 1 ? mapped : ch;
}

function isAsciiPunctuation(cp: number): boolean {
  return (cp >= 0x21 && cp <= 0x2f) || (cp >= 0x3a && cp <= 0x40) || (cp >= 0x5b && cp <= 0x60) || (cp >= 0x7b && cp <= 0x7e);
}

function compareCodePoints(a: string, b: string): number {
  const x = Array.from(a).map((c) => c.codePointAt(0) as number);
  const y = Array.from(b).map((c) => c.codePointAt(0) as number);
  for (let i = 0; i < Math.min(x.length, y.length); i++) {
    if (x[i] !== y[i]) return x[i] - y[i];
  }
  return x.length - y.length;
}

/** The normalised, punctuation-free, company-form-free, token-sorted key. */
function matchKey(value: string): string {
  let text = "";
  for (const ch of normaliseName(value)) {
    const cp = ch.codePointAt(0) as number;
    if (DROPPED.has(ch)) continue;
    if (ch === "&") text += " and ";
    else if (isAsciiPunctuation(cp) || ch === " ") text += " ";
    else text += lowerOne(ch);
  }
  const tokens = text.split(" ").filter((t) => t !== "" && !COMPANY_FORMS.has(t));
  const key = tokens.sort(compareCodePoints).join(" ");
  if (Array.from(key).length > MAX_KEY) {
    throw new RangeError(`names must be at most ${MAX_KEY} characters after normalising`);
  }
  return key;
}

/** Jaro-Winkler similarity in basis points, from exact integer fractions. */
function jaroWinkler(s1: string[], s2: string[]): number {
  const a = s1.length;
  const b = s2.length;
  if (a === 0 || b === 0) return 0;
  const window = Math.max(0, Math.floor(Math.max(a, b) / 2) - 1);
  const used = new Array<boolean>(b).fill(false);
  const order1: string[] = [];
  const matched1 = new Array<boolean>(a).fill(false);
  for (let i = 0; i < a; i++) {
    for (let j = Math.max(0, i - window); j <= Math.min(b - 1, i + window); j++) {
      if (!used[j] && s1[i] === s2[j]) {
        used[j] = true;
        matched1[i] = true;
        break;
      }
    }
  }
  for (let i = 0; i < a; i++) if (matched1[i]) order1.push(s1[i]);
  const m = order1.length;
  if (m === 0) return 0;
  let k = 0;
  let outOfOrder = 0;
  for (let j = 0; j < b; j++) {
    if (!used[j]) continue;
    if (s2[j] !== order1[k]) outOfOrder += 1;
    k += 1;
  }
  // Jaro = N / D exactly, with t = outOfOrder / 2.
  const n = 2 * m * m * (a + b) + a * b * (2 * m - outOfOrder);
  const d = 6 * a * b * m;
  let prefix = 0;
  while (prefix < 4 && prefix < a && prefix < b && s1[prefix] === s2[prefix]) prefix += 1;
  if (10 * n <= 7 * d) return roundDiv(10000 * n, d, "half-up");
  return roundDiv(10000 * (n * (10 - prefix) + prefix * d), 10 * d, "half-up");
}

/** Candidates similar to a name, best first, for a conflict check. */
export function conflictNameMatch(name: string, candidates: readonly string[], thresholdBasisPoints: number): readonly NameMatch[] {
  if (!Number.isInteger(thresholdBasisPoints) || thresholdBasisPoints < 0 || thresholdBasisPoints > 10000) {
    throw new RangeError(`thresholdBasisPoints must be between 0 and 10000, received ${thresholdBasisPoints}`);
  }
  const key = matchKey(name);
  if (key === "") {
    throw new RangeError("name is empty after normalising");
  }
  const s1 = Array.from(key);
  const matches: NameMatch[] = [];
  candidates.forEach((candidate, index) => {
    const other = matchKey(candidate);
    const score = jaroWinkler(s1, Array.from(other));
    if (score >= thresholdBasisPoints) matches.push({ index, candidate, matchKey: other, score });
  });
  return matches.sort((x, y) => y.score - x.score || x.index - y.index);
}

Install

fune build

With that line in your source, in a TypeScript project (language typescript in fune.project), fune build resolves it and its 2 dependencies, pins them in fune.lock, downloads only the TypeScript package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:

fune add legal.conflict-name-match
Download for TypeScript legal.conflict-name-match-1.0.0-typescript.fune · 16,523 bytes sha256 437f866d6639da32a041d937b8984a1892359b128d777e41e365b7a7df9973ba

The manifest, vectors and README with only the TypeScript implementation. Install it without the registry with fune add ./legal.conflict-name-match-1.0.0-typescript.fune, or fetch it from a terminal with fune pull legal.conflict-name-match@1.0.0:typescript.

The whole function, every language, is one file too: legal.conflict-name-match-1.0.0.fune, 25,026 bytes, sha256 630c751be9d9b1464d257483a124961fb1d9b35a735a0f28d73645e789fbb097. It installs into a project of any language.

Customise it in your app

The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.

before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.

// fune: before legal.conflict-name-match

after — your function gets the result and the arguments, and returns the final result.

// fune: after legal.conflict-name-match

replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.

// fune: replace math.round-div in legal.conflict-name-match
// fune: replace text.normalise-name in legal.conflict-name-match

step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show legal.conflict-name-match --steps.

// fune: step legal.conflict-name-match after <n|label>

Tests

A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.

CaseArgumentsExpected
Winkler's MARTHA/MARHTA: 0.9611; a company form and case do not stop an exact match Martha, marhta, Dwayne, MARTHA LTD, Jones, 90% → ×2
Winkler's DWAYNE/DUANE: exactly 0.84 Dwayne, Duane, 0% → ×1
Winkler's DIXON/DICKSONX: 0.8133, included at a threshold of 8133 Dixon, Dicksonx, 81.33% → ×1
and excluded one basis point higher Dixon, Dicksonx, 81.34% →
word order does not matter: 'Smith, John' is 'john smith' John Smith, Smith, John, SMITH John, 100% → ×2
Ltd, Limited and PLC are dropped; equal scores stay in input order Acme Holdings Limited, Acme Holdings PLC, ACME HOLDINGS LTD., Acme Holdings Ltd, 100% → ×3
'&' is 'and', and LLP is dropped Smith & Jones LLP, Smith and Jones, 100% → ×1
apostrophes and full stops vanish: O'Neill is ONeill O'Neill, ONEILL, O’Neill, 100% → ×2
below Jaro 0.7 there is no prefix boost: Amir/Amos is 6667, not 7333 Amir, Amos, 0% → ×1
a dropped letter: John/Jon 0.9333 John, Jon, 90% → ×1
Show the other 8 tests
CaseArgumentsExpected
an empty or form-only candidate scores 0, reported only at threshold 0 John, , Ltd, Jon, 0% → ×3
accented letters are kept, not folded: José/Jose 0.8833 José, Jose, 80% → ×1
best first: higher scores lead regardless of input order Martha, Dwayne, marhta, Martha, 90% → ×2
no candidates, no matches Anyone, , 0% →
a name that normalises to nothing is an error Ltd., Acme, 85% → error: name is empty after normalising
a threshold over 10000 is an error Acme, Acme, 100.01% → error: thresholdBasisPoints must be between 0 and 10000
a negative threshold is an error Acme, Acme, -0.01% → error: thresholdBasisPoints must be between 0 and 10000
a name over 500 characters is an error aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa… → error: names must be at most 500 characters after normalising

More from the author

1. **Normalise** with text.normalise-name (trim, collapse Unicode whitespace), then lower-case each character (one-to-one case mappings only, as text.normalise-name does, so all three languages agree). 2. **Punctuation.** Apostrophes (' and ’) and full stops are removed, so "O'Neill" is "oneill" and "J.P." is "jp"; "&" becomes the word "and"; every other ASCII punctuation character becomes a space. Letters outside ASCII are kept as they are: "José" and "Jose" are close, not equal. 3. **Company forms** are dropped as whole words: ltd, limited, plc, llp, llc, inc. "Acme Ltd" and "Acme Limited" are the same key, "acme". 4. **Token sort.** The words are sorted (by code point) and joined with single spaces, so word order does not matter: "Smith, John" and "John Smith" are both "john smith". This is the "token sort" idea of fuzzy matching.

The result for each candidate carries this `matchKey`, so a reviewer can see why two names matched.

## The score

Jaro-Winkler similarity of the two keys, over Unicode code points, as defined by Winkler (W. E. Winkler, "String Comparator Metrics and Enhanced Decision Rules in the Fellegi-Sunter Model of Record Linkage", Proceedings of the Section on Survey Research Methods, American Statistical Association, 1990, pp. 354-359; the definition as set out at https://en.wikipedia.org/wiki/Jaro%E2%80%93Winkler_distance):

- matching characters are equal characters no further apart than floor(max(|s1|, |s2|) / 2) - 1, each used once, scanning s1 left to right; - t is half the number of matched characters that are out of order; - Jaro = (m/|s1| + m/|s2| + (m - t)/m) / 3, or 0 when m is 0; - Winkler adds l x 0.1 x (1 - Jaro) for a common prefix of l characters (at most 4), **only when Jaro is above 0.7**, Winkler's boost threshold.

It is computed exactly as a fraction of integers and rounded half-up once to basis points (10000 = identical), so the three languages give the same score and a threshold means the same thing everywhere. The standard examples: MARTHA/MARHTA 9611, DWAYNE/DUANE 8400, DIXON/DICKSONX 8133.

Results are the candidates scoring at least `thresholdBasisPoints`, highest score first, and in input order among equal scores, so the output is deterministic. A candidate that normalises to nothing (an empty string, "Ltd") scores 0. A query that normalises to nothing is an error, and so is a key over 500 characters (the exact arithmetic is bounded).

## Limits

A similarity score is a screen, not a decision: it will miss a name that changed on marriage, a trading name, a transliteration ("Mohammed" and "Muhammad" score well; "Ivanov" and "Иванов" do not), and it will flag common surnames. Choosing the threshold is a risk decision for the firm; set it low enough that a person reviews near misses. The company-form list is short and English; "Co", "Company", "Group" and "The" are kept on purpose, because dropping them joins genuinely different names.

Files

PathBytes
README.md3,386
impl/python.py3,489
impl/rust.rs4,659
impl/typescript.ts4,011
vectors.json5,569