Functional Weave
Code in Python

legal.conflict-name-match Unreviewed

Fuzzy-match a name against existing clients and parties for a conflict check: normalised, token-sorted Jaro-Winkler.

1.0.1 · published 2026-10-03 by charlie · Anterra

Pinned by 18 tests, run in TypeScript, Python and Rust.

Unreviewed. This capability’s implementations agree in every language and pass its published test vectors, which were worked out from the official sources cited. But no qualified solicitor has yet checked those vectors, or confirmed that the capability covers the cases it claims. Treat it as a draft. Do not use it for real people, money or decisions without your own expert review. Once a qualified reviewer signs off, this notice is replaced with their name, qualification and the date. Each new version needs fresh sign-off.

Not professional advice. This capability calculates legal figures from published rules. It is a software component for developers, not legal advice. Rules change and every rate here has an effective date. Check that the dates cover your case. Verify results against the official sources listed in its README, and have a solicitor review how you use it, before anyone relies on the output. Provided “as is” under its licence, without warranty.

What it does

Before a firm takes on a client it checks the client and the other parties against everyone it has acted for or against. Exact matching misses "Smith, John" against "John Smith", "Acme Holdings Ltd" against "ACME HOLDINGS LIMITED" and a transposed "Marhta"; this returns every candidate whose name is similar enough, best first, so a person can review them.

## How a name is compared

For example

  • conflict_name_match(Martha, marhta, Dwayne, MARTHA LTD, Jones, 90%) → ×2 Winkler's MARTHA/MARHTA: 0.9611; a company form and case do not stop an exact match
  • conflict_name_match(Dwayne, Duane, 0%) → ×1 Winkler's DWAYNE/DUANE: exactly 0.84
  • conflict_name_match(Dixon, Dicksonx, 81.33%) → ×1 Winkler's DIXON/DICKSONX: 0.8133, included at a threshold of 8133

The function

The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.

def conflict_name_match(name: str, candidates: Sequence[str], threshold_basis_points: int) -> List[NameMatch]
namestringthe new client, counterparty or other name to check
candidatesstring[]existing names: clients, parties, former clients
threshold_basis_pointsintthe least similarity reported, 0 to 10000; 8500 is a common starting point, 0 reports every candidate
returnsNameMatch[]

The type it declares, generated into your project

@dataclass(frozen=True)
class NameMatch:
    """One candidate at or above the threshold."""

    #: its position in candidates
    index: int
    #: as given
    candidate: str
    #: the normalised, token-sorted form that was compared
    match_key: str
    #: Jaro-Winkler similarity of the match keys, 0 to 10000
    score: int

Your code names it in one line, in the file that uses it

from fune.legal.conflict_name_match import conflict_name_match  # legal.conflict-name-match@^1
impl/python.py · 97 lines · open · raw

Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.

from typing import List, Sequence

from .legal_conflict_name_match_types import NameMatch
from .math_round_div import round_div  ← from math.round-div ^1.0.0 · built alongside by fune
from .text_normalise_name import normalise_name  ← from text.normalise-name ^1.0.0 · built alongside by fune

_COMPANY_FORMS = frozenset(["ltd", "limited", "plc", "llp", "llc", "inc"])
_DROPPED = frozenset(["'", "’", "."])
_MAX_KEY = 500


def _lower_one(ch: str) -> str:
    # One-to-one case mappings only, as text.normalise-name, so all three languages agree.
    mapped = ch.lower()
    return mapped if len(mapped) == 1 else ch


def _is_ascii_punctuation(cp: int) -> bool:
    return 0x21 <= cp <= 0x2F or 0x3A <= cp <= 0x40 or 0x5B <= cp <= 0x60 or 0x7B <= cp <= 0x7E


def _match_key(value: str) -> str:
    """The normalised, punctuation-free, company-form-free, token-sorted key."""
    parts: List[str] = []
    for ch in normalise_name(value):
        if ch in _DROPPED:
            continue
        if ch == "&":
            parts.append(" and ")
        elif _is_ascii_punctuation(ord(ch)) or ch == " ":
            parts.append(" ")
        else:
            parts.append(_lower_one(ch))
    tokens = [t for t in "".join(parts).split(" ") if t != "" and t not in _COMPANY_FORMS]
    key = " ".join(sorted(tokens))
    if len(key) > _MAX_KEY:
        raise ValueError("names must be at most %d characters after normalising" % (_MAX_KEY,))
    return key


def _jaro_winkler(s1: str, s2: str) -> int:
    """Jaro-Winkler similarity in basis points, from exact integer fractions."""
    a = len(s1)
    b = len(s2)
    if a == 0 or b == 0:
        return 0
    window = max(0, max(a, b) // 2 - 1)
    used = [False] * b
    order1: List[str] = []
    for i in range(a):
        for j in range(max(0, i - window), min(b - 1, i + window) + 1):
            if not used[j] and s1[i] == s2[j]:
                used[j] = True
                order1.append(s1[i])
                break
    m = len(order1)
    if m == 0:
        return 0
    k = 0
    out_of_order = 0
    for j in range(b):
        if not used[j]:
            continue
        if s2[j] != order1[k]:
            out_of_order += 1
        k += 1
    # Jaro = n / d exactly, with t = out_of_order / 2.
    n = 2 * m * m * (a + b) + a * b * (2 * m - out_of_order)
    d = 6 * a * b * m
    prefix = 0
    while prefix < 4 and prefix < a and prefix < b and s1[prefix] == s2[prefix]:
        prefix += 1
    if 10 * n <= 7 * d:
        return round_div(10000 * n, d, "half-up")
    return round_div(10000 * (n * (10 - prefix) + prefix * d), 10 * d, "half-up")


def conflict_name_match(name: str, candidates: Sequence[str], threshold_basis_points: int) -> List[NameMatch]:
    """Candidates similar to a name, best first, for a conflict check."""
    if (
        isinstance(threshold_basis_points, bool)
        or not isinstance(threshold_basis_points, int)
        or threshold_basis_points < 0
        or threshold_basis_points > 10000
    ):
        raise ValueError("thresholdBasisPoints must be between 0 and 10000, received %s" % (threshold_basis_points,))
    key = _match_key(name)
    if key == "":
        raise ValueError("name is empty after normalising")
    matches: List[NameMatch] = []
    for index, candidate in enumerate(candidates):
        other = _match_key(candidate)
        score = _jaro_winkler(key, other)
        if score >= threshold_basis_points:
            matches.append(NameMatch(index=index, candidate=candidate, match_key=other, score=score))
    matches.sort(key=lambda x: (-x.score, x.index))
    return matches

Install

fune build

With that line in your source, in a Python project (language python in fune.project), fune build resolves it and its 2 dependencies, pins them in fune.lock, downloads only the Python package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:

fune add legal.conflict-name-match
Download for Python legal.conflict-name-match-1.0.1-python.fune · 17,112 bytes sha256 060cc31bf2a92ae0e6c705b6e1c679a4ab46afd2c493cae0117db259a17e4a46

The manifest, vectors and README with only the Python implementation. Install it without the registry with fune add ./legal.conflict-name-match-1.0.1-python.fune, or fetch it from a terminal with fune pull legal.conflict-name-match@1.0.1:python.

The whole function, every language, is one file too: legal.conflict-name-match-1.0.1.fune, 26,132 bytes, sha256 c27233e64cee06c1f59edb65c614a9919cbf8fb9f4683e8bcda553f267fa09b8. It installs into a project of any language.

Customise it in your app

The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.

before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.

# fune: before legal.conflict-name-match

after — your function gets the result and the arguments, and returns the final result.

# fune: after legal.conflict-name-match

replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.

# fune: replace math.round-div in legal.conflict-name-match
# fune: replace text.normalise-name in legal.conflict-name-match

step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show legal.conflict-name-match --steps.

# fune: step legal.conflict-name-match after <n|label>

Tests

A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.

CaseArgumentsExpected
Winkler's MARTHA/MARHTA: 0.9611; a company form and case do not stop an exact match Martha, marhta, Dwayne, MARTHA LTD, Jones, 90% → ×2
Winkler's DWAYNE/DUANE: exactly 0.84 Dwayne, Duane, 0% → ×1
Winkler's DIXON/DICKSONX: 0.8133, included at a threshold of 8133 Dixon, Dicksonx, 81.33% → ×1
and excluded one basis point higher Dixon, Dicksonx, 81.34% →
word order does not matter: 'Smith, John' is 'john smith' John Smith, Smith, John, SMITH John, 100% → ×2
Ltd, Limited and PLC are dropped; equal scores stay in input order Acme Holdings Limited, Acme Holdings PLC, ACME HOLDINGS LTD., Acme Holdings Ltd, 100% → ×3
'&' is 'and', and LLP is dropped Smith & Jones LLP, Smith and Jones, 100% → ×1
apostrophes and full stops vanish: O'Neill is ONeill O'Neill, ONEILL, O’Neill, 100% → ×2
below Jaro 0.7 there is no prefix boost: Amir/Amos is 6667, not 7333 Amir, Amos, 0% → ×1
a dropped letter: John/Jon 0.9333 John, Jon, 90% → ×1
Show the other 8 tests
CaseArgumentsExpected
an empty or form-only candidate scores 0, reported only at threshold 0 John, , Ltd, Jon, 0% → ×3
accented letters are kept, not folded: José/Jose 0.8833 José, Jose, 80% → ×1
best first: higher scores lead regardless of input order Martha, Dwayne, marhta, Martha, 90% → ×2
no candidates, no matches Anyone, , 0% →
a name that normalises to nothing is an error Ltd., Acme, 85% → error: name is empty after normalising
a threshold over 10000 is an error Acme, Acme, 100.01% → error: thresholdBasisPoints must be between 0 and 10000
a negative threshold is an error Acme, Acme, -0.01% → error: thresholdBasisPoints must be between 0 and 10000
a name over 500 characters is an error aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa… → error: names must be at most 500 characters after normalising

More from the author

1. **Normalise** with text.normalise-name (trim, collapse Unicode whitespace), then lower-case each character (one-to-one case mappings only, as text.normalise-name does, so all three languages agree). 2. **Punctuation.** Apostrophes (' and ’) and full stops are removed, so "O'Neill" is "oneill" and "J.P." is "jp"; "&" becomes the word "and"; every other ASCII punctuation character becomes a space. Letters outside ASCII are kept as they are: "José" and "Jose" are close, not equal. 3. **Company forms** are dropped as whole words: ltd, limited, plc, llp, llc, inc. "Acme Ltd" and "Acme Limited" are the same key, "acme". 4. **Token sort.** The words are sorted (by code point) and joined with single spaces, so word order does not matter: "Smith, John" and "John Smith" are both "john smith". This is the "token sort" idea of fuzzy matching.

The result for each candidate carries this `matchKey`, so a reviewer can see why two names matched.

## The score

Jaro-Winkler similarity of the two keys, over Unicode code points, as defined by Winkler (W. E. Winkler, "String Comparator Metrics and Enhanced Decision Rules in the Fellegi-Sunter Model of Record Linkage", Proceedings of the Section on Survey Research Methods, American Statistical Association, 1990, pp. 354-359; the definition as set out at https://en.wikipedia.org/wiki/Jaro%E2%80%93Winkler_distance):

- matching characters are equal characters no further apart than floor(max(|s1|, |s2|) / 2) - 1, each used once, scanning s1 left to right; - t is half the number of matched characters that are out of order; - Jaro = (m/|s1| + m/|s2| + (m - t)/m) / 3, or 0 when m is 0; - Winkler adds l x 0.1 x (1 - Jaro) for a common prefix of l characters (at most 4), **only when Jaro is above 0.7**, Winkler's boost threshold.

It is computed exactly as a fraction of integers and rounded half-up once to basis points (10000 = identical), so the three languages give the same score and a threshold means the same thing everywhere. The standard examples: MARTHA/MARHTA 9611, DWAYNE/DUANE 8400, DIXON/DICKSONX 8133.

Results are the candidates scoring at least `thresholdBasisPoints`, highest score first, and in input order among equal scores, so the output is deterministic. A candidate that normalises to nothing (an empty string, "Ltd") scores 0. A query that normalises to nothing is an error, and so is a key over 500 characters (the exact arithmetic is bounded).

## Limits

A similarity score is a screen, not a decision: it will miss a name that changed on marriage, a trading name, a transliteration ("Mohammed" and "Muhammad" score well; "Ivanov" and "Иванов" do not), and it will flag common surnames. Choosing the threshold is a risk decision for the firm; set it low enough that a person reviews near misses. The company-form list is short and English; "Co", "Company", "Group" and "The" are kept on purpose, because dropping them joins genuinely different names.

## Before you rely on this

**Not professional advice.** This capability calculates legal figures from published rules. It is a software component for developers, not legal advice. Rules change and every rate here has an effective date. Check that the dates cover your case. Verify results against the official sources listed above, and have a solicitor review how you use it, before anyone relies on the output. Provided "as is" under its licence, without warranty.

**Unreviewed.** This capability's implementations agree in every language and pass its published test vectors, which were worked out from the official sources cited. But no qualified solicitor has yet checked those vectors, or confirmed that the capability covers the cases it claims. Treat it as a draft. Do not use it for real people, money or decisions without your own expert review. Once a qualified reviewer signs off, this notice is replaced with their name, qualification and the date. Each new version needs fresh sign-off.

1.0.1 marks it unreviewed. The code and the tests are unchanged.

Files

PathBytes
README.md4,454
impl/python.py3,489
impl/rust.rs4,659
impl/typescript.ts4,011
vectors.json5,569