legal.conflict-name-match
Fuzzy-match a name against existing clients and parties for a conflict check: normalised, token-sorted Jaro-Winkler.
1.0.0 (not the latest) · published 2026-10-03 by charlie · Anterra
Pinned by 18 tests, run in TypeScript, Python and Rust.
Not professional advice. This capability calculates legal figures from published rules. It is a software component for developers, not legal advice. Rules change and every rate here has an effective date. Check that the dates cover your case. Verify results against the official sources listed in its README, and have a solicitor review how you use it, before anyone relies on the output. Provided “as is” under its licence, without warranty.
What it does
Before a firm takes on a client it checks the client and the other parties against everyone it has acted for or against. Exact matching misses "Smith, John" against "John Smith", "Acme Holdings Ltd" against "ACME HOLDINGS LIMITED" and a transposed "Marhta"; this returns every candidate whose name is similar enough, best first, so a person can review them.
## How a name is compared
For example
conflict_name_match(Martha, marhta, Dwayne, MARTHA LTD, Jones, 90%)→ ×2 Winkler's MARTHA/MARHTA: 0.9611; a company form and case do not stop an exact matchconflict_name_match(Dwayne, Duane, 0%)→ ×1 Winkler's DWAYNE/DUANE: exactly 0.84conflict_name_match(Dixon, Dicksonx, 81.33%)→ ×1 Winkler's DIXON/DICKSONX: 0.8133, included at a threshold of 8133
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
pub fn conflict_name_match(name: &str, candidates: &[String], threshold_basis_points: i64) -> Vec<NameMatch>
| name | string | the new client, counterparty or other name to check |
| candidates | string[] | existing names: clients, parties, former clients |
| threshold_basis_points | int | the least similarity reported, 0 to 10000; 8500 is a common starting point, 0 reports every candidate |
| returns | NameMatch[] |
The type it declares, generated into your project
/// One candidate at or above the threshold.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct NameMatch {
/// its position in candidates
pub index: i64,
/// as given
pub candidate: String,
/// the normalised, token-sorted form that was compared
pub match_key: String,
/// Jaro-Winkler similarity of the match keys, 0 to 10000
pub score: i64,
}
Your code names it in one line, in the file that uses it
fune!(legal.conflict-name-match@^1); // then call conflict_name_match(…)
Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.
use super::funejson::Value; ← the fune runtime: the JSON value the test vectors use; fune build keeps it only where a signature takes one
use super::math_round_div::round_div; ← from math.round-div ^1.0.0 · built alongside by fune
use super::text_normalise_name::normalise_name; ← from text.normalise-name ^1.0.0 · built alongside by fune
const COMPANY_FORMS: [&str; 6] = ["ltd", "limited", "plc", "llp", "llc", "inc"];
const MAX_KEY: usize = 500;
fn lower_one(ch: char) -> char {
// One-to-one case mappings only, as text.normalise-name, so all three languages agree.
let mapped: Vec<char> = ch.to_lowercase().collect();
if mapped.len() == 1 {
mapped[0]
} else {
ch
}
}
fn is_ascii_punctuation(cp: u32) -> bool {
(0x21..=0x2f).contains(&cp) || (0x3a..=0x40).contains(&cp) || (0x5b..=0x60).contains(&cp) || (0x7b..=0x7e).contains(&cp)
}
/// The normalised, punctuation-free, company-form-free, token-sorted key.
fn match_key(value: &str) -> String {
let mut text = String::new();
for ch in normalise_name(value).chars() {
if ch == '\'' || ch == '\u{2019}' || ch == '.' {
continue;
}
if ch == '&' {
text.push_str(" and ");
} else if is_ascii_punctuation(ch as u32) || ch == ' ' {
text.push(' ');
} else {
text.push(lower_one(ch));
}
}
let mut tokens: Vec<&str> = text.split(' ').filter(|t| !t.is_empty() && !COMPANY_FORMS.contains(t)).collect();
// Byte order of UTF-8 is code point order.
tokens.sort();
let key = tokens.join(" ");
if key.chars().count() > MAX_KEY {
panic!("names must be at most {} characters after normalising", MAX_KEY);
}
key
}
/// Jaro-Winkler similarity in basis points, from exact integer fractions.
fn jaro_winkler(s1: &[char], s2: &[char]) -> i64 {
let a = s1.len();
let b = s2.len();
if a == 0 || b == 0 {
return 0;
}
let window = (a.max(b) / 2).saturating_sub(1);
let mut used = vec![false; b];
let mut order1: Vec<char> = Vec::new();
for i in 0..a {
let lo = i.saturating_sub(window);
let hi = (b - 1).min(i + window);
for j in lo..=hi {
if !used[j] && s1[i] == s2[j] {
used[j] = true;
order1.push(s1[i]);
break;
}
}
}
let m = order1.len() as i64;
if m == 0 {
return 0;
}
let mut k = 0;
let mut out_of_order: i64 = 0;
for j in 0..b {
if !used[j] {
continue;
}
if s2[j] != order1[k] {
out_of_order += 1;
}
k += 1;
}
let (a, b) = (a as i64, b as i64);
// Jaro = n / d exactly, with t = out_of_order / 2.
let n = 2 * m * m * (a + b) + a * b * (2 * m - out_of_order);
let d = 6 * a * b * m;
let mut prefix: i64 = 0;
while prefix < 4 && prefix < a && prefix < b && s1[prefix as usize] == s2[prefix as usize] {
prefix += 1;
}
if 10 * n <= 7 * d {
return round_div(10000 * n, d, "half-up");
}
round_div(10000 * (n * (10 - prefix) + prefix * d), 10 * d, "half-up")
}
/// Candidates similar to a name, best first, for a conflict check.
///
/// # Panics
/// Panics on a threshold outside 0..=10000, a name that normalises to nothing,
/// or a key longer than 500 characters.
pub fn conflict_name_match(name: &str, candidates: &[String], threshold_basis_points: i64) -> Vec<NameMatch> {
if !(0..=10000).contains(&threshold_basis_points) {
panic!("thresholdBasisPoints must be between 0 and 10000, received {}", threshold_basis_points);
}
let key = match_key(name);
if key.is_empty() {
panic!("name is empty after normalising");
}
let s1: Vec<char> = key.chars().collect();
let mut matches: Vec<NameMatch> = Vec::new();
for (index, candidate) in candidates.iter().enumerate() {
let other = match_key(candidate);
let s2: Vec<char> = other.chars().collect();
let score = jaro_winkler(&s1, &s2);
if score >= threshold_basis_points {
matches.push(NameMatch { index: index as i64, candidate: candidate.clone(), match_key: other, score });
}
}
matches.sort_by(|x, y| y.score.cmp(&x.score).then(x.index.cmp(&y.index)));
matches
}
pub fn name_match_to_value(m: &NameMatch) -> Value {
Value::obj(vec![
("index", Value::Int(m.index)),
("candidate", Value::str(&m.candidate)),
("matchKey", Value::str(&m.match_key)),
("score", Value::Int(m.score)),
])
}
pub fn fune_vector(args: &[Value]) -> Value {
let candidates: Vec<String> = args[1].as_arr().iter().map(|v| v.as_str().to_string()).collect();
Value::Arr(conflict_name_match(args[0].as_str(), &candidates, args[2].as_i64()).iter().map(name_match_to_value).collect())
}Install
fune build
With that line in your source, in a Rust project (language rust in fune.project), fune build resolves it and its 2 dependencies, pins them in fune.lock, downloads only the Rust package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. A crate’s build.rs runs it before every compile. Or pin a range in fune.project and build in one step:
fune add legal.conflict-name-match
The manifest, vectors and README with only the Rust implementation. Install it without the registry with fune add ./legal.conflict-name-match-1.0.0-rust.fune, or fetch it from a terminal with fune pull legal.conflict-name-match@1.0.0:rust.
The whole function, every language, is one file too: legal.conflict-name-match-1.0.0.fune, 25,026 bytes, sha256 630c751be9d9b1464d257483a124961fb1d9b35a735a0f28d73645e789fbb097. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
// fune: before legal.conflict-name-match
after — your function gets the result and the arguments, and returns the final result.
// fune: after legal.conflict-name-match
replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.
// fune: replace math.round-div in legal.conflict-name-match
// fune: replace text.normalise-name in legal.conflict-name-match
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show legal.conflict-name-match --steps.
// fune: step legal.conflict-name-match after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| Winkler's MARTHA/MARHTA: 0.9611; a company form and case do not stop an exact match | Martha, marhta, Dwayne, MARTHA LTD, Jones, 90% | → | ×2 |
| Winkler's DWAYNE/DUANE: exactly 0.84 | Dwayne, Duane, 0% | → | ×1 |
| Winkler's DIXON/DICKSONX: 0.8133, included at a threshold of 8133 | Dixon, Dicksonx, 81.33% | → | ×1 |
| and excluded one basis point higher | Dixon, Dicksonx, 81.34% | → | |
| word order does not matter: 'Smith, John' is 'john smith' | John Smith, Smith, John, SMITH John, 100% | → | ×2 |
| Ltd, Limited and PLC are dropped; equal scores stay in input order | Acme Holdings Limited, Acme Holdings PLC, ACME HOLDINGS LTD., Acme Holdings Ltd, 100% | → | ×3 |
| '&' is 'and', and LLP is dropped | Smith & Jones LLP, Smith and Jones, 100% | → | ×1 |
| apostrophes and full stops vanish: O'Neill is ONeill | O'Neill, ONEILL, O’Neill, 100% | → | ×2 |
| below Jaro 0.7 there is no prefix boost: Amir/Amos is 6667, not 7333 | Amir, Amos, 0% | → | ×1 |
| a dropped letter: John/Jon 0.9333 | John, Jon, 90% | → | ×1 |
Show the other 8 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| an empty or form-only candidate scores 0, reported only at threshold 0 | John, , Ltd, Jon, 0% | → | ×3 |
| accented letters are kept, not folded: José/Jose 0.8833 | José, Jose, 80% | → | ×1 |
| best first: higher scores lead regardless of input order | Martha, Dwayne, marhta, Martha, 90% | → | ×2 |
| no candidates, no matches | Anyone, , 0% | → | |
| a name that normalises to nothing is an error | Ltd., Acme, 85% | → | error: name is empty after normalising |
| a threshold over 10000 is an error | Acme, Acme, 100.01% | → | error: thresholdBasisPoints must be between 0 and 10000 |
| a negative threshold is an error | Acme, Acme, -0.01% | → | error: thresholdBasisPoints must be between 0 and 10000 |
| a name over 500 characters is an error | aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa… | → | error: names must be at most 500 characters after normalising |
More from the author
1. **Normalise** with text.normalise-name (trim, collapse Unicode whitespace), then lower-case each character (one-to-one case mappings only, as text.normalise-name does, so all three languages agree). 2. **Punctuation.** Apostrophes (' and ’) and full stops are removed, so "O'Neill" is "oneill" and "J.P." is "jp"; "&" becomes the word "and"; every other ASCII punctuation character becomes a space. Letters outside ASCII are kept as they are: "José" and "Jose" are close, not equal. 3. **Company forms** are dropped as whole words: ltd, limited, plc, llp, llc, inc. "Acme Ltd" and "Acme Limited" are the same key, "acme". 4. **Token sort.** The words are sorted (by code point) and joined with single spaces, so word order does not matter: "Smith, John" and "John Smith" are both "john smith". This is the "token sort" idea of fuzzy matching.
The result for each candidate carries this `matchKey`, so a reviewer can see why two names matched.
## The score
Jaro-Winkler similarity of the two keys, over Unicode code points, as defined by Winkler (W. E. Winkler, "String Comparator Metrics and Enhanced Decision Rules in the Fellegi-Sunter Model of Record Linkage", Proceedings of the Section on Survey Research Methods, American Statistical Association, 1990, pp. 354-359; the definition as set out at https://en.wikipedia.org/wiki/Jaro%E2%80%93Winkler_distance):
- matching characters are equal characters no further apart than floor(max(|s1|, |s2|) / 2) - 1, each used once, scanning s1 left to right; - t is half the number of matched characters that are out of order; - Jaro = (m/|s1| + m/|s2| + (m - t)/m) / 3, or 0 when m is 0; - Winkler adds l x 0.1 x (1 - Jaro) for a common prefix of l characters (at most 4), **only when Jaro is above 0.7**, Winkler's boost threshold.
It is computed exactly as a fraction of integers and rounded half-up once to basis points (10000 = identical), so the three languages give the same score and a threshold means the same thing everywhere. The standard examples: MARTHA/MARHTA 9611, DWAYNE/DUANE 8400, DIXON/DICKSONX 8133.
Results are the candidates scoring at least `thresholdBasisPoints`, highest score first, and in input order among equal scores, so the output is deterministic. A candidate that normalises to nothing (an empty string, "Ltd") scores 0. A query that normalises to nothing is an error, and so is a key over 500 characters (the exact arithmetic is bounded).
## Limits
A similarity score is a screen, not a decision: it will miss a name that changed on marriage, a trading name, a transliteration ("Mohammed" and "Muhammad" score well; "Ivanov" and "Иванов" do not), and it will flag common surnames. Choosing the threshold is a risk decision for the firm; set it low enough that a person reviews near misses. The company-form list is short and English; "Co", "Company", "Group" and "The" are kept on purpose, because dropping them joins genuinely different names.
Files
| Path | Bytes |
|---|---|
| README.md | 3,386 |
| impl/python.py | 3,489 |
| impl/rust.rs | 4,659 |
| impl/typescript.ts | 4,011 |
| vectors.json | 5,569 |