text.normalise-name
Tidy a personal name: trim, collapse whitespace and title-case, with Mc, Mac, O', hyphen and van/de rules.
1.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 23 tests, run in TypeScript, Python and Rust.
What it does
Tidies a personal name typed into a form: " JOHN o'neill-MCDONALD " becomes "John O'Neill-McDonald". It is for display and for comparing records, not for deciding what someone's name "really" is.
## What it does
For example
normalise_name( john smith )→ John Smith trims and collapses runs of spacesnormalise_name(JOHN SMITH)→ John Smith all capitals are title-casednormalise_name(mary-jane o'neill)→ Mary-Jane O'Neill hyphen and apostrophe restart the capitals
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
pub fn normalise_name(value: &str) -> String
| value | string | a personal name as typed; may be all upper or all lower case |
| returns | string | single-spaced, title-cased; words typed in mixed case are kept as typed |
Your code names it in one line, in the file that uses it
fune!(text.normalise-name@^1); // then call normalise_name(…)
Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.
use super::funejson::Value; ← the fune runtime: the JSON value the test vectors use; fune build keeps it only where a signature takes one
/// Characters after which a new capital starts: hyphen, both apostrophes, full stop.
const PART_SEPARATORS: [char; 4] = ['-', '\'', '\u{2019}', '.'];
/// Lower-cased in the middle of a name: "Ludwig van Beethoven".
const PARTICLES: [&str; 20] = [
"van", "von", "de", "da", "di", "del", "della", "der", "den", "des", "du",
"la", "le", "ter", "ten", "dos", "das", "do", "bin", "ibn",
];
/// Names where "mac" is not a Gaelic prefix, so no capital follows it.
const NOT_MAC: [&str; 14] = [
"macari", "macaulay", "macedo", "macey", "machado", "machell", "machen",
"machin", "macias", "mackay", "mackey", "mackie", "mackin", "macklin",
];
/// The Unicode White_Space property, spelled out so all three languages agree.
fn is_whitespace(ch: char) -> bool {
let cp = ch as u32;
(0x09..=0x0d).contains(&cp)
|| cp == 0x20
|| cp == 0x85
|| cp == 0xa0
|| cp == 0x1680
|| (0x2000..=0x200a).contains(&cp)
|| cp == 0x2028
|| cp == 0x2029
|| cp == 0x202f
|| cp == 0x205f
|| cp == 0x3000
}
/// Case-map one character, but only when the mapping is itself one character.
/// 'ß' upper-cases to "SS"; leaving such characters alone keeps the three
/// languages identical.
fn lower_one(ch: char) -> char {
let mapped: Vec<char> = ch.to_lowercase().collect();
if mapped.len() == 1 {
mapped[0]
} else {
ch
}
}
fn upper_one(ch: char) -> char {
let mapped: Vec<char> = ch.to_uppercase().collect();
if mapped.len() == 1 {
mapped[0]
} else {
ch
}
}
fn is_mixed_case(word: &[char]) -> bool {
let upper = word.iter().any(|&c| lower_one(c) != c);
let lower = word.iter().any(|&c| upper_one(c) != c);
upper && lower
}
fn capitalise(chars: &[char]) -> String {
match chars.split_first() {
None => String::new(),
Some((first, rest)) => {
let mut out = String::new();
out.push(upper_one(*first));
out.extend(rest.iter());
out
}
}
}
/// Title-case one hyphen/apostrophe-delimited part, with the Mc and Mac rules.
fn case_part(part: &[char]) -> String {
let lower: Vec<char> = part.iter().map(|&c| lower_one(c)).collect();
let text: String = lower.iter().collect();
if lower.len() >= 3 && text.starts_with("mc") {
return format!("Mc{}", capitalise(&lower[2..]));
}
if lower.len() >= 6 && text.starts_with("mac") && !NOT_MAC.contains(&text.as_str()) {
return format!("Mac{}", capitalise(&lower[3..]));
}
capitalise(&lower)
}
fn title_word(word: &[char]) -> String {
let mut out = String::new();
let mut part: Vec<char> = Vec::new();
for &ch in word {
if PART_SEPARATORS.contains(&ch) {
out.push_str(&case_part(&part));
out.push(ch);
part.clear();
} else {
part.push(ch);
}
}
out.push_str(&case_part(&part));
out
}
/// Tidy a personal name: trim, collapse whitespace, and title-case words typed
/// all in one case, with the Mc, Mac, O', hyphen and particle rules.
pub fn normalise_name(value: &str) -> String {
let mut words: Vec<Vec<char>> = Vec::new();
let mut current: Vec<char> = Vec::new();
for ch in value.chars() {
if is_whitespace(ch) {
if !current.is_empty() {
words.push(std::mem::take(&mut current));
}
} else {
current.push(ch);
}
}
if !current.is_empty() {
words.push(current);
}
let count = words.len();
let out: Vec<String> = words
.iter()
.enumerate()
.map(|(i, word)| {
// Someone who typed their own capitals knows better than any rule.
if is_mixed_case(word) {
return word.iter().collect();
}
let lower: String = word.iter().map(|&c| lower_one(c)).collect();
if i > 0 && i + 1 < count && PARTICLES.contains(&lower.as_str()) {
return lower;
}
title_word(word)
})
.collect();
out.join(" ")
}
pub fn fune_vector(args: &[Value]) -> Value {
// Refuse what the typed signature cannot hold, with the wording TypeScript
// and Python use, rather than let the conversion below quietly change it.
if !matches!(args[0], Value::Str(_)) {
panic!("normaliseName needs a string, received {:?}", args[0]);
}
Value::Str(normalise_name(args[0].as_str()))
}Install
fune build
With that line in your source, in a Rust project (language rust in fune.project), fune build resolves it and nothing else, pins them in fune.lock, downloads only the Rust package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. A crate’s build.rs runs it before every compile. Or pin a range in fune.project and build in one step:
fune add text.normalise-name
The manifest, vectors and README with only the Rust implementation. Install it without the registry with fune add ./text.normalise-name-1.0.0-rust.fune, or fetch it from a terminal with fune pull text.normalise-name@1.0.0:rust.
The whole function, every language, is one file too: text.normalise-name-1.0.0.fune, 20,343 bytes, sha256 0bf63c83fe4c2e377d0cdaaa48a46a62df8d3f2ca06519e1f2eda52282043849. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
// fune: before text.normalise-name
after — your function gets the result and the arguments, and returns the final result.
// fune: after text.normalise-name
replace — it requires no other capability, so there is no dependency to replace.
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show text.normalise-name --steps.
// fune: step text.normalise-name after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| trims and collapses runs of spaces | john smith | → | John Smith |
| all capitals are title-cased | JOHN SMITH | → | John Smith |
| hyphen and apostrophe restart the capitals | mary-jane o'neill | → | Mary-Jane O'Neill |
| Mc gets a capital after it | MCDONALD | → | McDonald |
| Mc in a lower case full name | ronald mcdonald | → | Ronald McDonald |
| Mac gets a capital after it in a long enough name | ian macdonald | → | Ian MacDonald |
| a listed exception is not treated as Mac plus a name | MACEY | → | Macey |
| a short Mac word is never split | mack | → | Mack |
| O' and Mc across a hyphen | o'brien-mcarthur | → | O'Brien-McArthur |
| a typographic apostrophe works too | O’REILLY | → | O’Reilly |
Show the other 13 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| a particle in the middle is lower case | ludwig van beethoven | → | Ludwig van Beethoven |
| two particles in a row | JUAN DE LA CRUZ | → | Juan de la Cruz |
| a particle as the first word is capitalised | VAN MORRISON | → | Van Morrison |
| mixed case words are kept as typed, a naive title-case would say Dicaprio | leonardo DiCaprio | → | Leonardo DiCaprio |
| a capitalised particle typed deliberately is kept | Dick Van Dyke | → | Dick Van Dyke |
| full stops restart the capitals so initials survive | J.R.R. TOLKIEN | → | J.R.R. Tolkien |
| an apostrophe after a single letter | d'arcy | → | D'Arcy |
| accented Latin letters are cased | ZOË ÇELIK | → | Zoë Çelik |
| sharp s cannot be upper-cased to one letter so it is left as it is | STRAßER | → | Straßer |
| no-break spaces and tabs collapse to one space | anne marie | → | Anne Marie |
| the empty string stays empty | → | ||
| only whitespace becomes the empty string | → | ||
| a name that is not a string is an error | 42 | → | error: normaliseName needs a string |
More from the author
1. Trims, and collapses every run of whitespace to one ASCII space. Whitespace is the Unicode White_Space set, listed explicitly so all three languages agree (tab, line feed, vertical tab, form feed, carriage return, space, U+0085, U+00A0, U+1680, U+2000-U+200A, U+2028, U+2029, U+202F, U+205F, U+3000). 2. Leaves alone any word typed in **mixed case**. "DiCaprio", "MacKenzie", "DeVito" and "Ludwig Van Beethoven" come back exactly as typed. Someone who typed their own capitals knows better than a rule. Only words that are all upper case or all lower case are recased. 3. Title-cases each single-case word: the first letter upper case, the rest lower case, restarting after a hyphen, an apostrophe (' or ’) and a full stop, so "mary-jane" is "Mary-Jane", "o'neill" is "O'Neill", "d'arcy" is "D'Arcy" and "J.R.R." stays "J.R.R.". 4. **Mc**: a part starting "mc" gets a capital after it: "McDonald", "McArthur". 5. **Mac**: a part starting "mac" and at least six letters long gets a capital after it ("MacDonald", "MacIntyre"), except for a short list of names where "mac" is not a prefix: Macari, Macaulay, Macedo, Macey, Machado, Machell, Machen, Machin, Macias, Mackay, Mackey, Mackie, Mackin, Macklin. Shorter words ("Mack", "Macon") are never changed. 6. **Particles**: van, von, de, da, di, del, della, der, den, des, du, la, le, ter, ten, dos, das, do, bin and ibn are written in lower case when they are a whole word that is neither the first nor the last: "Ludwig van Beethoven", "Juan de la Cruz". As the first word they are capitalised ("Van Morrison").
## Honest limits
No rule gets every name right, and this one will be wrong for somebody. The Mac rule is a heuristic: "Mackenzie" becomes "MacKenzie", which some families spell and some do not, and a Mac-name missing from the exception list will gain a capital it should not have. Particles are capitalised by some families ("Dick Van Dyke") and not by others. Irish and Scottish Gaelic forms ("Mac an tSaoir", "Ó Briain" as two words) and Arabic "al-" names are not specially handled. Mixed-case words are trusted, so a genuinely mistyped "jOHN" is kept too. The mixed-case rule is the escape hatch: store what the person typed if they typed capitals, and show them the normalised form to confirm rather than silently rewriting it.
## Letters outside ASCII
Upper and lower case come from each language's own per-character Unicode mapping, used only when a character maps to exactly one character: "Ç" and "ç", "Ë" and "ë", Greek and Cyrillic all work. Characters whose case mapping would change the length ("ß" upper-cases to "SS", "İ" lower-cases to two code points) are left as they are, which keeps the three languages identical. Greek final sigma is not produced: an all-capitals Greek name ending in "Σ" comes back ending in "σ", not "ς".
The input must be a string; anything else is an error.
Files
| Path | Bytes |
|---|---|
| README.md | 3,180 |
| impl/python.py | 3,607 |
| impl/rust.rs | 4,569 |
| impl/typescript.ts | 3,942 |
| vectors.json | 2,486 |