Functional Weave
Code in Python

text.normalise-name

Tidy a personal name: trim, collapse whitespace and title-case, with Mc, Mac, O', hyphen and van/de rules.

1.0.0 · published 2026-10-03 by charlie · Anterra

Pinned by 23 tests, run in TypeScript, Python and Rust.

What it does

Tidies a personal name typed into a form: " JOHN o'neill-MCDONALD " becomes "John O'Neill-McDonald". It is for display and for comparing records, not for deciding what someone's name "really" is.

## What it does

For example

  • normalise_name( john smith ) → John Smith trims and collapses runs of spaces
  • normalise_name(JOHN SMITH) → John Smith all capitals are title-cased
  • normalise_name(mary-jane o'neill) → Mary-Jane O'Neill hyphen and apostrophe restart the capitals

The function

The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.

def normalise_name(value: str) -> str
valuestringa personal name as typed; may be all upper or all lower case
returnsstringsingle-spaced, title-cased; words typed in mixed case are kept as typed

Your code names it in one line, in the file that uses it

from fune.text.normalise_name import normalise_name  # text.normalise-name@^1
impl/python.py · 119 lines · open · raw
from typing import List

#: Characters after which a new capital starts: hyphen, both apostrophes, full stop.
_PART_SEPARATORS = ("-", "'", "’", ".")

#: Lower-cased in the middle of a name: "Ludwig van Beethoven".
_PARTICLES = frozenset([
    "van", "von", "de", "da", "di", "del", "della", "der", "den", "des", "du",
    "la", "le", "ter", "ten", "dos", "das", "do", "bin", "ibn",
])

#: Names where "mac" is not a Gaelic prefix, so no capital follows it.
_NOT_MAC = frozenset([
    "macari", "macaulay", "macedo", "macey", "machado", "machell", "machen",
    "machin", "macias", "mackay", "mackey", "mackie", "mackin", "macklin",
])


def _is_whitespace(ch: str) -> bool:
    """The Unicode White_Space property, spelled out so all three languages
    agree: str.isspace also accepts U+001C-U+001F, which JavaScript and Rust
    do not.
    """
    cp = ord(ch)
    return (
        0x09 <= cp <= 0x0D
        or cp == 0x20
        or cp == 0x85
        or cp == 0xA0
        or cp == 0x1680
        or 0x2000 <= cp <= 0x200A
        or cp == 0x2028
        or cp == 0x2029
        or cp == 0x202F
        or cp == 0x205F
        or cp == 0x3000
    )


def _lower_one(ch: str) -> str:
    # Only one-to-one mappings: "İ".lower() is two code points, and taking it
    # would make Python disagree with the length-preserving rule in all three.
    mapped = ch.lower()
    return mapped if len(mapped) == 1 else ch


def _upper_one(ch: str) -> str:
    # "ß".upper() is "SS": left alone, for the same reason.
    mapped = ch.upper()
    return mapped if len(mapped) == 1 else ch


def _lower_all(chars: str) -> str:
    return "".join(_lower_one(ch) for ch in chars)


def _is_mixed_case(word: str) -> bool:
    upper = any(_lower_one(ch) != ch for ch in word)
    lower = any(_upper_one(ch) != ch for ch in word)
    return upper and lower


def _capitalise(text: str) -> str:
    return _upper_one(text[0]) + text[1:] if text else ""


def _case_part(part: str) -> str:
    """Title-case one hyphen/apostrophe-delimited part, with the Mc and Mac rules."""
    lower = _lower_all(part)
    if len(lower) >= 3 and lower.startswith("mc"):
        return "Mc" + _capitalise(lower[2:])
    if len(lower) >= 6 and lower.startswith("mac") and lower not in _NOT_MAC:
        return "Mac" + _capitalise(lower[3:])
    return _capitalise(lower)


def _title_word(word: str) -> str:
    out = ""
    part = ""
    for ch in word:
        if ch in _PART_SEPARATORS:
            out += _case_part(part) + ch
            part = ""
        else:
            part += ch
    return out + _case_part(part)


def normalise_name(value: str) -> str:
    """Tidy a personal name: trim, collapse whitespace, and title-case words
    typed all in one case, with the Mc, Mac, O', hyphen and particle rules.
    """
    if not isinstance(value, str):
        raise TypeError("normaliseName needs a string, received %r" % (value,))

    words: List[str] = []
    current = ""
    for ch in value:
        if _is_whitespace(ch):
            if current:
                words.append(current)
            current = ""
        else:
            current += ch
    if current:
        words.append(current)

    out: List[str] = []
    for i, word in enumerate(words):
        # Someone who typed their own capitals knows better than any rule.
        if _is_mixed_case(word):
            out.append(word)
            continue
        lower = _lower_all(word)
        if 0 < i < len(words) - 1 and lower in _PARTICLES:
            out.append(lower)
            continue
        out.append(_title_word(word))
    return " ".join(out)

Install

fune build

With that line in your source, in a Python project (language python in fune.project), fune build resolves it and nothing else, pins them in fune.lock, downloads only the Python package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:

fune add text.normalise-name
Download for Python text.normalise-name-1.0.0-python.fune · 11,329 bytes sha256 a353080ffef6306220603e2df0a71f0d3a22c826ea4960f7a8c80b2e83977062

The manifest, vectors and README with only the Python implementation. Install it without the registry with fune add ./text.normalise-name-1.0.0-python.fune, or fetch it from a terminal with fune pull text.normalise-name@1.0.0:python.

The whole function, every language, is one file too: text.normalise-name-1.0.0.fune, 20,343 bytes, sha256 0bf63c83fe4c2e377d0cdaaa48a46a62df8d3f2ca06519e1f2eda52282043849. It installs into a project of any language.

Customise it in your app

The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.

before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.

# fune: before text.normalise-name

after — your function gets the result and the arguments, and returns the final result.

# fune: after text.normalise-name

replace — it requires no other capability, so there is no dependency to replace.

step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show text.normalise-name --steps.

# fune: step text.normalise-name after <n|label>

Tests

A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.

CaseArgumentsExpected
trims and collapses runs of spaces john smith → John Smith
all capitals are title-cased JOHN SMITH → John Smith
hyphen and apostrophe restart the capitals mary-jane o'neill → Mary-Jane O'Neill
Mc gets a capital after it MCDONALD → McDonald
Mc in a lower case full name ronald mcdonald → Ronald McDonald
Mac gets a capital after it in a long enough name ian macdonald → Ian MacDonald
a listed exception is not treated as Mac plus a name MACEY → Macey
a short Mac word is never split mack → Mack
O' and Mc across a hyphen o'brien-mcarthur → O'Brien-McArthur
a typographic apostrophe works too O’REILLY → O’Reilly
Show the other 13 tests
CaseArgumentsExpected
a particle in the middle is lower case ludwig van beethoven → Ludwig van Beethoven
two particles in a row JUAN DE LA CRUZ → Juan de la Cruz
a particle as the first word is capitalised VAN MORRISON → Van Morrison
mixed case words are kept as typed, a naive title-case would say Dicaprio leonardo DiCaprio → Leonardo DiCaprio
a capitalised particle typed deliberately is kept Dick Van Dyke → Dick Van Dyke
full stops restart the capitals so initials survive J.R.R. TOLKIEN → J.R.R. Tolkien
an apostrophe after a single letter d'arcy → D'Arcy
accented Latin letters are cased ZOË ÇELIK → Zoë Çelik
sharp s cannot be upper-cased to one letter so it is left as it is STRAßER → Straßer
no-break spaces and tabs collapse to one space anne  marie → Anne Marie
the empty string stays empty →
only whitespace becomes the empty string →
a name that is not a string is an error 42 → error: normaliseName needs a string

More from the author

1. Trims, and collapses every run of whitespace to one ASCII space. Whitespace is the Unicode White_Space set, listed explicitly so all three languages agree (tab, line feed, vertical tab, form feed, carriage return, space, U+0085, U+00A0, U+1680, U+2000-U+200A, U+2028, U+2029, U+202F, U+205F, U+3000). 2. Leaves alone any word typed in **mixed case**. "DiCaprio", "MacKenzie", "DeVito" and "Ludwig Van Beethoven" come back exactly as typed. Someone who typed their own capitals knows better than a rule. Only words that are all upper case or all lower case are recased. 3. Title-cases each single-case word: the first letter upper case, the rest lower case, restarting after a hyphen, an apostrophe (' or ’) and a full stop, so "mary-jane" is "Mary-Jane", "o'neill" is "O'Neill", "d'arcy" is "D'Arcy" and "J.R.R." stays "J.R.R.". 4. **Mc**: a part starting "mc" gets a capital after it: "McDonald", "McArthur". 5. **Mac**: a part starting "mac" and at least six letters long gets a capital after it ("MacDonald", "MacIntyre"), except for a short list of names where "mac" is not a prefix: Macari, Macaulay, Macedo, Macey, Machado, Machell, Machen, Machin, Macias, Mackay, Mackey, Mackie, Mackin, Macklin. Shorter words ("Mack", "Macon") are never changed. 6. **Particles**: van, von, de, da, di, del, della, der, den, des, du, la, le, ter, ten, dos, das, do, bin and ibn are written in lower case when they are a whole word that is neither the first nor the last: "Ludwig van Beethoven", "Juan de la Cruz". As the first word they are capitalised ("Van Morrison").

## Honest limits

No rule gets every name right, and this one will be wrong for somebody. The Mac rule is a heuristic: "Mackenzie" becomes "MacKenzie", which some families spell and some do not, and a Mac-name missing from the exception list will gain a capital it should not have. Particles are capitalised by some families ("Dick Van Dyke") and not by others. Irish and Scottish Gaelic forms ("Mac an tSaoir", "Ó Briain" as two words) and Arabic "al-" names are not specially handled. Mixed-case words are trusted, so a genuinely mistyped "jOHN" is kept too. The mixed-case rule is the escape hatch: store what the person typed if they typed capitals, and show them the normalised form to confirm rather than silently rewriting it.

## Letters outside ASCII

Upper and lower case come from each language's own per-character Unicode mapping, used only when a character maps to exactly one character: "Ç" and "ç", "Ë" and "ë", Greek and Cyrillic all work. Characters whose case mapping would change the length ("ß" upper-cases to "SS", "İ" lower-cases to two code points) are left as they are, which keeps the three languages identical. Greek final sigma is not produced: an all-capitals Greek name ending in "Σ" comes back ending in "σ", not "ς".

The input must be a string; anything else is an error.

Files

PathBytes
README.md3,180
impl/python.py3,607
impl/rust.rs4,569
impl/typescript.ts3,942
vectors.json2,486