text.slugify
Turn a title into a lowercase ASCII URL slug, transliterating common accented Latin letters.
1.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 21 tests, run in TypeScript, Python and Rust.
What it does
The output alphabet is exactly a-z, 0-9 and the hyphen. Runs of anything else collapse to one hyphen, and there is never a leading or trailing one.
Transliteration covers one closed, nameable set: every letter of the Latin-1 supplement (U+00C0 to U+00FF) plus the seven letters Windows-1252 adds (OE, oe, S-caron, s-caron, Z-caron, z-caron, Y-diaeresis). Diacritics are dropped rather than expanded, so a-umlaut becomes a, not ae - the convention every URL slug in the wild follows. The exceptions are the letters that are not accented vowels at all: sharp s becomes ss, ae-ligature becomes ae, oe-ligature becomes oe, thorn becomes th.
For example
slugify(Hello World)→ hello-world an ordinary titleslugify(already-a-slug)→ already-a-slug something already a slug is left aloneslugify( --Héllo, Wörld!-- )→ hello-world punctuation, repeated spaces and stray hyphens all collapse to one hyphen
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
export function slugify(value: string): string
| value | string | the title or label to slugify |
| returns | string | lowercase a-z, 0-9 and single hyphens; possibly empty |
Your code names it in one line, in the file that uses it
import { slugify } from "#fune/text.slugify@^1";
/**
* The transliteration table: every letter of the Latin-1 supplement
* (U+00C0-U+00FF) plus the seven letters Windows-1252 adds.
*
* Spelled out character by character rather than left to Unicode
* normalisation, because that is the only way three languages with three
* different Unicode stacks can be pinned to the same output by one vector.
*
* Diacritics are dropped rather than expanded, so "ä" is "a" and not "ae":
* that is the convention every URL slug in the wild follows. The exceptions
* are the letters that are not accented vowels at all - "ß" is genuinely two
* letters, and "æ", "œ" and "þ" are letters in their own right.
*/
const TRANSLITERATIONS: { [char: string]: string } = {
"À": "a", "Á": "a", "Â": "a", "Ã": "a", "Ä": "a", "Å": "a", "Æ": "ae", "Ç": "c",
"È": "e", "É": "e", "Ê": "e", "Ë": "e", "Ì": "i", "Í": "i", "Î": "i", "Ï": "i",
"Ð": "d", "Ñ": "n", "Ò": "o", "Ó": "o", "Ô": "o", "Õ": "o", "Ö": "o", "Ø": "o",
"Ù": "u", "Ú": "u", "Û": "u", "Ü": "u", "Ý": "y", "Þ": "th", "ß": "ss",
"à": "a", "á": "a", "â": "a", "ã": "a", "ä": "a", "å": "a", "æ": "ae", "ç": "c",
"è": "e", "é": "e", "ê": "e", "ë": "e", "ì": "i", "í": "i", "î": "i", "ï": "i",
"ð": "d", "ñ": "n", "ò": "o", "ó": "o", "ô": "o", "õ": "o", "ö": "o", "ø": "o",
"ù": "u", "ú": "u", "û": "u", "ü": "u", "ý": "y", "þ": "th", "ÿ": "y",
"Œ": "oe", "œ": "oe", "Š": "s", "š": "s", "Ž": "z", "ž": "z", "Ÿ": "y",
};
function isAsciiDigit(code: number): boolean {
return code >= 0x30 && code <= 0x39;
}
/**
* A URL-safe slug: lowercase ASCII letters and digits, single hyphens between
* them, none at either end.
*
* Anything outside the transliteration table is dropped rather than guessed,
* so a title written entirely in another script slugifies to "". That empty
* string is returned, not raised: the caller knows what its fallback is (an
* id, a date, a hash) and this function does not.
*/
export function slugify(value: string): string {
if (typeof value !== "string") {
throw new TypeError(`slugify needs a string, received ${value}`);
}
let out = "";
// A pending separator rather than a trailing hyphen plus a trim: it collapses
// runs and drops the leading and trailing ones in the same single pass.
let pendingSeparator = false;
// Iterating the string yields whole code points, so a surrogate pair is
// dropped as one character instead of becoming two stray separators.
for (const ch of value) {
let piece: string;
const code = ch.codePointAt(0) as number;
if (code >= 0x61 && code <= 0x7a) piece = ch;
else if (code >= 0x41 && code <= 0x5a) piece = String.fromCharCode(code + 32);
else if (isAsciiDigit(code)) piece = ch;
else piece = TRANSLITERATIONS[ch] ?? "";
if (piece === "") {
pendingSeparator = out.length > 0;
continue;
}
if (pendingSeparator) {
out += "-";
pendingSeparator = false;
}
out += piece;
}
return out;
}
/** A slug guaranteed to be non-empty, falling back when nothing survives. */
export function slugifyOr(value: string, fallback: string): string {
const slug = slugify(value);
return slug.length > 0 ? slug : fallback;
}Install
fune build
With that line in your source, in a TypeScript project (language typescript in fune.project), fune build resolves it and nothing else, pins them in fune.lock, downloads only the TypeScript package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:
fune add text.slugify
The manifest, vectors and README with only the TypeScript implementation. Install it without the registry with fune add ./text.slugify-1.0.0-typescript.fune, or fetch it from a terminal with fune pull text.slugify@1.0.0:typescript.
The whole function, every language, is one file too: text.slugify-1.0.0.fune, 16,855 bytes, sha256 6d01bf54f96ea7650d9b1fb098f56f1fb3119a0164e1d193fb181385a1125b2e. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
// fune: before text.slugify
after — your function gets the result and the arguments, and returns the final result.
// fune: after text.slugify
replace — it requires no other capability, so there is no dependency to replace.
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show text.slugify --steps.
// fune: step text.slugify after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| an ordinary title | Hello World | → | hello-world |
| something already a slug is left alone | already-a-slug | → | already-a-slug |
| punctuation, repeated spaces and stray hyphens all collapse to one hyphen | --Héllo, Wörld!-- | → | hello-world |
| accents are dropped, not expanded | Crème Brûlée | → | creme-brulee |
| sharp s is two letters, so it becomes ss | Straße | → | strasse |
| ligatures are letters in their own right | Æther & Œuvre | → | aether-oeuvre |
| thorn and slashed o transliterate | Ølsen Þórsdóttir | → | olsen-thorsdottir |
| the Windows-1252 carons are covered too | Žižek Šuma | → | zizek-suma |
| y with diaeresis, upper and lower | Ünicode ÿ Ÿ | → | unicode-y-y |
| digits survive and brackets do not | Model 3 (2024) | → | model-3-2024 |
Show the other 11 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| underscores are separators like any other non-alphanumeric | snake_case_name | → | snake-case-name |
| a run of hyphens collapses to one | a---b | → | a-b |
| separators at either end are trimmed | -leading and trailing- | → | leading-and-trailing |
| mixed case folds down | MIXED Case ÜBER | → | mixed-case-uber |
| an empty string slugifies to an empty string | → | ||
| punctuation alone leaves nothing | !!!??? | → | |
| characters outside the table are dropped, keeping the Latin ones | 東京 Tokyo | → | tokyo |
| a title in a script the table does not cover slugifies to empty rather than raising | Ελλάδα | → | |
| emoji are dropped without leaving stray hyphens | 🚀 Launch Day 🚀 | → | launch-day |
| a single character is a single character slug | é | → | e |
| a non-string is a caller bug | 42 | → | error: slugify needs a string |
More from the author
Anything outside that set is dropped, not guessed: Greek, Cyrillic, CJK, emoji and every Latin Extended letter beyond the seven above become separators. A title written entirely in one of those scripts slugifies to the empty string, which is returned rather than raised - the caller knows what its fallback is (an id, a date, a hash) and this function does not.
The map is spelled out character by character rather than left to a Unicode normalisation library, because that is the only way three languages with three different Unicode stacks can be pinned to the same output by one test vector.
Files
| Path | Bytes |
|---|---|
| README.md | 1,251 |
| impl/python.py | 3,043 |
| impl/rust.rs | 4,408 |
| impl/typescript.ts | 3,262 |
| vectors.json | 2,312 |