text.truncate
Shorten text to a maximum length on a word boundary and add an ellipsis, counting Unicode code points.
1.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 18 tests, run in TypeScript, Python and Rust.
What it does
Shortens text to at most `maxLength` characters, cutting at a word boundary and appending `ellipsis`: "Hello world" at 10 with "…" is "Hello…".
## The rules
For example
truncate(Hello world, 20, …)→ Hello world text that fits is returned unchangedtruncate(Hello world, 11, …)→ Hello world text exactly at the limit is returned unchanged, no ellipsistruncate(Hello world, 10, …)→ Hello… one over the limit backs up to the previous word
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
def truncate(value: str, max_length: int, ellipsis: str) -> str
| value | string | the text to shorten; returned unchanged if it already fits |
| max_length | int | the longest result allowed, in code points, ellipsis included |
| ellipsis | string | appended when text is cut, e.g. "…" or "..."; may be empty |
| returns | string |
Your code names it in one line, in the file that uses it
from fune.text.truncate import truncate # text.truncate@^1
from typing import Any
def _whole(value: Any) -> bool:
# bool is an int in Python; True is not a length of one here.
return isinstance(value, int) and not isinstance(value, bool)
def _is_whitespace(ch: str) -> bool:
"""The Unicode White_Space property, spelled out so all three languages
agree: str.isspace also accepts U+001C-U+001F, which JavaScript and Rust
do not.
"""
cp = ord(ch)
return (
0x09 <= cp <= 0x0D
or cp == 0x20
or cp == 0x85
or cp == 0xA0
or cp == 0x1680
or 0x2000 <= cp <= 0x200A
or cp == 0x2028
or cp == 0x2029
or cp == 0x202F
or cp == 0x205F
or cp == 0x3000
)
def _is_trailing_punctuation(ch: str) -> bool:
# Punctuation that reads badly directly before an ellipsis.
return ch in (",", ";", ":", "-", ".")
def truncate(value: str, max_length: int, ellipsis: str) -> str:
"""Shorten ``value`` to at most ``max_length`` code points, ellipsis
included, cutting at a word boundary where there is one.
"""
if not isinstance(value, str):
raise TypeError("truncate needs a string, received %r" % (value,))
if not _whole(max_length):
raise TypeError("maxLength must be a whole number, received %r" % (max_length,))
if max_length < 0:
raise ValueError("maxLength must be 0 or greater, received %d" % (max_length,))
if not isinstance(ellipsis, str):
raise TypeError("ellipsis must be a string, received %r" % (ellipsis,))
# A Python str is indexed by code point already.
if len(ellipsis) > max_length:
raise ValueError(
"maxLength must be at least the length of the ellipsis, received %d" % (max_length,)
)
if len(value) <= max_length:
return value
budget = max_length - len(ellipsis)
end = budget
# A cut just before whitespace is already a clean boundary.
if not _is_whitespace(value[budget]):
last_space = -1
for i in range(budget - 1, -1, -1):
if _is_whitespace(value[i]):
last_space = i
break
# No boundary to move back to: one long word is cut hard rather than
# reduced to a bare ellipsis.
if last_space > 0:
end = last_space
while end > 0 and (_is_whitespace(value[end - 1]) or _is_trailing_punctuation(value[end - 1])):
end -= 1
return value[:end] + ellipsisInstall
fune build
With that line in your source, in a Python project (language python in fune.project), fune build resolves it and nothing else, pins them in fune.lock, downloads only the Python package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:
fune add text.truncate
The manifest, vectors and README with only the Python implementation. Install it without the registry with fune add ./text.truncate-1.0.0-python.fune, or fetch it from a terminal with fune pull text.truncate@1.0.0:python.
The whole function, every language, is one file too: text.truncate-1.0.0.fune, 14,096 bytes, sha256 b7ad085ce8cf8efc872eace07a79adef908cc5061bb91d369cc64baf1819ae77. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
# fune: before text.truncate
after — your function gets the result and the arguments, and returns the final result.
# fune: after text.truncate
replace — it requires no other capability, so there is no dependency to replace.
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show text.truncate --steps.
# fune: step text.truncate after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| text that fits is returned unchanged | Hello world, 20, … | → | Hello world |
| text exactly at the limit is returned unchanged, no ellipsis | Hello world, 11, … | → | Hello world |
| one over the limit backs up to the previous word | Hello world, 10, … | → | Hello… |
| a cut that lands before a space keeps the whole word | Hello world again, 12, … | → | Hello world… |
| the ellipsis counts towards the limit | Hello world again, 14, ... | → | Hello world... |
| a trailing comma is dropped before the ellipsis | Hello, world, 10, ... | → | Hello... |
| one long word is cut hard rather than lost | Supercalifragilistic, 10, … | → | Supercali… |
| accented letters count once each | Crème brûlée à la carte, 14, … | → | Crème brûlée… |
| emoji count as one code point, not two UTF-16 units | 😀😀😀 abc, 5, … | → | 😀😀😀… |
| a line break is a word boundary | Line one Line two, 10, … | → | Line one… |
Show the other 8 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| a no-break space is a word boundary too | Mr Smithson, 9, … | → | Mr… |
| an empty ellipsis gives a plain cut | Hello world, 5, | → | Hello |
| room for only the ellipsis | Hello world, 1, … | → | … |
| the empty string fits any limit | , 5, … | → | |
| zero length with no ellipsis is the empty string | Hello, 0, | → | |
| a combining accent can be split from its letter: a documented limit | café, 4, | → | cafe |
| a limit shorter than the ellipsis is an error | Hello world, 2, ... | → | error: maxLength must be at least the length of the ellipsis |
| a negative limit is an error | Hello world, -1, | → | error: maxLength must be 0 or greater |
More from the author
1. If the text already fits in `maxLength`, it is returned unchanged, with no ellipsis. Nothing is trimmed or normalised. 2. Otherwise the text is cut to leave room for the ellipsis inside the limit: the result, ellipsis included, is never longer than `maxLength`. 3. If the cut lands exactly before whitespace, that is a clean word boundary and the cut stands. Otherwise it moves back to the last whitespace inside the kept part, so no word is split. 4. If there is no whitespace to move back to (one very long word, a URL, text in a script written without spaces), the word is cut hard at the limit. Splitting a word is better than returning nothing but an ellipsis. 5. Trailing whitespace and the trailing punctuation `, ; : - .` are removed from the cut before the ellipsis goes on, so "Hello, world" gives "Hello..." and not "Hello,...".
Whitespace is the Unicode White_Space set, spelled out character by character (tab, line feed, vertical tab, form feed, carriage return, space, U+0085, U+00A0, U+1680, U+2000-U+200A, U+2028, U+2029, U+202F, U+205F, U+3000) so all three languages agree; each language's built-in notion of "space" differs.
## Counting: code points, not graphemes
Length is counted in Unicode code points in all three languages. JavaScript's `.length` counts UTF-16 units, so a naive TypeScript version would count an emoji as two and disagree with Python; this one does not.
Code points are still not what a reader sees as one character. An accented letter written as a base letter plus a combining accent ("e" + U+0301) is two code points and a hard cut can separate them; a flag or a family emoji is several code points and can be cut in the middle. Getting that right needs the Unicode grapheme cluster rules (UAX #29), which none of the three standard libraries provides the same way, so this capability does not attempt it. If your text may contain such sequences, normalise it to NFC first (which fixes the accented letters) and leave some slack in `maxLength`.
## Errors
`maxLength` must be 0 or more, and at least the length of the ellipsis, since otherwise no truncated result can fit.
Files
| Path | Bytes |
|---|---|
| README.md | 2,330 |
| impl/python.py | 2,456 |
| impl/rust.rs | 2,382 |
| impl/typescript.ts | 2,517 |
| vectors.json | 2,227 |