Functional Weave
Code in Python

text.truncate@1.0.0

README.md

2,330 bytes · view raw

# text.truncate

Shortens text to at most `maxLength` characters, cutting at a word boundary
and appending `ellipsis`: "Hello world" at 10 with "…" is "Hello…".

## The rules

1. If the text already fits in `maxLength`, it is returned unchanged, with no
   ellipsis. Nothing is trimmed or normalised.
2. Otherwise the text is cut to leave room for the ellipsis inside the limit:
   the result, ellipsis included, is never longer than `maxLength`.
3. If the cut lands exactly before whitespace, that is a clean word boundary
   and the cut stands. Otherwise it moves back to the last whitespace inside
   the kept part, so no word is split.
4. If there is no whitespace to move back to (one very long word, a URL, text
   in a script written without spaces), the word is cut hard at the limit.
   Splitting a word is better than returning nothing but an ellipsis.
5. Trailing whitespace and the trailing punctuation `, ; : - .` are removed
   from the cut before the ellipsis goes on, so "Hello, world" gives
   "Hello..." and not "Hello,...".

Whitespace is the Unicode White_Space set, spelled out character by character
(tab, line feed, vertical tab, form feed, carriage return, space, U+0085,
U+00A0, U+1680, U+2000-U+200A, U+2028, U+2029, U+202F, U+205F, U+3000) so all
three languages agree; each language's built-in notion of "space" differs.

## Counting: code points, not graphemes

Length is counted in Unicode code points in all three languages. JavaScript's
`.length` counts UTF-16 units, so a naive TypeScript version would count an
emoji as two and disagree with Python; this one does not.

Code points are still not what a reader sees as one character. An accented
letter written as a base letter plus a combining accent ("e" + U+0301) is two
code points and a hard cut can separate them; a flag or a family emoji is
several code points and can be cut in the middle. Getting that right needs
the Unicode grapheme cluster rules (UAX #29), which none of the three
standard libraries provides the same way, so this capability does not attempt
it. If your text may contain such sequences, normalise it to NFC first (which
fixes the accented letters) and leave some slack in `maxLength`.

## Errors

`maxLength` must be 0 or more, and at least the length of the ellipsis, since
otherwise no truncated result can fit.