Functional Weave
Code in TypeScript

encoding.utf8

Text to UTF-8 bytes and strictly back (RFC 3629), so hashing and signing see the same bytes in every language.

1.0.0 · published 2026-10-03 by charlie · Anterra

Pinned by 25 tests, run in TypeScript, Python and Rust.utf8Encode 10 · utf8Decode 15

What it does

`utf8Encode("€")` is `[226, 130, 172]` and `utf8Decode([226, 130, 172])` is `"€"`. Every hash, HMAC and signature works on bytes, so text has to become bytes the same way in every language before it is hashed: this is that step. `crypto.sha256(utf8Encode(password))` gives the same digest in TypeScript, Python and Rust.

Bytes are lists of integers 0 to 255, as everywhere in the registry (see `encoding.hex`).

The functions

A group: 2 functions that work together, each in its own file, each pinned by its own tests in TypeScript, Python and Rust. A project can install only the ones it calls.

  1. utf8Encode (text: string) -> int[]
  2. utf8Decode (bytes: int[]) -> string

Once installed, your code imports each one from the group's module.

utf8Encode 10 tests

export function utf8Encode(text: string): readonly number[]
textstringany string; a lone UTF-16 surrogate becomes U+FFFD, as TextEncoder does
returnsint[]the UTF-8 bytes, each an integer from 0 to 255

For example

  • utf8Encode() → the empty string
  • utf8Encode(A) → 65 ASCII is one byte per character
  • utf8Encode(£) → 194, 163 the pound sign takes two bytes
import { utf8Encode } from "#fune/encoding.utf8@^1";
impl/typescript/utf8_encode.ts · 33 lines · open · raw
/**
 * Text as UTF-8 bytes. Written out rather than using TextEncoder so the
 * registry's TypeScript needs no DOM or Node typings; it gives the same
 * answer, including U+FFFD for a lone surrogate.
 */
export function utf8Encode(text: string): readonly number[] {
  if (typeof text !== "string") {
    throw new TypeError("text must be a string");
  }
  const out: number[] = [];
  for (let i = 0; i < text.length; i++) {
    let cp = text.charCodeAt(i);
    if (cp >= 0xd800 && cp <= 0xdbff && i + 1 < text.length) {
      const next = text.charCodeAt(i + 1);
      if (next >= 0xdc00 && next <= 0xdfff) {
        cp = 0x10000 + ((cp - 0xd800) << 10) + (next - 0xdc00);
        i++;
      }
    }
    // Anything still in the surrogate range is half a pair with no partner.
    if (cp >= 0xd800 && cp <= 0xdfff) cp = 0xfffd;
    if (cp < 0x80) {
      out.push(cp);
    } else if (cp < 0x800) {
      out.push(0xc0 | (cp >> 6), 0x80 | (cp & 63));
    } else if (cp < 0x10000) {
      out.push(0xe0 | (cp >> 12), 0x80 | ((cp >> 6) & 63), 0x80 | (cp & 63));
    } else {
      out.push(0xf0 | (cp >> 18), 0x80 | ((cp >> 12) & 63), 0x80 | ((cp >> 6) & 63), 0x80 | (cp & 63));
    }
  }
  return out;
}

utf8Decode throws on bad input 15 tests

export function utf8Decode(bytes: readonly number[]): string
bytesint[]must be well-formed UTF-8: no overlong forms, surrogates or code points past U+10FFFF
returnsstringthe text; a leading byte order mark is kept, not stripped

For example

  • utf8Decode() → no bytes is the empty string
  • utf8Decode(65, 226, 137, 162, 206, 145, 46) → A≢Α. RFC 3629 section 7: A, NOT IDENTICAL TO, ALPHA, full stop
  • utf8Decode(230, 151, 165, 230, 156, 172, 232, 170, 158) → 日本語 RFC 3629 section 7: nihongo, Japanese
import { utf8Decode } from "#fune/encoding.utf8@^1";
impl/typescript/utf8_decode.ts · 58 lines · open · raw
const INVALID = "bytes are not valid UTF-8";

/**
 * UTF-8 bytes back to text, strictly (RFC 3629): overlong forms, encoded
 * surrogates, code points past U+10FFFF, truncated sequences and stray
 * continuation bytes are errors. A leading byte order mark is kept.
 */
export function utf8Decode(bytes: readonly number[]): string {
  if (!Array.isArray(bytes) && !(bytes instanceof Uint8Array)) {
    throw new TypeError("bytes must be a list of integers from 0 to 255");
  }
  for (let i = 0; i < bytes.length; i++) {
    const b = bytes[i];
    if (typeof b !== "number" || !Number.isInteger(b) || b < 0 || b > 255) {
      throw new RangeError("bytes must be a list of integers from 0 to 255");
    }
  }
  let out = "";
  let i = 0;
  while (i < bytes.length) {
    const b0 = bytes[i];
    let cp: number;
    let need: number;
    let min: number;
    if (b0 < 0x80) {
      out += String.fromCharCode(b0);
      i++;
      continue;
    } else if (b0 >= 0xc2 && b0 <= 0xdf) {
      cp = b0 & 0x1f;
      need = 1;
      min = 0x80;
    } else if (b0 >= 0xe0 && b0 <= 0xef) {
      cp = b0 & 0x0f;
      need = 2;
      min = 0x800;
    } else if (b0 >= 0xf0 && b0 <= 0xf4) {
      cp = b0 & 0x07;
      need = 3;
      min = 0x10000;
    } else {
      // 80-BF continuation with no lead, C0/C1 (always overlong), F5-FF.
      throw new RangeError(INVALID);
    }
    if (i + need >= bytes.length) throw new RangeError(INVALID);
    for (let k = 1; k <= need; k++) {
      const b = bytes[i + k];
      if ((b & 0xc0) !== 0x80) throw new RangeError(INVALID);
      cp = (cp << 6) | (b & 0x3f);
    }
    if (cp < min || cp > 0x10ffff || (cp >= 0xd800 && cp <= 0xdfff)) {
      throw new RangeError(INVALID);
    }
    out += String.fromCodePoint(cp);
    i += need + 1;
  }
  return out;
}

Install

fune build

With that line in your source, in a TypeScript project (language typescript in fune.project), fune build resolves it and nothing else, pins them in fune.lock, downloads only the TypeScript package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:

fune add encoding.utf8

That builds the whole group. To build only what you call, and whatever it uses inside the group:

fune add encoding.utf8 --only utf8Encode
Download for TypeScript encoding.utf8-1.0.0-typescript.fune · 10,992 bytes sha256 70d12ad13cfea672ff3aa811677528f95d37188cadf44e9c1cf725ef3628f5fe

The manifest, vectors and README with only the TypeScript implementation. Install it without the registry with fune add ./encoding.utf8-1.0.0-typescript.fune, or fetch it from a terminal with fune pull encoding.utf8@1.0.0:typescript.

The whole function, every language, is one file too: encoding.utf8-1.0.0.fune, 14,837 bytes, sha256 a476924f3a2599c51305953b1cb504a9ed4b90e0a92bd71556a9bb45bc0c9f86. It installs into a project of any language.

Customise it in your app

The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.

before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.

// fune: before encoding.utf8.utf8Encode
// fune: before encoding.utf8.utf8Decode

after — your function gets the result and the arguments, and returns the final result.

// fune: after encoding.utf8.utf8Encode
// fune: after encoding.utf8.utf8Decode

replace — it requires no other capability, so there is no dependency to replace.

step — your function runs at a numbered point inside a function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show encoding.utf8 --steps.

// fune: step encoding.utf8.<fn> after <n|label>

Tests

A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.

utf8Encode 10 tests

CaseArgumentsExpected
the empty string →
ASCII is one byte per character A → 65
the pound sign takes two bytes £ → 194, 163
the euro sign takes three bytes € → 226, 130, 172
RFC 3629 section 7: A, NOT IDENTICAL TO, ALPHA, full stop A≢Α. → 65, 226, 137, 162, 206, 145, 46
RFC 3629 section 7: hangugeo, Korean 한국어 → 237, 149, 156, 234, 181, 173, 236, 150, 180
RFC 3629 section 7: nihongo, Japanese 日本語 → 230, 151, 165, 230, 156, 172, 232, 170, 158
an emoji outside the BMP is one four-byte sequence, not two surrogates 😀 → 240, 159, 152, 128
a NUL character is a zero byte ab → 97, 0, 98
a lone surrogate becomes U+FFFD, as TextEncoder does a�b → 97, 239, 191, 189, 98

utf8Decode 15 tests

CaseArgumentsExpected
no bytes is the empty string →
RFC 3629 section 7: A, NOT IDENTICAL TO, ALPHA, full stop 65, 226, 137, 162, 206, 145, 46 → A≢Α.
RFC 3629 section 7: nihongo, Japanese 230, 151, 165, 230, 156, 172, 232, 170, 158 → 日本語
a four-byte sequence 240, 159, 152, 128 → 😀
the highest code point, U+10FFFF 244, 143, 191, 191 → 􏿿
a leading byte order mark is kept, not stripped 239, 187, 191, 65 → A
an overlong NUL (C0 80) is refused 192, 128 → error: bytes are not valid UTF-8
an overlong three-byte slash (E0 80 AF) is refused 224, 128, 175 → error: bytes are not valid UTF-8
an encoded surrogate (ED A0 80) is refused 237, 160, 128 → error: bytes are not valid UTF-8
a code point above U+10FFFF (F4 90 80 80) is refused 244, 144, 128, 128 → error: bytes are not valid UTF-8
Show the other 5 tests
CaseArgumentsExpected
a truncated euro sign 226, 130 → error: bytes are not valid UTF-8
a stray continuation byte 65, 128 → error: bytes are not valid UTF-8
FF never appears in UTF-8 255 → error: bytes are not valid UTF-8
a value above 255 is not a byte 65, 256 → error: bytes must be a list of integers from 0 to 255
a fraction is not a byte 65.5 → error: bytes must be a list of integers from 0 to 255

More from the author

**Encoding never fails.** A JavaScript or Python string can hold a lone UTF-16 surrogate (half of an emoji, usually from cutting a string in the wrong place), which has no UTF-8 form. It is encoded as U+FFFD, the replacement character, which is what `TextEncoder` does in every browser. A Rust `&str` cannot hold one at all.

**Decoding is strict**, per RFC 3629 section 3 and the table in section 4: an overlong form (`C0 80` for NUL), an encoded surrogate (`ED A0 80`), a code point above U+10FFFF, a truncated sequence or a stray continuation byte is an error, never replaced with U+FFFD. Bytes that are not text should be carried as bytes (or `encoding.base64`), not decoded and hoped for.

A byte order mark (`EF BB BF`) at the start is decoded as U+FEFF and kept: this is a codec, not a file reader, and stripping it would make `utf8Encode(utf8Decode(b))` differ from `b`. No Unicode normalisation is applied in either direction (`"é"` precomposed and `"é"` as `e` plus a combining accent are different bytes).

Source: RFC 3629, UTF-8, a transformation format of ISO 10646 (https://www.rfc-editor.org/rfc/rfc3629); the encoded examples in section 7 are among the vectors.

Files

PathBytes
README.md1,615
impl/python/utf8_decode.py906
impl/python/utf8_encode.py901
impl/rust/utf8_decode.rs1,207
impl/rust/utf8_encode.rs564
impl/typescript/utf8_decode.ts1,797
impl/typescript/utf8_encode.ts1,202
vectors.json3,305