Functional Weave
Code in TypeScript

encoding.utf8@1.0.0

README.md

1,615 bytes · view raw

# encoding.utf8

`utf8Encode("€")` is `[226, 130, 172]` and `utf8Decode([226, 130, 172])` is
`"€"`. Every hash, HMAC and signature works on bytes, so text has to become
bytes the same way in every language before it is hashed: this is that step.
`crypto.sha256(utf8Encode(password))` gives the same digest in TypeScript,
Python and Rust.

Bytes are lists of integers 0 to 255, as everywhere in the registry (see
`encoding.hex`).

**Encoding never fails.** A JavaScript or Python string can hold a lone UTF-16
surrogate (half of an emoji, usually from cutting a string in the wrong
place), which has no UTF-8 form. It is encoded as U+FFFD, the replacement
character, which is what `TextEncoder` does in every browser. A Rust `&str`
cannot hold one at all.

**Decoding is strict**, per RFC 3629 section 3 and the table in section 4: an
overlong form (`C0 80` for NUL), an encoded surrogate (`ED A0 80`), a code
point above U+10FFFF, a truncated sequence or a stray continuation byte is an
error, never replaced with U+FFFD. Bytes that are not text should be carried
as bytes (or `encoding.base64`), not decoded and hoped for.

A byte order mark (`EF BB BF`) at the start is decoded as U+FEFF and kept:
this is a codec, not a file reader, and stripping it would make
`utf8Encode(utf8Decode(b))` differ from `b`. No Unicode normalisation is
applied in either direction (`"é"` precomposed and `"é"` as `e` plus a
combining accent are different bytes).

Source: RFC 3629, UTF-8, a transformation format of ISO 10646
(https://www.rfc-editor.org/rfc/rfc3629); the encoded examples in section 7
are among the vectors.