collections.diff
Compare two lists of records by key: which were added, which removed, and which changed and in what fields.
1.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 18 tests, run in TypeScript, Python and Rust.
What it does
Compares two lists of records matched by `key` and says what happened between them: records that were added, records that were removed, and records present in both whose contents changed, with the names of the fields that differ. It is the heart of a sync job, an audit log or a "review changes before import" screen.
ORDER: `added` and `changed` follow the order of the `after` list, `removed` follows the order of the `before` list. `unchanged` is only a count, because a caller showing a diff never wants the untouched rows back.
For example
diff_by_key(before ×3, after ×3, id)→ added ×1, removed ×1, changed ×1, unchanged 1 one of each: an addition, a removal, a change and an untouched recorddiff_by_key(before ×2, after ×2, k)→ added , removed , changed , unchanged 2 identical lists: nothing added, removed or changeddiff_by_key(before ×2, after ×2, k)→ added , removed , changed , unchanged 2 reordering is not a change
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
pub fn diff_by_key(before: &[Value], after: &[Value], key: &str) -> RecordDiff<Value>
| before | record[] | the old list; every record needs a unique value at key |
| after | record[] | the new list; every record needs a unique value at key |
| key | string | the field that identifies a record in both lists |
| returns | RecordDiff<record> |
The types it declares, generated into your project
/// What changed between two lists of records matched by key.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct RecordDiff<T> {
/// in after but not before, in after's order
pub added: Vec<T>,
/// in before but not after, in before's order
pub removed: Vec<T>,
/// in both but different, in after's order
pub changed: Vec<RecordChange<T>>,
/// how many records are in both and identical
pub unchanged: i64,
}
/// One record present in both lists whose contents differ.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct RecordChange<T> {
/// the key value, rendered as text
pub key: String,
pub before: T,
pub after: T,
/// names of the fields that differ, sorted by code point
pub fields: Vec<String>,
}
Your code names it in one line, in the file that uses it
fune!(collections.diff@^1); // then call diff_by_key(…)
Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.
use super::funejson::Value; ← the fune runtime: the JSON value the test vectors use; fune build keeps it only where a signature takes one
use std::collections::HashMap;
/// Largest integer JavaScript can hold exactly; beyond it the three languages disagree.
const SAFE_INTEGER: i64 = 9007199254740991;
/// A record's key as text, rendered the way collections.group-by-key renders a
/// group name so the two agree on what "the same key" means. None when absent.
///
/// # Panics
/// Panics on a float, a list or a map, or an integer outside the safe range.
fn key_of(value: &Value, key: &str) -> Option<String> {
match value {
Value::Null => None,
Value::Str(s) => Some(s.clone()),
Value::Bool(b) => Some((if *b { "true" } else { "false" }).to_string()),
Value::Int(i) => {
if i.abs() > SAFE_INTEGER {
panic!("cannot diff by the out-of-range number {} at \"{}\"", i, key);
}
Some(i.to_string())
}
Value::Float(f) => panic!("cannot diff by the fractional number {} at \"{}\"", f, key),
_ => panic!("cannot diff by the list or map at \"{}\"", key),
}
}
/// Deep JSON equality, spelled out so all three languages agree: numbers by
/// value (1 equals 1.0), no coercion between types (true is not 1, "1" is not
/// 1), lists in order, and a missing map field equal to a null one.
fn same(a: &Value, b: &Value) -> bool {
match (a, b) {
(Value::Null, Value::Null) => true,
(Value::Bool(x), Value::Bool(y)) => x == y,
(Value::Int(x), Value::Int(y)) => x == y,
(Value::Int(_), Value::Float(_)) | (Value::Float(_), Value::Int(_)) | (Value::Float(_), Value::Float(_)) => {
a.as_f64() == b.as_f64()
}
(Value::Str(x), Value::Str(y)) => x == y,
(Value::Arr(x), Value::Arr(y)) => x.len() == y.len() && x.iter().zip(y).all(|(p, q)| same(p, q)),
(Value::Obj(x), Value::Obj(y)) => {
// `get` answers Null for a missing field, which is the rule.
x.iter().all(|(k, v)| same(v, b.get(k))) && y.iter().all(|(k, v)| same(a.get(k), v))
}
_ => false,
}
}
fn field_names(record: &Value) -> Vec<String> {
match record {
Value::Obj(pairs) => pairs.iter().map(|(k, _)| k.clone()).collect(),
_ => Vec::new(),
}
}
/// Key every record of one list, refusing missing and duplicate keys.
///
/// # Panics
/// Panics on a record with no key and on a duplicate key.
fn index(records: &[Value], key: &str, side: &str) -> (Vec<String>, HashMap<String, usize>) {
let mut order: Vec<String> = Vec::new();
let mut by_key: HashMap<String, usize> = HashMap::new();
for (i, record) in records.iter().enumerate() {
let k = match key_of(record.get(key), key) {
Some(k) => k,
None => panic!("record {} in {} has no value at \"{}\"", i, side, key),
};
// Picking one of two duplicates would report changes that never happened.
if by_key.contains_key(&k) {
panic!("duplicate key \"{}\" in {}", k, side);
}
by_key.insert(k.clone(), i);
order.push(k);
}
(order, by_key)
}
/// Compare two lists of records matched by `key`: what was added, what was
/// removed, and what changed and in which fields.
///
/// # Panics
/// Panics on an empty key, a record without a key, a duplicate key, or a key
/// value that cannot be rendered.
pub fn diff_by_key(before: &[Value], after: &[Value], key: &str) -> RecordDiff<Value> {
if key.is_empty() {
panic!("diff_by_key needs a non-empty key name");
}
let (before_order, before_keys) = index(before, key, "before");
let (after_order, after_keys) = index(after, key, "after");
let mut added: Vec<Value> = Vec::new();
let mut changed: Vec<RecordChange<Value>> = Vec::new();
let mut unchanged: i64 = 0;
for (i, k) in after_order.iter().enumerate() {
let next = &after[i];
let prev = match before_keys.get(k) {
Some(j) => &before[*j],
None => {
added.push(next.clone());
continue;
}
};
let mut names = field_names(prev);
for name in field_names(next) {
if !names.contains(&name) {
names.push(name);
}
}
let mut fields: Vec<String> = names
.into_iter()
.filter(|name| !same(prev.get(name), next.get(name)))
.collect();
// String ordering in Rust is byte order, which for UTF-8 is code point order.
fields.sort();
if fields.is_empty() {
unchanged += 1;
} else {
changed.push(RecordChange {
key: k.clone(),
before: prev.clone(),
after: next.clone(),
fields,
});
}
}
let removed: Vec<Value> = before_order
.iter()
.enumerate()
.filter(|(_, k)| !after_keys.contains_key(*k))
.map(|(j, _)| before[j].clone())
.collect();
RecordDiff {
added,
removed,
changed,
unchanged,
}
}
pub fn record_change_to_value(change: &RecordChange<Value>) -> Value {
Value::obj(vec![
("key", Value::str(&change.key)),
("before", change.before.clone()),
("after", change.after.clone()),
("fields", Value::Arr(change.fields.iter().map(|f| Value::str(f)).collect())),
])
}
pub fn record_diff_to_value(diff: &RecordDiff<Value>) -> Value {
Value::obj(vec![
("added", Value::Arr(diff.added.clone())),
("removed", Value::Arr(diff.removed.clone())),
("changed", Value::Arr(diff.changed.iter().map(record_change_to_value).collect())),
("unchanged", Value::Int(diff.unchanged)),
])
}
pub fn fune_vector(args: &[Value]) -> Value {
record_diff_to_value(&diff_by_key(args[0].as_arr(), args[1].as_arr(), args[2].as_str()))
}Install
fune build
With that line in your source, in a Rust project (language rust in fune.project), fune build resolves it and nothing else, pins them in fune.lock, downloads only the Rust package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. A crate’s build.rs runs it before every compile. Or pin a range in fune.project and build in one step:
fune add collections.diff
The manifest, vectors and README with only the Rust implementation. Install it without the registry with fune add ./collections.diff-1.0.0-rust.fune, or fetch it from a terminal with fune pull collections.diff@1.0.0:rust.
The whole function, every language, is one file too: collections.diff-1.0.0.fune, 27,655 bytes, sha256 0416ed8558f592216df058b45d0fc42a83ef3f616cee47d44df8507dcd1a0546. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
// fune: before collections.diff
after — your function gets the result and the arguments, and returns the final result.
// fune: after collections.diff
replace — it requires no other capability, so there is no dependency to replace.
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show collections.diff --steps.
// fune: step collections.diff after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| one of each: an addition, a removal, a change and an untouched record | before ×3, after ×3, id | → | added ×1, removed ×1, changed ×1, unchanged 1 |
| identical lists: nothing added, removed or changed | before ×2, after ×2, k | → | added , removed , changed , unchanged 2 |
| reordering is not a change | before ×2, after ×2, k | → | added , removed , changed , unchanged 2 |
| from nothing: everything is added, in after's order | , after ×2, k | → | added ×2, removed , changed , unchanged 0 |
| to nothing: everything is removed, in before's order | before ×2, , k | → | added , removed ×2, changed , unchanged 0 |
| changed fields are sorted by code point, including a field added and a field removed | before ×1, after ×1, id | → | added , removed , changed ×1, unchanged 0 |
| changed records follow after's order | before ×2, after ×2, k | → | added , removed , changed ×2, unchanged 0 |
| 1 and 1.0 are the same number, so a float round trip is not a change | before ×1, after ×1, k | → | added , removed , changed , unchanged 1 |
| a number that became a string did change | before ×1, after ×1, k | → | added , removed , changed ×1, unchanged 0 |
| true is not 1 | before ×1, after ×1, k | → | added , removed , changed ×1, unchanged 0 |
Show the other 8 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| a null field and a missing field are the same, so dropping nulls is not a change | before ×1, after ×1, k | → | added , removed , changed , unchanged 1 |
| nested lists and maps are compared deeply; list order matters | before ×1, after ×1, k | → | added , removed , changed ×1, unchanged 0 |
| a key of 1 and a key of "1" match; the key field itself is reported as changed type | before ×1, after ×1, id | → | added , removed , changed ×1, unchanged 0 |
| a duplicate key in before is an error, not a guess | before ×2, , k | → | error: duplicate key "a" in before |
| a duplicate key in after is an error | , after ×2, k | → | error: duplicate key "7" in after |
| a record with no value at the key is an error | before ×1, after ×1, k | → | error: record 0 in after has no value at "k" |
| a fractional key is an error | before ×1, , k | → | error: cannot diff by the fractional number |
| an empty key name is an error | , , | → | error: needs a non-empty key name |
More from the author
CHANGED FIELDS are listed by name, sorted by Unicode code point. Sorting, rather than keeping the records' own field order, is deliberate: JavaScript reorders integer-like object keys ("2" before "a") and the three languages would otherwise disagree about the same records.
EQUALITY is deep JSON equality, spelled out the same way in all three languages rather than inherited from each one's `==`:
- numbers compare by value, so `1` and `1.0` are equal (JSON does not distinguish them, and a round trip through a float column must not show up as a change); - `true` is not `1`, and `"1"` is not `1`: a value that changed type did change; - lists are equal when they have the same length and equal items in the same order; - maps are equal when every field is equal, where a missing field and a null field count as the same. That is the house rule across the collections capabilities, and it stops a serialiser that drops nulls from making every record look changed. It applies to the top-level fields too: a field that goes from null to absent is not reported.
KEYS are rendered the way `collections.group-by-key` renders them: strings are themselves, whole numbers their digits, booleans `true`/`false`. So a record keyed `1` in one list and `"1"` in the other is the same record (and the key field itself is then reported as changed, because its type did change). A fractional number, a list or a map at the key is an error.
ERRORS, not guesses: a record with no value (or null) at the key, and a key that appears twice in the same list. A diff that silently picked one of two duplicates would report changes that never happened. Also an empty key name and a non-list argument. Inputs are never modified; the records in the result are the records passed in.
The result types are generic (`RecordDiff<T>`, `RecordChange<T>`) and this function returns them over `record`. That keeps the door open for a typed variant, and it is also how the manifest gets a record-valued field past the Python type generator, which does not yet import `Any` for a bare `record` field.
Files
| Path | Bytes |
|---|---|
| README.md | 2,642 |
| impl/python.py | 4,233 |
| impl/rust.rs | 5,898 |
| impl/typescript.ts | 4,816 |
| vectors.json | 5,433 |