# collections.dedupe-by-key
Removes records that share the value at `key`, keeping either the first or the
last record of each set of duplicates. `first` is "the original wins" (an
import that must not overwrite); `last` is "the latest wins" (a change feed
where later rows supersede earlier ones).
WHERE SURVIVORS SIT: every survivor keeps its own position relative to the
other survivors. With `keep = last`, the surviving record sits where the last
occurrence was, not where the first one was:
[a#1, b#1, a#2] keep first -> [a#1, b#1]
[a#1, b#1, a#2] keep last -> [b#1, a#2]
That is what a change feed means by "latest wins": the survivor is the latest
row, in the latest row's place. A caller who wants the latest values in the
original slot can dedupe with `last` and re-sort.
WHAT COUNTS AS THE SAME KEY: values are compared by the same rendering
`collections.group-by-key` uses, so the two agree on what a key is. Strings are
themselves, whole numbers are their decimal digits, booleans are `true` and
`false`. That means the number `1` and the string `"1"` are the same key, as
they are in a CSV, a query string and every JSON API that is loose about
types. A fractional number, a list or a map at the key is an error, because
there is no rendering of them all three languages agree on; so is a whole
number beyond 2^53, which JavaScript cannot hold exactly.
MISSING KEYS ARE NEVER DUPLICATES: a record whose key is absent or null is
always kept. Two records that both lack an id are two things we know nothing
about, not one thing seen twice, and collapsing them would silently lose data.
(This is where dedupe deliberately differs from group-by-key, which puts them
in one "" group.)
Errors: a non-list, an empty key name, a `keep` other than `first` or `last`,
and an ungroupable value at the key. The input is never modified.