Functional Weave
Code in TypeScript

stats.standard-deviation@1.0.0

README.md

1,816 bytes · view raw

# stats.standard-deviation

How spread out a set of numbers is. The caller says which one they mean,
because the two differ and both are called "the standard deviation":

- `population` divides the sum of squared deviations by n. Use it when the
  values are the whole population (every invoice this month). Excel `STDEV.P`,
  NumPy `std()`.
- `sample` divides by n - 1 (Bessel's correction). Use it when the values are
  a sample standing in for a larger population. Excel `STDEV.S`, R `sd()`,
  Python `statistics.stdev`. It needs at least two values.

**Algorithm.** Two passes: the mean first, then the sum of squared deviations
from it. The one-pass textbook shortcut, mean of squares minus square of the
mean, cancels catastrophically when the values are large and close together
(four readings near one billion come out as garbage or even a negative
variance); the vectors include that case.

**Precision.** The result is rounded to `decimals` places, half away from
zero, applied to the binary64 value: y = x x 10^decimals, r = floor(y), plus
one if y - r >= 0.5, divided back by 10^decimals. So a deviation of exactly
0.5 rounds to 1 at zero places, where Python's `round()` gives 0.

**Why the three languages agree to the bit.** Sums run left to right, and only
IEEE-754 +, -, x, /, square root and floor are used. All of these are
correctly rounded by the standard (square root included, unlike sin or exp,
whose last bit varies between math libraries), so TypeScript, Python and Rust
hold the same double before rounding and return the same rounded value.

Sources: NIST/SEMATECH e-Handbook of Statistical Methods, section 1.3.5.6
"Measures of Scale"; B. P. Welford, "Note on a Method for Calculating Corrected
Sums of Squares and Products", Technometrics 4(3), 1962, on why the one-pass
formula fails.