Functional Weave
Code in TypeScript

stats.standard-deviation@2.0.0

README.md

2,843 bytes · view raw

# stats.standard-deviation

How spread out a set of numbers is. The caller says which one they mean,
because the two differ and both are called "the standard deviation":

- `population` divides the sum of squared deviations by n. Use it when the
  values are the whole population (every invoice this month). Excel `STDEV.P`,
  NumPy `std()`.
- `sample` divides by n - 1 (Bessel's correction). Use it when the values are
  a sample standing in for a larger population. Excel `STDEV.S`, R `sd()`,
  Python `statistics.stdev`. It needs at least two values.

**Algorithm.** Two passes: the mean first, then the sum of squared deviations
from it. The one-pass textbook shortcut, mean of squares minus square of the
mean, cancels catastrophically when the values are large and close together
(four readings near one billion come out as garbage or even a negative
variance); the vectors include that case.

**Precision.** The result is rounded to `decimals` places (0 to 12) by
`math.round-float`: half away from zero, decided on the exact value of the
double. So a deviation of exactly 0.5 rounds to 1 at zero places, where
Python's `round()` gives 0, and a deviation of 2.675 (stored as
2.67499999999999982...) rounds to 2.67 at two places.

**Why the three languages agree to the bit.** Sums run left to right, and only
IEEE-754 +, -, x, /, square root and floor are used. All of these are
correctly rounded by the standard (square root included, unlike sin or exp,
whose last bit varies between math libraries), so TypeScript, Python and Rust
hold the same double before rounding, and `math.round-float` rounds it the
same way in all three.

## Changes in 2.0.0

2.0.0 rounds on the exact value of the double, so a deviation of 2.675 (the
population deviation of 0 and 5.35) now gives 2.67. 1.x rounded the scaled
product instead (floor of x x 10^decimals, compared with a half), and
2.675 x 100 is exactly 267.5 in floating point, so 1.x said 2.68. The private
rounding helper is gone: this version requires `math.round-float ^1.0.0` and
rounds with it, so every float capability in the registry rounds alike.

None of the 1.x vectors changed answers (each was recomputed from the exact
value of its double). New vectors pin the difference: the population
deviation of 0 and 5.35 gives 2.67 (1.x 2.68), of -2.675 and 2.675 gives
2.67 (1.x 2.68), and of 1.45 and -1.45 at one place gives 1.4 (1.x 1.5).

`decimals` is still 0 to 12, and still refused up front with the same message
("decimals must be a whole number from 0 to 12"), before an empty list or a
one-value sample is reported, as in 1.x.

Sources: NIST/SEMATECH e-Handbook of Statistical Methods, section 1.3.5.6
"Measures of Scale"; B. P. Welford, "Note on a Method for Calculating Corrected
Sums of Squares and Products", Technometrics 4(3), 1962, on why the one-pass
formula fails.