Functional Weave
Code in Python

stats.percentile@2.0.0

README.md

3,158 bytes · view raw

# stats.percentile

"The 95th percentile" is not one number: statistics packages disagree, and
Hyndman and Fan catalogue nine definitions. This capability makes the caller
name the one they mean.

- `nearest-rank`: the smallest value with at least p% of the sample at or
  below it. Rank = ceil(p/100 x n), and p = 0 gives the minimum. The answer is
  always a member of the list, which is what SLAs usually mean by "p95
  latency". This is Hyndman and Fan's type 1 (the inverse of the empirical
  distribution function).
- `linear`: interpolate between the two closest ranks at position
  h = (n - 1) x p/100 (zero-based), giving x[floor h] + (h - floor h) x
  (x[floor h + 1] - x[floor h]). This is Hyndman and Fan's type 7, the default
  in R (`quantile(type = 7)`) and NumPy (`method="linear"`), and Excel's
  `PERCENTILE.INC` (with p as a fraction there).

Both methods give the minimum at p = 0 and the maximum at p = 100. The input
is sorted numerically on a copy and never mutated.

**Precision.** The result is rounded to `decimals` places (0 to 12) by
`math.round-float`: half away from zero, decided on the exact value of the
double. So 0.5 rounds to 1 and -0.5 to -1 (Python's `round()` would give 0),
-0 is returned as 0, and 2.675, which is stored as 2.67499999999999982...,
rounds to 2.67 at two places.

**Why the three languages agree to the bit.** Only IEEE-754 addition,
subtraction, multiplication, division and floor are used, in the same order in
every language, and each of those is correctly rounded by the standard. No
library function whose last bit can differ between platforms (exp, log, sin)
is involved, so the unrounded result is the same double everywhere and the
rounding step, `math.round-float`, which is itself bit-identical in all three,
cannot split them. Position and rank are computed as (n - 1) x p / 100 and
p x n / 100, in that order.

## Changes in 2.0.0

2.0.0 rounds on the exact value of the double, so a percentile of 2.675 now
gives 2.67 at two places. 1.x rounded the scaled product instead (floor of
|x| x 10^decimals, compared with a half), and 2.675 x 100 is exactly 267.5 in
floating point, so 1.x said 2.68. The private rounding helper is gone: this
version requires `math.round-float ^1.0.0` and rounds with it, so every float
capability in the registry rounds alike.

None of the 1.x vectors changed answers (each was recomputed from the exact
value of its double). New vectors pin the difference: the linear 50th
percentile of 2.6 and 2.75 gives 2.67 (1.x 2.68), of -2.75 and -2.6 gives
-2.67 (1.x -2.68), the nearest-rank member 2.675 gives 2.67 (1.x 2.68), and
1.45 at one place gives 1.4 (1.x 1.5); 8.345, stored just above its tie,
still gives 8.35.

`decimals` is still 0 to 12, and still refused up front with the same message
("decimals must be a whole number from 0 to 12"), before the method is
looked at, as in 1.x.

Sources: R. J. Hyndman and Y. Fan, "Sample Quantiles in Statistical Packages",
The American Statistician 50(4), 1996, pp. 361-365; Microsoft, "PERCENTILE.INC
function" (support.microsoft.com); NIST/SEMATECH e-Handbook of Statistical
Methods, section 7.2.6.2 "Percentiles".