Set the conditions
Choose a ready situation or set the conditions yourself. The browser generates hundreds of data sets with a known true value, applies ten estimators to each and compares their errors.
Where the PM is the most accurate
Where other estimators are more accurate
Conditions
Method settings
Why
Estimator errors — further left is better
What the data look like under these conditions
One of the generated data sets: outliers in red, the green line is the true value, next to it — the estimates on this set.
Reproduce locally (Python, same numbers; files benchmark.py and ifpa.py):
How the picture changes
One point of conditions is not yet an answer. Here the browser goes through a range of values and shows where the PM wins and where it loses. The other conditions are taken from the “Conditions” panel above.
Highlighted are the values where the PM is the best of all ten estimators. Hover over the chart to see the values of all lines. 150 data sets per point.
Each cell shows the best estimator under those conditions and the ratio of the PM error to the error of the best of the other nine: below 1 — the PM is more accurate. Hover over a cell for the numbers. 80 data sets per cell, so neighbouring cells may differ slightly by chance.
Summary: which estimator to choose
| Situation | Signs in your data | What to choose | PM vs the best of the others | |
|---|---|---|---|---|
| The PM is more accurate | ||||
| Some results are shifted and state too small an uncertainty | narrow intervals far from the main group, not overlapping it | PM | 36–44 % more accurate | |
| Few results, one of them a confident outlier | 4–6 results, one narrow interval off to the side | PM | 41–52 % more accurate | |
| Confident outliers in both directions | narrow intervals on both sides of the main group | PM | 15–30 % more accurate | |
| Heavy tails and confident outliers | many distant values, some with narrow intervals | PM | 43–49 % more accurate | |
| Everyone understates the error, and some also fail | intervals overlap less than they should; narrow intervals off to the side | PM | 37 % more accurate | |
| The error is bounded, some results are shifted | the error limit is known: accuracy class, resolution, tolerance | PM | 37–41 % more accurate | |
| Bounded error, faults both ways with ordinary uncertainty | the same, with shifted results on both sides | PM | 48–62 % more accurate | |
| Bounded error, no outliers | all intervals overlap, the error limit is known | PM | 22–25 % more accurate | |
| The error is skewed, outliers go one way | deviations to one side are longer than to the other | PM | 27–29 % more accurate | |
| Other estimators are more accurate | ||||
| No outliers, normal errors, honest uncertainties | all intervals overlap, the χ² of agreement is normal | Weighted mean | 27–28 % worse | |
| Outliers in both directions with ordinary uncertainty | distant results with wide, “honest” intervals | Median, ISO 13528 Algorithm A | 1.5–1.6 times worse | |
| The spread exceeds the stated one for all results (dark uncertainty) | intervals systematically fail to overlap, no clear outliers | DerSimonian–Laird | 1.5–1.7 times worse | |
| Heavy tails without confident outliers | distant values occur, but with ordinary uncertainty | Mixture model mode, largest consistent subset | 1.8–1.9 times worse | |
| About four results, an outlier with ordinary uncertainty | very few results | Median, Hodges–Lehmann | 17–21 % worse | |
| Results split into two groups of equal strength | two groups, each consistent inside | None: sort out the groups | — | Open example |
The numbers compare the RMS errors of the PM and of the best of the other nine estimators, 500 data sets, seeds 1 and 7. Hover over a situation to see why.
How the simulation works
The configurator is a Monte Carlo experiment: the browser repeatedly generates a set of results with a known true value, applies all ten estimators to each set and measures how wrong each one is. Everything runs in your browser; benchmark.py repeats the same computation with the same numbers.
Data
The true value is zero; all quantities are in units of σ, the typical standard deviation of a measurement. For each of the m results:
- the stated uncertainty uk = 1, or, if “uncertainties differ” is on, random between 0.5 and 2, uniform on a log scale (precise and rough results are equally common);
- the result xk = r · uk · ε + τ · N(0, 1), where ε is an error of the chosen distribution, r is how many times the actual error exceeds the stated one, and τ is the “dark” uncertainty, common to all and not reflected in uk;
- with probability “share of outliers” the result becomes an outlier: a shift is added — always positive if outliers are one-sided, otherwise with a random sign. The outlier's stated uncertainty is multiplied by the chosen factor while its actual spread stays the same: a factor of 0.3 means “off by 6σ and sure to be three times more precise than usual”.
Error distributions ε: normal; uniform and triangular (bounded — a digital instrument, rounding, a tolerance); Laplace (sharp peak, exponential tails); Student's t with 3 degrees of freedom (heavy tails); lognormal with shape 0.6 (skewed). All are scaled to zero mean and unit variance, so only the shape differs.
Which estimators are compared
| Uses | Estimators |
|---|---|
| values x only | arithmetic mean, median, Hodges–Lehmann estimator, ISO 13528 robust mean (Algorithm A) |
| x and u as a weight | weighted mean and weighted median (weights 1/u²), DerSimonian–Laird, largest consistent subset, mixture model mode |
| x and u as an interval width | preferential median: each result is one vote, a narrow interval gets no extra weight |
Formulas and code of all estimators are on the Code page.
The preferential median in the simulation
A result “votes” with the interval x ± k·u; by default k = 2, as for the expanded uncertainty in metrology (≈ 95 %). The RAV partition is fine, with a norm h of at most 0.05 of the median u: it stands in for the self-refinement of IF&PA-A in the papers. A medium partition (h ≤ 0.2u) or the partition of the light version IF&PA-L (h by the narrowest interval) can be chosen. If the maximum coverage is reached in two non-adjacent regions, the widest one is taken, and for equal widths the one closest to the median of the results. The other estimators get the original x and u.
Below the ranking the configurator shows PM variants with other partitions and k and the place each would take. They are not ranked: picking the best variant while knowing the truth would be an unfair advantage.
How estimators are compared
- rank — by RMS error (the root of the mean squared deviation from the truth): it includes both bias and spread and punishes rare large misses heavily;
- median error — a typical miss: in half of the data sets the error is smaller;
- 90 % of errors below — the miss in bad but not exceptional data sets;
- ±2u coverage — the share of data sets where the estimate ±2u of its own uncertainty covered the truth; an honest uncertainty gives about 95 %. A dash means the estimator has no uncertainty of its own.
How precise the simulation itself is
The RMS error over N data sets is itself random: its relative error is about 1/√(2N), i.e. roughly
3 % for 500 data sets and 5 % for 200. A change of rank between errors a few per cent apart is chance; treat a gap of
10 % or more as meaningful. Checking stability is easy: change the seed or increase the number of data sets.
The random generator is deterministic: with the same seed the result repeats in the browser and in benchmark.py.
The “How the picture changes” charts use 150 data sets per point, the map 80 per cell.
What the simulation does not cover
- correlations between results (a shared standard, a shared procedure);
- several groups of results of comparable strength — this case is worked through in the calculator;
- asymmetric intervals and a drifting measurand.
So the configurator answers “which estimator is more reliable under conditions of this type”, not “which value is right for my data” — that is what the calculator is for.