Instruments
Evaluations whose best achievable score is derived rather than estimated — and standard statistical tools graded against an answer known exactly.
- Ceilings — evaluations whose optimal score is derived, not estimated
- Stated uncertainty — what a stated interval does when the truth is known exactly
- Deletion audits — what published tools do when the answer is known exactly
- Prediction sets — what a coverage guarantee covers, and what its rule abandons
- The swap — the trades that serve an atom some optimum abandons, graded by the lightest mass
A statistical tool is ordinarily graded against an estimate: the best score any predictor could reach on a task — the Bayes-optimal score, the expected score of a predictor that knows the distribution generating the data — is itself approximated from data, and whatever is measured against it inherits that approximation's error. The pages here remove the estimate. Each family of tasks is designed: the generating distribution is chosen so that the optimum comes out in closed form — the family's ceiling — and what a tool does against a known ceiling is then a computation with an exact answer: how far its score sits from optimal, whether a stated interval covers the truth, the size of the error it cannot see.
The designs run on the same object as the rest of these notes — exact residue arithmetic with the archimedean place deleted — and the ceilings are closed forms in the design's own parameters: nothing about the optimum is estimated from data. Read inward, the families chart their own dials against the ceiling. Read outward, published tools take the same tasks — deletion audits, Bayes-error estimators, set-valued predictors — and each miss is read as a law: the sign, the rate, or the exact limit of what the tool converges to instead of the truth.