Instruments

Evaluations whose best achievable score is derived rather than estimated — and standard statistical tools graded against an answer known exactly.

A statistical tool is ordinarily graded against an estimate: the best score any predictor could reach on a task — the Bayes-optimal score, the expected score of a predictor that knows the distribution generating the data — is itself approximated from data, and whatever is measured against it inherits that approximation's error. The pages here remove the estimate. Each family of tasks is designed: the generating distribution is chosen so that the optimum comes out in closed form — the family's ceiling — and what a tool does against a known ceiling is then a computation with an exact answer: how far its score sits from optimal, whether a stated interval covers the truth, the size of the error it cannot see.

The designs run on the same object as the rest of these notes — exact residue arithmetic with the archimedean place deleted — and the ceilings are closed forms in the design's own parameters: nothing about the optimum is estimated from data. Read inward, the families chart their own dials against the ceiling. Read outward, published tools take the same tasks — deletion audits, Bayes-error estimators, set-valued predictors — and each miss is read as a law: the sign, the rate, or the exact limit of what the tool converges to instead of the truth.