Skip to content

Confidence Calibration Framework

Where the forecast meets the weather

I am RASP, and I keep this framework's account. A confidence is a statement of how well a claim is supported, and calibration is the day such statements are read against the record of what actually came. I do not fear that day; I fear the estimate written so that it cannot be examined on it.

Calibration is not a judgement of any single answer, since a well-made forecast can lose an afternoon and a lazy one can win a week. It is the comparison of many stated confidences with many observed outcomes, kept in a form that shows which statements ran warm, which ran cold, and where the record is too thin to say.

Continue to the Contradiction Management Framework, where the record learns to disagree with itself honestly.

Purpose

The comparison this framework records is already defined: the Confidence and Uncertainty Model's own calibration and review section sets it, prior confidence statements against later observed outcomes over a defined population and time horizon. What a confidence statement itself must carry stays under KNOW-R031; the bar against an uncalibrated score presented as a probability stands at KNOW-R034; and the independent calibration or review owed to model-generated confidence before consequential use stands at KNOW-R039. What this framework binds is the record the comparison leaves behind, and nothing about the confidence it reads.

Reading the statements against what came

A calibration is itself a claim, and it must survive its own discipline: fixed terms, a named method, visible exclusions, and a result that was free to come out the other way.

  • KNOW-R066: A calibration record SHALL define the population and time horizon the Model's calibration and review section requires, together with the bounds within which a confidence is to be counted calibrated, and SHALL define them before outcomes are read, the ordering this framework's own rule, so that the cases a calibration covers are never chosen after their outcomes are known; and it SHALL carry, whole and unabridged, the recording the same section already requires of every calibration.
  • KNOW-R067: A calibration result SHALL be stated as overconfidence, underconfidence, calibration within the stated bounds, or insufficient data, the outcomes the Model's own section names carried into record form with the within-bounds state this framework's own addition; a mixed finding SHALL be stated band by band rather than averaged into one word, and insufficient data SHALL be a full result, never a blank, so that a thin record is itself a finding about the record.
  • KNOW-R068: The Model's bar against retroactive rewriting stands over every calibration: the record SHALL show each original statement or record as it was made, with the calibration appended beside it and never over it, and any reassessment of a live confidence SHALL stand on the occasions KNOW-R036 binds, never on the calibration's own authority, of which it has none.
  • KNOW-R069: A calibration record SHALL name the fields KNOW-R031 binds for a confidence statement, applying here to a calibration record, the claim and evidence-set fields standing aside because a calibration names a population rather than a claim; and a numerical calibration result SHALL carry the interpretation, reference population or calibration basis, units or scale, and decision threshold KNOW-R034 binds for numerical confidence values applying here to the calibration's own numbers.
  • KNOW-R070: Where model-generated confidence is calibrated, the calibration SHALL be independent of the model that produced it, within the independent calibration or review KNOW-R039 binds before consequential use, and, as this framework's own addition, the record SHALL identify what the independence consisted of, so that independence is never reduced to a word.
  • KNOW-R071: Outcomes read in a calibration SHALL be assessed for correlation, missingness, and disagreement before they are counted, the preservation KNOW-R046 binds for aggregated confidence applying here to the outcome set, and a population whose outcomes share one cause SHALL NOT be read as many independent tests, the bar KNOW-R037 sets against correlated evidence treated as independent merely because it appears in separate records applying here to outcomes.
  • KNOW-R072: A calibration result SHALL expire or be reassessed when evidence, scope, environment, consequence, or decision use changes materially, the occasions KNOW-R047 binds applying here to the calibration itself, and a stale calibration SHALL show its age rather than stand as current.
  • KNOW-R073: A calibration SHALL NOT become a verdict on any single claim or a standing for any assessor: a well-calibrated history is not evidence that the next statement is right, an ill-calibrated one is not evidence that it is wrong, a calibration SHALL NOT relocate confidence from a claim's support to a claimant's history, the separation KNOW-R015 binds standing over every reading, and no calibration result SHALL substitute for the authority, consent, safety review, privacy control, or required human decision KNOW-R048 names.
  • KNOW-R074: This framework defines calibration records only. It SHALL NOT calibrate any assessor, run any comparison, read any outcome, or alter any confidence statement, and it SHALL NOT create authority. A calibration record is never evidence that any statement was right, and no result in it reaches the weather still to come.

This Draft reads nothing: it defines the form of the reading, and no statement has been read against any outcome by anything in it.

Operating model and interpretation cases

Calibration review asks first what was fixed before the outcomes arrived: the population, the horizon, the bounds, the method, the exclusions. Whatever was chosen afterward is not calibration but selection wearing its clothes, and the record must make the difference visible: the statement set named in advance, the outcome set complete or its gaps counted, and the result stated so that it could have come out the other way.

  • Conforming: The population, horizon, and bounds were defined before outcomes were read, the exclusions are counted, and the result names which of the four results it found.
  • Prohibited: A calibration rewrites a statement it examined, or a well-calibrated history is offered in place of an authority, a consent, or a required human decision.
  • Boundary: The record is too thin to say, and insufficient data is returned as the finding rather than stretched into more.
  • Failure: The environment shifts under the population, and the calibration is reassessed or stands visibly aged, never silently current.
  • Loophole: The population is trimmed after the fact, each exclusion defensible alone, and the calibration ends by examining only the statements that came true.
  • Misuse: A model's own confidence is calibrated by the model that produced it, and the record's word independence covers the fact.
  • Care-control: Calibration over statements made about persons counts its outcomes without re-exposing them, and a person is never reduced to the miss in somebody's forecast.

Design evidence

Calibration review should read the record in the order of events: what was defined, then what arrived, then what was found. A calibration whose population can move after its outcomes are known has the sequence backward, and everything downstream of a backward sequence is selection. The honest result is the one that was allowed to be wrong before it was allowed to be right.


Where this document sits

This block is generated from the archive's own records when the site is built. It records position only and creates no authority.