This isn't the Codex itself — it's the answer to the question every fair-minded reader of the Codex should ask: how did you decide what counts as a "hit," a "miss," and an "open question"? If our standards are sound, the conclusions follow from the evidence. If they're not, nothing else in the document matters. So here they are, stated plainly, before you take our word for anything.
A claim only counts as evidence if it was stated before the outcome was known, in a form specific enough that it could have failed. Anything invoked only after the fact to explain a miss — a new calendar, a hidden meaning, a "delta" — is not evidence. It's a rescue, and we score it as a rescue, not as a hit.
Not everything posted is a testable claim. A mood, a meme, a wink, "get ready" with no date attached — none of that can be scored, because it can't fail. We only scored something as a prediction if it met all three of these:
Everything that didn't meet this bar — most of the corpus — was left uncategorized as prophecy. It's still analyzed as art, symbolism, or cipher, just not scored as a prediction.
Every decoding claim in the Codex carries one of three labels, and we were as strict about handing out the good label as the bad one:
| Label | What it requires |
|---|---|
| VERIFIED | Decoded programmatically from an original-resolution file, by a standard, published algorithm (Base64, hex, QR, etc.), reproducible by anyone with the same file. No judgment calls. |
| OPEN | Recognized as probably meaningful — a pattern, a string, a repeated motif — but not solved, or solved only partially. Never dressed up as more than it is. |
| MISSED | The decode itself is solid (would be VERIFIED on its own), but the resulting claim, once decoded, did not come true when checked against the record. |
Notice what's missing from that list: there's no "probably means" or "clearly refers to" tier. If we couldn't reproduce a decode mechanically, we labeled it OPEN and said so, rather than asserting a reading with confidence we didn't earn.
This is the one that matters most, because it's where almost every disagreement with the riddler community actually lives. We didn't just check whether an outcome happened — plenty of vague statements will eventually brush up against something true by chance. We checked whether the claim, as originally stated, specifically predicted that outcome.
The Christmas 2017 post said only that "something extremely huge" was coming. It named no company, no deal, no date. When Ripple's MoneyGram partnership was announced weeks later, the account linked back to the Christmas post and said, in effect, "see, I told you." We do not score this as a hit — the original claim was not specific enough to have failed, which means it also isn't specific enough to have succeeded. A claim that cannot fail cannot pass a test, by definition, no matter what happens afterward.
We applied the identical standard to claims made for the account by the community after the fact — including calendar reinterpretations, alternate numerologies, and "it actually meant this the whole time" readings. See the box below.
This project encountered many claims that a miss wasn't really a miss, because some adjustment scheme — a "delta," an alternate calendar, a different reference point — reveals the real, correct date or meaning underneath. We take these seriously enough to test them, and the test is one question:
Was this rule stated, in fixed form, before the outcome was known?
If yes — if someone said in advance "apply this exact offset and you'll get the true date" and then that method produced a testable result — we would score it like any other claim. If no — if the rule only appears after a specific miss already exists to justify it — it fails the test regardless of how elegant or internally consistent it looks. A rule invented to explain one outcome, after that outcome is known, cannot be falsified by definition, and something that cannot be falsified cannot be evidence. This is not a special standard we invented for this project — it's the ordinary definition of a testable claim in any field.
For the question of coordinated timing — did his posts predict price moves rather than just describe them — we didn't rely on impression. We ran a statistical test, and we committed to the test's design before running it, so we couldn't unconsciously shape the outcome once we saw the data:
This is the same basic logic as a clinical trial: decide what would count as a real effect before you look at the results, so the results can surprise you.
A methodology is only meaningful if you can say what would have broken it. All of the following would have overturned or substantially revised the Codex's findings, and none of them occurred anywhere in 215 posts across nine years:
We went looking for all four. We're reporting that we didn't find them — not asserting in advance that we wouldn't.
This methodology proves what the evidence shows and nothing more. It cannot prove a negative with total certainty — there is always, in principle, an unfalsifiable version of any theory that survives by construction (see Section XVII of the full Codex on this exact point). What it can do, and what we believe it does, is show that on every test that could have produced disconfirming evidence, none appeared — which is the most any fair-minded analysis of an anonymous, ongoing body of work can honestly claim.
Every VERIFIED decode in this project is reproducible from the original file using standard, published tools — nothing proprietary, nothing that requires taking our word for it. And the same battery applies to any other riddle account making similar claims: assemble the dated corpus from primary sources, apply the same three-tier system, run the same advance-statement test on any proposed rescue, and report what you find, including if it disagrees with us.