ATHENA’S PALACE The Codex

ScienceA Codex article

What Is Pre-Registration, and Can It Work for History?

Say what would prove you wrong, publicly and timestamped, before you look at the evidence.

Every claim in this article is in the ring. The numbered lines below are its claims. Think one is wrong? Argue it against the house, dig into it with a friend and the Owl, or propose a rewrite. Open a war, and the result stays on the line for everyone who reads this after you.
In this article
  1. The problem it solves is not lying
  2. What was deposited, and when
  3. What held
  4. What failed: four losses, named
  5. Why nobody in this field does it
  6. What would make this worthless
  7. Common questions
  8. Sources

Short answer. Pre-registration means stating what would prove you wrong (publicly, with a timestamp) before you look at the evidence. It is routine in medicine and psychology and essentially absent from the study of religious history. On 10 August 2026 this Foundation deposited twelve falsifiable predictions to Zenodo with a DOI. Four days later it built the instruments and tested them. One prediction held at odds of roughly ten million to one. Four came back against the house. This page lists all of them, because a page that listed only the survivor would be the thing pre-registration exists to prevent.

Tier: this page is a method statement, not a historical claim. Every number in it is checkable against the deposit record and the published code.

The problem it solves is not lying

Almost nobody in this field fabricates evidence. The problem is subtler and far more common: the decisions about which analysis counts get made after the data are in view. Which parallels are close enough to list. Which manuscripts are representative. Where to stop counting. Each choice feels reasonable on its own, and made downstream of the answer they reliably converge on the answer.

Statisticians call the general shape of this the garden of forking paths, and its signature is a field full of confident, incompatible, individually well-argued conclusions. That is a fair description of the Zoroastrian-influence debate, and of a good deal else.

Pre-registration removes the choice rather than asking you to resist it. If the rule is fixed and public before the data are seen, it is not available to be bent later, not by a dishonest researcher, and more importantly not by an honest one.

What was deposited, and when

On 10 August 2026 the Foundation deposited a volume containing twelve original theses to Zenodo, each closed with an explicit statement of what would kill it, under doi.org/10.5281/zenodo.21879071. The deposit is third-party, immutable and timestamped, which matters more than anything the volume says: the order is verifiable by a stranger without trusting us about anything.

One of those twelve carried a prediction labelled, in its own text, computational, now. It said that Old Iranian loanwords in the Hebrew Bible would cluster in legal and administrative semantic fields, and it named its own kill condition in advance: if Iranian loans turn out evenly distributed, or ritual-heavy, the carrier claim dies.

On 14 August the instrument was built and run. The prediction preceded the test by four days and the record proves it.

What held

Of 38 Old Iranian loanwords in Biblical Hebrew and Aramaic, 39.5% fall in the semantic fields covering law and political office, against an 8.3% expectation if they were spread evenly across the taxonomy. The exact one-sided binomial probability is about 1.6 in ten million.

Two features of that result matter more than the number. First, the semantic taxonomy was borrowed from linguists who have no stake in this question, chosen precisely so its categories could not be shaped to fit. Second, the effect gets stronger when you throw away the weak evidence: restricted to the etymologies nobody disputes, the concentration rises to 50%. An artifact of wishful lexicography would move the other way.

The full dataset is published, verse by verse, as The Persian Passband. It is the only result from that week that survives every correction applied to it.

What failed: four losses, named

A pre-registered statistical test returned a result against us, and it is still in the report. One measure of distributional concentration came back at p = 0.998, that is, the opposite of the predicted direction, decisively. The diagnosis is that the measure was direction-blind and could not distinguish concentration in the predicted fields from concentration anywhere else. That diagnosis was written after seeing the number and is worth less than the failure, which is why the failing test still runs and still prints.

A thesis was retracted the same week it was written. A claim that Western Christian doctrine sits on the grammatical gaps of Latin (the missing definite article, the collapsed perfect tense) demanded its own audit, got one, and produced no signal whatever: divergence occurred at 50% of the loci where the gap applied and 50% where it did not.

An attractive argument collapsed under its own model. Isaiah 45 calls the Persian king Cyrus a mashiach, the only gentile in the Hebrew Bible to receive the title. It is the most quotable exhibit in the whole influence literature. Modelled properly it returns 84.5% for a much duller explanation: Judah was participating in a standing imperial propaganda genre, and the Cyrus Cylinder has the god Marduk make the identical move.

And an announcement was withdrawn for a reason nobody had checked. Across six instruments, 53 statistical tests had been reported. At the conventional threshold, 2.7 false positives are expected from chance alone. Nothing had been corrected for that. Seven nominally significant results do not survive the correction, including one this Foundation had announced as a headline days earlier. They were retracted.

Why nobody in this field does it

Not because historians are careless. Because there is no norm requiring it and no penalty for its absence. Medicine adopted pre-registration after discovering how many trials quietly changed their endpoints; psychology adopted it after a replication crisis made the cost visible. History of religion has had neither reckoning, so the practice never arrived.

It is also genuinely harder here. You cannot randomise the Achaemenid empire, and most historical questions cannot be settled by one clean test. But some can, and the ones that can are exactly the ones the field argues about endlessly without resolving.

That gap is the reason this method is worth using. It is not that the technique is clever. It is that the position is empty.

What would make this worthless

One rigged result. The value of everything above depends entirely on the losses being real, and a single quietly amended prediction or dropped inconvenient test would zero it, not reduce it, zero it. Worse, it would make this Foundation an instance of the exact behaviour its books document in other institutions.

So the boring rules are load-bearing. Analysis plans are fixed in code before results are viewed. Amendments are published as dated amendments beside the original, never as silent edits. Failed tests appear in the verdict, not in an appendix.

And one instrument built that week is currently barred from use. It failed its own hostile re-score, and its results will not be cited for anything until people outside this house have scored its rubric independently. That study does not exist yet. Saying so is part of the method.

Common questions

Is pre-registration just a formality?

No. It removes the single largest source of false findings in any field that tests many hypotheses: choosing which analysis to report after seeing the data. If the decision rule is fixed and public beforehand, that choice is not available.

Why does no one in religious history do this?

Because the field runs on argument, authority and peer review after the fact rather than on stated predictions. There is no norm requiring it and no penalty for its absence. That is exactly why it is available as a method.

Doesn't publishing your failures just discredit you?

It costs a thesis and buys the only thing that makes the surviving ones worth reading. A body of work in which nothing ever failed is evidence that the falsification conditions were decorative.

What is the difference between this and simply being careful?

A timestamp. Care is invisible and unfalsifiable; a third-party deposit record is neither. Anyone can verify that the prediction preceded the test without trusting the person who made it.

Can a single researcher pre-register honestly?

Partly. Fixing the analysis plan in code before running it is real, and so is publishing the losses. What a single researcher cannot supply is independent scoring, which is why one instrument here is barred from use until other people have scored its rubric.

Sources

The deposit under test: doi.org/10.5281/zenodo.21879071, 10 August 2026 · Andrew Gelman & Eric Loken, "The Garden of Forking Paths" (2013) · Brian Nosek et al., "The Preregistration Revolution," PNAS 115 (2018) · Yoav Benjamini & Yosef Hochberg, "Controlling the False Discovery Rate," JRSS B 57 (1995) · Samuel Sandmel, "Parallelomania," JBL 81 (1962) · the datasets, code and failing tests described here are published in full.

The volume this page reports on: doi.org/10.5281/zenodo.21879071

First published in The Fire and the Veil by Asher Wilder, under CC BY 4.0. Original publication. Adapted for the Codex with numbered claims, updated internal links and punctuation.

The claims on trial

Each line is one claim this article makes, word for word the motion if somebody disputes it. Its standing changes only when a war over it is settled.

Still not settled?

Dig into it together: you, whoever you bring, and the Owl work out what is actually true, and share the findings. Or take the other side and spar it out in front of a neutral judge.

Add a claim

Anyone may add. Nobody may edit somebody else's line: if you think it is wrong, fight it, or rewrite it and win a debate for your words. Keep it to one proposition, because that exact sentence becomes the motion when somebody disputes it.

Opening a war costs 25 gold and wins 50; a reader with no gold gets their first war on the house, with no purse. You pick the arms: argue it in the Academy, or settle it at a game. Losing burns the deposit, which is why the button is not free. "Rewrite" is the same war fought for your own words: they become the motion, and if you win they replace the line, the old words kept beneath. Whoever defends the line is paid by whichever engine judged them, and the house takes the podium if nobody else does, so no claim goes undefended. "Argue it now" is the free door: the line becomes the motion of a judged duel against the house, with no account and no gold, and the page keeps no record of it. "Fight for it at a game" is its twin: the line rides into the games hall as your banner, and the page counts the matches finished under it this week. A count is not a verdict. It moves neither the standing nor the evidence tier.

Dispute this claim

Your words become the motion, and you argue for them in a judged debate. Whoever defends the line argues against; if nobody takes that podium within a minute or so, the house does. Win, and your words replace the line with your name on them, the old words kept beneath as the record. Lose, and the line has held.

No gold yet? Your first war is on the house. After that a war costs 25 gold, and winning it pays 50.