Scaling reliable science
The O-Ring Benchmark
How far can automated verification scale
without making any error?
Here, an error means failing to find a known error in a paper. False alarms and other review mistakes are not measured.
Science is cumulative. New research builds on earlier findings. If AI enables research production at unprecedented scale, errors that escape verification can become foundations for later claims. Scaling scientific production therefore requires verification and correction that can support the accumulation of reliable knowledge.
For Project APE, this is about accumulating reliable evidence on public policy. Consider a policy evaluated first in one context, then in another. Suppose its true effect is zero in both. A critical error in the first paper leads it to falsely report a positive effect, and the verifier misses the error. The follow-up paper correctly finds a null effect. Taken at face value, the findings could suggest that the policy works in one context but not the other, raising concerns about external validity. Yet the apparent heterogeneity comes from an error, not a genuine difference in policy effects.
Now imagine policy-evaluation papers being produced sequentially, with a verification step after each paper. The O-Ring benchmark examines one part of this challenge: how the probability of missing no errors changes as the number of papers produced in a sequence increases.
The homepage benchmark identifies the cost–recall frontier on papers with one known error. Here we take those frontier models and ask a harder question: how reliably do they find every error across a growing sequence of policy-evaluation papers, when a paper can contain more than one? We study this using shuffled, recorded reviews—not newly generated sequences of papers.
Reading the benchmark
Reliable evaluations. Reliable foundations.
What the curve means
At paper 10, a value of 80% means 800 of the 1,000 sampled chains found every benchmark error in their first 10 papers. Each paper has one verification step. After a chain misses an error, it never recovers. The curve starts at 100% before any papers are attempted. All cards share the same axes.
Each figure groups one model and harness, with separate lines for its completed thinking efforts. Line patterns identify effort; colors identify cumulative verification cost. Both encodings are consistent across figures. The paper selector compares all efforts at the same point in the chain, including the percentage-point difference between the highest and lowest available efforts. Higher effort is not assumed to perform better.
How far can we trust the chain?
The reliable scaling horizon is the largest number of consecutive papers for which the probability of finding every known error is at least 90% or 99%. A horizon of 0 means even one paper falls short. These are empirical probabilities for this fixed review pool, not confidence guarantees about future papers.
These are the first two steps in the “march of nines” described by Andrej Karpathy: 90% and 99% reliability. Here, each threshold applies to the entire chain, not just one paper.
The headline numbers show the best completed effort at each threshold, with the achieving effort labelled and ties shown. They are not averages or a combination of different efforts within one chain. Pending efforts are excluded from curves and headline numbers.
For each paper, we calculate the fraction of its three variants verified without a miss. The exact probability at step n averages the product of these fractions across all possible sets of n distinct papers. Horizons use this exact probability, not simulation noise. The curves summarize the 1,000 sampled chains. A horizon of 100 reaches the end of our paper pool; it does not establish reliability beyond it.
How are papers sampled?
Finding every error is harder than finding one. The Verifiers working paper also documents lower detection recall when errors coexist. The two-error observations therefore add evidence that the single-error frontier alone cannot supply.
This is not simply the single-error recall number multiplied across papers. The pool combines one-error and two-error papers, and a two-error review succeeds only if both errors are found. We use the recorded outcomes of those reviews—not an assumption that detecting one error and detecting another are independent.
CRED uses controlled error injection to place known research errors into papers so that detection can be measured against a known reference. There are 100 A-only variants, 100 B-only variants, and 100 A+B variants: 300 observations containing 400 error instances.
All three arms use the same 100 underlying papers. For each chain we shuffle these papers and independently select one variant per paper: A-only, B-only, or A+B, each with probability one-third. Every paper appears exactly once in each 100-step chain. The mix is two-thirds one-error and one-third two-error papers on average; it varies across chains. Every verifier uses the same paper orders and variant choices.
Both errors must be localized in an A+B review. We use CRED's frozen localization scorer; chapter-label correctness is not required and false positives are not penalized. Here, “miss-free” means all errors in the benchmark reference were found, not a flawless review or a guarantee that the paper contains no other errors.
What does the color cost?
At paper n, color shows the mean cumulative verification cost of papers 1 through n, including paper n, across all 1,000 chains—not just those still miss-free. Costs continue accumulating after a miss. This measures the budget for verifying n papers, not spending until the first failure. The selector shows the dollar estimate for each effort.
Costs are reconstructed from recorded token usage at the same benchmark-registry list rates used for the frontier. These are mixed-date pricing snapshots, not a fresh uniform price quote. Subscription runs show API-equivalent costs, not actual per-paper charges. Paper generation, retries, and infrastructure are excluded. Unknown costs are gray, never silently treated as zero. Dollar bands are fixed across models and zoom levels.
What this does not measure
Inclusion follows the latest single-error cost–recall frontier: a model and harness qualify if at least one effort lies on that frontier, and its other recorded efforts are retained to show the effect of thinking effort. Mistral Small 3.2 is omitted. GPT-OSS 120B is retained as a comparison even though it is not on the current frontier. Historical frontier members that are no longer on the current frontier are not shown.
These are shuffled, previously recorded outcomes, not fresh model calls or an agent performing a connected workflow. Order changes when failure arrives; it does not change the stored review outcome. This does not measure context accumulation or errors propagating through citations and dependent research. One missed error need not invalidate a paper or an entire literature.
A missed error is observable to our benchmark scorer, not necessarily to a deployed verifier. The simulated first failure is therefore not an operational signal telling a production system when to stop.
Only configurations with all 300 valid observations receive a curve. Missing, failed, malformed, or mismatched runs stay pending. More shuffles reduce simulation noise; they do not add evidence about the model.