August 26 update: Explore the live benchmark of verifier cost-effectiveness.

Can we automate
policy evaluation?

AI may soon be capable of producing rigorous economic research. If that happens, policy evaluation could scale dramatically around the world: identifying what works, what fails, and what causes harm. To that end, we set up Project APE at the Social Catalyst Lab, led by Prof. David Yanagizawa-Drott, as an experiment testing whether an autonomous system can reliably generate, replicate, and verify empirical policy research. Everything remains public: papers, code, data, and failures. For a 3D visualisation of the project's global scope, click here. Project APE has also been by invitation at research and policy institutions internationally.

Read our latest research: Verifying the Verifiers

3,598Ideas
1000Papers
18k+Matches
5%Win Rate
⚠️ Warning: We are learning how to build a reliable, autonomous research system. Expect bugs, errors, hallucinations, and trashable papers. None of the generated papers have been peer-reviewed and should not be used for evidence-based policy making.
What does "autonomous" even mean?

Living benchmark

How reliable is automated verification—and at what cost?

Automated verification means an AI reads a finished research paper—its data, code, and claims—and tries to catch the errors in it, the way a careful referee would. The animation follows that capability's cost–recall frontier from 2024 to the latest benchmark release: each point is a tested verifier configuration on CRED's controlled error-injection benchmark, and the frontier moves whenever a model or harness delivers more recall at lower cost. For examples of how errors are missed by specific verifiers, see these reports on Gemini 3.8 Flash, Sonnet 5, and GPT-4o mini.

80 configurations APE-CRED v0.92
Most recent addition: Codex 6 Astra, 03-SEP-2026

Where the experiment stands

From generation to verification

01

First 1,000 papers

From January to April 2026, Project APE generated its first 1,000 papers and then paused production to evaluate the system. We then developed the CRED taxonomy to classify errors and defects in empirical research. In Verifying the Verifiers, the working paper released in July 2026, we used it to create a controlled test by inserting known errors into 100 papers, then measured how reliably—and at what cost—different AI verifiers found them. The rankings below preserve the automated tournament as it stood when production paused.

02

CRED evaluation in progress

In a separate sample of 100 unaltered papers, human reviewers agreed that 97% of the selected verifier's reported findings were genuine errors or defects. Based on that verifier's reports, we estimate that every paper contains errors or defects under Project APE's CRED taxonomy—often several, including some classified as severe. These are automated findings; humans did not check all 1,000 papers individually. We are still analyzing the results and writing the report.

Evaluation context
03

Next experiment

We are now prototyping an experiment with a second set of 1,000 paper ideas, testing whether and how the system can learn to improve itself.

The loop and its output. The loop itself ran without human input; the harness around it was re-engineered throughout — paper formats were deliberately varied, from longer and more ambitious to shorter no-revision runs, and a two-model co-authored variant with both Claude Code and Codex CLI joined the solo pipeline. Review in this era meant automated LLM referees — like journal referees, they read the paper but could not run the code — and they were themselves unverified. Independent verification (CRED) came after the freeze.

The first 1,000 papers

LLM-as-a-Judge Tournament

The leaderboard below records the first 1,000 papers as they stood when production was frozen for evaluation. It reflects a noisy triage signal from an LLM judge that read only each paper's PDF and assessed it as an economics journal editor might. It is therefore not peer review, a CRED-based score, or evidence that a paper is correct. Because the judge could not inspect code or data, it could not systematically assess the papers' underlying validity. We are now writing up findings from a separate CRED-based analysis of errors and defects in these papers. Future experiments will update both the leaderboard and how it is constructed, including by incorporating CRED-based verification.

How the Tournament Worked

Ranking Metrics

Review Status

Swipe to see more columns

Rank Paper μEstimated skill rating (μ). Higher values indicate better research quality based on pairwise comparisons. σUncertainty (σ). Lower values mean higher confidence in the rating. Cons.Conservative Rating (μ - 3σ), adjusted for integrity penalties. Used for ranking. EloElo rating. Standard chess-like rating where 400 points difference = 90% win probability. MPMatches Played. Valid head-to-head comparisons, excluding annulled matches against papers flagged with severe issues during automated code review. Status✅ Peer reviewed · 🔎 Awaiting review · 🧐 Issues detected · 🚫 Critical errors
140.11.635.22104190
237.21.532.81989171
335.51.231.91920185
435.81.431.61932158
535.31.331.61914155
635.01.231.41900165
734.91.231.41896195
834.91.231.31897188
934.81.231.31892191
1034.81.231.21893185
1134.61.231.01883151
1233.91.230.31856168
1333.81.230.11851169
1433.31.129.91830166
1540.03.429.821006
1633.01.129.61820164
1732.61.129.41803180
1832.61.129.31802180
1932.61.129.21803171
2032.41.129.21795173
2132.31.129.11792179
2232.21.128.81788194
2331.81.128.7177450
2431.91.128.6177447
2531.51.128.41761199
2631.31.128.11750177
2731.11.127.91742186
2831.01.127.81741203
2930.71.027.61727175
3035.82.827.419327
3137.13.327.119847
3230.31.027.11710168
3330.01.026.9169955
3429.51.026.61681188
3529.30.926.51672213
3629.41.026.41677180
3729.21.026.21669192
3829.11.026.11663198
3929.81.126.0169336
4028.91.026.01657186
4129.11.125.8166542
4235.72.725.419298
4339.33.425.320727
4428.11.025.21625184
4528.71.324.8164734
4627.61.024.71605201
4727.30.924.51593223
4828.01.224.5161937
4928.71.424.5164829
5028.61.224.5164532
5136.43.124.219556
5228.11.324.2162331
5327.81.224.1161239
5433.52.623.818398
5528.41.423.8163625
5628.31.523.7163227
5726.91.123.6157750
5837.53.423.620025
5928.51.623.6164020
6031.72.723.517696

Total tokens used for tournament (excludes paper generation tokens): 1,476,570,316

Papers competed in head-to-head matchups judged by Gemini 3.1 Flash Lite, prompted as a senior editor at a top economics journal. Each comparison used position-swapping to control for bias. Rankings are a noisy and potentially biased signal of quality, not ground truth.