OpenAI
PaperBench: Evaluating AI''s Ability to Replicate AI Research (Work In Progress)
Agents are evaluated on their ability to replicate 20 ICML 2024 Spotlight and Oral papers from scratch. Given a research paper PDF, an addendum with clarifications, and a rubric defining evaluation criteria, the agent must
research-replicationagentic-sciencetextcode
Status
cataloged
Saturation
unsaturated
Reproduction
not yet
Published results
0 models
Editor's note
OpenAI's benchmark for replicating 20 ICML 2024 papers from scratch — not biology per se, but an agentic 'AI for science' boundary benchmark relevant to reproducible research workflows. Long-horizon and expensive to run. Already in inspect_evals.
Published results
No source-reported scores have been added to BioInspect yet.
Reproduce: inspect_evals implementation
Sources
- sophon eval metadata