Evals by capability
21 bio evaluations, grouped by the capability they probe. Tags come from each eval's bio taxonomy.
Single-cell & spatial omics
4Analyzing single-cell and spatially-resolved biology data.
scBench: A Benchmark for Single-Cell RNA-seq Analysis
curatedPractical single-cell RNA-seq analysis tasks — directly in the dbverse / spatial-omics wheelhouse and a high-value target for a bio-specialist leaderboard. The live verified leaderboard now has GPT-5.6 Sol at 62.1% with Pi and Opus 5 at 60.1% with Claude Code.
top: GPT-5.6 Sol · 62.1% · single-cell, transcriptomics
SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
catalogedReal-world spatial biology analysis tasks across five technologies and seven task categories, with a 115-evaluation expert-verified subset. The live verified leaderboard has GPT-5.6 Sol at 74.5% with Codex and Opus 5 at 71.0% with Claude Code.
top: GPT-5.6 Sol · 74.5% · spatial-biology, spatial-transcriptomics
SpatialBench-Long: Verifiable Long-Horizon Spatial Biology
curated24 long-horizon spatial biology evaluations spanning CosMx, Visium, Xenium, MERFISH, Slide-seq, histology, and lineage data. The live leaderboard now has GPT-5.6 Sol at 38.9% with Codex and Opus 5 at 25.0% with Pi.
top: GPT-5.6 Sol · 38.9% · spatial-biology, spatial-transcriptomics
scBench-Long: Verifiable Long-Horizon Single-Cell Biology
curated21 long-horizon single-cell evaluations across five research areas, deterministically graded with trajectory rubrics. The live leaderboard now has Opus 5 at 41.3% with Pi and GPT-5.6 Sol at 38.1% with Pi.
top: Opus 5 · 41.3% · single-cell, transcriptomics
Bioinformatics agents & pipelines
2End-to-end computational-biology workflows.
BixBench
curatedOpen-answer agentic bioinformatics: agents must run multi-step analyses over real datasets. Scores are low and the open-answer vs multiple-choice gap is large, so treat MCQ numbers with skepticism. Best public agents still miss roughly half the questions.
top: Biomni Lab (2026-02-03) · 52.2% · computational-biology, bioinformatics-agents
BioAgent Bench: AI Agent Evaluation Suite for Bioinformatics
catalogedEnd-to-end bioinformatics pipelines (RNA-seq, variant calling, metagenomics) with a robustness twist: stress-tests agents under corrupted inputs and decoy files, LLM-graded for pipeline progress and outcome validity. Finds frontier agents complete pipelines without heavy scaffolding but break under perturbation. ICML 2026.
scores pending · bioinformatics, agentic-pipelines
Genomics & sequence
1Reasoning over genomic and sequence data.
Clinical & medical
2Healthcare, medical imaging, and translational tasks.
HealthBench: Evaluating Large Language Models Towards Improved Human Health
catalogedOpenAI's rubric-graded evaluation of medical/health capabilities across realistic healthcare conversations. Model-graded against physician-written rubrics, so scoring cost is nontrivial. Already in inspect_evals.
scores pending · clinical, healthcare
VQA-RAD: Visual Question Answering for Radiology
catalogedClinician-authored visual QA over radiology images — a multimodal medical benchmark. Tests vision-language models, not text-only. Already in inspect_evals.
scores pending · radiology, medical-imaging
Molecular & chemical
2Molecular biology and chemistry.
BioMysteryBench
catalogedMolecular-biology reasoning with a hard split; reports accuracy on the human-solved subset. Newer and unsaturated — a useful difficulty signal for frontier bio reasoning.
scores pending · molecular-biology, reasoning
ChemBench: Are large language models superhuman chemists?
catalogedBroad chemistry knowledge and reasoning over 2,786 QA pairs. Chemistry-adjacent to biology (drug discovery, molecular properties); included as a boundary domain. Already in inspect_evals.
scores pending · chemistry, drug-discovery-adjacent
Literature & knowledge
3Biomedical QA and literature synthesis.
PubMedQA: A Dataset for Biomedical Research Question Answering
catalogedtop: Claude Opus 4.5 · 77.5% · biomedical, literature-qa
LAB-Bench: Measuring Capabilities of Language Models for Biology Research
catalogedBroad biology-research QA (literature QA, protocols, figure reading, sequence manipulation). Already implemented in inspect_evals, so it is the natural first target to reproduce ourselves with inspect_ai. Verified per-model scores pending our own runs.
scores pending · biology-research, literature-qa
LAB-Bench2: Improved Benchmark for AI Systems Performing Biology Research
catalogedFutureHouse's successor to LAB-Bench: ~1,900 tasks in more realistic research contexts, and markedly harder — model accuracy drops 26–46% across subtasks vs the original. Supersedes LAB-Bench (v1) as the current version.
scores pending · biology-research, literature-qa
Scientific reasoning & discovery
6Expert reasoning, research workflows, and replication.
BiomniBench-DataAnalysis (Biomni-DA-v0)
curatedPreliminary 15-task preview of Phylo's trace-based evaluation for long-horizon biological data analysis. It scores analytical process—including data handling, method selection, statistical rigor, source reliability, and reasoning—rather than only final-answer accuracy. Results are preliminary and the benchmark remains a work in progress.
top: Senior pharmaceutical scientists · 68.5% · biology-research, biomedical-data-analysis
LifeSciBench: Realistic, Expert-Level Life Science Tasks
curatedOpenAI's 750 expert-authored tasks spanning 7 scientific workflows × 7 life-science domains, each with an expert rubric — built to capture the ambiguity and judgment calls that knowledge-QA benchmarks miss (complex artifacts, situational ambiguity, open-ended answers). Far from saturated: no model passes 22.8% of tasks, and 34.8% have a best-model pass rate under 20%. Score shown is the problem-weighted normalized rubric score; top pass rate is 36.1% (GPT-Rosalind).
top: GPT-Rosalind · 57.6% · life-sciences, research-workflows
BixBench-Verified-50
curatedPhylo's expert-reviewed 50-task subset of BixBench removes, clarifies, or corrects ambiguous questions and flawed ground truths. Scores are substantially higher than on the full benchmark and must not be treated as directly comparable to full-BixBench results.
top: Biomni Lab (2026-02-03) · 88.7% · computational-biology, bioinformatics-agents
BioMedArena: Toolkit for Biomedical Deep-Research Agents
catalogedLess a single benchmark than an open-source harness aggregating 166 biomedical benchmarks and 75 tools across 9 functional families, decoupling tool exposure / harness mode / context / scoring. Across 12 backbones, tool-equipped agents gain ~15 pts on average over prior SOTA on 8 representative benchmarks. Useful infrastructure to track.
scores pending · biomedical, deep-research
FrontierScience: Expert-Level Scientific Reasoning
catalogedExpert / olympiad-level reasoning across physics, chemistry, and biology (160 problems). Only the biology subset is in-scope for us; unsaturated and hard. Already in inspect_evals.
scores pending · scientific-reasoning, biology-subset
PaperBench: Evaluating AI''s Ability to Replicate AI Research (Work In Progress)
catalogedOpenAI's benchmark for replicating 20 ICML 2024 papers from scratch — not biology per se, but an agentic 'AI for science' boundary benchmark relevant to reproducible research workflows. Long-horizon and expensive to run. Already in inspect_evals.
scores pending · research-replication, agentic-science
Biosecurity & safety
1Dual-use and hazardous-knowledge evaluations.