Biology-focused AI evaluations
BioInspect organizes AI benchmarks by scientific domain, modality, and task, with published results linked directly to their original sources. Each entry also includes biological taxonomy, reproducibility information, and concise editorial guidance.
21 evaluations · 10 with published results · press ⌘K to search
Evaluation catalog
| Eval | Domain | Status | # scored | Top result |
|---|---|---|---|---|
| scBench: A Benchmark for Single-Cell RNA-seq Analysis Practical single-cell RNA-seq analysis tasks — directly in the dbverse / spatial-omics wheelhouse and a high-value target for a bio-specialist leaderboard. The live verified leaderboard now has GPT-5.6 Sol at 62.1% with Pi and Opus 5 at 60.1% with Claude Code. | single-cell | curated | 10 | GPT-5.6 Sol · 62.1% |
| SpatialBench: Can Agents Analyze Real-World Spatial Biology Data? Real-world spatial biology analysis tasks across five technologies and seven task categories, with a 115-evaluation expert-verified subset. The live verified leaderboard has GPT-5.6 Sol at 74.5% with Codex and Opus 5 at 71.0% with Claude Code. | spatial-biology | cataloged | 9 | GPT-5.6 Sol · 74.5% |
| BiomniBench-DataAnalysis (Biomni-DA-v0) Preliminary 15-task preview of Phylo's trace-based evaluation for long-horizon biological data analysis. It scores analytical process—including data handling, method selection, statistical rigor, source reliability, and reasoning—rather than only final-answer accuracy. Results are preliminary and the benchmark remains a work in progress. | biology-research | curated | 6 | Senior pharmaceutical scientists · 68.5% |
| BixBench Open-answer agentic bioinformatics: agents must run multi-step analyses over real datasets. Scores are low and the open-answer vs multiple-choice gap is large, so treat MCQ numbers with skepticism. Best public agents still miss roughly half the questions. | computational-biology | curated | 5 | Biomni Lab (2026-02-03) · 52.2% |
| LifeSciBench: Realistic, Expert-Level Life Science Tasks OpenAI's 750 expert-authored tasks spanning 7 scientific workflows × 7 life-science domains, each with an expert rubric — built to capture the ambiguity and judgment calls that knowledge-QA benchmarks miss (complex artifacts, situational ambiguity, open-ended answers). Far from saturated: no model passes 22.8% of tasks, and 34.8% have a best-model pass rate under 20%. Score shown is the problem-weighted normalized rubric score; top pass rate is 36.1% (GPT-Rosalind). | life-sciences | curated | 5 | GPT-Rosalind · 57.6% |
| SpatialBench-Long: Verifiable Long-Horizon Spatial Biology 24 long-horizon spatial biology evaluations spanning CosMx, Visium, Xenium, MERFISH, Slide-seq, histology, and lineage data. The live leaderboard now has GPT-5.6 Sol at 38.9% with Codex and Opus 5 at 25.0% with Pi. | spatial-biology | curated | 5 | GPT-5.6 Sol · 38.9% |
| BixBench-Verified-50 Phylo's expert-reviewed 50-task subset of BixBench removes, clarifies, or corrects ambiguous questions and flawed ground truths. Scores are substantially higher than on the full benchmark and must not be treated as directly comparable to full-BixBench results. | computational-biology | curated | 4 | Biomni Lab (2026-02-03) · 88.7% |
| GeneBench-Pro: Multistage Statistical Reasoning in Genomics & Biology OpenAI's June 2026 benchmark for scientific 'research taste': 129 synthetic genomics / quantitative-biology / translational-medicine problems, each pairing a noisy dataset with a target estimand, graded deterministically from a known causal structure. Extremely hard and unsaturated. | genomics | curated | 4 | GPT-5.6 Sol Pro · 31.5% |
| scBench-Long: Verifiable Long-Horizon Single-Cell Biology 21 long-horizon single-cell evaluations across five research areas, deterministically graded with trajectory rubrics. The live leaderboard now has Opus 5 at 41.3% with Pi and GPT-5.6 Sol at 38.1% with Pi. | single-cell | curated | 4 | Opus 5 · 41.3% |
| PubMedQA: A Dataset for Biomedical Research Question Answering | biomedical | cataloged | 2 | Claude Opus 4.5 · 77.5% |
| BioAgent Bench: AI Agent Evaluation Suite for Bioinformatics End-to-end bioinformatics pipelines (RNA-seq, variant calling, metagenomics) with a robustness twist: stress-tests agents under corrupted inputs and decoy files, LLM-graded for pipeline progress and outcome validity. Finds frontier agents complete pipelines without heavy scaffolding but break under perturbation. ICML 2026. | bioinformatics | cataloged | 0 | — |
| BioMedArena: Toolkit for Biomedical Deep-Research Agents Less a single benchmark than an open-source harness aggregating 166 biomedical benchmarks and 75 tools across 9 functional families, decoupling tool exposure / harness mode / context / scoring. Across 12 backbones, tool-equipped agents gain ~15 pts on average over prior SOTA on 8 representative benchmarks. Useful infrastructure to track. | biomedical | cataloged | 0 | — |
| BioMysteryBench Molecular-biology reasoning with a hard split; reports accuracy on the human-solved subset. Newer and unsaturated — a useful difficulty signal for frontier bio reasoning. | molecular-biology | cataloged | 0 | — |
| ChemBench: Are large language models superhuman chemists? Broad chemistry knowledge and reasoning over 2,786 QA pairs. Chemistry-adjacent to biology (drug discovery, molecular properties); included as a boundary domain. Already in inspect_evals. | chemistry | cataloged | 0 | — |
| FrontierScience: Expert-Level Scientific Reasoning Expert / olympiad-level reasoning across physics, chemistry, and biology (160 problems). Only the biology subset is in-scope for us; unsaturated and hard. Already in inspect_evals. | scientific-reasoning | cataloged | 0 | — |
| HealthBench: Evaluating Large Language Models Towards Improved Human Health OpenAI's rubric-graded evaluation of medical/health capabilities across realistic healthcare conversations. Model-graded against physician-written rubrics, so scoring cost is nontrivial. Already in inspect_evals. | clinical | cataloged | 0 | — |
| LAB-Bench: Measuring Capabilities of Language Models for Biology Research Broad biology-research QA (literature QA, protocols, figure reading, sequence manipulation). Already implemented in inspect_evals, so it is the natural first target to reproduce ourselves with inspect_ai. Verified per-model scores pending our own runs. | biology-research | cataloged | 0 | — |
| LAB-Bench2: Improved Benchmark for AI Systems Performing Biology Research FutureHouse's successor to LAB-Bench: ~1,900 tasks in more realistic research contexts, and markedly harder — model accuracy drops 26–46% across subtasks vs the original. Supersedes LAB-Bench (v1) as the current version. | biology-research | cataloged | 0 | — |
| PaperBench: Evaluating AI''s Ability to Replicate AI Research (Work In Progress) OpenAI's benchmark for replicating 20 ICML 2024 papers from scratch — not biology per se, but an agentic 'AI for science' boundary benchmark relevant to reproducible research workflows. Long-horizon and expensive to run. Already in inspect_evals. | research-replication | cataloged | 0 | — |
| VQA-RAD: Visual Question Answering for Radiology Clinician-authored visual QA over radiology images — a multimodal medical benchmark. Tests vision-language models, not text-only. Already in inspect_evals. | radiology | cataloged | 0 | — |
| WMDP: Measuring and Reducing Malicious Use With Unlearning Proxy measurement of hazardous knowledge; the WMDP-bio subset is the biology slice, used for unlearning research. A safety benchmark — higher is not 'better'. Include with that framing. | biosecurity | cataloged | 0 | — |