Biology-focused AI evaluations

BioInspect organizes AI benchmarks by scientific domain, modality, and task, with published results linked directly to their original sources. Each entry also includes biological taxonomy, reproducibility information, and concise editorial guidance.

21 evaluations · 10 with published results · press ⌘K to search

Evaluation catalog

EvalDomainStatus# scoredTop result
scBench: A Benchmark for Single-Cell RNA-seq Analysis
Practical single-cell RNA-seq analysis tasks — directly in the dbverse / spatial-omics wheelhouse and a high-value target for a bio-specialist leaderboard. The live verified leaderboard now has GPT-5.6 Sol at 62.1% with Pi and Opus 5 at 60.1% with Claude Code.
single-cellcurated10GPT-5.6 Sol · 62.1%
SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
Real-world spatial biology analysis tasks across five technologies and seven task categories, with a 115-evaluation expert-verified subset. The live verified leaderboard has GPT-5.6 Sol at 74.5% with Codex and Opus 5 at 71.0% with Claude Code.
spatial-biologycataloged9GPT-5.6 Sol · 74.5%
BiomniBench-DataAnalysis (Biomni-DA-v0)
Preliminary 15-task preview of Phylo's trace-based evaluation for long-horizon biological data analysis. It scores analytical process—including data handling, method selection, statistical rigor, source reliability, and reasoning—rather than only final-answer accuracy. Results are preliminary and the benchmark remains a work in progress.
biology-researchcurated6Senior pharmaceutical scientists · 68.5%
BixBench
Open-answer agentic bioinformatics: agents must run multi-step analyses over real datasets. Scores are low and the open-answer vs multiple-choice gap is large, so treat MCQ numbers with skepticism. Best public agents still miss roughly half the questions.
computational-biologycurated5Biomni Lab (2026-02-03) · 52.2%
LifeSciBench: Realistic, Expert-Level Life Science Tasks
OpenAI's 750 expert-authored tasks spanning 7 scientific workflows × 7 life-science domains, each with an expert rubric — built to capture the ambiguity and judgment calls that knowledge-QA benchmarks miss (complex artifacts, situational ambiguity, open-ended answers). Far from saturated: no model passes 22.8% of tasks, and 34.8% have a best-model pass rate under 20%. Score shown is the problem-weighted normalized rubric score; top pass rate is 36.1% (GPT-Rosalind).
life-sciencescurated5GPT-Rosalind · 57.6%
SpatialBench-Long: Verifiable Long-Horizon Spatial Biology
24 long-horizon spatial biology evaluations spanning CosMx, Visium, Xenium, MERFISH, Slide-seq, histology, and lineage data. The live leaderboard now has GPT-5.6 Sol at 38.9% with Codex and Opus 5 at 25.0% with Pi.
spatial-biologycurated5GPT-5.6 Sol · 38.9%
BixBench-Verified-50
Phylo's expert-reviewed 50-task subset of BixBench removes, clarifies, or corrects ambiguous questions and flawed ground truths. Scores are substantially higher than on the full benchmark and must not be treated as directly comparable to full-BixBench results.
computational-biologycurated4Biomni Lab (2026-02-03) · 88.7%
GeneBench-Pro: Multistage Statistical Reasoning in Genomics & Biology
OpenAI's June 2026 benchmark for scientific 'research taste': 129 synthetic genomics / quantitative-biology / translational-medicine problems, each pairing a noisy dataset with a target estimand, graded deterministically from a known causal structure. Extremely hard and unsaturated.
genomicscurated4GPT-5.6 Sol Pro · 31.5%
scBench-Long: Verifiable Long-Horizon Single-Cell Biology
21 long-horizon single-cell evaluations across five research areas, deterministically graded with trajectory rubrics. The live leaderboard now has Opus 5 at 41.3% with Pi and GPT-5.6 Sol at 38.1% with Pi.
single-cellcurated4Opus 5 · 41.3%
PubMedQA: A Dataset for Biomedical Research Question Answering
biomedicalcataloged2Claude Opus 4.5 · 77.5%
BioAgent Bench: AI Agent Evaluation Suite for Bioinformatics
End-to-end bioinformatics pipelines (RNA-seq, variant calling, metagenomics) with a robustness twist: stress-tests agents under corrupted inputs and decoy files, LLM-graded for pipeline progress and outcome validity. Finds frontier agents complete pipelines without heavy scaffolding but break under perturbation. ICML 2026.
bioinformaticscataloged0
BioMedArena: Toolkit for Biomedical Deep-Research Agents
Less a single benchmark than an open-source harness aggregating 166 biomedical benchmarks and 75 tools across 9 functional families, decoupling tool exposure / harness mode / context / scoring. Across 12 backbones, tool-equipped agents gain ~15 pts on average over prior SOTA on 8 representative benchmarks. Useful infrastructure to track.
biomedicalcataloged0
BioMysteryBench
Molecular-biology reasoning with a hard split; reports accuracy on the human-solved subset. Newer and unsaturated — a useful difficulty signal for frontier bio reasoning.
molecular-biologycataloged0
ChemBench: Are large language models superhuman chemists?
Broad chemistry knowledge and reasoning over 2,786 QA pairs. Chemistry-adjacent to biology (drug discovery, molecular properties); included as a boundary domain. Already in inspect_evals.
chemistrycataloged0
FrontierScience: Expert-Level Scientific Reasoning
Expert / olympiad-level reasoning across physics, chemistry, and biology (160 problems). Only the biology subset is in-scope for us; unsaturated and hard. Already in inspect_evals.
scientific-reasoningcataloged0
HealthBench: Evaluating Large Language Models Towards Improved Human Health
OpenAI's rubric-graded evaluation of medical/health capabilities across realistic healthcare conversations. Model-graded against physician-written rubrics, so scoring cost is nontrivial. Already in inspect_evals.
clinicalcataloged0
LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Broad biology-research QA (literature QA, protocols, figure reading, sequence manipulation). Already implemented in inspect_evals, so it is the natural first target to reproduce ourselves with inspect_ai. Verified per-model scores pending our own runs.
biology-researchcataloged0
LAB-Bench2: Improved Benchmark for AI Systems Performing Biology Research
FutureHouse's successor to LAB-Bench: ~1,900 tasks in more realistic research contexts, and markedly harder — model accuracy drops 26–46% across subtasks vs the original. Supersedes LAB-Bench (v1) as the current version.
biology-researchcataloged0
PaperBench: Evaluating AI''s Ability to Replicate AI Research (Work In Progress)
OpenAI's benchmark for replicating 20 ICML 2024 papers from scratch — not biology per se, but an agentic 'AI for science' boundary benchmark relevant to reproducible research workflows. Long-horizon and expensive to run. Already in inspect_evals.
research-replicationcataloged0
VQA-RAD: Visual Question Answering for Radiology
Clinician-authored visual QA over radiology images — a multimodal medical benchmark. Tests vision-language models, not text-only. Already in inspect_evals.
radiologycataloged0
WMDP: Measuring and Reducing Malicious Use With Unlearning
Proxy measurement of hazardous knowledge; the WMDP-bio subset is the biology slice, used for unlearning research. A safety benchmark — higher is not 'better'. Include with that framing.
biosecuritycataloged0