Evals by capability

21 bio evaluations, grouped by the capability they probe. Tags come from each eval's bio taxonomy.

Single-cell & spatial omics

4

Analyzing single-cell and spatially-resolved biology data.

Bioinformatics agents & pipelines

2

End-to-end computational-biology workflows.

Genomics & sequence

1

Reasoning over genomic and sequence data.

Clinical & medical

2

Healthcare, medical imaging, and translational tasks.

Molecular & chemical

2

Molecular biology and chemistry.

Literature & knowledge

3

Biomedical QA and literature synthesis.

Scientific reasoning & discovery

6

Expert reasoning, research workflows, and replication.

BiomniBench-DataAnalysis (Biomni-DA-v0)
curated
Preliminary 15-task preview of Phylo's trace-based evaluation for long-horizon biological data analysis. It scores analytical process—including data handling, method selection, statistical rigor, source reliability, and reasoning—rather than only final-answer accuracy. Results are preliminary and the benchmark remains a work in progress.
top: Senior pharmaceutical scientists · 68.5% · biology-research, biomedical-data-analysis
LifeSciBench: Realistic, Expert-Level Life Science Tasks
curated
OpenAI's 750 expert-authored tasks spanning 7 scientific workflows × 7 life-science domains, each with an expert rubric — built to capture the ambiguity and judgment calls that knowledge-QA benchmarks miss (complex artifacts, situational ambiguity, open-ended answers). Far from saturated: no model passes 22.8% of tasks, and 34.8% have a best-model pass rate under 20%. Score shown is the problem-weighted normalized rubric score; top pass rate is 36.1% (GPT-Rosalind).
top: GPT-Rosalind · 57.6% · life-sciences, research-workflows
BixBench-Verified-50
curated
Phylo's expert-reviewed 50-task subset of BixBench removes, clarifies, or corrects ambiguous questions and flawed ground truths. Scores are substantially higher than on the full benchmark and must not be treated as directly comparable to full-BixBench results.
top: Biomni Lab (2026-02-03) · 88.7% · computational-biology, bioinformatics-agents
BioMedArena: Toolkit for Biomedical Deep-Research Agents
cataloged
Less a single benchmark than an open-source harness aggregating 166 biomedical benchmarks and 75 tools across 9 functional families, decoupling tool exposure / harness mode / context / scoring. Across 12 backbones, tool-equipped agents gain ~15 pts on average over prior SOTA on 8 representative benchmarks. Useful infrastructure to track.
scores pending · biomedical, deep-research
FrontierScience: Expert-Level Scientific Reasoning
cataloged
Expert / olympiad-level reasoning across physics, chemistry, and biology (160 problems). Only the biology subset is in-scope for us; unsaturated and hard. Already in inspect_evals.
scores pending · scientific-reasoning, biology-subset
PaperBench: Evaluating AI''s Ability to Replicate AI Research (Work In Progress)
cataloged
OpenAI's benchmark for replicating 20 ICML 2024 papers from scratch — not biology per se, but an agentic 'AI for science' boundary benchmark relevant to reproducible research workflows. Long-horizon and expensive to run. Already in inspect_evals.
scores pending · research-replication, agentic-science

Biosecurity & safety

1

Dual-use and hazardous-knowledge evaluations.