Phylo

BixBench-Verified-50

Phylo's expert-reviewed 50-task subset of BixBench removes, clarifies, or corrects ambiguous questions and flawed ground truths. Scores are substantially higher than on the full benchmark and must not be treated as directly comparable to full-BixBench results.

computational-biologybioinformatics-agentsbenchmark-qualitytabularcodetext
Status
curated
Saturation
partial
Reproduction
not yet
Published results
4 models

Editor's note

Phylo's expert-reviewed 50-task subset of BixBench removes, clarifies, or corrects ambiguous questions and flawed ground truths. Scores are substantially higher than on the full benchmark and must not be treated as directly comparable to full-BixBench results.

Published results

ModelScoreVariantProvenance
Biomni Lab (2026-02-03)88.7% accuracyverified 50-task subsetexternal
Edison Analysis78.0% accuracyverified 50-task subsetexternal
Claude Code (Opus 4.6)65.3% accuracyverified 50-task subsetexternal
OpenAI Agents SDK (GPT-5.2)61.3% accuracyverified 50-task subsetexternal

Sources

  • external Phylo — Evaluating AI Agents in Biology
  • dataset verified benchmark dataset