Phylo
BixBench-Verified-50
Phylo's expert-reviewed 50-task subset of BixBench removes, clarifies, or corrects ambiguous questions and flawed ground truths. Scores are substantially higher than on the full benchmark and must not be treated as directly comparable to full-BixBench results.
computational-biologybioinformatics-agentsbenchmark-qualitytabularcodetext
Status
curated
Saturation
partial
Reproduction
not yet
Published results
4 models
Editor's note
Phylo's expert-reviewed 50-task subset of BixBench removes, clarifies, or corrects ambiguous questions and flawed ground truths. Scores are substantially higher than on the full benchmark and must not be treated as directly comparable to full-BixBench results.
Published results
| Model | Score | Variant | Provenance |
|---|---|---|---|
| Biomni Lab (2026-02-03) | 88.7% accuracy | verified 50-task subset | external |
| Edison Analysis | 78.0% accuracy | verified 50-task subset | external |
| Claude Code (Opus 4.6) | 65.3% accuracy | verified 50-task subset | external |
| OpenAI Agents SDK (GPT-5.2) | 61.3% accuracy | verified 50-task subset | external |