—
BixBench
BixBench - A benchmark for evaluating AI agents on bioinformatics and computational biology tasks.
computational-biologybioinformatics-agentstabularcodetext
Status
curated
Saturation
unsaturated
Reproduction
not yet
Published results
5 models
Editor's note
Open-answer agentic bioinformatics: agents must run multi-step analyses over real datasets. Scores are low and the open-answer vs multiple-choice gap is large, so treat MCQ numbers with skepticism. Best public agents still miss roughly half the questions.
Published results
| Model | Score | Variant | Provenance |
|---|---|---|---|
| Biomni Lab (2026-02-03) | 52.2% accuracy | agent | external |
| Edison Analysis | 42.4% accuracy | agent | external |
| Claude Code (Opus 4.6) | 39.5% accuracy | agent | external |
| OpenAI Agents SDK (GPT-5.2) | 38.5% accuracy | agent | external |
| GPT-4o / Claude 3.5 Sonnet | 17.0% accuracy | open-answer (original paper) | paper |