BixBench

BixBench - A benchmark for evaluating AI agents on bioinformatics and computational biology tasks.

computational-biologybioinformatics-agentstabularcodetext
Status
curated
Saturation
unsaturated
Reproduction
not yet
Published results
5 models

Editor's note

Open-answer agentic bioinformatics: agents must run multi-step analyses over real datasets. Scores are low and the open-answer vs multiple-choice gap is large, so treat MCQ numbers with skepticism. Best public agents still miss roughly half the questions.

Published results

ModelScoreVariantProvenance
Biomni Lab (2026-02-03)52.2% accuracyagentexternal
Edison Analysis42.4% accuracyagentexternal
Claude Code (Opus 4.6)39.5% accuracyagentexternal
OpenAI Agents SDK (GPT-5.2)38.5% accuracyagentexternal
GPT-4o / Claude 3.5 Sonnet17.0% accuracyopen-answer (original paper)paper

Sources

  • sophon eval metadata
  • external Phylo — Evaluating AI Agents in Biology
  • paper BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology