Results
Compare model results reported by benchmark papers and official leaderboards. BioInspect reruns are attached to their individual evaluation records.
Source-reported scores are shown as published and are not combined into a single index because benchmarks may use different metrics, model versions, prompts, and harnesses.
Results within one benchmark
GPT-5.6 Sol
62.1%
Opus 5
60.1%
Claude Opus 4.6
52.8%
Claude Opus 4.5
49.9%
GPT-5.2
45.2%
Claude Sonnet 4.5
44.2%
GPT-5.1
37.9%
Grok-4.1
35.6%
Grok-4
33.9%
Gemini 2.5 Pro
29.2%
Cross-benchmark results matrix
Source-reported results are shown without aggregation. Blank cells indicate that the source did not report that model on the benchmark. Scores may use different metrics, variants, and evaluation harnesses and should not be treated as directly comparable across columns.
| Model | scBench: A Benchmark for Single-Cell RNA-seq Analysis | SpatialBench: Can Agents Analyze Real-World Spatial Biology Data? | BiomniBench-DataAnalysis (Biomni-DA-v0) | BixBench | LifeSciBench: Realistic, Expert-Level Life Science Tasks | SpatialBench-Long: Verifiable Long-Horizon Spatial Biology | BixBench-Verified-50 | GeneBench-Pro: Multistage Statistical Reasoning in Genomics & Biology | scBench-Long: Verifiable Long-Horizon Single-Cell Biology | PubMedQA: A Dataset for Biomedical Research Question Answering |
|---|---|---|---|---|---|---|---|---|---|---|
| Biomni Lab (2026-02-03) | — | — | 65.0% | 52.2% | — | — | 88.7% | — | — | — |
| Claude Code (Opus 4.6) | — | — | 47.8% | 39.5% | — | — | 65.3% | — | — | — |
| Claude Opus 4.5 | 49.9% | 38.4% | — | — | — | — | — | — | — | 77.5% |
| Claude Opus 4.6 | 52.8% | — | — | — | — | — | — | — | — | — |
| Claude Opus 4.8 | — | — | — | — | — | — | — | 16.0% | 25.4% | — |
| Claude Sonnet 4.5 | 44.2% | 28.3% | — | — | — | — | — | — | — | 0.0% |
| Edison Analysis | — | — | 52.7% | 42.4% | — | — | 78.0% | — | — | — |
| GPT-4o / Claude 3.5 Sonnet | — | — | — | 17.0% | — | — | — | — | — | — |
| GPT-5.1 | 37.9% | 27.4% | — | — | — | — | — | — | — | — |
| GPT-5.2 | 45.2% | 34.0% | — | — | — | — | — | — | — | — |
| GPT-5.4 | — | — | — | — | 47.9% | — | — | — | — | — |
| GPT-5.5 | — | — | — | — | 51.9% | 11.1% | — | — | — | — |
| GPT-5.6 Sol | 62.1% | 74.5% | — | — | — | 38.9% | — | 28.7% | 38.1% | — |
| GPT-5.6 Sol Pro | — | — | — | — | — | — | — | 31.5% | — | — |
| GPT-Rosalind | — | — | — | — | 57.6% | — | — | — | — | — |
| Gemini 2.5 Pro | 29.2% | 20.1% | — | — | — | — | — | — | — | — |
| Gemini 3.1 Pro | — | — | — | — | 51.5% | — | — | — | — | — |
| Gemini 3.5 Flash | — | — | — | — | — | 11.1% | — | 8.1% | 22.2% | — |
| Grok 4.3 | — | — | — | — | 39.9% | — | — | — | — | — |
| Grok-4 | 33.9% | 22.8% | — | — | — | — | — | — | — | — |
| Grok-4.1 | 35.6% | 24.7% | — | — | — | — | — | — | — | — |
| Junior scientists | — | — | 48.5% | — | — | — | — | — | — | — |
| OpenAI Agents SDK (GPT-5.2) | — | — | 51.4% | 38.5% | — | — | 61.3% | — | — | — |
| Opus 5 | 60.1% | 71.0% | — | — | — | 25.0% | — | — | 41.3% | — |
| Senior pharmaceutical scientists | — | — | 68.5% | — | — | — | — | — | — | — |