scBench: A Benchmark for Single-Cell RNA-seq Analysis

Evaluates whether models can solve practical single-cell RNA-seq analysis tasks with deterministic grading. Tasks require empirical interaction with .h5ad data files - agents must load and analyze the data to produce correct answers. Covers 30 canonical tasks across 5 sequencing

single-celltranscriptomicstabularcode
Status
curated
Saturation
unsaturated
Reproduction
not yet
Published results
10 models

Editor's note

Practical single-cell RNA-seq analysis tasks — directly in the dbverse / spatial-omics wheelhouse and a high-value target for a bio-specialist leaderboard. The live verified leaderboard now has GPT-5.6 Sol at 62.1% with Pi and Opus 5 at 60.1% with Claude Code.

Published results

ModelScoreVariantProvenance
GPT-5.6 Sol62.1% accuracyPi harnessexternal
Opus 560.1% accuracyClaude Code harnessexternal
Claude Opus 4.652.8% accuracydefaultpaper
Claude Opus 4.549.9% accuracydefaultpaper
GPT-5.245.2% accuracydefaultpaper
Claude Sonnet 4.544.2% accuracydefaultpaper
GPT-5.137.9% accuracydefaultpaper
Grok-4.135.6% accuracydefaultpaper
Grok-433.9% accuracydefaultpaper
Gemini 2.5 Pro29.2% accuracydefaultpaper

Sources

  • sophon eval metadata
  • external benchmarks.bio — scBench verified leaderboard
  • paper scBench (arXiv:2602.09063)