Wu et al. (Oxford / Clifton lab)

BioMedArena: Toolkit for Biomedical Deep-Research Agents

Less a single benchmark than an open-source harness aggregating 166 biomedical benchmarks and 75 tools across 9 functional families, decoupling tool exposure / harness mode / context / scoring. Across 12 backbones, tool-equipped agents gain ~15 pts on average over prior SOTA on 8 representative benchmarks. Useful infrastructure to track.

biomedicaldeep-researchtexttools
Status
cataloged
Saturation
partial
Reproduction
not yet
Published results
0 models

Editor's note

Less a single benchmark than an open-source harness aggregating 166 biomedical benchmarks and 75 tools across 9 functional families, decoupling tool exposure / harness mode / context / scoring. Across 12 backbones, tool-equipped agents gain ~15 pts on average over prior SOTA on 8 representative benchmarks. Useful infrastructure to track.

Published results

No source-reported scores have been added to BioInspect yet.

Sources