FutureHouse
LAB-Bench2: Improved Benchmark for AI Systems Performing Biology Research
FutureHouse's successor to LAB-Bench: ~1,900 tasks in more realistic research contexts, and markedly harder — model accuracy drops 26–46% across subtasks vs the original. Supersedes LAB-Bench (v1) as the current version.
biology-researchliterature-qatextfigures
Status
cataloged
Saturation
unsaturated
Reproduction
not yet
Published results
0 models
Editor's note
FutureHouse's successor to LAB-Bench: ~1,900 tasks in more realistic research contexts, and markedly harder — model accuracy drops 26–46% across subtasks vs the original. Supersedes LAB-Bench (v1) as the current version.
Published results
No source-reported scores have been added to BioInspect yet.
Sources
- paper benchmark paper