FutureHouse

LAB-Bench: Measuring Capabilities of Language Models for Biology Research

Tests LLMs and LLM-augmented agents abilities to answer questions on scientific research workflows in domains like chemistry, biology, materials science, as well as more general science tasks

biology-researchliterature-qatextfigures
Status
cataloged
Saturation
partial
Reproduction
not yet
Published results
0 models

Editor's note

Broad biology-research QA (literature QA, protocols, figure reading, sequence manipulation). Already implemented in inspect_evals, so it is the natural first target to reproduce ourselves with inspect_ai. Verified per-model scores pending our own runs.

Published results

No source-reported scores have been added to BioInspect yet.

Reproduce: inspect_evals implementation

Sources