OpenAI

HealthBench: Evaluating Large Language Models Towards Improved Human Health

A comprehensive evaluation benchmark designed to assess language models' medical capabilities across a wide range of healthcare scenarios.

clinicalhealthcaretext
Status
cataloged
Saturation
partial
Reproduction
not yet
Published results
0 models

Editor's note

OpenAI's rubric-graded evaluation of medical/health capabilities across realistic healthcare conversations. Model-graded against physician-written rubrics, so scoring cost is nontrivial. Already in inspect_evals.

Published results

No source-reported scores have been added to BioInspect yet.

Reproduce: inspect_evals implementation

Sources