OpenAI
HealthBench: Evaluating Large Language Models Towards Improved Human Health
A comprehensive evaluation benchmark designed to assess language models' medical capabilities across a wide range of healthcare scenarios.
clinicalhealthcaretext
Status
cataloged
Saturation
partial
Reproduction
not yet
Published results
0 models
Editor's note
OpenAI's rubric-graded evaluation of medical/health capabilities across realistic healthcare conversations. Model-graded against physician-written rubrics, so scoring cost is nontrivial. Already in inspect_evals.
Published results
No source-reported scores have been added to BioInspect yet.
Reproduce: inspect_evals implementation
Sources
- sophon eval metadata