Center for AI Safety (CAIS)

WMDP: Measuring and Reducing Malicious Use With Unlearning

A dataset of 3,668 multiple-choice questions developed by a consortium of academics and technical consultants that serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security.

biosecuritysafetytext
Status
cataloged
Saturation
partial
Reproduction
not yet
Published results
0 models

Editor's note

Proxy measurement of hazardous knowledge; the WMDP-bio subset is the biology slice, used for unlearning research. A safety benchmark — higher is not 'better'. Include with that framing.

Published results

No source-reported scores have been added to BioInspect yet.

Reproduce: inspect_evals implementation

Sources