
OpenAI News
· 1 min read
Introducing HealthBench
An evaluation for AI systems and human health.
Improving human health will be one of the defining impacts of AGI. If developed and deployed effectively, large language models have the potential to expand access to health information, support clinicians in delivering high-quality care, and help people advocate for their health and that of their communities.
To get there, we need to ensure models are useful and safe. Evaluations are essential to understanding how models perform in health settings. Significant efforts have already been made across academia and industry, yet many existing evaluations do not reflect realistic scenarios, lack rigorous validation against expert medical opinion, or leave no room for state-of-the-art models to improve.
Today, we’re introducing HealthBench: a new benchmark designed to better measure capabilities of AI systems for health. Built in partnership with 262 physicians who have practiced in 60 countries, HealthBench includes 5,000 realistic health conversations, each with a custom physician-created rubric to grade model responses.
HealthBench is grounded in our belief that evaluations for AI systems in health should be:
Meaningful: Scores reflect real-world impact. This should go beyond exam questions to capture complex, real-life scenarios and workflows that mirror the ways individuals and clinicians interact with models.
Trustworthy: Scores are faithful indicators of physician judgment. Evaluations should reflect the standards and priorities of healthcare professionals, providing a rigorous foundation for improving AI systems.
Unsaturated: Benchmarks support progress. Current models should show substantial room for improvement, offering model developers incentives to continuously improve performance.
Alongside the HealthBench benchmark, we’re also sharing how several of our models perform, setting a new baseline to improve upon.
Dataset description
Themes
Axes
Countries of practice
Performance of models
Cost
Reliability
The HealthBench family
Original source
This story was published by OpenAI News. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on openai.com


