Medical + biological AI
Independent analysis of medical and biological performance.
Compare recent AI models on published medical and biology benchmarks. We show reported scores and leave a gap when a model has not been evaluated.
Updated 4 September 2026
Leaderboard
Compare models by benchmark.
Medicine · Clinical work
HealthBench Professional · Astra
A same-source comparison on 525 physician-authored conversations across clinical consultation, documentation and research.
- Metric
- Length-adjusted rubric score (%)
- Published
- 2026
Note: OpenAI reports the best score at any effort. It independently evaluated the Claude rows with GPT-5.4 grading; this is OpenAI's comparison table, not each provider's own run.
Results
Higher scores are better. Effort appears in brackets when reported.
Public results by model.
These counts show how many charts on this page include a published score for each current model. “None yet” means we have not found a comparable public result.
| Model | Medical | Biology | Total |
|---|---|---|---|
| GPT-5.6 SolOpenAI · API | 15 | 18 | 33 |
| Claude Opus 5Anthropic · API | 12 | 14 | 26 |
| GPT-6 AstraOpenAI · API | 6 | 9 | 15 |
| Claude Sonnet 5Anthropic · API | 3 | 11 | 14 |
| Kimi K3Moonshot AI · Open weights | 5 | 6 | 11 |
| Gemini 3.8 FlashGoogle · API | 1 | 3 | 4 |
| Claude Fable 5.1Anthropic · API | 3 | None yet | 3 |
| DeepSeek V4 Pro 0813DeepSeek · API | 3 | None yet | 3 |
| Grok 4.6SpaceXAI · API | 1 | 2 | 3 |
| GPT-RosalindOpenAI · API | None yet | 1 | 1 |
| Qwen3.8-MaxAlibaba · Open weights | 1 | None yet | 1 |
| GLM-5.3Z.ai · API | None yet | None yet | 0 |
| Muse Spark 1.3Meta · API | None yet | None yet | 0 |