Skip to content

Medical + biological AI

Independent analysis of medical and biological performance.

Compare recent AI models on published medical and biology benchmarks. We show reported scores and leave a gap when a model has not been evaluated.

Updated 4 September 2026

Leaderboard

Compare models by benchmark.

Compare medical knowledge with the harder work: navigating EHRs, calling tools, writing SQL, diagnosing cases, handling prior authorization and running biological analyses. You can remove models that are still waiting for an evaluation.
Area
Models

Medicine · Clinical work

HealthBench Professional · Astra

A same-source comparison on 525 physician-authored conversations across clinical consultation, documentation and research.

Metric
Length-adjusted rubric score (%)
Published
2026

Note: OpenAI reports the best score at any effort. It independently evaluated the Claude rows with GPT-5.4 grading; this is OpenAI's comparison table, not each provider's own run.

Results

Higher scores are better. Effort appears in brackets when reported.

GPT-6 Astra
Claude Fable 5
GPT-5.6 Sol
Claude Fable 5.1
Claude Opus 5
Gemini 3.8 Flash
Length-adjusted rubric score (%) · OpenAI Astra launch comparison · 3 Sep 2026OpenAI · GPT-6 Astra launch report

Public results by model.

These counts show how many charts on this page include a published score for each current model. “None yet” means we have not found a comparable public result.

ModelMedicalBiologyTotal
GPT-5.6 SolOpenAI · API151833
Claude Opus 5Anthropic · API121426
GPT-6 AstraOpenAI · API6915
Claude Sonnet 5Anthropic · API31114
Kimi K3Moonshot AI · Open weights5611
Gemini 3.8 FlashGoogle · API134
Claude Fable 5.1Anthropic · API3None yet3
DeepSeek V4 Pro 0813DeepSeek · API3None yet3
Grok 4.6SpaceXAI · API123
GPT-RosalindOpenAI · APINone yet11
Qwen3.8-MaxAlibaba · Open weights1None yet1
GLM-5.3Z.ai · APINone yetNone yet0
Muse Spark 1.3Meta · APINone yetNone yet0