Headline results
Every score is agreement with the clinical review panel, computed from what the plans share. Frontier models saw the patient's data only; Diadia has the panel's clinical knowledge built in.
Five ways a treatment plan can hurt a patient: starting a treatment too early, adding one the clinicians never prescribed, leaving out one they did, getting a dose wrong, and putting steps in the wrong order. Lower is better.
% · lower is better · treatments the clinicians held back, started on day one
| Measure | Diadia | Grok 4.5 | Claude Opus 5 | Gemini 3.6 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Diadia | 16% | — | — | — | — | — |
| Grok 4.5 | — | 80% | — | — | — | — |
| Claude Opus 5 | — | — | 72% | — | — | — |
| Gemini 3.6 | — | — | — | 59% | — | — |
| GPT-5.6 Sol | — | — | — | — | 61% | — |
| GPT-5.6 Terra | — | — | — | — | — | 80% |
% · lower is better · share of the plan the clinicians never prescribed
| Measure | Diadia | Grok 4.5 | Claude Opus 5 | Gemini 3.6 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Diadia | 20% | — | — | — | — | — |
| Grok 4.5 | — | 51% | — | — | — | — |
| Claude Opus 5 | — | — | 61% | — | — | — |
| Gemini 3.6 | — | — | — | 48% | — | — |
| GPT-5.6 Sol | — | — | — | — | 73% | — |
| GPT-5.6 Terra | — | — | — | — | — | 70% |
% · lower is better · share of the clinicians' core treatments left out
| Measure | Diadia | Grok 4.5 | Claude Opus 5 | Gemini 3.6 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Diadia | 21% | — | — | — | — | — |
| Grok 4.5 | — | 42% | — | — | — | — |
| Claude Opus 5 | — | — | 40% | — | — | — |
| Gemini 3.6 | — | — | — | 56% | — | — |
| GPT-5.6 Sol | — | — | — | — | 59% | — |
| GPT-5.6 Terra | — | — | — | — | — | 63% |
% · higher is better · shared treatments dosed as prescribed
| Measure | Diadia | Grok 4.5 | Claude Opus 5 | Gemini 3.6 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Diadia | 70% | — | — | — | — | — |
| Grok 4.5 | — | 46% | — | — | — | — |
| Claude Opus 5 | — | — | 43% | — | — | — |
| Gemini 3.6 | — | — | — | 47% | — | — |
| GPT-5.6 Sol | — | — | — | — | 46% | — |
| GPT-5.6 Terra | — | — | — | — | — | 42% |
% · higher is better · shared treatments kept in the clinicians' order
| Measure | Diadia | Grok 4.5 | Claude Opus 5 | Gemini 3.6 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Diadia | 84% | — | — | — | — | — |
| Grok 4.5 | — | 54% | — | — | — | — |
| Claude Opus 5 | — | — | 58% | — | — | — |
| Gemini 3.6 | — | — | — | 59% | — | — |
| GPT-5.6 Sol | — | — | — | — | 54% | — |
| GPT-5.6 Terra | — | — | — | — | — | 59% |
Patient-safety measures, Diadia against the frontier
Scores are agreement with the panel of practising clinicians who reviewed and signed off each plan, computed from the treatments the plans share. Thin lines show the likely range.
Every score is agreement with the clinical review panel, computed from what the plans share. Frontier models saw the patient's data only; Diadia has the panel's clinical knowledge built in.
How closely each model matches a clinician-reviewed treatment plan. Score out of 100 for finding the same causes, choosing the same treatments and following the clinicians' own rules; 100 is an exact match.
© 2026 Diadia. All rights reserved.
© 2026 Diadia. All rights reserved.