Diadia Functional Medicine Benchmark

Functional Medicine Benchmark

Diadia and five frontier models wrote treatment plans for the same patients from the same data. Each plan was compared with the plan a panel of practising clinicians reviewed and signed off: the same causes, the same treatments, the clinicians' own rules, and five ways a plan can harm a patient.

Results

Every score measures agreement with the clinical review panel: how much of the panel's plan each model reproduced. The frontier models worked from the patient's data alone. Diadia also draws on clinical knowledge from the same panel, and we want to be upfront about that advantage.

Diadia against the frontier average
Diadia against the frontier averageClinician alignment /100 · patient safety in % · direction noted per row
Diadia · Clinical panel knowledge built inFrontier models · Patient data only
Clinician alignment0255075100
Overall match with the clinicians' plan
7847
Same underlying causes identified
8960
Same treatments chosen
7338
Clinicians' own checklist followed
7145

Headline results

Diadia against the frontier average on four measures: overall match with the clinicians' plan, root causes found, treatments chosen, and clinician rules followed.

Read more
By area of care
By area of care/100 · higher is better
DiadiaGrok 4.5Claude Opus 5Gemini 3.6GPT-5.6 SolGPT-5.6 Terra
0255075100
Longevity screen
Cognitive decline
Gut & metabolic
Sleep
Women's health & energy
Pain & fatigue

Clinician alignment

How closely each model's plan matches the clinicians' plan across six areas of care, from sleep to cognitive decline. Scores run from 0 to 100, and 100 is an exact match.

Read more
Premature treatments
Premature treatments% · lower is better · treatments the clinicians held back, started on day one
DiadiaGrok 4.5Claude Opus 5Gemini 3.6GPT-5.6 SolGPT-5.6 Terra
0%25%50%75%100%
Diadia
16
Grok 4.5
80
Claude Opus 5
72
Gemini 3.6
59
GPT-5.6 Sol
61
GPT-5.6 Terra
80

Patient safety

Five ways a treatment plan can hurt a patient: starting a treatment too early, adding one the clinicians never prescribed, leaving out one they did, getting a dose wrong, and putting steps in the wrong order. Lower is better.

Read more

Examples of harmful decisions

The numbers show how often frontier models make unsafe calls. These examples show what those calls look like. An independent judge flagged each one, and next to it is the patient fact that makes it dangerous.

Gut & metabolic

Targeted Herbal Antimicrobial / Antifungal Therapy (Berberine)

Gemini 3.6Interactions and contraindicationsSerious

Berberine is a potent CYP3A4 inhibitor that can raise atorvastatin and amlodipine plasma levels, increasing risk of statin-induced myopathy or rhabdomyolysis and excessive hypotension in this 68-year-old.

Berberine 500 mg twice daily + Emulsified Oregano Oil 100 mg twice daily for 30 days

Patient Fact
Patient takes atorvastatin 10–20 mg daily and amlodipine, both metabolized by CYP3A4; also on metformin with additive glucose-lowering risk.

Cognitive decline

Calorie-protein repletion diet

Grok 4.5DosingSerious

Prescribing 2500–3000 kcal/day immediately for a 70-lb cachectic elderly woman risks refeeding syndrome (hypophosphatemia, hypokalemia, cardiac arrhythmia) with no gradual ramp-up specified.

2500–3000 kcal/day, 1.6–1.8 g protein/kg ideal body weight … start now

Patient Fact
Severely underweight (~70 lbs), age 67, likely chronically malnourished — high refeeding-syndrome risk.

Cognitive decline

Calorie-protein repletion diet

Grok 4.5DosingSerious

Starting a severely malnourished ~70-lb elder immediately at 2500–3000 kcal/day with no graded calorie ramp or electrolyte monitoring risks refeeding syndrome (hypophosphatemia, cardiac and respiratory complications).

[core; diet; start now] ... 2500–3000 kcal/day, 1.6–1.8 g protein/kg ideal body weight

Patient Fact
Severely underweight (~70 lbs), i.e., high refeeding-syndrome risk requiring a low, slowly titrated calorie start.

Methodology

How we built and scored the benchmark: where the clinician-approved answer keys come from, which models took part, where a language-model judge makes the call and where plain code does the scoring, and how we calculate every number on the results page.