Insight

Diadia Puts Transparency to Work on Clinical AI’s Confidence Problem.

As AI gains momentum in the clinical arena, it’s crucial that we understand the logic behind its conclusions. Diadia’s Transparency Lab helps physicians bring a more thoughtful approach to evaluating results.

a wall shelf with neatly organized books

By now, most of us are using AI in some form. From personal organization and everyday problem solving to more substantial applications in our professional lives, we’re discovering where it helps us and where it doesn’t. The results are often surprising. We sometimes see AI perform very well in situations where we might not expect it, yet make strange mistakes on what seem like simpler requests.

The problem isn’t just that AI gets things wrong. It’s that AI tends to serve up results with a similar degree of confidence, whether it’s pulling from substantial evidence or inventing a point to fit a particular pattern it recognizes. On day-to-day tasks, those errors can be annoying. In complex clinical cases, they can have serious consequences.

Research we performed at Diadia turned up exactly this example [3]. When asked to explain a patient’s low testosterone level, a frontier AI model offered what appeared to be a logical analysis. It reasoned that DHEA is a precursor to both cortisol and testosterone, so that when cortisol demand increases, the body diverts DHEA to make cortisol, leaving less available for testosterone.

While the AI’s rationale might sound convincing at first, it contains an erroneous leap in logic. DHEA is a precursor to testosterone, through androstenedione, but it is not a precursor to cortisol, which is produced through a separate pathway. The model invented a false connection between two legitimate points of biology to complete its story and provide its conclusion. AI analysts call these instances faithful hallucinations: answers that may be partially accurate and sound plausible, but contain unsupported information that may skew or invalidate conclusions.

Faithful hallucinations - partially true statements

We aren’t the only ones tracking these errors. A 2025 analysis published in Nature Communications by Wu and colleagues looked at 800 medical questions across seven leading AI models and found that 50% to 90% of responses were not fully supported by the sources they cited [1]. Even GPT-4o with web search left roughly 30% of individual statements unsupported. Physicians are experiencing the issue, too. In a 2025 global survey available on arXiv, 91.8% of respondents reported encountering medical hallucinations in clinical practice [2].

So, how can we harness the undeniable access to information and processing power of AI to inform clinical insights without falling victim to errors in its reasoning? This was the challenge we took on at Diadia, and it led to the development of our new, publicly available resource, the Transparency Lab.

What is the Transparency Lab?

The Transparency Lab doesn’t approach clinical conclusions as singular or standalone answers. Instead, it breaks down results into a series of individual elements, checking each step against the published evidence. Every data point is graded for quality, and every link within the logic is tested along the way. It’s a way of evaluating the process as well as the product.

What is a structured reasoning graph?

The Transparency Lab produces what we call a structured reasoning graph, a helpful visualization of a clinical conclusion. Points on the graph represent data for biomarkers, biological processes, conditions, and interventions, while the connections between points show relationships and logic. By examining each connection individually, clinicians can identify instances that a simple right-or-wrong assessment can miss. They can see when a conclusion is well reasoned, even if it has yet to be fully proven.

The distinction is especially important in clinical care, where cases can be complex and available evidence doesn’t always lend itself to a simple yes or no. To provide more thorough and nuanced insights, Diadia labels conclusions in one of three ways: Supported, Unsupported, or Plausible.

What do the designations Supported, Unsupported and Plausible mean?

Supported means that each step of the reasoning process is backed by direct evidence. Unsupported means a critical step is contradicted or cannot be verified. Plausible is the new category we created for situations where the biological reasoning is sound and the evidence points in a direction, but the full chain has yet to be established in a particular clinical context. The Transparency Lab allows users to surface potentially valuable clinical insights without presenting them as more certain than the evidence warrants.

Consider a practical use-case. One claim submitted to the Transparency Lab asks whether vitamin D can lower thyroid antibodies in Hashimoto’s disease. The conclusion is labeled Plausible. There is established biological reasoning behind it, along with clinical evidence linking low vitamin D levels to elevated thyroid antibodies and showing reductions following supplementation. But the available research has not yet established the full chain with an RCT trial or any other study that proves the end to end conclusion to consider the conclusion Supported.

How is the Transparency Lab different from general AI?

That is precisely the distinction a confident AI answer can obscure. Rather than oversimplifying the evidence and presenting a binary yes or no conclusion, Diadia shows clinicians where established evidence ends and plausible reasoning begins.

Our goal is not to make AI sound more certain, but to make its reasoning more transparent to the clinical professionals who use it. Every report in the Transparency Lab reflects the same reasoning framework applied to real patient data, including the evidence behind each step, the quality of that evidence, and places where analysis comes up short.

We don’t believe clinicians should take AI at its word. They should be able to see its reasoning, examine all evidence presented, and decide for themselves how much confidence it deserves. That’s why we’ve given the public access to Diadia’s Transparency Lab. Today, it includes more than 3,400 clinical claims backed by over 52,000 citations, with more being added as the evidence evolves. Explore the Lab, pressure-test the claims you know best, and see what added context reveals.

Frequently asked questions

What is the Transparency Lab?

The Transparency Lab is a free, public library from Diadia Health that shows the evidence and reasoning behind AI-generated medical claims from real patient cases. Each entry decomposes a claim into its mechanisms, checks each against the peer-reviewed literature, and grades it Supported, Plausible, or Unsupported.

What is a structured reasoning graph?

It is the verification method behind every report. Diadia maps a claim as a graph of mechanistic steps, checks each step independently against the evidence, and assembles the verdict with fixed rules rather than another language model, so the same inputs always produce the same result.

What does a plausible verdict mean?

Plausible means the biological reasoning is sound and the evidence points the right way, but no trial has confirmed the full chain for that specific context. It is a precise middle ground between Supported and Unsupported, not a soft maybe, and it is where much of precision and functional medicine reasoning operates.

How is this different from asking ChatGPT?

A general model predicts the most likely-sounding answer and presents every claim with the same confidence. Diadia traces each claim to its evidence and shows you exactly where that evidence stops, so you can verify the reasoning instead of trusting it.

Citations

  1. Wu et al. (2025). An automated framework for assessing how well LLMs cite relevant medical references. Nature Communications. nature.com/articles/s41467-025-58551-6

  2. Kim et al. (2025). Medical Hallucinations in Foundation Models and Their Impact on Healthcare. arxiv.org/abs/2503.05777

  3. Diadia Health. Claim-Level Transparency Analysis of LLM-Generated Diagnostic Reports. biorxiv.org/content/10.64898/2026.05.03.721751v1