Off the clock, off the P&L: Your best clinician works 14 hours a week for free
The hidden economics of complex cases in a cash-pay precision practice, and how AI can change it.
AI cannot reason through a complex case. It can hand you a confident answer in seconds, but underneath the fluency it is pattern-matching, predicting the most likely-sounding response instead of working up from the evidence.

Closing that gap is what Diadia was built to do. More than tracing a claim back to a source, it performs a logical linking of the evidence behind each step, grades every piece for quality, and labels where the reasoning holds and where it breaks, all in the open, one claim at a time. We call it the Transparency Lab.
Here is one from our own study. Asked to explain a patient's low testosterone, a frontier model produced a tidy story: DHEA is a precursor to both cortisol and testosterone, so when cortisol demand runs high the body diverts DHEA to make cortisol and leaves less for testosterone. Half of it is true. DHEA is a direct precursor to testosterone, by way of androstenedione. But DHEA is not a precursor to cortisol at all, and there is no pathway between the two; cortisol is built on a separate branch of steroidogenesis.
The model could not see that, because it pattern-matches language rather than mechanism, so it invented the missing link to make the story hold together. This is a faithful hallucination: fluent, partly correct, and wrong in the one place that changes the plan. Run the same claim through Diadia and it is split into its steps, each checked against the evidence. The DHEA-to-testosterone link holds. The DHEA-to-cortisol link has nothing behind it and breaks, so the fabricated step is flagged rather than passed along as fact.

This is not a quirk of one claim. A general model delivers its best-supported findings and its shakiest guesses in the same confident voice, with nothing to tell them apart, and the research bears that out. The research bears this out. A 2025 analysis in Nature Communications by Wu and colleagues found that between 50 and 90 percent of large language model medical responses are not fully supported by the sources the model itself cites, and that even GPT-4o with web search leaves roughly 30 percent of individual statements unsupported. A separate global survey of clinicians (Kim et al., 2025) found that 91.8 percent had already encountered a medical hallucination in their own practice. A slicker answer does not fix that. What helps is an answer that shows its reasoning, so you can see for yourself which step holds and which one breaks.
The Transparency Lab is where Diadia puts that reasoning in the open. It is a free, public library where each entry takes a single clinical claim, decomposes it into the mechanisms it depends on, checks each one against the published literature, and shows you where the evidence holds and where it runs out. Instead of asking you to take our word for it, we ran it across thousands of health claims and published the results, citations attached, including the places our own analysis stops short. It holds over 3,400 published claims backed by more than 52,000 citations, and it states plainly that more than half are not yet fully established. That honesty is the whole point. We mark exactly where the evidence gives out on every claim, because that is what a clinician needs to see before trusting any of it on a patient.
A structured reasoning graph is the method behind every report. It is a directed graph in which the nodes are biological entities, the biomarkers, processes, conditions, and interventions in a claim, and the edges are the specific causal or logical steps between them. Diadia builds one for each claim and verifies every edge on its own against the peer-reviewed literature, rather than judging the statement as a single block.
That is exactly what happens with the DHEA claim above. Because each link is checked on its own, the supported one, DHEA to testosterone, stands while the fabricated one, DHEA to cortisol, breaks and gets flagged instead of riding along inside a confident answer. Judge the claim as one block and you lose that resolution entirely.
Three design choices make that verdict defensible. Evidence is graded by strength, so a recent systematic review or a well-powered trial outweighs a small cross-sectional study or a preprint, and a single supportive paper is not enough to mark a step Supported without independent replication by another high quality paper. The reasoning follows one of three patterns the graph names explicitly: direct mechanism, chain inference, or pathway convergence, where several independent lines meet on the same conclusion and logical linking is established. The final verdict is computed by fixed rules, so the same inputs resolve the same way every time, and every conclusion carries a citation trail you can follow back to the source.
We force a logical link between each piece of evidence, then trace that link to make sure the conclusion actually comes from the evidence. Most AI doesn't do that. It pattern-matches, and guesses what's most likely to be true.
The Transparency Lab reveals the mechanism behind the product so you can have confidence in the insights generated and the decisions you make from the results. Every report you read is the identical reasoning Diadia runs on a real patient's data, not a demo staged for the website.
Every claim resolves to one of three verdicts. Supported means each step in the chain has direct evidence. Unsupported means a critical step is contradicted or cannot be verified. Plausible, the verdict that carries the most weight in complex care, means the biology is sound and the evidence points the right way, but no trial has closed the full chain for that specific context. We go a step further than a simple label: every piece of evidence behind a claim is graded for quality, and the verdict itself is set by deterministic rules rather than a model's opinion, so the same evidence always lands on the same result.
Take a claim live in the Lab now: does vitamin D lower thyroid antibodies in Hashimoto's? It sits at Plausible, and the report walks the mechanism step by step. Active vitamin D binding the vitamin D receptor suppresses Th1 and Th17 activity and the inflammatory cytokines they drive, and it induces FoxP3 regulatory T cells and tolerogenic dendritic cells, both well supported in the literature. The clinical link, an inverse relationship between low 25(OH)D and elevated TPOAb and TgAb titers, with autoantibody reductions reported over three to six months of supplementation, is real but rests on evidence the trials have not fully closed. Read that report and you can see the exact step where established biology gives way to reasoning, which is where a great deal of your own work on a train wreck already lives.
A general chatbot gives you an answer and asks you to trust it. Diadia gives you the reasoning and invites you to check it. AI is already in your exam room, arriving in the differentials your patients read online and in the tools moving into how complex cases get built, so the question is not whether you will use it but whether you can trust the reasoning when you do. Pull up the claims you hold a strong opinion on and pressure-test them against your own read. If the reasoning holds in the Lab, that is the same reasoning you would get on your own patients, and the next step is to see it run on one of them.
What is the Transparency Lab?
The Transparency Lab is a free, public library from Diadia Health that shows the evidence and reasoning behind AI-generated medical claims from real patient cases. Each entry decomposes a claim into its mechanisms, checks each against the peer-reviewed literature, and grades it Supported, Plausible, or Unsupported.
What is a structured reasoning graph?
It is the verification method behind every report. Diadia maps a claim as a graph of mechanistic steps, checks each step independently against the evidence, and assembles the verdict with fixed rules rather than another language model, so the same inputs always produce the same result.
What does a plausible verdict mean?
Plausible means the biological reasoning is sound and the evidence points the right way, but no trial has confirmed the full chain for that specific context. It is a precise middle ground between Supported and Unsupported, not a soft maybe, and it is where much of precision and functional medicine reasoning operates.
How is this different from asking ChatGPT?
A general model predicts the most likely-sounding answer and presents every claim with the same confidence. Diadia traces each claim to its evidence and shows you exactly where that evidence stops, so you can verify the reasoning instead of trusting it.
Wu et al. (2025). An automated framework for assessing how well LLMs cite relevant medical references. Nature Communications. nature.com/articles/s41467-025-58551-6
Kim et al. (2025). Medical Hallucinations in Foundation Models and Their Impact on Healthcare. arxiv.org/abs/2503.05777
Diadia Health. Claim-Level Transparency Analysis of LLM-Generated Diagnostic Reports. docsend.com/view/yqbfbs8ur87ypdza
© 2026 Diadia. All rights reserved.
© 2026 Diadia. All rights reserved.