This website uses cookies

Read our Privacy policy and Terms of use for more information.

Ten frontier models, one incomplete chart, 140 treatment plans. Forty-one percent invented a fact that wasn't there. One added sentence dropped confident errors from 24 to 6.

Here is the sentence. Paste it into any prompt where AI touches a patient's record:

"If your plan depends on information that is not in the chart, say what is missing and ask for it instead of assuming it."

That line comes from Michael Hobbs, MD, a practicing pediatrician who has become one of the clearest voices translating frontier AI for clinicians in the field. In his write-up, "The Anatomy of an AI Clinical Error", he gave ten frontier models the same toddler ear-infection chart, complete enough to diagnose, but missing the two facts that decide treatment, and asked each for a plan.

Of the 140 responses, 41 percent asserted at least one fact that never existed in the chart: an antibiotic course nobody took, a follow-up nobody arranged. Then he added the sentence above and reran the experiment. Confident inaccuracies fell from 24 of 56 responses to 6. Two models stopped inventing entirely.

You are not treating ear infections. But you are running AI against incomplete charts every day, and the failure he found lives in the gap between what your record says and what the model needs to decide. More on that below.

The errors have shapes

The most common class was confident gap-filling. One model wrote: "High-dose amoxicillin is appropriate - no antibiotics in the past 30 days, no concurrent purulent conjunctivitis, no penicillin allergy." All three claims carried equal confidence. Two were in the chart. The middle one was invented, and it was the premise the drug choice depended on.

A second class is the model equivalent of a colleague who trained in 2003 and never updated: stale knowledge, and no amount of better charting fixes it. One model kept dosing amoxicillin at 45 mg/kg per day, the standard until roughly 2004. Another misread a documented 24-month-old as "under 2," citing an academy rule that does not exist.

Models have personalities

Three findings worth holding onto:

  • Treatment choice was "primarily a property of the model rather than the patient." One model prescribed antibiotics on every run; three chose watchful waiting every time.

  • The longest, most complete-looking plans fabricated the most. The three tersest models invented facts once each. The long outputs, he writes, "prioritized an aesthetic of completeness, and completeness requires data points that the record did not provide."

  • Demographics shifted the reasoning. A mother charted as a pediatric nurse was credited as a reliable observer in six of six runs. Charted as unemployed: zero of six. Same child, same plan, different justification.

Hallucination behaves like memory error

Hobbs's framing is that AI fabrication looks a lot like human memory error: predictable, clustered in specific situations, and something you can design around the way you design around your own fallibility. He points to a 2019 JAMA Network Open study that audio-recorded resident encounters and found that only 38.5 percent of documented review-of-systems findings could be confirmed on the recording. Models trained on millions of our charts, he writes, "inherited our documentation habits alongside the medical knowledge."

You would not stop consulting a trusted colleague because their memory is fallible. You would ask a follow-up question. That is what the sentence does.

He is plain about the limits. The sentence stops invented history; it does nothing for what a model learned wrong, like the pre-2004 dosing. It is one synthetic case, and the broader literature shows real but modest gains from prompting, with a December preprint finding the benefit uneven across tasks. What his data supports is narrower and more useful: the failures cluster in ambiguity, and a prompt that names the ambiguity closes most of that gap.

Where this shows up in your practice

Our charts are more incomplete than a pediatric ear-infection note, not less. Outside labs arrive as PDFs. Intake forms skip the question that matters. The med list is the patient's memory of it. Every one of these is a gap a model will fill for you if you let it:

  • The HRT plan that depends on whether she still has a uterus and the intake never asked. A model drafting the plan will assume one answer and write the progesterone decision around it.

  • The peptide protocol drafted against an incomplete med list. Missing anticoagulant, missing GLP-1, missing thyroid dose. The model writes a clean schedule because nothing in the chart told it not to.

  • The supplement plan built from labs the AI never saw. The ferritin and the B12 are in the outside PDF, not in the structured record. The plan says "no evidence of deficiency."

  • "No history of clots" is documented because nobody said otherwise, not because anyone asked.

The sentence catches every one of these. Instead of a confident plan, you get: "This plan depends on hysterectomy status, current anticoagulant use, and recent ferritin, none of which are in the chart. Please provide before I proceed." That is a better colleague.

A prompt for each class of error

Drawn from the five practice principles in his write-up:

  • Confident gap-filling → Add the sentence above to every prompt. In his rerun it cut confident inaccuracies by three quarters.

  • Silent assumptions → Follow every AI-drafted plan with: "What missing information, if any, would have changed this plan?" When asked, 97 percent of responses named their assumptions.

  • Stale knowledge → Prompting will not fix this class. Spot-check doses and thresholds against the current guideline, especially after a model update.

  • Shifted reasoning → Audit the justification and the defaults, not just the prescription line.

  • All of the above → Error patterns were stable per model, and models keep changing. Rerun your hardest cases whenever the model behind a tool changes.

Run it against your own stack

Hobbs published his full test packet on GitHub: the chart, the prompts, the scoring. It takes about ten minutes to run the case through whatever you are using today, and it will tell you more about your tool than any vendor demo. He has also built a parent-facing tool, Wren, that shows how far a practicing clinician can take this.

Hold the specifics loosely: models change fast enough that some of this will be out of date within weeks, which is why the habits matter more than the model names. Keep it all inside BAA-covered tools, with a human between AI output and the chart.

Start with the one sentence on Monday's first note.

Be a modern clinician with the help of Ultralight, the AI-native EHR built specifically for functional, integrative, and longevity medicine. We've recently launched wearables integrations and improved AI-native clinical workflows. Get in touch to see the latest updates.

One more introduction: Dr. Michael Hobbs, whose testing anchors this issue, recently joined Ultralight as our clinical advisor. His error-first way of evaluating AI is already shaping how we build.

Join us in Coronado: MVMNT Longevity Medicine Summit

On September 22 and 23, Sunita Mohanty, co-founder and CEO of Ultralight, joins the faculty at the MVMNT Longevity Medicine Summit at Loews Coronado Bay Resort. Two days of evidence-graded longevity science, hands-on labs, and physician-built protocols, capped at 300 clinicians. If you are going, come find us.

In the news

A GLP-1 extended lifespan in aging mice. In Nature this week, semaglutide started at 20 months of age, late life for a mouse, extended median lifespan by nearly 100 days, improved muscle and cognitive function, and outperformed calorie restriction on spatial memory and blood-sugar regulation in the NIH-funded Berkeley study. The mice were female and the authors say plainly that none of this transfers to humans yet, but your longevity patients will bring it up this month, and "promising in animals, unproven in us" is the conversation to be ready for.

Hormone therapy timed to the menopause transition tracked with 22% fewer cardiovascular events. In JAMA Internal Medicine this week, twenty years of SWAN data on more than 2,700 women with vasomotor symptoms showed 22 percent lower cardiovascular event risk when hormone therapy started in peri- or early postmenopause, and 27 percent when started within ten years of onset. It is observational and the authors say so plainly, so read it as support for the timing hypothesis rather than proof, and as one more data point for the counseling conversation with symptomatic patients.

Probiotic use nearly quadrupled while the diet that feeds the microbiome sat flat. A University of Michigan analysis of NHANES data from 2009 to 2023, published this week in Clinical Gastroenterology and Hepatology, found adult probiotic supplement use rose 287 percent to about 7 percent of adults while fiber intake stayed essentially flat and high-microbial food consumption barely moved. Patients are buying the capsule and skipping the food, which is a counseling opening every functional practice can use this week.

Upcoming Conferences & Events

Sept 18-19, Dr. Kara Fitzgerald Masterclass “Functional Medicine Is Longevity Medicine”· Online · Led by Kara Fitzgerald, this year’s Masterclass will continue under the theme that functional medicine is longevity medicine. Not as a trend, but as the most science-backed, clinically effective path to extending healthspan. Ultralight will be there!

Sept 2223, MVMNT Longevity Medicine Summit · Coronado, CA · Evidence-graded longevity science, hands-on labs, and clinical frameworks you can implement the week after. Capped at 300 clinicians. Ultralight will be there!

Oct 8–10, A4M Women's Health Summit · San Antonio, TX ·  The best clinical education on hormone, metabolic, and midlife women's health you will see this year. The room to be in if you are growing the perimenopause and menopause side of your practice.

Oct 17-18, Roundtable of Longevity Clinics · Buck Institute, Novato, CA · Some 250 longevity clinic leaders, physicians, and researchers working toward shared standards for longevity testing and interventions. In-person and virtual tickets are open. Ultralight will be there!

Oct 21-23, DOC (Living Room Lab) · Sonoma, CA · Salon-style sessions on longevity science and medical AI with faculty from UCSF, Stanford, and the Buck, plus validated diagnostics in the Living Room Lab. Ultralight will be there!

Oct 21–24, NAMS Annual Meeting · San Diego, CA  · The single most practice-changing meeting of the year for midlife women's health. Your protocols will look different after this one.

Nov 5–8, Eudēmonia Summit · West Palm Beach, FL ·  One of the most talked-about longevity gatherings in the U.S. Experientials, hands-on demos, and the best place to try the emerging frameworks your patients will ask you about next year. OvationLab and Ultralight will be there!

Nov 5-7, Private Physicians Alliance Annual Meeting · St. Petersburg, FL · The gathering for independent, cash-pay, and concierge physicians navigating practice independence. Practical and peer-driven. Ultralight will be there!

Nov 8-11, American College of Lifestyle Medicine Conference · Orlando, FL · Lifestyle medicine's main annual event — evidence-based approaches to behavior change, chronic disease, and healthspan. Growing overlap with the longevity medicine community.

Nov 10-13, Valley Forum 2026 · Napa Valley, CA · Invitation-only forum where healthcare leaders work through the industry's hard problems, AI's role in care among them. Ultralight will be there!

Dec 11–13, A4M Longevity Fest · Las Vegas, NV  · The biggest longevity event in the U.S. The room spans clinicians, industry, founders, and the people building next year's platforms, and the connections from this one tend to compound through the rest of your year. OvationLab and Ultralight will be there!

Know of an event we should add? Reply and tell us.

Until next week

Hobbs proved the biggest AI failure in the chart is cheap to fix, and the fix is a sentence you can paste before Monday's first visit. The models will keep changing. The habit of making them name what's missing is yours to keep.

Reply and tell us the sentence you use. We are compiling the prompts our readers actually run and the best ones go in a future issue, credited.

Know a clinician who runs AI against their charts? Forward this to them. They can subscribe here.

Until next week, keep building the practice you imagined when you started.

— Dr. G and Sunita