
On August 17, JAMA published a Viewpoint arguing that AI alone, with no physician in the loop, will deliver better medical care than physicians, and better than physicians working with AI.
The authors are not fringe: bioethicist Zeke Emanuel, Abe Baker-Butler, and Vinod and Neal Khosla. Their claim covers five cognitive tasks at the center of what you do all day: eliciting patient information, building a differential, ordering tests, prescribing, and managing chronic disease. Their timeline: real-world deployment in some workflows by 2030.
Days later, the disagreement happened out loud. On Chrissy Farr's Lifers podcast, Emanuel debated AMA CEO Dr. John Whyte, and it got heated, down to whether an AI should ever be the one to tell you that you have cancer.
This week: both sides' best arguments, the variable they are arguing past (who holds the assembled patient), and why this debate lands differently in a functional and longevity practices.
The debate
Emanuel's case: the evidence, not the vibes. Emanuel's position is blunt: "AI is already equivalent to humans in many areas and superior. It's just going to get better and better." The JAMA piece backs it with numbers that deserve a straight look. In one cited study, ChatGPT o3 identified the correct final diagnosis first in 60% of complex real-world cases; internal medicine physicians managed roughly 16% on a comparable subset. Google's AMIE system was rated better than physicians at eliciting patient complaints, 97% versus 50% favorable, in text-based encounters with simulated patient actors.
His strongest ground is chronic disease: "In clinical practice, both for diabetes and for hypertension, less than a quarter of patients are managed to guidelines today." That number should sting, because it's true. His vision follows from it: "We can decant a lot of those interactions to autonomous AI and the physician can focus on patients that really need the time," starting with patients who are uninsured, poor, and rural, and currently get no care at all. He also claims patients are more honest with the AI, which will sound plausible to anyone who has watched a patient round down their drinks per week.
Disclosure belongs in the picture: Neal Khosla runs Curai Health, an AI virtual care company, and Khosla Ventures backs AI health startups including Curai. That doesn't make the evidence wrong. It does mean the 2030 timeline comes from people with a position on it.
Whyte's case: the job is bigger than the benchmark. Whyte's counter is that the whole framing is "almost a reductionist view of what it means to be a doctor and practice medicine." His line that stuck: "There's much more to medicine than reading an X-ray. It's that relationship with the patient, on trust." His model is unambiguous, "AI should be part of the care team that's managed by the physician," and he draws the line at autonomy even for mundane calls like medication refills. He also raises the machinery nobody has built: liability frameworks, patients who can't parse predictive values, and the delivery question that made the debate flare: "Do you want to be told by an LLM that you have cancer through a text message? Absolutely not." His aviation analogy does a lot of work: people still won't board a plane without a pilot, even though the plane mostly flies itself.
What both sides take for granted. The sharpest response to the paper came from an ER physician, not a policy figure. Commenting on Emanuel's post, Graham Walker, MD (co-director of advanced development at TPMG, co-founder of Offcall, founder of MDCalc) pointed at the conditions everyone else argued past: we ask physicians to do this job while juggling ten other problems, insurance, documentation, and interruptions, "backwards, in heels," and then we hand AI "one patient, all the relevant information with a bow on top," one defined task, and unlimited time.
Be clear about what that argument does and doesn't do. It does not explain away the results; in the o3 study, the physicians worked from the same assembled cases and still lost decisively. On pure reasoning over complete information, concede the point. What the bow-on-top condition explains is why benchmark performance hasn't shown up in clinics: complete information is the one input real medicine almost never has. A real patient's picture is scattered across an EHR, lab portals, a wearables app, a supplement platform, and a message thread, and someone assembles it by hand before any reasoning starts. Every study in the JAMA piece begins after that work is done.
Even the paper's most unsettling finding, that hybrids underperform AI alone, reads differently in this light. The mechanism in the research is trust calibration: humans are poor at deciding when to trust AI, especially once it exceeds them. And in a working clinic, the human in the loop is usually second-guessing outputs against a picture they don't fully have, with attention mostly spent moving information between systems. Nobody designed that loop for the human to add anything.
Our take: both debaters argue over who should hold the reasoning while treating the assembled patient as a given. It isn't a given. Assembly is the unsolved, unbenchmarked, expensive half of the job, and it's the half a practice actually controls. Take the evidence seriously and the timeline with salt. Then ask the question that matters this week: who in your practice holds the assembled patient?
Why this lands differently in your practice
You practice the most context-heavy medicine there is. A functional or longevity panel carries more specialty labs, more longitudinal biomarkers, more wearable feeds, more supplements, and more protocols per patient than any conventional clinic. That puts you on both extremes of this debate at once: no benchmark looks less like your workday, and no one pays a higher assembly tax while waiting for someone to solve it.
What gets automated first is not what you sell. Emanuel names the deployment order himself: transactional encounters, and patients the system currently ignores. This isn't hypothetical. Utah has been running autonomous prescription renewals with Doctronic since January, with the AI recommending renewal in 72% of cases under physician review. The refill visit and the routine follow-up are being priced toward zero right now. What your patients pay cash for is the other half: synthesis across a year of data, and a relationship strong enough to change behavior. The JAMA authors exempt physical procedures from their claim. The deeper exemption is the one no benchmark has measured: whether patients actually get better over years.
You already agree with Emanuel about the disease. His under-a-quarter-managed-to-guidelines number describes the system your patients left. Your practice model exists because seven-minute medicine cannot manage chronic disease. He proposes fixing the seven minutes with autonomy; you fixed them with time. That's a real disagreement about the cure, and a far stronger position than pretending his evidence is wrong.
Two warnings, though. First, the protection only holds if your context is actually assembled. A practice running six disconnected systems is delivering fragmented care with a longer visit, closer to the thing being automated than to the thing that can't be. Second, your patients are the most AI-forward in medicine. They asked ChatGPT before they booked with you, they will quote the 60% number at you, and as of last week their bloodwork is being assembled without you: Whoop just opened Advanced Labs to non-members and added Grail's Galleri cancer screen, putting clinician-reviewed labs and biometrics in one consumer app. If their app's synthesis is faster and clearer than yours, your differentiation erodes from below. And if Emanuel is right that patients are more honest with the AI, the relationship premium is something you earn visit by visit.
The Playbook: Run a Context Integrity Audit
Run this on one stable, non-urgent follow-up. It audits your workflow, not the accuracy of the AI.
1. Define one decision. Write the question first: What information could change whether I continue, modify, or stop X? Star the three to five facts you expect to be decision-critical.
2. Build a one-page Context Brief. Include only what can change the decision:
Patient goal and relevant timeline
Medications, supplements, allergies, and recent changes
Relevant labs, specialty tests, and wearable trends
Prior response or intolerance
Contraindications, safety variables, and red flags
Missing, contradictory, stale, or unverified information
Date each item, label its source, and mark unknowns as unknown.
3. Measure the assembly burden. Record:
Assembly minutes
Distinct apps or portals opened
Manual copy-pastes or re-entry
Decision-critical items that were missing or conflicting
4. Compare safely.
Use only an AI environment approved by your practice for PHI and covered by a signed BAA when required. Otherwise, use a synthetic case. Removing the patient’s name alone is not de-identification.
Use the same model, a fresh session, the same clinical question, and the same requested output. Run it once with the fragment you usually provide and once with the Context Brief.
Using only the supplied information, list established facts with source and date; flag missing or contradictory information; and identify decision-relevant considerations and clinician questions. Separate fact from inference. Do not diagnose, prescribe, order, chart, or message the patient.
Compare starred facts missed, unsupported claims, contradictions missed, and clinician verification time.
Differences are signals to inspect, not proof of superior clinical performance.
5. Repair one seam.
Choose one change, one owner, and one retest date. Use the repaired workflow for five complex visits, then remeasure assembly time and missing or conflicting information.
Clinical guardrail: This is a workflow audit, not patient-care automation. Verify every statement that could change care, and do not let the experimental output independently generate orders, prescriptions, chart entries, or patient communication.
Run it before you trust it
Run the audit, then reply with two aggregate numbers only: assembly minutes and distinct systems opened.
Please do not include patient details or PHI.
Those numbers reveal one often-missed component of AI readiness: the burden of assembling the patient before reasoning begins.
And if Step 1 gives you a number you don't like, that is the job Ultralight was built for: labs, wearables, supplement lists, protocols, notes, and patient messages living in one longitudinal record, so the picture is assembled before any reasoning starts, yours or the AI's. Discount us accordingly and run the test anyway.
If the audit shows that your clinicians are rebuilding the patient before every complex visit, that is the problem Ultralight is designed to reduce: helping bring labs, wearables, supplements, protocols, notes, and messages into a more coherent longitudinal record before the visit begins.
Join us: Functional Medicine Is Longevity Medicine
Three weeks out: on September 18 and 19, Dr. Kara Fitzgerald, functional medicine clinician, researcher, and host of New Frontiers in Functional Medicine, runs her annual virtual masterclass under this year's theme: functional medicine is longevity medicine. Two days, fully online, free for practicing clinicians. Ultralight will be there.

In the news
The FDA is taking up testosterone for menopausal women. An Aug 18 Federal Register notice set a September 17 hybrid public workshop on a potential approved female indication: physiology across the lifespan, assay and measurement problems, candidate indications, and long-term cardiovascular and breast safety, with a comment docket open through October 19. Off-label female testosterone is daily work in this readership; this is the first formal step toward prescribing it on-label, and the docket is open to your comments.
AI can now estimate the biological age of 40 tissue types, and read much of it from a blood draw. A Nature Medicine study published August 14 trained deep-learning models on 25,712 histology images across 29 organs, hit a mean error under five years, then built blood-based predictors of tissue-specific aging validated across nine independent cohorts. Organ-age panels are already landing in your patients' inboxes; this is the peer-reviewed version of the concept, and a preview of what the commercial clocks will claim next.
Coffee tracked with lower visceral fat, more muscle, and higher total testosterone in men. In 2,264 46-year-olds from the Northern Finland Birth Cohort, higher intake ran dose-dependently with less visceral fat and higher skeletal muscle mass, plus higher total testosterone and SHBG in men and lower free androgens in women. Cross-sectional, so no causality, but patients will ask, and the SHBG detail is the teachable part: total hormone numbers can rise while the free fraction falls.
Upcoming Conferences & Events
Sept 18-19, Dr. Kara Fitzgerald Masterclass “Functional Medicine Is Longevity Medicine”· Online · Led by Kara Fitzgerald, this year’s Masterclass will continue under the theme that functional medicine is longevity medicine. Not as a trend, but as the most science-backed, clinically effective path to extending healthspan. Ultralight will be there!
Sept 22–23, MVMNT Longevity Medicine Summit · Coronado, CA · Evidence-graded longevity science, hands-on labs, and clinical frameworks you can implement the week after. Capped at 300 clinicians. Ultralight will be there!
Oct 8–10, A4M Women's Health Summit · San Antonio, TX · The best clinical education on hormone, metabolic, and midlife women's health you will see this year. The room to be in if you are growing the perimenopause and menopause side of your practice.
Oct 17-18, Roundtable of Longevity Clinics · Buck Institute, Novato, CA · Some 250 longevity clinic leaders, physicians, and researchers working toward shared standards for longevity testing and interventions. In-person and virtual tickets are open. Ultralight will be there!
Oct 21-23, DOC (Living Room Lab) · Sonoma, CA · Salon-style sessions on longevity science and medical AI with faculty from UCSF, Stanford, and the Buck, plus validated diagnostics in the Living Room Lab. Ultralight will be there!
Oct 21–24, NAMS Annual Meeting · San Diego, CA · The single most practice-changing meeting of the year for midlife women's health. Your protocols will look different after this one.
Nov 5–8, Eudēmonia Summit · West Palm Beach, FL · One of the most talked-about longevity gatherings in the U.S. Experientials, hands-on demos, and the best place to try the emerging frameworks your patients will ask you about next year. OvationLab and Ultralight will be there!
Nov 5-7, Private Physicians Alliance Annual Meeting · St. Petersburg, FL · The gathering for independent, cash-pay, and concierge physicians navigating practice independence. Practical and peer-driven. Ultralight will be there!
Nov 8-11, American College of Lifestyle Medicine Conference · Orlando, FL · Lifestyle medicine's main annual event — evidence-based approaches to behavior change, chronic disease, and healthspan. Growing overlap with the longevity medicine community.
Dec 11–13, A4M Longevity Fest · Las Vegas, NV · The biggest longevity event in the U.S. The room spans clinicians, industry, founders, and the people building next year's platforms, and the connections from this one tend to compound through the rest of your year. OvationLab and Ultralight will be there!
Know of an event we should add? Reply and tell us.
Until next week
Emanuel and Whyte are arguing about who should hold the reasoning. The assembled patient is the variable your practice controls, and the context test tells you where you stand. Forward this issue to the colleague who quoted the 60% number at you.
Reply and tell us your assembly time from Step 1. The best ideas in this newsletter come from clinicians doing the work.
Until next week, keep building the practice you imagined when you started.
— Dr. G and Sunita