AI helped physicians perform better in a new randomized trial. That's worth paying attention to. But it shouldn't end the conversation. A study published today in npj Digital Medicine randomized 249 physicians in Kenya, Indonesia and the Netherlands to work either with or without access to GPT-4o. The differences were significant: → Kenya: +18 percentage points → Indonesia: +10.7 percentage points → Netherlands: +7.2 percentage points That is encouraging evidence for a future where AI augments physicians rather than attempts to replace them. But there are important qualifications. The control physicians were not allowed to use traditional resources such as clinical guidelines or the internet. The study did not assess harms. And the researchers caution that the findings came from controlled conditions—not routine clinical practice. That distinction matters. Real medicine rarely presents itself as a neatly bounded problem. A complex patient may arrive with years of symptoms, multiple medications, conflicting labs, previous diagnoses, incomplete records, failed interventions and changes that only become meaningful when viewed across time. There may not be one clean answer. There may not even be enough information yet to reach one. So the next question in clinical AI shouldn't simply be: Does AI improve physician performance? It should also be: What kind of AI-physician relationship produces better clinical judgment? Can the physician inspect the evidence behind a recommendation? See where uncertainty exists? Disagree with the system? Change the reasoning path? Identify what might be missing? And when physician and AI disagree, who retains clinical authority? At Physician Logic Squared, we believe those questions belong at the center of clinical AI design. The goal isn't an AI that thinks so the physician doesn't have to. It's a system that helps the physician synthesize more, question more intelligently and see connections that might otherwise remain buried in a patient's timeline. AI can strengthen clinical performance. Now we need to make sure it strengthens clinical judgment too. Because the best outcome isn't a physician who can work faster only while the AI is present. It's a physician whose capacity to think has been amplified. What should the next generation of physician-AI trials measure beyond accuracy?
Physician Logic Squared (PL2)
Technology, Information and Internet
san franciso, san franciso 183 followers
Physician-led AI clinical decision support for functional medicine.
About us
Physician Logic Squared (PL2) was created to address that gap. Built by physicians, PL2 is grounded in a simple idea: Clinical intelligence isn’t about having more information, it’s about seeing relationships clearly.
- Website
-
https://www.physicianlogic2.com
External link for Physician Logic Squared (PL2)
- Industry
- Technology, Information and Internet
- Company size
- 2-10 employees
- Headquarters
- san franciso, san franciso
- Type
- Privately Held
- Founded
- 2026
Locations
-
Primary
Get directions
san franciso, san franciso, US
Employees at Physician Logic Squared (PL2)
Updates
-
A physician can become faster while gradually being asked to think less. That possibility deserves more attention as clinical AI becomes part of everyday practice. The American College of Physicians new ethics position paper puts three ideas at the center of responsible AI use: Rationality, Self-Governance and Competence. We already spend a lot of time asking whether clinical AI is accurate, secure and efficient. Competence raises a different question: what happens to the physician's own reasoning after using the system every day? Does the technology help them examine evidence, consider a broader differential, recognize uncertainty and challenge the reasoning? Or does it increasingly turn clinical judgment into accepting or rejecting a finished answer? Those are very different models of clinical support. Good technology should reduce the burden of finding and organizing information without reducing the physician's responsibility to interpret it. It should create more room for thinking, not quietly make thinking optional. Clinical AI should not simply make physicians faster. It should preserve—and ideally strengthen—the habits that make good clinical judgment possible. How should we measure whether clinical AI is strengthening or weakening physician competence over time? #ClinicalAI #MedicalEthics #PhysicianLeadership
-
-
A clinical AI system does not have to invent a diagnosis to create a dangerous error. Sometimes it only has to invent the missing context. Michael Hobbs MD recently tested 10 frontier AI models using a straightforward pediatric case. Two pieces of clinical history were deliberately left out. In 57 of 140 responses — 41% — the models filled in the missing information anyway. They didn't necessarily fabricate wild medical facts. They made assumptions that allowed the clinical plan to continue. More interestingly, when Hobbs later asked what information was missing, 97% of the responses could identify the assumptions they had just made. That distinction deserves more attention. The model may "know" that information is missing and still proceed as though it were known. Hobbs then tested seven clinician-facing AI and decision-support tools. UpToDate Expert AI, for example, surfaced its assumptions and asked the clinician to clarify missing variables in all three runs. But the experiment raises a larger question for clinical AI: What should happen when the system does not have enough patient context to reason safely? Retrieving more papers is not the answer. Adding citations is not enough either. A citation can tell you where the medical evidence came from. It cannot tell you whether that evidence was applied to facts that were actually true about this patient. Clinical intelligence should make the boundary visible: What do we know? What are we assuming? What is missing? How confident are we? What does the physician need to verify before proceeding? Sometimes the most clinically intelligent answer is not another recommendation. It is: I need more information before I can answer this safely. That is the kind of clinical AI worth building.
-
-
Medicine keeps asking physicians to become more resilient. But resilience is a strange answer to a workflow problem. A physician should not need more resilience to find the lab result from eight months ago, search five notes for the treatment that worked, reconstruct what changed between three visits, or repeat a plan because the patient left without something clear they could refer back to. This work is necessary. But it is also consuming the same attention physicians need for interpretation, judgment, uncertainty and the person sitting in front of them. That is where decades of medical training create value. The goal of technology in medicine should be simple: take away the work surrounding clinical reasoning without taking the physician out of the reasoning itself. Organize the history, surface what changed, help reconstruct the timeline and make the plan easier to communicate. Then let the physician decide what it means and what happens next. We do not need physicians becoming better at carrying unnecessary cognitive load. We need better systems around them. Physicians: if you could permanently remove one manual task from your workday, what would it be? #PhysicianBurnout #ClinicalWorkflow #HealthcareTechnology
-
-
Three advanced AI models were asked to judge the same clinical reasoning. They disagreed 62.2% to 74.3% of the time. That should get the attention of anyone building clinical AI. A study published in the Journal of Medical Systems evaluated AI-generated clinical reasoning across 1,000 hospital cases. The researchers weren't simply asking whether the final diagnosis was correct. They examined something more important: Did the evidence presented in the reasoning actually support the conclusion? Three independent LLMs were used as judges. Agreement was low. In some cases, the same rationale could effectively be viewed as supported by one verifier and unsupported by another. A preliminary physician review of 50 cases added another important finding: agreement between the individual AI verifiers and the physician also varied. This doesn't mean AI cannot help evaluate clinical reasoning. The researchers themselves identify opportunities for selective automation when verifiers strongly agree. But it challenges a tempting assumption: If one AI might be wrong, simply ask another AI to verify it. Medicine is not that simple. A model can retrieve the right information and still connect it incorrectly. It can cite relevant evidence and still reach a conclusion the evidence does not justify. Another model may not reliably recognize the difference. For clinical decision support, physician review cannot become a ceremonial click at the end of an automated process. The physician needs to be able to inspect the reasoning, question the evidence, see uncertainty, modify the interpretation and retain final authority. At Physician Logic Squared, this is the direction we believe clinical intelligence needs to move: Review → Edit → Approve. AI should strengthen clinical judgment. It should not quietly inherit it. Study: Byun H, Lee D, Jung M, Jang B. Multi-Axial Analysis of Clinical Reasoning in Large Language Models: Inter-Verifier Disagreement and Its Implications for Automated Evaluation. Journal of Medical Systems, 2026. DOI: 10.1007/s10916-026-02440-y. #ClinicalAI #ClinicalReasoning #HealthcareAI #PatientSafety
-
-
The right treatment can still come at the wrong time. One thing the toughest cases have taught me is that finding several problems does not mean treating all of them at once. When I was younger in practice, I spent more time asking, “What can I treat here?” After more than two decades with difficult cases, I spend much more time asking, “What deserves attention first?” A patient may have GI dysfunction, inflammation, hormonal changes, poor sleep and nutritional deficiencies at the same time. You can identify all five correctly and still have an important clinical decision left to make: sequence. What needs stabilizing first? What may be downstream? What should be monitored while something else is addressed? What intervention might make sense later, but not yet? That is where experience changes clinical reasoning. Knowing a treatment is information. Knowing when it belongs in the course of a complex case requires context, treatment history, response and judgment. Sometimes the most sophisticated thing you can do is resist adding the next intervention until you understand what the patient needs first. For physicians managing complex cases: what most often makes you reconsider the order of a treatment plan? #FunctionalMedicine #ClinicalReasoning #IntegrativeMedicine #ComplexCare
-
-
Most clinical AI is still built around a moment. A physician asks a question. The AI searches the available information and returns an answer. A new perspective in npj Health Systems proposes a different model: bounded AI agents that continuously review the evolving EHR across notes, labs, orders, imaging and other documentation. Instead of waiting for someone to know what to ask, the system could assemble dispersed evidence into a source-linked brief for human review. Importantly, the authors are not proposing autonomous clinical decisions. Their model includes the timeline, supporting evidence, missing information and uncertainty — with a human reviewer deciding whether anything should be escalated. That distinction is important. In complex patient care, physicians often have the information they need. The problem is that it may be spread across six months of notes, multiple lab results, treatment changes and follow-ups. One abnormality rarely tells the whole story. What changed first? What happened after treatment? What persisted? What disappeared? What has not been explained? Clinical AI becomes much more useful when it helps physicians follow that story across time — without taking the final judgment away from them. A symptom is a snapshot. The timeline is the story. #ClinicalAI #HealthcareAI #ClinicalDecisionSupport #HealthTech #PhysicianLeadership
-
-
Physician Logic Squared (PL2) reposted this
In a new Nature Medicine study, researchers followed 53 patients with Crohn’s disease or ulcerative colitis who also had oral thrush. Patients received either oral nystatin or systemic fluconazole. Fluconazole reduced intestinal Candida. But it didn’t stop there. Researchers also saw changes in bacterial diversity, increases in short-chain-fatty-acid-producing bacteria, changes in microbial metabolites and improvement in disease activity over the eight-week follow-up. The easy conclusion is: maybe we should be treating more Candida. I don’t think the evidence gets us there. This was a small observational study in a very specific group of patients. It wasn’t randomized or placebo-controlled. It certainly doesn’t establish that Candida causes IBD or that fluconazole should become routine IBD treatment. What interests me is the ecosystem response. Earlier in my career, I spent much more time asking, “What do we need to kill?” After years of treating complicated GI patients, I’ve become much more interested in another question: What allowed this organism to become dominant in the first place? Because when you change one part of the microbiome, you may change much more than the organism you targeted. Fungi change. Bacteria change. Metabolites change. The environment changes. That’s systems biology in practice. A microbiome test gives us a snapshot. An intervention creates a new timeline. For physicians treating complex GI patients: after targeting one organism or pathway, what do you re-check to determine whether the rest of the ecosystem actually moved in the direction you expected? #MicrobiomeResearch #SystemsBiology #InflammatoryBowelDisease
-
-
The most dangerous clinical AI answer may be the one that sounds complete. Healthcare AI safety conversations tend to focus on hallucinations. Did the model invent a study? Fabricate a diagnosis? State something medically false? Those failures matter. But the NOHARM benchmark is drawing attention to a different problem: omission. An AI response can sound medically reasonable while failing to surface something important: a diagnosis that should have been considered, a test that may have changed the differential, a treatment response buried three years back in the chart, a contradiction between symptoms and laboratory findings, or a piece of history that changes the entire interpretation. That creates a harder safety problem. When AI fabricates something, a knowledgeable physician may recognize the error. When AI leaves something out, there may be nothing obviously wrong to challenge. The answer simply looks finished. This is why clinical AI cannot be evaluated only on whether its outputs are accurate. Physicians also need to ask: What evidence produced this conclusion? What alternatives were considered? Where is the uncertainty? What information could change the interpretation? What am I missing? For complex patients, the information may already exist across years of symptoms, labs, medications, specialist visits and treatment responses. The challenge is not simply retrieving it. The challenge is determining what matters, what connects, what conflicts and what has not yet been considered. Clinical AI should support that process without quietly taking ownership of the decision. Surface the evidence. Expose the uncertainty. Show the gaps. Let the physician review the reasoning. The goal should not be an AI system that always has an answer. It should be a clinical intelligence system that knows when the picture may still be incomplete. #ClinicalAI #HealthcareAI #ClinicalReasoning #PatientSafety #PhysicianAI
-