Dr Artin Entezarjou, PhD

Medical Operations, Sweden

The Clinician’s Perspective

We asked doctors & AI to judge 760 clinical facts. Here's what happened when they disagreed.

Dr Artin Entezarjou, PhD

Medical Operations, Sweden

Last year I wrote that AI cannot replace our clinical judgment yet.

Two things informed that position. In a study my colleagues and I published in BMJ Open, we found that GPT-4 underperformed compared to clinicians when assessing complex primary care cases with multiple comorbidities and ambiguous presentations (Arvidsson et al., BMJ Open, 2024). And my former research group PETRA was conducting systematic reviews of clinician and patient perspectives on AI in primary care. Based on studies published up to February 2024, the qualitative literature on AI in primary care was still relatively scarce, and research on large language model capabilities specifically was limited (Mundzic et al., JMIR AI, 2026; Bogdanffy et al., JMIR AI, 2026).

That was the state of things roughly eighteen months ago. Since then, the models have progressed immensely. Academic publications are just now reaching print with findings based on last year's reasoning models, while we in the industry are already working with the next generation. The gap between what clinical AI companies observe in practice and what the peer-reviewed literature reports is growing.

In this article, I want to lay out what is emerging from both sides: recently published research from leading clinical AI groups, and internal evaluation data from Tandem Health, where I lead clinical AI evaluation and continue to practise part-time. More than 70% of our commercial team have clinical backgrounds. We work with these models daily. We intend to publish our findings formally in a peer-reviewed journal, but given the pace of development and the gap between industry observation and academic publication, we believe there is value in sharing some of our early data now.

From fragmented performance to consistent reasoning

The progression in model capability over the past year follows a pattern worth understanding.

When my colleagues and I tested GPT-4 on complex primary care scenarios in our BMJ Open study, the model struggled with exactly the kind of cases that define general practice: patients with overlapping conditions, ambiguous symptoms, and contextual factors that shape clinical decisions. This was consistent with what most clinicians experienced when they first tried these tools. Impressive on structured questions, unreliable on the messy reality of clinical work.

An earlier randomised trial showed a more nuanced picture. On structured diagnostic reasoning tasks, GPT-4 working alone actually outperformed both physician groups, including physicians who had access to GPT-4 as an aid. The model was strong where the task was well-defined, but the gap with human judgment on complex, contextual reasoning remained (Goh et al. JAMA Network Open, 2024).

Then came the reasoning models. A study evaluated OpenAI's o1 across multiple clinical reasoning tasks, including blinded second opinions on real emergency department cases. This was qualitatively different from what GPT-4 had achieved. The model matched or exceeded attending physicians in diagnostic accuracy, including in high-uncertainty triage moments with incomplete information. Physician raters asked to distinguish AI-generated differentials from human ones could not reliably do so (Brodeur et al. Science, 2025).

Chen et al., authors on several of these studies, have reflected that these results challenge a principle long taught in informatics: that human + computer should always outperform either alone. They argue that what remains uniquely human clusters into three qualities: competence (judgment under genuine uncertainty), communication (influence and advocacy), and character (professional accountability).

These are useful distinctions, but they are moving boundaries, not fixed ones. Each model generation narrows what counts as "genuine uncertainty." The question is no longer whether AI should assist clinicians, but where exactly human work still positively changes the outcome, and how we prove it. That requires measurement infrastructure, not assumption. It is what we have been building at Tandem.

The trajectory from GPT-4 to reasoning models represents a step change, not incremental improvement. And the models we are evaluating internally at Tandem have moved beyond what has been published so far.

Testing AI as a clinical audit partner

At Tandem, our AI medical scribe is used by thousands of clinicians across Europe. As part of ensuring clinical safety, we maintain an evaluation suite of automated tests that runs whenever changes are made to the product. This suite uses synthetic consultation cases designed to verify that the system continues to capture clinically significant information reliably.

The evaluation described here is one of many experiments we run internally as part of operating Europe's first MDR Class IIa certified AI medical scribe. It is not a standalone clinical trial. It's ongoing work to understand where the system performs well and where it does not.

One emerging approach in AI evaluation is often called "LLM as judge": using a language model not to produce the clinical output, but to help assess whether that output meets a defined standard. In medicine, that cannot simply be assumed. Human clinicians remain the benchmark for clinical quality, so the question is whether AI can help clinicians perform the audit itself - and where human judgment still changes the result.

As part of developing this evaluation infrastructure, we wanted to understand how well humans and AI respectively identify which facts in a consultation are clinically significant. Our clinicians, myself included, independently assessed 760 clinical facts from synthetic but clinically realistic consultations for clinical significance: does this fact matter for patient safety? We then compared these blinded human assessments against the AI's independent assessment of the same facts.

We built it for scribe safety work, but the same measurement problem shows up anywhere you rely on a model's clinical output: you need a way to prove when human oversight still changes the result.

For this clinical-significance labelling task, the AI's sensitivity was 95.8%. Clinicians, working without the AI's input, reached 55% sensitivity. When we then revealed the AI's assessments and reasoning to the clinicians, they revised their judgment to align with the AI in roughly 95% of disagreements. In the remaining 5%, clinicians reviewed the AI's reasoning and maintained their original assessment, having identified risks the AI had missed.

To be precise about what these numbers describe: this is not a measure of scribe accuracy or product error rates. It is a measure of audit quality: how well humans and AI respectively judge clinical significance when evaluating the same set of facts. When the two disagreed, the AI's judgment held up 95% of the time and the clinician's held up 5% of the time. Both directions matter.

The mechanism is straightforward. The AI processes the full clinical context simultaneously. Clinicians apply heuristics, as we always do. Those heuristics work well at the bedside, where speed and pattern recognition are essential. Under systematic evaluation, they proved less granular than an AI weighing every available detail.

A legitimate question: does the 95% revision rate reflect genuine improvement, or are clinicians being anchored by AI reasoning? We cannot fully separate the two from this evaluation alone. Automation bias is well documented in clinical decision support research. But clinicians did not defer uniformly, as the 5% demonstrates, and the pattern is consistent with findings from Brodeur et al. and Goh et al. using entirely different study designs with no reveal step and no anchoring opportunity.

The 5% and why systems matter

The 5% of cases where clinicians identified a risk the AI missed deserve attention.

In healthcare, 5% is not a small number. Across thousands of clinical encounters, it represents real patients whose safety depends on a second set of eyes.

We found those cases because we built the evaluation infrastructure to look for them. Under the EU Medical Device Regulation, our coding tool is classified as a Class IIa medical device. That classification requires systematic clinical evaluation as a condition of market access.

But the point goes deeper than regulatory compliance. Today, human review demonstrably improves outcomes - our 5% proves it. We design our products accordingly. The honest observation is that this boundary is moving. Building the infrastructure to track exactly when oversight still changes the result matters more than assuming it always will.

This is also what generates the data to understand where human oversight adds the most value. Without structured evaluation, the 5% remains invisible. Products that have not invested in this infrastructure are asking clinicians and patients to accept an unknown level of risk. The regulatory frameworks exist for good reason. They are the starting point for clinical AI that earns trust.

A study from the PETRA research group, which I co-founded, surveyed physicians and the general population in Sweden on what accuracy threshold they expect AI to meet before they trust it (Arvidsson et al., BMJ Health & Care Informatics, 2026). The answer: performance at or above human level. That expectation is reasonable. For defined clinical reasoning tasks, the evidence suggests we are approaching it. The 5% is a reminder that approaching it is not the same as having arrived, and that the system around the AI matters as much as the AI itself.

The clinician of the future

Consider the typical general practitioner in Sweden today. Responsible for roughly 1,750 patients. Managing prescription renewals, lab result explanations, triage messages, referral letters, documentation. Reactive and time-pressed. The administrative overhead erodes the thing that makes primary care valuable: the longitudinal relationship between clinician and patient.

Now consider what becomes possible when AI reliably handles routine cognitive tasks. That same GP, still responsible for 1,750 patients, able to deliver care as if the list were 1,100. That's the benchmark Sweden has set, and only 8% of primary care centres meet it today (Swedish Medical Association, 2025). Not by working harder, but because the AI manages the administrative load that currently consumes the majority of clinical time: renewals, lab explanations, triage, scheduling, follow-up. Not through digital portals that exclude patients who cannot navigate them. Through AI voice agents that meet patients in conversation, in their language, on their terms.

The GP's role shifts from processing volume to full presence. The father and daughter trying to understand a suspected Alzheimer's diagnosis. The patient who needs to hear, face to face, that a rapid cancer workup is necessary. The moment where the right decision is to deviate from the guideline, because this patient's values and circumstances require a different path, and that decision only works inside a relationship built on trust.

AI can surface the evidence, rank the options, and summarise the clinical picture. But no one should receive difficult news from a system that has never experienced what it means to grieve, or to hope, or to be afraid.

Which loops matter

Human oversight still improves outcomes in clinical AI. Our own 5% is the proof, and our regulatory process is how we found it.

What is becoming clear is that the field needs to map, task by task, where human oversight still improves outcomes, and where it does not. Goh et al. already showed that physicians with AI access underperformed the AI working alone on structured reasoning. That finding should not be comfortable, but it should not be ignored. For some cognitive tasks, adding a human checkpoint may already reduce system performance. For others, like the 5% we identified, human review remains essential. The distinction will sharpen with each year of evaluation data, and building the infrastructure to track it - structured, clinician-driven assessment against clinical standards - matters more than defaulting to oversight as an article of faith.

This is the work we are doing at Tandem. And the clinicians using AI tools daily are among the most important voices in this conversation, because they know which moments truly require their judgment and which are consumed by tasks a well-evaluated AI could handle safely. Organisations that bring clinicians into the design of these workflows, alongside operational and technical leadership, will build systems that are both safer and more effective.

After Garry Kasparov lost to Deep Blue in 1997, he promoted "centaur chess," where humans played alongside AI engines. For a time, the best results came from human-AI teams. Then the engines matured enough that human input became a liability, and centaur tournaments quietly disappeared.

The lesson was not that humans became irrelevant. The lesson was that understanding roles made both humans and machines better.

Last year I wrote that AI cannot replace our clinical judgment. What I have come to believe since then is that much of what we call clinical judgment is analytical, and AI is rapidly matching it. But there is a different set of skills, uniquely human, that no model can replicate: the ability to sit with someone in uncertainty, to build trust over years, to advocate for a patient whose needs do not fit the algorithm. These are not leftovers after AI takes the cognitive work. They are the core of what medicine has always been. The sooner we make space for them, the better healthcare becomes for everyone.

Artin Entezarjou, MD, PhD, is a specialist physician and Medical Product Lead at Tandem Health, where he leads clinical evaluation and correctness across the company's AI products. He holds a PhD in applied AI and telemedicine and co-founded the PETRA research group during his academic career. He continues to practise clinically part-time.

Get started with Tandem today

Join thousands of clinicians enjoying stress-free documentation.

Get started with Tandem today

Join thousands of clinicians enjoying stress-free documentation.

Get started with Tandem today

Join thousands of clinicians enjoying stress-free documentation.