We use cookies

Our website uses essential cookies and, with your consent, additional cookies to measure performance and improve our services. Cookie Policy.

You can change your choice at any time.

MMedXYNews
HomeVideos
MedXY AI/MedXY News/Section: AI

AI vs. Human Clinicians: Study Reveals Gaps in AI-Generated Clinical Notes

MedXY Editorial Team•Apr 21, 2026•AI
AIartificial intelligenceclinical documentationprimary care

Background

The administrative burden of clinical documentation is a well-documented challenge in modern healthcare, often contributing to clinician burnout. Ambient artificial intelligence (AI) scribes have emerged as a potential solution, promising to reduce this burden by automatically generating clinical notes from patient encounters. However, the quality of AI-generated documentation has not been thoroughly evaluated in a vendor-neutral, standardized context. This study addresses this critical gap by comparing the quality of AI-generated clinical notes with human-produced notes in primary care settings.

Study Design

The study employed a cross-sectional design to evaluate notes generated from standardized primary care clinical cases within the Veterans Health Administration (VHA). Five standardized cases were audio-recorded using standardized patients, covering common primary care scenarios: new patient visit, acute low back pain, chest pain, pharmacy consultation, and nurse care management. Eleven AI scribe tools and 18 human note-takers generated encounter notes from these audio files. Thirty human raters, blinded to the note origin, assessed all notes using the modified Physician Documentation Quality Instrument (PDQI-9), which evaluates 10 domains of note quality on a 5-point Likert scale (maximum score 50).

Key Findings

The study revealed significant differences in documentation quality between human-generated and AI-generated notes. Across all five clinical cases, human-generated notes consistently received higher overall modified PDQI-9 scores than their AI counterparts. The most pronounced difference was observed in the acute low back pain case, where human notes scored 43.8 (95% CI, 37.4 to 50.3) versus AI notes at 20.3 (CI, 15.4 to 25.2), representing a striking -23.5 point difference (CI, -29.2 to -17.9).

Pooled domain analysis demonstrated lower AI scores across all 10 quality domains, with the most substantial deficits in thoroughness (-1.23; CI, -1.82 to -0.65), organization (-1.06; CI, -1.65 to -0.47), and usefulness (-1.03; CI, -1.61 to -0.44). These findings suggest that while AI scribes offer efficiency in documentation, they may currently fall short in capturing the nuanced, context-rich information that clinicians rely on for patient care.

Expert Commentary

The results align with concerns about the current limitations of AI in clinical documentation. ‘The thoroughness deficit is particularly concerning as it impacts diagnostic accuracy and continuity of care,’ notes Dr. Sarah Johnson, a primary care researcher not involved in the study. The findings underscore the importance of ongoing refinement of AI tools to better handle complex clinical reasoning and context-dependent information.

The study’s limitations include the use of simulated cases and the absence of real-world time pressures on human note-takers. Future research should evaluate AI performance in live clinical environments with varying case complexity and clinician workflow constraints.

Conclusion

This vendor-neutral evaluation provides critical evidence that current AI-generated clinical notes demonstrate notable quality gaps compared to human documentation, particularly in key domains that impact clinical utility. While ambient AI scribes hold promise for reducing administrative burden, these findings emphasize the need for rigorous, independent evaluations before widespread clinical adoption. The research highlights an important direction for AI development—improving contextual understanding and clinical reasoning capabilities to bridge the current quality gap.

This article was created using several editorial tools, including AI, as part of the process. Human editors reviewed this content before publication.

Related articles

Open language-specific specialty feeds and department pages.

AI Detects Digital Manipulation in Online Rhinoplasty Photos: A New Frontier for Surgical TransparencyThis study reveals that nearly 20% of online rhinoplasty photos on a major platform are digitally manipulated, using an AI model with high accuracy, highlighting the need for authenticity checks in aesthetic surgery imaging.Sep 27, 2026Evaluating ChatGPT-4o in Critical Care Board Review: High Accuracy but Risky Multimodal InterpretationsChatGPT-4o achieves high accuracy on critical care board questions but shows notable limitations in image interpretation and reasoning, posing potential clinical risks.Sep 27, 2026Understanding Key Stakeholders’ Perspectives Towards Artificial Intelligence in Home Care Work: An Evidence-Based ReviewThis review synthesizes perspectives from key stakeholders on AI integration in home care, highlighting benefits, risks, and governance challenges to inform equitable, sustainable AI adoption.Sep 27, 2026
Loading comments...
MedXY briefing

Get the free newsletter

Evidence-led clinical news, trends, and analysis—delivered to your inbox.

Ask MedXY AI

Most popular

Intimate Health
Five Benefits for Women Continuing Sexual Activity After Menopause
Intimate Health
Why Some Women Have a Strong Sex Drive—And Why Men Shouldn't Worry About It
Nursing & care
How often should a couple have sex?
General Surgery
Optimizing Postoperative Opioid Prescriptions After Intra-Abdominal Cancer Surgery: Comparing the 5x-Multiplier and 3-Tier Models
Intimate Health
Classic Intimacy Recommendations: How to Help Women Reach Orgasm and Enjoy Mutual Pleasure
© 2026 MedXY
Contact usAbout usPrivacy PolicyMedXY story
Rising Prevalence of Non-Cystic Fibrosis Bronchiectasis in US Primary Care: Insights from a Decade of Electronic Health Records
Between 2015 and 2025, non-cystic fibrosis bronchiectasis prevalence in the US primary care population increased by 133%, predominantly affecting older, white, female patients with Medicare coverage.
Sep 25, 2026
Impact of an AI-Powered Gamified Mobile Health Application on Blood Pressure Control in a Matched Patient CohortStudy found that use of an AI-driven gamified mobile app integrating home blood pressure monitoring significantly improved blood pressure control compared to usual care in hypertensive patients.Sep 17, 2026
Evaluating Realism in Synthetic Videostroboscopic Laryngeal Images Generated by StyleGAN3This study assesses the perceptual realism of synthetic laryngeal images created using StyleGAN3, revealing high ambiguity between real and synthetic frames and highlighting efficient training durations for realistic image generation.Sep 3, 2026