
Every time a person sees a doctor, the visit produces two kinds of records. There are the tidy checkbox fields: the billing codes, the lab values, the prescription entries. And there is the note the clinician actually writes, describing what the patient said, how they are coping, what side effects they mentioned and why a medication was changed.
That written note is where most of the medical record lives. It is also where most medical research has never looked.
In a study published today in Nature Medicine, researchers at RespondHealth, co-led with physicians and statisticians at Drexel University and collaborators at Stanford University, the University of Miami, the University of Pennsylvania and the Icahn School of Medicine at Mount Sinai, describe a system that accurately reads those written notes at scale. It then converts what it finds into data that can be computed and analyzed and keeps every extracted fact linked back to the exact sentence in the chart it came from.
Checking the machine’s work
The central question with any AI system reading medical records is whether to believe it. A system that misreads “denies chest pain” as “chest pain” does not just make an error; it manufactures a patient history that never happened.
The researchers built the study around that problem. Board-certified physicians at independent institutions reviewed the system’s output line by line against the original notes. One physician reviewed 100 patient visits drawn from 96 different providers and 86 practices across four different electronic health record systems. A second physician independently reviewed a subset of those visits, and a third physician, who did not know which reviewer had said what, settled the disagreements to produce an adjudicated standard.
Measured against that adjudicated standard, the system scored 99.4% on a combined measure of how often it was right about what it reported and how much of what was in the note it caught. The physicians themselves agreed with each other on 94.7% of the statements they both reviewed, which is a reminder that expert judgment has its own variability. The time difference is stark in a different way. The five physician reviewers needed roughly 70 hours to work through 120 charts. The system reviewed each chart in seconds.
To show what the approach makes possible, the team applied it to one of the most widely discussed drug classes in medicine: GLP-1 receptor agonists, the class that includes semaglutide and tirzepatide.
What the researchers found about GLP-1 medications
The team followed more than 16,000 adults who started a GLP-1 medication in routine outpatient care. Some were using the medication to treat diabetes and others for weight loss. The study tracked what happened to them over the following months.
Two patterns stood out, and they ran in opposite directions.
People who started with healthier blood sugar lost more weight and lost it faster. At 12 months, people with normal blood sugar at the start had lost about 7.7% of their body weight, compared with about 2.7% among people whose diabetes was poorly controlled when they began. The time it took to reach 5% weight loss followed the same pattern: about 210 days for people with normal blood sugar and about 413 days for people with poorly controlled diabetes.
Blood sugar improvement went the other way. The people who started with the worst blood sugar control improved the most and improved the fastest. Weight loss also depended on who was taking the medication. At 12 months, women had lost about 6.1% of body weight compared with about 4.0% for men, and adults ages 20 to 39 had lost about 8.1% compared with about 5.1% for those 40 and older.
The researchers stress that this was an observational study of people in ordinary clinical care, not a controlled trial, and that it cannot by itself establish cause and effect. What it can do is show how these medications actually behave outside the tightly managed conditions of a trial.
The part that has been invisible
The study also looked at things that appear almost nowhere in coded medical data because no billing code captures them.
Among patients with a documented baseline, depression scores improved most in those who started out most depressed. In the group with moderate to severe depressive symptoms at the outset, scores fell by about six points on a standard nine-question screening scale over 12 months. Reported pain intensity fell as well, again most clearly among patients who started with the highest pain. Waist circumference declined by about 2.7 inches (6.9 centimeters) over a year in the smaller group of patients for whom it was recorded.
These findings are preliminary and based on modest numbers of patients, and the authors present them as observations worth pursuing rather than conclusions. But they would not have been observable at all without the ability to read clinical notes because these measurements were written down in prose, not entered into a database field.
The same is true of much of the core data. Across the full study population, 70% of the blood sugar readings used in the analysis were found only in written notes. About a third of the evidence that a patient was actually taking a GLP-1 medication came only from notes. Thousands of patients could be included in the analysis only because their starting measurements were recovered from narrative text; using coded fields alone, they would simply have been dropped.
Why it matters
“Most of what a clinician knows about a patient is written in sentences, not stored in fields,” said Vicki Seyfert-Margolis, Ph.D., chief executive officer and founder of RespondHealth and a senior author on the study. “For decades we have built evidence about how medicines work in the real world while looking at a fraction of the record. What this work shows is that the rest of it can be made analyzable, and made analyzable in a way that a physician can check.”
Charles B. Cairns, M.D., dean of the College of Medicine at Drexel University and a senior author, added, “The point is not that the machine replaces the reviewer. The point is that every number this system produces can be traced back to the sentence in the chart that supports it. A clinician can look at the evidence and disagree. That is what makes it usable in medicine.”
The authors note that the method is not specific to GLP-1 medications or diabetes. The same approach can be applied to any condition where the important details are recorded in narrative form, which is to say most of them.
Publication details
Edward Kim et al, Computable longitudinal patient journeys from structured and unstructured EHR data, Nature Medicine (2026). DOI: 10.1038/s41591-026-04695-x
Journal information:
Nature Medicine
Key medical concepts
Provided by
RespondHealth
Citation:
AI reads doctors’ notes at scale, revealing data absent from coded medical records (2026, September 18)
retrieved 19 September 2026
from https://medicalxpress.com/news/2026-09-ai-doctors-scale-revealing-absent.html
This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no
part may be reproduced without the written permission. The content is provided for information purposes only.

