experiment · iScience · la publicación, 10 sep 2026 · gratis
Una herramienta de IA podría ayudar a decidir quién necesita una tomografía urgente
En un estudio chino, un programa que hace tres preguntas seguidas detectó más casos que un modelo anterior, pero los propios autores piden que nadie lo use todavía en clínicas.
Versión breve · la versión detallada sigue, unos 6 min
- El estudio, de un vistazo
- Quiénes
- Pacientes con sospecha de disección aguda de aorta
- Cuántos
- 869 pacientes de tres hospitales chinos y 200 de otros dos
- Dónde
- China central
- Cuándo
- De enero de 2022 a octubre de 2023
- Tipo de estudio
- análisis de lo que ya estaba en historias clínicas
- Quién lo hizo
- Hospital Tongji y Universidad Normal de China Central, Wuhan
- El límite que importa
- Los autores piden no usarlo en clínicas hasta que haya estudios con pacientes nuevos.
Médicos de atención primaria frente a los programas de IA, en 50 casos reales
Los médicos no detectaron el problema el 19.6% de las veces y dieron falsas alarmas el 53.3%; los programas de IA tuvieron menos casos no detectados.
GPT-4o tuvo el resultado más equilibrado: 8.0% de casos no detectados y 4.0% de falsas alarmas.
Los propios autores son claros: no debe usarse clínicamente hasta que se hagan estudios prospectivos, es decir, con pacientes nuevos seguidos en tiempo real y no con historias clínicas ya cerradas.
El artículo recoge un estudio en el que, cada hora que pasa sin diagnosticar una disección aguda de aorta, la muerte aumenta entre 1% y 2% entre quienes la sufren. La tomografía con contraste (CTA) es el examen que confirma el diagnóstico, pero no siempre está disponible de inmediato, sobre todo en hospitales pequeños o durante un traslado en ambulancia.
Investigadores de Wuhan y de la Universidad Normal de China Central construyeron AAD-Agent, un asistente de inteligencia artificial que pide la información en tres pasos: primero la edad, el sexo y los síntomas; después los conteos de glóbulos rojos y blancos; y por último el D-dímero. El programa da una respuesta después de cada paso y se detiene si ya sospecha el caso.
El estudio, retrospectivo y multicéntrico, reunió a 869 pacientes de tres hospitales chinos y a 200 de otros dos. En el grupo externo, el programa mantuvo la capacidad de detectar casos reales por encima de 0,80, mientras que un modelo estándar de aprendizaje automático se derrumbó: su puntuación F1 fue de 0,964 en el grupo interno y de 0,600 en el externo. En comparación con 18 médicos de atención primaria evaluados en 50 casos, los programas de IA tuvieron menos casos no detectados.
Sin embargo, el estudio infló a propósito la proporción de enfermos: entre 70% y 78% de los pacientes tenían disección, cuando en una sala de urgencias real con dolor de pecho esa cifra es menor al 1%. Por eso la exactitud y la capacidad de acertar un positivo están infladas, aunque los autores calcularon que, con una proporción realista de enfermos, la capacidad de descartar un caso negativo sería muy alta.
Además, todos los datos provienen de cinco hospitales del centro de China. Nadie sabe cómo funcionaría en América Latina, Estados Unidos o Canadá, donde los pacientes, los registros médicos y los protocolos son distintos.
Los propios autores son claros: no debe usarse clínicamente hasta que se hagan estudios prospectivos, es decir, con pacientes nuevos seguidos en tiempo real y no con historias clínicas ya cerradas. La herramienta está pensada solo para descartar, no para confirmar. Una predicción negativa no reemplaza el criterio del médico ni la tomografía.
Si esos estudios confirman lo que este trabajo sugiere, un chequeo barato y rápido podría ayudar un día al personal de urgencias de clínicas sin escáner las 24 horas a decidir quién necesita traslado o imágenes urgentes.
Si usted o un familiar llega a urgencias con dolor repentino de pecho o espalda, pregunte si el centro tiene tomografía disponible y, si no la tiene, cómo se decide el traslado. Esa pregunta sigue siendo suya, no del programa.
Qué significa para usted
Por ahora no hay nada que pedir ni usar: ningún hospital de América Latina, Estados Unidos o Canadá ha probado esta herramienta, y los propios autores piden esperar estudios con pacientes nuevos antes de aplicarla. Si le preocupa un dolor repentino de pecho o espalda, la pregunta útil sigue siendo si el centro tiene tomografía y cómo se decide un traslado, y eso conviene consultarlo con el profesional que lo atiende.
Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503
Quién pagó: Financiado por la Fundación Nacional de Ciencias Naturales de China (subvenciones 72574075 y 72274067) otorgadas a L.Y.; los financiadores no tuvieron ningún papel en el diseño del estudio, la recogida de datos, el análisis, la interpretación, la redacción ni la decisión de publicar, y no se prestó ningún equipo ni software: los modelos de lenguaje se usaron a través de sus API públicas.
No tome esto como consejo médico profesional.
Los hallazgos de otros estudios que aquí se mencionan los conocemos por este documento, que fue el que leímos; no abrimos cada uno de esos estudios.
Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1220 palabras · unos 6 minLeerla →Cerrar
Un agente de IA para detectar un desgarro de aorta: qué probó el estudio y qué le falta
Se trata de un estudio retrospectivo en cinco hospitales de China central; los propios autores piden que no se use en pacientes todavía.

La aorta es la arteria grande que sale del corazón. Cuando su pared interna se desgarra, la sangre se filtra entre sus capas y la persona puede morir en horas. El artículo recoge un estudio en el que, cada hora que pasa sin tratarse, el riesgo de morir sube entre un 1% y un 2%1. El artículo recoge un estudio en el que el diagnóstico a tiempo sigue siendo difícil, sobre todo donde el examen que la confirma —una tomografía con contraste, la CTA— no está disponible2.
El artículo recoge un estudio en el que casi todas las tomografías que se piden por sospecha de este desgarro resultan innecesarias: hasta un 97.3% de ellas3. El artículo recoge un estudio en el que eso significa radiación, contraste que puede dañar los riñones y gastos que se podrían evitar3. Y el artículo recoge un estudio en el que, en China, donde se hizo este estudio, la carga es mayor porque en la atención primaria hay menos experiencia y menos recursos para detectar el problema a tiempo4.
Frente a eso, un equipo de médicos de urgencias del Hospital Tongji y de informáticos de la Universidad Normal de China Central, en Wuhan, construyeron una herramienta llamada AAD-Agent. No es un aparato: es un programa que usa un modelo grande de lenguaje —de los mismos que hay detrás de los chatbots— y lo obliga a razonar por pasos, como lo haría un médico en la sala de urgencias5.
Funciona así: primero le entregan la edad, el sexo, los síntomas y la presión arterial; el programa dice si sospecha o no. Si no sospecha, le añaden los resultados del hemograma; si sigue sin sospechar, le añaden el dímero D, un análisis de sangre. En cada etapa puede detenerse y, cuando sospecha, recomienda la tomografía5. La idea no es reemplazar ese examen, sino ayudar a descartar el problema en quien tiene poco riesgo y a priorizar a quien sí lo tiene6.
El estudio comparó esa herramienta con un modelo de aprendizaje automático —el tipo clásico, entrenado con datos— usando 869 pacientes de tres hospitales chinos para la validación interna y 200 de otros dos hospitales para la validación externa. El modelo clásico fue excelente en casa y se derrumbó afuera. Los agentes de lenguaje, en cambio, se mantuvieron más estables: DeepSeek-R1 acertó el 71.5% de los casos, GPT-3.5 el 71.5% y GPT-4o el 71.0%7.
Hay un número que los autores destacan y que conviene entender bien. La mayoría de los pacientes del estudio sí tenía el desgarro: alrededor del 70% en el primer grupo y el 78% en el segundo. En una sala de urgencias real, en cambio, menos del 1% de quienes llegan con dolor de pecho lo tiene. Cuando los autores rehicieron la cuenta suponiendo ese 1% realista, el valor de un resultado negativo —es decir, la confianza de que quien da negativo de verdad no tiene el problema— llegó a cerca del 99.6%8. Traducido: si el programa dice que no, es muy poco probable que se equivoque.

Esta familia de estudios ya tiene antecedentes. Un análisis de tres grandes programas europeos de mamografía evaluó si la inteligencia artificial podía integrarse al doble lectura que usan esos programas para detectar cáncer de mama. Según el resumen de ese trabajo —solo pudimos leer el resumen, el artículo completo está detrás de una suscripción—, en 597,419 exámenes la detección aumentó en alrededor de 1 caso por cada 1,000, sin un aumento consistente de las llamadas a repetir estudios910. Es decir, la IA como lector complementario, no como sustituto, y con controles de calidad explícitos11.
Volvamos al estudio chino, porque sus límites son tan importantes como sus resultados. Sus propios autores advierten que la herramienta no sirve para confirmar el diagnóstico, solo para descartarlo, y que eso queda pendiente de estudios prospectivos12. Y son tajantes: recomiendan fuertemente no usarla en la práctica clínica hasta que esos estudios se hagan13.
¿Por qué tanta cautela? Porque todos los datos vinieron de cinco hospitales de China central. El desempeño en otras regiones, otras poblaciones y otros sistemas de salud no se ha probado. Además, el estudio excluyó al 59% de los pacientes elegibles, y los excluidos tenían una proporción de hombres y una frecuencia del desgarro distintas, lo que indica un sesgo de selección. Tampoco se pudo comparar de frente con la estrategia establecida —el puntaje ADD-RS combinado con el dímero D— porque faltaban variables en los registros. Y no había datos sobre si la falsa luz de la disección estaba abierta o cerrada, así que no se pudo evaluar el desempeño por subtipo.
La comparación con médicos también merece cuidado. Dieciocho médicos de atención primaria en formación, con unos 4.2 años de experiencia, evaluaron 50 casos reales presentados como relatos clínicos estructurados. Fallaron en no detectar el problema el 19.6% de las veces y dieron falsas alarmas el 53.3%14. Los agentes lo hicieron mejor: GPT-4o tuvo la tasa más equilibrada, con 8.0% de casos no detectados y 4.0% de falsas alarmas15. Pero son 50 casos, en un entorno simulado, no en el ajetreo real de una urgencia.

Los autores declaran que el trabajo lo pagó la Fundación Nacional de Ciencias Naturales de China, con dos subsidios16. Y aclaran que los financiadores no tuvieron papel en el diseño, el análisis ni la decisión de publicar16.
Así lo leemos nosotros. Lo que este estudio muestra no es que la IA sea mejor médico, sino algo más específico y más útil: que un sistema entrenado en un lugar puede volverse frágil cuando cambia el mundo donde se lo usa. El modelo clásico aprendió de memoria los números de tres hospitales; cuando los números cambiaron —otra presión arterial, otro dímero D—, sus reglas dejaron de servir. El agente de lenguaje, en cambio, trabaja sobre descripciones escritas, con el contexto, y eso lo hace menos dependiente de cifras exactas. Si eso se confirma, lo que deberíamos esperar es que herramientas así toleren mejor las diferencias entre hospitales, regiones y poblaciones. Sabríamos que nos equivocamos si al probarlas en hospitales muy distintos mantuvieran la misma sensibilidad y especificidad, o si sus errores no cambiaran al cambiar la población.
Lo que usted puede hacer con esto es concreto. Si un familiar llega a urgencias con dolor súbito de pecho o de espalda, pregunte si la decisión de hacer o no la tomografía se basa en un protocolo probado en ese mismo hospital, y pida que le expliquen cómo se ha medido la seguridad de cualquier herramienta nueva en pacientes como usted. Y cuando oiga hablar de un avance médico con inteligencia artificial, pregunte si ya se usa en hospitales reales y si hay estudios que midan sus resultados; comparta esa pregunta con sus vecinos, para que nadie confunda una promesa con una prueba.
Lo que este estudio deja claro es que el camino está abierto, pero no recorrido: los autores mismos piden validación prospectiva en poblaciones reales de urgencias antes de cualquier uso clínico17. Mientras tanto, la pregunta que usted puede llevar consigo a cualquier consulta es sencilla: ¿este protocolo se probó en pacientes como yo?
De dónde sale cada dato de contexto, y cuánto leímos de cada documento
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "As a time-sensitive disease, the mortality rate rises 1%–2% per h post-onset."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Despite recent advances, timely diagnosis remains difficult, particularly in primary care and resource-limited settings where computed tomography angiography (CTA), the confirmatory standard, is not readily available."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Moreover, unnecessary CTA scans, which account for up to 97.3% of CTAs ordered for suspected AAD, incur radiation exposure, contrast-induced nephropathy, and excess costs."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "In China, the clinical burden of AAD is compounded by high mortality, substantial costs, and the frequent misdiagnosis or delayed detection, especially in primary care settings with limited diagnostic expertise and resources."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "To address this gap, we developed AAD-Agent, an LLM-enhanced autonomous AI agent that emulates hierarchical emergency reasoning. It actively acquires data in three stages (demographics/symptoms→routine blood tests→D-dimer), dynamically updates risk estimates, and provides interpretable outputs."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Designed as a rapid pre-imaging triage aid, AAD-Agent aims not to replace CTA but to safely rule out low-risk patients and defer unnecessary imaging, particularly in resource-limited settings or during initial triage to prioritize high-risk patients for confirmatory CTA."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "In contrast, AAD-Agent maintained external stability across LLMs: DeepSeek-R1 (accuracy 0.715, F1-score 0.820), GPT-3.5 (0.715, 0.825), and GPT-4o (0.710, 0.803)."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Although the case-enriched design inflates these metrics, prevalence-adjusted negative predictive value (≈99.6% at 1% prevalence) supports AAD-Agent as a safe, low-cost rule-out aid to defer CTA in low-risk patients."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I"
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Its low external specificity precludes confirmatory diagnosis, limiting use to rule out pending prospective validation."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "We strongly caution against clinical use until such prospective studies are completed."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "As shown in Figure 6 , primary care physicians exhibited a high miss rate of 19.6% and an extremely elevated misdiagnosis rate of 53.3%. In stark contrast, all evaluated LLMs demonstrated substantially superior performance across both metrics."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "GPT-4o presented the most balanced performance, with a false-negative rate of 8.0% (sensitivity = 92.0%) and the lowest false-positive rate of 4.0%."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The work was funded by the 10.13039/501100001809 National Natural Science Foundation of China (grant nos. 72574075 and 72274067 ), both awarded to L.Y."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Prospective validation in unselected, real-world emergency populations is essential before any clinical application."
Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503
Quién pagó: Financiado por la Fundación Nacional de Ciencias Naturales de China (subvenciones 72574075 y 72274067) otorgadas a L.Y.; los financiadores no tuvieron ningún papel en el diseño del estudio, la recogida de datos, el análisis, la interpretación, la redacción ni la decisión de publicar, y no se prestó ningún equipo ni software: los modelos de lenguaje se usaron a través de sus API públicas.
No tome esto como consejo médico profesional.
Los hallazgos de otros estudios que aquí se mencionan los conocemos por este documento, que fue el que leímos; no abrimos cada uno de esos estudios.
experiment · iScience · the paper, 10 Sep 2026 · free
An AI Triage Tool for a Killer Heart Emergency Held Up Where a Standard Model Fell Apart — But Its Real-World Accuracy Is Still Unknown
In a small, case-heavy study from central China, an AI agent kept catching aortic dissection when a conventional model collapsed. The catch: the patients were chosen so that most already had the disease, and the authors say it must not be used in clinics yet.
Short version · the longer version follows, about 7 min
- The study at a glance
- Who
- patients suspected of acute aortic dissection
- How many
- 869 patients from 3 Chinese hospitals and 200 from 2 additional hospitals
- Where
- China
- When
- January 2022 to October 2023
- Kind of study
- analysis of what people did
- Who did it
- Tongji Hospital and Central China Normal University
- The limit that matters
- Patients were chosen so most already had the disease; real-world accuracy is unknown.
These are shares of 50 test cases judged by 18 primary care doctors and by one AI agent (GPT-4o); the cases were picked so half already had the disease, so the numbers would differ in a real emergency room.
A staged AI agent against a standard machine-learning model, on hospitals where neither was developed
The standard model's combined score for catching cases and avoiding false alarms fell from 0.964 where it was developed to 0.600 on outside hospitals; the AI agent held at 0.820.
The standard model's combined score fell from 0.964 to 0.600 on outside hospitals; the AI agent held at 0.825.
The standard model's combined score fell from 0.964 to 0.600 on outside hospitals; the AI agent held at 0.803.
The authors write that people should not use it clinically until prospective studies — following patients forward in real time — are completed.
Acute aortic dissection is a tear in the body's main artery. The article describes a study in which the risk of death rises by 1 to 2 percent for every hour that passes after symptoms begin, and notes that the scan that confirms it — computed tomography angiography, or CTA — is often not available right away. In a busy emergency room, that hour is everything.
Researchers in Wuhan, working with the School of Computer Science at Central China Normal University, built a tool they call AAD-Agent. It is not a doctor and not a scanner. It works in three steps, asking for a little more information each time: first a patient's age, sex, history and symptoms; then routine blood counts; then a blood test called D-dimer. It flags suspected cases and recommends the scan. The study tested the software, not a new way of working in a clinic.
The small study used records from 869 patients at three Chinese hospitals and 200 more from two additional hospitals. On the outside hospitals, the AI agent caught between about 59 and 87 in every 100 true cases, depending on which version was used, while a standard machine-learning model built on the same information fell apart — its combined score for catching cases and avoiding false alarms was 0.964 on the hospitals where it was developed but only 0.600 on the outside hospitals, a collapse the researchers trace to differences between hospitals.
The authors state a problem plainly. They deliberately packed their sample with people who had the disease: about 70 to 78 percent, against under 1 percent of real chest-pain patients in an emergency room. That makes the tool's overall accuracy and its share of correct alarms look better than they would in real life. The real-world numbers are not known.
All the data came from five hospitals in central China. How it would perform in Latin America, the United States or Canada is untested. Inconsistent records and missing data also meant that 59 percent of eligible patients were left out of the study, and those left out were more often women and less often had the disease.
The authors write that people should not use it clinically until prospective studies — following patients forward in real time — are completed. This is not something to ask for at a clinic today.
If those studies confirm it, a fast, cheap rule-out check could one day help staff in clinics without round-the-clock scanners decide who needs transfer or an urgent scan. The next time you hear about an AI triage tool, ask: was it tested on patients like the ones in my emergency room?
What this means for you
For now, nothing about this reaches your own clinic, and no one should ask for it there. What you can watch for is the next study: one that follows real emergency patients forward, in places outside central China, before anyone calls this a working tool. Until then, a professional is the one to ask.
Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503
Who paid: Funded by the National Natural Science Foundation of China (grants 72574075 and 72274067) awarded to L.Y.; the funders had no role in study design, data collection, analysis, interpretation, writing, or the decision to submit, and no equipment or software was lent — the LLMs were used through their public APIs.
Do not take this as professional medical advice.
The findings of other studies mentioned here are known to us through this document, which is the one we read; we did not open each of those studies.
The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1437 words · about 7 minRead it →Close
An AI Triage Aid for a Fast-Moving Emergency Held Up Better Than Older Models Across Hospitals
In a five-hospital Chinese study, a staged AI agent kept its footing where a conventional model collapsed — but its real-world accuracy is still unknown, and the authors say it must not be used yet.

Acute aortic dissection is a tear in the wall of the aorta, the body's main artery. It is a cardiovascular emergency with a high fatality rate and rapid progression1. As a time-sensitive disease, the mortality rate rises 1%–2% per hour after onset2. The confirmatory test is computed tomography angiography, or CTA. Timely diagnosis remains difficult, particularly in primary care and resource-limited settings where CTA is not readily available3. Unnecessary CTA scans — which account for up to 97.3% of CTAs ordered for suspected AAD — carry radiation exposure, contrast-induced nephropathy and excess costs4.
To address this gap, a team led by Li Yan of Tongji Hospital in Wuhan and Yutao Ma of Central China Normal University developed AAD-Agent, an AI agent built on large language models. It acquires data in three stages — demographics and symptoms, then routine blood tests, then D-dimer — updates its risk estimate after each stage, and stops early if it judges the case suspicious5. It was designed as a rapid pre-imaging triage aid, not to replace CTA, but to safely rule out low-risk patients and defer unnecessary imaging, particularly in resource-limited settings or during initial triage to prioritize high-risk patients for confirmatory CTA6. The work was funded by the National Natural Science Foundation of China, with both grants awarded to L.Y.7.
The study compared AAD-Agent against a conventional machine-learning ensemble model. The ML model performed well on internal data but deteriorated on external data. AAD-Agent maintained external stability across several LLMs: DeepSeek-R1 (accuracy 0.715, F1-score 0.820), GPT-3.5 (0.715, 0.825), and GPT-4o (0.710, 0.803)8. Although the case-enriched design inflates these metrics, a prevalence-adjusted negative predictive value of approximately 99.6% at 1% prevalence supports AAD-Agent as a safe, low-cost rule-out aid to defer CTA in low-risk patients9.
The study also compared five AI agents with 18 primary care physicians at community hospitals on a subset of 50 cases. The physicians showed a high miss rate of 19.6% and an extremely elevated misdiagnosis rate of 53.3%. All evaluated LLMs performed substantially better on both measures10. GPT-4o had the most balanced performance, with a false-negative rate of 8.0% and the lowest false-positive rate of 4.0%11. However, the agent's low external specificity precludes confirmatory diagnosis, limiting its use to ruling out cases pending prospective validation12. The authors strongly caution against clinical use until such prospective studies are completed13. Prospective validation in unselected, real-world emergency populations is essential before any clinical application14.
The problem this addresses is not abstract. In China, the clinical burden of AAD is compounded by high mortality, substantial costs, and frequent misdiagnosis or delayed detection, especially in primary care settings with limited diagnostic expertise and resources15. The study drew on 869 patients from three Chinese hospitals for internal validation and 200 from two more for external validation. About 70% of the internal cohort and about 78% of the external cohort had confirmed AAD — a rate far higher than the under-1% seen in real emergency chest-pain populations, a design choice that inflates accuracy and positive predictive value while understating negative predictive value.

The technology works differently from a standard prediction model. A conventional machine-learning system relies on fixed patterns learned from its training data; when a new hospital's patients have different blood pressure readings, different D-dimer levels or different white blood cell counts, its decision boundaries can fail. The AI agent instead processes narrative clinical descriptions — a patient's story, in words — and reasons through them in stages, recalibrating as each new piece of data arrives5. That design appears to make it less brittle when the data shift from one hospital to another, though the study cannot yet prove that in real time.
This is not the first attempt to bring AI into cancer screening or emergency triage. According to the summary of a study of AI-supported double reading in European mammography programs — we could read only the summary, the full paper is behind a subscription — most European population screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers16. AI is being evaluated to support or optimize these established screening pathways17. That evidence synthesis covered three large studies embedded within routine population screening programs: MASAI, ScreenTrustCAD, and PRAIM18. Across 597,419 examinations, the pooled cancer detection rate showed a small absolute increase19. In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection — about 1 in 1,000 — without a consistent increase in recall, alongside improved positive predictive value and efficiency signals20. Those findings suggest AI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution21. The AAD-Agent study is earlier-stage and smaller by comparison — a retrospective, case-enriched design rather than a prospective program-embedded trial.
Here is how we read it. The pattern in studies like this one is that an AI tool performs well where it was developed and then stumbles somewhere else. That is not a failure of the idea; it is a warning about the gap between a research dataset and a real emergency room. What we expect, if a tool like this ever reaches a clinic near you, is that its output will carry more weight than its evidence warrants — that a "low risk" label will feel like a clearance, even though the study says the tool only helps defer a scan and never rules dissection out on its own. You would know we are wrong if emergency teams kept the scan decision with the clinician and used the tool as one input among many. What you can do with this now: if you or a relative ever face a decision about a scan, ask who made the final call and whether the scan decision still rests with the doctor.
A second thing we notice: the people who know least about a rare emergency may be the most likely to trust a confident-sounding tool, while a specialist would question it. That is an expectation, not a finding. You could test it by asking whether the staff using such a tool have had training in when it can be wrong, and whether a senior clinician reviews borderline results. And if you are ever told your symptoms do not fit the usual pattern, ask what else it could be and whether a scan has been considered.

What would have to happen next is clear from the authors themselves: prospective validation in unselected, real-world emergency populations is essential before any clinical application14. The study had no head-to-head comparison with the established ADD-RS plus D-dimer strategy because required variables were inconsistently documented. About 59% of eligible patients were excluded, and those excluded differed in sex ratio and AAD prevalence, indicating selection bias. The physician comparison used only 50 cases and 18 trainees in a simulated setting, not real workflow conditions. False lumen patency data were unavailable, so performance by AAD subtype could not be assessed. And all data came from five hospitals in central China, so performance in Latin America, the United States or Canada is untested.
The trade-off the study names is this: sensitivity over specificity. A tool that flags more people as possibly having AAD will catch more real cases, but it will also send more people for scans they may not need12. In a pre-imaging triage context, the authors argue, that trade-off is acceptable because missing a dissection is far more harmful than performing an unnecessary scan. Whether that reasoning holds in your local emergency room depends on whether the scan is actually available and whether the staff can act on the result in time.
The most useful thing you can carry away is not the number. It is the question. If a tool like this ever enters a clinic where you live, the question to ask is not "What did the machine say?" but "Who decided, and on what evidence?" That question works whether the tool is present or not. It is the same question you would ask about any test, any scan, any recommendation. And it is the one that keeps the decision where it belongs — with a person who can see you, not just your data.
Where each piece of context comes from, and how much of it we read
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "Acute aortic dissection (AAD) is a life-threatening cardiovascular emergency, posing a significant challenge to doctors worldwide due to its high fatality rate and rapid progression."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "As a time-sensitive disease, the mortality rate rises 1%–2% per h post-onset."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "Despite recent advances, timely diagnosis remains difficult, particularly in primary care and resource-limited settings where computed tomography angiography (CTA), the confirmatory standard, is not readily available."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "Moreover, unnecessary CTA scans, which account for up to 97.3% of CTAs ordered for suspected AAD, incur radiation exposure, contrast-induced nephropathy, and excess costs."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "To address this gap, we developed AAD-Agent, an LLM-enhanced autonomous AI agent that emulates hierarchical emergency reasoning. It actively acquires data in three stages (demographics/symptoms→routine blood tests→D-dimer), dynamically updates risk estimates, and provides interpretable outputs."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "Designed as a rapid pre-imaging triage aid, AAD-Agent aims not to replace CTA but to safely rule out low-risk patients and defer unnecessary imaging, particularly in resource-limited settings or during initial triage to prioritize high-risk patients for confirmatory CTA."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "The work was funded by the 10.13039/501100001809 National Natural Science Foundation of China (grant nos. 72574075 and 72274067 ), both awarded to L.Y."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "In contrast, AAD-Agent maintained external stability across LLMs: DeepSeek-R1 (accuracy 0.715, F1-score 0.820), GPT-3.5 (0.715, 0.825), and GPT-4o (0.710, 0.803)."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "Although the case-enriched design inflates these metrics, prevalence-adjusted negative predictive value (≈99.6% at 1% prevalence) supports AAD-Agent as a safe, low-cost rule-out aid to defer CTA in low-risk patients."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "As shown in Figure 6 , primary care physicians exhibited a high miss rate of 19.6% and an extremely elevated misdiagnosis rate of 53.3%. In stark contrast, all evaluated LLMs demonstrated substantially superior performance across both metrics."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "GPT-4o presented the most balanced performance, with a false-negative rate of 8.0% (sensitivity = 92.0%) and the lowest false-positive rate of 4.0%."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "Its low external specificity precludes confirmatory diagnosis, limiting use to rule out pending prospective validation."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "We strongly caution against clinical use until such prospective studies are completed."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "Prospective validation in unselected, real-world emergency populations is essential before any clinical application."
- Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503 - the article this story is about — the whole article — the passage: "In China, the clinical burden of AAD is compounded by high mortality, substantial costs, and the frequent misdiagnosis or delayed detection, especially in primary care settings with limited diagnostic expertise and resources."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "Most European population mammography screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "Artificial intelligence (AI) is being evaluated to support or optimize these established European screening pathways."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "We performed a prespecified, focused evidence synthesis of three large studies embedded within routine population screening programs operating under European-relevant workflows: MASAI (randomized AI-supported risk triage within a national program), ScreenTrustCAD (prospective paired-reader evaluation with AI as an independent reader in a double-reading framework), and PRAIM (nationwide decision-referral implementation)."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I"
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution."
Li, H., Pei, Y., Hu, T. et al. (2026). An LLM-enhanced AI agent for rapid diagnosis of acute aortic dissection in multi-center settings. iScience. https://doi.org/10.1016/j.isci.2026.117503
Who paid: Funded by the National Natural Science Foundation of China (grants 72574075 and 72274067) awarded to L.Y.; the funders had no role in study design, data collection, analysis, interpretation, writing, or the decision to submit, and no equipment or software was lent — the LLMs were used through their public APIs.
Do not take this as professional medical advice.
