weeklyAI · Week of 27 September 2026weeklyAI · Semana del 27 de septiembre de 2026

← Your rights← Sus derechos

analysis of texts · Nature · la publicación, 24 jun 2026 · gratis

La IA médica puede delatar a algunos pacientes, no a todos

Un estudio con siete bases de datos clínicas muestra que ciertos pacientes son identificables casi con certeza, mientras el promedio oculta ese riesgo.

Versión breve · la versión detallada sigue, unos 7 min

Pregunte a weeklyAI

Pregúnteme por este estudio: a quiénes se estudió, qué encontró y qué no dice.

Las conversaciones se guardan mientras exista weeklyAI, para mejorar la publicación. Se responde en el idioma en que usted escribe.

El estudio, de un vistazo
Quiénes
Pacientes cuyos datos médicos se usaron para entrenar modelos de IA
Cuántos
siete grandes conjuntos de datos clínicos reales
Dónde
Hospitales y poblaciones concretas (CheXpert, MIMIC-CXR, Fitzpatrick 17k, FairVision, EMBED, PTB-XL y MIMIC-IV-ED)
Cuándo
2026
Tipo de estudio
experiment
Quién lo hizo
Universidad Técnica de Múnich, Imperial College de Londres y otras instituciones
El límite que importa
Probó modelos en investigación, no sistemas ya instalados en hospitales; no mide daños clínicos

Frecuencia de registros de ciertos grupos entre los más vulnerables frente a lo esperado (MIMIC-IV-ED)

Pacientes negrosfrente aLo esperado según el conjunto general

31% más de lo esperado

Pacientes con seguro Medicaidfrente aLo esperado según el conjunto general

126% más de lo esperado

Pacientes diagnosticados con cáncerfrente aLo esperado según el conjunto general

18% más de lo esperado

Un promedio bajo puede dar una falsa tranquilidad, porque esconde que algunos pacientes quedan muy expuestos.
Lectura de weeklyAI
Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

Cuando usted entrega una radiografía, un electrocardiograma o su historial a un hospital, esos datos pueden terminar alimentando un programa de inteligencia artificial. Muchas veces usted no sabe que eso ocurre ni para qué se usan después.

Un equipo de investigadores de la Universidad Técnica de Múnich, el Imperial College de Londres y otras instituciones revisó siete grandes conjuntos de datos clínicos reales. Entrenaron cientos de programas, unos doscientos por cada conjunto, para probar un tipo de ataque llamado "inferencia de pertenencia". La idea es simple: dado un programa ya entrenado, alguien intenta averiguar si los datos de una persona determinada se usaron para entrenarlo.

El resultado: para un grupo pequeño de pacientes, el ataque acertaba casi siempre. Para la mayoría, no. Y cuando los investigadores miraron el promedio de todos los registros, el ataque parecía apenas mejor que adivinar al azar.

Un promedio bajo puede dar una falsa tranquilidad, porque esconde que algunos pacientes quedan muy expuestos. En un conjunto de dermatología, al aumentar el tamaño del programa, la proporción de pacientes con aciertos casi perfectos pasó de cero a uno de cada diez mil, luego a uno de cada mil y después a uno de cada diez.

También aparecieron diferencias entre grupos. En una base de datos de urgencias, los registros de pacientes negros, de pacientes con seguro Medicaid y de pacientes con diagnóstico de cáncer aparecían más de lo esperado entre los más vulnerables: un 31%, un 126% y un 18% más, respectivamente.

El estudio probó programas de diagnóstico en un entorno de investigación, no sistemas ya instalados en hospitales. No demuestra que la IA de su hospital haya sido atacada ni que sus datos hayan quedado expuestos. Los datos vienen de hospitales y poblaciones concretas. Tampoco midió daños clínicos, solo qué tan bien funcionaba el ataque.

Los autores recomiendan que las revisiones de privacidad informen el riesgo paciente por paciente, no solo en promedio, y que los hospitales consideren protecciones verificables, como la privacidad diferencial, antes de usar estas herramientas. Señalan además que si el riesgo se extiende a otros ataques distintos sigue siendo una pregunta abierta.

Usted puede preguntar a su hospital o a su autoridad de salud si, cuando usan IA con datos de pacientes, miden y publican ese riesgo individual.

Qué significa para usted

Por ahora no hay una norma general que obligue a su hospital a medir este riesgo, así que lo útil es preguntar si lo evalúan y si publican esos resultados. El estudio solo describe siete bases de datos concretas y no demuestra que su atención haya sido vulnerada; ante cualquier duda sobre sus datos clínicos, consulte a un profesional.

Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0

Quién pagó: El estudio fue financiado parcialmente por la Konrad Zuse School of Excellence in Reliable AI (relAI), la subvención ERC Deep4MI (n.º 884622) y la Fundación Alemana de Investigación (proyecto n.º 532139938); B.G. recibió apoyo de la Royal Academy of Engineering, y los recursos computacionales fueron proporcionados por el Leibniz Supercomputing Centre; el artículo no indica si los financiadores tuvieron algún papel en el estudio.

No tome esto como consejo médico profesional.

Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1463 palabras · unos 7 minLeerla →Cerrar

Un modelo médico puede delatar a un paciente: quién queda más expuesto

Un estudio entrenó cientos de versiones de sistemas de diagnóstico para medir, paciente por paciente, qué tan fácil es saber si sus datos fueron usados. El promedio decía que casi no había riesgo. Algunos casos decían lo contrario.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

Los sistemas de inteligencia artificial (IA) para diagnóstico médico se entrenan con datos de pacientes reales, y esos datos pueden filtrarse a través de ataques de privacidad1. El estudio que leemos se concentra en uno de esos ataques: el de inferencia de pertenencia, que busca determinar si los datos de una persona en particular se usaron para entrenar un modelo2. Así lo leemos nosotros: no se trata de adivinar un diagnóstico a partir del nombre, sino de algo más silencioso y más exacto. Muchos de estos sistemas se ofrecen como un servicio que recibe una imagen o un registro y devuelve una probabilidad, por ejemplo, 78% de probabilidad de neumonía3. El ataque se apoya en un detalle: los modelos suelen mostrarse un poco más seguros cuando predicen sobre datos que ya vieron durante su entrenamiento que sobre datos nuevos4.

Los autores, encabezados por Moritz A. Knolle, auditaron siete conjuntos de datos clínicos reales: CheXpert, MIMIC-CXR, Fitzpatrick 17k, FairVision, EMBED, PTB-XL y MIMIC-IV-ED. Para cada conjunto entrenaron 200 modelos sobre subconjuntos aleatorios de pacientes, de modo que cada registro quedara incluido en unos modelos y excluido de otros. Con eso midieron el éxito del ataque registro por registro, y luego, para cada paciente, tomaron su caso más expuesto. El hallazgo central es incómodo: en esos conjuntos, el ataque puede alcanzar tasas de éxito casi perfectas para pacientes individuales, incluso cuando el desempeño agregado del ataque no se aparta mucho de una simple adivinanza5. Dicho de otro modo, el promedio puede verse inofensivo mientras un puñado de personas queda prácticamente al descubierto.

El estudio también encontró que el número de pacientes con éxito de ataque alto crece de forma sustancial con la capacidad del modelo6. Y encontró que los grupos subrepresentados —por estado de enfermedad, raza autodeclarada, seguro de salud, sexo o protocolo de imagen— enfrentan un éxito de ataque desproporcionadamente alto6. En el conjunto de urgencias MIMIC-IV-ED, los registros de pacientes negros, de pacientes con seguro Medicaid o de pacientes diagnosticados con cáncer aparecieron con más frecuencia de la esperada entre los más vulnerables: 31%, 126% y 18% por encima de lo que correspondería al conjunto general, respectivamente7. En términos llanos, los grupos que ya están poco representados en los datos de entrenamiento suelen estar sobrerrepresentados entre los registros más fáciles de identificar8.

Por qué importa tanto la pertenencia: si un modelo se entrenó con una población estrecha, saber que alguien estuvo en ese entrenamiento equivale a saber algo médico sobre esa persona. Un ataque exitoso contra un modelo que predice la eficacia de una inmunoterapia contra el cáncer a partir de análisis de sangre revela que esa persona tiene cáncer9. Los propios autores advierten que las métricas de privacidad agregadas pueden subestimar gravemente el riesgo individual10. Y dejan abierto un punto: si esos perfiles desiguales de riesgo se extienden a ataques distintos del de pertenencia es una pregunta sin responder11.

Hay límites que conviene retener. El estudio probó modelos de diagnóstico en un entorno de investigación, no sistemas hospitalarios ya desplegados; no muestra que el sistema de un hospital concreto haya sido atacado ni que los datos de alguien hayan quedado expuestos. Los conjuntos provienen de hospitales y poblaciones específicas. La estimación del riesgo por registro exigió 200 modelos por conjunto, y los autores señalan que su margen de error es pequeño en todos los registros estudiados. Las comparaciones por subgrupos usaron un umbral del percentil 99 y pruebas estadísticas con corrección; algunas comparaciones no pudieron evaluarse porque las categorías no eran mutuamente excluyentes, y los autores señalan que las comparaciones por raza en el conjunto de mamografías probablemente estén confundidas por diferencias en la densidad mamaria. Sobre la mitigación: la protección formal de privacidad redujo el éxito del ataque, pero para eliminarlo en todos los pacientes haría falta contabilizarla por paciente y no por registro, y solo con protección fuerte (ε = 1) se observaron excepciones en un subconjunto de pacientes.

El terreno donde esto ocurre no está ordenado. Según el resumen de un estudio sobre gobernanza de la IA en salud en la región europea de la OMS, solo 8% de los Estados que respondieron una encuesta tenía una estrategia de IA específica para salud y 14% estaba desarrollándola12; apenas 14 habían emitido guías sobre las implicaciones éticas de usar IA en salud13, y menos de 10% había desarrollado estándares de responsabilidad14. De ese trabajo pudimos leer solo el resumen. La conclusión de sus autores es que la gobernanza de la IA en salud está poco desarrollada15 y que hacen falta mecanismos legales y de política adaptables16. Traducido a la vida diaria: las reglas que deberían responder a un hallazgo como este todavía se están escribiendo.

La IA se usa cada vez más para apoyar decisiones, planificar operaciones, guiar procedimientos y cumplir funciones autónomas, y eso introduce riesgos técnicos, humanos, legales y éticos17. Sus autores sostienen que los marcos regulatorios y legales actuales no están del todo preparados para estos desafíos18 y que la responsabilidad puede recaer tanto en los desarrolladores por diseño defectuoso o falta de advertencia como en las instituciones por implementación u oversight negligentes19. De ese trabajo también pudimos leer solo el resumen.

Según el resumen de un experimento con 768 estudiantes universitarios, dos mensajes breves en la interfaz —una garantía de privacidad y una advertencia sobre los límites profesionales— cambiaron la percepción de protección y la conciencia de esos límites20, y la confianza mejor calibrada apareció cuando ambos mensajes estaban presentes21. Sus autores concluyen que esos mensajes pueden ayudar a tratar a los asistentes como herramientas limitadas y a reconocer cuándo hace falta ayuda profesional22, aunque advierten que no deben entenderse como una invitación a usarlos más22. De ese estudio pudimos leer solo el resumen.

Aquí es donde entra nuestra lectura. El patrón que vemos es viejo y conocido: quienes diseñan y operan un sistema técnico suelen estar lejos de los lugares donde ese sistema se aplica. Desde esa distancia, la forma habitual de medir el riesgo promedia a todos y no distingue a nadie en particular. Por eso cabe esperar que los equipos que entrenan y despliegan modelos médicos no tengan manera de notar, por sí mismos, que ciertos pacientes quedan mucho más expuestos que otros. También cabe esperar que las instituciones que se presentan como abiertas a todos no repartan por igual lo que ofrecen, y que las personas ya subrepresentadas —por enfermedad poco frecuente, por raza, por tipo de seguro, por sexo o por tipo de estudio— carguen con la mayor parte del riesgo sin que eso aparezca en las métricas que se publican. Lo que nos haría pensar que nos equivocamos: que los informes de privacidad de esos modelos ya reportaran el riesgo paciente por paciente y grupo por grupo, con los grupos pequeños pesando igual que los grandes, y que los equipos consultaran de manera sistemática a las comunidades afectadas antes de desplegar.

Nuestra expectativa es que la mayoría de los pacientes no sepa que un modelo entrenado con sus estudios puede delatar si sus datos fueron usados, ni que exista un derecho a preguntarlo o a oponerse; y que quienes lo saben no encuentren un canal claro para ejercerlo. Lo que nos haría pensar que nos equivocamos: que los hospitales y las empresas informaran de manera comprensible, antes o al momento de usar los datos, que un modelo puede revelar la pertenencia de un paciente al conjunto de entrenamiento, y que existiera un procedimiento sencillo y gratuito para consultar, oponerse o reclamar. Esto es lo que usted puede hacer con todo esto: cuando un hospital anuncie que usa IA sobre imágenes o historias clínicas, pregunte por escrito qué modelos se usan sobre sus datos, con qué finalidad y a quién se comunican, y si el informe de privacidad distingue por paciente y por grupo en lugar de mostrar solo un promedio general. Pida también saber quién revisó ese informe y con qué criterio. Guarde copia de su solicitud y de la respuesta: sirve como punto de partida si más adelante quiere reclamar ante una autoridad de protección de datos. Si usted pertenece a un grupo poco frecuente en los datos de un hospital o de un estudio, pida de manera concreta cómo se protege su caso y exija que la respuesta no sea un promedio. Y cuando le digan que un modelo es más potente o más preciso, pregunte qué pasó con la privacidad de los pacientes que lo hicieron posible, porque el estudio sugiere que esas dos cosas crecen juntas.

De dónde sale cada dato de contexto, y cuánto leímos de cada documento

  1. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Medical artificial intelligence (AI) models hold the promise to improve global access to high-quality diagnostics 1 . However, the training data underlying these models often contain sensitive patient information that may be exposed through privacy attacks 2 – 7 ."
  2. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "We focus on membership inference attacks 2 – 4 (MIAs), which seek to determine whether the data of a given individual were used to train a model."
  3. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "A popular deployment strategy for AI models gives users access to a model through a prediction interface, which, for a given input (for example, the chest radiograph of a patient), returns a corresponding prediction (for example, a 78% chance of pneumonia)."
  4. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "To infer membership status, MIAs typically make use of the fact that AI models are often slightly more confident about their predictions on training than on non-training data."
  5. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Across a diverse range of medical datasets, we show that MIAs can achieve near-perfect success rates for individual patients, even when the aggregate performance does not substantially deviate from random guessing."
  6. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "We further find that the number of patients with high attack success increases substantially with model capacity, and that underrepresented groups—stratified by disease status, self-reported race, insurance, sex or imaging protocol—face disproportionately high attack success."
  7. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For example, in MIMIC-IV-ED, records from Black patients, patients with Medicaid insurance or patients diagnosed with cancer were observed more frequently than expected among the most vulnerable records (+31%, +126%, and +18% relative change to the overall dataset, respectively)."
  8. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Groups of patients that are underrepresented in a model training dataset are often overrepresented among the records most susceptible to MIAs. By contrast, the opposite often holds for majority groups."
  9. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For example, a successful MIA against the model in ref. 10 , which predicts anti-cancer immunotherapy efficacy from routine blood test data, reveals that an individual has cancer."
  10. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Together, our findings show that aggregate privacy metrics can severely underestimate individual privacy risk."
  11. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Whether the disparate risk profiles we observe extend to attacks beyond MIAs remains an open question, motivating the further development of risk assessment and mitigation techniques that cater to all data-contributing patients."
  12. Adib K, Letchford N, Dunning HE, Salama N, Tolias Y, De Barros J, et al. (2026). Governance of artificial intelligence for health systems, WHO European Region. Bulletin of the World Health Organization. 10.2471/blt.25.294978 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Of the 50 Member States responding to the survey, 8% (4/50) have a health-specific AI strategy and 14% (7) are developing one."
  13. Adib K, Letchford N, Dunning HE, Salama N, Tolias Y, De Barros J, et al. (2026). Governance of artificial intelligence for health systems, WHO European Region. Bulletin of the World Health Organization. 10.2471/blt.25.294978 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Only 14 Member States have issued guidelines to address the ethical implications of using AI in health or across sectors."
  14. Adib K, Letchford N, Dunning HE, Salama N, Tolias Y, De Barros J, et al. (2026). Governance of artificial intelligence for health systems, WHO European Region. Bulletin of the World Health Organization. 10.2471/blt.25.294978 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Less than 10% (4) of Member States have developed liability standards for AI or guidance on the application of existing liability standards."
  15. Adib K, Letchford N, Dunning HE, Salama N, Tolias Y, De Barros J, et al. (2026). Governance of artificial intelligence for health systems, WHO European Region. Bulletin of the World Health Organization. 10.2471/blt.25.294978 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "In the WHO European Region, governance of AI in health care is underdeveloped."
  16. Adib K, Letchford N, Dunning HE, Salama N, Tolias Y, De Barros J, et al. (2026). Governance of artificial intelligence for health systems, WHO European Region. Bulletin of the World Health Organization. 10.2471/blt.25.294978 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Adaptive legal and policy mechanisms are needed to respond effectively to the complex and evolving challenges of AI integration in health systems."
  17. Hashimoto DA, Marwaha JS, Lee SA, Schwaitzberg S, Duffourc MN. (2026). Risk and liability in the deployment of AI systems for surgery: a SAGES white paper. Surgical Endoscopy. 10.1007/s00464-026-12881-8 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Artificial intelligence (AI) is increasingly utilized in surgical care for decision support, operative planning, intraoperative guidance, and autonomous functions. While these systems can enhance efficiency and clinical performance, they also introduce risks related to technology, human factors, legal issues, and ethics."
  18. Hashimoto DA, Marwaha JS, Lee SA, Schwaitzberg S, Duffourc MN. (2026). Risk and liability in the deployment of AI systems for surgery: a SAGES white paper. Surgical Endoscopy. 10.1007/s00464-026-12881-8 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Current regulatory and legal frameworks are not fully equipped to address the challenges of AI-assisted surgery."
  19. Hashimoto DA, Marwaha JS, Lee SA, Schwaitzberg S, Duffourc MN. (2026). Risk and liability in the deployment of AI systems for surgery: a SAGES white paper. Surgical Endoscopy. 10.1007/s00464-026-12881-8 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Although surgeons remain the ultimate clinical decision-makers, liability may also extend to developers for defective design or failure to warn, and to institutions for negligent implementation or oversight."
  20. Zhang Z, Lu X, Zhang Y, Zhang H, Zhang M. (2026). Privacy assurances and professional-boundary warnings in generative AI mental health chatbots: a randomized vignette experiment on calibrated trust, overreliance risk, and professional help-seeking intentions. Frontiers in Psychology. 10.3389/fpsyg.2026.1934264 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Privacy assurance increased perceived privacy protection, F (1, 764) = 159.30, p 2 = 0.172, d = 0.91. Professional-boundary warning increased boundary awareness, F (1, 764) = 176.40, p 2 = 0.188, d = 0.96."
  21. Zhang Z, Lu X, Zhang Y, Zhang H, Zhang M. (2026). Privacy assurances and professional-boundary warnings in generative AI mental health chatbots: a randomized vignette experiment on calibrated trust, overreliance risk, and professional help-seeking intentions. Frontiers in Psychology. 10.3389/fpsyg.2026.1934264 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Calibrated trust was highest when both messages were present."
  22. Zhang Z, Lu X, Zhang Y, Zhang H, Zhang M. (2026). Privacy assurances and professional-boundary warnings in generative AI mental health chatbots: a randomized vignette experiment on calibrated trust, overreliance risk, and professional help-seeking intentions. Frontiers in Psychology. 10.3389/fpsyg.2026.1934264 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Such messages should not be understood as prompts for greater use. Rather, they may help users treat chatbots as limited tools for information and support navigation and recognize when professional help is needed."

Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0

Quién pagó: El estudio fue financiado parcialmente por la Konrad Zuse School of Excellence in Reliable AI (relAI), la subvención ERC Deep4MI (n.º 884622) y la Fundación Alemana de Investigación (proyecto n.º 532139938); B.G. recibió apoyo de la Royal Academy of Engineering, y los recursos computacionales fueron proporcionados por el Leibniz Supercomputing Centre; el artículo no indica si los financiadores tuvieron algún papel en el estudio.

No tome esto como consejo médico profesional.

analysis of texts · Nature · the paper, 24 Jun 2026 · free

Your Scan Trained the AI. A Study Says It Could Also Point Back to You.

A research audit found that for a small number of patients, a model's answers were enough to reveal whether their data had been used to train it — even when the average result looked like a coin toss.

Short version · the longer version follows, about 5 min

Ask weeklyAI

Ask me about this study: who was studied, what it found, and what it does not say.

Conversations are saved for as long as weeklyAI exists, to improve the publication. Answers come in the language you write in.

The study at a glance
Who
Patients in seven large medical datasets (chest X-rays, skin images, eye scans, mammograms, heart traces, emergency-department records)
How many
seven large datasets; 200 target models per dataset
Where
Stanford, Beth Israel Deaconess, Emory University and other institutions
When
not stated in the passages
Kind of study
experiment
Who did it
Technical University of Munich, Imperial College London and other institutions
The limit that matters
Only membership inference attacks were tested; whether the pattern holds for other attacks is an open question.
How much more often some groups appeared among the most attackable records (one emergency-department dataset)
Black patients31%
Patients on Medicaid126%
Patients with a cancer diagnosis18%

Relative change in how often each group appeared among the most vulnerable records compared with its share of the overall dataset, in MIMIC-IV-ED; these are relative changes, not shares of patients.

The average number looked harmless; the individual picture did not.
weeklyAI's reading
How it could look · illustration generated by weeklyAI.watch, not a photograph

People hand over medical images, heart tracings and health records every day. Those files often become training material for AI.

A group of researchers wanted to know what that means for the person inside the file.

They audited seven large datasets of real-world clinical data, including medical images, heart tracings and health records, involving tens of thousands of patients. They trained hundreds of AI models — roughly 200 for each dataset — to copy a standard diagnostic task, such as reading a chest X-ray.

Then they ran an attack that needs nothing more than the model's own predictions. It asks a simple question: was this person's data part of the training set?

For most patients, the answer came out near a coin toss. For a small subset, it was close to certain. The average number looked harmless; the individual picture did not.

The study, from the Technical University of Munich, Imperial College London and other institutions, tested diagnostic AI in a research setting — not any hospital's live system. It does not show that your data have been exposed. It measured how often an attack succeeded, not harm to any patient.

Two more things. The findings apply only to these seven datasets, not to every hospital or country. And the study looked only at this one type of attack; the authors say whether the same uneven pattern holds for other attacks is an open question.

The vulnerable group grew as models got bigger. On one skin-image dataset, patients facing near-certain attack success went from zero, to about one in 10,000, to about one in 1,000, to about one in 10 as the model changed.

Underrepresented groups also appeared more often among the most exposed — by race, insurance, sex or disease. In one emergency-department dataset, records from Black patients, patients on Medicaid and patients with cancer were all over-represented.

The authors want privacy checked patient by patient, not averaged away, and want hospitals to consider verifiable protections such as differential privacy before deploying these tools.

So the question worth asking your hospital or health authority is a plain one: when an AI tool is used on patients like me, is anyone measuring the risk to each person — or only the average?

What this means for you

For now, nothing changes about your own records, and the study does not say your data were exposed or that anyone was harmed. What you can watch for is whether the hospitals and health authorities you deal with start asking about privacy per patient rather than on average, and whether the rules they cite cover tools like these. If a medical AI question touches your own care, a professional is the one to ask.

Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0

Who paid: The study was partially funded by the Konrad Zuse School of Excellence in Reliable AI (relAI), ERC Grant Deep4MI (grant no. 884622) and the German Research Foundation (project no. 532139938); B.G. received support from the Royal Academy of Engineering, and computational resources were provided by the Leibniz Supercomputing Centre; the article does not state whether funders had any role in the study.

Do not take this as professional medical advice.

The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1009 words · about 5 minRead it →Close

Medical AI Can Give Away Who Was In the Training Data — Some Patients Almost Every Time

A patient-level audit of seven real-world clinical datasets finds that a privacy attack can identify some individuals' records almost perfectly, even when the standard scorecard says the model is barely leaking at all.

How it could look · illustration generated by weeklyAI.watch, not a photograph

Give a medical AI model a chest X-ray, and it will answer with something like "78% chance of pneumonia"1. That single answer is all a stranger needs to start guessing whether your X-ray was part of the data the model was trained on1. The article describes a study in which this is called a membership inference attack, and it works because models tend to be a little more confident about the records they learned from than the ones they never saw23.

The new study, by Moritz A. Knolle and Georgios Kaissis with colleagues at the Technical University of Munich, Imperial College London and other institutions, audited diagnostic AI models across seven large datasets of real clinical data — chest X-rays from Stanford and Beth Israel Deaconess, skin images, eye scans, mammograms from Emory University, heart traces and emergency-department records4. For each dataset they trained 200 target models on random patient subsets, which let them estimate attack success for every single record and every single patient, not just on average.

What they found: the attack can work almost perfectly on some individual patients, even when its average performance across the whole dataset is barely better than a coin toss5. In other words, the number hospitals and researchers normally report — the aggregate score — can hide the fact that a handful of people are sitting completely exposed6.

And the exposure is not spread evenly. People in groups that are a small share of the training data — by disease, self-reported race, insurance, sex or imaging protocol — turn up far more often among the most vulnerable records78. In one emergency-department dataset, records from Black patients, patients on Medicaid, or patients with a cancer diagnosis appeared among the most attackable records 31%, 126% and 18% more often than their share of the data would predict9.

The authors also found that bigger, more capable models make this worse: the number of patients with high attack success increases substantially with model capacity7. They tested a privacy protection called differential privacy and found it reduced attack success, but noted that some patients still exceeded the promised bound at strong protection levels, and that this could be mitigated by patient-level rather than record-level accounting.

The study has clear limits. Whether the uneven pattern it found extends to attacks beyond membership inference is an open question the authors name themselves10. The findings cover membership inference only; whether the same uneven pattern holds for other kinds of privacy attacks is an open question the authors name themselves10. The attack success estimates relied on 200 models per dataset. The authors estimate that reproducing all of their experimental data would require around 900 hours on a single A100 GPU.

This matters beyond the lab because membership itself can be the secret. The article describes a study in which a successful attack against a model that predicts anti-cancer immunotherapy response from routine blood tests reveals that a person has cancer11. The article describes a study in which medical data are a prime target for cybercriminals, and stripping names from a dataset is increasingly understood to be insufficient to prevent re-identification in large, detailed datasets12.

The pattern the authors describe — underrepresented groups carrying the heaviest exposure — fits what they say could become a vicious cycle: if minority groups see worse model performance and higher privacy risk, they may trust these systems less and contribute less data, which makes the next model worse for them again13. The authors write that current trends in medical AI development and deployment could widen existing health inequalities13.

Their recommendation is concrete: report attack success at the level of individual patients or records, not just averages14, and protect vulnerable models with verifiable techniques such as differential privacy or strict access controls1516. Differential privacy works by carefully perturbing parameter updates with white noise during training or fine-tuning, which limits the contribution of any individual's data to the final model.

Here is how we read it. The pattern is familiar from other corners of life: when a system is built around the majority, the people at the edges stand out — and standing out is exactly what makes you findable. We expect that as AI tools spread through clinics in our region, the patients most likely to be singled out by a model will often be the same patients already watched most closely at the front desk. You could show we are wrong if the patients a model can identify turn out to be spread evenly across every group, or if the most exposed turn out to be the largest and best-represented ones.

We also read this against the way hospitals adopt technology. Tools arrive as the normal thing to do, and nobody remembers choosing them. If you ask your clinic who signed off on an AI tool, what it is used for, and whether you can be seen without it, and nobody can name a decision-maker, that itself is your answer. When you are told your data is anonymised, ask what test was run to show that, and whether a person could still be picked out from the model's own answers. Under many data protection laws you can ask what is held about you and why; use that request rather than accepting the word "anonymous."

What the authors want next is straightforward: audits that report risk per patient, and privacy protections that are mathematically checkable rather than promises1514. Until then, the scorecard most institutions rely on may be telling them — and you — that everything is fine when, for some patients, it is not6.

The next time you hear that a hospital or clinic is using AI on scans or records, ask one question: who checked which patients the model can be made to reveal?

Where each piece of context comes from, and how much of it we read

  1. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "A popular deployment strategy for AI models gives users access to a model through a prediction interface, which, for a given input (for example, the chest radiograph of a patient), returns a corresponding prediction (for example, a 78% chance of pneumonia)."
  2. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "We focus on membership inference attacks 2 – 4 (MIAs), which seek to determine whether the data of a given individual were used to train a model."
  3. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "To infer membership status, MIAs typically make use of the fact that AI models are often slightly more confident about their predictions on training than on non-training data."
  4. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Medical artificial intelligence (AI) models hold the promise to improve global access to high-quality diagnostics 1 . However, the training data underlying these models often contain sensitive patient information that may be exposed through privacy attacks 2 – 7 ."
  5. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Across a diverse range of medical datasets, we show that MIAs can achieve near-perfect success rates for individual patients, even when the aggregate performance does not substantially deviate from random guessing."
  6. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Together, our findings show that aggregate privacy metrics can severely underestimate individual privacy risk."
  7. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "We further find that the number of patients with high attack success increases substantially with model capacity, and that underrepresented groups—stratified by disease status, self-reported race, insurance, sex or imaging protocol—face disproportionately high attack success."
  8. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Groups of patients that are underrepresented in a model training dataset are often overrepresented among the records most susceptible to MIAs. By contrast, the opposite often holds for majority groups."
  9. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "For example, in MIMIC-IV-ED, records from Black patients, patients with Medicaid insurance or patients diagnosed with cancer were observed more frequently than expected among the most vulnerable records (+31%, +126%, and +18% relative change to the overall dataset, respectively)."
  10. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Whether the disparate risk profiles we observe extend to attacks beyond MIAs remains an open question, motivating the further development of risk assessment and mitigation techniques that cater to all data-contributing patients."
  11. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "For example, a successful MIA against the model in ref. 10 , which predicts anti-cancer immunotherapy efficacy from routine blood test data, reveals that an individual has cancer."
  12. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Given that medical data are a key target for cybercriminals 12 , 13 , and pseudonymization alone is increasingly recognized as insufficient to prevent the re-identification of individuals in large, high-dimensional datasets 14 – 16 , there is a need to improve our understanding of the threat that AI privacy attacks pose to individual patients."
  13. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Our findings suggest that current trends in medical AI development and deployment could exacerbate these health inequalities."
  14. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Audits should report the success of privacy attacks at the level of individual data contributors or, if the necessary patient- or person-level identifiers are unavailable, at the record level."
  15. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "This suggests that current AI privacy risk reporting practices may underestimate individual-level risk and thus motivates the integration of mathematically verifiable risk mitigation strategies such as differential privacy (DP) into medical AI model development workflows."
  16. Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "To prevent privacy harm, we recommend that vulnerable models be protected by verifiable risk mitigation strategies and/or strict access controls."

Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0

Who paid: The study was partially funded by the Konrad Zuse School of Excellence in Reliable AI (relAI), ERC Grant Deep4MI (grant no. 884622) and the German Research Foundation (project no. 532139938); B.G. received support from the Royal Academy of Engineering, and computational resources were provided by the Leibniz Supercomputing Centre; the article does not state whether funders had any role in the study.

Do not take this as professional medical advice.

The findings of other studies mentioned here are known to us through this document, which is the one we read; we did not open each of those studies.