analysis of texts · Nature · la publicación, 24 jun 2026 · gratis
La IA médica delató a algunos pacientes con casi total certeza
Un estudio con siete bases de datos clínicas midió ataques a modelos de diagnóstico. No evaluó leyes, ni hospitales, ni países.
Versión breve · la versión detallada sigue, unos 8 min
Pregunte al chat de weeklyAI
Pregúnteme por este estudio: a quiénes se estudió, qué encontró y qué no dice.
Las conversaciones se guardan mientras exista weeklyAI, para mejorar la publicación. Se responde en el idioma en que usted escribe.
- El estudio, de un vistazo
- Quiénes
- Modelos de IA para diagnóstico médico y los pacientes cuyos datos los entrenaron
- Cuántos
- siete conjuntos de datos clínicos reales; 200 modelos por conjunto
- Dónde
- Estados Unidos (Stanford, Beth Israel Deaconess, Emory, Harvard) y Alemania (PTB-XL)
- Cuándo
- Publicado el 24 de junio de 2026
- Tipo de estudio
- experiment
- Quién lo hizo
- Universidad Técnica de Múnich, Imperial College de Londres y Universidad de Potsdam
- El límite que importa
- Es una auditoría de laboratorio: no muestra filtraciones reales ni evalúa leyes de ningún país.
Cambio relativo respecto del conjunto de datos completo en la frecuencia con que estos grupos aparecen entre los registros más vulnerables al ataque; son tres grupos distintos, no partes de un mismo total.
Riesgo de ataque: promedio del conjunto frente a pacientes individuales
El promedio no se aparta sustancialmente del azar, pero contra algunos pacientes el ataque alcanza tasas de éxito casi perfectas.
Los grupos que aportan menos registros suelen estar sobrerrepresentados entre los más susceptibles al ataque; con los mayoritarios suele ocurrir lo contrario.
En promedio, el ataque casi no superó el azar. Pero en algunos pacientes fue casi perfecto: el programa acertaba casi siempre.

Los modelos de inteligencia artificial para diagnóstico médico se entrenan con datos de pacientes, como radiografías, electrocardiogramas y fichas clínicas. Sólo mide un riesgo técnico.
Investigadores de la Universidad Técnica de Múnich y de instituciones asociadas, incluido el Imperial College de Londres, auditaron modelos de inteligencia artificial para diagnóstico médico. Usaron siete bases de datos clínicas reales: radiografías de tórax de Stanford y del Beth Israel Deaconess Medical Center, mamografías de Emory University, imágenes de fondo de ojo de Harvard, electrocardiogramas, fichas de urgencias y un conjunto de imágenes de dermatología.
Para cada base entrenaron 200 modelos sobre grupos al azar de pacientes. Luego probaron si un programa podía adivinar, mirando sólo las respuestas del modelo, si los datos de una persona habían servido para entrenarlo.
En promedio, el ataque casi no superó el azar. Pero en algunos pacientes fue casi perfecto: el programa acertaba casi siempre. En un conjunto de dermatología, la proporción de pacientes con acierto casi perfecto pasó de ninguno, con el modelo más chico, a uno de cada diez, con el más grande. También observaron que los grupos menos representados en los datos aparecían más de lo esperado entre los pacientes más expuestos. En urgencias, los registros de pacientes negros, de pacientes con seguro Medicaid y de pacientes con cáncer aparecieron con más frecuencia que en el conjunto general.
El estudio midió ataques en un ambiente de investigación. No muestra que haya ocurrido una filtración real en un hospital, ni que alguien haya sido identificado. Los autores señalan que el mismo tipo de ataque puede usarse contra modelos generativos, y que probar su método en esos modelos queda como trabajo futuro.
Los propios investigadores proponen pedir protección matemática llamada privacidad diferencial, aplicada a cada paciente y no a cada registro. Pero el estudio no evalúa si eso mejora la salud de nadie, y no examina ninguna ley, ningún tribunal ni ningún trámite de reclamo en América.
Qué exigirán las autoridades de su país, y con qué pruebas, todavía no está escrito en ninguna parte.
Qué significa para usted
Usted puede pedir hoy, si su hospital usa estos sistemas, que le digan si sus datos sirvieron para entrenarlos y con qué protección. El estudio no prueba que alguien haya sido identificado ni que su hospital haya fallado, así que no saque conclusiones sobre su atención. Vigile qué exigen las autoridades de su país y qué responde su clínica cuando pregunte.
Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0
Quién pagó: El estudio fue financiado en parte por la Escuela de Excelencia Konrad Zuse en IA Confiable, el Consejo Europeo de Investigación (subvención n.º 884622), la Fundación Alemana de Investigación (proyecto n.º 532139938), la Real Academia de Ingeniería y la Universidad Técnica de Múnich, sin declaraciones sobre la intervención de los financiadores.
No tome esto como consejo médico profesional.
Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1583 palabras · unos 8 minLeerla →Cerrar
Un modelo médico puede delatar a un paciente concreto aunque el promedio diga que no hay riesgo
Un estudio audita siete bases de datos clínicas reales y encuentra que el ataque funciona casi perfecto contra algunas personas, mientras la cifra global apenas se distingue del azar.

El artículo describe un estudio según el cual los modelos de inteligencia artificial (IA) para diagnóstico médico prometen ampliar el acceso a diagnósticos de alta calidad, pero los datos con los que se entrenan suelen contener información sensible de pacientes que podría quedar expuesta mediante ataques de privacidad1. El artículo describe un estudio que se concentra en uno de esos ataques: la inferencia de pertenencia, que busca determinar si los datos de una persona concreta se usaron para entrenar un modelo2.
La idea es más simple de lo que suena. Cuando usted usa un modelo entrenado, normalmente solo ve la respuesta que le devuelve: por ejemplo, "78% de probabilidad de neumonía" ante una radiografía de tórax. El artículo describe un estudio según el cual ese acceso de caja negra puede ser explotado por un usuario no confiable para averiguar si un registro determinado formó parte del conjunto de entrenamiento del modelo3. El artículo describe un estudio según el cual el ataque se apoya en un detalle: los modelos de IA suelen mostrarse algo más seguros de sus predicciones sobre los datos con los que fueron entrenados que sobre los que no4.
¿Por qué importa eso? Porque saber que alguien pertenece a un conjunto de entrenamiento puede equivaler a saber algo íntimo. El artículo describe un estudio en el que se da un ejemplo: contra un modelo que predice la eficacia de la inmunoterapia contra el cáncer a partir de análisis de sangre rutinarios, un ataque exitoso de pertenencia revela que esa persona tiene cáncer5.
Lo nuevo del trabajo es la escala del examen. El artículo describe un estudio en el que se presenta una de las primeras auditorías de privacidad a nivel de paciente de modelos de IA para aplicaciones de diagnóstico médico2. Usaron siete conjuntos de datos clínicos reales y, para cada uno, entrenaron 200 modelos sobre subconjuntos aleatorios de pacientes, de modo que pudieron medir el éxito del ataque registro por registro y paciente por paciente, y no solo en promedio.
El artículo describe un estudio cuyo hallazgo central es incómodo: en una variedad de conjuntos de datos médicos, los ataques de inferencia de pertenencia pueden alcanzar tasas de éxito casi perfectas contra pacientes individuales, incluso cuando el desempeño agregado del ataque no se aparta sustancialmente del azar6. Dicho de otro modo: la métrica promedio puede indicar que casi no hay vulnerabilidad mientras algunas personas concretas quedan prácticamente expuestas.
El artículo describe un estudio en el que se lo formula sin rodeos: en conjunto, sus hallazgos muestran que las métricas de privacidad agregadas pueden subestimar gravemente el riesgo individual7. El artículo describe un estudio en el que se agrega un segundo patrón: el número de pacientes con alto éxito de ataque aumenta sustancialmente con la capacidad del modelo, y los grupos subrepresentados —estratificados por estado de enfermedad, raza autoinformada, seguro de salud, sexo o protocolo de imagen— enfrentan un éxito de ataque desproporcionadamente alto8.
El artículo describe un estudio según el cual los grupos que aportan menos registros a los datos de entrenamiento suelen estar sobrerrepresentados entre los registros más susceptibles al ataque, y lo contrario suele ocurrir con los grupos mayoritarios9. El artículo describe un estudio en el que se señala que este hallazgo complementa la literatura existente sobre desigualdades en salud, que ha reportado peores resultados de salud y menor esperanza de vida para grupos marginados y minoritarios10, y se advierte que las tendencias actuales de desarrollo y despliegue de IA médica podrían agravar esas desigualdades11.
Hay un ejemplo concreto en una de las bases de datos. El artículo describe un estudio en el que, en MIMIC-IV-ED, registros de pacientes negros, de pacientes con seguro Medicaid o de pacientes diagnosticados con cáncer aparecieron con más frecuencia de la esperada entre los registros más vulnerables: +31%, +126% y +18% de cambio relativo respecto del conjunto de datos completo, respectivamente12.
Los propios autores marcan los límites. El artículo describe un estudio en el que queda abierto si estos perfiles desiguales de riesgo se extienden a ataques distintos de la inferencia de pertenencia13. El artículo describe un estudio en el que también se observó que las violaciones de la garantía de privacidad diferencial a nivel de registro aparecen solo en un subconjunto de pacientes bajo protección fuerte, y que mitigar por completo estos ataques para todos los pacientes requeriría implementar la protección a nivel de paciente y no a nivel de registro14.
El equipo que firma el trabajo incluye a Moritz A. Knolle, Martin J. Menten, Friederike Jungmann, Felix Meissen, Ben Glocker, Daniel Rueckert y Georgios Kaissis, con afiliaciones en la Universidad Técnica de Múnich, el Imperial College de Londres y la Universidad de Potsdam. El financiamiento proviene de la escuela Konrad Zuse para IA confiable, una beca del Consejo Europeo de Investigación y la Fundación Alemana de Investigación; Ben Glocker recibió apoyo de la Real Academia de Ingeniería del Reino Unido. Dos autores declaran empleo parcial fuera de la academia: Glocker en DeepHealth y Kaissis en Google DeepMind.
¿Qué es, en palabras llanas, un ataque de inferencia de pertenencia? Es una prueba de hipótesis disfrazada de consulta. Usted pregunta al modelo por un registro y observa cuán seguro se muestra; si se muestra más seguro de lo habitual, sospecha que ese registro estuvo en el entrenamiento4. No se necesita acceso a los parámetros del modelo ni capacidad de modificarlo: basta con consultarlo, como lo haría cualquier usuario3.
Conviene distinguir este tipo de ataque de otros más agresivos, que exigen acceso a los parámetros del modelo, a las actualizaciones durante el entrenamiento o incluso la capacidad de modificar su arquitectura. Los autores no los consideran en este estudio porque sus supuestos no son realistas para escenarios de despliegue cuidadosos. El ataque que sí estudian requiere una sola consulta al modelo y podría ejecutarlo cualquiera que se presente como usuario real de un sistema de IA.
Tampoco protegen aquí las técnicas de gobernanza de datos como el aprendizaje federado o en enjambre: como los ataques se ejecutan contra modelos ya entrenados, esas estrategias no ofrecen defensa. Es un punto que suele sorprender, porque se presentan a menudo como solución de privacidad.
¿Qué se ha intentado? La mitigación que los autores consideran más prometedora es la privacidad diferencial, un enfoque matemáticamente verificable. El artículo describe un estudio según el cual consiste en perturbar con ruido blanco las actualizaciones de parámetros durante el entrenamiento o el ajuste fino, limitando la contribución de los datos de cualquier individuo al modelo final15. No es una promesa comercial: es una garantía formal, y por eso los autores la respaldan.
Pero la privacidad diferencial tiene un costo y un umbral. Los resultados del estudio indican que mitigar por completo estos ataques para todos los pacientes exige protección a nivel de paciente, no de registro14. El artículo describe un estudio en el que se recomienda que los estándares de reporte de auditorías de privacidad cambien: las auditorías deberían informar el éxito de los ataques al nivel de cada persona que aportó datos o, si no hay identificadores disponibles, al nivel de cada registro16.
Vale la pena detenerse en lo que este trabajo no dice. No afirma que estos ataques hayan causado filtraciones reales en hospitales en funcionamiento: es una auditoría de investigación que midió el éxito de ataques en laboratorio, no un daño ocurrido. Tampoco evalúa si la privacidad diferencial mejora los resultados clínicos de los pacientes: solo midió el éxito de los ataques y el desempeño diagnóstico de los modelos. Y no establece que los modelos generativos de IA —los que producen texto o imágenes— enfrenten el mismo riesgo; los propios autores dejan eso como pregunta abierta13.
Para el lector en América Latina y en América del Norte, hay una consecuencia práctica que este estudio no aborda y que conviene nombrar con claridad: el artículo no contiene información sobre derechos legales, marcos legales de privacidad, decisiones judiciales ni mecanismos de exigencia en ningún país. Es un trabajo técnico. Lo que sí ofrece es una pregunta que usted puede llevar a cualquier discusión sobre IA médica en su país: ¿los reportes de privacidad de un sistema que usa sus datos hablan de promedios o hablan de usted?
Esa pregunta tiene sentido porque las cifras agregadas son las que suelen publicarse. Un informe puede decir que el riesgo global es bajo y, al mismo tiempo, ocultar que un pequeño grupo de personas queda prácticamente expuesto7. Si usted aporta datos a un sistema de salud que entrena modelos, tiene motivos para preguntar cómo se mide el riesgo de una persona concreta y no solo el del conjunto.
Los autores también señalan un efecto de círculo vicioso que vale la pena considerar: si los grupos minoritarios perciben menor rendimiento y menor seguridad en los modelos de IA, podría disminuir su disposición a contribuir con datos para entrenarlos. Eso afectaría justamente a los grupos cuya representación ya es escasa, y con ella la calidad de los modelos para todos.
Nos quedamos con una idea que el estudio no puede resolver y que usted sí puede usar: cuando alguien le presente un sistema de IA médica como seguro, la pregunta útil no es cuál es el riesgo promedio, sino cuál es el riesgo para la persona que está enfrente.
De dónde sale cada dato de contexto, y cuánto leímos de cada documento
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Medical artificial intelligence (AI) models hold the promise to improve global access to high-quality diagnostics 1 . However, the training data underlying these models often contain sensitive patient information that may be exposed through privacy attacks 2 – 7 ."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Here we present one of the first patient-level privacy audits of AI models for medical diagnostic applications. We focus on membership inference attacks 2 – 4 (MIAs), which seek to determine whether the data of a given individual were used to train a model."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This black-box access to a model can be exploited by an untrusted user to conduct a MIA that shows the membership status of a target record, that is, whether the target record was a member of the training dataset of a model or not (Fig. 1a )."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "To infer membership status, MIAs typically make use of the fact that AI models are often slightly more confident about their predictions on training than on non-training data."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For example, a successful MIA against the model in ref. 10 , which predicts anti-cancer immunotherapy efficacy from routine blood test data, reveals that an individual has cancer."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Across a diverse range of medical datasets, we show that MIAs can achieve near-perfect success rates for individual patients, even when the aggregate performance does not substantially deviate from random guessing."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Together, our findings show that aggregate privacy metrics can severely underestimate individual privacy risk."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "We further find that the number of patients with high attack success increases substantially with model capacity, and that underrepresented groups—stratified by disease status, self-reported race, insurance, sex or imaging protocol—face disproportionately high attack success."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Groups of patients that are underrepresented in a model training dataset are often overrepresented among the records most susceptible to MIAs. By contrast, the opposite often holds for majority groups."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This finding—that a disproportionately large share of the AI privacy risk burden rests on underrepresented groups—complements the existing literature on health inequalities, which has reported worse health outcomes and life expectancy for marginalized and minority groups 46 ."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Our findings suggest that current trends in medical AI development and deployment could exacerbate these health inequalities."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For example, in MIMIC-IV-ED, records from Black patients, patients with Medicaid insurance or patients diagnosed with cancer were observed more frequently than expected among the most vulnerable records (+31%, +126%, and +18% relative change to the overall dataset, respectively)."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Whether the disparate risk profiles we observe extend to attacks beyond MIAs remains an open question, motivating the further development of risk assessment and mitigation techniques that cater to all data-contributing patients."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Specifically, our results indicate that fully mitigating MIAs for all data-contributing patients requires implementing DP protection at the patient level rather than at the record level."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "To this end, mathematically verifiable approaches to risk mitigation, such as DP 48 , are emerging as the most promising solution. DP, by carefully perturbing parameter updates with white noise during model training or fine-tuning 49 , limits the contribution of the data of any individual to the parameter update and, by extension, to the final model."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Thus, reporting standards for AI privacy audits need to change. Audits should report the success of privacy attacks at the level of individual data contributors or, if the necessary patient- or person-level identifiers are unavailable, at the record level."
Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0
Quién pagó: El estudio fue financiado en parte por la Escuela de Excelencia Konrad Zuse en IA Confiable, el Consejo Europeo de Investigación (subvención n.º 884622), la Fundación Alemana de Investigación (proyecto n.º 532139938), la Real Academia de Ingeniería y la Universidad Técnica de Múnich, sin declaraciones sobre la intervención de los financiadores.
No tome esto como consejo médico profesional.
Los hallazgos de otros estudios que aquí se mencionan los conocemos por este documento, que fue el que leímos; no abrimos cada uno de esos estudios.
analysis of texts · Nature · the paper, 24 Jun 2026 · free
Medical AI Can Recognize Who Was in the Training Data
A research audit of seven clinical datasets found some patients could be identified from a diagnostic model, though the study measured attack success, not real breaches.
Short version · the longer version follows, about 5 min
Ask the weeklyAI chat
Ask me about this study: who was studied, what it found, and what it does not say.
Conversations are saved for as long as weeklyAI exists, to improve the publication. Answers come in the language you write in.
- The study at a glance
- Who
- patients in seven large sets of real clinical data
- How many
- seven large sets of real clinical data
- Where
- hospitals including Stanford Hospital, Beth Israel Deaconess Medical Center, Emory University in Atlanta, and Harvard Medical School
- When
- not stated in the passages
- Kind of study
- analysis of what people did
- Who did it
- Technical University of Munich and partner institutions
- The limit that matters
- It measured how well an attack could work in a research setting, not real breaches.
These are relative changes in how often each group appeared among the most vulnerable records compared with the overall dataset in MIMIC-IV-ED; they are not shares of patients and not measures of actual breaches.
The average attack looked like a coin toss.

You give your health data to a hospital. A company builds an AI tool with it. What protects you?
This study does not answer that. It is not a law, a court ruling, or a regulator's decision. It is a privacy audit by researchers at the Technical University of Munich and partner institutions, published in Nature. They tested whether someone using only a model's predictions could tell if a particular patient's records helped train it.
They studied seven large sets of real clinical data from hospitals and other sources. The sets held chest X-rays, mammograms, eye images, skin images, heart recordings, and emergency room records. The sets came from hospitals including Stanford Hospital, Beth Israel Deaconess Medical Center, Emory University in Atlanta, and Harvard Medical School.
For each dataset, the researchers trained 200 versions of a diagnostic AI model on random groups of patients. Then they checked whether an attacker could guess, from the model's answers alone, which patients were in the training group.
The average attack looked like a coin toss. But that average hid a great deal. For some individual patients, the attack worked almost perfectly. The researchers say these were often people in groups that were small in the data: Black patients, patients on Medicaid, or patients with cancer in one emergency room dataset appeared more often than expected among the most exposed records. Bigger models put more patients at high risk.
The study cannot show that any patient's privacy was actually breached in a hospital. It measured how well an attack could work in a research setting, not what happened in real care. It also cannot say whether the uneven risks it found would show up in other kinds of attacks on these models.
For readers in the Americas, the study offers no legal shield, no complaint form, no phone number. It found no law or rule that requires a company to protect you this way.
What is still unknown is whether any government will require the strongest privacy protection, called differential privacy, at the level of the patient. The researchers say that is what would be needed to protect everyone.
What this means for you
For now, no law or court ruling names the harm described here, so there is nothing specific to demand from a company yet. If a clinic or a hospital offers you an AI tool, you can ask whether it was trained with the strongest protections, at the level of the patient.
Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0
Who paid: The study was supported by the Konrad Zuse School of Excellence in Reliable AI, the European Research Council (grant no. 884622), the German Research Foundation (project no. 532139938), the Royal Academy of Engineering, and Technische Universität München, with no statement on funder control.
Do not take this as professional medical advice.
The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1063 words · about 5 minRead it →Close
Medical AI Can Give Up a Patient’s Secret—and the Average Says Nothing
A Munich-led audit of seven clinical datasets finds that some patients are near-perfectly identifiable from a diagnostic model, while the usual privacy score looks like a coin toss.

Medical artificial intelligence promises better access to high-quality diagnostics, but the data behind these models often carry sensitive patient information that privacy attacks can expose.1 A team including Moritz A. Knolle and Georgios Kaissis, with the Technical University of Munich, Imperial College London and the Hasso Plattner Institute, ran what they describe as one of the first privacy audits aimed at individual patients rather than whole datasets.2 Their target was the membership inference attack: a way of asking a trained model, through nothing but its ordinary predictions, whether a particular person’s record was part of the data it learned from.2
The idea is simple. A model is often slightly more confident about the cases it was trained on than about cases it has never seen.3 An outsider who can only send in a record and read the answer—what the authors call black-box access—can use that small difference to guess whether the record belonged to the training set.4 In a narrow group, that guess is not trivia. The authors give an example: a model that predicts whether immunotherapy will work, built from routine blood tests, gives away that a person has cancer if membership can be established.5
Their central result is that this danger is not spread evenly. Across a diverse range of medical datasets, the attacks reached near-perfect success for individual patients even when the overall performance was barely better than guessing.6 To see this, they trained 200 models per dataset on random subsets of patients and scored each record separately. The authors’ conclusion is blunt: privacy measured as an average can severely underestimate the risk to an individual.7
The datasets were real clinical material, not laboratory toys: chest radiographs from Stanford and from Beth Israel Deaconess, a dermatology image collection drawn from open-source atlases, an ophthalmology set from Harvard, a mammography set from Emory, electrocardiograms and emergency-department records. They trained the models with the standard techniques that keep them from memorising—augmentation, weight decay, learning-rate schedules—and still found a small group of patients who were highly exposed.
The risk grew with the size of the model.8 In the dermatology set, the number of patients with near-perfect attack success rose from none, to one in ten thousand, to one in a thousand, to one in ten as the models got larger.8 And it fell unevenly across groups: patients who are underrepresented in the training data—by disease, self-reported race, insurance, sex or imaging protocol—were overrepresented among the most exposed records.8 In the emergency-department data, records from Black patients, from patients on Medicaid and from patients diagnosed with cancer appeared among the most vulnerable more often than their share of the data would suggest.9 The authors put the increases at 31 percent, 126 percent and 18 percent respectively.9
Two limits matter here. The study measured how well an attack *could* work in a research audit; it did not find that anyone’s records were actually stolen in a hospital. And whether the same uneven pattern shows up for other kinds of attack remains an open question the authors themselves flag.10 This is not a finding about generative AI or about any health system the paper does not name.
What would have to change? The authors argue that privacy audits should stop reporting a single average and start reporting success for individual data contributors, or at least record by record.11 They point to differential privacy—a technique that adds carefully calibrated noise during training so that no single person’s data can shape the final model—as the most promising verifiable fix.12 Their experiments suggest that protecting every patient would require applying it at the level of the patient, not the individual record.13 One caution belongs with that: the study measured attack success and diagnostic accuracy, nothing more. It did not test whether privacy protection improves anyone’s health.
Two other strands of research, which we could read only in summary, sit alongside this one. A small interview study of fifteen people who use companion chatbots such as Replika and Character.AI found that the human-like design lowers the threshold for intimate disclosure, and that users who recognise the corporate machinery underneath still treat what they share as jointly held.1415 A separate comparison of six leading large language models against European data-protection principles found strong security controls but clear gaps in limiting purpose, minimising data and documenting lawful grounds—and mixed support for a person’s right to see, correct or delete what a model holds.1617 And a third line of work, also read in summary only, warns that removing identifiers one prompt at a time may not be enough: across a long clinical conversation, small details accumulate, and in a simulated cohort about four in five patients fell below the small-cell threshold before the conversation ended.1819
In the United States, that last point touches a specific rule. Under the HIPAA Safe Harbor method, eighteen categories of identifier must be stripped from each disclosure independently, but the method does not assess risk that builds up across turns.20 The authors of that separate work note that clinicians have no tool to judge that cumulative risk in real time.21 The present study offers no legal remedy of its own—it is a technical audit, not a ruling, a statute or a regulator’s decision.
Why bring this to ordinary readers now? Because the question it raises is not abstract. If a hospital, a clinic or a company trains a diagnostic tool on your scans or your records, the average privacy number it publishes may tell you very little about what that model reveals about *you*. The people most exposed, in this study, were the ones already least represented.
What the authors recommend is narrow and practical: decide, model by model, what an attacker could learn from membership alone; protect the vulnerable ones with verifiable methods or strict access controls; and report risk at the level of the person.1112
For a reader with no legal training, the useful move is to ask a direct question of whoever holds your medical data—your hospital, your insurer, the vendor selling them a model. Not “what is your average privacy score,” but: for the patients most exposed by this model, what have you done?
Where each piece of context comes from, and how much of it we read
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Medical artificial intelligence (AI) models hold the promise to improve global access to high-quality diagnostics 1 . However, the training data underlying these models often contain sensitive patient information that may be exposed through privacy attacks 2 – 7 ."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Here we present one of the first patient-level privacy audits of AI models for medical diagnostic applications. We focus on membership inference attacks 2 – 4 (MIAs), which seek to determine whether the data of a given individual were used to train a model."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "To infer membership status, MIAs typically make use of the fact that AI models are often slightly more confident about their predictions on training than on non-training data."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "This black-box access to a model can be exploited by an untrusted user to conduct a MIA that shows the membership status of a target record, that is, whether the target record was a member of the training dataset of a model or not (Fig. 1a )."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "For example, a successful MIA against the model in ref. 10 , which predicts anti-cancer immunotherapy efficacy from routine blood test data, reveals that an individual has cancer."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Across a diverse range of medical datasets, we show that MIAs can achieve near-perfect success rates for individual patients, even when the aggregate performance does not substantially deviate from random guessing."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Together, our findings show that aggregate privacy metrics can severely underestimate individual privacy risk."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "We further find that the number of patients with high attack success increases substantially with model capacity, and that underrepresented groups—stratified by disease status, self-reported race, insurance, sex or imaging protocol—face disproportionately high attack success."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "For example, in MIMIC-IV-ED, records from Black patients, patients with Medicaid insurance or patients diagnosed with cancer were observed more frequently than expected among the most vulnerable records (+31%, +126%, and +18% relative change to the overall dataset, respectively)."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Whether the disparate risk profiles we observe extend to attacks beyond MIAs remains an open question, motivating the further development of risk assessment and mitigation techniques that cater to all data-contributing patients."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Thus, reporting standards for AI privacy audits need to change. Audits should report the success of privacy attacks at the level of individual data contributors or, if the necessary patient- or person-level identifiers are unavailable, at the record level."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "To this end, mathematically verifiable approaches to risk mitigation, such as DP 48 , are emerging as the most promising solution. DP, by carefully perturbing parameter updates with white noise during model training or fine-tuning 49 , limits the contribution of the data of any individual to the parameter update and, by extension, to the final model."
- Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0 - the article this story is about — the whole article — the passage: "Specifically, our results indicate that fully mitigating MIAs for all data-contributing patients requires implementing DP protection at the patient level rather than at the record level."
- Chiu HC, Foote J. (2026). Chatting with confidants or corporations? Privacy management with AI companions. Journal of Computer-Mediated Communication. https://doi.org/10.1093/jcmc/zmag014 — only the abstract - the full text could not be fetched — the passage: "We find that anthropomorphic design lowers disclosure thresholds, encouraging increasingly intimate self-disclosure."
- Chiu HC, Foote J. (2026). Chatting with confidants or corporations? Privacy management with AI companions. Journal of Computer-Mediated Communication. https://doi.org/10.1093/jcmc/zmag014 — only the abstract - the full text could not be fetched — the passage: "As relationships develop, some users engage in simulated co-ownership, treating shared information as relationally held despite recognizing corporate control."
- Alzoubi YI, Mishra A. (2026). Assessing privacy risks of proprietary and open-weight large language models: A GDPR-oriented comparative analysis. Journal of Responsible Technology. https://doi.org/10.1016/j.jrt.2026.100187 — only the abstract - the full text could not be fetched — the passage: "Findings indicate strong performance in security and risk-based controls, particularly among proprietary models, while notable gaps remain in Purpose Limitation, Data Minimization, and documentation of lawful bases."
- Alzoubi YI, Mishra A. (2026). Assessing privacy risks of proprietary and open-weight large language models: A GDPR-oriented comparative analysis. Journal of Responsible Technology. https://doi.org/10.1016/j.jrt.2026.100187 — only the abstract - the full text could not be fetched — the passage: "Data Subject Rights exhibited mixed support across all evaluated models, reflecting technical challenges in implementing access, rectification, and erasure within machine learning contexts."
- Weatherhead J, Hasan A, Weatherhead J, Golovko G, Grant B, Garcia JD, et al. (2026). K-anonymity decay in multi-turn clinical large language model conversations. Frontiers in Digital Health. 10.3389/fdgth.2026.1832168 — only the abstract - the full text could not be fetched — the passage: "Under progressive disclosure, 79.9% of simulated patients fell below the small-cell threshold ( k<5 ) by the end of their disclosure sequence, with a median of seven disclosure steps to reach this threshold; when rare attributes were disclosed first, the median decreased to four steps."
- Weatherhead J, Hasan A, Weatherhead J, Golovko G, Grant B, Garcia JD, et al. (2026). K-anonymity decay in multi-turn clinical large language model conversations. Frontiers in Digital Health. 10.3389/fdgth.2026.1832168 — only the abstract - the full text could not be fetched — the passage: "Even when no individual step disclosed a direct identifier, the cumulative quasi-identifier profile degraded k-anonymity below accepted safety thresholds."
- Weatherhead J, Hasan A, Weatherhead J, Golovko G, Grant B, Garcia JD, et al. (2026). K-anonymity decay in multi-turn clinical large language model conversations. Frontiers in Digital Health. 10.3389/fdgth.2026.1832168 — only the abstract - the full text could not be fetched — the passage: "Under the Health Insurance Portability and Accountability Act (HIPAA) Safe Harbor method, eighteen specified identifier categories must be removed from each disclosure independently, but this approach does not assess cumulative re-identification risk: successive turns reveal additional quasi-identifiers that progressively narrow the set of matching records, eroding the privacy protection that any single turn appeared to provide."
- Weatherhead J, Hasan A, Weatherhead J, Golovko G, Grant B, Garcia JD, et al. (2026). K-anonymity decay in multi-turn clinical large language model conversations. Frontiers in Digital Health. 10.3389/fdgth.2026.1832168 — only the abstract - the full text could not be fetched — the passage: "These findings suggest an operational limitation in per-prompt de-identification as applied to multi-turn clinical AI and large language model conversations: although HIPAA Safe Harbor includes a provision requiring no actual knowledge that remaining information could identify an individual, clinicians lack tools to assess cumulative re-identification risk in real time."
Knolle, M. A., Menten, M. J., Jungmann, F. et al. (2026). Disparate privacy risks from medical AI. Nature. https://doi.org/10.1038/s41586-026-10688-0
Who paid: The study was supported by the Konrad Zuse School of Excellence in Reliable AI, the European Research Council (grant no. 884622), the German Research Foundation (project no. 532139938), the Royal Academy of Engineering, and Technische Universität München, with no statement on funder control.
Do not take this as professional medical advice.