analysis of texts · PLOS digital health · la publicación, 27 jul 2026 · gratis
La precisión de una IA hospitalaria puede desvanecerse sin que nadie lo note
Un estudio siguió cuatro sistemas de IA clínica tras su puesta en marcha: el rendimiento validado no se mantuvo, y las señales que avisaban antes no eran las que se suelen vigilar.
Versión breve · la versión detallada sigue, unos 9 min
- El estudio, de un vistazo
- Quiénes
- cuatro sistemas de inteligencia artificial clínica ya en uso en hospitales
- Cuántos
- cuatro sistemas
- Dónde
- una misma organización de salud
- Cuándo
- periodos de 18 a 24 meses
- Tipo de estudio
- análisis de lo que hicieron los sistemas en uso real
- Quién lo hizo
- Georgy Kopanitsa, publicado en PLOS Digital Health
- El límite que importa
- Solo cuatro sistemas en una organización: muestra un patrón, no una regla universal.
Qué se desvió primero: la calibración o la capacidad de ordenar pacientes
La calibración se deterioró de forma consistente y a menudo antes de que se notara cualquier cambio en la capacidad de ordenar
La capacidad de ordenar se mantuvo cerca de los niveles de validación durante los primeros meses, mientras la calibración ya se había desviado
Las señales operativas dieron indicios tempranos de fragilidad; la vigilancia por desenlaces llegaba tarde por los retrasos en las etiquetas y la documentación
El estudio encontró que el problema temprano más consistente no fue la caída en la capacidad de ordenar, sino la pérdida de calibración: el riesgo previsto dejó de coincidir con el riesgo observado.

Muchos hospitales ya usan programas de inteligencia artificial que asignan una puntuación a cada paciente: riesgo de deterioro, de sepsis, de ciertos diagnósticos. La alerta llega a una lista o a una pantalla, y alguien decide si actúa.
El estudio, publicado en PLOS Digital Health y firmado por Georgy Kopanitsa, examinó cuatro de esos sistemas en uso real dentro de una misma organización de salud. Los comparó con los resultados que cada uno había obtenido antes de instalarse y observó cómo se comportaron después, entre 18 y 24 meses, con datos clínicos y registros operativos recogidos de rutina.
El hallazgo central: el buen desempeño de la validación no se conservó como una propiedad estable. Lo que se desvió primero y con más constancia fue la calibración, es decir, la correspondencia entre el riesgo anunciado y el riesgo observado. En dos de los sistemas, la capacidad de ordenar a los pacientes por riesgo seguía pareciéndose a la de la validación mientras la calibración ya se había alejado de ella. Las señales del flujo de trabajo —datos de entrada faltantes y retrasos en la llegada de la información— se asociaron con ese deterioro de forma más fuerte y más constante que los indicadores de población, como la demografía o la frecuencia de la enfermedad.
Es una advertencia sobre cuándo se mira, no una condena de ninguna herramienta concreta. El estudio abarcó solo cuatro sistemas en una organización, de modo que muestra un patrón, no una regla sobre el hospital que usted frecuenta. Y midió el comportamiento de los modelos y de sus datos operativos, no los resultados de los pacientes. Tampoco midió la confianza de los médicos ni fijó un umbral a partir del cual convenga dejar de confiar en un sistema.
Las etiquetas que confirman lo que pasó —el diagnóstico, el desenlace— llegaron con días o hasta meses de retraso, según la tarea. Si la vigilancia depende solo de ellas, el problema se detecta tarde, cuando ya lleva tiempo ocurriendo. Las señales disponibles en el momento mismo de la predicción, en cambio, se pueden observar de inmediato.
Los autores no evalúan leyes ni regulaciones de ningún país, y este estudio no sirve para aprobar o rechazar un producto. Lo que ofrece es una pregunta concreta para llevar a una reunión, a una consulta o a una oficina de atención al usuario: ¿qué señales se revisan mes a mes en este sistema, y quién las mira antes de que lleguen los resultados definitivos?
Qué significa para usted
Si usted usa o depende de un sistema así, pregunte qué señales operativas se revisan mes a mes y quién las mira antes de que lleguen los diagnósticos definitivos. No concluya todavía que la herramienta falla ni que es segura: el estudio solo siguió cuatro sistemas en una organización y no midió daños a pacientes. Ante dudas sobre su atención, consulte a un profesional.
Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534
Quién pagó: El autor no recibió financiación específica para este trabajo, y el artículo no indica que ningún financiador tuviera participación alguna ni que se prestara equipo o software.
No tome esto como consejo médico profesional.
Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1794 palabras · unos 9 minLeerla →Cerrar
Un sistema de IA aprobado en un hospital puede fallar en silencio, según un estudio de cuatro sistemas
El estudio midió el comportamiento de cuatro herramientas clínicas ya en uso, no los resultados en pacientes.

Los sistemas de inteligencia artificial que se usan hoy en hospitales para decidir a quién vigilar de cerca, a quién mandar primero a una cama o a quién alertar por posible sepsis se aprueban con pruebas hechas antes de instalarlos. Un estudio retrospectivo siguió a cuatro de esos sistemas, ya en funcionamiento dentro de una misma organización de salud, durante periodos de 18 a 24 meses. Encontró que el buen desempeño de las pruebas de validación no se mantuvo como una propiedad estable después del despliegue: la calibración se deterioró de forma consistente y a menudo antes de que se notara cualquier cambio en la capacidad de ordenar pacientes por riesgo1.
Vale la pena entender qué significa "calibración", porque es el corazón del asunto. Imagine que un sistema dice que un paciente tiene 20 por ciento de probabilidad de deteriorarse en las próximas 24 horas. Calibración es que esa cifra se parezca a la realidad: que de cada cien pacientes con ese puntaje, aproximadamente veinte terminen deteriorándose. Discriminación, en cambio, es la capacidad de ordenar: poner a los más graves por delante de los menos graves, aunque los números absolutos estén mal. El estudio encontró que el problema temprano más consistente no fue la caída en la capacidad de ordenar, sino la pérdida de calibración: el riesgo previsto dejó de coincidir con el riesgo observado2.
En dos de los sistemas, ambos modelos de predicción sobre historia clínica electrónica en atención aguda, la capacidad de ordenar pacientes se mantuvo cerca de los niveles de validación durante los primeros meses, mientras que las medidas de calibración ya se habían desviado3. El estudio llama a esto "falla silenciosa de calibración": los modelos conservaron la capacidad de ordenar pero subestimaron o sobreestimaron sistemáticamente el riesgo absoluto. Como las decisiones de umbral y de escalamiento se apoyan en esas cifras absolutas, vigilar solo la capacidad de ordenar habría retrasado la detección4.
El estudio también siguió señales operativas que no dependen de conocer el desenlace clínico: cuántos datos de entrada faltaban en el momento de la predicción y cuánto tardaban en llegar. Esas señales dieron indicios tempranos de fragilidad, mientras que la vigilancia basada en desenlaces llegaba tarde por los retrasos en las etiquetas y en los procesos de documentación5. La razón de ese retraso es concreta: esas señales operativas se podían observar en el momento mismo de la predicción, mientras que las etiquetas de desenlace solo estaban disponibles tras demoras que, según la tarea, iban de días a meses6. Bajo una vigilancia basada únicamente en desenlaces, el deterioro se habría detectado solo después de esa demora, lo que deja una brecha de gobernanza: los sistemas pueden seguir operando aunque ya haya evidencia operativa de fragilidad emergente7.
El estudio tiene límites que conviene no saltarse. Cubrió solo cuatro sistemas de IA en una sola organización de salud, así que sus cifras no son universales: lo que muestra es un patrón, no una regla sobre la herramienta de ningún hospital en particular8. Midió el desempeño de los modelos y señales operativas, no resultados en pacientes: no mostró que nadie haya resultado dañado y no es evidencia de que una herramienta específica haya perjudicado a alguien. Tampoco midió la confianza de los médicos ni definió cuándo deberían dejar de apoyarse en un sistema, de modo que no puede decir cuál es un umbral seguro para ninguna herramienta concreta. El estudio abstrae deliberadamente tareas clínicas y modalidades de datos muy distintas, lo que simplifica matices propios de cada tarea, como el peso distinto que tiene la descalibración en una alerta temprana frente a una herramienta diagnóstica. No modeló las respuestas de la organización ante el deterioro detectado, ni el efecto de retroalimentación que ocurre cuando las propias predicciones del sistema influyen en las decisiones y, con el tiempo, en las etiquetas con que se audita. Y no evalúa el cumplimiento de marcos regulatorios específicos ni prescribe umbrales regulatorios.
Esta no es una preocupación aislada. En la literatura sobre aprendizaje automático, el problema se formaliza como cambio de conjunto de datos o deriva de concepto, y abarca cambios en la distribución conjunta de entradas, desenlaces y sus relaciones8. En salud, varios autores han subrayado que esos cambios no son eventos excepcionales sino una propiedad estructural de los datos clínicos8. El autor del trabajo que aquí resumimos cita a Finlayson y colaboradores, quienes sostienen que el cambio de conjunto de datos en medicina es "la regla más que la excepción", impulsado por cambios en la prestación de cuidados, los incentivos clínicos y la documentación, más que por variación aleatoria8.
Ya existe evidencia empírica que apoya esa inquietud. Análisis longitudinales han mostrado que los modelos predictivos pueden perder calibración y discriminación con el tiempo dentro de una misma institución, incluso cuando la demografía de los pacientes y la gravedad nominal de los casos se mantienen relativamente estables8. Minne y colaboradores mostraron que el desempeño de un modelo SAPS-II adaptado para mortalidad en cuidados intensivos cambió con el tiempo, lo que ilustra cómo los cambios temporales en la mezcla de casos y en los procesos de atención pueden afectar la fiabilidad de los modelos de ajuste por gravedad8. Los cambios en los procesos de atención, sobre todo en la solicitud de pruebas y en las prácticas de documentación, pueden alterar tanto la disponibilidad de las variables como el significado de que falten datos, lo que puede afectar el desempeño de un modelo de predicción de deterioro incluso cuando la mezcla de pacientes es relativamente estable8.
Los fracasos de generalización de alto perfil ilustran la misma fragilidad. Zech y colaboradores mostraron que un modelo de detección de neumonía aprendió factores propios de cada centro y ajenos a la clínica, lo que produjo grandes diferencias de desempeño entre hospitales pese a una validación interna sólida8. Evaluaciones externas de sistemas propietarios ampliamente desplegados, como el Epic Sepsis Model, han reportado mala discriminación y mala calibración en el uso clínico rutinario pese a su amplia adopción8. El estudio que nos ocupa se sitúa junto a esos sistemas reales en una tabla comparativa y señala que el reto central no es la falta de datos, sino la falta de alineación conceptual y operativa entre validación, monitoreo y gobernanza8.
Hay un aporte menos explorado: la deriva inducida por el flujo de trabajo. Los datos clínicos no se observan pasivamente; se producen mediante procesos sociotécnicos en los que intervienen clínicos, conjuntos de órdenes, plantillas, interfaces y políticas organizacionales8. Cambios en esos procesos pueden alterar no solo la distribución de las variables sino también su interpretación semántica8. Subbaswamy y Saria enfatizan que los sistemas predictivos que no toman en cuenta explícitamente la estructura causal son frágiles ante cambios de política y de flujo de trabajo, incluso cuando las propiedades estadísticas parecen similares8. Trabajos complementarios sobre fatiga de alertas y sesgo de automatización muestran que la interacción del clínico con los sistemas de apoyo puede influir recursivamente tanto en las entradas como en los desenlaces, creando bucles de retroalimentación que invalidan los supuestos originales de desempeño8.
Las guías regulatorias y de reporte reconocen cada vez más la necesidad de evaluar los sistemas de IA a lo largo de su ciclo de vida8. Marcos como TRIPOD-AI y CONSORT-AI han mejorado sustancialmente la transparencia en el desarrollo y la validación de modelos, y los organismos reguladores insisten en el monitoreo continuo de los dispositivos médicos habilitados con IA o aprendizaje automático8. Sin embargo, esos marcos se centran sobre todo en generar evidencia en el punto de evaluación y ofrecen orientación limitada sobre cómo detectar y gestionar el deterioro del desempeño durante la operación clínica rutinaria, en particular cuando las etiquetas de desenlace llegan tarde, están incompletas o están influidas por el uso del propio modelo8.
En cuanto a lo que sigue, el propio autor es explícito: la gobernanza posterior al despliegue debería incluir monitoreo continuo atento a la calibración y revisión de la telemetría operativa, sobre todo cuando las etiquetas de desenlace llegan tarde o están incompletas9. Y añade algo que reparte responsabilidades: si el deterioro lo impulsan la evolución del flujo de trabajo, los cambios en la documentación o las fallas en el conducto de datos, entonces la responsabilidad de la seguridad no puede recaer solo en quienes desarrollaron el modelo; debe distribuirse entre las operaciones clínicas, la tecnología de la información y las funciones de gobernanza de calidad10.
Así lo leemos nosotros. Imagine que le dicen que una herramienta fue aprobada una vez, hace dos años, y que con eso basta para confiar en ella hoy. Esa es la lógica que este estudio pone en duda: la aprobación inicial es una foto de un momento, no una garantía que se renueve sola. Lo que esperaríamos, en un hospital de nuestra región que use una herramienta así, es que su precisión se vaya desviando poco a poco sin que nadie lo note, porque nadie está mirando los números que sí cambian a diario: cuántos datos faltan, cuánto tardan en llegar. Si en cambio el sistema mantuviera exactamente el mismo desempeño tras uno o dos años de uso, sin que nadie lo haya ajustado ni revisado, entonces esta forma de verlo no aplicaría. Usted podría saber que nos equivocamos si el hospital le mostrara revisiones periódicas con resultados estables.
Y hay algo que usted puede hacer con esto, sin necesitar formación técnica. La próxima vez que una clínica, una aseguradora o una autoridad de salud le diga que su sistema de IA "es seguro porque fue validado", pregunte tres cosas concretas: cada cuánto lo revisan, quién es responsable de detectar que dejó de funcionar bien, y si pueden mostrarle esa revisión por escrito. Si solo le muestran la aprobación inicial, usted puede señalar, con razón, que un solo examen no demuestra que algo siga funcionando. Esa pregunta no requiere abogados ni conocimientos de estadística, y cambia la conversación: de "confíe" a "muéstreme cómo lo comprueban".
El estudio concluye que la validación debería tratarse como una referencia operativa de base y no como una garantía durable de seguridad, y que la gobernanza posterior al despliegue debería incluir monitoreo continuo atento a la calibración y revisión de la telemetría operativa, en particular cuando los desenlaces llegan tarde o incompletos9. Nada de esto significa que usted o alguien de su familia deba evitar, retrasar o reemplazar una atención indicada por un profesional. Significa que la pregunta sobre cómo se vigila una herramienta después de instalarla es una pregunta legítima que cualquier persona puede hacer. ¿Cada cuánto revisan el sistema de IA que se usa donde usted se atiende, y quién responde si deja de funcionar bien?
De dónde sale cada dato de contexto, y cuánto leímos de cada documento
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Across all systems, validation-era performance did not persist as a stable operational property after deployment. Calibration drift emerged consistently and often preceded detectable changes in discrimination."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The most consistent early problem was not always a drop in the model's ability to rank patients by risk, but a loss of calibration, meaning that the predicted risk no longer matched the observed risk."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "In Systems A and B, AUROC remained close to validation levels during early deployment, whereas calibration slopes declined and calibration intercepts shifted."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This represents silent calibration failure: the models preserved ranking ability but systematically misestimated absolute risk. Because absolute risk estimates inform thresholds and escalation decisions, monitoring focused primarily on AUROC would delay detection of this failure mode."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Label-independent operational signals, including input missingness and data latency, provided early indication of emerging fragility, whereas outcome-based monitoring was delayed by label latency and documentation processes."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Operational telemetry, including missingness and latency, was observable at inference time, while outcome labels became available only after task-dependent delays ranging from days to months."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Under outcome-based monitoring alone, degradation would therefore be detected only after a delay. This creates a governance gap in which systems may continue operating despite telemetry evidence of emerging performance degradation."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "While this study is grounded in real-world deployment data, the empirical scope is limited to a finite set of clinical AI systems operating within specific institutional contexts. As a result, quantitative effect sizes and degradation rates should not be interpreted as universal."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Post-deployment governance should therefore include continuous, calibration-aware monitoring and operational telemetry review, particularly when outcome labels are delayed or incomplete."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - el artículo del que trata esta nota — el artículo completo — el pasaje: "If degradation is driven by workflow evolution, documentation changes, or data pipeline failures, then responsibility for safety cannot reside solely with model developers. Instead, accountability must be distributed across clinical operations, information technology, and quality governance functions."
Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534
Quién pagó: El autor no recibió financiación específica para este trabajo, y el artículo no indica que ningún financiador tuviera participación alguna ni que se prestara equipo o software.
No tome esto como consejo médico profesional.
Los hallazgos de otros estudios que aquí se mencionan los conocemos por este documento, que fue el que leímos; no abrimos cada uno de esos estudios.
analysis of texts · PLOS digital health · the paper, 27 Jul 2026 · free
Hospital AI Can Quietly Lose Accuracy After Launch, Small Study Finds
Four patient-scoring systems drifted in the months after deployment. The warning signs appeared before the outcome data doctors usually rely on.
Short version · the longer version follows, about 7 min
- The study at a glance
- Who
- four clinical AI systems already running in routine care
- How many
- four
- Where
- one large healthcare organization
- When
- followed for 18 to 24 months each
- Kind of study
- analysis of what people did
- Who did it
- Almazov National Medical Research Centre and ITMO University, Saint Petersburg, Russia
- The limit that matters
- Only four systems in one organization; the pattern is not a rule for any hospital's tool.
Two ways of watching a deployed AI tool
The second kind of signal was available immediately, before any outcome was known, and moved earlier and tracked the slippage more closely than broad population measures.
Workflow signals were more strongly tied to the decline than population measures.
When a hospital says its AI was validated, ask the only question that word leaves open: checked when—and by whom?

A patient arrives at the emergency department. Software reads the chart and produces a risk score—how likely is this person to get worse, to develop sepsis, to need intensive care?
Hospitals in many countries now run tools that do exactly this. The article reports on four of them, all deployed in routine care inside one large healthcare organization, all previously checked against historical data and cleared for use.
Then researchers watched what happened next.
The study compared each system's pre-deployment results with its behavior after deployment, using routinely collected clinical data and operational records. Observation periods ranged from 18 to 24 months. The models were not retrained, and no thresholds were changed.
Validation-era performance did not hold. Calibration drifted early and consistently—meaning the predicted risk stopped matching the risk actually observed. Often this happened before the system's ability to rank patients by risk fell. Signals from the workflow itself, such as missing inputs and delayed data reaching the system, were more strongly tied to the decline than broad population measures.
The study covered only four systems in one organization and focused on model performance and operational behavior rather than downstream patient outcomes. It shows a pattern, not a rule about any particular hospital's tool, and it is not evidence that any specific AI tool has harmed anyone.
There is also a timing problem built into how hospitals check these systems. Outcome labels—whether the patient truly deteriorated, truly had sepsis—arrived days to months late. Monitoring based only on outcomes would therefore catch trouble late. The operational signals the study examined were visible at the moment the prediction was made.
What the study did not do: it did not measure whether doctors trusted these systems, and it did not define when clinicians should stop relying on one. It cannot tell you what a safe threshold is for any given tool.
That leaves a practical question anyone can carry into a hospital board meeting, a regulator's office or a public consultation: for the AI already running in your hospital, who is watching its inputs and timing—right now, before the outcome data arrive?
What this means for you
For the AI already running in your hospital, the study offers no rule about your own care or your country's law, and a doctor is the one to ask about any medical decision. What it does support is a question worth carrying to a board meeting or a regulator: who is watching the system's missing inputs and delayed data, at the moment each prediction is made, rather than waiting months for outcomes.
Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534
Who paid: The author received no specific funding for this work, and the article states no funders had any role and no equipment or software was lent.
Do not take this as professional medical advice.
The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1349 words · about 7 minRead it →Close
A Hospital AI Can Pass Every Test—Then Quietly Lose Its Grip
New research followed four clinical AI tools for up to two years after launch. Their accuracy drifted, and the first warning signs were not the ones hospitals usually watch.

A hospital's artificial intelligence tool can pass every test before launch and still go wrong afterward—without anyone touching the software. That is the central finding of a study published in July 2026 in *PLOS Digital Health* by Georgy Kopanitsa, a researcher at the Almazov National Medical Research Centre and ITMO University in Saint Petersburg, Russia. Kopanitsa followed four AI systems already running in one large healthcare organization for 18 to 24 months each. The models, their features, their decision thresholds and their intended uses were never modified during the study. Their performance still slid. "Across all systems, validation-era performance did not persist as a stable operational property after deployment," the paper states, adding that "calibration drift emerged consistently and often preceded detectable changes in discrimination"1.
To understand what slipped, you need one plain distinction. A model can be good at sorting patients into higher-risk and lower-risk piles—that is ranking—while still being wrong about what the actual number means. "The most consistent early problem was not always a drop in the model's ability to rank patients by risk, but a loss of calibration, meaning that the predicted risk no longer matched the observed risk"2. Picture a tool that tells a ward team a patient has a 20 percent chance of deteriorating within a day. If that number has quietly become 8 percent, the ranking may still look fine on paper while the number that drives the decision is wrong. In two of the systems, the ranking measure stayed close to its pre-launch level during the first months, "whereas calibration slopes declined and calibration intercepts shifted"3. Kopanitsa calls this "silent calibration failure: the models preserved ranking ability but systematically misestimated absolute risk," and notes that "monitoring focused primarily on AUROC would delay detection of this failure mode"4.
The study is small and specific, and its author says so. It covered four systems in one organization, and "quantitative effect sizes and degradation rates should not be interpreted as universal"5. The four tools spanned different jobs: two high-frequency models predicting deterioration and sepsis from electronic health records in acute care, one diagnostic support tool for specialty clinics, and one imaging-based risk model in radiology. The study measured how the models performed and what the operational data looked like—not what happened to patients. It did not measure whether doctors trusted the tools, did not set a threshold for when a tool should be pulled, and did not model how a hospital might respond once it spotted trouble. It also does not evaluate compliance with any specific regulation or set regulatory limits5.
The early warning signals Kopanitsa found were not about patients at all. They were about plumbing. "Label-independent operational signals, including input missingness and data latency, provided early indication of emerging fragility, whereas outcome-based monitoring was delayed by label latency and documentation processes"6. In plain terms: how often expected input fields were empty when the model ran, and how late the data arrived. Those signals "were observable at inference time, while outcome labels became available only after task-dependent delays ranging from days to months"7. In the study, one system's label was a clinical diagnosis that took 3 to 14 days to confirm, and another's was an outcome that took 7 to 30 days. So a hospital waiting for confirmed outcomes to check its tool is waiting weeks to learn what its own data pipeline already showed on day one. Kopanitsa calls this a governance gap, in which "systems may continue operating despite telemetry evidence of emerging performance degradation"8.
The comparison that matters here is between two ways of watching a deployed model. One watches what the model predicted versus what actually happened—accurate, but slow, because outcomes arrive late. The other watches the data feeding the model—empty fields, delayed results, features that stop showing up—and it is available immediately, before any outcome is known. The study found the second kind of signal moved earlier and tracked the slippage more closely than broad population measures like shifting patient demographics. This is the trade-off: the fast signal is not a measure of whether the tool is helping anyone, only of whether the conditions it was built for still hold. It tells you to look, not what you will find.
Here is how we read it. The pattern is not that hospitals buy bad software. It is that approval and ongoing reality are two different things, and the people who signed off have little incentive to be the ones who reopen the question. Everyone downstream—nurses, ward staff, technicians—sees the outputs daily, but the tool carries an official stamp, and questioning a stamped thing costs something. The doubts that never get said out loud are often the ones that would have caught the drift. What would prove us wrong is easy to describe and hard to find: a hospital that revises or pauses an AI tool the moment its own post-launch numbers wobble, before any outside pressure and before any visible harm. If you are told a hospital tool was "validated," that word describes a moment in the past. The useful question is what has been checked since, and by whom. That question is yours to ask, and no one has to grant you permission to ask it.
There is a second thing worth carrying out of this. When a complex system misbehaves, attention lands on the most visible part—usually the software or the company that sold it. But the study's own framing points elsewhere: if the drift comes from changed workflows, changed documentation, or a data pipeline that quietly shifted, then the vendor is not the only party holding the problem. So if an AI-assisted decision affects you or someone you care for, ask two questions in the same breath: who is responsible for watching this tool after launch, and where do I take a complaint? If the only answer you get is the name of the company, you have learned something about the hospital's arrangements, and you can say so.
What would change the picture, according to the study, is not a better model. It is a different habit: "Post-deployment governance should therefore include continuous, calibration-aware monitoring and operational telemetry review, particularly when outcome labels are delayed or incomplete"9. And responsibility, Kopanitsa argues, cannot sit with one party: "If degradation is driven by workflow evolution, documentation changes, or data pipeline failures, then responsibility for safety cannot reside solely with model developers. Instead, accountability must be distributed across clinical operations, information technology, and quality governance functions"10. That is a description of what should happen, not a report of what does.
Read the limits before you carry this anywhere. Four systems, one organization—a pattern, not a verdict on the tool at your local hospital5. The study measured model behavior and data signals, not patients; it is not evidence that any particular tool has hurt anyone. It did not measure clinician trust or define when a doctor should stop relying on a system, so it cannot tell you what a safe cutoff looks like for any given tool. And the author notes that the findings are conceptual guidance for evaluation and governance, "rather than as prescriptive thresholds applicable to all clinical AI systems"5.
So keep one thing. The next time a clinic, a health plan or a government office tells you an AI system was validated, you now know that word covers a test that happened once, under conditions that have since changed—and that the tool's own data pipeline will show the drift before any outcome does. You can ask when it was last checked, what the numbers were, and who is accountable for watching it. That is a question a patient, a family member or a hospital worker can put to an administrator without any legal training at all.
When a hospital says its AI was validated, ask the only question that word leaves open: checked when—and by whom?
Where each piece of context comes from, and how much of it we read
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "Across all systems, validation-era performance did not persist as a stable operational property after deployment. Calibration drift emerged consistently and often preceded detectable changes in discrimination."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "The most consistent early problem was not always a drop in the model's ability to rank patients by risk, but a loss of calibration, meaning that the predicted risk no longer matched the observed risk."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "In Systems A and B, AUROC remained close to validation levels during early deployment, whereas calibration slopes declined and calibration intercepts shifted."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "This represents silent calibration failure: the models preserved ranking ability but systematically misestimated absolute risk. Because absolute risk estimates inform thresholds and escalation decisions, monitoring focused primarily on AUROC would delay detection of this failure mode."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "While this study is grounded in real-world deployment data, the empirical scope is limited to a finite set of clinical AI systems operating within specific institutional contexts. As a result, quantitative effect sizes and degradation rates should not be interpreted as universal."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "Label-independent operational signals, including input missingness and data latency, provided early indication of emerging fragility, whereas outcome-based monitoring was delayed by label latency and documentation processes."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "Operational telemetry, including missingness and latency, was observable at inference time, while outcome labels became available only after task-dependent delays ranging from days to months."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "Under outcome-based monitoring alone, degradation would therefore be detected only after a delay. This creates a governance gap in which systems may continue operating despite telemetry evidence of emerging performance degradation."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "Post-deployment governance should therefore include continuous, calibration-aware monitoring and operational telemetry review, particularly when outcome labels are delayed or incomplete."
- Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534 - the article this story is about — the whole article — the passage: "If degradation is driven by workflow evolution, documentation changes, or data pipeline failures, then responsibility for safety cannot reside solely with model developers. Instead, accountability must be distributed across clinical operations, information technology, and quality governance functions."
Kopanitsa, G. (2026). Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001534
Who paid: The author received no specific funding for this work, and the article states no funders had any role and no equipment or software was lent.
Do not take this as professional medical advice.