experiment · Quantitative imaging in medicine and surgery · la publicación, 10 ago 2026 · gratis
La IA que leyó 404,502 mamografías en Moscú: lecturas más precisas, con una advertencia
Un estudio de casi dos años y medio en un solo sistema de salud encontró que el programa mejoró con el tiempo. No midió si las pacientes vivieron más.
Versión breve · la versión detallada sigue, unos 7 min
Pregunte a weeklyAI
Pregúnteme por este estudio: a quiénes se estudió, qué encontró y qué no dice.
Las conversaciones se guardan mientras exista weeklyAI, para mejorar la publicación. Se responde en el idioma en que usted escribe.
- El estudio, de un vistazo
- Quiénes
- Mamografías de mujeres de 18 años o más, leídas por un sistema de inteligencia artificial y revisadas por radiólogos
- Cuántos
- 404,502 mamografías de 206 organizaciones médicas; participaron 336 radiólogos
- Dónde
- Moscú, Rusia
- Cuándo
- El período de pruebas y seguimiento duró dos años y cinco meses; el sistema se incorporó al seguro médico obligatorio en junio de 2023
- Tipo de estudio
- analysis of what people did
- Quién lo hizo
- Centro de Diagnóstico y Telemedicina del Departamento de Salud de Moscú, junto con la plataforma desarrolladora
- El límite que importa
- No midió si las pacientes vivieron más; se hizo en una sola ciudad y un solo sistema de salud
Son los cambios porcentuales que el estudio reportó en su propio sistema de inteligencia artificial para mamografías durante dos años y cinco meses en Moscú; miden el desempeño del programa, no si las pacientes vivieron más.
Coincidencia entre lecturas de mamografías
Los dos radiólogos humanos coincidieron solo en 49.9% de los casos; la discrepancia crítica entre la inteligencia artificial y los radiólogos fue de 0.08% en promedio
Cuando le toque hacerse una mamografía, pregunte: ¿este centro usa inteligencia artificial para leer las imágenes? ¿Y con qué pruebas locales se evaluó antes de usarla en pacientes como yo?

En Moscú, un sistema de inteligencia artificial revisó 404,502 mamografías provenientes de 206 organizaciones médicas. Participaron 336 radiólogos. El período de pruebas y seguimiento duró dos años y cinco meses.
El programa se llama Third Opinion Mammograms, desarrollado por Third Opinion Platform, y funciona como una red neuronal que busca y clasifica señales de cáncer de mama en las imágenes. No trabaja solo: los radiólogos hacían la lectura final, y los resultados de la IA llevaban la etiqueta de "solo para investigación".
Los investigadores probaron el sistema por etapas y lo volvieron a evaluar después de cada actualización del programa. Además, un radiólogo experto independiente revisó una muestra de casos al inicio, al final y después de cada actualización, y comparó sus lecturas con las de la IA. Los investigadores probaron el sistema por etapas, lo volvieron a probar después de cada actualización del programa y compararon los resultados con las lecturas de radiólogos expertos e independientes. Cuando encontraban fallas, los desarrolladores corregían y el sistema se evaluaba otra vez.
Con ese proceso, la precisión del sistema subió de 0,77 a 0,90. La capacidad de detectar casos positivos pasó de 0,84 a 0,88, y la de descartar correctamente los negativos, de 0,70 a 0,91. Los defectos técnicos bajaron de 3.0% a 0.8%, y la calificación clínica que le dieron los radiólogos subió de 54.38% a 80.36%. En junio de 2023, el sistema se incorporó al programa de seguro médico obligatorio de Moscú.
Todo ocurrió en una sola ciudad y un solo sistema de salud, así que los mismos resultados no están garantizados en una clínica donde usted vive. Las imágenes venían de equipos de solo tres fabricantes —General Electric, Fujifilm y Medical Technologies Ltd.—, de modo que el desempeño en otras máquinas es desconocido. Y el estudio no midió si las pacientes vivieron más o tuvieron mejores resultados: comparó si las lecturas de la IA coincidían con las de los radiólogos, no si un cáncer detectado antes cambió el desenlace. La muestra con la que se calibró el sistema fue pequeña: 100 mamografías.
El estudio tampoco evaluó de forma formal el impacto en los flujos de trabajo de las organizaciones ni en la eficiencia económica. Eso queda para investigaciones futuras.
En un centro de salud pública, una mamografía con apoyo de IA podría significar detecciones más tempranas y menos falsas alarmas. Pero eso dependerá de si el sistema fue probado en su ciudad, con sus equipos y sus pacientes.
Cuando le toque hacerse una mamografía, pregunte: ¿este centro usa inteligencia artificial para leer las imágenes? ¿Y con qué pruebas locales se evaluó antes de usarla en pacientes como yo?
Qué significa para usted
Si en su centro le dicen que usan inteligencia artificial para leer mamografías, pregunte qué pruebas locales se hicieron antes de aplicarla a pacientes como usted. No dé por hecho que un resultado bueno en Moscú se repita en su ciudad, con otros equipos y otras pacientes. Consulte a un profesional sobre su propia lectura.
Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864
Quién pagó: Este trabajo fue apoyado por el Departamento de Salud de Moscú como parte de un esfuerzo de investigación y desarrollo; cuatro autores son empleados de Third Opinion Platform, el desarrollador del sistema de IA evaluado, y proporcionaron información sobre el sistema pero no participaron en la realización del estudio, el análisis de datos ni la interpretación de los resultados.
No tome esto como consejo médico profesional.
Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1373 palabras · unos 7 minLeerla →Cerrar
En Moscú, un sistema de inteligencia artificial para mamografías mejoró tras tres años de pruebas continuas
El estudio siguió más de 400,000 mamografías y terminó integrándose en el seguro público de salud de esa ciudad. No midió si las pacientes vivieron más.

Un estudio observacional realizado en Moscú siguió durante dos años y cinco meses un sistema de inteligencia artificial que lee mamografías, con más de 400,000 imágenes de 206 organizaciones médicas y la participación de 336 radiólogos1. Al final del período, la capacidad del sistema para distinguir un caso con signos de cáncer de uno sin ellos —medida como AUC— subió de 0,83 a 0,92; eso equivale a un aumento del 10.8%. La exactitud general pasó de 0,77 a 0,90; la capacidad de detectar los casos positivos, de 0,84 a 0,88; y la de descartar correctamente los negativos, de 0,70 a 0,912. La tasa promedio de fallas técnicas bajó de 3.0% a 0.8%, y la calificación clínica que le dieron los radiólogos subió de 54.38% a 80.36%3. El estudio terminó con la incorporación del sistema al programa regional de seguro médico obligatorio3.
El sistema se llama Third Opinion Mammograms y hace dos cosas a la vez: localiza hallazgos sospechosos en la imagen y los clasifica, analizando únicamente las imágenes4. Está registrado ante la autoridad sanitaria de Rusia, lo que equivale a una aprobación oficial de seguridad y eficacia clínica en ese país5. En total, se analizaron 404,502 mamografías de mujeres de 18 años o más1. Después de la incorporación al seguro público, la discrepancia crítica entre la lectura de la inteligencia artificial y la de los radiólogos fue de 0.08% en promedio6. Y algo notable: cuando se comparó a un radiólogo experto con los radiólogos de planta que hacían la segunda lectura, coincidieron solo en 49.9% de los casos7.
Ahora las limitaciones, que son importantes. El conjunto de imágenes usado en la calibración retrospectiva —la parte que mide si el sistema acierta— fue pequeño: 100 mamografías8. No hubo validación externa, es decir, no se probó con datos independientes de otras regiones o países8. El estudio se hizo en un solo centro, aunque ese centro coordina 238 instalaciones, y sus autores advierten que los resultados podrían no ser generalizables a otros sistemas de salud9. Tampoco se evaluó formalmente si las pacientes vivieron más o mejor, ni el impacto en la organización del trabajo o en los costos10. Y los propios autores advierten que, como se hicieron varias actualizaciones al sistema a lo largo del tiempo, no se puede aislar cuánto mejoró por cada actualización y cuánto por otros cambios simultáneos.
Para entender la magnitud del asunto: el cáncer de mama es el tumor más frecuente en mujeres y la principal causa de muerte por cáncer en esa población; por eso muchos países tienen programas de tamizaje, que buscan detectarlo temprano para reducir la mortalidad11. Pero esos programas exigen mucho recurso humano, incluyendo la doble lectura por radiólogos, y los radiólogos escasean mientras el volumen de mamografías crece12. Además, factores humanos contribuyen a errores como diagnósticos tardíos o biopsias innecesarias en mujeres sin cáncer13.
La tecnología funciona así: el sistema recibe la imagen digital, la convierte y la procesa, y produce un reporte con las zonas que considera sospechosas, más imágenes de apoyo. No emite un diagnóstico final; señala y clasifica4. Durante la fase de seguimiento, la lectura clínica final siempre la hizo un radiólogo humano, y los resultados de la inteligencia artificial estaban etiquetados como "solo para investigación". Es decir, la herramienta no reemplazó al médico: lo asistió.
Hay otros trabajos que apuntan en la misma dirección. Una revisión sistemática que solo pudimos leer en su resumen —el texto completo estaba detrás de una suscripción— reunió once estudios con aproximadamente 148,170 participantes y encontró que los sistemas de diagnóstico con inteligencia artificial mostraron mejor exactitud, sensibilidad, especificidad y eficiencia que los enfoques convencionales1415. Esa misma revisión señala que las aplicaciones en mamografía y ecografía redujeron la carga de trabajo de los radiólogos y los costos sanitarios, y mejoraron la detección, sobre todo en mujeres con mamas densas16. Pero también concluye que se necesitan más estudios de validación clínica a gran escala y de implementación en el mundo real antes de generalizar su uso17.
Otro estudio, también leído solo en su resumen, probó un dispositivo de reconocimiento mamario no invasivo en 6,817 personas en 107 hospitales y encontró una consistencia alta con las evaluaciones clínicas18. Un tercer trabajo, cuyo resumen pudimos leer pero cuya versión completa está detrás de una suscripción, evaluó un algoritmo comercial en un programa de tamizaje y halló que el rendimiento variaba considerablemente entre tres equipos de mamografía, al punto de que los autores sugieren umbrales específicos por equipo o incluso retirar el uso de la inteligencia artificial en ciertos aparatos19. Ese mismo trabajo encontró que la carga de trabajo en las reuniones de consenso bajó casi un tercio cuando se usaron umbrales ajustados por equipo20.
Así lo leemos nosotros. En un programa público de mamografías, un sistema automático podría rendir de manera desigual según el tipo de equipo, la zona o el grupo de mujeres, y esa diferencia podría no aparecer en el promedio general de aciertos. Si al separar los resultados por fabricante de equipo, región y características de las pacientes el rendimiento se mantiene parecido en todos los grupos, esa sospecha pierde fuerza. Usted puede preguntar en su centro de salud si el sistema que lee las mamografías se evalúa por separado según el equipo y la población atendida, y pedir que la lectura final siempre la haga una persona.
Así lo leemos nosotros. Los equipos que revisan el sistema una y otra vez, con casos confirmados por biopsia y con la opinión de quienes leen mamografías a diario, corregirán antes los errores que quienes confían en una sola prueba inicial. Si un sistema aprobado con una única evaluación inicial mantiene su rendimiento sin revisiones posteriores durante años, la insistencia en comprobar y volver a probar no sería necesaria. Usted puede preguntar cada cuánto se vuelve a evaluar el sistema y si las mejoras se midieron con casos reales confirmados, no solo con pruebas de laboratorio.
Y una tercera cosa que pensamos: un sistema de análisis de imágenes dentro de un programa público podría acumular datos de las pacientes y usarlos para fines que van más allá de la lectura clínica, como mejorar el producto comercial del desarrollador. Si los datos anonimizados se destruyen o se guardan bajo reglas que impiden su uso comercial, y si el proveedor no obtiene ningún beneficio adicional con ellos, la sospecha se debilita. Usted puede preguntar quién es dueño de sus imágenes, cuánto tiempo se guardan y si se usan para entrenar sistemas comerciales; puede pedir que le expliquen eso en palabras sencillas antes de firmar.
Vale la pena mirar quiénes hicieron esto y quién lo pagó. El estudio fue financiado por el Departamento de Salud de Moscú. Cuatro de los autores son empleados de la empresa que desarrolló el sistema evaluado; ellos aportaron información sobre el sistema, pero no participaron en la conducción del estudio, el análisis de datos ni la interpretación de resultados. Los autores declaran no tener otros conflictos de interés.
¿Qué haría falta para que esto significara algo más? Los propios autores piden que se evalúe la reproducibilidad de este método en otros entornos y que se aclare cómo los resultados del monitoreo pueden guiar la recalibración del modelo, la gestión de actualizaciones y las decisiones de despliegue clínico21. Y el trabajo que solo pudimos leer en resumen pide más validación clínica a gran escala antes de la implementación generalizada17.
Lo que esto hace posible para usted es concreto. Si en su clínica pública le dicen que una inteligencia artificial lee las mamografías, ahora tiene preguntas que no tenía antes: ¿quién hace la lectura final?, ¿cada cuánto se vuelve a evaluar el sistema?, ¿se probó con los equipos que usan aquí?, ¿quién guarda mis imágenes y para qué? No son preguntas técnicas; son preguntas de paciente. Y la más importante: la lectura final, ¿la firma una persona? Si la respuesta es sí, la herramienta está haciendo lo que debe hacer: asistir, no sustituir. Si la respuesta es no, usted tiene derecho a pedir que alguien la revise.
De dónde sale cada dato de contexto, y cuánto leímos de cada documento
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The test dataset comprised 404,502 mammograms from 206 medical organizations and three mammography equipment manufacturers. A total of 336 radiologists participated. The testing and monitoring period lasted 2 years and 5 months."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Over this time, AUC increased by 10.8% (from 0.83 to 0.92), accuracy by 16.9% (from 0.77 to 0.90), sensitivity by 4.8% (from 0.84 to 0.88), and specificity by 30.0% (from 0.70 to 0.91)."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The average technical defect rate decreased by 26.7% (from 3.0% to 0.8%), and the clinical assessment score rose by 47.8% (from 54.38% to 80.36%). The study culminated in the integration of the AI system into the regional CMI program."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This solution simultaneously performs computer-aided detection (CADe) to localize target findings and computer-aided diagnosis (CADx) for classifying them. Its clinical task was detecting and classifying breast cancer signs on mammograms, analyzing images only."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The software was registered with Russia’s Federal Service for Surveillance in Healthcare (Roszdravnadzor) under certificate No. 2022/16534. This regulatory approval confirms that the medical device has met national standards for safety, quality, and clinical efficacy, serving as the local equivalent of the European Conformité Européenne (CE) mark."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Following the Moscow CMI implementation, the mean rate of critical discrepancies between AI and radiologist assessments of BI-RADS categories was 0.08% (range, 0.01–0.18%)."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Notably, inter-reader agreement between the expert and end-user radiologists reached only 49.9%, which aligns with established levels of reader variability and highlights the critical need for clinical decision support systems."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "A key limitation of this study is the relatively small retrospective calibration testing dataset (100 mammograms) and the lack of external validation on independent datasets from other regions or countries."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Second, although data were aggregated through the Moscow Reference Center for Diagnostic Imaging—which oversees 238 facilities and ensures population diversity—the study’s single-center design remains a critical limitation. Consequently, the findings may not be fully generalizable to other healthcare systems and geographic regions."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Finally, the study lacked a formal assessment of the AI system’s impact on patient-centered clinical outcomes, organizational workflows, or economic efficiency. While successful integration into the CMI program was achieved, these aspects warrant separate investigation."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Breast cancer is the most prevalent malignancy among women and the leading cause of cancer-related mortality in this population ( 2 ). Consequently, numerous countries implement mammography screening programs, facilitating early detection and timely intervention to reduce mortality ( 3 )."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "However, screening entails substantial resource demands, including double-reading by radiologists ( 4 ). Unfortunately, radiologists remain scarce facing widespread overwork and burnout as mammography volumes steadily increase ( 5 , 6 )."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Human factors contribute to diagnostic errors like delayed cancer diagnoses or unnecessary biopsies in malignancy-free women ( 7 , 8 )."
- Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Eleven studies involving approximately 148,170 participants were included."
- Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "AI-driven diagnostic systems demonstrated improved accuracy, sensitivity, specificity, and efficiency compared with conventional approaches."
- Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "AI applications in mammography and ultrasound reduced radiologists' workload and healthcare costs while enhancing cancer detection rates, particularly in women with high breast density."
- Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "However, further large-scale clinical validation and real-world implementation studies are required before widespread clinical implementation."
- Zhou J, Si P, Zhang Y, Song J, He T, Liu Q, et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. 10.1038/s41467-026-73170-5 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "BIRD is applied in breast cancer screening for 6,817 individuals and shows high consistency (Cohen's kappa: 0.702 (95% confidence interval: 0.628-0.777)) with clinical assessments in real-world application across 107 hospitals."
- Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "AI-based performance varied considerably among the three mammography devices, emphasizing the need for device-specific thresholds or even the withdrawal of AI use on certain devices."
- Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Workload for radiologists in consensus conferences was lowered by nearly one-third when using device-specific thresholds compared to a general threshold (p < 0.001)."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Future studies should further assess the reproducibility of this framework in other settings and clarify how monitoring results can guide model recalibration, update management, and clinical deployment decisions."
Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864
Quién pagó: Este trabajo fue apoyado por el Departamento de Salud de Moscú como parte de un esfuerzo de investigación y desarrollo; cuatro autores son empleados de Third Opinion Platform, el desarrollador del sistema de IA evaluado, y proporcionaron información sobre el sistema pero no participaron en la realización del estudio, el análisis de datos ni la interpretación de los resultados.
No tome esto como consejo médico profesional.
experiment · Quantitative imaging in medicine and surgery · the paper, 10 Aug 2026 · free
Moscow Study: AI Mammography Software Improved Cancer Detection Accuracy After Repeated Testing
After nearly three years of testing, the software's readings came much closer to radiologists' readings. The study does not show whether women lived longer.
Short version · the longer version follows, about 7 min
Ask weeklyAI
Ask me about this study: who was studied, what it found, and what it does not say.
Conversations are saved for as long as weeklyAI exists, to improve the publication. Answers come in the language you write in.
- The study at a glance
- Who
- women aged 18 and over who had mammograms
- How many
- 404,502 mammograms
- Where
- Moscow, Russia
- When
- 2 years and 5 months
- Kind of study
- analysis of what people did
- Who did it
- Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Department of Health
- The limit that matters
- No proof women lived longer; early careful testing used only 100 mammograms.
These are the AI's scores on four reading tasks at the start and end of testing; the study did not measure whether women lived longer.
An earlier catch, and fewer false alarms, are things worth asking for. They are not things to assume.

In Moscow, women get mammograms. X-rays of the breast, looking for cancer early, when it is easier to treat.
An AI system read 404,502 of those mammograms. The images came from 206 medical organizations. Three hundred thirty-six radiologists took part.
The software, called Third Opinion Mammograms and made by Third Opinion Platform, looks for signs of breast cancer on the images and classifies what it finds. It was not simply switched on. Over two years and five months, radiologists and developers tested it, found problems, sent them back for fixes, and tested again.
The early results were not impressive. In the first tests, the AI's overall accuracy was 0.77 — meaning it got a little over three out of four cases right. Its ability to correctly clear a woman with no cancer was only 0.70. Its ability to correctly clear a woman with no cancer was only 0.70.
That was the point of the testing. Each time the software was updated, it was tested again. By the end, accuracy had risen to 0.90. Its ability to correctly clear healthy women rose to 0.91. Technical failures — images the software could not fully process — dropped from 3.0% to 0.8%.
In June 2023, the system was folded into Moscow's compulsory medical insurance program, the city's public health coverage for residents.
The AI reads the images. A radiologist makes the final call. When the AI and radiologists disagreed about a serious finding, it happened in 0.08% of cases.
This was a study in one city and one health system. The results are not guaranteed in a clinic where you live. The AI was tested on images from only three equipment manufacturers — General Electric, Fujifilm and Medical Technologies Ltd. How it would perform on other machines is unknown.
And the study did not measure whether women lived longer or had better outcomes. It measured whether the AI's readings matched radiologists' readings. That is not the same thing.
The study was also small in one respect: the careful early testing used just 100 mammograms. A bigger set would give a firmer picture. The work was funded by the Moscow Health Care Department, and four of the authors work for the company that makes the software.
The proof has limits.
If a public clinic near you starts using AI to read mammograms, you can ask: Was it tested on patients here, or somewhere else? Who makes the final call — the software or a person? And has anyone checked whether it works as well here as it did in the place where it was built?
An earlier catch, and fewer false alarms, are things worth asking for. They are not things to assume.
What this means for you
In Moscow the software got closer to radiologists' readings after repeated testing, and that is all the study shows. If a clinic near you adopts something similar, you can ask whether it was tested on patients like you, and who makes the final call. Ask a professional what it means for your own mammogram.
Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864
Who paid: This work was supported by the Moscow Health Care Department as part of a research and development effort; four authors are employees of Third Opinion Platform, the developer of the AI system evaluated, and they provided information about the system but were not involved in conducting the study, data analysis, or interpretation of results.
Do not take this as professional medical advice.
The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1370 words · about 7 minRead it →Close
An AI Read 400,000 Mammograms in Moscow. Here's What It Found — and What It Didn't.
A three-year public insurance program put a breast-cancer screening algorithm through repeated testing. Performance climbed. Whether women lived longer was never measured.

In Moscow, a computer program that reads mammograms was put through years of testing and retesting before it was allowed into the city's public insurance system, and its readings got better along the way. That is the central finding of a study published in *Quantitative Imaging in Medicine and Surgery* by Yuriy Vasilev, Denis Rumyantsev, Anton Vladzymyrskyy and colleagues at the Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Department of Health. The work was paid for by the Moscow Health Care Department.
The numbers, in plain terms. Over the testing period — 2 years and 5 months — the system's overall ability to separate cancer from non-cancer (a score called AUC) rose from 0.83 to 0.92, a 10.8 percent gain. Its accuracy rose from 0.77 to 0.90. Its ability to catch cancers that were present (sensitivity) rose from 0.84 to 0.88. Its ability to correctly clear women who did not have cancer (specificity) rose from 0.70 to 0.91 — the largest improvement of the four. The rate of technical failures fell from 3.0 percent to 0.8 percent, and a clinical assessment score assigned by expert radiologists rose from 54.38 percent to 80.36 percent12.
The scale is worth pausing on. The test dataset held 404,502 mammograms from 206 medical organizations and three equipment manufacturers; 336 radiologists took part3. The technical monitoring dataset alone covered 429,027 mammograms. The clinical monitoring dataset — the one where experts sat down and judged the AI's reads case by case — was 480 mammograms. The program-monitoring dataset after rollout covered 547,909 mammograms3.
What the study did not do is just as important. It never measured whether women actually lived longer or had better outcomes; it measured how well the AI's readings matched radiologists' readings4. The calibration dataset — the small set used to check the AI's core diagnostic performance — was only 100 mammograms5. The work was done in one city, in one health system, with images from three equipment makers, and was never validated against independent datasets from other regions or countries56. The study also did not formally assess how the AI affected patient outcomes, how it changed the way work was organized, or whether it saved money4.
To understand what this study is about, start with the problem it addresses. The article describes a study in which breast cancer is the most common cancer in women and the leading cause of cancer death in that group7. The article describes a study in which many countries run screening programs precisely because finding it early saves lives7. The article describes a study in which screening is expensive in people, not just machines: it typically requires two radiologists to read each mammogram, and radiologists are scarce, overworked and burning out as volumes rise8. The article describes a study in which human readers also make mistakes — missing cancers, or calling back women who turn out to be fine9. This is the gap AI is meant to fill.
The technology itself is not a chatbot. This system — Third Opinion Mammograms® — does two jobs at once: it finds suspicious areas on the image (computer-aided detection) and then classifies what it found (computer-aided diagnosis). It reads images only, and outputs a structured report plus marked-up pictures10. The software was registered with Russia's health regulator, Roszdravnadzor, under certificate No. 2022/16534 — the local equivalent of the European CE mark11.
The approach the Moscow team used is unusual. Rather than test the software once and declare it fit, they tested it in stages: functional testing, then calibration testing against known cases, then prospective monitoring of every mammogram at participating sites, then software updates, then repeat testing. When the first two calibration tests failed to meet the 0.80 threshold, the developers went back and the third test passed. When updates came, the relevant tests were repeated.
Here is how other work in this area reads. We also read the summary of a systematic review of AI in breast cancer detection — we could read only the summary, the full paper is behind a subscription — which gathered eleven studies involving approximately 148,170 participants and reported that AI-driven systems showed improved accuracy, sensitivity and specificity compared with conventional approaches1213. The same summary noted that AI in mammography and ultrasound reduced radiologists' workload and costs, particularly in women with dense breast tissue14, but concluded that further large-scale validation and real-world studies are still needed before widespread use15. We also read the summary of a study of a hand-held breast-recognition device, tested across 107 hospitals on 6,817 people, which showed high agreement with clinical assessments — again, we could read only the summary16.
One more piece of context matters for anyone thinking about their own clinic. We read the summary of a study of a commercially available AI algorithm used in screening, where performance varied considerably across three different mammography machines — so much so that the authors argued for device-specific thresholds, or even withdrawing the AI's use on certain devices entirely17. That study also found that using AI lowered radiologists' workload in consensus meetings by nearly one-third compared with a general threshold18. We could read only the summary; the full paper is behind a subscription.
Here is how we read it. The pattern in this study is not "AI is good at mammography." It is that a tool got better because someone kept checking it. The testing was iterative: fail, fix, retest. The updates were triggered by results. The monitoring was continuous. That is a different claim from "the algorithm works," and it is the claim the evidence actually supports.
What this leads us to expect, in homes like yours: when a clinic adopts an AI reading tool and does not say how often it is checked, updated or retested, you should expect the performance to be whatever it was on the day it was installed — no more. If the clinic does describe an ongoing monitoring plan, with retesting after updates, that is the sign the tool is being treated as something that can drift. You would know we are wrong if AI systems turned out to hold their accuracy indefinitely without retesting. We have not seen evidence of that.
What you can do with it: when you or someone you love is told an AI helped read a mammogram, you can ask two questions. Was this system tested on people like me — my age, my breast density, the machine at this clinic? And who checks it, and how often, after today? Neither question requires technical training. Both are reasonable to ask of any clinic, anywhere.
The study's own authors are direct about what remains open. They call for future work to test whether this same testing-and-monitoring approach reproduces in other settings, and to clarify how monitoring results should guide retraining and deployment decisions19. They also note that the study was done in a single center, which limits how well the findings apply to other places6.
One number from the rollout deserves attention because it reframes the whole exercise. After the AI entered the Moscow insurance program, the average rate of serious disagreement between AI and radiologist assessments of BI-RADS categories was 0.08 percent20. But the agreement between two groups of human radiologists — experts and regular readers — reached only 49.9 percent21. In other words, the humans disagreed with each other far more than the AI disagreed with the humans. The study's authors read this as evidence of the need for decision support, not as a verdict on either party21.
What this makes possible for you is not a promise that AI will catch your cancer. It is a set of questions you now know how to ask. Ask whether the system was tested on patients like you. Ask who monitors it after it goes live. Ask whether the clinic can show you the results of that monitoring. The study from Moscow is one city's answer. The question is whether your clinic has one at all.
Where each piece of context comes from, and how much of it we read
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "Over this time, AUC increased by 10.8% (from 0.83 to 0.92), accuracy by 16.9% (from 0.77 to 0.90), sensitivity by 4.8% (from 0.84 to 0.88), and specificity by 30.0% (from 0.70 to 0.91)."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "The average technical defect rate decreased by 26.7% (from 3.0% to 0.8%), and the clinical assessment score rose by 47.8% (from 54.38% to 80.36%). The study culminated in the integration of the AI system into the regional CMI program."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "The test dataset comprised 404,502 mammograms from 206 medical organizations and three mammography equipment manufacturers. A total of 336 radiologists participated. The testing and monitoring period lasted 2 years and 5 months."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "Finally, the study lacked a formal assessment of the AI system’s impact on patient-centered clinical outcomes, organizational workflows, or economic efficiency. While successful integration into the CMI program was achieved, these aspects warrant separate investigation."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "A key limitation of this study is the relatively small retrospective calibration testing dataset (100 mammograms) and the lack of external validation on independent datasets from other regions or countries."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "Second, although data were aggregated through the Moscow Reference Center for Diagnostic Imaging—which oversees 238 facilities and ensures population diversity—the study’s single-center design remains a critical limitation. Consequently, the findings may not be fully generalizable to other healthcare systems and geographic regions."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "Breast cancer is the most prevalent malignancy among women and the leading cause of cancer-related mortality in this population ( 2 ). Consequently, numerous countries implement mammography screening programs, facilitating early detection and timely intervention to reduce mortality ( 3 )."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "However, screening entails substantial resource demands, including double-reading by radiologists ( 4 ). Unfortunately, radiologists remain scarce facing widespread overwork and burnout as mammography volumes steadily increase ( 5 , 6 )."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "Human factors contribute to diagnostic errors like delayed cancer diagnoses or unnecessary biopsies in malignancy-free women ( 7 , 8 )."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "This solution simultaneously performs computer-aided detection (CADe) to localize target findings and computer-aided diagnosis (CADx) for classifying them. Its clinical task was detecting and classifying breast cancer signs on mammograms, analyzing images only."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "The software was registered with Russia’s Federal Service for Surveillance in Healthcare (Roszdravnadzor) under certificate No. 2022/16534. This regulatory approval confirms that the medical device has met national standards for safety, quality, and clinical efficacy, serving as the local equivalent of the European Conformité Européenne (CE) mark."
- Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — only the abstract - the full text could not be fetched — the passage: "Eleven studies involving approximately 148,170 participants were included."
- Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — only the abstract - the full text could not be fetched — the passage: "AI-driven diagnostic systems demonstrated improved accuracy, sensitivity, specificity, and efficiency compared with conventional approaches."
- Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — only the abstract - the full text could not be fetched — the passage: "AI applications in mammography and ultrasound reduced radiologists' workload and healthcare costs while enhancing cancer detection rates, particularly in women with high breast density."
- Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — only the abstract - the full text could not be fetched — the passage: "However, further large-scale clinical validation and real-world implementation studies are required before widespread clinical implementation."
- Zhou J, Si P, Zhang Y, Song J, He T, Liu Q, et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. 10.1038/s41467-026-73170-5 — only the abstract - the full text could not be fetched — the passage: "BIRD is applied in breast cancer screening for 6,817 individuals and shows high consistency (Cohen's kappa: 0.702 (95% confidence interval: 0.628-0.777)) with clinical assessments in real-world application across 107 hospitals."
- Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — only the abstract - the full paper is behind a subscription — the passage: "AI-based performance varied considerably among the three mammography devices, emphasizing the need for device-specific thresholds or even the withdrawal of AI use on certain devices."
- Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — only the abstract - the full paper is behind a subscription — the passage: "Workload for radiologists in consensus conferences was lowered by nearly one-third when using device-specific thresholds compared to a general threshold (p < 0.001)."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "Future studies should further assess the reproducibility of this framework in other settings and clarify how monitoring results can guide model recalibration, update management, and clinical deployment decisions."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "Following the Moscow CMI implementation, the mean rate of critical discrepancies between AI and radiologist assessments of BI-RADS categories was 0.08% (range, 0.01–0.18%)."
- Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864 - the article this story is about — the whole article — the passage: "Notably, inter-reader agreement between the expert and end-user radiologists reached only 49.9%, which aligns with established levels of reader variability and highlights the critical need for clinical decision support systems."
Vasilev, Y., Rumyantsev, D., Vladzymyrskyy, A. et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. https://doi.org/10.21037/qims-2026-0864
Who paid: This work was supported by the Moscow Health Care Department as part of a research and development effort; four authors are employees of Third Opinion Platform, the developer of the AI system evaluated, and they provided information about the system but were not involved in conducting the study, data analysis, or interpretation of results.
Do not take this as professional medical advice.
The findings of other studies mentioned here are known to us through this document, which is the one we read; we did not open each of those studies.