weeklyAI · Week of 27 September 2026weeklyAI · Semana del 27 de septiembre de 2026

← The good news← Las buenas noticias

experiment · Nature communications · la publicación, 20 may 2026 · gratis

Un programa gratuito en China clasificó a 6,817 mujeres por riesgo de cáncer de mama con ayuda de una IA

La herramienta también mejoró la lectura de las ecografías por parte de los médicos, pero fuera de China su rendimiento bajó, y el seguimiento se centró en las mujeres de mayor riesgo.

Versión breve · la versión detallada sigue, unos 9 min

Pregunte a weeklyAI

Pregúnteme por este estudio: a quiénes se estudió, qué encontró y qué no dice.

Las conversaciones se guardan mientras exista weeklyAI, para mejorar la publicación. Se responde en el idioma en que usted escribe.

El estudio, de un vistazo
Quiénes
mujeres que se hicieron una ecografía de mama
Cuántos
6,817
Dónde
China
Cuándo
no lo dice el pasaje
Tipo de estudio
experiment
Quién lo hizo
hospitales chinos
El límite que importa
El seguimiento se centró en las mujeres de riesgo alto; las de riesgo bajo no tuvieron ese seguimiento.
Cómo quedaron clasificadas por riesgo las 6,817 mujeres del programa de detección
riesgo bajo78.82%
riesgo medio16.77%
riesgo alto4.42%

Son los porcentajes de las 6,817 mujeres clasificadas por el programa de detección en China; el seguimiento posterior se concentró en el grupo de riesgo alto, así que el estudio no midió cuántos cánceres pudieron haberse pasado por alto en el grupo de riesgo bajo.

Precisión de BIRD dentro y fuera de China

prueba interna en Chinafrente aconjunto de imágenes de Egipto

0,837 frente a 0,661

prueba interna en Chinafrente aconjunto de imágenes de España

0,837 frente a 0,604

Un resultado normal de la IA, en este estudio, no es prueba de que no haya cáncer.
Lectura de weeklyAI
Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

En diez sitios de China, 6,817 mujeres se hicieron una ecografía de mama sin pagar, como parte de un programa de detección temprana. Un sistema de inteligencia artificial llamado BIRD revisaba cada imagen y las ordenaba en tres grupos: riesgo bajo, medio y alto. De todas ellas, 5,373 quedaron en riesgo bajo, 1,143 en riesgo medio y 301 en riesgo alto. Esa clasificación sirvió para que los médicos decidieran a quién seguir de cerca y a quién no.

El sistema es un programa de inteligencia artificial. Se desarrolló con 18,201 imágenes de entrenamiento del Hospital Central de Dazhou, en Sichuan, China. Después se puso a prueba con imágenes internas, con imágenes de ocho hospitales chinos y con dos conjuntos internacionales, uno de Egipto y otro de España. Además, se hicieron tres estudios en los que radiólogos leyeron ecografías con y sin la ayuda del programa.

En la prueba interna, BIRD acertó en 0,837 de los casos, es decir, poco más de ocho de cada diez. En los ocho hospitales chinos, el porcentaje de aciertos fue de 0,708 a 0,913. En dos de esos estudios con radiólogos, el programa mejoró la precisión de los médicos de manera medible (P < 0,05). En uno de ellos, un grupo de radiólogos en formación pasó de 0,687 a 0,736 de aciertos cuando trabajó con la ayuda del programa.

El seguimiento del programa de detección se concentró en las mujeres clasificadas como de riesgo alto: a ellas se les hizo seguimiento durante unos trece meses, y tres fueron diagnosticadas con cáncer de mama. Las mujeres de riesgo bajo no recibieron ese mismo seguimiento, así que el estudio no midió cuántos cánceres pudieron haberse pasado por alto en ese grupo. Un resultado normal de la IA, en este estudio, no es prueba de que no haya cáncer.

Cuando BIRD se aplicó a los conjuntos de imágenes de Egipto y España, su precisión cayó a 0,661 y 0,604, bastante por debajo de lo que logró en China. Los propios autores escriben que el sistema necesita más validación antes de usarse clínicamente fuera del país donde se desarrolló y se validó.

El estudio también desarrolló otros modelos que, a partir de la misma ecografía, intentan distinguir entre lesiones benignas y malignas, entre cáncer invasor y un tipo no invasor llamado DCIS, y entre subtipos moleculares del tumor. Sus resultados de clasificación fueron de 0,753 a 0,965. Son modelos de investigación, no herramientas que usted pueda pedir hoy.

Lo que este trabajo muestra, medido y con límites, es que la ecografía asistida por IA puede ordenar a muchas mujeres por riesgo y ayudar al médico que lee la imagen, en los hospitales chinos donde se probó. Si en su país se propone un programa así, pregunte si fue evaluado en su población, quién revisa los resultados y cómo se hace el seguimiento de las mujeres clasificadas como de riesgo bajo.

Qué significa para usted

Para usted, que solo oye la alarma, lo medido es esto: en hospitales de China, esa asistencia mejoró la lectura de los médicos y ordenó por riesgo a miles de mujeres, aunque el seguimiento se centró en las de riesgo alto. Si aquí proponen algo parecido, pregunte si se evaluó en su población y cómo se sigue a las de riesgo bajo.

Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5

Quién pagó: El estudio fue financiado por el Programa Nacional Clave de Investigación y Desarrollo de China, la Comisión de Salud de la Provincia de Henan, la Comisión de Salud de la Provincia de Sichuan, el Departamento de Ciencia y Tecnología de la Provincia de Sichuan y un proyecto especial de cooperación entre la ciudad de Dazhou y universidades; el artículo no indica si los financiadores tuvieron alguna influencia, y no se mencionan prestadores de equipos o software.

No tome esto como consejo médico profesional.

Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1723 palabras · unos 9 minLeerla →Cerrar

La IA que ayudó a leer ecografías de mama en China: lo que el estudio midió y lo que todavía no

Un programa gratuito clasificó por riesgo a 6,817 mujeres y mejoró la lectura de los radiólogos en pruebas controladas. Fuera de China, la misma herramienta todavía no está probada.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

El cáncer de mama es el tumor maligno más frecuente entre las mujeres en el mundo, y detectarlo temprano es la base para mejorar la supervivencia1. En mujeres asiáticas, cuyo tejido mamario suele ser denso, la mamografía pierde sensibilidad y la ecografía se ha vuelto el método principal de tamizaje2. Pero leer una ecografía depende mucho de quién la lee: aunque existe un sistema estandarizado de categorías, llamado BI-RADS, la clasificación sigue siendo subjetiva y varía de un lector a otro, lo que puede llevar a biopsias innecesarias o a diagnósticos tardíos3.

Para atacar ese problema, un grupo de investigadores de hospitales chinos desarrolló y probó un sistema de inteligencia artificial llamado BIRD, pensado para mujeres y evaluado en instituciones y poblaciones diversas4. El sistema usa un algoritmo llamado YOLOv5 para identificar las lesiones en la imagen de ecografía y un algoritmo llamado EfficientNet para asignar a cada lesión identificada su categoría BI-RADS5. En su conjunto de prueba interno, la precisión de BIRD fue de 0,837, y en dos estudios de lectura mejoró de manera significativa la precisión de los radiólogos6.

El programa también se llevó a la práctica. BIRD se aplicó en tamizaje de cáncer de mama a 6,817 mujeres y mostró una consistencia alta con las evaluaciones clínicas en un despliegue real en 107 hospitales7. De esas mujeres, 5,373 (el 78.82%) quedaron clasificadas como riesgo bajo, 1,143 (16.77%) como riesgo medio y 301 (4.42%) como riesgo alto8. En un seguimiento posterior del grupo de riesgo alto, con una mediana de 13 meses, se diagnosticaron tres casos de cáncer de mama9. La clasificación por riesgo orientó la decisión clínica y permitió a los médicos armar planes de diagnóstico y tratamiento según el nivel de cada paciente10. El propio artículo lo resume así: el tamizaje derivó de forma segura a las 5,373 mujeres de riesgo bajo mediante informes automáticos, alivió en parte la carga de trabajo de los radiólogos y confirmó tres cánceres en el grupo de riesgo alto11.

Aquí conviene frenar. El programa de tamizaje siguió sobre todo al grupo de riesgo alto, lo cual es razonable para un estudio de viabilidad, pero limita la caracterización completa del desempeño y deja claro que hace falta un ensayo formal de tamizaje con seguimiento completo12. Las variaciones de rendimiento en los conjuntos internacionales indican que BIRD necesita validación adicional antes de usarse clínicamente fuera del ámbito donde se desarrolló y se validó dentro de China13. Y las métricas definitivas de eficacia tendrán que establecerse en futuros ensayos controlados aleatorizados14.

Vale la pena mirar este trabajo dentro de una familia de investigaciones. Una revisión sobre inteligencia artificial en cáncer de mama, de la que solo pudimos leer el resumen porque el texto completo no estaba disponible, señala que el cáncer de mama sigue siendo una de las principales causas de muerte por cáncer en el mundo y que la detección temprana es esencial para mejorar la supervivencia y los resultados del tratamiento15. Esa misma revisión indica que la inteligencia artificial, en particular el aprendizaje profundo, se ha vuelto una herramienta prometedora para mejorar la precisión y la eficiencia del diagnóstico16, y que los sistemas diagnósticos basados en IA mostraron mejor precisión, sensibilidad, especificidad y eficiencia frente a los enfoques convencionales17. Según ese resumen, las aplicaciones en mamografía y ecografía redujeron la carga de trabajo de los radiólogos y los costos de salud, y mejoraron la detección sobre todo en mujeres con mama densa18. La misma revisión advierte que hacen falta estudios de validación clínica a gran escala y de implementación en el mundo real antes de una adopción generalizada19.

Otros trabajos que leímos, también solo en su resumen, muestran que este camino tiene obstáculos concretos. Un estudio sobre un sistema de IA para mamografía, cuyo texto completo estaba detrás de una suscripción, cuenta que ese sistema se entrenó con unos 4,000 mamogramas20 y que se probó con 404,502 mamogramas de 206 organizaciones médicas y tres fabricantes de equipos, con 336 radiólogos participantes21. Según ese resumen, con el tiempo la precisión del sistema subió y los defectos técnicos bajaron2223, y el trabajo terminó integrando la herramienta en un programa regional de tamizaje24. Pero sus propios autores señalan como limitación clave que la calibración retrospectiva se hizo con un conjunto pequeño, de 100 mamogramas, y sin validación externa en datos de otras regiones o países25. Además, concluyen que las pruebas iterativas con monitoreo prospectivo en el mundo real, actualizaciones del desarrollador y retroalimentación de los radiólogos mejoraron sustancialmente el rendimiento26, y que esa metodología de ciclo de vida demuestra que es factible integrar y desplegar este tipo de sistemas equilibrando validación rigurosa y mejora continua27.

Un tercer trabajo, igualmente leído solo en su resumen, muestra hasta qué punto el equipo importa. Ese estudio evaluó un algoritmo comercial que da a cada mamografía un puntaje numérico28. La lectura doble estándar detectó 5,07 cánceres por cada mil mujeres examinadas; con IA y un umbral general, la detección fue de 4,81; y con umbrales ajustados a cada equipo, de 4,9029. La carga de trabajo en las reuniones de consenso bajó casi un tercio al usar umbrales específicos por equipo en lugar de un umbral general30. Sin embargo, el rendimiento fue pobre en uno de los equipos incluso con umbrales ajustados31. Los autores concluyen que el rendimiento varió considerablemente entre los tres equipos y que esto exige umbrales específicos por dispositivo o incluso retirar el uso de la IA en ciertos aparatos32, y que los datos de entrenamiento deben representar no solo a la población objetivo sino también a los distintos equipos usados en la práctica33.

Así lo leemos nosotros. La pregunta que este trabajo deja abierta no es si la máquina acierta, sino qué pasa con la mujer que recibe un resultado. Cuando una herramienta automática ordena a las personas en categorías de riesgo, quienes reciben esa decisión no ven cómo se tomó ni tienen un camino claro para preguntar por qué les tocó esa etiqueta. Es razonable esperar que en un programa de tamizaje con miles de participantes muchas mujeres hayan recibido su clasificación sin una explicación en palabras simples ni un canal sencillo para pedir una segunda mirada. Sabríamos que nos equivocamos si el programa hubiera entregado a cada mujer una explicación comprensible de su categoría y un mecanismo accesible y sin costo para pedir revisión humana. Lo que usted puede hacer con esto: cuando una herramienta automática le dé un resultado de salud, pida que le expliquen en palabras simples por qué le tocó esa categoría y quién revisa el resultado, y pregunte si hay una persona a la que pueda pedir una segunda opinión.

Hay algo más que nos parece importante. La gente común depende de aparatos cuyo funcionamiento no entiende, y esa dependencia puede sentirse como comodidad frente a la pantalla cuando en realidad es desconocimiento de cómo opera el sistema. Es probable que muchas participantes hayan recibido un informe sin saber qué miraba el sistema, en qué se basaba su puntaje de riesgo ni qué pasaría con sus imágenes después. Sabríamos que estamos equivocados si las participantes entendían, en términos sencillos, qué hacía la IA, qué significaba su nivel de riesgo y qué se haría con su información, y si ese entendimiento se hubiera medido o documentado. Lo que usted puede hacer: antes de aceptar un tamizaje o estudio con IA, pregunte qué parte del resultado la dio la máquina y qué parte una persona, y pregunte también quién guarda sus imágenes y por cuánto tiempo.

Y una tercera cosa, que es la que más nos importa para su casa. El buen juicio se educa practicando la duda y la verificación, y no basta con tener razón en el punto de partida si uno no revisa. Es de esperar que los equipos que solo confiaron en el puntaje automático, sin compararlo con su propia lectura clínica, hayan dejado pasar casos dudosos, y que los que mantuvieron la duda y revisaron caso por caso hayan detectado mejor las discrepancias. Sabríamos que estamos equivocados si el programa hubiera medido y reportado cuántas veces el criterio humano cambió la clasificación de la IA, y si esa discrepancia se hubiera usado para revisar casos y no solo para calcular un porcentaje de acuerdo. Lo que usted puede hacer: cuando un resultado automático le parezca raro o distinto a lo que siente, pida que una persona lo revise, y pregunte si el centro compara sistemáticamente lo que dice la máquina con lo que dice el personal de salud.

Sobre quién pagó y qué intereses hay en juego: el trabajo fue financiado por el Programa Nacional de Investigación y Desarrollo de China, la Comisión de Salud de la provincia de Henan, el Departamento de Ciencia y Tecnología de Henan, el Programa de Ciencia y Tecnología Médica de la Comisión de Salud de Sichuan, el fondo de Proyectos Clave del Departamento de Ciencia y Tecnología de Sichuan y un proyecto especial de cooperación de la ciudad de Dazhou. El artículo declara que todos los demás autores no tienen conflictos de interés34.

Lo que esto hace posible para usted, hoy. Este sistema no está disponible ni validado en América Latina, Estados Unidos ni Canadá: ninguna clínica de nuestra región lo ha probado, y el propio estudio dice que necesita más validación fuera de China. Lo que sí puede llevarse a casa es la pregunta, no el aparato. La próxima vez que en un control le digan que "la inteligencia artificial revisó su estudio", pregunte tres cosas: qué parte del resultado la dio la máquina y qué parte una persona, qué se hace con las imágenes después, y quién revisa el resultado si usted quiere una segunda mirada. Y si le dicen que todo está bien porque la máquina lo dijo, recuerde que en este estudio la herramienta no siguió de cerca a las mujeres clasificadas como riesgo bajo, de modo que un resultado normal no es prueba de que no haya nada. ¿Puede pedir hoy, en su centro de salud, que le expliquen en palabras simples quién leyó su estudio y cómo?

De dónde sale cada dato de contexto, y cuánto leímos de cada documento

  1. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Breast cancer is the most prevalent malignancy among women worldwide 1 , 2 , with early detection being the cornerstone of improved survival."
  2. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "In Asian women, who tend to have dense breast tissue, mammography’s sensitivity is markedly reduced, and ultrasound has emerged as the primary screening modality 3 ."
  3. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "However, ultrasound interpretation is highly operator-dependent, and although the Breast Imaging Reporting and Data System (BI-RADS) classification is standardized, it remains subjective 4 , 5 , leading to inter-reader variability that may cause unnecessary biopsies or delayed diagnosis."
  4. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Here, we develop and test an end-to-end breast intelligent recognition device (BIRD) for women across diverse institution/population datasets."
  5. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The BIRD utilized the YOLOv5 localization algorithm to identify lesions in breast ultrasound images and the EfficientNet classification algorithm for BI-RADS classification of the identified lesions"
  6. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The accuracy of BIRD in the internal test set is 0.837 (95% confidence interval: 0.827–0.846), which significantly improves radiologists’ accuracy in 2 reader studies ( P  < 0.05)."
  7. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "BIRD is applied in breast cancer screening for 6,817 individuals and shows high consistency (Cohen’s kappa: 0.702 (95% confidence interval: 0.628-0.777)) with clinical assessments in real-world application across 107 hospitals."
  8. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Among them, 5373 individuals (78.82%) were classified as low risk (BI-RADS 1-2), 1143 (16.77%) as medium risk (BI-RADS 0 and 3), and 301 (4.42%) as high risk (BI-RADS 4-5) (Fig.  4b )."
  9. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "A subsequent follow-up of the high-risk group ( n  = 301), which had a median follow-up of 13 months, found that 3 patients were diagnosed with breast cancer."
  10. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The BIRD risk stratification offered crucial guidance for clinical decision-making, which enabled doctors to formulate diagnostic and treatment plans according to different risk levels."
  11. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "breast cancer screening of 6817 individuals safely triaged 5373 low-risk cases (78.8%) through AI-automated reporting, partially alleviating radiologist workload while confirming 3 cancers in the high-risk cohort (4.4% prevalence), thereby demonstrating pragmatic utility beyond accuracy metrics alone."
  12. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The 10-site screening program employed risk-stratified follow-up focused on high-risk participants, which is appropriate for feasibility assessment but limits comprehensive performance characterization, highlighting the distinct need for a formal screening trial with complete follow-up."
  13. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Performance variations across international datasets indicate that BIRD requires further validation before clinical deployment outside the development and domestic validation setting."
  14. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "definitive efficacy metrics will be established through future randomized controlled trials."
  15. Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Breast cancer (BC) remains one of the leading causes of cancer-related deaths worldwide, with early detection being essential for improving survival rates, treatment outcomes, and preventive women's healthcare strategies."
  16. Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Artificial Intelligence (AI), particularly deep learning (DL) and machine learning (ML) algorithms, has emerged as a promising tool for improving the accuracy and efficiency of BC diagnosis."
  17. Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "AI-driven diagnostic systems demonstrated improved accuracy, sensitivity, specificity, and efficiency compared with conventional approaches."
  18. Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "AI applications in mammography and ultrasound reduced radiologists' workload and healthcare costs while enhancing cancer detection rates, particularly in women with high breast density."
  19. Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "However, further large-scale clinical validation and real-world implementation studies are required before widespread clinical implementation."
  20. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "The mammography AI system utilized U-Net++ and Mask2Former architectures, trained on ~4,000 mammograms."
  21. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "The test dataset comprised 404,502 mammograms from 206 medical organizations and three mammography equipment manufacturers. A total of 336 radiologists participated."
  22. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Over this time, AUC increased by 10.8% (from 0.83 to 0.92), accuracy by 16.9% (from 0.77 to 0.90), sensitivity by 4.8% (from 0.84 to 0.88), and specificity by 30.0% (from 0.70 to 0.91)."
  23. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "The average technical defect rate decreased by 26.7% (from 3.0% to 0.8%), and the clinical assessment score rose by 47.8% (from 54.38% to 80.36%)."
  24. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "The study culminated in the integration of the AI system into the regional CMI program."
  25. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "A key limitation of this study is the relatively small retrospective calibration testing dataset (100 mammograms) and the lack of external validation on independent datasets from other regions or countries."
  26. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Iterative testing with prospective real-world monitoring, interleaved developer updates, and radiologist feedback substantially enhanced mammography AI performance."
  27. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "This lifecycle testing methodology demonstrates feasibility for clinical integration and CMI program deployment, balancing rigorous validation with continuous improvement."
  28. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "A commercially available AI algorithm assessed each mammography with a numeric case score. Optimal thresholds were determined across all devices and for each device separately using the Youden-Index."
  29. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Standard double-reading had a CDR of 5.07 (95% CI: 4.19-6.09). Using AI with a general threshold achieved a CDR of 4.81 (95% CI: 3.95-5.80), which was enhanced using device-specific thresholds to 4.90 (95% CI: 4.03-5.89)."
  30. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Workload for radiologists in consensus conferences was lowered by nearly one-third when using device-specific thresholds compared to a general threshold (p < 0.001)."
  31. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Nonetheless, AI screening performance for one of the mammography devices was poor even with device-specific thresholds."
  32. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "AI-based performance varied considerably among the three mammography devices, emphasizing the need for device-specific thresholds or even the withdrawal of AI use on certain devices."
  33. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Lastly, training datasets must be representative not only of the target population but also of the different screening devices used in practice."
  34. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - el artículo del que trata esta nota — el artículo completo — el pasaje: "All other authors have no competing interests."

Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5

Quién pagó: El estudio fue financiado por el Programa Nacional Clave de Investigación y Desarrollo de China, la Comisión de Salud de la Provincia de Henan, la Comisión de Salud de la Provincia de Sichuan, el Departamento de Ciencia y Tecnología de la Provincia de Sichuan y un proyecto especial de cooperación entre la ciudad de Dazhou y universidades; el artículo no indica si los financiadores tuvieron alguna influencia, y no se mencionan prestadores de equipos o software.

No tome esto como consejo médico profesional.

experiment · Nature communications · the paper, 20 May 2026 · free

Free Breast Scans in China, Sorted by AI — With One Number Still Missing

A 6,817-woman screening program used software to rank risk. The follow-up did not check everyone.

Short version · the longer version follows, about 8 min

Ask weeklyAI

Ask me about this study: who was studied, what it found, and what it does not say.

Conversations are saved for as long as weeklyAI exists, to improve the publication. Answers come in the language you write in.

The study at a glance
Who
women screened for breast cancer with ultrasound
How many
6,817 individuals
Where
China
When
March 2022 to May 2023
Kind of study
experiment
Who did it
Hospitals and universities across China
The limit that matters
The follow-up checked mainly high-risk women, not the thousands told they were low risk.
How the software sorted 6,817 women by risk
low risk78.82%
medium risk16.77%
high risk4.42%

Share of 6,817 screened women placed in each risk group; the follow-up then checked mainly the high-risk group, so what happened to the low-risk women was not fully measured.

Radiologists reading scans with and without the software

radiologists reading aloneagainstradiologists reading with the software

accuracy rose from about 69 in 100 to about 74 in 100 in one group of trainees

A low-risk sorting is not proof that nothing is there.
weeklyAI's reading
How it could look · illustration generated by weeklyAI.watch, not a photograph

In ten places across China, 6,817 women received free breast ultrasound screening. A computer program called BIRD looked at each scan and sorted the women into three groups: low, medium and high risk. Low risk held 5,373 women — nearly 79 percent. High risk held 301.

Doctors used those groups to decide what to do next. They followed the high-risk women for a median of 13 months. Three were diagnosed with breast cancer.

The screening ran from March 2022 to May 2023, and the study was published in May 2026 in *Nature Communications*.

BIRD was built using ultrasound images from Dazhou Central Hospital, split into training, optimization and test sets. Researchers then tested it elsewhere: eight other hospitals in China, plus public datasets from Egypt and Spain. They also ran three studies in which 19 radiologists read scans with and without the software.

On the hospital's own test images, BIRD matched the correct category 0.837 of the time — about 84 in 100. Across the eight other Chinese hospitals, that figure ranged from about 71 in 100 to about 91 in 100.

In a study at Dazhou with 1,000 images, four trainee radiologists read scans twice, 20 days apart. The second time, BIRD's reading sat beside theirs. One group's accuracy rose from about 69 in 100 to about 74 in 100. In a separate study of 527 images, radiologists also read more accurately with the software than without it.

Now the part the screening numbers do not cover. The follow-up focused on the high-risk women. How many cancers were missed among the thousands told they were low risk was not fully measured. A low-risk sorting is not proof that nothing is there. The researchers call the screening a feasibility study, not a completed trial.

The international results point the same way. On the Egypt dataset, BIRD's accuracy was 0.661 — about 66 in 100. On the Spain dataset, 0.604 — about 60 in 100. The authors write that BIRD needs more testing before it is used outside the hospitals where it was developed and checked in China.

The study also built three related models: one to tell benign from malignant growths, one to tell invasive cancer from an early, non-invasive form, and one to guess a tumor's molecular subtype. Their scores ranged from 0.753 to 0.965.

No clinic in Latin America, the United States or Canada tested BIRD. None of this makes it available or approved there.

What the study does show is a working picture of what AI-assisted ultrasound could offer in places where mammography reads dense breast tissue poorly — earlier sorting, a second set of eyes for a trainee. That picture has to be tested where you live before it means anything for you.

The question to carry into your next appointment — in any country — is not whether a computer looked at the image, but whether a doctor did, and what happens to the women whose scan comes back clear.

What this means for you

For you, the practical weight of this is small and specific: no clinic outside China tested this software, so it is not something to ask your own doctor about yet. What you can watch for is whether ultrasound results in your area come with a clear plan for the women told nothing was found.

Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5

Who paid: The study was funded by the National Key Research and Development Program of China, the Health Commission of Henan Province, the Health Commission of Sichuan Province, the Science and Technology Department of Sichuan Province, and a Dazhou City School Cooperation Special project; the article does not state whether funders had any say, and no equipment or software lenders are mentioned.

Do not take this as professional medical advice.

The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1644 words · about 8 minRead it →Close

A Free Breast Screening Program in China Used AI to Sort 6,817 Women by Risk. Here's What It Actually Measured.

The system helped doctors read scans more accurately in controlled tests. Outside the hospitals that built it, its accuracy dropped sharply — and the screening program never checked what happened to the women it called low risk.

How it could look · illustration generated by weeklyAI.watch, not a photograph

Breast cancer is the most common cancer among women worldwide, and finding it early is the cornerstone of surviving it1. In Asian women, who tend to have dense breast tissue, mammography often misses tumors, so ultrasound has become the primary screening tool2. But reading an ultrasound is highly dependent on who is holding the probe, and even the standardized reporting system doctors use remains subjective — leading to differences between readers that can cause unnecessary biopsies or delayed diagnosis3.

A team of researchers across China developed and tested an end-to-end breast intelligent recognition device, or BIRD, for women across diverse institutions and populations4. The system uses one algorithm — YOLOv5 — to find lesions in breast ultrasound images, and another — EfficientNet — to classify those lesions according to BI-RADS, the standard scale radiologists use5. It runs in real time beside an ultrasound machine and gives each lesion a BI-RADS classification.

On the internal test set, BIRD's accuracy was 0.837 — meaning about 84 out of 100 images were classified correctly. The article reports that this accuracy significantly improved radiologists' performance in two reader studies6. In a prospective reader study of 527 images from three hospitals, six radiologists read the same scans with and without BIRD's help. When BIRD was correct, the accuracy of the six radiologists improved. When BIRD was wrong, three of them got worse.

The system was then deployed across 107 hospitals in China. In a consistency analysis of 38 hospitals and 9,320 images, the average agreement between BIRD and local radiologists was 0.702 on Cohen's kappa — a measure of how often two readers agree beyond chance7. That is substantial agreement, but the article is explicit: this reflects implementation concordance, not diagnostic accuracy.

In a separate screening program at 10 sites, BIRD sorted 6,817 women into risk groups. Of them, 5,373 (78.82%) were classified as low risk, 1,143 (16.77%) as medium risk, and 301 (4.42%) as high risk8. The high-risk group was followed for a median of 13 months; three patients were diagnosed with breast cancer9. The authors say the risk stratification offered crucial guidance for clinical decision-making, enabling doctors to formulate diagnostic and treatment plans according to different risk levels10. They also describe the screening as safely triaging the low-risk cases through AI-automated reporting, partially alleviating radiologist workload, and confirming three cancers in the high-risk cohort11.

Here is the limit that matters most. The screening program followed up mainly the high-risk group. The women told they were low risk were not systematically checked again. The article states that this risk-stratified follow-up focused on high-risk participants is appropriate for feasibility assessment but limits comprehensive performance characterization, and highlights the need for a formal screening trial with complete follow-up12. It also notes that definitive efficacy metrics will be established through future randomized controlled trials13.

On international datasets — one from Egypt, one from Spain — BIRD's accuracy dropped to 0.661 and 0.604. The article states plainly that performance variations across international datasets indicate BIRD requires further validation before clinical deployment outside the development and domestic validation setting14. The study was supported by Chinese government research programs and health commissions.

This study belongs to a family of efforts to bring AI into breast imaging, and the pattern across that family is consistent: promise in controlled settings, caution before routine use. According to the summary of a review of AI in breast cancer diagnosis — we could read only the summary, the full paper is behind a subscription — AI-driven diagnostic systems have demonstrated improved accuracy, sensitivity, specificity, and efficiency compared with conventional approaches15. The same summary reports that AI applications in mammography and ultrasound reduced radiologists' workload and healthcare costs while enhancing cancer detection rates, particularly in women with high breast density16. But it also states that further large-scale clinical validation and real-world implementation studies are required before widespread clinical implementation17.

Another study, this one in Russia, took a different approach to the same problem. According to its summary — we could read only the summary, the full paper is behind a subscription — researchers integrated an AI system into a regional mammography program and tracked its performance over time. The test dataset comprised 404,502 mammograms from 206 medical organizations and three equipment manufacturers, with 336 radiologists participating18. Over the study period, the system's accuracy improved from 0.77 to 0.90, and sensitivity from 0.84 to 0.8819. The technical defect rate dropped from 3.0% to 0.8%, and the clinical assessment score rose from 54.38% to 80.36%20. The study culminated in the integration of the AI system into the regional program21. But a key limitation, the authors note, is the relatively small retrospective calibration testing dataset and the lack of external validation on independent datasets from other regions or countries22.

A third study, from Switzerland, tested a commercial AI algorithm on mammograms and found that performance varied considerably among three different mammography devices. According to its summary — we could read only the summary, the full paper is behind a subscription — standard double-reading had a cancer detection rate of 5.07 per thousand, while AI with a general threshold achieved 4.81, improved to 4.90 with device-specific thresholds23. Workload for radiologists in consensus conferences was lowered by nearly one-third when using device-specific thresholds24. But the authors also report that AI screening performance for one of the devices was poor even with device-specific thresholds25, and they emphasize the need for device-specific thresholds or even the withdrawal of AI use on certain devices26. They add that training datasets must be representative not only of the target population but also of the different screening devices used in practice27.

Here is how we read it. The pattern across these studies is not that AI fails or that AI works. The pattern is that a tool built and tested in one kind of place tends to perform worse in another — and the drop shows up in the people who were never in the original sample. The BIRD study is unusually honest about this: it reports the drop on international data, states that further validation is needed outside China, and says the screening program's follow-up was not designed to catch what it missed. What we expect, in homes like yours, is that a clinic adopting a tool like this will be able to show you the accuracy numbers from the hospital that built it, and will not be able to show you what happened to the people it called low risk. We would be wrong if the screening program later publishes complete follow-up data showing that the low-risk group was tracked and that missed cancers were rare. What you can do with this: at any clinic that uses a computer to help read your scan, ask who reviews the computer's sorting and what happens to the people placed in the lowest group. Ask whether you can be told, in words you understand, why you were put in one group rather than another. If you are told you are low risk by a machine-assisted reading, ask what follow-up is planned and when you should come back.

We also read something else in these numbers. When a doctor and a machine agree, it is tempting to treat that agreement as proof the machine is right. But agreement only means they said the same thing. The BIRD study measured agreement between the tool and radiologists across 38 hospitals, and the authors themselves note this reflects concordance rather than accuracy. In the reader studies, when BIRD was wrong, three radiologists got worse — they followed the machine into its mistake. The worry is not that this happened once in a study. The worry is that it grows with familiarity. What would show we are wrong: if hospitals tracked the cases where the tool was wrong and measured whether doctors caught them anyway, and if that catch rate held steady over time. What you can ask: when a computer gives your clinician a result, ask what the clinician saw with their own eyes and whether they disagree. Ask whether the clinic tracks the cases where the computer was wrong, not only the cases where it was right.

For a reader in Latin America, the United States, or Canada, the practical situation is this: no clinic in your country has tested BIRD. The study was conducted entirely in China, with Chinese hospitals, Chinese radiologists, and Chinese populations. The international datasets that were tested — from Egypt and Spain — showed lower accuracy, and the authors state that the system requires further validation before clinical deployment outside its development and domestic validation setting14. This does not mean the technology is not coming. It means that if you hear about a similar system being used where you live, the questions above are the ones worth asking.

The last thing we would say is this. The study's most useful contribution may not be its accuracy numbers. It may be that it shows what a serious evaluation looks like: reader studies with and without the tool, a randomized comparison, deployment across more than a hundred hospitals, and a public accounting of where the tool fell short. That is not the norm. Most AI tools reach clinics with far less. If a hospital near you adopts one, you now know what to ask for — and what to ask about.

What to ask at your next breast ultrasound: If a computer helped read your scan, who reviews its sorting, what happens to the women it calls low risk, and when should you come back?

Where each piece of context comes from, and how much of it we read

  1. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "Breast cancer is the most prevalent malignancy among women worldwide 1 , 2 , with early detection being the cornerstone of improved survival."
  2. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "In Asian women, who tend to have dense breast tissue, mammography’s sensitivity is markedly reduced, and ultrasound has emerged as the primary screening modality 3 ."
  3. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "However, ultrasound interpretation is highly operator-dependent, and although the Breast Imaging Reporting and Data System (BI-RADS) classification is standardized, it remains subjective 4 , 5 , leading to inter-reader variability that may cause unnecessary biopsies or delayed diagnosis."
  4. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "Here, we develop and test an end-to-end breast intelligent recognition device (BIRD) for women across diverse institution/population datasets."
  5. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "The BIRD utilized the YOLOv5 localization algorithm to identify lesions in breast ultrasound images and the EfficientNet classification algorithm for BI-RADS classification of the identified lesions"
  6. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "The accuracy of BIRD in the internal test set is 0.837 (95% confidence interval: 0.827–0.846), which significantly improves radiologists’ accuracy in 2 reader studies ( P  < 0.05)."
  7. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "BIRD is applied in breast cancer screening for 6,817 individuals and shows high consistency (Cohen’s kappa: 0.702 (95% confidence interval: 0.628-0.777)) with clinical assessments in real-world application across 107 hospitals."
  8. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "Among them, 5373 individuals (78.82%) were classified as low risk (BI-RADS 1-2), 1143 (16.77%) as medium risk (BI-RADS 0 and 3), and 301 (4.42%) as high risk (BI-RADS 4-5) (Fig.  4b )."
  9. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "A subsequent follow-up of the high-risk group ( n  = 301), which had a median follow-up of 13 months, found that 3 patients were diagnosed with breast cancer."
  10. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "The BIRD risk stratification offered crucial guidance for clinical decision-making, which enabled doctors to formulate diagnostic and treatment plans according to different risk levels."
  11. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "breast cancer screening of 6817 individuals safely triaged 5373 low-risk cases (78.8%) through AI-automated reporting, partially alleviating radiologist workload while confirming 3 cancers in the high-risk cohort (4.4% prevalence), thereby demonstrating pragmatic utility beyond accuracy metrics alone."
  12. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "The 10-site screening program employed risk-stratified follow-up focused on high-risk participants, which is appropriate for feasibility assessment but limits comprehensive performance characterization, highlighting the distinct need for a formal screening trial with complete follow-up."
  13. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "definitive efficacy metrics will be established through future randomized controlled trials."
  14. Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5 - the article this story is about — the whole article — the passage: "Performance variations across international datasets indicate that BIRD requires further validation before clinical deployment outside the development and domestic validation setting."
  15. Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — only the abstract - the full text could not be fetched — the passage: "AI-driven diagnostic systems demonstrated improved accuracy, sensitivity, specificity, and efficiency compared with conventional approaches."
  16. Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — only the abstract - the full text could not be fetched — the passage: "AI applications in mammography and ultrasound reduced radiologists' workload and healthcare costs while enhancing cancer detection rates, particularly in women with high breast density."
  17. Bothou A, Bolou A, Dinas K, Kyrkou G, Hardy D, Pappou P, et al. (2026). Artificial Intelligence in Early Breast Cancer Detection: A Systematic Review of Innovations in Preventive Women’s Healthcare. Healthcare. 10.3390/healthcare14121674 — only the abstract - the full text could not be fetched — the passage: "However, further large-scale clinical validation and real-world implementation studies are required before widespread clinical implementation."
  18. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — only the abstract - the full text could not be fetched — the passage: "The test dataset comprised 404,502 mammograms from 206 medical organizations and three mammography equipment manufacturers. A total of 336 radiologists participated."
  19. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — only the abstract - the full text could not be fetched — the passage: "Over this time, AUC increased by 10.8% (from 0.83 to 0.92), accuracy by 16.9% (from 0.77 to 0.90), sensitivity by 4.8% (from 0.84 to 0.88), and specificity by 30.0% (from 0.70 to 0.91)."
  20. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — only the abstract - the full text could not be fetched — the passage: "The average technical defect rate decreased by 26.7% (from 3.0% to 0.8%), and the clinical assessment score rose by 47.8% (from 54.38% to 80.36%)."
  21. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — only the abstract - the full text could not be fetched — the passage: "The study culminated in the integration of the AI system into the regional CMI program."
  22. Vasilev Y, Rumyantsev D, Vladzymyrskyy A, Omelyanskaya O, Arzamasov K, Bazhin A, et al. (2026). Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study. Quantitative Imaging in Medicine and Surgery. 10.21037/qims-2026-0864 — only the abstract - the full text could not be fetched — the passage: "A key limitation of this study is the relatively small retrospective calibration testing dataset (100 mammograms) and the lack of external validation on independent datasets from other regions or countries."
  23. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — only the abstract - the full paper is behind a subscription — the passage: "Standard double-reading had a CDR of 5.07 (95% CI: 4.19-6.09). Using AI with a general threshold achieved a CDR of 4.81 (95% CI: 3.95-5.80), which was enhanced using device-specific thresholds to 4.90 (95% CI: 4.03-5.89)."
  24. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — only the abstract - the full paper is behind a subscription — the passage: "Workload for radiologists in consensus conferences was lowered by nearly one-third when using device-specific thresholds compared to a general threshold (p < 0.001)."
  25. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — only the abstract - the full paper is behind a subscription — the passage: "Nonetheless, AI screening performance for one of the mammography devices was poor even with device-specific thresholds."
  26. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — only the abstract - the full paper is behind a subscription — the passage: "AI-based performance varied considerably among the three mammography devices, emphasizing the need for device-specific thresholds or even the withdrawal of AI use on certain devices."
  27. Blum M, Morant R, Eichenberger A, Geissler A, Subelack J, Vogel J, et al. (2026). AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study. Insights into Imaging. 10.1186/s13244-026-02383-5 — only the abstract - the full paper is behind a subscription — the passage: "Lastly, training datasets must be representative not only of the target population but also of the different screening devices used in practice."

Zhou, J., Si, P., Zhang, Y. et al. (2026). A non-invasive end-to-end intelligent assistance system for breast ultrasound. Nature Communications. https://doi.org/10.1038/s41467-026-73170-5

Who paid: The study was funded by the National Key Research and Development Program of China, the Health Commission of Henan Province, the Health Commission of Sichuan Province, the Science and Technology Department of Sichuan Province, and a Dazhou City School Cooperation Special project; the article does not state whether funders had any say, and no equipment or software lenders are mentioned.

Do not take this as professional medical advice.