weeklyAI · Week of 20 September 2026weeklyAI · Semana del 20 de septiembre de 2026

← The good news← Las buenas noticias

experiment · PLOS digital health · la publicación, 15 sep 2026 · gratis

La IA revisó primero las fotos de ojos: menos filas para el especialista, con fallos que nadie vuelve a mirar

En un hospital de Corea del Sur, el software descartó la mayoría de las imágenes como normales. También dejó pasar más casos de retinopatía diabética y degeneración macular que la lectura solo humana.

Versión breve · la versión detallada sigue, unos 6 min

El estudio, de un vistazo
Quiénes
Fotografías de fondo de ojo de pacientes, leídas por tres oftalmólogos especialistas en retina
Cuántos
6,904 fotografías de 2,593 pacientes
Dónde
Hospital universitario de Seúl, Corea del Sur
Cuándo
Fotografías tomadas entre 2004 y 2024
Tipo de estudio
analysis of what people did
Quién lo hizo
VUNO Inc. y el Hospital Nacional Universitario de Seúl
El límite que importa
Fue una simulación retrospectiva en un solo centro; nadie ha medido qué pasa en una clínica real
Cuántas imágenes llegaron a revisión del oftalmólogo cuando el programa filtró primero
Retinopatía diabética7.3%
Retinopatía diabética8%
Oclusión de la vena retinal11.7%
Oclusión de la vena retinal11.9%
Degeneración macular relacionada con la edad13.5%
Degeneración macular relacionada con la edad16.8%

Porcentaje de imágenes que sí pasaron por un oftalmólogo cuando el programa revisó primero, según la enfermedad; son los casos en que el programa marcó la imagen como sospechosa, y el resto se archivó como normal sin que nadie más la viera.

Lo que el programa detectó frente a lo que detectó la lectura solo humana

Lectura solo humanafrente aFlujo con el programa revisando primero

En retinopatía diabética, la detección de casos positivos bajó de 0.9168 a 0.8838

Lectura solo humanafrente aFlujo con el programa revisando primero

En degeneración macular, la detección de casos positivos bajó de 0.8570 a 0.8445

Lectura solo humanafrente aFlujo con el programa revisando primero

En oclusión de la vena retinal, el rendimiento no cambió

Y como las imágenes que el software marca como negativas nunca las ve un médico, cada caso que el software deja pasar se queda sin una segunda mirada dentro de ese flujo.
Lectura de weeklyAI
Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

Cada vez que alguien se hace una foto del fondo del ojo en una clínica, esa imagen suele quedar en una fila esperando a que un especialista en retina la mire. En un estudio, tres oftalmólogos especialistas en retina revisaron por separado 6,904 fotografías de 2,593 pacientes.

Los autores del estudio simularon después otro flujo de trabajo: el software revisaba cada imagen primero. Cuando la marcaba como negativa, se aceptaba como negativa sin que ningún médico la viera. Cuando la marcaba como positiva, pasaba a un oftalmólogo.

El resultado: la revisión directa por el especialista bajó a alrededor del 7% al 8% de las imágenes para retinopatía diabética, alrededor del 12% para oclusión de la vena retinal y alrededor del 14% al 17% para degeneración macular relacionada con la edad.

Para retinopatía diabética, la capacidad del flujo de trabajo de detectar correctamente los casos positivos cayó de 0.9168 a 0.8838. Para degeneración macular, de 0.8570 a 0.8445. Para oclusión de la vena retinal no cambió. Y como las imágenes que el software marca como negativas nunca las ve un médico, cada caso que el software deja pasar se queda sin una segunda mirada dentro de ese flujo.

Este fue un ejercicio retrospectivo: se simularon las decisiones del software sobre lecturas que ya existían, no se mostró ninguna salida de la IA a los médicos en tiempo real. Nadie ha medido todavía qué le pasa a un paciente en una clínica de verdad. Además, se excluyeron las imágenes ambiguas, entre el 6.1% y el 11.6% en el caso de la degeneración macular, así que los resultados describen los casos más claros y podrían verse más seguros de lo que sería una población completa sin filtrar.

El estudio apunta a filas más cortas y atención especializada más rápida para las imágenes que sí la necesitan. Los propios autores señalan que el ajuste debe hacerse enfermedad por enfermedad, que un umbral más estricto recuperó buena parte de la pérdida en retinopatía diabética pero no en degeneración macular, y que conviene mantener a un médico dentro del proceso.

La pregunta que puede llevar a su próxima cita: si en su clínica usan un programa para filtrar imágenes, ¿quién revisa los casos que el programa descarta?

Qué significa para usted

Lo que este estudio permite decir es limitado: fue una simulación con 6,904 fotografías de un solo centro en Corea del Sur, no una clínica ni un resultado clínico. Si en su próxima cita usan un programa que descarta imágenes, pregunte quién revisa las que el programa deja pasar, y pida que un oftalmólogo siga interpretando los casos marcados como positivos.

Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700

Quién pagó: VUNO Inc., financiador comercial y proveedor del sistema de IA evaluado, apoyó el estudio con la subvención SNUH 0620242810, con dos autores empleados por VUNO y el Instituto Médico de Corea financiando a un coautor; el artículo afirma que el financiador no tuvo ningún papel adicional en el diseño, la recolección de datos, el análisis ni las decisiones de publicación.

No tome esto como consejo médico profesional.

Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1258 palabras · unos 6 minLeerla →Cerrar

La inteligencia artificial ya puede descartar la mayoría de las fotos de fondo de ojo. El estudio coreano midió cuántos casos se le escapan.

Un programa revisó 6,904 fotografías de retina antes que los oftalmólogos. El trabajo de revisión bajó hasta un 92%. En dos de las tres enfermedades, también bajó la capacidad de detectar los casos positivos.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

En un consultorio con cámara de fondo de ojo, alguien tiene que mirar cada fotografía y decidir si hay enfermedad. El estudio probó otro orden: primero mira el programa. Las imágenes que el programa marca como normales se archivan como normales sin que ningún oftalmólogo las vea; solo las que marca como sospechosas llegan al especialista, que tiene la última palabra1. La diferencia con dejar que el programa decida solo es esa: aquí el médico sigue siendo el único que diagnostica todo lo que el programa señala, y el programa solo saca de la cola lo que parece sano2.

Los números del estudio son grandes y concretos. Con 6,904 fotografías de fondo de ojo tomadas entre 2004 y 2024 en un hospital universitario de Seúl, cada imagen fue leída por separado por tres oftalmólogos especialistas en retina. Con ese flujo, la revisión directa por el especialista se redujo al 7.3%–8.0% de las imágenes para retinopatía diabética, al 11.7%–11.9% para oclusión venosa retiniana y al 13.5%–16.8% para degeneración macular asociada a la edad3. Dicho de otra forma: aproximadamente el 92% de las imágenes de retinopatía diabética, el 88% de las de oclusión venosa y el 85% de las de degeneración macular ya no llegaron a ojos humanos4.

Ahí está el hallazgo y ahí está el problema. Para la oclusión venosa retiniana el rendimiento no cambió: el programa no empeoró la detección5. Para la retinopatía diabética, la capacidad de detectar los casos positivos bajó de 0,9168 a 0,8838, una caída pequeña pero clara, mientras la mejora en descartar sanos no llegó a ser concluyente6. Para la degeneración macular pasó algo parecido: la detección de positivos bajó de 0,8570 a 0,8445, también de forma pequeña pero sostenida7. Cuando se juntaron las tres enfermedades en un solo filtro —¿esta foto necesita que la vea un especialista, sí o no?—, la revisión bajó al 27.7%–30.1% de las imágenes, una reducción de alrededor del 70%–72%8, y la detección de positivos bajó de 0,9122 a 0,90399.

Ese descenso importa porque en este circuito no hay segunda mirada. Si el programa dice que una imagen es normal, esa imagen no la ve nadie más, así que cada error del programa se convierte en un caso positivo que se pierde, y las ganancias de eficiencia tienen que pesarse contra la seguridad de cada enfermedad y confirmarse en estudios prospectivos10. Los propios autores probaron si se arreglaba subiendo la exigencia del programa. En retinopatía diabética, subirla redujo la pérdida de detección de unos 3 puntos porcentuales a cerca de 1, a cambio de un aumento apenas marginal del trabajo de revisión11. En degeneración macular, la misma maniobra aumentó bastante la carga de revisión y casi no recuperó la detección12. Y hay que decirlo con todas las letras: esto fue una simulación retrospectiva en un solo centro de referencia terciario13, y la reducción de carga se refiere solo a cuántas imágenes pasan por el oftalmólogo, no al tiempo real de lectura ni a la carga clínica total14.

Esto no es un experimento aislado. En los programas europeos de mamografía, la mayoría trabaja con doble lectura y arbitraje, un modelo que salva vidas pero que está cada vez más presionado por la carga de trabajo de los radiólogos, la especificidad variable y los cánceres de intervalo, según el resumen de un estudio sobre lectura doble asistida por inteligencia artificial; solo pudimos leer el resumen, el artículo completo está detrás de una suscripción15. La inteligencia artificial se está evaluando justamente para apoyar u optimizar esos circuitos ya establecidos16. Ese trabajo reunió tres estudios grandes metidos dentro de programas nacionales de tamizaje: uno aleatorizado de triaje por riesgo, otro prospectivo con el programa como lector independiente dentro de la doble lectura, y otro de implementación nacional17. En conjunto, sobre 597,419 exámenes, la detección de cáncer subió de forma leve, alrededor de un caso por cada mil, sin un aumento consistente de los llamados a repetir estudios1819. Y la conclusión de ese resumen es que el programa funciona mejor como lector complementario, con controles explícitos de calidad y vigilancia de los cánceres de intervalo20.

La comparación es útil, pero también marca la diferencia. En la mamografía europea el programa entra en un circuito donde ya hay dos lecturas humanas, y las investigaciones se hicieron dentro de programas reales de tamizaje17. Aquí, en cambio, las imágenes que el programa descarta no las ve nadie1, y todo se reconstruyó después sobre lecturas que los médicos ya habían hecho sin conocer al programa13. Es la diferencia entre apoyar a un lector humano y sustituir su mirada en los casos que el programa considera resueltos.

Así lo leemos nosotros. Cuando una tarea se entrega a un sistema automático, la persona que antes la hacía deja de revisar los casos que el sistema ya dio por resueltos, y esos casos se quedan sin segunda mirada. Eso es lo que pasa aquí: lo que el programa marca como normal no vuelve a pasar por ojos humanos, y quien recibe un resultado "normal" no tiene manera de saber que hubo una duda. Esperamos que en la práctica, en consultorios y campañas de tamizaje, los casos descartados por el programa tampoco se revisen después, salvo que alguien lo exija. Nos equivocaríamos si encontráramos que los casos marcados como normales igual se revisan, o que existe una segunda revisión para los dudosos, o que los positivos omitidos se detectan por otra vía antes de que haya daño.

Y hay un segundo punto que nos parece igual de importante. El estudio se hizo en un solo centro de referencia terciario, y los autores señalan que la utilidad del programa puede variar en poblaciones con distinta prevalencia de enfermedades y en centros que no son de referencia. Esperamos que en poblaciones con enfermedades oculares distintas a las del grupo donde se probó, o con menos acceso a especialistas, la proporción de casos que el programa deja pasar sea mayor que la reportada aquí, y que nadie lo note porque no se mide. Sabríamos que nos equivocamos si el mismo programa mantuviera su rendimiento en grupos diversos de varios países y con distintas cámaras, y si los casos omitidos se registraran y se investigaran abiertamente.

Para el lector, la pregunta práctica es qué pedir. Si en su centro de salud hay cámara de fondo de ojo y el resultado llega con un sello del programa, vale preguntar si además de la máquina alguien miró su imagen, y qué debe hacer si el papel dice normal pero usted nota cambios en la vista. Vale preguntar también con qué pacientes se probó el programa, si se parece a la gente de su comunidad, y si llevan registro de los casos en que el programa y el médico no coinciden. Y conviene saber qué enfermedades busca: este sistema solo está aprobado para retinopatía diabética, oclusión venosa retiniana y degeneración macular; otras cosas que se ven en una foto de fondo de ojo, como el glaucoma, quedaron fuera del circuito.

Si algún día le entregan un resultado de tamizaje hecho con software, la frase que puede usar es simple: "¿Alguien revisó mi imagen además del programa, y qué hago si el programa dijo que estaba bien pero yo noto algo raro?"

De dónde sale cada dato de contexto, y cuánto leímos de cada documento

  1. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "In this workflow, AI-negative images were classified as negative without ophthalmologist review, whereas AI-positive images were referred for human interpretation."
  2. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Unlike autonomous referral, in which the AI issues the final referral decision, this design retains the ophthalmologist as the sole diagnostician for all AI-positive images and uses AI only to remove likely normal images from the review queue."
  3. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The workflow reduced direct ophthalmologist review to 7.3%–8.0% of images for DR, 11.7%–11.9% for RVO, and 13.5%–16.8% for AMD."
  4. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Residual reader workload decreased to 7.3%–8.0% for DR, 11.7%–11.9% for RVO, and 13.5%–16.8% for AMD, corresponding to approximate reductions of 92%, 88%, and 85%, respectively."
  5. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For RVO, performance was essentially unchanged compared with human-only reading. For DR and AMD, the workflow missed slightly more positive cases, with improved or unchanged specificity."
  6. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For DR, mean sensitivity decreased from 0.9168 to 0.8838 (difference, -0.0330; 95% CI, -0.0495 to -0.0165; p  < 0.001), whereas the increase in specificity did not reach statistical significance."
  7. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For AMD, the AI-assisted workflow was associated with a small but statistically significant reduction in sensitivity, from 0.8570 to 0.8445 (difference, -0.0124; 95% CI, -0.0215 to -0.0033; p  = 0.008)"
  8. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Under the AI-assisted workflow, the proportion of images requiring direct ophthalmologist review decreased to 27.7%–30.1% across readers, corresponding to an approximate 70%–72% reduction in review burden."
  9. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Composite sensitivity decreased slightly but significantly under the AI-assisted workflow (from 0.9122 to 0.9039; p < 0.001), whereas the increase in specificity (0.9659 to 0.9804) was not statistically significant."
  10. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "because AI-negative images are not reviewed, any false-negative AI decision becomes a missed positive, so efficiency gains must be weighed against disease-specific safety and confirmed prospectively."
  11. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For DR, raising the target sensitivity from 99.0% to 99.5% reduced the per-reader sensitivity decrement relative to human-only reading from approximately 3 percentage points to about 1 percentage point, recovering most of the loss at the cost of only a marginal increase in residual review"
  12. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "for AMD the same change substantially increased residual review with little sensitivity recovery and left the decrement statistically significant."
  13. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This was a retrospective simulation study conducted at a single tertiary referral center."
  14. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - el artículo del que trata esta nota — el artículo completo — el pasaje: "the reported workload reduction reflects direct ophthalmologist review burden for DR, RVO, and AMD only. It should not be interpreted as actual reading time or total clinical workload reduction in routine fundus interpretation."
  15. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Most European population mammography screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers."
  16. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Artificial intelligence (AI) is being evaluated to support or optimize these established European screening pathways."
  17. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "We performed a prespecified, focused evidence synthesis of three large studies embedded within routine population screening programs operating under European-relevant workflows: MASAI (randomized AI-supported risk triage within a national program), ScreenTrustCAD (prospective paired-reader evaluation with AI as an independent reader in a double-reading framework), and PRAIM (nationwide decision-referral implementation)."
  18. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I"
  19. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals."
  20. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution."

Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700

Quién pagó: VUNO Inc., financiador comercial y proveedor del sistema de IA evaluado, apoyó el estudio con la subvención SNUH 0620242810, con dos autores empleados por VUNO y el Instituto Médico de Corea financiando a un coautor; el artículo afirma que el financiador no tuvo ningún papel adicional en el diseño, la recolección de datos, el análisis ni las decisiones de publicación.

No tome esto como consejo médico profesional.

experiment · PLOS digital health · the paper, 15 Sep 2026 · free

AI Could Spare Eye Specialists Most of the Work — With a Few Misses Nobody Would Catch

A Korean study simulated letting software clear eye photos as normal. It cut specialist reviews by roughly 85% to 92%. It also missed a few real cases.

Short version · the longer version follows, about 6 min

The study at a glance
Who
Retinal photographs of patients' eyes
How many
6,904
Where
Seoul National University Hospital, Korea
When
2004 to 2024
Kind of study
analysis of what people did
Who did it
Seoul National University Hospital and VUNO Inc.
The limit that matters
Simulation only; no doctor saw the software's answer, and unclear images were set aside
Share of eye photos a specialist still had to open, after software cleared the rest
Diabetic retinopathy7.3%
Diabetic retinopathy8%
Retinal vein occlusion11.7%
Retinal vein occlusion11.9%
Age-related macular degeneration13.5%
Age-related macular degeneration16.8%

Out of every 100 eye photos, this is how many still went to a specialist for a look, for each of the three diseases. The rest were filed as normal by the software and never seen by a doctor.

How many real cases the workflow caught, with and without the software clearing photos first

Diabetic retinopathy, doctors aloneagainstDiabetic retinopathy, with software clearing first

Fell from 0.9168 to 0.8838

Age-related macular degeneration, doctors aloneagainstAge-related macular degeneration, with software clearing first

Fell from 0.8570 to 0.8445

Retinal vein occlusion, doctors aloneagainstRetinal vein occlusion, with software clearing first

Essentially unchanged

In this setup, nothing the software calls negative is ever seen by a doctor.
weeklyAI's reading
How it could look · illustration generated by weeklyAI.watch, not a photograph

At a clinic, a camera photographs the back of the eye. Those photographs then wait for a specialist to look at each one. In a study at Seoul National University Hospital, three retina specialists graded 6,904 such photographs, taken from 2,593 patients between 2004 and 2024.

Researchers at the hospital and at VUNO Inc., a Korean company that makes the software, then simulated a different arrangement. Software looked at every image first. Anything the software called negative was accepted as negative, with no doctor looking at it. Anything it called positive went to an ophthalmologist, whose reading decided the final answer.

The software is approved in Korea for three conditions: diabetic retinopathy, retinal vein occlusion and age-related macular degeneration.

The result was a large cut in the number of images a specialist had to open. Reviews fell to about 7% to 8% of images for diabetic retinopathy, about 12% for retinal vein occlusion, and about 14% to 17% for age-related macular degeneration. That is a reduction of roughly 92%, 88% and 85%.

For diabetic retinopathy, the workflow caught fewer true cases than reading by a specialist alone. For age-related macular degeneration, it also caught fewer true cases than reading by a specialist alone. For retinal vein occlusion, the workflow performed the same as reading by a specialist alone.

In this setup, nothing the software calls negative is ever seen by a doctor. There is no second look. Every case the software misses is a missed case within that workflow.

Two limits shape what these numbers can claim. This was a retrospective simulation at a single Korean hospital. The doctors graded the images without seeing the software's output, so nobody has yet measured what happens in a live clinic, where knowing the software's answer might change a doctor's own reading. And ambiguous images were set aside — 6.1% to 11.6% of images for age-related macular degeneration — so the results describe clearer cases and may look safer than a full screening population would.

What it could make possible is shorter queues: specialists spending their attention on the images that need it, with clinics choosing a stricter, more cautious setting for each disease and keeping a human deciding every case the software flags. For diabetic retinopathy, a stricter setting recovered most of the missed cases at little extra cost in reviews; for age-related macular degeneration it did not.

If a clinic near you starts using software like this, ask one question: which diseases is it cleared for, and who looks at the images it clears?

What this means for you

For now, nothing changes in your own eye care, and no one should conclude this software makes screening safe everywhere. If a clinic near you adopts something like it, ask which diseases it is cleared for and who reviews the images it clears as normal.

Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700

Who paid: VUNO Inc., the commercial funder and provider of the AI system evaluated, supported the study under grant SNUH 0620242810, with two authors employed by VUNO and the Korea Medical Institute funding a co-author; the article states the funder had no additional role in design, data collection, analysis, or publication decisions.

Do not take this as professional medical advice.

The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1138 words · about 6 minRead it →Close

A Korean Hospital Tested Letting Software Clear Eye Photos Before a Doctor Saw Them

In a simulation of 6,904 retinal images, specialists were left with roughly a tenth of the work. But because no one reviews what the software clears, every miss becomes a missed case.

How it could look · illustration generated by weeklyAI.watch, not a photograph

The experiment works like a sieve placed before a specialist. A photograph of the back of the eye goes into the software first. If the software calls the image negative, it is filed as negative and no ophthalmologist ever looks at it. If the software calls it positive, a human reads it and makes the final call1. The authors are careful to say this is not a machine diagnosing on its own: the eye doctor remains the only diagnostician for every image the software flags, and the software's only job is to take likely-normal pictures out of the queue2.

The study covered three common retinal diseases: diabetic retinopathy, retinal vein occlusion, and age-related macular degeneration. Across 6,904 photographs, human review dropped to 7.3%–8.0% of images for diabetic retinopathy, 11.7%–11.9% for retinal vein occlusion, and 13.5%–16.8% for age-related macular degeneration3 — reductions of roughly 92%, 88% and 85%4. That is the headline gain, and it is large.

The cost is unevenly distributed. For retinal vein occlusion, performance was essentially unchanged compared with reading by people alone. For diabetic retinopathy and age-related macular degeneration, the workflow missed slightly more positive cases, while specificity — the ability to correctly clear a healthy eye — improved or stayed the same5. In the pooled multi-reader analysis, sensitivity for diabetic retinopathy fell from 0.9168 to 0.8838, a difference of -0.03306, and for age-related macular degeneration from 0.8570 to 0.8445, a difference of -0.01247. When the three diseases were combined into a single question — does this image show any of them? — review fell to 27.7%–30.1% of images, roughly a 70%–72% reduction8, and sensitivity slipped slightly from 0.9122 to 0.90399.

The authors state the central limit themselves: because AI-negative images are never reviewed, any false-negative decision becomes a missed positive, so the efficiency gain has to be weighed against disease-specific safety and confirmed prospectively10. They also tested a stricter setting. Raising the software's target from 99.0% to 99.5% cut the sensitivity loss for diabetic retinopathy from about 3 percentage points to about 1, recovering most of the loss for only a marginal increase in review11. For age-related macular degeneration the same change substantially increased review with little recovery, and the loss stayed statistically significant12. This was a retrospective simulation at a single tertiary referral center13, and the workload figure means images reviewed, not reading time or total clinical workload14.

Anyone who has sat in a screening line knows the arithmetic underneath this: there are more photographs than there are specialists to read them, and the queue is where delays live. A summary of a review of mammography screening in Europe describes the same pressure — programs built on double reading by two radiologists, delivering mortality benefit but strained by workload, variable specificity and cancers found between screening rounds; we could read only the summary, the full paper is behind a subscription15. That review describes AI being evaluated to support or optimize those established pathways16, and it looked at three large studies embedded inside routine national screening programs, including one randomized trial of AI-supported risk triage and one nationwide implementation17. Across 597,419 examinations, the pooled increase in cancer detection was +0.9 per 1,00018, and the authors conclude that AI integration may yield a small absolute increase in detection — about one in a thousand — without a consistent increase in recalls19. That is the same shape of result as the eye study: a modest shift in what gets caught, set against a large change in who has to look.

The mechanism is worth picturing, because it explains why the misses matter more here than in a system where a human still sees everything. In this design the software is a gate, not a second opinion. A gate that opens for the wrong person is correctable; a gate that stays shut is not, because no one downstream knows it stayed shut10. The authors suggest the fix may not be a stricter gate alone but additional safeguards — referring cases the software is unsure about, or having someone selectively review borderline negatives — and say these should be tested prospectively, in studies where AI output reaches the clinician in real time, because behavior with a machine in the room (over-reliance, under-reliance, anchoring) was not measured here at all20. They also call for testing in populations with different rates of disease and outside big referral hospitals, since whether this helps depends on the local balance between missed cases and saved work21.

Here is how we read it. The number that will travel is the 92%. The number that should travel with it is the one nobody measured: what a specialist's eye used to catch on the images that now get cleared without anyone looking. In a clinic that adopts this, we would expect the workflow to feel like an unambiguous improvement — faster, cheaper, more consistent — while the habit of a person glancing at every photograph quietly thins out, and nothing visibly breaks to announce it. You would know we were wrong if clinics running this kind of triage kept auditing a sample of the images the software passed as negative, and those audits kept finding the same rate of disease as before. That audit is the thing to ask about.

A second thing we notice. A tool built to sort eyes into three named diseases will misplace the eyes that do not fit those three names — the unusual, the overlapping, the ones a doctor would describe with a shrug. The authors' own look back at the missed cases points that way, describing atypical and overlapping appearances rather than only faint ones. So if you or someone in your family has a condition that does not fit neatly into one label, or a history of unusual eye problems, the useful move is not to trust the silence of a cleared image. Say it out loud, and ask that a person still look.

What you can do with this is small and concrete. At your next eye appointment, ask whether a human looked at your photograph or only the software, and whether anyone ever rechecks the images the software clears. When a clinic or a health service tells you it has become faster through automation, ask what used to be checked by a person and is no longer — and who is accountable if that check was the one that mattered. You do not need to refuse the technology to ask that question. You only need to ask it out loud.

What was cleared without anyone looking — and who checks?

Where each piece of context comes from, and how much of it we read

  1. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "In this workflow, AI-negative images were classified as negative without ophthalmologist review, whereas AI-positive images were referred for human interpretation."
  2. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "Unlike autonomous referral, in which the AI issues the final referral decision, this design retains the ophthalmologist as the sole diagnostician for all AI-positive images and uses AI only to remove likely normal images from the review queue."
  3. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "The workflow reduced direct ophthalmologist review to 7.3%–8.0% of images for DR, 11.7%–11.9% for RVO, and 13.5%–16.8% for AMD."
  4. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "Residual reader workload decreased to 7.3%–8.0% for DR, 11.7%–11.9% for RVO, and 13.5%–16.8% for AMD, corresponding to approximate reductions of 92%, 88%, and 85%, respectively."
  5. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "For RVO, performance was essentially unchanged compared with human-only reading. For DR and AMD, the workflow missed slightly more positive cases, with improved or unchanged specificity."
  6. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "For DR, mean sensitivity decreased from 0.9168 to 0.8838 (difference, -0.0330; 95% CI, -0.0495 to -0.0165; p  < 0.001), whereas the increase in specificity did not reach statistical significance."
  7. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "For AMD, the AI-assisted workflow was associated with a small but statistically significant reduction in sensitivity, from 0.8570 to 0.8445 (difference, -0.0124; 95% CI, -0.0215 to -0.0033; p  = 0.008)"
  8. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "Under the AI-assisted workflow, the proportion of images requiring direct ophthalmologist review decreased to 27.7%–30.1% across readers, corresponding to an approximate 70%–72% reduction in review burden."
  9. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "Composite sensitivity decreased slightly but significantly under the AI-assisted workflow (from 0.9122 to 0.9039; p < 0.001), whereas the increase in specificity (0.9659 to 0.9804) was not statistically significant."
  10. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "because AI-negative images are not reviewed, any false-negative AI decision becomes a missed positive, so efficiency gains must be weighed against disease-specific safety and confirmed prospectively."
  11. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "For DR, raising the target sensitivity from 99.0% to 99.5% reduced the per-reader sensitivity decrement relative to human-only reading from approximately 3 percentage points to about 1 percentage point, recovering most of the loss at the cost of only a marginal increase in residual review"
  12. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "for AMD the same change substantially increased residual review with little sensitivity recovery and left the decrement statistically significant."
  13. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "This was a retrospective simulation study conducted at a single tertiary referral center."
  14. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "the reported workload reduction reflects direct ophthalmologist review burden for DR, RVO, and AMD only. It should not be interpreted as actual reading time or total clinical workload reduction in routine fundus interpretation."
  15. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "Most European population mammography screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers."
  16. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "Artificial intelligence (AI) is being evaluated to support or optimize these established European screening pathways."
  17. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "We performed a prespecified, focused evidence synthesis of three large studies embedded within routine population screening programs operating under European-relevant workflows: MASAI (randomized AI-supported risk triage within a national program), ScreenTrustCAD (prospective paired-reader evaluation with AI as an independent reader in a double-reading framework), and PRAIM (nationwide decision-referral implementation)."
  18. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I"
  19. Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals."
  20. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "These behavioral effects, including automation bias and trust calibration, should be evaluated in prospective studies incorporating real-time AI interaction"
  21. Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700 - the article this story is about — the whole article — the passage: "External validation in populations with different prevalence profiles and in non-tertiary settings will also be important, because the clinical usefulness of this strategy depends on the context-specific trade-off between missed positives and workload reduction."

Park, D., Oh, R., Baek, J. et al. (2026). Evaluating an AI-assisted triage workflow for retinal diseases. PLOS Digital Health. https://doi.org/10.1371/journal.pdig.0001700

Who paid: VUNO Inc., the commercial funder and provider of the AI system evaluated, supported the study under grant SNUH 0620242810, with two authors employed by VUNO and the Korea Medical Institute funding a co-author; the article states the funder had no additional role in design, data collection, analysis, or publication decisions.

Do not take this as professional medical advice.