other · Nature medicine · la publicación, 13 sep 2026 · gratis
La herramienta de IA que ayudó a los médicos a predecir quién se beneficia de la inmunoterapia
Un estudio internacional con 2,396 pacientes halló que los datos clínicos y de sangre de rutina superaron al marcador estándar PD-L1. Pero la herramienta aún no está aprobada ni se usa en la práctica, y su desempeño bajó en un hospital con pacientes distintos a los estudiados.
Versión breve · la versión detallada sigue, unos 7 min
- El estudio, de un vistazo
- Quiénes
- Pacientes con cáncer de pulmón de células no pequeñas avanzado tratados con inmunoterapia
- Cuántos
- 2.396 pacientes
- Dónde
- Seis centros en Italia, Alemania, Grecia, Israel, España y Estados Unidos
- Cuándo
- Entre septiembre de 2012 y octubre de 2023
- Tipo de estudio
- analysis of what people did
- Quién lo hizo
- Consorcio I3LUNG, publicado en Nature Medicine
- El límite que importa
- Es retrospectivo: revisó datos ya recogidos, no decisiones nuevas, y su desempeño bajó en un grupo externo.
Son las cifras de un estudio pequeño y retrospectivo con 20 médicos que evaluaron cada uno a 10 pacientes; miden qué tan bien acertaron antes y después de recibir la sugerencia de la herramienta, no el efecto en decisiones nuevas.
La herramienta de IA frente a las pruebas que se usan hoy para decidir quién recibe inmunoterapia
Los modelos de IA superaron de forma significativa a todas esas pruebas en el grupo de prueba independiente.
La ventaja de sumar imágenes apareció en las pruebas internas, pero no se sostuvo de forma consistente en el grupo de prueba ni en el externo.
En un estudio de uso clínico, 20 médicos —10 especialistas en pulmón y 10 no especialistas— evaluaron cada uno a 10 pacientes, primero solos y luego con la herramienta de IA explicable.

En el cáncer de pulmón avanzado, la inmunoterapia logra un beneficio duradero en apenas 20 a 30 de cada 100 pacientes. El único marcador aprobado para decidir quién la recibe, el PD-L1, es imperfecto. Elegir mal significa toxicidad y gastos innecesarios para quien no responderá.
El consorcio I3LUNG reunió datos del mundo real de 2,396 pacientes con cáncer de pulmón de células no pequeñas en seis centros de seis países: Italia, Alemania, Grecia, Israel, España y Estados Unidos. Recogieron datos clínicos y de sangre de rutina, tomografías, láminas de patología y análisis genómicos entre septiembre de 2012 y octubre de 2023.
Con eso construyeron modelos de aprendizaje automático y profundo, incluida una herramienta de inteligencia artificial explicable, y los compararon con los marcadores que se usan hoy. Los modelos que emplearon solo datos clínicos y de sangre de rutina —sexo, estado funcional, tabaquismo, expresión de PD-L1, metástasis, y valores de laboratorio como la proporción entre neutrófilos y linfocitos y la deshidrogenasa láctica— alcanzaron un desempeño de hasta 0,77 en el grupo de prueba independiente y superaron de forma significativa al PD-L1, al estado funcional ECOG, a esos dos valores de laboratorio y al índice LIPI.
Pero hay que mirar dónde se probó la herramienta. En un grupo externo de validación, con pacientes de características distintas, el desempeño cayó a un rango de 0,55 a 0,72. Es decir, el mismo programa puede funcionar peor en un hospital cuyos pacientes no se parecen a los del estudio. Vale la pena preguntar si se ha probado donde a usted lo atienden.
La ventaja de sumar imágenes, además de los datos de sangre, apareció en las pruebas internas, pero no se sostuvo de forma consistente en el grupo de prueba ni en el externo.
En un estudio de uso clínico, 20 médicos —10 especialistas en pulmón y 10 no especialistas— evaluaron cada uno a 10 pacientes, primero solos y luego con la herramienta de IA explicable. Con la herramienta, los médicos detectaron mejor a los pacientes que responderían: pasaron de identificar correctamente 72 de cada 100 a 87 de cada 100. La exactitud de los médicos para predecir la respuesta también mejoró, de 57 de cada 100 a 65 de cada 100. En cambio, el aumento general de predicciones correctas de respuesta no alcanzó significación estadística. Este estudio fue pequeño y retrospectivo: revisó datos ya recogidos, no puso la herramienta por delante de decisiones nuevas.
Por eso, hoy no se puede pedir como atención estándar. Se está haciendo una validación prospectiva en más de 2,000 pacientes y se planea un ensayo aleatorizado. Si los resultados se sostienen, algún día su médico podría usar un programa que solo necesita análisis de sangre y datos de rutina para orientar quién debe recibir inmunoterapia.
Pregunte en su centro de salud: ¿este hospital participa en la validación de la herramienta y sus pacientes se parecen a los del estudio?
Qué significa para usted
Todavía no hay nada que pedir como atención estándar: la herramienta solo se probó con datos ya recogidos, nunca frente a decisiones nuevas, y funcionó peor en pacientes distintos a los estudiados. Pregunte en su centro si participa en la validación con más de 2,000 pacientes, y si los suyos se parecen a los del estudio. La decisión sigue siendo de su médico.
Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2
Quién pagó: El proyecto I3LUNG recibió financiación del programa Horizonte 2020 de la Unión Europea en virtud del acuerdo de subvención número 101057695; el financiador no tuvo ningún papel en el diseño del estudio, la recopilación y el análisis de datos, la decisión de publicar ni la preparación del manuscrito.
No tome esto como consejo médico profesional.
Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1364 palabras · unos 7 minLeerla →Cerrar
Una herramienta de inteligencia artificial ayudó a los médicos a predecir quién se beneficia de la inmunoterapia en cáncer de pulmón avanzado
En un estudio internacional con casi 2,400 pacientes, un programa que solo usa datos de sangre y ficha clínica superó a la prueba estándar. Todavía no está aprobado ni disponible: la validación en curso sigue reclutando pacientes.

La inmunoterapia cambió el tratamiento del cáncer de pulmón de células no pequeñas avanzado. El artículo describe un estudio en el que solo entre 20 y 30 de cada 100 pacientes obtienen un beneficio duradero1. El artículo describe un estudio en el que la mayoría desarrolla resistencia al tratamiento: entre 5 y 20 de cada 100 desde el inicio, y entre 60 y 85 de cada 100 más adelante2. El artículo describe un estudio en el que urge una forma confiable de saber, desde el diagnóstico, quién se beneficiará y quién no, para evitar toxicidad y gastos innecesarios; hoy la única prueba aprobada es la de PD-L1, y su capacidad de predicción es limitada3.
El estudio I³LUNG, publicado en *Nature Medicine* y financiado por el programa Horizonte 2020 de la Unión Europea, reunió a 2,396 pacientes de seis centros en seis países, tratados entre septiembre de 2012 y octubre de 2023. Los investigadores combinaron datos clínicos y de sangre, imágenes de tomografía computarizada, muestras digitalizadas de tejido y datos genómicos en dos tipos de modelos de inteligencia artificial4. El hallazgo central es que los modelos que usan únicamente datos de sangre y ficha clínica —edad, sexo, estado funcional, sitios de metástasis, recuentos de células y enzimas— funcionaron de manera consistente en distintos desenlaces, con un rendimiento de hasta 0,77 en el conjunto de prueba5, y superaron de forma significativa a la prueba de PD-L1, al estado funcional ECOG, a la razón neutrófilos-linfocitos, a la enzima lactato deshidrogenasa y al índice pronóstico LIPI en ese mismo conjunto independiente6.
La parte más llamativa para una familia es la prueba con médicos reales. Veinte oncólogos —diez expertos en pulmón y diez no expertos, incluidos residentes y oncólogos generales— evaluaron 100 casos reales, primero solos y después con la ayuda del programa, que además explica en palabras sencillas por qué recomienda lo que recomienda7. Al predecir qué pacientes responderían al tratamiento, la colaboración entre médico y herramienta subió la sensibilidad de 0,72 a 0,87 y la exactitud de 0,57 a 0,65, aunque bajó un poco la especificidad8. En la estimación de supervivencia, la herramienta aumentó la probabilidad de acertar un 36% en el conjunto de todos los médicos; un 14% entre los expertos y un 61% entre los no expertos9.
Conviene entender qué es esto y qué no es. No es un aparato que mira al paciente ni un escáner nuevo. Es un programa que toma datos que ya existen en la ficha clínica y en un examen de sangre de rutina, los compara con los de miles de pacientes anteriores y devuelve una probabilidad. En este estudio también se probaron versiones más ambiciosas que sumaban tomografías, láminas de patología y genética4. En la validación cruzada, esas versiones multimodales parecían mejores, pero esa ventaja no se reprodujo de manera consistente en el conjunto de prueba ni en la validación externa10. Los propios autores advierten que el número reducido de casos con todas las modalidades completas limita la fiabilidad de estos análisis11.
El rendimiento cayó cuando la herramienta se probó en un grupo externo, el de la Universidad de Chicago, con pacientes distintos: allí el rango bajó a entre 0,55 y 0,72, y los autores lo atribuyen a diferencias entre poblaciones12. Además, el estudio es retrospectivo, es decir, se hizo mirando hacia atrás sobre datos ya recogidos, con toda la heterogeneidad de la vida real11. Los datos genómicos eran escasos y las características extraídas de las imágenes se limitaron al tumor principal. Esto no invalida el hallazgo, pero marca su alcance: lo que funcionó en estos seis centros podría funcionar distinto en otro hospital con otros pacientes.
No es el primer intento. El artículo describe estudios que usaban solo genética, solo radiología o solo patología digital, pero quedaban limitados por su tamaño, su alcance y el hecho de hacerse en un solo centro13. El artículo describe dos trabajos recientes que combinaron varias modalidades y superaron a los marcadores tradicionales, pero se apoyaban en cohortes pequeñas, de unos 250 a 300 pacientes, y solo unos 80 tenían todos los datos completos14. Este estudio es, según sus autores, el mayor estudio multimodal internacional en cáncer de pulmón e inmunoterapia, y sus autores destacan que incluyó una prueba con médicos reales.
Hay un antecedente útil para calibrar la magnitud del asunto: un análisis de tres grandes estudios europeos de tamizaje de mama con inteligencia artificial, del que solo pudimos leer el resumen porque el artículo completo está detrás de una suscripción, encontró que la integración de la inteligencia artificial en programas de doble lectura produjo un aumento pequeño pero real en la detección de cáncer, de aproximadamente un caso por cada mil exámenes, sin un aumento consistente en las revisiones adicionales15. Ese es el tipo de magnitud que suele tener un avance real en medicina: no un salto, sino una mejora modesta y medida.
¿Qué tendría que pasar para que esto llegue a un hospital cerca de usted? Los propios autores lo dicen: una validación prospectiva del sistema, tanto en su versión de datos clínicos como en la multimodal, está en curso con más de 2,000 pacientes16. Eso significa que hoy la herramienta no está aprobada ni se usa de rutina; no se puede pedir como parte del tratamiento estándar. Los autores señalan que ya está en marcha un camino hacia la práctica clínica, con un estudio de uso en curso, una validación silenciosa en más de 2,000 pacientes y un ensayo aleatorizado previsto, y que el despliegue se espera en un plazo de dos a tres años.
Así lo leemos nosotros. Este estudio pertenece a una familia de historias que ya conocemos: la de la tecnología que promete acercar la medicina especializada a donde no la hay. Si una herramienta así se ofrece como servicio gratuito o de bajo costo dentro de un sistema público, una persona con cáncer de pulmón avanzado en una zona sin oncólogo de pulmón podría recibir una orientación más parecida a la de un centro especializado. Si en cambio queda detrás de un pago o dentro de un hospital privado, la distancia entre quienes ya tienen acceso y quienes no lo tienen podría crecer en lugar de cerrarse.
También conviene mirar quién responde cuando la recomendación falla. La herramienta no decide: sugiere, y el médico tratante sigue siendo quien firma la decisión. Pero conviene preguntar si esa herramienta fue evaluada en personas parecidas a usted o a su familiar, y si sus resultados se revisan por separado en distintos grupos de pacientes. Un programa puede funcionar bien en promedio y fallar más en un centro, en un país o en un grupo de personas, y eso no siempre se nota a simple vista. Sabremos que nos equivocamos si con el tiempo vemos que el acceso a estas herramientas depende de la capacidad de pago y no mejora la orientación para quienes hoy solo cuentan con un médico general.
Lo que usted puede hacer con esto, hoy, es concreto. Puede preguntar en su centro de salud o en su servicio de oncología si existe algún programa que use datos clínicos y de laboratorio para orientar el tratamiento de inmunoterapia, si ese servicio está cubierto o tiene costo, y si los datos de los pacientes se usan de forma protegida y se comparten con el equipo tratante. Puede preguntar también si la herramienta que se usa en su hospital fue evaluada en personas parecidas a usted, y pedir que la decisión final la tome un médico que conozca su caso, no solo el resultado de un programa. No se trata de pedir que la herramienta reemplace a nadie, sino de saber si existe, si está medida y si alguien responde por ella.
La pregunta que puede llevar a su próxima consulta: ¿este hospital usa alguna herramienta de apoyo para decidir el tratamiento de inmunoterapia, y quién revisa sus resultados?
De dónde sale cada dato de contexto, y cuánto leímos de cada documento
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "IO targeting programmed cell death protein 1 (PD-1) 1 – 3 , PD-L1 (ref. 4 ) and cytotoxic T-lymphocyte-associated protein 4 (ref. 5 ) has transformed metastatic NSCLC care, achieving long-term benefit in 20–30% of patients 6 ."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "IO alone or in combination with chemotherapy (IO/CHT) currently forms the backbone of advanced non-oncogene-addicted NSCLC treatment; however, most patients face primary (5–20%) or secondary (60–85%) resistance 7 ."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This underscores the need for robust predictive biomarkers to identify likely IO responders at diagnosis, avoiding unnecessary toxicity and unnecessary cost. Despite its limited predictive power, PD-L1 remains the only clinically approved biomarker 8 ."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "We integrated real-world clinical and blood (CB) data, computed tomography (CT) images, digital pathology (DP), and genomics into machine learning early fusion (MLEF) and deep learning intermediate fusion (DLIF) models."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Machine learning (ML) and deep learning (DL) CB-only models achieved consistent performance across outcomes with area under the curve (AUC) up to 0.77 in the test (TEST) set."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "AI models significantly surpassed PD-L1, Eastern Cooperative Oncology Group performance status (ECOG PS), neutrophil-to-lymphocyte ratio (NLR), lactate dehydrogenase (LDH) and Lung Immune Prognostic Index (LIPI) score in the independent TEST set."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The clinical usability study showed that lung expert and nonexpert physicians improved their prediction with the explainable AI (XAI) ML CB-only based tool."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For DCR prediction, physician–XAI collaboration improved sensitivity for all physicians from 0.72 (95% CI: 0.64–0.80) to 0.87 (95% CI: 0.79–0.92), P = 0.0011, and accuracy from 0.57 to 0.65, P = 0.0431, at the expense of slightly lower specificity"
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "In the OS estimation task, XAI increased the probability of correct prediction by 36% across all physicians—14% for experts and 61% for nonexperts."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Although multimodal integration with MLEF (CB+CT+DP) was associated with higher performance, its incremental benefit remains uncertain, not translated in TEST and EXVAL."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This study has several limitations. First, the retrospective design and heterogeneity of real-world data may affect model performance. Second, the relatively small number of complete multimodal cases limits the reliability of multimodal analyses, particularly for DLIF."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Performance drop in external validation (EXVAL) likely reflects population differences (AUC range: 0.55–0.72)."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Although unimodal studies (for example, genomics 9 , radiomics 10 and DP 11 ) show promise, they remain limited by size, scope and single-center design 12 ."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Although both models outperformed traditional biomarkers, they were limited by small cohorts (approximately 250–300 patients), with only approximately 80 patients having complete multimodal data."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — solo el resumen - el texto completo está tras una suscripción — el pasaje: "In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - el artículo del que trata esta nota — el artículo completo — el pasaje: "A prospective validation of the decision support system (both CB and multimodal) is currently undergoing in more than 2,000 patients."
Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2
Quién pagó: El proyecto I3LUNG recibió financiación del programa Horizonte 2020 de la Unión Europea en virtud del acuerdo de subvención número 101057695; el financiador no tuvo ningún papel en el diseño del estudio, la recopilación y el análisis de datos, la decisión de publicar ni la preparación del manuscrito.
No tome esto como consejo médico profesional.
Los hallazgos de otros estudios que aquí se mencionan los conocemos por este documento, que fue el que leímos; no abrimos cada uno de esos estudios.
other · Nature medicine · the paper, 13 Sep 2026 · free
A Computer Using Routine Data Beat the Standard Lung Cancer Test. It Still Has to Prove Itself.
A large study of 2,396 patients found an AI tool using routine data predicted who would benefit from immunotherapy better than the test doctors use now. The gain did not hold up everywhere.
Short version · the longer version follows, about 6 min
- The study at a glance
- Who
- People with advanced non-small cell lung cancer treated with immunotherapy
- How many
- 2,396 patients
- Where
- Six hospitals in six countries
- When
- Treated between September 2012 and October 2023
- Kind of study
- analysis of what people did
- Who did it
- I3LUNG consortium, six hospitals in six countries
- The limit that matters
- It looked backward at records, not a forward test
These are the doctors' correct calls in a small usability study of 20 physicians reviewing 10 cases each; the rise in overall correct calls was not statistically significant.
The AI tool against the standard measures doctors use now
The AI models significantly surpassed all of them in the independent test group
The tool helped the least experienced physicians most.

People with advanced non-small cell lung cancer often get immunotherapy, a treatment that rallies the body's own defenses against a tumor. It works durably for only about 20 to 30 percent of them. Most have primary or secondary resistance to it. The one biomarker approved to guide that choice, PD-L1, is imperfect.
So the I3LUNG consortium gathered real-world records from 2,396 patients treated between September 2012 and October 2023 at six hospitals in six countries. The material was ordinary: sex, smoking status, performance status, PD-L1 results, where the cancer had spread, and two blood values, plus CT scans, pathology slides and gene tests.
They built machine learning and deep learning models, including ones designed to explain their reasoning to a doctor. Then they compared the models against the standard measures.
When the models used only routine clinical and blood data, they outperformed PD-L1, performance status, the neutrophil-to-lymphocyte ratio, LDH and the combined LIPI score in the independent test group. The result reached up to 0.77 on one common measure for models using only routine clinical and blood data, and the AI models significantly surpassed PD-L1 and other standard measures in the independent test group.
Then the tool met a hospital whose patients looked different. Its accuracy fell, landing between 0.55 and 0.72. If you or a relative are treated somewhere unlike the six study centers, ask whether the tool has been tested there.
Adding scans and pathology slides looked promising in the internal comparisons, but those gains did not reliably survive the test and external groups, likely because so few patients had all their data complete.
To see whether doctors could actually use it, the team ran a small usability study: 20 physicians, 10 lung cancer experts and 10 non-experts, each reviewing 10 patient cases. With the tool and its explanations, doctors became better at catching patients who would benefit—from 0.72 to 0.87. Both experts and non-experts improved. But the rise in overall correct calls was not statistically significant (odds ratio 1.37, P = 0.1). Real, modest, unproven.
This was a look backward at records, not a trial. A prospective validation in more than 2,000 patients is running now, with a randomized trial planned.
If immunotherapy comes up for you or someone you love, ask the oncologist what evidence exists that any decision-support tool performs well at your own hospital.
What this means for you
What you can watch for is the prospective validation now running in more than 2,000 patients, since nothing measured here happened during live care and no tool has been shown to change outcomes. If immunotherapy comes up for you or a relative, ask the oncologist whether any decision-support tool has been tested at your own hospital, and treat the reported gains as promising but unproven.
Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2
Who paid: The I3LUNG project received funding from the European Union's Horizon 2020 program under grant agreement number 101057695; the funder had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.
Do not take this as professional medical advice.
The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1205 words · about 6 minRead it →Close
A Blood Test, an AI, and the Doctors Who Took Its Advice
The largest study of its kind tested whether a computer could help oncologists pick who benefits from immunotherapy in advanced lung cancer. It could — modestly, and not yet everywhere.

Immunotherapy has changed advanced non-small cell lung cancer. The article describes a study in which it produces long-lasting benefit in 20–30% of patients1. But the article describes a study in which most patients face resistance: 5–20% never respond at all, and 60–85% respond and then stop responding2. The article describes a study in which the only approved test for choosing who gets it is PD-L1, and it is weak3. That gap is the reason this study exists.
The I³LUNG project enrolled 2,396 patients with stage IIIC–IVB lung cancer treated with immunotherapy, alone or with chemotherapy, at six clinical centers in six countries. The researchers fed routine clinical and blood data — sex, performance status, smoking, PD-L1, where the cancer had spread, and two common blood counts — into machine-learning and deep-learning models4. They also tested whether adding CT scans, digital pathology slides and genomic data improved things4. Blood-and-clinical-only models reached an area under the curve up to 0.77 in the test set5, and they beat PD-L1, performance status, the neutrophil-to-lymphocyte ratio, LDH and the composite LIPI score6.
What is the technology, in plain words? It is not a new scanner or a new drug. It is pattern recognition on numbers your oncologist already collects: how well you are functioning, whether the cancer has reached bone, liver or brain, how your immune cells and lactate dehydrogenase look on a blood draw. The model learns, across thousands of past patients, which combinations preceded long survival and which preceded early death, then produces a probability for a new patient. The explainable part matters: the tool shows which of those factors pushed its answer up or down, in a chart a doctor can read next to the chart of the patient in front of them.
Nothing here is the first attempt. The article describes studies in which single-modality approaches using genomics, radiomics or digital pathology have shown promise but stayed small, narrow or confined to one hospital7. The article describes two earlier efforts that combined CT, histology and genomic or molecular data and beat traditional biomarkers, but each rested on roughly 250–300 patients, only about 80 of whom had complete data across all modalities8. This study's contribution is scale and breadth, not a new idea.
Twenty oncologists — ten lung-cancer specialists and ten non-specialists — each reviewed ten patient cases, first with the clinical data and images alone, then again with the model's output and its explanations9. For predicting disease control, the doctors' ability to correctly identify patients who would benefit rose from 0.72 to 0.87, and their overall accuracy from 0.57 to 0.65, at the cost of slightly more false alarms10. For estimating how long a patient would live, the probability of a correct estimate rose by 36% overall — 14% for the experts, 61% for the non-experts11.
Read that last pair of numbers again, because it is where the real story is. The tool helped the least experienced physicians most. That is the shape of a technology that might matter in a general oncology clinic, a district hospital, a place without a thoracic specialist on staff — which describes most of where people actually get treated, in Latin America as much as in rural North America.
Performance fell in the external validation cohort — a separate group of patients at the University of Chicago — to a range of 0.55 to 0.72, which the authors say likely reflects differences in the populations12. Adding CT scans, digital pathology and genomic data looked better during internal testing but the advantage did not carry over into the independent test and external validation groups13. Only a small number of patients had complete data across all modalities, which limits how much can be concluded from the multimodal work14. And this study was retrospective — built from records of treatment that already happened.
Here is how we read it. When a tool arrives with the authority of a model behind it, people tend to defer — and the doctors in this study were more likely to follow the tool's incorrect suggestions than the non-experts were. That is worth knowing before you sit in an exam room. When a clinician shows you a risk score, ask what it is based on and whether it changes anything your doctor already suspected. Ask, plainly, whether your doctor feels free to disagree with it. A good doctor will have an answer; a good tool will survive the question.
And when someone offers you a more complex version — the one that also reads your scans and your biopsy slides — ask what it adds over the simpler information already in your file, and whether that added value has been shown outside the place where it was built. That question is not skepticism for its own sake. It is the same question the researchers themselves are asking.
What happens next is already scheduled. Prospective validation — testing the tool forward, on patients whose outcomes are not yet known — is under way in more than 2,000 patients15. A usability study of the multimodal version is already under way, followed by a silent prospective validation in 2,000 patients and a randomized trial intended to support regulatory approval. The authors put deployment two to three years out.
A parallel case shows what that road looks like. A separate review of three large studies embedded in European mammography screening programs — MASAI, ScreenTrustCAD and PRAIM — examined AI added to the double-reading that those programs already use; we could read only the summary of that review, the full paper is behind a subscription. Across 597,419 examinations, the pooled increase in cancer detection was about one additional cancer per thousand women screened, without a consistent rise in recalls1617. The reviewers concluded that AI works best as a complementary reader inside established workflows, and that adoption requires explicit quality assurance and monitoring of interval cancers and stage distribution18. That is the standard this lung-cancer tool will eventually have to meet: not "does it look good in development," but "does it find or spare something that matters, in a program that watches for its mistakes."
So here is what to watch for where you live. If you or a relative are facing immunotherapy for advanced lung cancer, the practical question is not whether AI is coming to oncology — it is whether this particular tool has been tested on patients like the ones in your hospital, and whether anyone can tell you how often it is wrong for different groups of patients. Those two questions are answerable, and asking them is not distrust. It is the same oversight the researchers built into their own study. The tool cannot be requested today as standard care, and it is not a reason to skip, delay or change any treatment your oncologist recommends. But the next time a prediction appears on a screen in the room, you will know what it is made of, what it was measured against, and what to ask about it.
Where each piece of context comes from, and how much of it we read
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "IO targeting programmed cell death protein 1 (PD-1) 1 – 3 , PD-L1 (ref. 4 ) and cytotoxic T-lymphocyte-associated protein 4 (ref. 5 ) has transformed metastatic NSCLC care, achieving long-term benefit in 20–30% of patients 6 ."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "IO alone or in combination with chemotherapy (IO/CHT) currently forms the backbone of advanced non-oncogene-addicted NSCLC treatment; however, most patients face primary (5–20%) or secondary (60–85%) resistance 7 ."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "This underscores the need for robust predictive biomarkers to identify likely IO responders at diagnosis, avoiding unnecessary toxicity and unnecessary cost. Despite its limited predictive power, PD-L1 remains the only clinically approved biomarker 8 ."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "We integrated real-world clinical and blood (CB) data, computed tomography (CT) images, digital pathology (DP), and genomics into machine learning early fusion (MLEF) and deep learning intermediate fusion (DLIF) models."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "Machine learning (ML) and deep learning (DL) CB-only models achieved consistent performance across outcomes with area under the curve (AUC) up to 0.77 in the test (TEST) set."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "AI models significantly surpassed PD-L1, Eastern Cooperative Oncology Group performance status (ECOG PS), neutrophil-to-lymphocyte ratio (NLR), lactate dehydrogenase (LDH) and Lung Immune Prognostic Index (LIPI) score in the independent TEST set."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "Although unimodal studies (for example, genomics 9 , radiomics 10 and DP 11 ) show promise, they remain limited by size, scope and single-center design 12 ."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "Although both models outperformed traditional biomarkers, they were limited by small cohorts (approximately 250–300 patients), with only approximately 80 patients having complete multimodal data."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "The clinical usability study showed that lung expert and nonexpert physicians improved their prediction with the explainable AI (XAI) ML CB-only based tool."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "For DCR prediction, physician–XAI collaboration improved sensitivity for all physicians from 0.72 (95% CI: 0.64–0.80) to 0.87 (95% CI: 0.79–0.92), P = 0.0011, and accuracy from 0.57 to 0.65, P = 0.0431, at the expense of slightly lower specificity"
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "In the OS estimation task, XAI increased the probability of correct prediction by 36% across all physicians—14% for experts and 61% for nonexperts."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "Performance drop in external validation (EXVAL) likely reflects population differences (AUC range: 0.55–0.72)."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "Although multimodal integration with MLEF (CB+CT+DP) was associated with higher performance, its incremental benefit remains uncertain, not translated in TEST and EXVAL."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "This study has several limitations. First, the retrospective design and heterogeneity of real-world data may affect model performance. Second, the relatively small number of complete multimodal cases limits the reliability of multimodal analyses, particularly for DLIF."
- Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2 - the article this story is about — the whole article — the passage: "A prospective validation of the decision support system (both CB and multimodal) is currently undergoing in more than 2,000 patients."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I"
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals."
- Ferre R, Benefield T, Kuzmiak CM. (2026). Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs. Clinical Imaging. 10.1016/j.clinimag.2026.110923 — only the abstract - the full paper is behind a subscription — the passage: "These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution."
Prelaj, A., Miskovic, V., Sacco, M. et al. (2026). Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC. Nature Medicine. https://doi.org/10.1038/s41591-026-04488-2
Who paid: The I3LUNG project received funding from the European Union's Horizon 2020 program under grant agreement number 101057695; the funder had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.
Do not take this as professional medical advice.
The findings of other studies mentioned here are known to us through this document, which is the one we read; we did not open each of those studies.