experiment · Frontiers in Psychiatry · la publicación, 11 sep 2026 · gratis
Los chatbots dieron respuestas aceptables en la mayoría de los casos, pero no en los urgentes
Un estudio con 90 preguntas sobre depresión en la vejez halló respuestas aceptables en la mayoría de los casos, pero menos de la mitad en los escenarios de mayor riesgo.
Versión breve · la versión detallada sigue, unos 5 min
- El estudio, de un vistazo
- Quiénes
- Tres programas conversacionales: ChatGPT, Gemini y Doubao
- Cuántos
- 90 preguntas sobre depresión en la vejez
- Dónde
- China (autores de universidades y hospitales de China)
- Cuándo
- julio de 2026
- Tipo de estudio
- experiment
- Quién lo hizo
- Universidades y hospitales de China
- El límite que importa
- Las preguntas las escribieron los investigadores, no personas mayores ni cuidadores reales
Porcentaje de las 90 preguntas en que cada programa dio una respuesta clínicamente aceptable; las preguntas las escribieron los investigadores, no personas mayores ni cuidadores reales.
Errores graves de seguridad por programa
5.6% frente a 10.0% de respuestas con error grave de seguridad
5.6% frente a 16.7% de respuestas con error grave de seguridad
10.0% frente a 16.7% de respuestas con error grave de seguridad
Tome lo que le diga un chatbot como el punto de partida de una conversación con un profesional de salud, nunca como la decisión misma.
Cada vez más personas mayores, y los hijos y hijas que las acompañan, consultan a un chatbot sobre el ánimo bajo, los cambios de memoria o un medicamento. La respuesta llega en segundos, con tono seguro y bien redactada.
Para medir qué tan confiables son esas respuestas, un equipo de investigadores escribió 90 preguntas sobre depresión en la vejez, planteadas desde la perspectiva del propio adulto mayor y desde la de un familiar cuidador. Las enviaron a ChatGPT, Gemini y Doubao, y dos psiquiatras evaluaron las respuestas sin saber cuál programa las había escrito.
ChatGPT y Gemini dieron respuestas clínicamente aceptables a alrededor de tres cuartas partes de las preguntas; Doubao, al 60%. En las preguntas de riesgo bajo, los tres salieron bien. El problema apareció en las de riesgo alto: la aceptabilidad cayó a 63.3% en ChatGPT, 50.0% en Gemini y 36.7% en Doubao. En las preguntas de riesgo alto, el consejo fue menos urgente de lo que la situación requería en 36.7% de las respuestas de Doubao y en 26.7% de las de Gemini.
Ahora bien: esas 90 preguntas las redactaron los investigadores, no las hicieron personas mayores ni cuidadores reales. Nadie midió si un lector común entendería la respuesta ni qué haría después. Tome lo que le diga un chatbot como el punto de partida de una conversación con un profesional de salud, nunca como la decisión misma. Además, todas las preguntas se hicieron en inglés y cada programa se probó en una sola versión pública durante julio de 2026. No sabemos si estos resultados se repiten en español ni en otras versiones, así que no dé por sentado que un chatbot es más seguro, ni menos, en su propio idioma. El estudio tampoco probó si repreguntar o usar otras herramientas mejora las respuestas: solo muestra cómo se comportaron estos sistemas en un uso común, sin supervisión.
El error grave más frecuente tuvo que ver con el riesgo de suicidio. Otros fallos: no considerar un delirium o un deterioro médico agudo, dar indicaciones sobre medicamentos sin revisión clínica y no responder con suficiente urgencia ante el abandono del autocuidado. Entre los elementos que más faltaron en las respuestas de crisis estuvo el consejo de que la persona no debía quedarse sola.
Usados con cuidado y junto a un profesional, estos programas todavía pueden servir: ayudan a ordenar lo que uno observa, a poner en palabras lo que preocupa y a preparar las preguntas para la próxima consulta.
Antes de seguir un consejo de un chatbot sobre ánimo, memoria o medicamentos, pregúntese: ¿esto lo comenté ya con un médico o una enfermera?
Qué significa para usted
Si usted o un hijo suyo consulta a un chatbot por ánimo bajo, memoria o medicamentos, recuerde que el estudio no midió lo que haría una persona real, solo cómo respondieron estos programas en inglés. Cuando la situación parece urgente, la respuesta puede quedarse corta o quitarle urgencia; ante señales como no querer comer, descuidarse o hablar de no seguir viviendo, consulte a un profesional.
Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736
Quién pagó: Los autores declararon que no recibieron apoyo financiero para este trabajo ni para su publicación, y que no tenían relaciones comerciales o financieras que pudieran constituir un conflicto de interés.
No tome esto como consejo médico profesional.
Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1076 palabras · unos 5 minLeerla →Cerrar
Un chatbot puede responder sobre la depresión en la vejez. Un estudio midió cuántas veces esa respuesta sería segura.
Tres modelos de lenguaje contestaron 90 preguntas escritas por investigadores. Ninguno fue del todo fiable cuando la situación era urgente.

Un equipo de investigadores de China puso a prueba tres programas conversacionales de uso público —ChatGPT, Gemini y Doubao— con 90 preguntas sobre depresión en personas mayores. Las preguntas eran de tres niveles de riesgo: bajo, medio y alto. Dos psiquiatras evaluaron cada respuesta sin saber qué programa la había escrito, comparándola con un estándar clínico definido de antemano para cada pregunta. El estudio se publicó en la revista Frontiers in Psychiatry y sus autores trabajan en universidades y hospitales de China. Los autores declararon no haber recibido financiamiento para este trabajo ni tener relaciones comerciales que pudieran sesgarlo.
El resultado principal: ChatGPT dio una respuesta clínicamente aceptable en 78.9% de las preguntas, Gemini en 72.2% y Doubao en 60.0%1. La diferencia entre ChatGPT y Gemini no fue estadísticamente significativa; ambos superaron a Doubao. Los errores graves de seguridad aparecieron en 5.6% de las respuestas de ChatGPT, 10.0% de Gemini y 16.7% de Doubao2.
La adecuación geriátrica completa —es decir, haber tomado en cuenta todos los factores propios de la vejez que la pregunta exigía— se logró en 64.4%, 55.6% y 43.3% de las respuestas, respectivamente3. Cuando los investigadores repitieron 30 preguntas en conversaciones nuevas, la coherencia clínica entre una respuesta y otra fue de 90.0% para ChatGPT, 83.3% para Gemini y 73.3% para Doubao4. En los casos de crisis, los elementos que más se omitieron fueron decir que la persona mayor no debe quedarse sola e involucrar explícitamente a un familiar o cuidador de confianza5.
Hay que leer estos números con cuidado. Las 90 preguntas las redactaron los investigadores, no las hicieron personas mayores reales ni sus cuidadores en una consulta6. Todas estaban en inglés7. Cada programa se probó a través de una sola versión pública de su aplicación, en julio de 2026. El estudio no midió qué entendería o haría una persona real ante esas respuestas, ni probó si hacer preguntas de seguimiento mejora el resultado. Tampoco es una muestra representativa de lo que la gente pregunta en la vida diaria: fue una prueba de esfuerzo, con los tres niveles de riesgo repartidos por igual a propósito.
¿Qué son estos programas? Los modelos de lenguaje de uso general se han vuelto una fuente cada vez más común de información de salud conversacional8. A diferencia de un buscador, no devuelven una lista de enlaces: generan una respuesta directa que puede sonar individualizada, coherente y con autoridad clínica8. Una persona mayor o un hijo puede preguntarles si un síntoma es depresión, si conviene suspender un medicamento, si el olvido es parte del envejecimiento normal o con qué urgencia hace falta ayuda profesional9. Es exactamente el tipo de pregunta en la que uno quiere una respuesta clara. Y es exactamente donde una respuesta clara puede estar equivocada.

El peso de fondo: la depresión en la vejez es una causa importante de discapacidad, dependencia funcional y pérdida de calidad de vida10. Y a pesar de eso, sigue siendo con frecuencia mal reconocida y mal tratada11. Ahí está la tentación del atajo: si el sistema de salud tarda o queda lejos, un chatbot responde al instante. El problema es que estos modelos no distinguen bien cuándo un síntoma depresivo es solo eso y cuándo es la punta de algo urgente.
Los propios autores piden prudencia. Dicen que, hasta que estas capacidades se demuestren de forma consistente en distintas versiones, idiomas y entornos reales, la supervisión de un clínico debe seguir siendo central en el uso de modelos de lenguaje para la depresión en la vejez12. Es decir: el estudio no dice que estos programas sirvan para decidir, sino que midió cuánto fallan cuando se los usa sin supervisión.
Así lo leemos nosotros. Una voz que suena segura y ordenada invita a dejar de preguntar. Si un chatbot responde con calma que los olvidos acompañan a la depresión, o que hay que esperar a que el antidepresivo haga efecto, es fácil quedarse ahí, sobre todo cuando la alternativa es conseguir una consulta. Lo que esperamos ver en las casas es justamente eso: respuestas tranquilizadoras aceptadas sin contraste, precisamente porque suenan bien. Nos equivocaríamos si las familias, después de leer una respuesta así, igual llamaran al médico, pidieran otra opinión o compararan con otra fuente antes de decidir.
Y hay un segundo patrón que conviene tener presente. Aceptar lo que ya viene masticado cuesta menos que revisarlo. Por eso es probable que, ante una respuesta sobre ánimo o memoria, nadie note lo que falta: el riesgo de caídas, la mezcla de medicamentos, la necesidad de que alguien acompañe. Si eso no ocurre —si la gente revisa la respuesta buscando lo que no dice, o le pide a otro que la lea antes de actuar—, entonces este patrón no aplica. Lo que usted puede hacer con esto es concreto: lea la respuesta preguntándose qué no mencionó. Caídas, confusión, medicamentos, soledad, falta de apetito. Si no aparecen, esa ausencia es la señal. Y si la respuesta suena demasiado tranquila ante un cambio repentino, insista en una evaluación real.

Una última cosa que conviene saber, porque no depende de usted. Estos programas no son autónomos: dependen de personas que los diseñan y los ajustan, y pueden contestar distinto la misma pregunta en momentos distintos. Por eso, si una respuesta sobre salud le parece importante, vuelva a preguntar lo mismo otro día y compare. Si la urgencia o el consejo cambian, quédese con la versión más prudente y consulte a un profesional.
¿Qué haría falta para que esto cambie? Los autores piden que las evaluaciones futuras se hagan con preguntas reales, en más idiomas, en más versiones y en entornos clínicos verdaderos, y que midan no solo si la información es correcta sino si la urgencia recomendada es la adecuada12. También habría que comprobar si la respuesta de un chatbot cambia lo que una persona entiende y hace. Nada de eso se midió aquí.
Nadie puede saltarse, retrasar ni reemplazar una consulta por lo que diga una aplicación. Lo que este estudio deja en sus manos es una costumbre simple: cuando un chatbot le responda sobre ánimo, memoria o medicamentos en una persona mayor, léalo buscando lo que no dice, y llévelo a la consulta como punto de partida. ¿Qué le faltó mencionar a esa respuesta?
De dónde sale cada dato de contexto, y cuánto leímos de cada documento
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Clinically acceptable responses were generated for 78.9% of questions by ChatGPT, 72.2% by Gemini, and 60.0% by Doubao (overall P<0.001)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Major safety errors occurred in 5.6% of ChatGPT responses, 10.0% of Gemini responses, and 16.7% of Doubao responses (raw P = 0.015; FDR-adjusted q=0.023)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Complete geriatric-specific appropriateness was achieved in 64.4%, 55.6%, and 43.3% of responses, respectively."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Test-retest clinical consistency was highest for ChatGPT (90.0%), followed by Gemini (83.3%) and Doubao (73.3%)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Advice that the older adult should not remain alone, explicit involvement of a trusted caregiver, and reduction of access to potentially lethal medication or other means of harm were the elements most frequently omitted."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The benchmark consisted of standardized scenarios rather than questions entered prospectively by older adults or caregivers during actual healthcare decision-making."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "All questions were presented in English. This improved standardization across models but may have influenced the performance of systems optimized for other languages or cultural contexts."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "General-purpose large language models are increasingly available as sources of conversational health information (16). Unlike conventional search tools, these systems generate direct responses that may appear individualized, coherent, and clinically authoritative (17)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "An older adult or caregiver can ask whether a symptom represents depression, whether a medicine should be stopped, whether memory decline is part of normal ageing, or how urgently professional help is required."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Late-life depression is a major source of disability, functional dependence, and diminished quality of life among older adults (1)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Despite these consequences, late-life depression remains frequently under-recognized and undertreated (3)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Until these capabilities are demonstrated consistently across model versions, languages, and real-world settings, clinician oversight should remain central to the use of large language models in late-life depression care."
Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736
Quién pagó: Los autores declararon que no recibieron apoyo financiero para este trabajo ni para su publicación, y que no tenían relaciones comerciales o financieras que pudieran constituir un conflicto de interés.
No tome esto como consejo médico profesional.
Los hallazgos de otros estudios que aquí se mencionan los conocemos por este documento, que fue el que leímos; no abrimos cada uno de esos estudios.
experiment · Frontiers in Psychiatry · the paper, 11 Sep 2026 · free
Chatbots answered 3 in 4 late-life depression questions safely, study finds
ChatGPT and Gemini did best on a set of 90 written questions about depression in later life. As the situations grew more dangerous, all three programs tested slipped.
Short version · the longer version follows, about 5 min
- The study at a glance
- Who
- Three chatbots answering questions about depression in later life
- How many
- 90 questions
- Where
- Not stated in the passages
- When
- July 2026
- Kind of study
- analysis of what people did
- Who did it
- Researchers publishing in Frontiers in Psychiatry
- The limit that matters
- Questions were written by researchers, not asked by real older adults or caregivers
These are the shares of 90 written questions each chatbot answered acceptably, judged by psychiatrists who did not know which program wrote which answer; the questions were written by researchers, not asked by real patients or families.
Serious safety errors, by chatbot
5.6% of ChatGPT answers had a serious safety error, against 10.0% of Gemini's
5.6% of ChatGPT answers had a serious safety error, against 16.7% of Doubao's
10.0% of Gemini answers had a serious safety error, against 16.7% of Doubao's
Older adults and the adult children who look out for them are turning to chatbots with some of the hardest questions they face: Why has my mother stopped eating? Is this low mood, or something else? Should she stop that medicine? The programs answer in seconds, in plain language, as if they know.
Researchers wrote 90 questions about depression in later life — some phrased as an older person would ask, some as a caregiver would — and put each one to ChatGPT, Gemini and Doubao. Two psychiatrists scored the answers without being told which program had produced them, and a third settled disagreements. The study appeared in the journal Frontiers in Psychiatry.
ChatGPT and Gemini gave clinically acceptable answers to about three-quarters of the questions; Doubao to 60 percent. ChatGPT scored highest, though its lead over Gemini was not statistically significant — a way of saying the gap could be chance.
The picture changed with risk. The 90 questions were split evenly into low, moderate and high risk. On the high-risk ones, acceptable answers fell to 63.3 percent for ChatGPT, 50 percent for Gemini and 36.7 percent for Doubao. Under-triage — sending someone to a less urgent level of care than the situation called for — rose to 16.7 percent, 26.7 percent and 36.7 percent. In high-risk scenarios, major safety errors affected 13.3 percent of ChatGPT responses, 20 percent of Gemini responses and 33.3 percent of Doubao responses.
The questions were written by researchers, not asked by real older adults or caregivers, so no one measured what a real person would understand or do with an answer, or whether they would act on it. All 90 were in English, and each program was tested through one public app version in July 2026 — so the results may not carry over to Spanish, or to other versions. Do not assume a chatbot is safer, or less safe, in your own language.
The most common serious error was failing to recognise or escalate suicide risk. Other errors: not considering delirium or a sudden medical problem when thinking changed quickly, and unsafe advice about prescribed medicines. Few answers told a caregiver to stay with the person while help was on the way. Even ChatGPT gave a fully complete crisis response only 40 percent of the time.
The repeated questions showed a further problem. When the same question was asked again in a new conversation, ChatGPT's answers agreed with themselves 90 percent of the time, Gemini's 83.3 percent and Doubao's 73.3 percent. The study also found that repeated answers differed in ways that could change what a person does in 10 to 20 percent of cases, and that the level of urgency changed between the two answers in 3.3 percent of ChatGPT pairs, 6.7 percent of Gemini pairs and 13.3 percent of Doubao pairs.
The study did not test whether asking follow-up questions, or using other tools, improves the answers. It shows only how the programs behaved in ordinary, unsupervised use.
Used carefully, these tools can still help. A chatbot's answer can help a family put words to what they are seeing, and gather questions worth bringing to a doctor or nurse. The researchers' own conclusion: these programs may help with low-risk education, but should not decide on their own how urgent a situation is, what to do about a medicine, or how to handle a crisis. Bring the answer to a clinician — and ask: "Does this fit my mother's situation, and how soon should she be seen?"
What this means for you
The study says nothing about whether a chatbot answered well for you or someone in your family, only how three programs handled 90 written questions in English. So bring its answer to a doctor or nurse and ask whether it fits the situation and how soon the person should be seen.
Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736
Who paid: The authors declared that no financial support was received for this work or its publication, and they declared no commercial or financial relationships that could be a conflict of interest.
Do not take this as professional medical advice.
The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1028 words · about 5 minRead it →Close
A Chatbot Answered 79% of Late-Life Depression Questions Acceptably. In an Emergency, That Fell to 63%.
Researchers wrote 90 questions an older adult or a caregiver might ask, put them to three chatbots, and had psychiatrists score the replies without knowing which machine wrote them.

The article describes a study in which late-life depression is a major source of disability, functional dependence, and diminished quality of life among older adults, and it remains frequently under-recognized and undertreated.12 That gap is exactly where a chatbot gets used: an older person or an adult child types a question into a phone and gets an answer in seconds.
The study, published in Frontiers in Psychiatry by Wei Xiao, Huanyu Zhang, Xiaoyi Chen, Jun Cai, Xuchen Luo and Jiehua Deng, put 90 written questions to three widely used chatbots — ChatGPT, Gemini and Doubao — and had two psychiatrists score every reply against standards fixed in advance, with a third senior psychiatrist settling disagreements. The authors declare no financial support and no competing interests.
The article describes a study in which these are not search engines. They generate a direct reply that reads as if it were written for you personally, in a coherent and confident voice.3 The article describes a study in which the questions people bring them are ordinary and human: is this sadness, or depression; should this medicine be stopped; is forgetting names just part of getting older; how soon does this need a doctor.4
Judged by the study's combined standard — accurate, safe, no serious safety error, urgency neither played down nor exaggerated, and the age-related factors that change the answer all taken into account — ChatGPT was acceptable on 78.9% of questions, Gemini on 72.2% and Doubao on 60.0%.5 The gap between the top two was not statistically significant; both were ahead of Doubao.
Serious safety errors — the kind that could change what a family does next — appeared in 5.6% of ChatGPT answers, 10.0% of Gemini's and 16.7% of Doubao's.6 Complete attention to the things that make an older patient different — memory change, several illnesses at once, several medicines at once, frailty, falls, poor eating, dependence on a caregiver — was reached in 64.4%, 55.6% and 43.3% of answers.7
Then the researchers split the questions by how urgent the correct answer was. On the 30 low-risk questions, all three did reasonably well. On the 30 high-risk ones — the scenarios where minutes matter — ChatGPT was acceptable on 63.3%, Gemini on 50.0% and Doubao on 36.7%.5

The most frequent serious error involved a failure to recognize or respond to risk.6 On the 30 crisis questions, the elements most often left out were advice that the older adult should not remain alone and explicit involvement of a trusted caregiver.8
Ask the same question twice in two separate chats and the answers should match. They did 90.0% of the time for ChatGPT, 83.3% for Gemini and 73.3% for Doubao.9 Another way to picture that: roughly one question in ten, one in six, and one in four came back differently the second time, in ways that could change what a family does.
Read the numbers with the design in view. The 90 questions were written by the research team, not collected from real older adults or caregivers, and nobody measured whether a real person understood the reply or acted on it.10 Every question was in English, and each model was reached through one public app version during July 2026 — so these results do not tell you how the same chatbot behaves in Spanish, or in a version released since.11 The question set was deliberately balanced across risk levels and clinical areas as a stress test, not drawn to match what people actually ask, so the overall percentages are not a picture of everyday use. The study had limited power to catch rare safety events or small differences between models, and the bar for "acceptable" was set conservatively by the authors, not validated outside this study.
Here is how we read it. A fluent, well-organised answer feels like expertise, and fluency is not the same thing as being right. In this study the weaker answers would not have announced themselves as weaker — they would have read just as smoothly as the good ones. So tone is no guide. What separates a usable answer from an unfinished one is whether it tells you what to do, how soon, and who else needs to be involved. You can test that yourself in about a minute: ask the same question twice in two separate chats and compare. If the urgency or the advice shifts between the two replies, you have learned something the reply itself will never tell you — that this particular answer is not settled, and belongs in front of a person.

What could change this picture is the technology itself. The authors' own conclusion is not that these systems are unsafe, but that the evidence for relying on them is not there yet: consistent performance across versions, languages and real-world settings, with clinician oversight kept central.12
Their recommendation for now is narrow. These tools may help with selected low-risk education — understanding what depression can look like, preparing questions for a visit, finding reasons to seek care. They should not be the thing that decides how urgent a situation is, whether a medicine changes, or how a crisis is handled in an older adult.12
For a household in Latin America or in the United States or Canada, the practical version is this: a chatbot is a decent place to start organising a worry, and a poor place to finish one. Keep a pharmacist, a nurse line or a clinician in the loop before anything about mood, memory or medicines turns into an action. If you are the adult child watching a parent's appetite, sleep or interest in things slip, use the tool to prepare what you will ask — then ask a person.
Next time a chatbot gives you a calm, tidy answer about your mother's memory or your father's medication, ask it one thing before you act: when should this be seen, and by whom?
Where each piece of context comes from, and how much of it we read
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "Late-life depression is a major source of disability, functional dependence, and diminished quality of life among older adults (1)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "Despite these consequences, late-life depression remains frequently under-recognized and undertreated (3)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "General-purpose large language models are increasingly available as sources of conversational health information (16). Unlike conventional search tools, these systems generate direct responses that may appear individualized, coherent, and clinically authoritative (17)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "An older adult or caregiver can ask whether a symptom represents depression, whether a medicine should be stopped, whether memory decline is part of normal ageing, or how urgently professional help is required."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "Clinically acceptable responses were generated for 78.9% of questions by ChatGPT, 72.2% by Gemini, and 60.0% by Doubao (overall P<0.001)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "Major safety errors occurred in 5.6% of ChatGPT responses, 10.0% of Gemini responses, and 16.7% of Doubao responses (raw P = 0.015; FDR-adjusted q=0.023)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "Complete geriatric-specific appropriateness was achieved in 64.4%, 55.6%, and 43.3% of responses, respectively."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "Advice that the older adult should not remain alone, explicit involvement of a trusted caregiver, and reduction of access to potentially lethal medication or other means of harm were the elements most frequently omitted."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "Test-retest clinical consistency was highest for ChatGPT (90.0%), followed by Gemini (83.3%) and Doubao (73.3%)."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "The benchmark consisted of standardized scenarios rather than questions entered prospectively by older adults or caregivers during actual healthcare decision-making."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "All questions were presented in English. This improved standardization across models but may have influenced the performance of systems optimized for other languages or cultural contexts."
- Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736 - the article this story is about — the whole article — the passage: "Until these capabilities are demonstrated consistently across model versions, languages, and real-world settings, clinician oversight should remain central to the use of large language models in late-life depression care."
Xiao, W., Zhang, H., Chen, X. et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry. https://doi.org/10.3389/fpsyt.2026.1956736
Who paid: The authors declared that no financial support was received for this work or its publication, and they declared no commercial or financial relationships that could be a conflict of interest.
Do not take this as professional medical advice.
The findings of other studies mentioned here are known to us through this document, which is the one we read; we did not open each of those studies.
