weeklyAI · Week of 2 October 2026weeklyAI · Semana del 2 de octubre de 2026

← Your rights← Sus derechos

experiment · Journal of Experimental Orthopaedics · la publicación, 26 jun 2026 · gratis

Un chatbot local dio contenido dañino en 9 de 20 respuestas e inventó información en 11 de 20

El estudio probó cuatro modelos con 20 preguntas escritas después de una cirugía ortopédica. Ninguna persona real participó, así que no dice cómo le iría a un paciente de verdad.

Versión breve · la versión detallada sigue, unos 8 min

Pregunte al chat de weeklyAI

Pregúnteme por este estudio: a quiénes se estudió, qué encontró y qué no dice.

Las conversaciones se guardan mientras exista weeklyAI, para mejorar la publicación. Se responde en el idioma en que usted escribe.

El estudio, de un vistazo
Quiénes
Cuatro programas de inteligencia artificial (dos comerciales y dos locales) evaluados con preguntas posoperatorias de ortopedia
Cuántos
20 preguntas posoperatorias escritas
Dónde
Hospital Universitario Balgrist, Zúrich, Suiza
Cuándo
No lo dice el pasaje
Tipo de estudio
experiment
Quién lo hizo
Departamento de Cirugía Ortopédica, Hospital Universitario Balgrist, Zúrich, Suiza
El límite que importa
Se usaron 20 preguntas escritas y un solo texto de instrucciones; no participaron pacientes reales
Calificación general de cada programa de inteligencia artificial en una escala de 0 a 1
Claude 4.5 Sonnet0.949puntos
GPT-50.937puntos
GPT-OSS0.873puntos
Apertus0.693puntos

Calificación media de desempeño general de cada modelo, en una escala de 0 a 1, según cuatro revisores cegados; se trata de respuestas a 20 preguntas escritas, no de la atención de pacientes reales.

Desempeño general de los modelos comerciales frente a los locales

Claude 4.5 Sonnet (comercial)frente aGPT-OSS (local)

Claude 4.5 Sonnet obtuvo una calificación significativamente más alta

Claude 4.5 Sonnet (comercial)frente aApertus (local)

Claude 4.5 Sonnet obtuvo una calificación significativamente más alta

GPT-5 (comercial)frente aGPT-OSS (local)

GPT-5 obtuvo una calificación significativamente más alta

GPT-5 (comercial)frente aApertus (local)

GPT-5 obtuvo una calificación significativamente más alta

Claude 4.5 Sonnet (comercial)frente aGPT-5 (comercial)

La diferencia entre ambos no resultó concluyente

Los problemas de seguridad se concentraron en Apertus: contenido dañino y contenido inventado aparecieron cada uno en el 22.5% de las calificaciones evaluadas.
Lectura de weeklyAI

Investigadores del Hospital Universitario Balgrist, en Zúrich, Suiza, escribieron veinte preguntas que un paciente suele hacer después de una operación de cadera, rodilla, hombro o columna. Se las hicieron a cuatro programas de inteligencia artificial y cuatro revisores las calificaron sin saber qué programa había respondido.

El modelo Apertus, que funciona dentro del hospital y no en internet, dio contenido dañino en 9 de las 20 respuestas. Inventó información en 11 de las 20. Son cifras que los autores del estudio reportan como respuestas del programa, no como daño a ningún paciente.

Los otros dos modelos que funcionan fuera del hospital, GPT-5 y Claude 4.5 Sonnet, obtuvieron las mejores calificaciones, por encima de 0.93 en una escala que va de 0 a 1. El cuarto modelo, GPT-OSS, también local, quedó en el medio, con 0.873. Apertus quedó último, con 0.693. Las diferencias entre GPT-5 y Claude no fueron claras; las diferencias entre los locales y los comerciales sí.

Entre los revisores hubo acuerdo moderado al calificar, lo que significa que no siempre coincidieron.

El estudio no usó ningún dato de paciente y no midió si un modelo local protege mejor la información médica que uno comercial. Tampoco probó que entrenar estos programas con material de ortopedia los mejore; eso queda como una idea de los autores.

Se trató de 20 preguntas escritas y de un solo texto de instrucciones, que pedía respuestas cortas, con señales de alarma y tono tranquilo. Sin esas instrucciones, los resultados podrían ser distintos. Las versiones de estos programas cambian rápido.

Lo que todavía no se sabe es si un modelo que funcione dentro de un hospital podría algún día dar consejos seguros sobre una recuperación, sin que la información del paciente salga de ahí.

Qué significa para usted

El estudio no usó ningún dato de pacientes y no midió si un modelo local protege mejor su información médica. Usted puede preguntar a su médico si el hospital planea usar estos programas y cómo verifica sus respuestas antes de que lleguen a un paciente. Mientras eso no esté claro, conviene no tomar como consejo médico lo que responda un chatbot sobre una recuperación.

Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813

Quién pagó: Los autores no tienen financiación que declarar; ningún financiador tuvo influencia, y no se prestó equipo ni software.

No tome esto como consejo médico profesional.

Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1606 palabras · unos 8 minLeerla →Cerrar

Un hospital probó cuatro chatbots con preguntas de recuperación posoperatoria. Los modelos comerciales respondieron mejor; uno de los locales inventó o resultó dañino en una de cada cinco respuestas evaluadas.

El estudio no usó pacientes reales ni datos clínicos, así que no dice cómo le iría a usted. Sí deja una pregunta útil para hacer donde vive.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

Un equipo del Departamento de Cirugía Ortopédica del Hospital Universitario Balgrist, en Zúrich, Suiza, escribió veinte preguntas posoperatorias con un caso clínico breve y una pregunta tal como la formularía un paciente. Las envió a cuatro modelos de lenguaje: dos comerciales, alojados en la nube, y dos locales, ejecutados dentro de la institución. Cuatro revisores cegados —dos cirujanos ortopédicos certificados y dos residentes de ortopedia— calificaron cada respuesta. Ningún paciente participó y no se introdujo ningún dato de salud identificable en ningún modelo.

Vale la pena entender qué es lo que se probó, porque no es un buscador ni una hoja de instrucciones. Un modelo de lenguaje genera texto abierto: no elige entre respuestas escritas de antemano, sino que arma cada frase en el momento1. Por eso puede sonar seguro de sí mismo y, a la vez, equivocarse. El artículo describe un estudio en el que la inteligencia artificial generativa se usa cada vez más en salud y puede aliviar la carga inicial del personal, pero una respuesta no validada puede poner en riesgo al paciente si es inexacta, incompleta o contradice las recomendaciones2. La diferencia entre un modelo alojado en la nube y uno local es dónde ocurre ese cálculo: los locales procesan dentro del perímetro de seguridad del hospital y permiten integrarse con sus sistemas y protocolos internos, aunque hasta ahora no estaba claro cómo rendían frente a los comerciales3.

El resultado principal fue una calificación de desempeño general. Claude 4.5 Sonnet obtuvo la media más alta, 0.949, seguido de cerca por GPT-5, con 0.937; después GPT-OSS, con 0.873, y bastante más atrás Apertus, con 0.6934. Los modelos difirieron de forma significativa. La distancia entre Claude 4.5 Sonnet y GPT-5 fue tan pequeña que no resultó concluyente. Las diferencias entre esos dos y los locales sí fueron claras.

El hallazgo que debería detener a cualquiera antes de instalar un chatbot en una consulta es otro. Los problemas de seguridad se concentraron en Apertus: contenido dañino y contenido inventado aparecieron cada uno en el 22.5% de las calificaciones evaluadas. Llevado al nivel de las respuestas mismas, hubo contenido dañino en 9 de las 20 respuestas únicas del modelo, y fabricación en 11 de las 205. En la discusión del estudio, los autores señalan que esa tasa de contenido dañino o inventado genera preocupaciones serias de seguridad y va en contra de usarlo sin validación en la comunicación posoperatoria con pacientes6.

El mecanismo que produce ese tipo de error no es misterioso. Cuando un modelo genera texto abierto en lugar de recuperar una respuesta ya escrita, puede completar un hueco con algo plausible pero falso1. En salud, ese hueco puede ser una dosis, un plazo o una señal de alarma. Y las consecuencias no son teóricas: un trabajo que solo pudimos leer en su resumen documenta que, entre intervenciones con chatbots de inteligencia artificial generativa, se registraron fallas de seguridad, como entregar información clínica inexacta7. En ese mismo conjunto de estudios, la supervisión humana resultó limitada, los protocolos de derivación en crisis variaban en rigor y en general estaban poco desarrollados, y el monitoreo sistemático de eventos adversos era escaso89. Hay que decirlo con la misma claridad con que lo dicen los autores: el contenido dañino se observó en las respuestas de un modelo, no en la atención de pacientes.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

El otro documento que leímos, también solo en su resumen, describe un caso clínico en el que la interacción con un chatbot pareció corroborar y reforzar el contenido delirante de un paciente y contradecir el consejo médico10. El equipo tratante incorporó, como parte del manejo, un plan de cuidado que restringía el uso del chatbot11. No podemos saber en qué medida, si acaso alguna, esas respuestas contribuyeron al cuadro que el paciente ya presentaba12. Lo traemos porque muestra el otro extremo de la misma moneda: un sistema que responde con seguridad sobre temas de salud no es lo mismo que un sistema que responde bien.

Hay una razón práctica por la que los hospitales miran los modelos locales, y es la privacidad. Los modelos comerciales suelen alojarse en la nube, lo que puede dejar información de salud protegida fuera del control de la institución y complicar la retención de datos, la auditabilidad y el cumplimiento normativo13. Para escenarios clínicos futuros con datos identificables de pacientes, habría que considerar requisitos regulatorios como la ley suiza de protección de datos y el reglamento europeo14. Ahora el detalle que importa: en este estudio no se usó ningún dato identificable ni información de salud protegida, así que las implicaciones de gobernanza y privacidad no se probaron empíricamente15. Dicho de otro modo: nadie midió aquí si un modelo local protege mejor sus datos que uno comercial.

Los límites del estudio son varios y conviene tenerlos a mano. Se usaron veinte preguntas escritas, un conjunto fijo, que puede no capturar la complejidad de las consultas reales, incluidas las que incluyen una foto de la herida16. Los modelos evolucionan rápido y los resultados pueden cambiar con versiones futuras17. Todos se evaluaron con la configuración predeterminada; ajustar parámetros podría dar perfiles distintos18. Y el formato fijo de la pregunta, aunque mejoró la consistencia interna, pudo introducir sesgos en varias dimensiones evaluadas19. Los propios autores advierten que el límite de cinco frases pudo favorecer la brevedad y la claridad a costa de la completitud y los matices.

Así lo leemos nosotros. Un hospital puede encerrar un modelo dentro de sus paredes y sentir que con eso ya protegió a sus pacientes. No es lo mismo. Lo que queda adentro son los datos; lo que sale hacia el paciente es el consejo, y ese consejo no mejora por el hecho de haberse generado en un servidor propio. Si un modelo local no fue entrenado con las guías y los protocolos de esa institución, sus respuestas no reflejarán las necesidades reales de sus pacientes, y ahí es donde puede fallar en temas de seguridad. Esperaríamos ver esto en cualquier clínica que instale un chatbot sin ajustarlo antes: respuestas fluidas, seguras de sí mismas, y equivocadas en el detalle que importa. Sabríamos que nos equivocamos si un modelo local, sin entrenamiento específico, igualara a los comerciales en precisión y en ausencia de invenciones. Mientras eso no se demuestre, cuando un hospital o una clínica le ofrezca un chatbot para preguntas después de una cirugía, pregunte si ese sistema fue entrenado con los protocolos y las guías de esa institución, si fue validado para su tipo de cirugía, y quién responde por sus errores. Si no hay respuestas claras, lleve sus dudas directamente a su equipo de salud.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

También leímos, y esto es nuestra lectura, que la comodidad de tener todo dentro de casa puede volverse una falsa sensación de protección. Que los datos no salgan no garantiza que el consejo sea bueno. Y hay un patrón conocido en esta familia de sistemas: quien aconseja debería demostrar que estudió para hacerlo. Un modelo local que no fue entrenado a fondo con literatura ortopédica y protocolos institucionales no debería usarse para aconsejar a pacientes, aunque sea fácil de instalar. Lo que sigue, según los propios autores, es avanzar en los modelos locales mediante ajuste con datos específicos, validación rigurosa y barreras de seguridad robustas20. Eso es una hipótesis de trabajo, no un resultado: nadie probó aquí que ajustar un modelo local lo vuelva mejor.

Otros trabajos que leímos coinciden en el rumbo. Entre intervenciones con chatbots de inteligencia artificial generativa, la mayoría incorporó al menos un mecanismo técnico de seguridad, sobre todo ajuste con datos específicos e ingeniería de instrucciones21, y un subconjunto menor armó capas de protección combinando sistemas de recuperación de información, filtros de contenido o clasificadores de riesgo22. La conclusión de ese conjunto de estudios es que estas herramientas requieren un enfoque que integre salvaguardas técnicas con el diseño junto a los usuarios, controles de procedimiento y supervisión humana23, y que hace falta evaluar su eficacia, mejorar las protecciones y estandarizar cómo se miden los resultados de seguridad24. Traducido a la vida diaria: la pregunta no es si el chatbot existe, sino quién lo revisa.

En el caso concreto que leímos, el equipo clínico consideró apropiado evaluar el uso de chatbots dentro de la evaluación clínica y, donde estuviera indicado, implementar intervenciones para mitigar los efectos asociados25. Es un criterio que suena razonable para cualquier servicio de salud, y no depende de que el modelo sea local o comercial.

Aquí está lo que este estudio le deja hacer, y es poco y es mucho a la vez. No puede usar estas cifras para juzgar a su propio médico ni para decidir si un chatbot de su país es seguro: se probaron veinte preguntas de texto en un solo hospital suizo, con instrucciones fijas y orientadas a la seguridad. Lo que sí puede hacer es llevar tres preguntas a la próxima vez que alguien le ofrezca una herramienta así. ¿Con qué guías de esta institución fue entrenada? ¿Fue validada para mi tipo de cirugía? ¿Quién revisa sus respuestas y quién responde si se equivoca? Y una más, para usted mismo: si la respuesta de un chatbot le preocupa, no reemplace con ella a su equipo de salud. La próxima vez que vea un anuncio de un asistente de salud "que funciona dentro del hospital", pregunte qué se validó y con qué.

De dónde sale cada dato de contexto, y cuánto leímos de cada documento

  1. Olisaeloka L, Richardson CG, Wang AY, Munthali RJ, Vigo DV. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare. https://doi.org/10.3390/healthcare14101395 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Unlike earlier rule-based and retrieval-based systems, GenAI chatbots generate open-ended outputs that can be inaccurate and unsafe."
  2. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Generative artificial intelligence (AI), particularly large language models (LLMs), are increasingly utilised in healthcare. While this may reduce the initial workload of healthcare professionals, unvalidated model outputs can pose a relevant risk to patient safety if they are inaccurate, incomplete or inconsistent with recommendations."
  3. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Locally deployed (on‐premise) LLMs may address these barriers by keeping processing within the hospital security perimeter while enabling integration with internal systems and approved clinical protocols; however, their performance relative to commercial models remains unclear."
  4. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The highest overall performance scores were observed for Claude 4.5 Sonnet (mean 0.949, SD 0.124) and GPT‐5 (0.937, 0.142), followed by GPT‐OSS (0.873, 0.189), while Apertus performed worst (0.693, 0.317)."
  5. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Safety‐relevant issues were concentrated in Apertus, with harmful content and fabrication each occurring in 22.5% of evaluated ratings (18/80). At the output level, this corresponded to harmful content in 9/20 unique model responses, and fabrication in 11/20 unique model responses."
  6. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The observed rate of harmful or fabricated content in Apertus outputs raises substantial safety concerns and argues against unvalidated deployment in patient‐facing postoperative communication."
  7. Olisaeloka L, Richardson CG, Wang AY, Munthali RJ, Vigo DV. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare. https://doi.org/10.3390/healthcare14101395 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Documented safety failures included missed suicidal ideation and provision of inaccurate clinical information."
  8. Olisaeloka L, Richardson CG, Wang AY, Munthali RJ, Vigo DV. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare. https://doi.org/10.3390/healthcare14101395 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "During intervention delivery, detailed onboarding with role clarification was common, but human oversight was limited."
  9. Olisaeloka L, Richardson CG, Wang AY, Munthali RJ, Vigo DV. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare. https://doi.org/10.3390/healthcare14101395 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Crisis referral protocols varied in rigour but were mostly underdeveloped, and systematic adverse event monitoring was sparse."
  10. Shah S, Morrin H. (2026). Substance-induced manic psychosis in which delusions were corroborated by a chatbot - case report. BMC Psychiatry. https://doi.org/10.1186/s12888-026-08137-3 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "This case describes a substance-induced manic episode with psychotic features in which interaction with an AI (artificial intelligence) chatbot appeared to corroborate and reinforce the patient's delusional thought content and to contradict medical advice."
  11. Shah S, Morrin H. (2026). Substance-induced manic psychosis in which delusions were corroborated by a chatbot - case report. BMC Psychiatry. https://doi.org/10.1186/s12888-026-08137-3 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Behavioural management included implementation of a care plan restricting AI chatbot use, as a form of environmental containment."
  12. Shah S, Morrin H. (2026). Substance-induced manic psychosis in which delusions were corroborated by a chatbot - case report. BMC Psychiatry. https://doi.org/10.1186/s12888-026-08137-3 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "The AI chatbot reportedly affirmed his perceived "spiritual awakening," minimised the possibility that his presentation represented a manic episode, and provided medical advice, including discouragement of prescribed antipsychotic medication, though it cannot be determined to what extent, if any, these statements contributed to his existing presentation."
  13. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Moreover, commercially available LLMs are usually cloud‐hosted, potentially placing protected health information outside institutional governance and complicating data retention, auditability and compliance."
  14. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "For future clinical deployment scenarios involving identifiable patient data, regulatory requirements under the Swiss nFADP and the EU GDPR will need to be considered."
  15. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "No patient‐identifiable or protected health information was used in this study; therefore, data‐governance and privacy implications were not empirically tested."
  16. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This study has several limitations. First, the study used a static set of 20 text‐based postoperative questions, which may not fully capture the complexity of real‐world patient inquiries, including multimodal concerns such as wound photographs."
  17. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Second, given the rapid evolution and potential obsolescence of LLM versions, the reported benchmarks may change with future model updates."
  18. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Third, all models were evaluated using default runtime settings; modification of parameters such as temperature or compute allocation could yield different performance profiles."
  19. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Finally, the standardised prompt, while improving internal consistency, may have introduced prompt‐induced bias across several QUEST domains."
  20. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Achieving privacy‐compliant, real‐time clinical integration will require advancing local LLMs through fine‐tuning, rigorous validation and robust safety guardrails."
  21. Olisaeloka L, Richardson CG, Wang AY, Munthali RJ, Vigo DV. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare. https://doi.org/10.3390/healthcare14101395 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Twenty-one studies across 11 countries were included. Most interventions incorporated at least one technical safety mechanism, most commonly fine-tuning and prompt engineering."
  22. Olisaeloka L, Richardson CG, Wang AY, Munthali RJ, Vigo DV. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare. https://doi.org/10.3390/healthcare14101395 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "A smaller subset implemented layered safety architectures combining retrieval systems, content filters or risk classifiers, and rule-based algorithms."
  23. Olisaeloka L, Richardson CG, Wang AY, Munthali RJ, Vigo DV. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare. https://doi.org/10.3390/healthcare14101395 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "GenAI chatbot interventions require a robust sociotechnical approach that integrates technical safeguards with user co-design, procedural controls, and human oversight."
  24. Olisaeloka L, Richardson CG, Wang AY, Munthali RJ, Vigo DV. (2026). Safety Mechanisms and Risk Mitigation in Generative AI Mental Health Chatbots: A Systematic Scoping Review. Healthcare. https://doi.org/10.3390/healthcare14101395 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Future research is needed to evaluate efficacy, improve safeguards and standardise safety outcome measurement."
  25. Shah S, Morrin H. (2026). Substance-induced manic psychosis in which delusions were corroborated by a chatbot - case report. BMC Psychiatry. https://doi.org/10.1186/s12888-026-08137-3 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "As AI chatbot use becomes increasingly widespread, clinicians should consider assessing their use and impact within clinical assessments and, where clinically indicated, implementing interventions to mitigate asso"

Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813

Quién pagó: Los autores no tienen financiación que declarar; ningún financiador tuvo influencia, y no se prestó equipo ni software.

No tome esto como consejo médico profesional.

Los hallazgos de otros estudios que aquí se mencionan los conocemos por este documento, que fue el que leímos; no abrimos cada uno de esos estudios.

experiment · Journal of Experimental Orthopaedics · the paper, 26 Jun 2026 · free

A Chatbot Gave Made-Up Advice in 11 of 20 Test Questions

Swiss researchers tested four AI models on questions patients ask after surgery. One local model did badly, and no one tested real patients or your privacy.

Short version · the longer version follows, about 5 min

Ask the weeklyAI chat

Ask me about this study: who was studied, what it found, and what it does not say.

Conversations are saved for as long as weeklyAI exists, to improve the publication. Answers come in the language you write in.

The study at a glance
Who
Four AI chatbots answering made-up questions about recovery after orthopaedic surgery
How many
20 questions
Where
Zurich, Switzerland
When
2026
Kind of study
experiment
Who did it
Department of Orthopaedic Surgery, Balgrist University Hospital, Zurich
The limit that matters
Only 20 written questions, one fixed prompt, no patients, privacy not tested
Average scores of four chatbots on 20 questions about recovery after surgery
Claude 4.5 Sonnet0.949score out of 1
GPT-50.937score out of 1
GPT-OSS0.873score out of 1
Apertus0.693score out of 1

Average scores given by four reviewers to each chatbot's answers; higher is better, but the test used only 20 written questions and one fixed prompt.

Commercial chatbots compared with locally hosted chatbots

Commercial models (Claude 4.5 Sonnet and GPT-5)againstLocally hosted models (GPT-OSS and Apertus)

Commercial models scored higher overall and had fewer safety-relevant ratings, but they remain externally controlled.

The study did not test whether a local model would keep your health information more private than a commercial one.
weeklyAI's reading

One local model, Apertus, gave made-up information in 11 of the 20 questions. It gave harmful content in 9 of the 20.

Researchers at Balgrist University Hospital in Zurich wrote 20 questions that patients commonly ask after orthopaedic surgery, on the hip, knee, shoulder and spine. Then they put each question to four chatbots and asked four blinded reviewers to grade the answers. Two were senior surgeons, two were residents in training.

The two commercial models did best. Claude 4.5 Sonnet scored 0.949 out of 1 and GPT-5 scored 0.937. GPT-OSS, hosted locally, scored 0.873. Apertus, also hosted locally, scored 0.693. The scores are averages of the reviewers' grades. The four models differed by a clear margin overall.

Harmful content and made-up information were concentrated in Apertus. The study counted these as harmful answers, not harm done to patients. No patient was involved. No patient data was used.

The study did not test whether a local model would keep your health information more private than a commercial one. The authors note that commercial models sit outside a hospital's control, and that local ones could stay inside it. What that means for your own medical information is not known.

The study also used only 20 written questions and one fixed prompt that told every model to answer briefly, safely and in a set order. Without that prompt, the answers could look different.

Whether a locally hosted AI could ever match the commercial models for safe, private patient advice is not known. The authors suggest fine-tuning local models on orthopaedic literature and hospital guidelines. They did not test that.

What this means for you

The study was done in one Swiss hospital with made-up questions, so it says nothing about what a chatbot would tell you about your own recovery, and nothing about your privacy. If you hear that a local AI keeps your health information safer, remember that part was not tested. Ask your surgeon your questions, and treat any chatbot answer as something to check.

Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813

Who paid: The authors have no funding to report; no funders had any say, and no equipment or software was lent.

Do not take this as professional medical advice.

The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1097 words · about 5 minRead it →Close

One chatbot made things up in eleven of the twenty test questions about recovery.

The test was small and used no patients. It still shows how far apart the systems now answering health questions can be.

How it could look · illustration generated by weeklyAI.watch, not a photograph

Four artificial-intelligence chatbots were given the same twenty made-up questions about life after orthopaedic surgery — hip, knee, shoulder and spine — and four reviewers scored the answers without knowing which machine wrote which. The two commercial models, Claude 4.5 Sonnet and GPT-5, came out on top, with average scores of 0.949 and 0.937 out of a possible 1. The locally hosted GPT-OSS followed at 0.873, and the locally hosted Apertus came last at 0.693. The gap between the two leaders was not statistically significant after correction; every other gap was1.

The safety picture is the part a family should sit with. Harmful content and made-up content each appeared in 22.5 percent of Apertus's evaluated ratings — 18 of 80. Counted by answer rather than by rating, that meant harmful content in 9 of 20 answers and fabrication in 11 of 202. The authors write plainly that this rate "raises substantial safety concerns and argues against unvalidated deployment in patient-facing postoperative communication"3. No patients were involved anywhere in the study; the harm was in the text, not in anyone's care.

The authors are affiliated with the Department of Orthopaedic Surgery, Balgrist University Hospital, in Zurich, Switzerland. The reviewers were two board-certified, fellowship-trained orthopaedic surgeons and two orthopaedic surgery residents, all blinded to which model produced which answer4. No funding was reported; one author has received speaker fees from OPED AG, and the others declare no conflicts of interest. The paper appears in the *Journal of Experimental Orthopaedics*.

What the technology actually is matters here. A large language model does not look up an answer in a hospital manual. It produces text by predicting what should come next, which is why the same machine can sound equally confident when it is right and when it is inventing. The authors describe the risk in exactly those terms: unvalidated outputs "can pose a relevant risk to patient safety if they are inaccurate, incomplete or inconsistent with recommendations"4. The questions were all text-based, so the study never tested things like photographs of a wound5.

How it could look · illustration generated by weeklyAI.watch, not a photograph

The trade-off the study names is not speed against accuracy. It is control against performance. Commercial models won this test, but they run on someone else's computers, which the authors say can place protected health information "outside institutional governance and complicating data retention, auditability and compliance"6. A hospital that keeps the model on its own machines keeps the data inside its walls — that is the promise of local hosting7 — but in this test the local models were the ones making things up2.

One more thing the authors flag for anyone thinking about rules and rights: deploying these tools on identifiable patient data would bring Switzerland's revised data protection law and Europe's GDPR into play8. That is a statement about what would have to be considered, not a test the study ran — the authors say directly that privacy implications were not empirically examined because no real patient information was used9.

The limits are as important as the scores. Twenty text-based questions cannot capture what real patients actually send — including photographs of a wound, which the study never tested5. Model versions change fast, so these benchmarks may not survive the next update10. Everything ran at default settings, and different settings could produce different results11. And the prompt itself — a fixed structure of acknowledgement, answer, actions, warning signs and reassurance, capped at five sentences — may have shaped how reviewers judged tone, safety and trust12. This is model behaviour under one carefully written prompt, not what happens when a worried person types a question at midnight.

Here is how we read it. When a machine answers your question about your own recovery, you cannot see how it chose that answer, and its mistakes are hard to spot from the outside — a confident sentence and a correct one look identical on a screen. That is why the difference between these four systems is worth your attention even though the test was small: the two that scored best are the two you cannot inspect, and the two you could inspect were the ones inventing things. We expect that in homes like yours this will show up first as a quiet substitution — a chatbot consulted before the discharge sheet, or instead of the phone call to the surgical team — and that the confidence of the answer, not its accuracy, will be what decides whether it is believed. You would know we are wrong if people using these tools kept checking their own symptoms and kept calling their surgeons at least as often as before. The useful thing to carry away is a habit: ask who runs the tool, ask what it was built on, and ask whether a clinician has checked its advice. Write down what it told you, and confirm anything that matters with a qualified professional.

How it could look · illustration generated by weeklyAI.watch, not a photograph

We also read work on this class of system more broadly — we could read only the summaries of these papers, the full texts sit behind subscriptions. None of this is a verdict on the Zurich study. It is the neighbourhood the study lives in, and it is why the authors' own conclusion — that local models need fine-tuning, rigorous validation and robust safety guardrails before they can be trusted in clinical workflows13 — reads less like a research agenda than a warning label.

The authors also point out something easy to miss: the local models lost here, but that does not settle the question. Their hypothesis is that a smaller model trained specifically on orthopaedic literature, hospital protocols and rehabilitation guidelines could outperform a general commercial one — untested, and stated as a hypothesis13. The commercial advantage, meanwhile, comes with the condition the study itself names: those models "remain externally controlled"14.

For a reader in the United States, Canada or anywhere in Latin America, the practical question is not which model won a Swiss benchmark. It is what you are entitled to know when a health system starts using one of these tools on you. The study gives you no rights and no procedures — it was not designed to. What it gives you is a question worth asking your own hospital, clinic or insurer before the tool arrives: which system are you using, where does it run, who reviewed its answers, and who is accountable when it is wrong?

Where each piece of context comes from, and how much of it we read

  1. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "The highest overall performance scores were observed for Claude 4.5 Sonnet (mean 0.949, SD 0.124) and GPT‐5 (0.937, 0.142), followed by GPT‐OSS (0.873, 0.189), while Apertus performed worst (0.693, 0.317)."
  2. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "Safety‐relevant issues were concentrated in Apertus, with harmful content and fabrication each occurring in 22.5% of evaluated ratings (18/80). At the output level, this corresponded to harmful content in 9/20 unique model responses, and fabrication in 11/20 unique model responses."
  3. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "The observed rate of harmful or fabricated content in Apertus outputs raises substantial safety concerns and argues against unvalidated deployment in patient‐facing postoperative communication."
  4. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "Generative artificial intelligence (AI), particularly large language models (LLMs), are increasingly utilised in healthcare. While this may reduce the initial workload of healthcare professionals, unvalidated model outputs can pose a relevant risk to patient safety if they are inaccurate, incomplete or inconsistent with recommendations."
  5. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "This study has several limitations. First, the study used a static set of 20 text‐based postoperative questions, which may not fully capture the complexity of real‐world patient inquiries, including multimodal concerns such as wound photographs."
  6. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "Moreover, commercially available LLMs are usually cloud‐hosted, potentially placing protected health information outside institutional governance and complicating data retention, auditability and compliance."
  7. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "Locally deployed (on‐premise) LLMs may address these barriers by keeping processing within the hospital security perimeter while enabling integration with internal systems and approved clinical protocols; however, their performance relative to commercial models remains unclear."
  8. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "For future clinical deployment scenarios involving identifiable patient data, regulatory requirements under the Swiss nFADP and the EU GDPR will need to be considered."
  9. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "No patient‐identifiable or protected health information was used in this study; therefore, data‐governance and privacy implications were not empirically tested."
  10. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "Second, given the rapid evolution and potential obsolescence of LLM versions, the reported benchmarks may change with future model updates."
  11. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "Third, all models were evaluated using default runtime settings; modification of parameters such as temperature or compute allocation could yield different performance profiles."
  12. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "Finally, the standardised prompt, while improving internal consistency, may have introduced prompt‐induced bias across several QUEST domains."
  13. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "Achieving privacy‐compliant, real‐time clinical integration will require advancing local LLMs through fine‐tuning, rigorous validation and robust safety guardrails."
  14. Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813 - the article this story is about — the whole article — the passage: "In this structured evaluation setting, commercial LLMs showed higher overall performance and fewer safety‐relevant ratings, but they remain externally controlled."

Lanter, L., Masel, S., Hochreiter, B. et al. (2026). Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis. Journal of Experimental Orthopaedics. https://doi.org/10.1002/jeo2.70813

Who paid: The authors have no funding to report; no funders had any say, and no equipment or software was lent.

Do not take this as professional medical advice.