weeklyAI · Week of 2 October 2026weeklyAI · Semana del 2 de octubre de 2026

← Your students← Sus estudiantes

experiment · Frontiers in Education · la publicación, 1 sep 2026 · gratis

Una IA calificó los mismos exámenes cinco días seguidos y no se dio la misma nota

En un curso de salud pública de Noruega, ChatGPT se acercó a los profesores la mitad de las veces, pero cambió de criterio entre sesiones y casi no puso notas extremas.

Versión breve · la versión detallada sigue, unos 6 min

Pregunte al chat de weeklyAI

Pregúnteme por este estudio: a quiénes se estudió, qué encontró y qué no dice.

Las conversaciones se guardan mientras exista weeklyAI, para mejorar la publicación. Se responde en el idioma en que usted escribe.

El estudio, de un vistazo
Quiénes
Exámenes escritos de una maestría en salud pública, corregidos por cuatro programas de inteligencia artificial y dos docentes humanos
Cuántos
32 exámenes
Dónde
Noruega, Universidad Metropolitana de Oslo
Cuándo
No lo dice el pasaje
Tipo de estudio
analysis of what people did
Quién lo hizo
Universidad Metropolitana de Oslo, Noruega
El límite que importa
Solo 32 exámenes y dos correctores humanos: no se puede generalizar
Cuánto coincidió cada programa con la nota exacta de los docentes
ChatGPT50%
Kimi28.1%

Porcentaje de exámenes en que la nota puesta por el programa fue exactamente la misma que la del docente, con un solo archivo a la vez; son solo dos programas de los cuatro probados y el estudio es pequeño.

Qué tan cerca quedó cada programa de la nota humana

ChatGPTfrente aDocentes humanos

Coincidió en la nota exacta en la mitad de los casos y quedó a lo sumo a un escalón en nueve de cada diez

Kimifrente aDocentes humanos

Coincidió en la nota exacta en poco más de uno de cada cuatro casos y quedó a lo sumo a un escalón en casi ocho de cada diez

Programas en modo rápidofrente aProgramas en modo de razonamiento

El acuerdo con los docentes fue escaso en modo rápido y mejoró en modo de razonamiento

Programas de inteligencia artificialfrente aDocentes humanos

Los programas repartieron menos notas extremas que los docentes

ChatGPT en la parte baja de la escalafrente aDocentes humanos

ChatGPT tendía a poner notas más altas que los docentes

ChatGPT en la parte alta de la escalafrente aDocentes humanos

ChatGPT tendía a poner notas más bajas que los docentes

La IA puede servir como un segundo lector bajo supervisión: alguien que revisa después, no quien decide.
Lectura de weeklyAI

Usted corrige exámenes largos y conoce el problema: mantener el mismo criterio entre un trabajo y otro, entre un día y otro, entre un colega y usted. Investigadores de la Universidad Metropolitana de Oslo, en Noruega, probaron si una inteligencia artificial puede compartir esa carga.

Tomaron 32 exámenes anonimizados de un curso de maestría en salud pública, con respuestas largas, y los entregaron a ChatGPT, Gemini, LeChat y Kimi junto con la misma guía de calificación que usaron los profesores.

Cuando ChatGPT trabajó en su modo de razonar, coincidió con la nota humana exacta en la mitad de los casos. Si se aceptaba una nota de diferencia, llegó a nueve de cada diez. Kimi quedó más atrás: coincidió exactamente en poco más de uno de cada cuatro casos y con una nota de margen en casi ocho de cada diez. En el modo rápido, todos se alejaron bastante de los profesores.

El detalle incómodo vino después. El mismo ChatGPT calificó los mismos exámenes en cinco días distintos y coincidió consigo mismo sólo de manera moderada. Y hay un patrón que se repite: la IA casi no puso las notas más altas ni las más bajas, mientras los profesores sí las usaron. La IA tiende a quedarse en el medio.

El estudio es pequeño: 32 exámenes y dos profesores como referencia. La guía estaba escrita para personas, no para máquinas, y los archivos en PDF traían texto de más que pudo distraer al programa. Los autores tampoco controlaron los ajustes internos del modelo, así que no pueden descartar cambios en el sistema. Nada de esto permite decir que la IA califique bien en su curso, ni que reemplace a un profesor.

Lo que sí sugiere es un uso acotado. La IA puede servir como un segundo lector bajo supervisión: alguien que revisa después, no quien decide. Si va a probarla, busque un modelo que muestre su razonamiento paso a paso y compare sus notas con las suyas antes de confiar en ellas.

Vale la pena preguntar, en su institución, quién revisa las calificaciones que produce una máquina y cómo se le avisa al estudiante que una IA participó.

Qué significa para usted

Ese resultado no le dice que la inteligencia artificial pueda calificar en su curso ni que reemplace a un profesor. Antes de confiar en ella, pida que muestre su razonamiento y compare sus notas con las suyas durante varias sesiones, porque el mismo programa cambió de criterio entre un día y otro. Pregunte también quién revisa las calificaciones que produce una máquina y cómo se le avisa al estudiante.

Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452

Quién pagó: Los autores declararon que no recibieron apoyo financiero para este trabajo ni para su publicación, y declararon no tener relaciones comerciales o financieras que pudieran constituir un conflicto de interés.

Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1275 palabras · unos 6 minLeerla →Cerrar

La IA que corrige exámenes escritos acierta a medias, y cambia de opinión según el día

Un estudio noruego comparó cuatro modelos de lenguaje con dos correctores humanos en pruebas de ensayo de una maestría en salud pública. La coincidencia fue moderada, y no se mantuvo igual entre sesiones.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

Cuatro programas de inteligencia artificial —ChatGPT, Gemini, LeChat y Kimi— recibieron la instrucción de poner una nota final en una escala de la A a la F a exámenes escritos de una maestría en salud pública, usando la misma guía de corrección que usan los examinadores humanos1. No se les mostró ninguna nota humana antes: cada modelo trabajó a ciegas, con las mismas instrucciones y el mismo formato de archivo, y debía justificar cada puntaje, contar palabras y señalar fortalezas y debilidades del trabajo en una frase2. Los resultados se compararon después con las notas que ya habían puesto dos docentes de la casa de estudios.

En la primera prueba, con un solo archivo a la vez, ChatGPT fue el que más se pareció a los correctores humanos, seguido de Kimi3. Traducido a algo que se puede imaginar: de cada diez exámenes, ChatGPT coincidió exactamente en la nota con los docentes en cinco, y en nueve de cada diez su nota quedó a lo sumo a un escalón de la humana, ni mejor ni peor4. Kimi acertó la nota exacta en menos de tres de cada diez, y quedó dentro de un escalón en casi ocho de cada diez4. Con el resto de los modelos la coincidencia fue menor.

La forma de trabajar del modelo también importó. En el modo rápido, el acuerdo con los humanos fue escaso; cuando los modelos pasaron a un modo de razonamiento más pausado, el acuerdo mejoró5. Aun así, la estabilidad se rompió en cuanto cambiaron las condiciones: repetir la misma corrección no dio siempre el mismo resultado5. Y hubo un patrón constante: los modelos repartieron menos notas extremas —menos A y menos F— que los docentes humanos6. En la parte baja de la escala tendían a subir la nota; en la parte alta, a bajarla, como si comprimieran todo hacia el medio7. El promedio general casi no se movió: los humanos promediaron 3,63 y cinco rondas de ChatGPT promediaron 3,73, una diferencia que los autores no consideran significativa8.

Estas cifras hay que leerlas con cuidado. El estudio usó solo 32 exámenes y dos correctores humanos, y los propios autores advierten que eso limita cuánto se puede generalizar9. La configuración interna de los modelos —el llamado parámetro de temperatura, que regula cuánto varía una respuesta— quedó en su valor por defecto y no se controló, lo que puede explicar por qué la misma corrección cambia de un día a otro10. Además, la guía de corrección estaba escrita para docentes humanos, no para una máquina, y puede no haber sido la más clara para el modelo11. Los exámenes se subieron como PDF con portadas, formatos y texto sobrante, lo que también pudo meter ruido.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

No es el primer intento de poner a una máquina a corregir. Otro estudio, del que solo pudimos leer el resumen porque el texto completo está detrás de una suscripción, comparó tres modelos con las notas de dos expertos en exámenes prácticos de anatomía, histología y fisiología, con imágenes y respuestas breves, en 309 estudiantes de medicina12. Ahí el orden de los puntajes coincidió bastante, pero a la hora de decidir quién aprueba y quién no la coincidencia se derrumbó: desde poco más de la mitad de los casos con un modelo hasta más de ocho de cada diez con otro1314. Y en un ítem con un modelo tridimensional de pelvis, casi todos los estudiantes aprobaron con la corrección humana mientras los tres modelos les pusieron cero15. La conclusión de ese trabajo es parecida a la de este: sirven como asistentes para una primera pasada, no para decidir solos16.

Hay otra familia de investigaciones que va en la dirección opuesta: detectar si un texto lo escribió una máquina. Un estudio del que también leímos solo el resumen comparó cuatro herramientas de detección con cuatro tipos de trabajos académicos, entre ellos textos escritos por humanos con pasajes generados por IA y textos de IA "humanizados" a propósito17. Una de las herramientas anduvo bastante mejor que las demás, mientras que las otras subestimaron el contenido generado por IA, sobre todo con el modelo más avanzado1819. Todas identificaron bien los textos plenamente humanos, y los falsos positivos fueron raros2021. Los propios autores advierten que estas herramientas no deberían ser la única prueba en decisiones de peso22.

Y está el problema de fondo que empuja toda esta carrera: la evaluación escrita a gran escala es lenta, cara y difícil de estandarizar entre docentes, momentos y grupos23. Los modelos de lenguaje aparecen entonces como una promesa de eficiencia, pero su desempeño cambia según el modelo, la tarea y la forma de pedirle las cosas23. Lo que este estudio muestra es que, en exámenes de ensayo de humanidades y salud, esa promesa todavía no se cumple de manera confiable.

Así lo leemos nosotros. Lo que deja huella cuando alguien con autoridad evalúa un trabajo no es solo el número: es el reconocimiento que acompaña al juicio, el gesto que sostiene la confianza de quien aprendió algo y lo empuja a seguir produciendo. Si esa tarea se delega a una máquina, es razonable esperar que el estudiante reciba una calificación parecida pero pierda algo que no aparece en ninguna planilla. ¿Cómo sabríamos que nos equivocamos? Si los alumnos calificados por máquina reportaran la misma confianza y las mismas ganas de seguir estudiando que los calificados por su docente, o si el docente conservara un gesto explícito de reconocimiento además del número automático. Usted puede probarlo en su propia aula: si usa apoyo automático para corregir, resérvese un momento para comentarle a cada estudiante algo de su trabajo, y pregúnteles qué recuerdan de las evaluaciones que los marcaron.

Así podría verse · ilustración generada por weeklyAI.watch, no es una fotografía

También leímos otros trabajos sobre este tipo de sistemas, y hay un riesgo que conviene mirar de frente: cuando una decisión importante se apoya en herramientas cuyos criterios no se explican, quien recibe una nota inesperada queda sin argumentos para discutirla. Es una preocupación de categoría, no un hallazgo de este estudio. Y hay un límite de alcance que conviene decir con todas las letras: este estudio se hizo en una sola universidad, en Noruega, con un solo curso, así que nada de esto dice qué pasaría en otro formato de examen. Que algo funcione en un formato no garantiza que funcione en otro.

¿Qué puede hacer usted con esto? Si es docente, antes de confiar una corrección a una máquina, compare lo que usted dice valorar en clase con lo que la herramienta premia, y anote las diferencias; esa lista es material para una reunión de colegas, no para un cajón. Si es estudiante, guarde sus borradores y pida saber con qué criterios se lo evaluó. Y si es directivo, pida que cualquier nota asistida por máquina venga con una justificación legible y un canal real de apelación: la diferencia entre un promedio razonable y un caso injusto suele estar en el examen individual, no en la estadística del curso.

Los autores concluyen que estos programas pueden servir como asistentes de corrección bajo supervisión humana, pero que hoy no son confiables para calificar solos en evaluaciones de alto impacto. El estudio se hizo en una sola universidad, en Noruega, con un solo curso, así que nada de esto dice qué pasaría en su institución. La pregunta que queda abierta es simple y urgente: cuando la nota de un examen escrito llega sin explicación, ¿quién responde por ella?

De dónde sale cada dato de contexto, y cuánto leímos de cada documento

  1. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Four LLMs—ChatGPT, Gemini, LeChat, and Kimi—were prompted to assign final grades on an A-F scale using the same grading guidance as human examiners."
  2. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The LLMs were also instructed to 4) summarize the total number of words and points from all three responses. Finally, the prompt instructed the LLMs to 5) provide a short overall justification for the final grade and 6) point out the major strengths and weaknesses of the submission in one sentence."
  3. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "In the single-file upload setting, ChatGPT showed the strongest agreement with human examiners (weighted kappa 0.718, 95% CI 0.578–0.859), followed by Kimi (weighted kappa 0.571, 95% CI 0.373–0.768)."
  4. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "ChatGPT achieved 50.0% exact agreement and 90.6% agreement within ±1 grade, with corresponding values of 28.1% and 78.1% for Kimi."
  5. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Agreement with human examiners was limited in fast mode and improved in reasoning or thinking modes; however, reproducibility across sessions and implementation conditions remained limited."
  6. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "LLMs used fewer extreme grades than did human examiners."
  7. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "At the lower end of the grading scale, the differences were predominantly negative, indicating that ChatGPT tended to assign higher grades than human examiners. Conversely, at the upper end, the differences were predominantly positive, indicating that ChatGPT tended to assign lower grades."
  8. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The average numeric grade of human graders was 3.63, and the average of the five ChatGPT parallels was 3.73; the almost identical average grades represent a non-significant difference."
  9. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Second, the sample size of the exam submissions was limited (32), which may have affected generalizability."
  10. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Third, model parameters, such as temperature and randomness, were left to their defaults and not systematically controlled, which may contribute to variability in repeated grading sessions."
  11. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Fourth, grading guidelines were originally designed for human evaluators and may not have been optimally structured for LLM."
  12. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "This retrospective inter-rater reliability study compared three LLMs: ChatGPT-4o, Gemini 2.5 Flash, and Claude 3.5 Haiku with original human reference scores on 19 integrated anatomy, histology, and physiology objective structured practical examination items completed by 309 pre-medical students."
  13. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Rank-order correlations were strong across all models ( ρ = 0.784-0.921), but categorical agreement diverged substantially."
  14. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Pass/fail agreement ranged from 52.1% (ChatGPT; κ = 0.067, slight) to 84.1% (Claude; κ = 0.680, substantial)."
  15. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "The most extreme case was a pelvic three-dimensional model item on which 93.3% of students succeeded by human grading but all three LLMs assigned a mean score of zero."
  16. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "LLMs may serve as assessment assistants for first-pass ranking and discrepancy flagging, but item-level human review of flagged discordances, error-profile monitoring, and preserved human authority over pass/fail decisions are required before high-stakes deployment."
  17. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — solo el resumen - no se pudo obtener el texto completo — el pasaje: "This study compares the accuracy of four popular detection tools: GPTZero, Pangram, Copyleaks, and Turnitin on four kinds of academic papers: fully human-written, fully AI-written, hybrid (human with GenAI-inserted passages), and humanised GenAI (AI-generated passages were humanised using a prompt designed to resemble possible student behaviour)."
  18. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Results show that Pangram consistently performed better than the other tools, achieving high accuracy in detecting fully AI-generated, hybrid, and humanised texts."
  19. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — solo el resumen - no se pudo obtener el texto completo — el pasaje: "In contrast, the other tools significantly underestimated GenAI content, particularly for texts generated with the most advanced model."
  20. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — solo el resumen - no se pudo obtener el texto completo — el pasaje: "All tools correctly identified fully human texts."
  21. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — solo el resumen - no se pudo obtener el texto completo — el pasaje: "False positives were rare across all tools, suggesting improvement compared to earlier studies."
  22. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Findings show that while detection tools can provide useful initial flags, they should not be used as sole evidence in high-stakes decision-making but should be implemented in a broader evaluation strategy."
  23. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — solo el resumen - no se pudo obtener el texto completo — el pasaje: "Medical education faces increasing demand for scalable grading solutions. Large language models have been proposed as automated graders for high-stakes assessments, but evidence for their reliability in multimodal practical examinations remains limited."

Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452

Quién pagó: Los autores declararon que no recibieron apoyo financiero para este trabajo ni para su publicación, y declararon no tener relaciones comerciales o financieras que pudieran constituir un conflicto de interés.

experiment · Frontiers in Education · the paper, 1 Sep 2026 · free

An AI Can Read Your Students' Essays. It Grades Them Its Own Way.

In a small study at one Norwegian university, the best chatbot matched human examiners half the time, and gave different grades to the same paper on different days.

Short version · the longer version follows, about 7 min

Ask the weeklyAI chat

Ask me about this study: who was studied, what it found, and what it does not say.

Conversations are saved for as long as weeklyAI exists, to improve the publication. Answers come in the language you write in.

The study at a glance
Who
Four large language models (ChatGPT, Gemini, LeChat, Kimi) grading master's-level public health essays
How many
32 exam submissions
Where
OsloMet – Oslo Metropolitan University, Norway
When
January 2026
Kind of study
analysis of what people did
Who did it
OsloMet – Oslo Metropolitan University, Norway
The limit that matters
Only 32 exam submissions and two human graders, so findings may not travel far
How often ChatGPT and Kimi matched the human examiners' grades
ChatGPT, exact same grade50%
ChatGPT, within one grade90.6%
Kimi, exact same grade28.1%
Kimi, within one grade78.1%

These are the shares of 32 exam papers where each model's grade matched the human examiners exactly, or came within one step, in the single-file upload setting.

ChatGPT and Kimi against the human examiners

ChatGPTagainstHuman examiners

Exact match on 50.0% of papers and within one grade on 90.6%

KimiagainstHuman examiners

Exact match on 28.1% of papers and within one grade on 78.1%

The researchers say these tools may work as supervised grading assistants, not as the one who decides.
weeklyAI's reading

You grade long essay answers. You know the work of keeping one standard across a stack of papers, a term of days and a room of colleagues.

Researchers at OsloMet, a university in Oslo, Norway, put that work to four chatbots. They took 32 anonymized exam papers from a master's course in public health and asked ChatGPT, Gemini, LeChat and Kimi to grade them. The chatbots got the same written guidelines the human examiners used. The humans had already graded the papers and the students had their results.

In its thinking mode, ChatGPT gave the exact same grade as the human examiners on half the papers. Within one grade, it matched on about nine in ten. Kimi came second: about 28 percent exact, about 78 percent within one grade. In fast mode, the agreement with human examiners was not high.

Then the researchers ran the same papers through ChatGPT again, once a day for five days. The model agreed with itself only moderately across those days. The chatbots also handed out fewer top grades and fewer failing grades than the humans did. The chatbots leaned toward the middle of the scale.

The study is small: 32 papers, two human graders. It ran at one university, on one course, in one country. The guidelines were written for people, not machines, and the papers arrived as formatted PDFs. So half is a benchmark from a single study, not a promise about your course.

The researchers say these tools may work as supervised grading assistants, not as the one who decides. If you try one, look for a model that shows its reasoning, and check its grades against your own.

What this means for you

Keep this in mind if you sit down with one of these tools before your next stack of papers: it agreed with human examiners on half the papers and with itself only moderately over five days, so its judgment shifts even when the paper does not. Ask it to show its reasoning, compare its grades with yours, and treat the result as a second opinion you still have to weigh.

Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452

Who paid: The authors declared that no financial support was received for this work or its publication, and they declared no commercial or financial relationships that could be a conflict of interest.

The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1436 words · about 7 minRead it →Close

AI Graders Marked 32 Public Health Essays. Half the Time, They Agreed With the Humans.

A Norwegian master's course let four chatbots grade real exams. One matched the examiners exactly on half the papers — and quietly moved its own marks when it graded the same pile again days later.

How it could look · illustration generated by weeklyAI.watch, not a photograph

The study comes from OsloMet – Oslo Metropolitan University, in Norway, and it was not a trial of a new teaching method. It was a test of whether artificial intelligence can do a job that eats university teachers alive: reading long, open-text essay answers and deciding what they are worth. Two human examiners had already graded the papers and published the results to students. Only then did the researchers feed the same anonymized submissions to four large language models — ChatGPT, Gemini, LeChat and Kimi — with the same grading guidance the humans had used, asking each to hand out a final grade on the A–F scale1. Each model was also told to score every one of the three exam questions, give a short reason for each score, count the words in the answer, total everything up, and name the paper's main strength and weakness in a single sentence2.

The results were not a triumph. Agreement with the human examiners was weak when the models answered quickly and improved when they were switched into reasoning or "thinking" modes — but even then, running the same task again, or changing how many files were handed over at once, shifted the answers3. In the cleanest condition, one file at a time, ChatGPT agreed with the human examiners best, followed by Kimi4. On 32 exam submissions, ChatGPT gave exactly the same grade as the humans on 50.0 percent of papers and a grade within one step on 90.6 percent. Kimi managed 28.1 percent exact and 78.1 percent within one step5.

Here is the mechanism, in plain words. These models do not consult a rubric the way a person does. They were built by training on enormous amounts of text to predict what comes next, so when you paste in an exam paper and a set of marking guidelines, the model produces the grade that looks most plausible given everything it has seen before1. That is why the wording of the instructions and the order of the files matter so much, and why the same paper can come back with a different mark on a different day.

That instability is the part worth pausing on. When ChatGPT graded the same submissions on five separate days, the variation was moderate rather than trivial6. On average the human examiners gave 3.63 and the five ChatGPT runs averaged 3.73 — a difference the authors describe as not significant7. But the average hides the real behaviour. At the bottom of the scale ChatGPT tended to give higher marks than the humans; at the top it tended to give lower ones8. In other words, it pulled everything toward the middle. Across the board, the models handed out fewer extreme grades than the human examiners did9.

This is not the only study of its kind, and the picture from the others is mixed. A retrospective study of a practical medical exam compared three different models — ChatGPT-4o, Gemini 2.5 Flash and Claude 3.5 Haiku — against human reference scores on 19 anatomy, histology and physiology items answered by 309 pre-medical students; we could read only the summary, the full paper is behind a subscription10. The models ranked students in broadly the same order, but their pass/fail calls diverged sharply: agreement ran from 52.1 percent for one model to 84.1 percent for another11. On one item — a three-dimensional pelvic model — 93.3 percent of students passed under human grading while all three models gave a mean score of zero12. The authors of that study concluded that ranking well is not the same as being able to make a pass/fail decision in a high-stakes setting13.

How it could look · illustration generated by weeklyAI.watch, not a photograph

A second line of work asks a different question: not whether AI can grade, but whether anyone can tell what AI wrote. A study of four detection tools — GPTZero, Pangram, Copyleaks and Turnitin — tested them against fully human-written, fully AI-written, hybrid and "humanised" academic papers; we could read only the summary14. One tool performed consistently better than the others15, the rest significantly underestimated AI content, especially from the most advanced models16, and all of them correctly identified the fully human texts17. The authors' advice was that such tools can raise an initial flag but should never be the sole evidence in a high-stakes decision18.

Meanwhile, students are already using these tools whether or not their courses are ready. A study at a large public university in the United Arab Emirates found widespread student use of AI for summarizing, writing, proofreading, lab work and coding, alongside uneven understanding of when that use crosses a line; instructors were aware but course-level policies were inconsistent; we could read only the summary1920. That study argues that a single AI literacy course is not enough and calls for each course to say clearly which tasks permit AI and which do not, with strictly invigilated, AI-free assessments to build foundational skills first21.

The OsloMet authors are careful about what their numbers mean. Only 32 exam submissions were used, which they say may limit how far the findings travel22. Model settings such as temperature — the dial that controls how much randomness the model introduces — were left at their defaults and never systematically controlled, which may itself explain why repeated runs disagreed23. The grading guidelines were written for human examiners who knew the course, not for a machine24. And because only two human graders served as the reference, human grading was treated as a practical benchmark rather than an absolute truth. The authors also flag two risks they did not measure: that teachers may over-trust an AI mark, and that a model may reproduce biases buried in its training data.

Here is how we read it. When a tool arrives that can produce the decision people used to make, the people who formally hold the authority tend to lean on it — especially when they are tired — while the tool itself has no way of knowing which of its judgments are least trustworthy. We would expect that in a department under pressure, the AI grade quietly becomes the grade, and the cases where the model is shakiest are exactly the ones it will not flag. You would know we are wrong if you saw graders routinely overruling the model, writing down why, and treating its mark as one input among several even in a busy week. If you help grade, keep the final mark as your own decision.

How it could look · illustration generated by weeklyAI.watch, not a photograph

We also read that a model asked to judge long interpretive writing will fall back on what it handles well — length and fluency — and be most confident where it is weakest, at the very top and very bottom of the scale. In this study one model's final grade was strongly linked to how many words the answer contained25, and the human examiners showed a similar pull. You would know we are wrong if a model's grades tracked the substance of an argument rather than its length. Before trusting any AI mark, read the reasons it gives for the highest and lowest papers yourself.

We also read that this kind of agreement has not been shown to carry over everywhere, so what works in one exam format is not evidence that it works in yours. A second thing to watch for is that students graded by a machine may start writing for the machine instead of for a reader, and treat the mark as something a system issued rather than something to argue with a teacher. Keep at least one piece of assessment where a person reads the work and explains the mark in their own words.

For now, the authors' own conclusion is narrow: some models, ChatGPT and Kimi in particular, come moderately close to human examiners, but not reliably enough to grade alone. They suggest AI is currently better suited to being a supervised assistant than an autonomous grader. They also point to what would have to change: better-written rubrics and clearer prompts, possibly showing the model a few examples of top, middle and low answers so it can calibrate, and averaging several runs rather than trusting a single one2627.

Next time a grade comes home with a machine's fingerprints on it, ask the one question that matters: who read this paper, and can they tell you why it got the mark it did?

Where each piece of context comes from, and how much of it we read

  1. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "Four LLMs—ChatGPT, Gemini, LeChat, and Kimi—were prompted to assign final grades on an A-F scale using the same grading guidance as human examiners."
  2. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "The LLMs were also instructed to 4) summarize the total number of words and points from all three responses. Finally, the prompt instructed the LLMs to 5) provide a short overall justification for the final grade and 6) point out the major strengths and weaknesses of the submission in one sentence."
  3. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "Agreement with human examiners was limited in fast mode and improved in reasoning or thinking modes; however, reproducibility across sessions and implementation conditions remained limited."
  4. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "In the single-file upload setting, ChatGPT showed the strongest agreement with human examiners (weighted kappa 0.718, 95% CI 0.578–0.859), followed by Kimi (weighted kappa 0.571, 95% CI 0.373–0.768)."
  5. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "ChatGPT achieved 50.0% exact agreement and 90.6% agreement within ±1 grade, with corresponding values of 28.1% and 78.1% for Kimi."
  6. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "Repeated grading by ChatGPT across 5 days showed moderate variation."
  7. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "The average numeric grade of human graders was 3.63, and the average of the five ChatGPT parallels was 3.73; the almost identical average grades represent a non-significant difference."
  8. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "At the lower end of the grading scale, the differences were predominantly negative, indicating that ChatGPT tended to assign higher grades than human examiners. Conversely, at the upper end, the differences were predominantly positive, indicating that ChatGPT tended to assign lower grades."
  9. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "LLMs used fewer extreme grades than did human examiners."
  10. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — only the abstract - the full text could not be fetched — the passage: "This retrospective inter-rater reliability study compared three LLMs: ChatGPT-4o, Gemini 2.5 Flash, and Claude 3.5 Haiku with original human reference scores on 19 integrated anatomy, histology, and physiology objective structured practical examination items completed by 309 pre-medical students."
  11. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — only the abstract - the full text could not be fetched — the passage: "Pass/fail agreement ranged from 52.1% (ChatGPT; κ = 0.067, slight) to 84.1% (Claude; κ = 0.680, substantial)."
  12. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — only the abstract - the full text could not be fetched — the passage: "The most extreme case was a pelvic three-dimensional model item on which 93.3% of students succeeded by human grading but all three LLMs assigned a mean score of zero."
  13. Akhund SA, Qazi S, Mazhar MA, Shaikh AA, Atif E, Shaikh A, et al. (2026). Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis. Medical Education Online. https://doi.org/10.1080/10872981.2026.2684837 — only the abstract - the full text could not be fetched — the passage: "Strong rank-order correlation alone does not support LLM grading for categorical decisions in high-stakes assessment contexts."
  14. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — only the abstract - the full text could not be fetched — the passage: "This study compares the accuracy of four popular detection tools: GPTZero, Pangram, Copyleaks, and Turnitin on four kinds of academic papers: fully human-written, fully AI-written, hybrid (human with GenAI-inserted passages), and humanised GenAI (AI-generated passages were humanised using a prompt designed to resemble possible student behaviour)."
  15. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — only the abstract - the full text could not be fetched — the passage: "Results show that Pangram consistently performed better than the other tools, achieving high accuracy in detecting fully AI-generated, hybrid, and humanised texts."
  16. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — only the abstract - the full text could not be fetched — the passage: "In contrast, the other tools significantly underestimated GenAI content, particularly for texts generated with the most advanced model."
  17. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — only the abstract - the full text could not be fetched — the passage: "All tools correctly identified fully human texts."
  18. Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00226-w — only the abstract - the full text could not be fetched — the passage: "Findings show that while detection tools can provide useful initial flags, they should not be used as sole evidence in high-stakes decision-making but should be implemented in a broader evaluation strategy."
  19. Nazir A. (2026). Unauthorized intelligence? Investigating AI use among undergraduate students toward curriculum redesign. Social Sciences & Humanities Open. https://doi.org/10.1016/j.ssaho.2026.103033 — only the abstract - the full text could not be fetched — the passage: "Results indicate widespread student use of AI for summarizing, writing, proofreading, laboratory tasks, and coding, alongside uneven understanding of when such use becomes unauthorized."
  20. Nazir A. (2026). Unauthorized intelligence? Investigating AI use among undergraduate students toward curriculum redesign. Social Sciences & Humanities Open. https://doi.org/10.1016/j.ssaho.2026.103033 — only the abstract - the full text could not be fetched — the passage: "Instructors reported high awareness of student AI use but inconsistent course-level policies, resulting in fragmented expectations."
  21. Nazir A. (2026). Unauthorized intelligence? Investigating AI use among undergraduate students toward curriculum redesign. Social Sciences & Humanities Open. https://doi.org/10.1016/j.ssaho.2026.103033 — only the abstract - the full text could not be fetched — the passage: "The study argues that a stand-alone AI literacy course is insufficient. Instead, it calls for systematic, curriculum-wide reform in which each course clearly defines AI-permitted and AI-restricted tasks and incorporates strictly invigilated, AI-free assessments to ensure that foundational skills are developed before AI assistance is allowed."
  22. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "Second, the sample size of the exam submissions was limited (32), which may have affected generalizability."
  23. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "Third, model parameters, such as temperature and randomness, were left to their defaults and not systematically controlled, which may contribute to variability in repeated grading sessions."
  24. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "Fourth, grading guidelines were originally designed for human evaluators and may not have been optimally structured for LLM."
  25. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "ChatGPT, the LLM most aligned with human assessors, showed a moderate correlation between the total number of words in the exam submissions and the suggested final grade (Table 4)."
  26. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "The additional time invested in optimizing grading guidelines and unambiguous prompt design could be the difference between close alignment with human graders or unacceptable randomness in the LLM grades."
  27. Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452 - the article this story is about — the whole article — the passage: "Perhaps the LLM could be asked to develop its own instructions or rubrics if it were presented with a few examples of top-, middle-, and low-grade submissions. The course organizer can create artificial examples for LLM calibration."

Brevik, A., Jerpseth, H., Lafontan, S. R. (2026). Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment. Frontiers in Education. https://doi.org/10.3389/feduc.2026.1904452

Who paid: The authors declared that no financial support was received for this work or its publication, and they declared no commercial or financial relationships that could be a conflict of interest.