experiment · Scientific reports · la publicación, 20 abr 2026 · gratis
Un estudio midió cuatro modelos de IA para detectar videos falsos, pero acertó poco.
Versión breve · la versión detallada sigue, unos 6 min
Pregunte a weeklyAI
Pregúnteme por este estudio: a quiénes se estudió, qué encontró y qué no dice.
Las conversaciones se guardan mientras exista weeklyAI, para mejorar la publicación. Se responde en el idioma en que usted escribe.
- El estudio, de un vistazo
- Quiénes
- Cuatro modelos de inteligencia artificial para detectar videos falsos
- Cuántos
- 3,502 videos reales y falsos
- Dónde
- No lo dice el estudio
- Cuándo
- Publicado el 20 de abril de 2026
- Tipo de estudio
- experimento
- Quién lo hizo
- Universidades de India y Malasia
- El límite que importa
- Se probó con videos de investigación, no con material electoral real ni con su video.
El mejor modelo frente a los otros tres, con las tres colecciones mezcladas
La red tridimensional acertó más
La red tridimensional acertó más
La red tridimensional acertó más
Cuando un programa acierta sesenta y cinco de cada cien, usted no tiene un detector: tiene una moneda cargada, y la moneda no sabe nada de su país.

Un estudio publicado el 20 de abril de 2026 comparó cuatro modelos de inteligencia artificial (3DCNN, 3DResNet, TCN y VAE) para detectar videos falsos. Usó 3,502 videos reales y falsos de tres bases de investigación (FF++, DFDC y CDF), no de campañas ni elecciones. El mejor modelo acertó 64.68% en datos mezclados. El VAE acertó 51.71%. El estudio no probó los modelos en material electoral real ni mide daño alguno.
El documento no nombra ninguna forma para que usted verifique esto.
El estudio es un ejercicio de comparación entre modelos, no una regla ni una auditoría electoral. No existe en el documento ninguna autoridad electoral que se haya pronunciado, ni resolución, ni dictamen sobre estos modelos. El propio estudio señala que la generalización falla: un detector así no prueba que un video sea genuino ni falso.
Qué significa para usted
Antes de votar, si alguien le muestra un video para convencerlo, sepa que ningún detector de este estudio podría confirmarle si es genuino o falso: el mejor acertó 64.68% y el estudio admite que falla al enfrentar material distinto al de su entrenamiento. No hay todavía autoridad electoral que haya revisado estas herramientas, así que pida las resoluciones, auditorías y actas de su autoridad y compárelas usted mismo.
Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1
Quién pagó: La financiación de acceso abierto fue proporcionada por Symbiosis International (Deemed University); los autores no recibieron financiación específica para este estudio y el artículo declara no tener intereses en competencia.
Versión detalladaLos pasajes copiados del artículo, las ilustraciones y cada fuente con cuánto leímos de ella · 1126 palabras · unos 6 minLeerla →Cerrar
La regla y la evidencia: un estudio midió cuatro modelos de IA para detectar videos falsos, pero acertó poco
Un estudio comparó cuatro programas que buscan caras falsas en videos. El mejor no llega a tres aciertos de cada cinco cuando los videos vienen de fuentes distintas.

### El acto
No hay un acto que citar esta semana. Lo que tenemos es un estudio publicado el 20 de abril de 2026, que no dicta reglas ni obliga a nadie. Sus autores entrenaron y evaluaron cuatro programas de computadora para distinguir videos reales de videos falsos, y publicaron cuántas veces cada uno acertó1.
El estudio no impone deberes a ninguna autoridad electoral, no fija plazos, no dice quién debe vigilarlo y no tiene fuerza sobre nadie. Es un informe de laboratorio con fecha de publicación, y su valor para usted depende de que sus números se lean tal como están escritos.
### El caso que responde
Esta semana no se encontró ningún caso real que este estudio responda, y no se inventa ninguno.
### Lo que se midió
Los cuatro programas se probaron con 3,502 videos reales y falsos tomados de tres colecciones ya existentes, hechas para investigar: FaceForensics++, el reto de detección de deepfakes y Celeb-DF. La mitad eran reales y la mitad falsos, y cada video se redujo a 16 cuadros de 128 por 128 píxeles2.
El mejor de los cuatro, la red tridimensional, acertó casi dos de cada tres videos cuando se probó con las tres colecciones mezcladas3. Los otros tres quedaron por debajo: la red residual en 57.02%, la red temporal en 60.43% y el autoencoder variacional en 51.71%.
El autoencoder variacional merece un párrafo aparte, porque su número no es un empate técnico con el azar, sino algo peor. De cada cien videos falsos que se le pusieron delante, detectó menos de cuatro: su recall fue de 3.88%4. Cuando dijo "falso", acertó poco más de un tercio de las veces4.
El propio estudio reconoce su límite central: los modelos pierden capacidad cuando se los saca de las condiciones donde se entrenaron. Los autores escriben que hay "una degradación significativa en la generalización" y que cada modelo falla de manera distinta cuando los datos vienen mezclados de fuentes heterogéneas5. La revisión recoge un estudio en el que la mayoría de los sistemas de detección andan bien en el entrenamiento y mal cuando se los prueba con trucos que no vieron antes6.
Nada de esto se midió en una elección. Se midió en videos de investigación, no en material de campaña, no en un conteo, no en un registro de votantes. El estudio no dice una palabra sobre ningún país, y no puede decirse que sus resultados valgan para el suyo12.
### La respuesta de la otra región al mismo asunto
No encontramos esta semana ningún tribunal, ley o autoridad de otro país que se haya pronunciado sobre este método en particular.
### La lectura de la casa
Aquí es donde hablamos nosotros. Lo que acabamos de leer no es una herramienta que le devuelva la vista. Es una herramienta que le pide confiar en una etiqueta, y esa etiqueta se equivoca mucho más de lo que la palabra "inteligencia" deja imaginar. Cuando un programa acierta sesenta y cinco de cada cien, usted no tiene un detector: tiene una moneda cargada, y la moneda no sabe nada de su país.
Lo que nos preocupa no es el número. Es lo que suele pasarle al número cuando entra en circulación. Un resultado de laboratorio, con sus tasas de error impresas al lado, se convierte en la calle en una frase corta: "un estudio dice que es falso". Esa frase viaja sin el 35% de error, sin el 3.88% de recall, sin la advertencia de que los videos eran de investigación. Y en una campaña, una frase así pesa más que el video mismo.
Cabe esperar que en los próximos meses alguien muestre un video, diga que un sistema lo analizó y pida que usted lo crea o lo descarte por eso. Si eso pasa, la expectativa que tenemos es simple: el número se usará como veredicto y no como medición. Sabremos que nos equivocamos si cada vez que alguien invoca un detector automático muestra también cuántas veces ese detector se equivoca, y si su autoridad electoral exige que un ojo humano revise antes de tratar una etiqueta como prueba.
La revisión recoge un estudio en el que se compara su red tridimensional con otros métodos en cada colección por separado, y en cada una de esas comparaciones su modelo queda por encima. Lo decimos como algo que conviene vigilar: cuando dos métodos dicen números tan distintos sobre el mismo material, la diferencia no la resuelve el entusiasmo, la resuelve la repetición por manos independientes. Y aquí hay un problema práctico: los videos que se usaron no están disponibles libremente por restricciones de licencia, así que repetir el experimento no está al alcance de cualquiera.
Lo que usted puede hacer con esto no requiere saber de computadoras. Cuando alguien le muestre un video sospechoso y le diga que una máquina ya lo revisó, pregunte quién lo revisó, con qué método, sobre qué videos se probó ese método y cuántas veces se equivoca. Si no le pueden contestar esas cuatro cosas, no tiene una prueba: tiene una opinión con vocabulario técnico. Y esa pregunta, hecha en voz alta en una conversación familiar o en la puerta de un local, vale más que cualquier detector.
### Lo que no se sabe
No se sabe si estos programas funcionarían con el video concreto que usted vea en su teléfono. Se probaron con videos de investigación, de 128 por 128 píxeles, y el propio estudio advierte que esa resolución baja puede borrar las señales finas de la cara que sirven para detectar un montaje25.
No se sabe cómo se comportarían ante un video dirigido a un electorado específico, en un idioma específico, con la compresión de una red social específica. Nadie midió eso5.
No se sabe si el resultado se sostiene cuando el material falso lo produce un método que no existía cuando se armaron las colecciones. La revisión recoge un estudio en el que la generalización es justamente el problema sin resolver7.
Y no se sabe qué haría su autoridad electoral con una etiqueta automática, porque el estudio no menciona ninguna autoridad electoral, ningún país y ninguna elección1.
Lo que sí sabemos es lo que usted ya sabía antes de leer esto: que un video puede ser falso, y que alguien va a querer que usted actúe como si no lo fuera. La próxima vez que vea uno, antes de reenviarlo, hágase una sola pregunta: ¿quién me está pidiendo que lo crea, y qué me está pidiendo hacer después?
De dónde sale cada dato de contexto, y cuánto leímos de cada documento
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - el artículo del que trata esta nota — el artículo completo — el pasaje: "This research paper investigates how four deep learning models including 3D Convolutional Neural Network (3DCNN), 3D Residual Network (3DResNet), Temporal Convolutional Network (TCN), and Variational Autoencoder (VAE) perform in detecting and classifying deepfake videos."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The models received training and evaluation using FaceForensics++ (FF++) and Deepfake Detection Challenge (DFDC) and Celeb-DF (CDF) datasets which contained more than 3500 real and fake video samples."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The experimental findings indicate that 3DCNN can reach the highest test accuracy of 64.68% which is higher than the results of 3DResNet, TCN, and VAE under the cross-dataset conditions."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The model fails to identify fake videos effectively because its recall rate reaches only 3.88%."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - el artículo del que trata esta nota — el artículo completo — el pasaje: "It analyses indicate that there is significant degradation in generalization and different failure behavior with models when presented with heterogeneous distributions of data."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - el artículo del que trata esta nota — el artículo completo — el pasaje: "Most deepfake detection systems show strong training results but perform badly when tested on untested manipulation techniques."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - el artículo del que trata esta nota — el artículo completo — el pasaje: "The generalization ability of deepfake detection models remains a significant challenge because they struggle to detect deepfakes across various datasets and deepfake generation techniques."
Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1
Quién pagó: La financiación de acceso abierto fue proporcionada por Symbiosis International (Deemed University); los autores no recibieron financiación específica para este estudio y el artículo declara no tener intereses en competencia.
Los hallazgos de otros estudios que aquí se mencionan los conocemos por este documento, que fue el que leímos; no abrimos cada uno de esos estudios.
experiment · Scientific reports · the paper, 20 Apr 2026 · free
A study of four AI deepfake detectors found the best one was right on 64.68 percent of mixed videos.
Short version · the longer version follows, about 6 min
Ask weeklyAI
Ask me about this study: who was studied, what it found, and what it does not say.
Conversations are saved for as long as weeklyAI exists, to improve the publication. Answers come in the language you write in.
- The study at a glance
- Who
- four deep learning models (3DCNN, 3DResNet, TCN, VAE) tested on deepfake videos
- How many
- 3,502 real and fake videos
- Where
- no country named; research benchmark datasets only
- When
- published 20 April 2026
- Kind of study
- experiment
- Who did it
- Rungta International Skills University, INTI International University, SRM University, Symbiosis International
- The limit that matters
- Tested only on research benchmark videos, never on a real campaign or election feed
The four detectors against each other on the same mixed pile of videos
3DCNN reached the highest test accuracy of 64.68%, higher than the other three under cross-dataset conditions
VAE performed worst, catching only 3.88% of the fakes
On the combined data, 3DCNN reached the highest test accuracy at 64.68 percent and the best F1-score at 0.6420.

Prateek Agrawal, Dharmendra Pathak, Vishu Madaan and colleagues trained and evaluated four deep learning models — 3DCNN, 3DResNet, TCN and a Variational Autoencoder — for deepfake video detection. The study, published on 20 April 2026, used 3,502 real and fake videos drawn from three research benchmark datasets (FaceForensics++, DFDC and Celeb-DF), split 80–20 per dataset, with models trained for 30 epochs. On the combined data, 3DCNN reached the highest test accuracy at 64.68 percent and the best F1-score at 0.6420. 3DResNet reached 57.02 percent, TCN 60.43 percent and the VAE 51.71 percent, with 3.88 percent recall, which the authors call unsuitable for classification.
The document names no way for a voter to check a specific video.
The document states no electoral authority rule, ruling or audit. It reports only that generalization degraded across heterogeneous data and lists limits including low-resolution 128x128 frames.
What this means for you
If you are checking a video before you vote, this study offers no way to verify that particular clip, and no electoral authority rule, ruling or audit appears in it. What you can watch for is the gap between a detector's score on one dataset and its much lower score on mixed videos, so treat any claim that a machine reliably sorts real from fake as unsettled for now.
Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1
Who paid: Open access funding was provided by Symbiosis International (Deemed University); the authors received no specific funding for this study, and the article declares no competing interests.
The longer versionThe passages copied from the paper, the pictures, and every source with how much of it we read · 1186 words · about 6 minRead it →Close
The rule and the evidence: A study of four AI deepfake detectors found the best one was right on 64.68 percent of mixed videos
On a mixed pile of 3,502 research videos, the strongest of four detectors still missed a third of the fakes — and the study that tested them never touched a real campaign.

### THE ACT
This week's document is not a law, a ruling or an audit. It is a study, published on 20 April 2026, and what it puts in force is a measurement. The authors state their own scope plainly: "This work does not present a new architecture of detection but rather poses itself as a systematic benchmarking and generalization study."
What was actually done: four deep learning models were trained and evaluated on deepfake video. "This research paper investigates how four deep learning models including 3D Convolutional Neural Network (3DCNN), 3D Residual Network (3DResNet), Temporal Convolutional Network (TCN), and Variational Autoencoder (VAE) perform in detecting and classifying deepfake videos"1.
The material: "The models received training and evaluation using FaceForensics++ (FF++) and Deepfake Detection Challenge (DFDC) and Celeb-DF (CDF) datasets which contained more than 3500 real and fake video samples"2. The paper gives the count as 3,502 videos, split evenly between real and fake, and split across three research benchmarks.
Who is bound by it: no one. There is no authority enforcing anything here, no deadline for a person to meet, no filing a reader can make. What a reader can do is read the result and decide how much weight to give any tool that claims to tell them a video is fake.
### THE CASE IT ANSWERS
No real instance was found this week.
### WHAT WAS MEASURED
The effect is accuracy, and it is small. "The experimental findings indicate that 3DCNN can reach the highest test accuracy of 64.68% which is higher than the results of 3DResNet, TCN, and VAE under the cross-dataset conditions"3. In plain terms: on the combined pile, the best of the four was right roughly 65 times in 100 — about two in three, and wrong the rest of the time.
The other three did worse. The paper reports the training accuracy of 70.74% for the 3DCNN against 62.98% for 3DResNet and 65.43% for TCN, and a test accuracy of 51.71% for VAE4. On the mixed test, 3DResNet reached 57.02% and TCN 60.43%.
The VAE barely worked as a classifier at all. "The model fails to identify fake videos effectively because its recall rate reaches only 3.88%"5. Recall is the share of actual fakes it caught. Three or four in a hundred.
The sample was 3,502 videos, split evenly between real and fake, and split equally across three research benchmarks, with about 1,167 videos per source; an 80–20 split gave 934 videos for training and 233 for testing per dataset, frames at 128×128 pixels, 16 frames per video, 30 training rounds.
No country is named. No election is named. The videos are research benchmarks, not campaign material.
The measured thing is accuracy on a fixed pile of research videos. The paper says nothing about whether any of this would hold on a live election feed. It was never tested that way.
### THE OTHER REGION'S ANSWER TO THE SAME THING
We found none this week: no court, law or authority in another country, in the passages we read, ruled on or regulated this detection method.
### THE LIBRARY'S READING
The library gave no reading this week.
Here is how we read it. Two things are being confused in public talk about fake video, and the study shows the gap. One is whether a machine can be built that sorts real from fake. The other is whether you, looking at a clip on your phone the night before you vote, can get a straight answer. This study is about the first. It says the best of four machines, on a mixed pile, is right about two times in three.
So expect this in homes like yours: a clip will arrive, someone will say a tool checked it, and the tool's answer will be treated as a verdict. It is not a verdict. A detector at this level of accuracy would flag many real videos as fake and wave through many fakes as real, and the paper reports that these models get notably worse when the videos come from sources they did not train on: "there is significant degradation in generalization and different failure behavior with models when presented with heterogeneous distributions of data"6, and "The article describes a study in which most deepfake detection systems show strong training results but perform badly when tested on untested manipulation techniques"7.
We also read, in other work on this kind of system, that a different hybrid model reports very high accuracy on two of these same benchmarks — far above what this paper found — which tells us the numbers swing widely depending on how a tool is built and what it is tested on, and that a single accuracy figure is not a settled fact about deepfake detection. And we read that a finding like this one has not been shown to carry over to a country, a language or a broadcast style it was never tested on: that is a caution about carrying the result to your country, not a doubt about what the study found.
How you could tell we are wrong: if a detector kept the same accuracy on faces, languages and video styles it never trained on, with no drop when the source changed, then the worry would be misplaced.
What you can do with it: when an authority, a newsroom or a friend says a video is fake, ask what the tool was tested on, and whether anything from your country's broadcasts was in that test. If the answer is nothing, the label is not yet evidence about your country. Then look for the original post — who put it up first, and when.
### WHAT IS NOT KNOWN
The paper does not say how these models would behave on a real campaign video, because they were never tested on one.
The models were fed low-resolution frames — 128×128 pixels — and the authors name this as a limit, warning that small facial details useful for spotting a fake may simply be lost. They also note the models use no attention mechanisms and no explicit adaptation to new video sources, and that the training and testing data are not publicly available because of the benchmark licences, which limits anyone else from checking the work.
The VAE was judged in a role it was not built for — sorting real from fake — rather than the anomaly-detection role the authors say suits it.
And no one has measured whether a detector like this changes what a voter believes, or how they vote. That was not the question asked.
So: before you trust a machine's word on what you see, ask what it was trained on — and if the answer is not your country, treat it as a question, not a verdict.
Where each piece of context comes from, and how much of it we read
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - the article this story is about — the whole article — the passage: "This research paper investigates how four deep learning models including 3D Convolutional Neural Network (3DCNN), 3D Residual Network (3DResNet), Temporal Convolutional Network (TCN), and Variational Autoencoder (VAE) perform in detecting and classifying deepfake videos."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - the article this story is about — the whole article — the passage: "The models received training and evaluation using FaceForensics++ (FF++) and Deepfake Detection Challenge (DFDC) and Celeb-DF (CDF) datasets which contained more than 3500 real and fake video samples."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - the article this story is about — the whole article — the passage: "The experimental findings indicate that 3DCNN can reach the highest test accuracy of 64.68% which is higher than the results of 3DResNet, TCN, and VAE under the cross-dataset conditions."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - the article this story is about — the whole article — the passage: "The experimental findings show that 3DCNN produced the highest accuracy rate of 70.74% when compared to 3DResNet (62.98%), TCN (65.43%), and VAE (51.71%)."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - the article this story is about — the whole article — the passage: "The model fails to identify fake videos effectively because its recall rate reaches only 3.88%."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - the article this story is about — the whole article — the passage: "It analyses indicate that there is significant degradation in generalization and different failure behavior with models when presented with heterogeneous distributions of data."
- Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1 - the article this story is about — the whole article — the passage: "Most deepfake detection systems show strong training results but perform badly when tested on untested manipulation techniques."
Agrawal, P., Pathak, D., Madaan, V. et al. (2026). Spatiotemporal deep learning for real-time video-based deepfake detection using 3DCNN, 3DResNet, TCN, and VAE. Scientific Reports. https://doi.org/10.1038/s41598-026-49090-1
Who paid: Open access funding was provided by Symbiosis International (Deemed University); the authors received no specific funding for this study, and the article declares no competing interests.
The findings of other studies mentioned here are known to us through this document, which is the one we read; we did not open each of those studies.