Surrey and NVIDIA training fix could let AI-generated scenes respond properly to your controls
en-GBde-DEes-ESfr-FR

Surrey and NVIDIA training fix could let AI-generated scenes respond properly to your controls


AI systems that generate video frame by frame could follow a user’s camera commands far more accurately, thanks to a new training method developed by researchers from the University of Surrey and NVIDIA.

This improvement matters most where someone steers a generated scene rather than simply watching it. That covers video games built on worlds the AI creates, virtual production sets a director can walk a camera around and simulated environments used to train robots.

Getting an AI-generated world to turn left when it is told has consistently been difficult for the technology partly because of a flaw in how such models are typically trained. Fast video models that produce content one frame at a time (called students) are trained by a second, much slower model (teachers) that marks the work after the fact. This marker is typically what is allowed to look at the whole finished clip at once, so it can judge an early frame using knowledge of frames and camera moves that had not yet happened when the faster model produced it.

The research team call this the teacher–student context mismatch. In reality, the student model is being graded against a standard it can never meet once deployed.

Hmrishav Bandyopadhyay, lead author of the study from SketchX in the Surrey Institute for People-Centred AI at the University of Surrey, said:

"If you grade an AI student model’s work using a marker who can already see what happens next, you are teaching it to lean on information it will never have in the real world. What we have developed is simple – the teacher only ever sees what the student saw when it made each decision. That alignment turns out to matter far more than we expected, and it matters most when someone is actively steering the camera.”

The team’s method, Context-Matched Distillation, rebuilds the teacher so it can only look backwards, then grades each generated frame against the actual history the student produced during its own trial run rather than a reconstruction. Because early attempts tend to wander off course, the method also adds a controlled amount of noise to that history, so the teacher is not distracted by rough patches in the student’s early work.

Tested against seven existing pipelines on standard video generation benchmarks, the approach produced the highest overall quality scores and a substantial improvement in camera accuracy, specifically recording the lowest camera position errors of any method tested, on both the easy and hard test sets.

Generating roughly 30 seconds of video, a much harder task because small errors accumulate into visible drift, the method scored highest on overall quality while producing more movement than rival systems, several of which achieved stability largely by generating less motion in the first place.

In a blind comparison where an AI judge was shown pairs of videos without being told which system made them, the researchers’ models were preferred in between 60 and 88 per cent of matchups against each of six competing systems.

Professor Yi-Zhe Song, Director of SketchX Lab, Professor of Computer Vision and AI at the University of Surrey and Co-Director of the Surrey Institute for People-Centred AI, said:

"Very soon we will be moving from generating clips to generating entire large-scale places and “worlds” – somewhere you can enter, move through and change, built as you go rather than made in advance. That shift reaches well past entertainment; it changes how we prototype a building before anyone breaks ground, how we teach machines to operate in spaces too dangerous or too rare to practise in, and how much of what we watch is generated on demand rather than filmed.”

The method was built on NVIDIA’s Cosmos-Predict2.5-2B video model and works for both single-frame and multi-frame generation. The researchers note it also avoids an expensive preparation stage that competing approaches require, and that extending it to longer videos does not force the teacher to process more footage at once.

The study has been posted as a preprint.

[ENDS]

Bandyopadhyay, H., Ren, X., Huang, Z., Wu, J. Z., Cao, T., Li, R., Chu, B., Fidler, S., Song, Y.-Z., & Wang, Z. (2026). Context-matched distillation: Teacher causality for autoregressive video distillation. arXiv preprint arXiv:2608.13391.
Archivos adjuntos
  • Credit: University of Surrey
  • Credit: University of Surrey
Regions: Europe, United Kingdom
Keywords: Applied science, Artificial Intelligence, Computing, Technology

Disclaimer: AlphaGalileo is not responsible for the accuracy of content posted to AlphaGalileo by contributing institutions or for the use of any information through the AlphaGalileo system.

Testimonios

We have used AlphaGalileo since its foundation but frankly we need it more than ever now to ensure our research news is heard across Europe, Asia and North America. As one of the UK’s leading research universities we want to continue to work with other outstanding researchers in Europe. AlphaGalileo helps us to continue to bring our research story to them and the rest of the world.
Peter Dunn, Director of Press and Media Relations at the University of Warwick
AlphaGalileo has helped us more than double our reach at SciDev.Net. The service has enabled our journalists around the world to reach the mainstream media with articles about the impact of science on people in low- and middle-income countries, leading to big increases in the number of SciDev.Net articles that have been republished.
Ben Deighton, SciDevNet
AlphaGalileo is a great source of global research news. I use it regularly.
Robert Lee Hotz, LA Times

Trabajamos en estrecha colaboración con...


  • The Research Council of Norway
  • SciDevNet
  • Swiss National Science Foundation
  • iesResearch
Copyright 2026 by DNN Corp Terms Of Use Privacy Statement