When AI enters the physical world, safety gets real
en-GBde-DEes-ESfr-FR

When AI enters the physical world, safety gets real

12.08.2026 TranSpread

Vision-language models (VLMs) connect images with text, while vision-language-action models (VLAs) extend that connection to robot plans and control signals. This makes natural-language instruction, scene understanding and flexible task execution possible, but it also creates a chain of dependency: flawed data can distort perception, weak visual-language alignment can produce hallucinations, and malicious inputs can redirect decisions. In a chatbot, such errors may generate misinformation; in an autonomous vehicle or industrial robot, they may lead to collisions, damaged equipment or failed missions. Existing safeguards are often benchmark-specific, fragmented across system layers or too computationally costly for real-time use. Because of these challenges, deeper research is needed into unified, adaptive safeguards for multimodal agents operating under uncertain physical conditions.

Published (DOI: 10.1007/s11633-025-1626-x) online on July 13, 2026, in Machine Intelligence Research, the review was conducted by researchers from the Institute of Automation, Chinese Academy of Sciences; University College London (UCL); Minzu University of China; and the China Academy of Electronics and Information Technology. The team examined how VLMs and VLAs are used in embodied intelligence (EI), organized the major security threats and defensive approaches, and connected technical safety with accountability, fairness, privacy, environmental sustainability and human oversight. The article appears in the journal’s special issue on the security and ethics of generative AI.

The survey first tracks VLM and VLA use across four linked functions: perception, planning, instruction following and human-robot interaction (HRI). It then shows how failures can cascade. Biased training data, weak visual encoders or poor cross-modal alignment can make a model describe objects that are not present. Forged traffic signs, altered labels, cloned voices or deceptive captions can misguide perception and planning. Tiny adversarial perturbations, hidden backdoor triggers and multimodal jailbreak prompts may bypass safety controls, while persistent sensing can expose identity, location, possessions and social behavior. The authors organize countermeasures into equally connected layers. These include hallucination filtering and vision-grounded alignment; cross-modal forgery detection, watermarking and provenance tracing; defenses against perturbations, backdoors and jailbreaks; differential privacy (DP), secure multi-party computation (SMPC) and homomorphic encryption (HE); and safeguards for navigation, communications and physical control. A further strand uses causal explanations, intent alignment and risk assessment so robots can interpret ambiguous instructions, anticipate hazards and correct actions. The review's central insight is that no single filter can secure an embodied agent: protection must follow the entire path from sensor input to model reasoning, system architecture and physical execution.

The authors said the central challenge is not simply making models more accurate, but ensuring that a system remains safe when its sensors, language inputs and operating conditions are imperfect. They said defenses should be combined rather than deployed as isolated patches, with transparent risk metrics, continuous monitoring and human oversight for critical decisions. A trustworthy robot must also explain what it is doing, recognize when it is uncertain and fall back safely instead of acting with false confidence. The authors added that technical progress must move alongside privacy protection, fairness, accountability and responsible governance.

For developers and regulators, the survey provides a practical checklist for evaluating embodied systems before large-scale deployment. Future platforms could combine interpretable reasoning, attack detection, privacy-preserving computation and dynamic safety controls under reproducible, open evaluation protocols. The authors call for designs that address four dimensions together: technical robustness, regulatory alignment, social equity and environmental sustainability. Such an approach could support safer autonomous transport, healthcare assistance, warehouse automation, industrial inspection and collaborative robotics, while making responsibility easier to trace when failures occur. The review also warns that strong laboratory results may not transfer cleanly to noisy, culturally diverse and resource-constrained environments. Progress will therefore depend on cross-disciplinary cooperation and testing that measures not only task success, but safe behavior under stress.

###

References

DOI

10.1007/s11633-025-1626-x

Original Source URL

https://doi.org/10.1007/s11633-025-1626-x

Funding information

This work was partially supported by the National Natural Science Foundation of China (Nos. 62506362, 62302539 and U21B2045), the Strategic Priority Research Program of Chinese Academy of Sciences, China (No. XDA0480302), and Engineering and Physical Sciences Research Council (EPSRC) Funded Grant, UK (No. EP/Y028805/1).

About Machine Intelligence Research

Machine Intelligence Research (original title: International Journal of Automation and Computing) is published by Springer and sponsored by the Institute of Automation, Chinese Academy of Sciences. The journal publishes high-quality papers on original theoretical and experimental research, targets special issues on emerging topics, and strives to bridge the gap between theoretical research and practical applications.

Paper title: Embodied Intelligence Security with Vision-language Models: A Survey
Angehängte Dokumente
  • Overview of VLM integration and security strategies in EI.
12.08.2026 TranSpread
Regions: North America, United States
Keywords: Applied science, Artificial Intelligence

Disclaimer: AlphaGalileo is not responsible for the accuracy of content posted to AlphaGalileo by contributing institutions or for the use of any information through the AlphaGalileo system.

Referenzen

We have used AlphaGalileo since its foundation but frankly we need it more than ever now to ensure our research news is heard across Europe, Asia and North America. As one of the UK’s leading research universities we want to continue to work with other outstanding researchers in Europe. AlphaGalileo helps us to continue to bring our research story to them and the rest of the world.
Peter Dunn, Director of Press and Media Relations at the University of Warwick
AlphaGalileo has helped us more than double our reach at SciDev.Net. The service has enabled our journalists around the world to reach the mainstream media with articles about the impact of science on people in low- and middle-income countries, leading to big increases in the number of SciDev.Net articles that have been republished.
Ben Deighton, SciDevNet
AlphaGalileo is a great source of global research news. I use it regularly.
Robert Lee Hotz, LA Times

Wir arbeiten eng zusammen mit...


  • The Research Council of Norway
  • SciDevNet
  • Swiss National Science Foundation
  • iesResearch
Copyright 2026 by DNN Corp Terms Of Use Privacy Statement