Add BrightSurf on Google Email

When AI enters the physical world, safety gets real

08.12.26 | Maximum Academic Press
Fluke 87V Industrial Digital Multimeter

Fluke 87V Industrial Digital Multimeter is a trusted meter for precise measurements during instrument integration, repairs, and field diagnostics.


Artificial intelligence (AI) is moving beyond screens into cars, drones, service robots and collaborative machines that can perceive, reason and act. A new survey maps the security and ethical risks created when vision-language models guide these embodied systems, where a mistaken description or manipulated command can become a physical action. The review connects failures across perception, planning, instruction following and human-robot interaction, covering hallucinations, synthetic forgeries, adversarial attacks, privacy leakage and unsafe execution. It also shows how the same multimodal capabilities can support defense through contextual checking, forgery detection, privacy protection and risk-aware reasoning, offering a roadmap toward embodied intelligence that is not only capable, but dependable in real-world settings.

Vision-language models (VLMs) connect images with text, while vision-language-action models (VLAs) extend that connection to robot plans and control signals. This makes natural-language instruction, scene understanding and flexible task execution possible, but it also creates a chain of dependency: flawed data can distort perception, weak visual-language alignment can produce hallucinations, and malicious inputs can redirect decisions. In a chatbot, such errors may generate misinformation; in an autonomous vehicle or industrial robot, they may lead to collisions, damaged equipment or failed missions. Existing safeguards are often benchmark-specific, fragmented across system layers or too computationally costly for real-time use. Because of these challenges, deeper research is needed into unified, adaptive safeguards for multimodal agents operating under uncertain physical conditions.

Published (DOI: 10.1007/s11633-025-1626-x) online on July 13, 2026, in Machine Intelligence Research , the review was conducted by researchers from the Institute of Automation, Chinese Academy of Sciences; University College London (UCL); Minzu University of China; and the China Academy of Electronics and Information Technology. The team examined how VLMs and VLAs are used in embodied intelligence (EI), organized the major security threats and defensive approaches, and connected technical safety with accountability, fairness, privacy, environmental sustainability and human oversight. The article appears in the journal’s special issue on the security and ethics of generative AI.

The survey first tracks VLM and VLA use across four linked functions: perception, planning, instruction following and human-robot interaction (HRI). It then shows how failures can cascade. Biased training data, weak visual encoders or poor cross-modal alignment can make a model describe objects that are not present. Forged traffic signs, altered labels, cloned voices or deceptive captions can misguide perception and planning. Tiny adversarial perturbations, hidden backdoor triggers and multimodal jailbreak prompts may bypass safety controls, while persistent sensing can expose identity, location, possessions and social behavior. The authors organize countermeasures into equally connected layers. These include hallucination filtering and vision-grounded alignment; cross-modal forgery detection, watermarking and provenance tracing; defenses against perturbations, backdoors and jailbreaks; differential privacy (DP), secure multi-party computation (SMPC) and homomorphic encryption (HE); and safeguards for navigation, communications and physical control. A further strand uses causal explanations, intent alignment and risk assessment so robots can interpret ambiguous instructions, anticipate hazards and correct actions. The review's central insight is that no single filter can secure an embodied agent: protection must follow the entire path from sensor input to model reasoning, system architecture and physical execution.

The authors said the central challenge is not simply making models more accurate, but ensuring that a system remains safe when its sensors, language inputs and operating conditions are imperfect. They said defenses should be combined rather than deployed as isolated patches, with transparent risk metrics, continuous monitoring and human oversight for critical decisions. A trustworthy robot must also explain what it is doing, recognize when it is uncertain and fall back safely instead of acting with false confidence. The authors added that technical progress must move alongside privacy protection, fairness, accountability and responsible governance.

For developers and regulators, the survey provides a practical checklist for evaluating embodied systems before large-scale deployment. Future platforms could combine interpretable reasoning, attack detection, privacy-preserving computation and dynamic safety controls under reproducible, open evaluation protocols. The authors call for designs that address four dimensions together: technical robustness, regulatory alignment, social equity and environmental sustainability. Such an approach could support safer autonomous transport, healthcare assistance, warehouse automation, industrial inspection and collaborative robotics, while making responsibility easier to trace when failures occur. The review also warns that strong laboratory results may not transfer cleanly to noisy, culturally diverse and resource-constrained environments. Progress will therefore depend on cross-disciplinary cooperation and testing that measures not only task success, but safe behavior under stress.

###

References

DOI

10.1007/s11633-025-1626-x

Original Source URL

https://doi.org/10.1007/s11633-025-1626-x

Funding i nformation

This work was partially supported by the National Natural Science Foundation of China (Nos. 62506362, 62302539 and U21B2045), the Strategic Priority Research Program of Chinese Academy of Sciences, China (No. XDA0480302), and Engineering and Physical Sciences Research Council (EPSRC) Funded Grant, UK (No. EP/Y028805/1).

About Machine Intelligence Research

Machine Intelligence Research (original title: International Journal of Automation and Computing) is published by Springer and sponsored by the Institute of Automation, Chinese Academy of Sciences. The journal publishes high-quality papers on original theoretical and experimental research, targets special issues on emerging topics, and strives to bridge the gap between theoretical research and practical applications.

Machine Intelligence Research

Not applicable

Embodied Intelligence Security with Vision-language Models: A Survey

13-Jul-2026

The authors declare that they have no competing interests.

Keywords

Article Information

Contact Information

Licheng Ou
Machine Intelligence Research
mir_official@ia.ac.cn

Source

This article is based on a news release from Maximum Academic Press. BrightSurf curates and republishes science news from research institutions worldwide; the original release is linked below.

How to Cite This Article

APA:
Maximum Academic Press. (2026, August 12). When AI enters the physical world, safety gets real. Brightsurf News. https://www.brightsurf.com/news/1ZZY6271/when-ai-enters-the-physical-world-safety-gets-real.html
MLA:
"When AI enters the physical world, safety gets real." Brightsurf News, Aug. 12 2026, https://www.brightsurf.com/news/1ZZY6271/when-ai-enters-the-physical-world-safety-gets-real.html.