What is a vision-language-action model (VLA)?
A vision-language-action model (VLA) is an AI model that takes in camera images and a plain-language instruction. It directly outputs the movements a robot should make to carry out the task.
Also known as: VLA, VLA model
Researched and fact-checked by AI, with no human review. 7 sources listed below. How we verify
Last updated
How it works
A VLA starts from a vision-language model, the kind of AI that can describe a photo or answer questions about it. It adds a third skill: producing robot commands. RT-2 is a 2023 paper presented at the Conference on Robot Learning. Its authors gave that name to the category of models they proposed. They wrote robot actions as text tokens, the small units of text a language model reads and writes. So one model could learn from web images and text and from recorded robot movements.
Google DeepMind's July 2023 account of RT-2 said the robot data had been collected with 13 robots over 17 months in an office kitchen. It said success in situations new to the robot rose from 32% for the earlier RT-1 system to 62%.
Some later designs split the work in two. The robot maker Figure said in February 2025 that its Helix system combines two parts. A 7-billion-parameter vision-language model runs 7 to 9 times a second, it said, and a smaller 80-million-parameter controller issues commands 200 times a second. Parameters are the adjustable values a model learns in training.
Why it matters
The appeal is generality. A 2025 review is listed on arXiv as accepted by the journal IEEE Access. It says VLAs aim to learn control policies that carry over to different tasks, objects, robot bodies and settings. The authors expect this to let robots take on new jobs with little or no extra task-specific data. OpenVLA was presented at the 2024 Conference on Robot Learning. Its authors describe it as a 7-billion-parameter open-source model trained on 970,000 real-world robot demonstrations.
Where things stand in 2026
As of October 2026, Google DeepMind's product page lists Gemini Robotics 2 as its most advanced VLA, alongside a lightweight version that runs on the robot's own hardware. The company says the model can control entire humanoid robots. It says it is working with more than 100 trusted testers. These are the company's descriptions, not independent test results.
Whether more data will be enough is unsettled. Rest of World reported in June 2026 on robot training data, without naming VLAs. It said there is not yet enough evidence that teleoperation recordings and first-person videos will yield robots that function in arbitrary environments. Oregon State University robotics professor Alan Fern told the outlet the idea was plausible but unproven.
Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Proceedings of Machine Learning Research (Conference on Robot Learning 2023)
- RT-2: New model translates vision and language into action, Google DeepMind
- Helix: A Vision-Language-Action Model for Generalist Humanoid Control, Figure AI
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications, arXiv (Kawaharazuka et al., accepted to IEEE Access)
- OpenVLA: An Open-Source Vision-Language-Action Model, Proceedings of Machine Learning Research (Conference on Robot Learning 2024)
- Gemini Robotics, Google DeepMind
- China is training a robot future — one folded shirt at a time, Rest of World