Vision-Language-Action & Imitation Learning Companies
Vision-language-action models take camera images and an instruction and output robot actions directly. Together with imitation learning from teleoperated demonstrations, they are the dominant approach to general-purpose robot manipulation right now.
These are the companies and labs building such models, from open-weight releases you can fine-tune yourself to closed systems inside humanoid programs.
Last updated 2026-08-09 · 10 companies listed
Physical Intelligence
United States
Robot foundation model lab behind the pi-series VLA models (open-sourced pi0 and pi0.5), trained on large-scale teleoperation data.
Google DeepMind
United Kingdom / United States
Pioneered the VLA category with RT-2 and continues it with the Gemini Robotics model family.
NVIDIA
United States
Publishes the open GR00T N robot foundation models and the surrounding Isaac training stack.
Hugging Face
United States / France
Maintains LeRobot and released SmolVLA, a small open VLA model trainable on a single consumer GPU.
Figure AI
United States
Humanoid company whose Helix system drives the robot end-to-end from vision and language.
Skild AI
United States
Builds a general-purpose robot foundation model intended to run across many robot bodies.
1X Technologies
Norway / United States
Humanoid maker (NEO) betting on end-to-end learned control trained from demonstration data.
Tesla
United States
Trains the Optimus humanoid with imitation learning from teleoperation and video at scale.
Dyna Robotics
United States
Trains dexterous manipulation foundation models aimed at commercial deployment of single tasks.
Covariant
United States
Early robotic foundation model company for warehouse picking; core team joined Amazon in 2024.
Need robot training data instead of a list?
AY-Robots collects demonstration data on real robot arms through a global teleoperator network, and sells ready-made datasets on the marketplace.