Vision-Language-Action & Imitation Learning Companies

Vision-language-action models take camera images and an instruction and output robot actions directly. Together with imitation learning from teleoperated demonstrations, they are the dominant approach to general-purpose robot manipulation right now.

These are the companies and labs building such models, from open-weight releases you can fine-tune yourself to closed systems inside humanoid programs.

Last updated 2026-08-09 · 10 companies listed

Robot foundation model lab behind the pi-series VLA models (open-sourced pi0 and pi0.5), trained on large-scale teleoperation data.

Google DeepMind

United Kingdom / United States

Pioneered the VLA category with RT-2 and continues it with the Gemini Robotics model family.

NVIDIA

United States

Publishes the open GR00T N robot foundation models and the surrounding Isaac training stack.

Hugging Face

United States / France

Maintains LeRobot and released SmolVLA, a small open VLA model trainable on a single consumer GPU.

Figure AI

United States

Humanoid company whose Helix system drives the robot end-to-end from vision and language.

Skild AI

United States

Builds a general-purpose robot foundation model intended to run across many robot bodies.

1X Technologies

Norway / United States

Humanoid maker (NEO) betting on end-to-end learned control trained from demonstration data.

Tesla

United States

Trains the Optimus humanoid with imitation learning from teleoperation and video at scale.

Dyna Robotics

United States

Trains dexterous manipulation foundation models aimed at commercial deployment of single tasks.

Covariant

United States

Early robotic foundation model company for warehouse picking; core team joined Amazon in 2024.

Need robot training data instead of a list?

AY-Robots collects demonstration data on real robot arms through a global teleoperator network, and sells ready-made datasets on the marketplace.