Every vision language action model, side by side
Parameters, GPU memory, inference latency, licence and what each model actually scored in simulation, on real hardware and after fine tuning. Every value links to the paper, model card or repository it came from. Nothing here is estimated unless it says so.
- 85 models
- 57 with open weights
- 32 developers
- 332 sourced results
The table
Sort by any column, filter by category, licence or whether this platform can train the model for you. GPU memory is the weights only footprint at bf16 computed from the parameter count, unless a published figure exists, in which case that one is shown and marked as reported.
Showing 85 of 85 models. Every number links to the source on the model page. Values from different benchmark suites are not comparable, so sort inside one column rather than across columns.
- ACEACE-Brain-0.5ACE-Brain, July 2026
- Parameters
- 8.8 B
- GPU memory
- 16 GB weights, 24 GB card
- Latency
- not published
- Weights
- see model card
- Simulation
- 98.2% LIBERO
- Real world
- 86.3% MindCube
- Gemini Robotics 2Google DeepMind, July 2026
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Real world
- 92% Multi-finger dexterity (Google DeepMind internal)
- Gemini Robotics ER 2Google DeepMind, July 2026
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Gemini Robotics On-Device 2Google DeepMind, July 2026
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- RBYLingBot-VLA 2.0Robbyant, July 2026
- Parameters
- 6 B
- GPU memory
- 11 GB weights, 16 GB card
- Latency
- 130 ms
- Weights
- Apache-2.0
- Simulation
- 93.52% RoboTwin 2.0
- Real world
- 66.2% GM-100
- AI2MolmoAct2Allen Institute for AI, May 2026
- Parameters
- 5.5 B
- GPU memory
- 10 GB weights, 16 GB card
- Latency
- 180 ms
- Weights
- Apache-2.0
- Simulation
- 100% LIBERO
- Real world
- 44.3% RoboEval
- AI2MolmoAct2-ThinkAllen Institute for AI, May 2026
- Parameters
- 5.5 B
- GPU memory
- 10 GB weights, 16 GB card
- Latency
- 790 ms
- Weights
- Apache-2.0
- Simulation
- 98.8% LIBERO
- XSRWall-OSS-0.5X Square Robot, May 2026
- Parameters
- 4 B
- GPU memory
- 7.5 GB weights, 12 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 96.5% LIBERO
- Real world
- 61.1% Real robot, 15-task suite (Wall-OSS-0.5 report)
- Gemini Robotics-ER 1.6Google DeepMind, April 2026
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Real world
- 93% Instrument reading (Google DeepMind internal)
- GENGEN-1Generalist AI, April 2026
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Real world
- 99% Generalist AI internal real-robot evaluation
- AGIGO-2 (Genie Operator-2)AgiBot, April 2026
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Simulation
- 98.5% LIBERO
- NVIDIA Isaac GR00T N1.7NVIDIA, April 2026
- Parameters
- 3.1 B
- GPU memory
- 16 GB reported
- Latency
- 27.9 ms
- Weights
- NVIDIA Open Model
- Simulation
- 98.45% LIBERO
Trainable here - ππ0.7Physical Intelligence, April 2026
- Parameters
- 5 B
- GPU memory
- 9.3 GB weights, 16 GB card
- Latency
- not published
- Weights
- closed
- Cosmos PolicyNVIDIA, January 2026
- Parameters
- 2 B
- GPU memory
- 6.8 GB reported
- Latency
- not published
- Weights
- Code Apache-2.0
- Simulation
- 98.5% LIBERO
- FIGHelix 02Figure AI, January 2026
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- RBYLingBot-VLARobbyant, January 2026
- Parameters
- 4 B
- GPU memory
- 7.5 GB weights, 12 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 88.56% RoboTwin 2.0
- Real world
- 35.41% GM-100
- UNIUnifoLM-VLA-0Unitree, January 2026
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- CC BY-NC-SA
- Simulation
- 100% LIBERO
- NVIDIA Isaac GR00T N1.6NVIDIA, December 2025
- Parameters
- 3.3 B
- GPU memory
- 6.1 GB weights, 12 GB card
- Latency
- 44 ms
- Weights
- NVIDIA OneWay NC
- Simulation
- 98.45% LIBERO
- SUTNORA-1.5SUTD, November 2025
- Parameters
- 4 B
- GPU memory
- 7.5 GB weights, 12 GB card
- Latency
- not published
- Weights
- MIT
- Simulation
- 95% LIBERO
- RynnVLA-002Alibaba, November 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 97.4% LIBERO
- Real world
- 50% Real robot (LeRobot SO100)
- ππ*0.6Physical Intelligence, November 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Real world
- 90% Physical Intelligence real-robot RECAP evaluation
- SAIInternVLA-M1Shanghai AI Laboratory, October 2025
- Parameters
- 4.1 B
- GPU memory
- 12 GB reported
- Latency
- not published
- Weights
- Code MIT
- Simulation
- 95.9% LIBERO
- THUX-VLA-0.9BTsinghua University, October 2025
- Parameters
- 880 M
- GPU memory
- 1.6 GB weights, 8 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 98.1% LIBERO
- Gemini Robotics 1.5Google DeepMind, September 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Gemini Robotics-ER 1.5Google DeepMind, September 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- THURDT2-FMTsinghua University, September 2025
- Parameters
- not published
- GPU memory
- 16 GB reported
- Latency
- not published
- Weights
- Apache-2.0
- THURDT2-VQTsinghua University, September 2025
- Parameters
- 8.3 B
- GPU memory
- 16 GB reported
- Latency
- not published
- Weights
- Apache-2.0
- XSRWALL-OSS (flow)X Square Robot, September 2025
- Parameters
- 4 B
- GPU memory
- 7.5 GB weights, 12 GB card
- Latency
- not published
- Weights
- Apache-2.0
- XSRWALL-OSS-FASTX Square Robot, September 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- Apache-2.0
- AI2MolmoAct-7B-DAllen Institute for AI, August 2025
- Parameters
- 8.1 B
- GPU memory
- 15 GB weights, 24 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 95.4% LIBERO
- AI2MolmoAct-7B-OAllen Institute for AI, August 2025
- Parameters
- 7.7 B
- GPU memory
- 14 GB weights, 24 GB card
- Latency
- not published
- Weights
- Apache-2.0
- RynnVLA-001Alibaba, August 2025
- Parameters
- 7 B
- GPU memory
- 13 GB weights, 24 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Real world
- 91.7% Real robot (LeRobot SO-100)
- GR-3ByteDance, July 2025
- Parameters
- 4 B
- GPU memory
- 7.5 GB weights, 12 GB card
- Latency
- not published
- Weights
- closed
- Real world
- 97.5% Real robot (ByteMini), table bussing (long-horizon)
- villa-XMicrosoft, July 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Simulation
- 90.1% LIBERO
- CASBitVLAChinese Academy of Sciences, June 2025
- Parameters
- 3 B
- GPU memory
- 1.4 GB reported
- Latency
- 73 ms
- Weights
- MIT
- Simulation
- 99% LIBERO
- Gemini Robotics On-DeviceGoogle DeepMind, June 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- SmolVLAHugging Face, June 2025
- Parameters
- 450 M
- GPU memory
- 0.8 GB weights, 8 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 87.3% LIBERO
- Real world
- 78.3% SO-100 real robot
Trainable here - BAIUniVLA (Unified Vision-Language-Action Model, BAAI)BAAI, June 2025
- Parameters
- 8.5 B
- GPU memory
- 16 GB weights, 24 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 95.5% LIBERO
- MIDChatVLA-2Midea Group, May 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- NVIDIA Isaac GR00T N1.5NVIDIA, May 2025
- Parameters
- 2.7 B
- GPU memory
- 5.1 GB weights, 8 GB card
- Latency
- not published
- Weights
- NVIDIA OneWay NC
- Simulation
- 47.5% RoboCasa
- Real world
- 93.3% Real GR-1 humanoid manipulation, language following
Trainable here - ODLUniVLA (task-centric latent actions)OpenDriveLab, May 2025
- Parameters
- 7.5 B
- GPU memory
- 14 GB weights, 24 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 95.2% LIBERO
- Real world
- 75% Real robot (UniVLA paper setup)
- SUTNORASUTD, April 2025
- Parameters
- 3 B
- GPU memory
- 8.3 GB reported
- Latency
- not published
- Weights
- MIT
- Simulation
- 87.9% LIBERO
- Real world
- 56.7% Real-world WidowX (authors' setup)
- ππ0.5Physical Intelligence, April 2025
- Parameters
- 3.6 B
- GPU memory
- 8 GB reported
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 98.8% LIBERO
Trainable here - Gemini RoboticsGoogle DeepMind, March 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- 250 ms
- Weights
- closed
- Real world
- 100% Gemini Robotics dexterous long-horizon specialist tasks
- Gemini Robotics-ERGoogle DeepMind, March 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Simulation
- 65% ALOHA 2 sim task suite
- Real world
- 65% ALOHA 2 real robot
- AGIGO-1 (Genie Operator-1)AgiBot, March 2025
- Parameters
- 3 B
- GPU memory
- 5.6 GB weights, 8 GB card
- Latency
- not published
- Weights
- CC BY-NC-SA
- Real world
- 60% Real robot (complex long-horizon and dexterous tasks)
- NVIDIA Isaac GR00T N1NVIDIA, March 2025
- Parameters
- 2.2 B
- GPU memory
- 4.1 GB weights, 8 GB card
- Latency
- 63.9 ms
- Weights
- NVIDIA OneWay NC
- Simulation
- 66.5% DexMimicGen (DexMG)
- Real world
- 76.8% Real-world GR-1 humanoid evaluation
- MIDChatVLAMidea Group, February 2025
- Parameters
- 3.4 B
- GPU memory
- 6.3 GB weights, 12 GB card
- Latency
- not published
- Weights
- MIT
- Real world
- 51.4% Real robot (authors' setup)
- MIDDexVLAMidea Group, February 2025
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Simulation
- 99.1% LIBERO
- FIGHelixFigure AI, February 2025
- Parameters
- 7.1 B
- GPU memory
- 13 GB weights, 24 GB card
- Latency
- not published
- Weights
- closed
- Real world
- 95% Figure internal logistics evaluation (parcel induction)
- Magma-8BMicrosoft, February 2025
- Parameters
- 8.9 B
- GPU memory
- 17 GB reported
- Latency
- 1100 ms
- Weights
- MIT
- Simulation
- 52.3% SimplerEnv
- Real world
- 67.3% AITW
- SUOpenVLA-OFTStanford University, February 2025
- Parameters
- 7.5 B
- GPU memory
- 14 GB weights, 24 GB card
- Latency
- 72.9 ms
- Weights
- MIT
- Simulation
- 97.1% LIBERO
- SAISpatialVLAShanghai AI Laboratory, January 2025
- Parameters
- 3.5 B
- GPU memory
- 8.5 GB reported
- Latency
- not published
- Weights
- MIT
- Simulation
- 78.1% LIBERO
- ππ0-FASTPhysical Intelligence, January 2025
- Parameters
- 3 B
- GPU memory
- 8 GB reported
- Latency
- 750 ms
- Weights
- Apache-2.0
- HKUMoto (Moto-GPT)HKU, December 2024
- Parameters
- 98 M
- GPU memory
- 0.2 GB weights, 8 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Simulation
- 74% SimplerEnv
- Real world
- 60% Real-world (authors' own three-task suite)
- SAISeerShanghai AI Laboratory, December 2024
- Parameters
- 316 M
- GPU memory
- 0.6 GB weights, 8 GB card
- Latency
- not published
- Weights
- closed
- Simulation
- 87.7% LIBERO
- Real world
- 78.4% Real-world (authors' own four-task suite)
- SAISeer-LargeShanghai AI Laboratory, December 2024
- Parameters
- 566 M
- GPU memory
- 1.1 GB weights, 8 GB card
- Latency
- not published
- Weights
- closed
- CogACT-BaseMicrosoft, November 2024
- Parameters
- 7.6 B
- GPU memory
- 30 GB reported
- Latency
- 181 ms
- Weights
- MIT
- Simulation
- 74.8% SIMPLER
- Real world
- 71.2% Real robot (Realman arm)
- CogACT-LargeMicrosoft, November 2024
- Parameters
- not published
- GPU memory
- 30 GB reported
- Latency
- not published
- Weights
- MIT
- Simulation
- 76.7% SIMPLER
- CogACT-SmallMicrosoft, November 2024
- Parameters
- not published
- GPU memory
- 30 GB reported
- Latency
- not published
- Weights
- MIT
- Simulation
- 73.3% SIMPLER
- GR-2ByteDance, October 2024
- Parameters
- 230 M
- GPU memory
- 0.4 GB weights, 8 GB card
- Latency
- not published
- Weights
- closed
- Real world
- 97.7% Real robot, multi-task (ByteDance platform)
- KSTLAPA (Latent Action Pretraining)KAIST, October 2024
- Parameters
- 7 B
- GPU memory
- 13 GB weights, 24 GB card
- Latency
- not published
- Weights
- MIT
- Simulation
- 62% Language Table (simulation)
- Real world
- 50.1% Real-world tabletop manipulation (Franka Emika Panda)
- THURDT-170MTsinghua University, October 2024
- Parameters
- 170 M
- GPU memory
- 0.3 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- THURDT-1BTsinghua University, October 2024
- Parameters
- 1.2 B
- GPU memory
- 2.2 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- Real world
- 100% Real robot (Mobile ALOHA, RDT paper protocol)
- ππ0 (pi-zero)Physical Intelligence, October 2024
- Parameters
- 3.3 B
- GPU memory
- 8 GB reported
- Latency
- 73 ms
- Weights
- Apache-2.0
- MITHPT-BaseMIT, September 2024
- Parameters
- 13 M
- GPU memory
- 0 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- Real world
- 70% Real-world (authors' own setup)
- MITHPT-LargeMIT, September 2024
- Parameters
- 51 M
- GPU memory
- 0.1 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- MITHPT-SmallMIT, September 2024
- Parameters
- 3 M
- GPU memory
- 0 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- MITHPT-XLargeMIT, September 2024
- Parameters
- 227 M
- GPU memory
- 0.4 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- Real world
- 76.7% Real-world (authors' own setup)
- MIDTinyVLAMidea Group, September 2024
- Parameters
- 1.3 B
- GPU memory
- 2.4 GB weights, 8 GB card
- Latency
- 14 ms
- Weights
- MIT
- Simulation
- 77.6% Meta-World
- Real world
- 94% Real-world Franka (authors' setup)
- UCBCrossFormerUC Berkeley, August 2024
- Parameters
- 130 M
- GPU memory
- 0.2 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- Real world
- 95% CrossFormer real-robot evaluation (paper Table 3)
- UCBEmbodied Chain-of-Thought (ECoT)UC Berkeley, July 2024
- Parameters
- 7.2 B
- GPU memory
- 13 GB weights, 24 GB card
- Latency
- not published
- Weights
- MIT
- Real world
- 72% ECoT accelerated inference study
- SUOpenVLAStanford University, June 2024
- Parameters
- 7.5 B
- GPU memory
- 16.8 GB reported
- Latency
- 239.6 ms
- Weights
- MIT
- Simulation
- 76.5% LIBERO
- Real world
- 85% Google robot (real RT-1 mobile manipulator)
- PKURoboMambaPeking University, June 2024
- Parameters
- 3.2 B
- GPU memory
- 6 GB weights, 12 GB card
- Latency
- not published
- Weights
- closed
- Simulation
- 63% SAPIEN manipulation
- UCBOcto-BaseUC Berkeley, May 2024
- Parameters
- 93 M
- GPU memory
- 0.2 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- Simulation
- 75.1% LIBERO
- Real world
- 72% Octo real-robot finetuning suite (paper Table I)
- UCBOcto-SmallUC Berkeley, May 2024
- Parameters
- 27 M
- GPU memory
- 0.1 GB weights, 8 GB card
- Latency
- not published
- Weights
- MIT
- UMA3D-VLAUMass Amherst, March 2024
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- closed
- Simulation
- 68% RLBench
- OXERT-1-XOpen X-Embodiment, October 2023
- Parameters
- not published
- GPU memory
- not published
- Latency
- not published
- Weights
- Apache-2.0
- Real world
- 50% Open X-Embodiment cross-lab evaluation
- OXERT-2-XOpen X-Embodiment, October 2023
- Parameters
- 55 B
- GPU memory
- 102 GB weights
- Latency
- not published
- Weights
- closed
- RT-2Google DeepMind, July 2023
- Parameters
- 55 B
- GPU memory
- 102 GB weights
- Latency
- not published
- Weights
- closed
- Simulation
- 90% Language-Table (simulation)
- RoboCatGoogle DeepMind, June 2023
- Parameters
- 1.2 B
- GPU memory
- 2.2 GB weights, 8 GB card
- Latency
- not published
- Weights
- closed
- SUACT (Action Chunking with Transformers)Stanford University, April 2023
- Parameters
- 80 M
- GPU memory
- 0.1 GB weights, 8 GB card
- Latency
- 10 ms
- Weights
- MIT
- Simulation
- 86% ALOHA simulation (MuJoCo)
- Real world
- 96% Real ALOHA hardware
Trainable here - CUDiffusion Policy (CNN-based, DiffusionPolicy-C)Columbia University, March 2023
- Parameters
- 278 M
- GPU memory
- 0.5 GB weights, 8 GB card
- Latency
- 100 ms
- Weights
- MIT
- Simulation
- 93% Robomimic
- CUDiffusion Policy (Transformer-based, DiffusionPolicy-T)Columbia University, March 2023
- Parameters
- 31 M
- GPU memory
- 0.1 GB weights, 8 GB card
- Latency
- not published
- Weights
- closed
- Simulation
- 90% Robomimic
- RT-1Google, December 2022
- Parameters
- 35 M
- GPU memory
- 0.1 GB weights, 8 GB card
- Latency
- not published
- Weights
- Apache-2.0
- Real world
- 97% RT-1 real-robot evaluation
Three different questions
A model that wins a simulation suite is not automatically the model that works on your bench, and neither of those tells you what happens after you fine tune it on fifty of your own episodes. The arena keeps the three apart on purpose.
Simulation
Reproducible benchmark suites. Everyone runs the same episodes, so the numbers can be compared inside a suite.
- 1AI2MolmoAct2100%LIBERO, Object
- 2UNIUnifoLM-VLA-0100%LIBERO, Object
- 3MIDDexVLA99.1%LIBERO, Object
- 4CASBitVLA99%LIBERO, object
- 5AI2MolmoAct2-Think98.8%LIBERO, Spatial
- 6ππ0.598.8%LIBERO, Spatial, fine-tuned checkpoint at 30k steps
Real world
Rollouts on physical hardware. The setups differ, so read these as evidence, not as a ranking.
- 1Gemini Robotics100%Gemini Robotics dexterous long-horizon specialist tasks, full long-horizon lunch-box packing
- 2THURDT-1B100%Real robot (Mobile ALOHA, RDT paper protocol), Pour Water-L-1/3, instruction following (correct amount poured)
- 3GENGEN-199%Generalist AI internal real-robot evaluation, aggregate over the reported tasks, about 1 hour of robot data per task
- 4GR-297.7%Real robot, multi-task (ByteDance platform), 105 manipulation tasks, simple setting
- 5GR-397.5%Real robot (ByteMini), table bussing (long-horizon), instruction following
- 6RT-197%RT-1 real-robot evaluation, seen tasks, over 700 instructions
Fine tuned tasks
What the model reaches after being adapted to a new task or a new robot, usually from a small number of demonstrations. This is the number that matters if you bring your own data.
- 1THURDT-1B100%Real robot (Mobile ALOHA, RDT paper protocol), Handover, 5-shot few-shot learning
- 2NVIDIA Isaac GR00T N1.598.8%Unitree G1 post-training, Place one of two fruits or objects, 1000 demonstrations
- 3AI2MolmoAct287.1%Real-world zero-shot (Franka DROID setup), 5 tasks, 15 trials per task, no per-task fine-tuning
- 4RoboCat86%RoboCat fine-tuning with 1000 demonstrations, KUKA lifting
- 5ππ0.785.6%Physical Intelligence real-robot evaluation, Zero-shot cross-embodiment shirt folding on bimanual UR5e, no UR5e laundry data in training
- 6UCBOcto-Small83%Octo ablation suite (real WidowX), 40 trials across two language-conditioned and two goal-conditioned tasks
These lists rank the best value each model reports on that axis, and the suite is printed next to it. A model at the top of the simulation list is not necessarily better than the one below it if the two ran different suites. Models that also report embodied reasoning benchmarks carry those on their own page, kept separate because no robot moves in them.
What practitioners say
Benchmarks measure what the authors chose to measure. The community rating is the other half: one vote per AY-Robots account, on any model, changeable at any time. Open a model page to cast yours.
No votes yet. If you have run any of these models on real hardware, you are exactly the person whose rating is worth something here. Create an account and open any model page.
Five kinds of model
The label VLA is used loosely. These groups separate what the models actually do, which matters more for your decision than the year they came out.
Foundation VLA
39 modelsLarge models pretrained on multi robot corpora, meant to be fine tuned rather than trained from scratch.
Compact VLA
10 modelsModels small enough to fine tune on a single consumer card, and in some cases to run on the robot itself.
Reasoning VLA
17 modelsModels that emit an intermediate plan, a subtask or a chain of thought before they emit actions.
Action policy
15 modelsPolicies without language conditioning. They learn one task from your demonstrations and are the baseline every VLA has to beat.
World model
4 modelsModels that predict future observations. They are used to generate training data or to plan, not to command joints directly.
5 of these run on your own arm from here
Record a dataset with an SO-100 class arm, pick a model, and the training pool rents the GPU by the hour. For those models the platform pages carry measured numbers from that pool rather than published ones: real defaults, real GPU tier, real inference latency.

How this list is built
A comparison table is only worth reading if you know where its numbers came from. Four rules, applied to every row.
Every number carries its source
Parameter counts, licences, latencies and benchmark results are taken from the paper, the model card or the official repository of the model in question, and the link sits next to the value. Nothing here is inferred from a similar model.
Missing is missing
Most publications do not report inference latency or GPU memory. Those fields stay empty instead of being filled with a plausible guess. An empty cell is a correct answer, an invented one is not.
Computed values are labelled
The only numbers the arena produces itself are the weights only memory footprint at a given precision and the smallest real GPU that fits it. Both are derived from the parameter count with a stated formula, and both are labelled as computed wherever they appear.
Comparison stays inside a suite
Results are grouped into simulation, real world and fine tuned tasks, and each value keeps its suite and split. Two models are only directly comparable when they report the same suite on the same split.
Last editorial update 2026-08-11. Company names and logos are the trademarks of their owners and appear here only to identify who published which model.
Questions
- What is a vision language action model?
- A vision language action model, usually shortened to VLA, takes camera images and a task written in plain language and outputs robot actions, normally joint positions or end effector deltas. The vision and language part is typically a pretrained vision language model, and an action head converts its representation into continuous commands. The point of the design is that a single policy can follow instructions it was never explicitly trained on, instead of one policy per task.
- Why does the arena not rank models with one overall score?
- Because the numbers come from different benchmark suites, and a LIBERO success rate and a real robot success rate do not measure the same thing. Averaging them would produce a number that looks authoritative and means nothing. The arena therefore sorts inside a column, always shows which suite a value came from, and links every value to the source it was taken from.
- How much GPU memory does a VLA need?
- Most papers do not publish a memory figure, so the arena computes the weights only footprint from the parameter count at bf16 precision and marks it as computed. Weights are a lower bound: activations, the image encoder intermediates, the KV cache and the CUDA context come on top, which in practice is roughly 30 to 50 percent more for inference. Fine tuning is a different order of magnitude again, because optimiser state and gradients usually cost several times the weights unless you use LoRA or a similar adapter.
- Which of these models can I actually train on AY-Robots?
- Five of them: GR00T N1.7, GR00T N1.5, Pi0.5, SmolVLA and ACT. Those are the models the training pool can fine tune on a LeRobot dataset that you record with your own arm, and the platform pages for them carry the real defaults, GPU tier and measured inference latency from that pool. Every other model in the arena is listed for comparison and links to its own repository.
- What does the community rating mean?
- It is a one to five rating from people with an AY-Robots account, one vote per account and per model, changeable at any time. It is a subjective signal about how a model behaved in practice, not a benchmark. Read it next to the measured numbers, not instead of them.
- How current is this list?
- Every entry carries its own editorial date and the sources it was built from. The field moves fast, so a model can gain a new checkpoint or a new licence between updates. Where a source contradicts what you read here, trust the source and tell us, the link is on every model page.
Drive a real arm before you pick a model
A physical SO-100 is online and free to control in the browser. No signup, no client, no hardware of your own.