Policies

Robot learning policy models, compared

Five models can be trained on a dataset you record with your own arm. They differ in what they cost to train, how fast they run on hardware, and how much data they need before they do anything useful. This page puts the numbers side by side.

  • 5 supported models
  • LeRobot v2.1 datasets
  • GPUs rented by the hour

All five models at a glance

The GPU tier is what the training pool requests, not a recommendation you can override. A model on the 80 GB tier cannot be fine-tuned on a consumer card, and inference latency is measured per action step in that same pool.

ModelVendorParametersGPU tierInferenceDefault stepsMin. episodes
GR00T N1.7NVIDIAabout 3 billionA100 80 GB or H100 80 GB152 ms per step20,00050
Pi0.5Physical Intelligenceabout 3 billionA100 80 GB or H100 80 GB485 ms per step30,00050
SmolVLAHugging Faceabout 450 millionRTX 4090 or any card with 24 GB245 ms per step20,00030
ACTStanford (ALOHA)about 80 millionRTX 4090 or any card with 24 GB20 ms per step100,00050
GR00T N1.5NVIDIAabout 3 billionA100 80 GB or H100 80 GB165 ms per step2,00050

Default steps and the minimum episode count are starting points the training form fills in for you, not limits. Every trainer counts steps, none of them counts epochs.

Try the arm before you pick a model

A real SO-100 is online and free to drive in the browser. No signup, no client, no hardware of your own.