Training guides

Train a policy on your arm

Every supported model on every supported arm, with the defaults the training backend really sends. Pick the row for your model and the column for your hardware.

  • 17 guides
  • 5 models, 4 arms
  • LeRobot v2.1 datasets

Find your combination

Rows are the models the training backend supports, columns are the arms the client can drive. A cell links to the guide for that exact pair. Where a cell has no link, that combination is not written up yet: the closest guide is the same model on another arm, since only the driver and the wiring change. If a whole row is empty, the model reference behind the model name has the parameters.

Training guides by policy and robot arm
ModelSO-100Full supportSO-101Full supportKoch v1.1CompatibleLeKiwiCompatible
GR00T N1.7A100 80 GB or H100 80 GBOpen guide90 min · intermediateOpen guide75 min · intermediateOpen guide90 min · advancedOpen guide100 min · advanced
GR00T N1.5A100 80 GB or H100 80 GBOpen guide75 min · intermediateNo guide yetNo guide yetNo guide yet
Pi0.5A100 80 GB or H100 80 GBOpen guide25 min · intermediateOpen guide18 min · intermediateOpen guide20 min · advancedOpen guide22 min · advanced
SmolVLARTX 4090 or any card with 24 GBOpen guide45 min · beginnerOpen guide35 min · beginnerOpen guide40 min · intermediateOpen guide50 min · advanced
ACTRTX 4090 or any card with 24 GBOpen guide25 min · intermediateOpen guide18 min · intermediateOpen guide20 min · advancedOpen guide22 min · advanced

Which model should I pick

The arm rarely decides this, the task and your budget do. All five models read the same LeRobot v2.1 datasets, so a dataset you record once can be trained more than once.

You just recorded a dataset and want to know whether it is worth scaling

Runs on a 24 GB card, so a full run costs a fraction of an A100 hour.

SmolVLA
RTX 4090 or any card with 24 GB
245 ms per action step
Typical run 2 to 5 hours, about 1 to 3 USD
You have a finished dataset and want the highest success rate

Pretrained on a multi robot corpus, so it copes with object positions your demonstrations never showed.

GR00T N1.7
A100 80 GB or H100 80 GB
152 ms per action step
Typical run 3 to 6 hours, about 4 to 12 USD
The task needs careful contact, for example inserting or stacking

Flow matching gives continuous trajectories, at the highest latency of the five models.

Pi0.5
A100 80 GB or H100 80 GB
485 ms per action step
Typical run 3 to 6 hours, about 4 to 12 USD
The arm has to react quickly and the task never changes

No pretraining and no language conditioning, but by far the shortest control loop.

ACT
RTX 4090 or any card with 24 GB
20 ms per action step
Typical run 2 to 5 hours, about 1 to 3 USD
You are reproducing a run that was started before N1.7 existed

Kept available for exactly that case. New projects should not start here.

GR00T N1.5
A100 80 GB or H100 80 GB
165 ms per action step
Typical run 3 to 6 hours, about 4 to 12 USD

Not sure your data is good enough yet

A real SO-100 is online and free to drive in the browser. Try the motion you want to teach before you record fifty episodes of it.