
An honest seven-day schedule for a first SO-100 or SO-101 build: print, motor ids, assembly, calibration, teleoperation, recording, training, rollout, and where each day loses people.
Week one, without the marketing
- •Seven days only works if the parts are already on your desk. The SO-ARM100 bill of materials sources most servos from Alibaba, and shipping is the longest single line item in the entire project.
- •Day two is the quiet killer. lerobot-setup-motors writes an id and a baudrate into each servo's EEPROM one motor at a time, and the SO-100 docs tell you to do it before assembly because the connectors are not easy to reach afterwards.
- •The recording day is longer than anyone expects. lerobot-record defaults to 50 episodes, 60 seconds per episode and 60 seconds of reset. That is 100 minutes of continuous clock time before you retake a single episode.
- •Training is the shortest day. The LeRobot compute guide puts ACT at roughly 30 to 60 minutes for 5 epochs over a 45k-frame dataset on one RTX 4090, and 6 to 14 hours on Apple Silicon.
- •Day seven is where the honest disappointment lives. A first policy usually works in the exact lighting, camera pose and object position you recorded in, and nowhere else.
- •On AY-Robots the pure waiting days collapse into a form: the backend rents the GPU by required VRAM and the inference pod destroys itself when idle. The building, calibrating and recording days do not shrink. Nobody records your episodes for you.
What a first week actually looks like
Most first-week guides for a low-cost arm are written backwards. They start from the finished demo, walk you through the commands that produced it, and quietly leave out the three evenings that went nowhere. This one is written forwards. It follows the same order you will actually hit it in: parts, print, motor ids, assembly, calibration, teleoperation, recording, training, and the first rollout that does not work. For each day there is a point where a predictable number of people stop, and naming that point in advance is most of what makes the difference.
The hardware is the SO-100 and its successor the SO-101, both built around Feetech STS3215 bus servos. The software is LeRobot. Everything below is checked against LeRobot v0.6.1, released 3 August 2026, and against the SO-ARM100 repository as it stood in August 2026. Upstream moves fast enough that command names change between releases, so if a flag does not exist for you, check which version you installed before assuming the article is wrong.
| Day | What you are doing | Hands-on time | Where people stop |
|---|---|---|---|
| Before day 1 | Ordering parts, waiting for servos | 1 hour, then weeks of waiting | Buying one arm's servo list for the other arm: the SO-101 leader mixes three gear ratios, the SO-100 leader is six identical motors with the gears taken out |
| Day 1 | 3D printing and parts inventory | 2 hours of work, a printer running much longer | Skipping the gauge print, then finding the servo does not fit the printed holder |
| Day 2 | Motor ids and baudrates | 1 to 2 hours for twelve motors | One motor does not answer, and the fault is cabling, power or a jumper on the wrong channel |
| Day 3 | Assembling follower and leader | 2 to 4 hours for both arms | Realising the wires should have gone in before the part was closed |
| Day 4 | Calibration and first teleoperation | 1 to 2 hours | The follower twitches and sags instead of following, or a joint stops short |
| Day 5 | Cameras and the recording session | 3 to 4 hours | Hour two of repeating the same motion, with 30 episodes still to go |
| Day 6 | Training a first policy | 1 hour of work, hours of waiting | No GPU, or a loss that falls beautifully while the arm does nothing |
| Day 7 | Running the policy on the arm | 2 hours | It works once, then fails after you move the camera five centimetres |
Before day one: the order that decides the week
The single most common way to lose a week is to order the wrong servos, and which servos are wrong depends on which arm you are building. The bill of materials on the front page of the SO-ARM100 repository is the SO-101 one, and it is not twelve identical motors: it lists seven STS3215 units with 1/345 gearing, two with 1/191 and three with 1/147, because the SO-101 leader has to be back-drivable by a human hand while the follower needs holding torque. That list totals $229.88 in the US and EUR 226.30 in the EU, with a single follower arm at $121.94. The older SO-100 list, which the repository now marks as deprecated, has the simpler shape: twelve identical STS3215 servos for the pair at $232 and EUR 244, or six at $123 and EUR 128 for a single follower, with the leader made back-drivable by removing its gears rather than by buying different ones. AY-Robots quotes the SO-100 at roughly 110 to 150 EUR in parts and the SO-101 at 130 to 170 EUR, which lands in the same place once you add local shipping.
| Arm | Servos | Voltage | Parts cost (AY-Robots figure) | Support level |
|---|---|---|---|---|
| SO-100 | Feetech STS3215 bus servos | 7.4 V | ~110 to 150 EUR | Full, reference arm |
| SO-101 | Feetech STS3215 | 7.4 V | ~130 to 170 EUR | Full |
| Koch v1.1 | Dynamixel XL330 / XL430 | 5 V and 12 V rails | ~250 to 350 EUR | Compatible |
| LeKiwi | Feetech STS3215 (arm) | 7.4 V arm, 12 V base | ~400 to 500 EUR | Compatible |
The STS3215 ships in a 7.4 V and a 12 V variant. The SO-ARM100 repository puts the 7.4 V stall torque at 16.5 kg.cm at 6 V and the 12 V version at 30 kg.cm, and states plainly that if you buy the 12 V motors you also need a 12 V 5 A or larger power supply instead of the 5 V one, and that on the SO-101 the leader arm is always 7.4 V. The reverse mistake is fatal: feeding 12 V into a 7.4 V bus destroys the servos, and there is no calibration step that recovers from it. Check the label on the barrel jack before the first power-on, every time. If a servo has already stopped answering, servo not responding walks through what is recoverable.
If you are choosing between the two arms, read SO-100 against SO-101 before ordering rather than after. The SO-ARM100 repository now marks its own SO-100 documentation as deprecated and points new builders at the SO-101, which the same README describes as having improved wiring, updated motors for the leader arm and no gear removal. AY-Robots still treats the SO-100 as its reference arm because so many of them exist, but if you are buying parts today there is very little reason to build the older one.
Day 1: print, and check the print before you trust it
The STL files ship pre-oriented so that a whole arm prints from a single file, with variants for a 220 x 220 mm bed and a 205 x 250 mm bed. The tested settings are PLA+, a 0.4 mm nozzle at 0.2 mm layer height (or 0.6 mm at 0.4 mm), and 15 percent infill. The repository does not publish an estimated print time, which is the honest answer: it depends entirely on your printer, and the printer is busy rather than you.
The repository includes a Gauges folder with two small test parts: one checks your printer against a standard 4x2 Lego block, the other against an actual STS3215 servo. They take minutes. Printing an entire arm and only then discovering that your printer runs 0.15 mm tight is a full day thrown away, and it happens often enough that upstream shipped the gauges specifically to prevent it.
Do the boring inventory on day one too. Lay out the motors and confirm the gear ratios are what you ordered, count the M2x6 and M3x6 screws, and confirm you have a Phillips #0 and #1. The upstream parts list recommends exactly those two sizes, and a stripped M2 head in a motor holder is a far worse afternoon than a trip to a hardware shop.
Day 2: motor ids, the step that eats an evening
Every servo on the bus needs a unique id, and every brand-new STS3215 arrives with the same default id of 1. Setting those ids is a physical, one-at-a-time ritual: connect exactly one motor to the controller board, press Enter, let the script write the id and baudrate into the motor's EEPROM, then move the cable to the next motor. Twelve motors across two arms means twelve rounds of that. There is no batch mode, because the bus cannot address two motors that share an id.
- 1Install LeRobot and the Feetech SDK
LeRobot needs Python 3.12 or newer. The base package is deliberately lightweight, so the robot workflow extras and the Feetech motor support are separate installs.
bashconda create -y -n lerobot python=3.12 conda activate lerobot conda install ffmpeg -c conda-forge git clone https://github.com/huggingface/lerobot.git cd lerobot pip install -e ".[core_scripts]" # record, replay, calibrate pip install -e ".[feetech]" # STS3215 bus servos - 2Find the serial port for each controller board
Run it once per arm. The script asks you to unplug the USB cable and press Enter, then prints the port that disappeared. On Linux you may need to open permissions on the device node first.
bashlerobot-find-port # Linux only, if the port is present but not writable: sudo chmod 666 /dev/ttyACM0 sudo chmod 666 /dev/ttyACM1 - 3Write ids and baudrates, one motor at a time
The script prompts for the gripper first and walks backwards down the arm. Connect only the named motor, and make sure it is not yet daisy-chained to any other. Repeat the whole procedure for the leader with the teleop flags.
bashlerobot-setup-motors \ --robot.type=so100_follower \ --robot.port=/dev/ttyACM0 # expected prompts, in order: # Connect the controller board to the 'gripper' motor only and press enter. # 'gripper' motor id set to 6 # Connect the controller board to the 'wrist_roll' motor only and press enter. lerobot-setup-motors \ --teleop.type=so100_leader \ --teleop.port=/dev/ttyACM1
The SO-100 documentation is explicit: unlike the SO-101, the motor connectors are not easily accessible once the arm is assembled, so motor configuration has to happen beforehand. Builders who assemble first and configure second end up taking the arm apart. The other half of the SO-100 tax is gear removal: all six leader motors need their gears removed so the leader can be moved by hand, which the SO-101 does not require at all. If a board simply will not talk to a motor, check the power supply, the USB cable, the 3-pin cable, and on a Waveshare board that both jumpers sit on the B channel. Arm not detected covers the rest.
Day 3: assembly, and the wires you forgot to route
The upstream guidance is that a first arm takes a bit more than an hour, and a second arm somewhat less once the sequence is familiar. That is accurate for the screwing. It is not accurate for a first build, because the assembly videos show the mechanical steps without showing when the cables go in, and the documentation says so directly: inserting the cables beforehand is much easier than afterwards. Every joint has a cable that has to be routed through the part before the part is closed, and every time you get that wrong you undo four M3 screws.
Two specific fits are tight by design and worth knowing about in advance. The shoulder motor holder on the SO-100 usually needs a workbench or a similar hard surface to press the part around the motor, and the fourth motor holder is the same story. If a part will not go on with hand pressure, that is expected, not a print failure. Force applied to a printed part in the right direction is fine; force applied to a motor horn is how you lose the position you just set.
- You know where every cable runs, which makes every later fault diagnosable instead of mysterious.
- EUR 226.30 for a two-arm SO-101 setup, or EUR 244 for the older twelve-servo SO-100 list, against several hundred for the assembled arms sold by the vendors the repository links.
- Reprinting a broken gripper finger costs an hour, not a support ticket.
- You learn what the six joints physically do, which turns out to matter when you later look at a joint trajectory and ask why it saturated.
- Two to four hours of assembly, plus printing, plus the motor-id evening, before you have written a single line of code.
- A tolerance problem in your printer shows up as a mechanical mystery three days later.
- Vendors listed in the SO-ARM100 repository sell parts kits and assembled arms, and for a team that only cares about data collection that is usually the cheaper choice in total hours.
- Nothing about building the arm teaches you anything about the policy, which is where the actual difficulty lives.
Day 4: calibration, then the first minute of teleoperation
Calibration is two gestures. Move every joint to the middle of its range and press Enter, then move every joint through its full range of motion. That is it. What it produces is a mapping between raw encoder counts and normalised joint values, stored under the id you passed on the command line. Use the same id later during recording and rollout, because that is how LeRobot finds the calibration file. The upstream docs make the point that matters: calibration is what lets a network trained on one arm work on another.
lerobot-calibrate \
--robot.type=so100_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm
lerobot-calibrate \
--teleop.type=so100_leader \
--teleop.port=/dev/ttyACM1 \
--teleop.id=my_leader_armThen teleoperate. This is the first moment the project feels real, and it is also the first moment a mechanical problem becomes visible. If the follower shakes and then droops, you are usually looking at a power or torque problem rather than a software one, and arm twitches then sags lists the usual causes. If one joint stops well short of where the leader is, the range of motion sweep did not cover that joint properly and joint stops early applies. The leader-follower scheme itself is simple: the leader publishes joint positions, the follower is commanded to match them.
lerobot-teleoperate \
--robot.type=so100_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm \
--teleop.type=so100_leader \
--teleop.port=/dev/ttyACM1 \
--teleop.id=my_leader_arm \
--display_data=trueIf your parts have not arrived yet, you do not have to wait to learn what teleoperation feels like. The live arm on AY-Robots is a physical SO-100 streaming in the browser with no signup, handed out through a queue. Ten minutes of driving a real arm teaches you more about why demonstrations are hard than an hour of reading, and it costs nothing. The try page collects the other things you can do without hardware.
Day 5: cameras, then the long recording session
Cameras first, because they are the part of the setup most likely to change under you. LeRobot discovers OpenCV and RealSense devices for you, and the documentation warns that the identifiers can change after a reboot or after re-plugging, depending on the operating system. That is not a small caveat. A dataset recorded with wrist and top swapped is a dataset that trains a policy to reach in the wrong direction, and nothing about the loss curve will tell you.
lerobot-find-cameras opencv # or: lerobot-find-cameras realsense
# --- Detected Cameras ---
# Camera #0:
# Name: OpenCV Camera @ 0
# Type: OpenCV
# Id: 0
# Backend api: AVFOUNDATION
# Default stream profile:
# Format: 16.0
# Width: 1920
# Height: 1080
# Fps: 15.0
#
# The Id field is what you pass as index_or_path.The Hugging Face guidance on what makes a good dataset is short and specific: two camera views, steady framing, neutral and stable lighting, sharp focus, and at least 480x640. The leader arm must not appear in frame, and the only things moving should be the follower and the object. Name your features by what they see rather than by the device, so images.top and images.wrist rather than images.laptop. Write the task description as a real sentence of roughly 25 to 50 characters, because a vision-language model will actually read it.

Now the number that surprises people. LeRobot dataset recording defaults, straight out of the config dataclass, are 50 episodes at 30 fps, 60 seconds per episode and 60 seconds of reset between episodes. Multiply it: 50 times 120 seconds is 6000 seconds, or 1 hour 40 minutes of unbroken clock time, assuming you never retake anything and never need a break. In practice you will retake, because the guidance is to keep grasping behaviour consistent, and inconsistent takes are worth deleting rather than keeping.
| lerobot-record flag | Default | What it actually costs you |
|---|---|---|
| --dataset.num_episodes | 50 | The upstream tip is at least 50 episodes, roughly 10 per object location |
| --dataset.episode_time_s | 60 | The compute guide sizes its own reference dataset at 30 s per episode; a narrow pick-and-place needs less, so shorten this or end the take early with the right arrow |
| --dataset.reset_time_s | 60 | Time to put the object back. Also 60 seconds by default, and also usually too long |
| --dataset.fps | 30 | Matches the camera guidance; the dataset fps and the requested fps must agree or recording refuses to start |
| --dataset.push_to_hub | true | Your dataset uploads to the Hub by default. Set it to false if the task is not public |
| --dataset.streaming_encoding | false | Off by default. Turning it on encodes during capture and makes saving an episode near-instant |
By default lerobot-record appends a date-time stamp to the repo id so each session gets a unique name, unless you pass --dataset.no_stamp. That is helpful once you know it and confusing the first time you go looking for a dataset that is not where you expected. During recording, right arrow saves the episode and moves on, left arrow deletes the current episode and retries, and Escape stops, encodes and uploads.
lerobot-record \
--robot.type=so100_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm \
--robot.cameras="{ top: {type: opencv, index_or_path: 1, width: 640, height: 480, fps: 30}, wrist: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30} }" \
--teleop.type=so100_leader \
--teleop.port=/dev/ttyACM1 \
--teleop.id=my_leader_arm \
--dataset.repo_id=${HF_USER}/so100_pick_place \
--dataset.num_episodes=50 \
--dataset.single_task="Put the red brick in the bowl" \
--dataset.episode_time_s=20 \
--dataset.reset_time_s=10 \
--dataset.streaming_encoding=true \
--display_data=trueThis is the day where people give up quietly. Not with an error message, just with boredom around episode 28. The counter-measure is scheduling: record in two blocks with a real break, decide the object positions in advance so you are not improvising, and keep the task narrow. Record your first dataset walks the same ground with the AY-Robots client, and our guide to collecting high-quality VLA training data goes deeper on what to vary and what to hold fixed.
Two routes through the same week
The manual route and the platform route are not different workflows, they are the same workflow with different amounts of infrastructure work attached. Here is the same goal, from recorded episodes to a policy running on the arm, done both ways.
You install LeRobot, record locally, find a GPU, run the trainer, and serve the checkpoint yourself. Everything is under your control and everything is your problem.
# 1. train, locally or on a rented box
lerobot-train \
--dataset.repo_id=${HF_USER}/so100_pick_place \
--policy.type=act \
--output_dir=outputs/train/act_so100 \
--job_name=act_so100 \
--policy.device=cuda \
--steps=20000 \
--policy.scheduler_decay_steps=20000
# 2. run it back on the arm
lerobot-rollout \
--strategy.type=base \
--policy.path=outputs/train/act_so100/checkpoints/last/pretrained_model \
--robot.type=so100_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm \
--robot.cameras="{ top: {type: opencv, index_or_path: 1, width: 640, height: 480, fps: 30}, wrist: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30} }" \
--task="Put the red brick in the bowl" \
--duration=60- You need a CUDA machine, or you pay per second on Hugging Face Jobs with --job.target, or you accept 6 to 14 hours on Apple Silicon for ACT.
- The trainer default is 100,000 steps, while the compute guide says imitation learning usually converges in 5 to 10 epochs over the dataset. On a 45k-frame dataset at batch 8 that is about 5,600 steps per epoch, so 5 to 10 epochs is roughly 28,000 to 56,000 steps and the default runs about two to four times longer than the guidance suggests.
- If you shorten the run, shorten the learning-rate schedule with --policy.scheduler_decay_steps as well, or the rate never decays.
- In LeRobot v0.6.x inference is lerobot-rollout. In v0.5.1 the same job was done by lerobot-record. This is exactly the kind of thing that makes older tutorials fail.
You pick a model, a dataset and hyperparameters in a form. The backend rents a GPU on a spot market sized by the VRAM the model needs, runs the trainer, and writes checkpoints to object storage. For inference, a pod is auto-provisioned to serve the policy and the local robot client talks to that endpoint.
- Five trainable policies: ACT, SmolVLA, GR00T N1.5, GR00T N1.7 and Pi0.5. The full comparison lives at the policies page.
- Datasets come from the public directory, from a Hugging Face repo id, or from your own machine via the desktop client.
- A run on the RTX 4090 tier (ACT, SmolVLA) takes 2 to 5 hours at 0.30 to 0.60 USD per hour, roughly 1 to 3 USD. The A100 or H100 tier (GR00T, Pi0.5) takes 3 to 6 hours at 1.20 to 2.00 USD per hour, roughly 4 to 12 USD.
- Inference pods carry an idle watchdog and destroy themselves after an idle period, so a forgotten pod does not bill overnight.
- The same operations are exposed through the CLI and the MCP server, so an agent can drive the same steps.
What it does not do: it does not print your parts, it does not remove the gears from your SO-100 leader motors, and it does not perform your calibration. Day 1 through day 4 are identical either way. The honest saving is on days 6 and 7.
Day 6: training, which is mostly waiting
Pick ACT for the first run. It is the smallest of the five trainable policies at roughly 80 M parameters, it trains from scratch rather than fine-tuning a foundation model, it runs on any 24 GB card, and it has the lowest inference cost of the group at 20 ms per action step. It is also the policy the LeRobot imitation-learning tutorial uses in its own worked training example. The trade is that ACT has no base model at all: it only exists after training on your task, so it will never generalise beyond what you recorded.
| Policy | Params | GPU tier | Inference per action step | Min episodes | Dataset format |
|---|---|---|---|---|---|
| ACT | ~80 M | RTX 4090 or any 24 GB card | 20 ms | 50 | LeRobot v3.0 |
| SmolVLA | ~450 M | RTX 4090 or any 24 GB card | 245 ms | 30 | LeRobot v3.0 |
| GR00T N1.7 | ~3 B, ~40 M trained | A100 80 GB or H100 80 GB | 152 ms | 50 | LeRobot v2.0 or v2.1 |
| GR00T N1.5 | ~3 B | A100 80 GB or H100 80 GB | 165 ms | 50 | LeRobot v2.0 or v2.1 |
| Pi0.5 | ~3 B, PaliGemma backbone | A100 80 GB or H100 80 GB | 485 ms | 50 | LeRobot v3.0 |
Current LeRobot writes datasets at codebase version v3.0. GR00T's loader wants v2.0 or v2.1 and crashes on a v3.0 dataset, so a week-one recording has to be converted down before a GR00T run will start. If you were planning to jump straight to a large VLA, this is where the plan stalls. Dataset rejected as v3 covers the conversion. ACT, SmolVLA and Pi0.5 take v3.0 directly, which is another reason ACT is the sane first run.
For wall-clock expectations, the LeRobot compute guide gives indicative figures for 5 epochs over a 50-episode dataset of about 45,000 frames at 640x480. ACT on a single RTX 4090 or 3090 lands at roughly 30 to 60 minutes. The same job on an L4 or A10G is 1 to 2 hours. SmolVLA on an A100 40 GB at batch 16 is 1 to 2 hours. Apple Silicon with MPS runs ACT in 6 to 14 hours, which is technically possible and practically a way to lose day six. The guide is explicit that these are order-of-magnitude figures and real runs deviate by plus or minus 50 percent.

The failure mode on this day is not a crash. It is a training curve that looks perfect while the policy does nothing useful, which is common enough that it has its own page at loss falls, policy does nothing. Behaviour-cloning loss measures how well the network predicts the action you recorded, not whether the resulting trajectory picks anything up, and while LeRobot v0.6 can compute a held-out loss on an evaluation split, that is off by default (eval_steps is 0) and a lower behaviour-cloning loss still does not rank checkpoints by task success. The only honest evaluation is running it on the arm. If the run itself dies, out of memory during training and training job stuck queued cover the two most common stops.
Day 7: the rollout, and the disappointment that follows it
Deployment is one command, and it needs to match the recording setup exactly. Same camera names, same resolutions, same robot id so the calibration file is found, same task string. LeRobot v0.6 exposes five rollout strategies: base for a plain autonomous run with no recording, sentry for continuous recording with auto-upload, highlight for ring-buffer recording saved by a keystroke, dagger for human-in-the-loop collection, and episodic for episode-oriented runs with reset phases. Every strategy also accepts --inference.type=rtc, real-time chunking for the slow VLA models; the default backend is sync, one policy call per control tick. Start with base and sync.
lerobot-rollout \
--strategy.type=base \
--policy.path=${HF_USER}/act_so100_pick_place \
--robot.type=so100_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=my_follower_arm \
--robot.cameras="{ top: {type: opencv, index_or_path: 1, width: 640, height: 480, fps: 30}, wrist: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30} }" \
--task="Put the red brick in the bowl" \
--duration=60What usually happens on a first attempt: the arm moves confidently towards roughly the right place, closes the gripper slightly early or slightly late, and either succeeds or nudges the object out of reach. Then you move the object 15 cm and it fails completely. This is normal and it is not a bug in your setup. ACT predicts a chunk of 100 actions at a time by default and executes all 100 of them, so once a chunk commits to the wrong approach there is no correction until the next one. Action chunking is what makes these policies smooth and it is also why they look stubborn.
- One task, one object, one background, with the object starting in positions you actually recorded.
- Smooth motion that looks like your own demonstrations, because that is exactly what it learned.
- Recovering from small position offsets, if you recorded roughly 10 episodes per object location as the upstream tip suggests.
- A concrete measurement: 20 attempts, count successes, and you have a number to improve against.
- New objects, new lighting, or a camera you moved after recording. See policy only works in one setup.
- Language instructions it never saw. ACT has no language grounding at all; that is what the VLA models are for.
- Recovering from its own mistakes. Failure recovery only appears in the policy if failure recovery appears in your data.
- Reliable gripper timing on transparent or thin objects, which is the single most common first-week failure. Gripper does not close is worth reading before you blame the model.
Inference has to sit next to the servos for fast tasks. The control loop is 20 to 485 ms per action step depending on the model, and adding public-internet round trips to that turns a working policy into a hesitant one. Remote inference on AY-Robots is genuinely viable for slow pick-and-place. It is not viable for fast reactive motion, and ACT at 20 ms per step is precisely the model that suffers most from a network hop. If your task needs tight reaction, run the policy locally and use the platform for training only.

What week two is for
The week-one goal is not a good policy. It is a closed loop: hardware that answers, a dataset that loads, a run that finishes, and a checkpoint that moves the arm. Once that loop exists, every improvement is cheap, because you are changing one variable against a known baseline instead of debugging four things at once.
- Measure before you change anything. Twenty attempts from the same start positions gives you a success rate, and without one you cannot tell an improvement from a good afternoon.
- Add the variation you actually need, one axis at a time. The upstream advice is explicit that adding too much variation too quickly hurts results.
- Record 20 more episodes covering exactly the failures you saw, rather than 200 more of what already works.
- Only then try a larger model. SmolVLA needs 30 episodes minimum on this platform and adds language conditioning at 245 ms per action step; GR00T N1.7 needs 50 and an A100-class card.
- Compare before you commit. The arena holds 85 VLA models with 332 benchmark results, each value linked to its paper or model card.
If you want the longer form of any single day, the SO-100 getting-started track covers hardware and calibration, train your first policy covers day six, and run your first policy covers day seven. The complete SO-100 guide is the single longest reference we have on the arm itself, and the VLA overview explains what the larger models buy you once ACT stops being enough.
Can a complete beginner really go from parts to a working policy in seven days?▾
Yes, if the parts are already on the desk and the definition of working is narrow: one task, one object position range, one lighting setup. The seven days of hands-on work are realistic. What is not realistic is a policy that generalises, which takes several more rounds of recording and evaluation. Treat week one as building the pipeline, not the product.
Should I build the SO-100 or the SO-101?▾
The SO-101 unless you already own SO-100 parts. The SO-ARM100 README marks the SO-100 documentation as deprecated and points new builders at the SO-101, which it describes as having improved wiring, updated motors for the leader arm and no gear removal. The practical differences show up at ordering time and at build time. Ordering: the SO-101 leader mixes three gear ratios (three 1/147, two 1/191, one 1/345), while the SO-100 pair is twelve identical STS3215 servos. Building: the SO-100 leader needs the gears removed from all six motors, and its connectors are hard to reach once assembled, so its motor configuration has to happen first. Both are fully supported on AY-Robots.
How many episodes do I actually need?▾
The LeRobot tutorial suggests at least 50 episodes with about 10 per object location for a grasp-and-place task. AY-Robots sets a minimum of 50 episodes for ACT, GR00T N1.5, GR00T N1.7 and Pi0.5, and 30 for SmolVLA. Below those numbers the run will complete and the policy will not be useful, which is a more expensive outcome than not starting.
What does the training day cost if I do not own a GPU?▾
On AY-Robots a run on the RTX 4090 or 24 GB tier, which covers ACT and SmolVLA, takes 2 to 5 hours at 0.30 to 0.60 USD per hour, so roughly 1 to 3 USD. The A100 80 GB or H100 tier, which GR00T N1.5, GR00T N1.7 and Pi0.5 need, takes 3 to 6 hours at 1.20 to 2.00 USD per hour, so roughly 4 to 12 USD. These are spot-market prices, hence the ranges.
My policy works in the morning and fails in the afternoon. What changed?▾
Almost always the light. A policy trained on 50 episodes recorded under one lighting condition has no reason to be invariant to another, and window light moving across a table is a bigger distribution shift than it looks. Move the setup away from a window, record under the light you will actually run under, and check the camera has not silently changed exposure. The failure-mode pages under /fix cover the rest.
Do I need two cameras, or is one enough?▾
The Hugging Face dataset guidance prefers two views, and in practice a wrist camera plus a fixed overhead or side view is the working default. One view is enough to train something, but depth ambiguity around the grasp is where single-view policies fail, and the wrist view is what tells the model when the gripper should close.
Day five is the one that needs a tool
The desktop client records LeRobot-format datasets straight out of a teleoperation session: episodes, camera streams and joint states, with the recording controls where you can reach them mid-take. It is the part of week one where the right software saves real hours.
Get the desktop clientSources
- TheRobotStudio/SO-ARM100: SO-101 bill of materials, STS3215 torque and voltage, print settings and gauges
- SO-ARM100: the SO-100 documentation, marked deprecated in the repository README (twelve-servo bill of materials)
- LeRobot docs: SO-100 assembly, gear removal and calibration
- LeRobot docs: SO-101 assembly, motor setup and calibration
- LeRobot docs: Installation, Python and extras
- LeRobot docs: Imitation Learning on Real-World Robots
- LeRobot docs: Compute and hardware guide, VRAM and wall-clock
- LeRobot docs: Cameras
- LeRobot Community Datasets: The "ImageNet" of Robotics - When and How? (dataset quality checklist)
- LeRobot v0.6.1 release, 3 August 2026
- lerobot v0.6.1: DatasetRecordConfig defaults (fps 30, episode_time_s 60, reset_time_s 60, num_episodes 50, push_to_hub true, streaming_encoding false, no_stamp)
- lerobot v0.6.1: ACTConfig defaults (chunk_size 100, n_action_steps 100)
- lerobot v0.6.1: TrainPipelineConfig defaults (seed 1000, batch_size 8, steps 100000, eval_steps 0)
- Zhao et al., Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)
- lerobot v0.6.1: lerobot-rollout, the five strategies and the sync / rtc inference backends
Ready for high-quality robotics data?
AY-Robots connects your robots to skilled operators worldwide.
Get Started