
Real numbers for a working manipulation setup: a $229.88 SO-101 bill of materials, what cameras and a host machine cost, GPU rental at 1 to 12 USD a run, and where small budgets leak first.
A manipulation lab is three budgets pretending to be one. There is the hardware budget, which people obsess over. There is the data budget, which is mostly your own hours and which nobody costs. And there is the compute budget, which is the one that is genuinely cheap now and that people still overpay for by buying a graphics card. This article puts real numbers on all three, taken from published bills of materials, vendor price pages and papers that were fetched while writing it, and it says where the money leaks first.
The short version: a working single-arm setup that can record a LeRobot dataset and run a trained policy costs a few hundred euros in parts and single-digit dollars per training run. The SO-100 class of arm is what made that true. Everything above that number is a choice you should be able to justify, not a default.
What you need to know
- •The upstream SO-101 bill of materials totals $229.88 in the US and 226.30 EUR in the EU for a leader and follower pair, before 3D printing. A follower-only build is $121.94 / 124.30 EUR.
- •AY-Robots lists SO-100 parts at about 110 to 150 EUR and SO-101 at about 130 to 170 EUR. Koch v1.1 runs 250 to 350 EUR, LeKiwi 400 to 500 EUR.
- •Research-grade reference points for scale: ALOHA states under $20k for the whole bimanual system, Mobile ALOHA states $32k including onboard power and compute.
- •Do not buy a training GPU. On AY-Robots a full fine-tuning run costs about 4 to 12 USD on the A100/H100 tier and about 1 to 3 USD on the 24 GB tier.
- •SmolVLA is the cheapest complete path: 30 episodes minimum instead of 50, and a 24 GB card instead of an 80 GB one.
- •The first thing that actually breaks a small budget is cameras: too many, too high a resolution, on one USB controller.
- •Feeding 12 V into 7.4 V STS3215 servos destroys them. That is the single most expensive five-second mistake in this hobby.
The three budgets, separated
Before any shopping, split the spend. The reason is that the three budgets scale completely differently. Hardware is a one-off with a long tail of breakage. Data is a recurring cost measured in hours of teleoperation. Compute is metered by the minute and is the only one you can turn off entirely between projects.
| Budget | What it is | How it scales | Typical share on a small setup |
|---|---|---|---|
| Hardware | Arm, servos, control board, power supply, cameras, host machine | One-off, plus replacement servos | Most of the money, once |
| Data | Your hours driving the arm and resetting the scene | Linear in episodes, and it never stops | Most of the calendar time |
| Compute | Rented GPU hours for fine-tuning and inference pods | Per run, per minute, zero when idle | Single-digit dollars per run |
It helps to anchor against what a funded lab spends. The ALOHA paper states the whole bimanual system costs under $20k, with the ViperX 6-DoF follower arm at around $5600 off the shelf and the smaller WidowX leader at $3300. Mobile ALOHA puts the mobile version at $32k including onboard power and compute. The DROID platform, which 13 institutions used to collect 76k trajectories over 350 hours across 564 scenes, is built on a Franka Panda 7-DoF arm with two Zed 2 stereo cameras, a wrist-mounted Zed Mini and an Oculus Quest 2 for teleop. Those are the numbers you are not paying.
A 3D-printed SO-101 pair replaces the $8900 of ALOHA arm hardware with a bill of materials of $229.88. It is a worse arm. It is also under 3 percent of the price, and for learning imitation learning on tabletop pick-and-place it is enough. That trade is the whole premise of a small-budget lab.
The arm: what 230 dollars actually buys
The SO-ARM100 repository publishes a per-region bill of materials. Below is the US column for the two-arm leader-follower setup, which is the configuration you want if you plan to record data rather than just wiggle a servo. Prices are as published in that README on 2026-08-23; the repo is maintained and the numbers move.
| Part | Qty | Unit (US) | Unit (EU) | Note |
|---|---|---|---|---|
| STS3215 servo 7.4 V, 1/345 gear (C001) | 7 | $13.89 | 12.20 EUR | 6 for the follower, 1 for the leader shoulder lift |
| STS3215 servo 7.4 V, 1/191 gear (C044) | 2 | $13.89 | 12.20 EUR | Leader shoulder pan and elbow flex |
| STS3215 servo 7.4 V, 1/147 gear (C046) | 3 | $13.89 | 12.20 EUR | Leader wrist flex, wrist roll, gripper |
| Motor control board (Waveshare) | 2 | $10.60 | 11.40 EUR | One per arm |
| Power supply | 2 | $10.00 | 15.70 EUR | 5 V for the 7.4 V servo variant |
| USB-C cable, 2 pcs | 1 | $7.00 | 7.00 EUR | Host to control board |
| Table clamps, 4 pcs | 1 | $9.00 | 9.70 EUR | Both arms have to be bolted down |
| Screwdriver set | 1 | $6.00 | 9.00 EUR | Phillips #0 and #1 is the actual requirement |
| Total | $229.88 | 226.30 EUR | Follower only: $121.94 / 124.30 EUR |
Twelve servos at $13.89 is $166.68 of that $229.88. The arm is essentially a pile of Feetech bus servos with plastic between them, and that is also its failure mode: when something breaks, it is a servo, and a servo is 14 dollars plus a disassembly evening. Budget for two spares from day one. That is the cheapest insurance in the build.
STS3215 servos ship in a 7.4 V variant and a 12 V variant, and they look identical. The SO-ARM100 README notes the 7.4 V version has a stall torque of 16.5 kg.cm at 6 V while the 12 V version reaches 30 kg.cm, and that choosing the 12 V motors means buying a 12 V 5A+ supply instead of a 5 V one. Putting that 12 V supply on 7.4 V servos destroys them, usually several at once because they are daisy-chained on one bus. The SO-100, SO-101 and the arm side of LeKiwi all run 7.4 V. Label the power brick. If an arm has gone quiet after a swap, start at servo not responding.
The four arms AY-Robots supports sit at different points on that curve. The parts costs below are the platform's own figures; you can put any two side by side on the arm comparison pages.
| Arm | Servos | Voltage | Parts cost | Support level |
|---|---|---|---|---|
| SO-100 | Feetech STS3215 bus servos | 7.4 V | ~110 to 150 EUR | Full, reference arm |
| SO-101 | Feetech STS3215 | 7.4 V | ~130 to 170 EUR | Full |
| Koch v1.1 | Dynamixel XL330 / XL430 | 5 V and 12 V rails | ~250 to 350 EUR | Compatible |
| LeKiwi | Feetech STS3215 (arm) | 7.4 V arm, 12 V base | ~400 to 500 EUR | Compatible |
Cross-check those against upstream and you get a useful lesson about BOM totals. The Koch v1.1 repository lists $199 for the leader and $278 for the follower in US pricing, so $477 for the pair, against the platform's 250 to 350 EUR for the same arm. The LeKiwi BOM lists $482 US for the 12 V build and $499 for the 5 V build, and states explicitly that it excludes the cost of 3D printing. Region, vendor and what is counted move these totals by a third or more. Treat any published BOM as an order of magnitude, not a quote.
If you have no printer, that is not a blocker and it is also not free. The SO-ARM100 README lists tested printers including the Prusa MINI+, Creality Ender 3 and Bambu Lab machines, with PLA+ at a 0.4 mm nozzle, 0.2 mm layer height and 15 percent infill. It also ships print gauges you check against a real STS3215 servo or a standard Lego brick before committing to a full plate. Vendors selling pre-printed frame kits and fully assembled arms are listed in the same README. Paying to skip a week of tolerance chasing is a defensible line item on a small budget; buying a printer to make one arm usually is not.
Cameras: the line item where money leaks first
This is where small budgets die. Cameras look cheap individually, they are the part people over-buy, and the failure is not financial at first. It is a recording session that silently drops frames and produces a dataset that trains a policy that does not work.
LeRobot's camera layer is deliberately boring. Per the camera documentation, OpenCVCamera covers phones, built-in laptop cameras and USB webcams, RealSenseCamera adds Intel RealSense with depth, and there are ZMQCamera and Reachy2Camera for networked and robot-native streams. Auto-discovery works for the first two:
lerobot-find-cameras opencv # or: lerobot-find-cameras realsense
# --- Detected Cameras ---
# Camera #0:
# Name: OpenCV Camera @ 0
# Type: OpenCV
# Id: 0
# Backend api: AVFOUNDATION
# Default stream profile:
# Format: 16.0
# Width: 1920
# Height: 1080
# Fps: 15.0USB 2.0 Hi-Speed provides roughly 196 Mbps for isochronous transfers, and UVC webcams reserve bandwidth by declaring what they think they need rather than what they use. The Good Penguin's writeup documents a camera declaring 3060 bytes per frame, about 195 Mbps, for a 320x240 MJPEG stream at 25 fps that genuinely needs around 46 Mbps. Two of those cannot coexist on one bus. What you see is No space left on device (ENOSPC) on the second camera, or the DWC2 variant Insufficient periodic bandwidth for periodic transfer. It is not a disk error and it is not your code. Fixes, cheapest first: drop resolution and fps, force MJPEG rather than raw YUYV, move the second camera to a different physical USB controller (not a hub hanging off the same one), or override the declared value with the uvcvideo.bandwidth_cap module parameter. Start at camera not detected.
The practical consequence for a budget: two 640x480 webcams at 30 fps on separate controllers beat one 1080p camera, and beat three cameras that fight. Note that the real-robot tutorial does not start you there. Its teleoperation example runs a single camera named front at 640x480 and 30 fps, its record example raises that same single camera to 1920x1080, and only the rollout example drives two streams at once, an OpenCV camera and a RealSense at 640x480. One 1080p stream is precisely the configuration that stops working the moment you add a second camera to the same controller. Two angles is the standard configuration for tabletop manipulation because the end effector needs a close view and the scene needs a wide one.
- The five trainable policies on this platform take RGB frames; none of them requires a depth stream.
- A 640x480 MJPEG stream is a fraction of the USB bandwidth of an uncompressed 1080p one, so you can actually run two.
- Replacing one that dies costs 20 to 40 EUR, not 300.
- More angles beat better angles: two viewpoints on a small budget usually outperform one expensive sensor.
- Depth genuinely helps for bin picking, clutter, and transparent or reflective objects where RGB is ambiguous.
- LeRobot stores depth as its own video stream, quantized to 12-bit and lossless by default, so the data path is supported rather than bolted on.
- RealSense on macOS is documented as unstable and can need sudo just to enumerate.
- It is a real cost multiplier: a RealSense usually costs more than the entire arm it is looking at.
The host machine, and why it is not a GPU
The computer next to the arm has one job during recording: pull two camera streams and a serial bus at a steady 30 Hz and write them to disk without stalling. That is an I/O problem, not a compute problem. A laptop you already own is the correct answer for most people, and the honest advice is to spend nothing here until something actually stalls.
| Host option | Good for | Will not do | Verified reference |
|---|---|---|---|
| Laptop you already own | Recording, teleop, ACT training on Apple silicon | Fine-tuning a 3 B VLA | lerobot-train accepts policy.device=mps |
| Raspberry Pi 5 | On-robot host, LeKiwi base controller | Any VLA inference at a useful rate | 16 GB launched at $120 in January 2025 |
| Jetson Orin Nano Super dev kit | Local inference next to the servos for small policies | Fine-tuning; 8 GB is below the 16 GB GR00T inference floor | $249, 67 sparse TOPS, 8 GB LPDDR5, 102 GB/s, 7 to 25 W |
| Rented cloud GPU | All fine-tuning, and remote inference for slow tasks | Fast reactive control over the public internet | See the compute table below |
The 16 GB Raspberry Pi 5 launched at $120 in January 2025. On 2 February 2026 Raspberry Pi announced further price rises: +$10 on 2 GB, +$15 on 4 GB, +$30 on 8 GB and +$60 on 16 GB, because the cost of some parts has more than doubled in a quarter on the back of memory demand from the AI infrastructure buildout. The same force that made renting an A100 cheap made your single-board computer expensive. Re-price the BOM before you order, not from a blog post.
The Jetson number is worth stating precisely because it is the honest ceiling for local inference. NVIDIA's Orin Nano Super announcement puts the dev kit at $249, down from $499, with 67 sparse TOPS, 33 dense TOPS, 8 GB of LPDDR5 at 102 GB/s in a 7 to 25 W envelope. Meanwhile the Isaac-GR00T repository lists inference as needing one GPU with 16 GB or more of VRAM and fine-tuning as wanting 40 GB or more, recommending H100 or L40 nodes. An 8 GB Orin Nano is under the inference floor for a 3 B vision-language-action model. It is a fine host for ACT, which is about 80 M parameters at 20 ms per action step.
Compute: rent it, and do not round up
This is the part of the budget that got cheap while everyone was not looking. Fine-tuning a VLA on your own data is a few hours of one GPU. Buying that GPU costs more than the arm, the cameras, the host and every spare servo combined, and it will idle 95 percent of the time.
Here are per-hour prices fetched on 2026-08-24 from two vendors that publish them. Runpod's pricing page carries a Community Cloud / Secure Cloud toggle and a per-second view; the rates below are the ones it serves by default. Both marketplaces move, so treat these as an order of magnitude and not a quote.
| GPU | Runpod, per hour | Hugging Face Jobs |
|---|---|---|
| RTX 4090 24 GB | $0.74/hr | not offered |
| RTX A6000 48 GB | $0.53/hr | not offered |
| L40S 48 GB | $0.99/hr | $1.80/hr (1x L40S) |
| A10G 24 GB | not listed | $1.00/hr (a10g-small) |
| A100 PCIe 80 GB | $1.39/hr | $2.50/hr (A100-large) |
| A100 SXM 80 GB | $1.59/hr | $2.50/hr (A100-large) |
| H100 PCIe 80 GB | $2.89/hr | not offered |
| H100 SXM 80 GB | $3.29/hr | not offered |
| H200 141 GB | $4.59/hr | $5.00/hr |
| Nvidia L4 24 GB | $0.49/hr | $0.80/hr (1x L4) |
| Nvidia T4 16 GB | not listed | $0.40/hr (T4-small) |
Two structural details matter more than the headline rate. First, billing granularity: Hugging Face Jobs bills by the minute and only while a job is Starting or Running, with no charge during build, and its default job timeout is 30 minutes unless you raise it. Vast.ai bills per second and offers an interruptible tier it describes as 50 percent or more cheaper, preemptible, and suited to checkpointed batch training. Second, idle time: an instance you forget about is the real budget risk, not the hourly rate.
The platform rents a GPU on a spot market sized by the VRAM the chosen model needs, so the price is a range rather than a number. On the A100 80 GB / H100 tier (GR00T N1.7, GR00T N1.5, Pi0.5) a run takes 3 to 6 hours at 1.20 to 2.00 USD per hour, so about 4 to 12 USD. On the RTX 4090 / 24 GB tier (SmolVLA, ACT) it is 2 to 5 hours at 0.30 to 0.60 USD per hour, so about 1 to 3 USD. Inference pods carry an idle watchdog and destroy themselves after an idle period, which is the part that stops a forgotten tab becoming a bill. Details on pricing and in the billing docs.

For scale on what you are not paying: the SmolVLA paper reports that the project consumed approximately 30,000 GPU hours, pretraining for 200,000 steps at a global batch size of 256 across 4 GPUs, on 481 community datasets totalling 22.9K episodes and 10.6M frames. At the 4090 rate above that pretraining bill alone would run into five figures. You are renting the last sliver of that curve, the fine-tuning on your own 30 to 50 episodes.
- A full run is 1 to 12 USD depending on tier, so a card you buy only pays for itself after hundreds of runs.
- You can use an 80 GB A100 for a 3 B model without owning one, which no consumer card gives you.
- Zero cost between projects, and no power, noise or heat in the room with the robot.
- Spot and interruptible tiers exist precisely because training is checkpointable.
- No egress or per-hour charges, and no dependency on marketplace availability.
- Local iteration on ACT and SmolVLA is genuinely faster when you are debugging a dataset rather than training seriously.
- A card you own can also serve inference next to the robot, which is where latency actually matters.
- If you already own the card for other reasons, the marginal cost of a run is electricity.
Costing one real project, end to end
Here is the actual sequence for a first pick-and-place project on an SO-101, with the money and the hours marked. Commands are from the LeRobot SO-101 guide and the imitation learning tutorial as read on 2026-08-23.
- 11. Parts and printing: about 230 USD, one week
Order the two-arm BOM, print the frames or buy a printed kit. Install LeRobot and the Feetech SDK. The install is the only software cost in the whole project.
bashpip install -e ".[feetech]" - 22. Ports, motor ids and calibration: 0 USD, one evening
Motors usually ship with a default id of 1, so ids and baudrate are written to EEPROM one motor at a time. Then calibrate both arms so leader and follower agree on joint positions. Skipping calibration is the classic way to burn a whole dataset.
bashlerobot-find-port lerobot-setup-motors \ --robot.type=so101_follower \ --robot.port=/dev/tty.usbmodem585A0076841 lerobot-calibrate \ --robot.type=so101_follower \ --robot.port=/dev/tty.usbmodem58760431551 \ --robot.id=my_awesome_follower_arm - 33. Cameras: about 25 to 60 USD, one hour, or one lost afternoon
Two webcams at 640x480 and 30 fps, on separate USB controllers. Verify with a teleop session before recording anything, because a stream that fails here fails silently later.
bashlerobot-find-cameras opencv lerobot-teleoperate \ --robot.type=so101_follower \ --robot.port=/dev/tty.usbmodem5AB90687491 \ --robot.id=my_follower_arm \ --robot.cameras="{front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, wrist: {type: opencv, index_or_path: 2, width: 640, height: 480, fps: 30}}" \ --teleop.type=so101_leader \ --teleop.port=/dev/tty.usbmodem5AB90689011 \ --teleop.id=my_leader_arm \ --display_data=true - 44. Record the dataset: 0 USD, two to four hours of your life
This is the real cost. LeRobot defaults are 50 episodes at episode_time_s=60 with reset_time_s=60, which is 100 minutes of wall clock before any re-records. The tutorial's own advice is at least 50 episodes with 10 per grasp location, cameras fixed, behaviour consistent.
bashlerobot-record \ --robot.type=so101_follower \ --robot.port=/dev/tty.usbmodem585A0076841 \ --robot.id=my_awesome_follower_arm \ --robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, wrist: {type: opencv, index_or_path: 2, width: 640, height: 480, fps: 30}}" \ --teleop.type=so101_leader \ --teleop.port=/dev/tty.usbmodem58760431551 \ --teleop.id=my_awesome_leader_arm \ --dataset.repo_id=${HF_USER}/record-test \ --dataset.num_episodes=50 \ --dataset.single_task="Grab the black cube" - 55. Train: 1 to 12 USD, 2 to 6 hours, unattended
Locally on a card you own, on Hugging Face Jobs by adding --job.target, or on AY-Robots by picking model and dataset in a form. All three run the same trainers.
bash# local lerobot-train \ --dataset.repo_id=${HF_USER}/so101_test \ --policy.type=act \ --output_dir=outputs/train/act_so101_test \ --policy.device=cuda # same command, rented hardware, billed by the minute lerobot-train \ --dataset.repo_id=${HF_USER}/so101_test \ --policy.type=act \ --policy.repo_id=${HF_USER}/my_policy \ --job.target=a10g-small hf jobs hardware # list flavors and prices hf jobs cancel <job-id> - 66. Run it back: cents, or a pod hour
Roll the policy out on the arm. On AY-Robots an inference pod is auto-provisioned and destroys itself when idle. Note that ACT has no base model, so this checkpoint only exists because you trained it.
bashlerobot-rollout \ --strategy.type=base \ --policy.path=${HF_USER}/my_policy \ --robot.type=so100_follower \ --robot.port=/dev/ttyACM1 \ --task="Put lego brick into the transparent box" \ --duration=60
Add it up: roughly 230 to 320 USD of hardware, about 12 hours of your time, and under 15 USD of compute for several training runs. The hardware is a one-off. The 12 hours is not, and that is the number people get wrong.
Two ways to get there
Source the parts, print the frames, install LeRobot, record locally, then rent a GPU yourself and wire the trainer to your dataset. Everything in this article is reproducible this way, and for ACT and SmolVLA you may not need to rent anything at all. Worth reading once before you record: the video encoding parameters page, because the defaults decide your disk footprint and your training throughput.
- You pay: the BOM at cost, plus GPU hours at marketplace rates (Runpod listed a 4090 at $0.74/hr and an A100 PCIe 80 GB at $1.39/hr on 2026-08-24).
- You own: every failure mode. Camera enumeration, USB bandwidth, EEPROM ids, dataset version mismatches, and the instance you forgot to stop.
- You get: total control, and no dependency on anyone's uptime.
- Hidden costs: your hours, and the fact that the GR00T fine-tuning entry point (launch_finetune.py, a tyro CLI) exposes no seed, so those runs are not bit-for-bit reproducible. lerobot's default seed is 1000.
A LeRobot v3.0 dataset crashes the GR00T loader. It has to be converted down to v2.1 first. GR00T N1.7 and N1.5 want v2.0 or v2.1; Pi0.5, SmolVLA and ACT want v3.0. Recording once and assuming it feeds every trainer is wrong, and you find out three hours into a rented pod. See dataset rejected as v3.
# datasets land here by default
ls ~/.cache/huggingface/lerobot/${HF_USER}/record-test
# LeRobot video encoding defaults, worth knowing before you blame the codec:
# vcodec = libsvtav1
# pix_fmt = yuv420p
# crf = 30
# g = 2 (keyframe every 2 frames)
# preset = 12Same trainers, same dataset format, someone else operating the GPU market. You pick a model and a dataset in a form on the training page, the backend rents a GPU sized by the VRAM the model needs, runs the trainer and writes checkpoints to object storage. The desktop client from the download page records LeRobot datasets straight from a teleop session.
| Policy | Params | GPU tier | Min episodes | Dataset format | Cost per run |
|---|---|---|---|---|---|
| GR00T N1.7 | ~3 B (~40 M trained) | A100 80 GB / H100 | 50 | v2.0 or v2.1 | 4 to 12 USD |
| GR00T N1.5 | ~3 B | A100 80 GB / H100 | 50 | v2.0 or v2.1 | 4 to 12 USD |
| Pi0.5 | ~3 B, PaliGemma backbone | A100 80 GB / H100 | 50 | v3.0 | 4 to 12 USD |
| SmolVLA | ~450 M | RTX 4090 / 24 GB | 30 | v3.0 | 1 to 3 USD |
| ACT | ~80 M | RTX 4090 / 24 GB | 50 | v3.0 | 1 to 3 USD |
- Cheapest complete path: SmolVLA on an SO-100. 30 episodes instead of 50 and a 24 GB card instead of an 80 GB one, which cuts both the data hours and the compute bill.
- GR00T N1.7, GR00T N1.5 and Pi0.5 are cloud-only here. SmolVLA and ACT also run locally, so you can debug a dataset without paying anything.
- Inference pods carry an idle watchdog and destroy themselves, so a forgotten tab does not keep billing.
- Base checkpoints are the vendors' own: nvidia/GR00T-N1.7-3B, nvidia/GR00T-N1.5-3B, lerobot/pi05_base. ACT has no base model at all.
If you do not have an arm yet, the live queue puts a physical SO-100 in your browser with no signup, and the try page lays out the three zero-hardware entry points. The same operations are available from the CLI and from the MCP server.
Where the money is wasted first
Ranked by how often it happens, not by size. Every item here is a thing that looks like an investment and behaves like a tax.
- A training GPU. At 1 to 3 USD a run on the 24 GB tier, which is the tier a consumer card actually competes with, buying one only pays off after hundreds of runs, and it still cannot hold a 3 B model comfortably. Buy the card for inference next to the robot if you buy it at all.
- Too many cameras. Three streams contending for USB bandwidth produce worse data than two that do not, and the failure is silent. Two viewpoints is the working default.
- A depth camera before you need depth. None of the five trainable policies here requires a depth stream. Buy it when RGB is provably ambiguous for your task, not before.
- 1080p everything. The LeRobot record example does use 1920x1080, but on a single camera, and that is the catch: the resolution the tutorial gets away with on one stream is the one that breaks you on two. Higher resolution costs USB bandwidth, disk, encode CPU and training throughput, and buys very little for tabletop manipulation.
- A second arm before the first task works. A leader-follower pair is not two robots, it is one recording rig. A second follower is a scaling decision, not a starting one.
- A 3D printer for one build. Printed frame kits and assembled arms are sold by the vendors listed in the SO-ARM100 README. A printer plus its learning curve is not cheaper for a single arm.
- Re-recording a dataset because calibration drifted. Free to avoid, expensive to fix. Calibrate, verify with lerobot-replay, then record.
- Idle cloud instances. The one genuinely unbounded cost in this list. Set a timeout, or use a platform whose pods kill themselves.
There is a matching list of things worth paying for: two spare servos, decent table clamps, a second USB controller (a PCIe card or a different port group, not a hub), and enough light on the workspace that the cameras are not fighting your ceiling lamp. All four together are a small fraction of the arm, and all four prevent a class of failure. If a policy is misbehaving and you cannot tell why, a policy that only works in one setup and loss falls but the policy does nothing are the two pages that usually explain it.

The parts that cost nothing
A surprising amount of a manipulation project can be done for zero. Not as a marketing claim, but because the expensive parts are the ones you can defer.
- Drive a real arm before buying one. The live page streams a physical SO-100 with no signup, queue-based. It answers 'do I actually want to do this' for free.
- Read other people's data. The dataset directory lists public datasets, and datasets can also come straight from a Hugging Face repo id or from your own machine.
- Compare models with published numbers rather than vibes. The arena holds 85 VLA models with 332 benchmark results, each value linked to its paper or model card.
- Learn the pipeline on data you did not record: train your first policy walks the whole flow, and a dataset can come straight from a Hugging Face repo id.
- Get paid instead of paying. The operator page is teleoperation work for people who want the hours without the hardware.

Where a small budget genuinely does not work
Being honest about this is more useful than another table. There are three places where the cheap setup is not a smaller version of the expensive one, it is a different thing.
The first is precision. The ALOHA paper specs its ViperX follower at 5 to 8 mm accuracy with 1 mm repeatability, on an arm that costs around $5600. A printed arm on plastic gearing with position-controlled hobby servos is looser than that, and it drifts with temperature and wear. Tasks that need repeatable sub-millimetre placement are out, and no amount of training data fixes a mechanical tolerance.
Inference has to sit next to the servos for fast tasks. The control loop on this platform is 20 ms per action step for ACT, 152 ms for GR00T N1.7, 165 ms for GR00T N1.5, 245 ms for SmolVLA and 485 ms for Pi0.5. Adding public-internet round trips on top turns a working policy into a hesitant one. Renting a cloud pod for inference is viable for slow pick-and-place; it is not viable for fast reactive motion. If that is your task, the money has to go into local compute, and the cheap remote path does not substitute. More on inference latency.
The second is throughput. One arm and one operator is about 50 episodes an evening at the documented 60 second episode and 60 second reset. DROID needed 13 institutions and 50 collectors over 12 months for its 76k trajectories. If your research question needs that scale, the bottleneck is people, and no hardware saving addresses it. The third is that a cheap arm makes debugging harder, not easier: when a policy fails you now have to rule out a slipping horn and a drifting calibration before you look at the model. That is a real tax on your time, paid to save several thousand euros, and for most people it is still worth it. Action chunking helps hide some of the jitter, but it does not add mechanical repeatability that the arm never had.
What is the absolute minimum to record a dataset and train a policy?▾
A follower-only SO-101 build at $121.94 / 124.30 EUR per the upstream BOM, one USB webcam, a laptop you already own, and about 1 to 3 USD for a SmolVLA or ACT run on the 24 GB tier. Without a leader arm you teleoperate with an on-screen pad rather than leader-follower, which is slower to record with but works. Adding the leader takes the parts total to $229.88 / 226.30 EUR and is the upgrade most people should make first.
Do I need to buy a GPU?▾
No, and for fine-tuning you should not. A full run on AY-Robots is about 4 to 12 USD on the A100 80 GB / H100 tier and about 1 to 3 USD on the 24 GB tier. Runpod listed an A100 PCIe 80 GB at $1.39/hr and a 4090 at $0.74/hr on 2026-08-24. The only good reason to own a card is local inference next to the arm, where latency is the constraint rather than cost.
Which of the five models is cheapest to actually get working?▾
SmolVLA. It needs 30 episodes rather than 50, which is about 40 minutes less recording at the documented 60 second episode and 60 second reset, and it trains on a 24 GB card rather than an 80 GB one, which is the difference between about 1 to 3 USD and 4 to 12 USD per run. ACT is cheaper per step at 20 ms inference but wants 50 episodes and has no base model, so it only exists after you have trained it on your own task.
How much does a camera actually need to cost?▾
Two ordinary USB webcams running 640x480 at 30 fps is the working configuration in the LeRobot tutorials. The constraint is not image quality, it is USB bandwidth: USB 2.0 Hi-Speed gives about 196 Mbps for isochronous transfers and UVC cameras reserve what they declare rather than what they use, so two cameras on one controller can fail with 'No space left on device'. Spend the money on a second USB controller before you spend it on a better sensor.
Is a Raspberry Pi or Jetson enough to run the policy on the robot?▾
For ACT, yes. For a 3 B VLA, no. The Isaac-GR00T repository lists 16 GB or more of VRAM for inference, and the Jetson Orin Nano Super dev kit has 8 GB of LPDDR5. It is a good host for recording and for small policies at $249 with 67 sparse TOPS, but it is under the floor for GR00T-class models. Note also that Raspberry Pi raised prices on 2 February 2026 by up to $60 on the 16 GB board because of memory costs, so re-check current pricing before budgeting.
Why do published BOM totals disagree with each other?▾
Region, vendor and what is counted. The Koch v1.1 repo lists $199 leader plus $278 follower in US pricing while AY-Robots lists 250 to 350 EUR for the same arm. The LeKiwi BOM totals $482 US for the 12 V build and states explicitly that it excludes 3D printing. Shipping, customs and whether you already own a screwdriver set move a 230 dollar build by 30 percent. Treat every BOM as an order of magnitude and re-price before ordering.
See what a run actually costs before you commit
Per-model GPU tiers, run times and price ranges, plus how inference pods bill and when they shut themselves down. Not an estimate, the same numbers the trainer uses.
Open the pricing pageWhere to go next
If you are buying hardware, start with the complete SO-100 setup guide and the SO-100 versus SO-101 comparison, since the newer arm is the better build for a few euros more. If you are about to record, read how to collect high-quality VLA training data first; the hours you save there are worth more than anything in this article. And if scale is what you are actually asking about, the DROID dataset writeup shows what 13 institutions and 12 months buys that one arm on a desk cannot.
For the mechanical side, SO-100 getting started and record your first dataset are the two pages that turn a BOM into a working rig. The training docs cover what the trainer actually sends, the policies page compares the five models side by side, and the failure-mode index is the page to bookmark, because on a small budget most of your losses are hours, not euros.
Sources
- TheRobotStudio/SO-ARM100: Standard Open Arm, bill of materials and print settings
- LeRobot docs: SO-101 assembly, motor setup and calibration
- LeRobot docs: imitation learning on real-world robots (record, train, rollout)
- LeRobot docs: cameras, lerobot-find-cameras and camera classes
- Hugging Face Jobs: hardware flavors, hourly pricing and per-minute billing
- Runpod GPU pricing: per-GPU hourly rates for Pods
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ALOHA)
- Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- NVIDIA Isaac-GR00T: hardware requirements and fine-tuning entry point
- Koch v1.1 low-cost robot arm: bill of materials
- LeKiwi bill of materials
- NVIDIA Jetson Orin Nano Super Developer Kit: $249, 67 TOPS
- Raspberry Pi: more memory-driven price rises (2 February 2026)
- Multiple UVC cameras on Linux: USB bandwidth and ENOSPC
Ready for high-quality robotics data?
AY-Robots connects your robots to skilled operators worldwide.
Get Started