The AY-Robots operator page reading Become a Robot Operator from anywhere in the world, with a photo of an SO-100 robot arm
robot data collectionteleoperation jobsLeRobot datasetsrobot training data marketSO-100

Robot Data Collection as a Business: Pay, Buyers, Reality

AY-Robots ResearchAugust 23, 202619 min read

What recording robot demonstrations for other people actually pays in 2026, who buys the data, how many episodes per hour you can really deliver, and where the business breaks.

Somebody is paying for robot demonstrations. In April 2026 MIT Technology Review followed gig workers in Nigeria and India who strap an iPhone to their forehead and film themselves folding laundry for micro1, a Palo Alto company that resells the footage to robotics labs. Its chief executive estimated that robotics companies spend more than 100 million USD a year buying real-world data from micro1 and firms like it. In June 2026 TechCrunch covered XDOF, a data company with roughly 60 employees and 20 customers that had raised 70 million USD. In August 2026 TechCrunch reported micro1 going from a 100 million USD gross annual run rate to 500 million USD in eight months, though that figure covers its whole AI training data business and not robot data alone. This is a real market with real invoices in it.

What you need to know before you buy an arm to sell data

  • Published operator pay in 2026 runs from about 15 USD per hour for at-home egocentric video (micro1, per MIT Technology Review) to 30 to 55 USD per hour for on-site robot teleoperation (an OpenTrain AI listing).
  • The published academic record is brutal about throughput: DROID produced 350 hours of interaction from 50 collectors over 12 months, which is about 7 recorded hours per person per year.
  • Buyers do not pay for episode count. The Data Scaling Laws paper found 32 environments with one unique object each and 50 demonstrations per pair reached about 90 percent success on unseen scenes; more demos in the same room bought almost nothing.
  • Supply is enormous and mostly free. Hugging Face listed more than 72,400 datasets tagged LeRobot on 24 August 2026, and the count climbs daily. Uploading another pick-and-place set is not a business.
  • The paid work is contract work with a spec: a named task, a fixed camera rig, a documented variation plan. Proof that your data trains is the actual product.
  • AY-Robots gives you the recording client, the dataset directory and the training loop to prove a dataset works. It does not run a marketplace and will not find you a buyer.

Who is buying, and what they are buying

The buyer side splits into three groups that pay differently. Frontier robotics labs want data on the exact robot they are shipping, and mostly build that in house or buy it as a managed program. Data vendors sit between those labs and the people holding cameras: Scale AI wrote in September 2025 that it had passed 100,000 production hours at its San Francisco prototyping laboratory and with contributors globally, for customers including Physical Intelligence, Generalist AI and Cobot. Then there is the state-funded tier, which barely behaves like a market.

Rest of World reported in January 2026 that more than 40 state-backed robot data collection centres had been announced across China by December, with about two dozen already running, according to the analyst firm Interact Analysis. One camp near Beijing, built by the Shijingshan government with the humanoid company Leju, covers more than 10,000 square metres and offers 16 scripted scenarios. In one Hubei facility close to 100 humanoids are driven through folding, ironing and wiping hundreds of times a day. UBTech sold 566 million yuan, about 80 million USD, of humanoids into three such centres. That is a lot of capacity that does not need to earn its money back this quarter.

Buyer typeWhat they wantPublished evidenceWhat it means for a solo operator
Frontier robot labTeleoperation data on the exact robot they shipXDOF describes teleoperation on the deployed robot as the most valuable tier (TechCrunch, June 2026)Closed to you unless you own their hardware
Data vendor, teleop programsHours on a standard rig, in a controlled site, on shiftScale AI: more than 100,000 production hours (Scale AI blog, Sept 2025)You apply as an operator, they own the rig
Data vendor, egocentric videoFirst-person chore video from many homesmicro1: thousands of contractors in 50+ countries (MIT Technology Review, April 2026)Open to anyone with a phone, priced accordingly
Academic lab or small startupA specific task on a cheap arm, in LeRobot formatDROID pooled 50 collectors across 13 institutions for 76k trajectoriesThis is the realistic contract for an SO-100 owner
State-funded centre (China)Standardised industry-wide data40+ centres announced, about 24 operating (Interact Analysis via Rest of World)Not a customer for a European or US freelancer
The value ladder buyers use

XDOF, covered by TechCrunch in June 2026, sorts robot data into three tiers by value. Top tier: teleoperation recorded on the actual robot being deployed. Middle tier: teleoperated robots collecting more general data, in their case through the low-cost GELLO teleoperation rig. Bottom tier: egocentric video of humans doing tasks with no robot involved. An SO-100 sits in the middle tier. That is not the cheapest rung, and it is the rung where a single person with 150 EUR of hardware can actually compete.

What the work pays in 2026

Published rates are thin because most of this is priced per project behind an NDA, but there are enough real datapoints to draw a range. One rate below comes from named reporting and one from a live job listing; the other two rows are there to mark where no published price exists.

RoleRateSourceNotes
At-home egocentric recorder (micro1)15 USD per hourMIT Technology Review, 1 April 2026Worker in Nigeria; clips reviewed by AI and a human, then accepted or rejected
On-site robot teleoperator (OpenTrain AI listing)30 to 55 USD per hourLive job posting, verified 24 August 2026US, on-site, 24/7 facility, four overlapping shift windows, leader-arm rigs and VR
Contract dataset for a named taskNegotiated per projectNo public price list foundYou are selling a deliverable, not hours
Publishing an open dataset0 USDHugging Face LeRobot tag: 72,400+ datasetsReputation, not revenue

The gap between the top and bottom of that table is not skill, it is capital. The on-site teleoperator earns twice to more than three times the at-home recorder because someone else bought the robot, the cameras, the network and the shift supervisor. If you own an SO-100 you have bought yourself into the middle of that range and taken on the equipment risk. Whether that pays depends entirely on how many recorded minutes you can actually deliver per hour of your life, which is the number almost nobody quotes.

The hours nobody pays for

MIT Technology Review's reporting includes a detail that should decide whether you take this work: a recorder in Delhi said it takes him about an hour to produce a 15-minute video, because he spends the rest of the time inventing new chores to film. That is a four to one ratio of effort to delivered footage. The 15 USD an hour in the same article is quoted for a different worker, in Nigeria, and the reporting never says whether that hour is clock time or accepted footage. That ambiguity is the entire margin. Before signing anything, get in writing whether you are paid for wall-clock time or for accepted output, and what the rejection rate has been for other operators.

The arithmetic of one operator

Take the best documented data collection effort in the open literature and divide. DROID, published at RSS 2024, contains 76,000 demonstration trajectories or 350 hours of interaction, collected across 564 scenes and 84 tasks by 50 data collectors in North America, Asia and Europe over 12 months. That is about 16.6 seconds of robot motion per trajectory, 1,520 trajectories per collector for the year, and roughly 7 hours of actual recorded interaction per collector across an entire year. RoboTurk's follow-up work is the other anchor: 111 hours from 54 remote users on three tasks in one week, which is about two recorded hours per person for the week.

Neither number means those people worked seven hours a year. It means recorded motion is a small fraction of the clock. Everything else is setup, resets, calibration, re-recording the take where the gripper missed, and moving the props. The defaults in LeRobot say the same thing in a config file rather than a paper.

MetricValueWhere it comes from
Default episode length in lerobot-record60 s--dataset.episode_time_s default
Default reset window between episodes60 s--dataset.reset_time_s default
Default episode count per run50--dataset.num_episodes default
Ceiling on episodes per hour at those defaults3060 s + 60 s per episode
Share of wall clock that is recorded dataabout 50 percentReset time equals episode time by default
DROID average trajectory lengthabout 16.6 s350 h divided by 76,000 trajectories
DROID recorded interaction per collector per yearabout 7 h350 h divided by 50 collectors
Scale AI production hours vs DROID plus Open X-Embodiment100,000 h vs about 5,000 hScale AI blog, 24 Sept 2025
bash
# 50 episodes, one task, leader-follower on an SO-101 follower
lerobot-record \
    --robot.type=so101_follower \
    --robot.port=/dev/ttyACM0 \
    --robot.id=my_follower_arm \
    --robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, wrist: {type: opencv, index_or_path: 1, width: 640, height: 480, fps: 30} }" \
    --teleop.type=so101_leader \
    --teleop.port=/dev/ttyACM1 \
    --teleop.id=my_leader_arm \
    --dataset.repo_id=${HF_USER}/pick-cube-kitchen-01 \
    --dataset.num_episodes=50 \
    --dataset.single_task="Grab the black cube" \
    --dataset.episode_time_s=30 \
    --dataset.reset_time_s=15 \
    --display_data=true

# during the run:
#   n or right arrow  -> end this episode early, move on
#   r or left arrow   -> throw this episode away and re-record it
#   q or ESC          -> stop, encode the videos, upload
The command that defines your hourly output. Version: LeRobot as documented on 24 August 2026.

Cutting the episode to 30 seconds and the reset to 15 seconds takes the ceiling from 30 episodes per hour to 80. Whether you can actually hit it depends on how long it takes to put the cube back and how often you press r. Track that ratio for a week before you quote anyone a price. If you are new to this, the record your first dataset walkthrough and the episode glossary entry are the places to start, and the SO-100 data collection page covers the rig side.

The AY-Robots public dataset directory listing recorded LeRobot datasets with their episode counts
The public dataset directory at /directory. Browsing what already exists for your task is the cheapest market research available: if there are twenty free datasets of a cube going into a bin, that is not the thing to sell.

What buyers actually check

The single most useful published result for anyone selling demonstrations is the Data Scaling Laws paper (ICLR 2025). The authors collected more than 40,000 demonstrations and ran over 15,000 real robot rollouts to answer one question: what do you get for another demonstration? The answer is that generalisation follows roughly a power law in the number of environments and objects, and that once you pass about 50 demonstrations per environment-object pair, more demonstrations in that pair do almost nothing. Their recommended recipe is 32 diverse environments, each with its own object and 50 demonstrations, which reached around 90 percent success in unseen environments with unseen objects.

Read that as a pricing statement. A buyer who understands the literature is not paying you per episode, they are paying you per distinct scene-object pair with enough coverage to be usable. Five hundred episodes of the same cube on the same table is worth less than the same five hundred split across ten rooms with ten different objects. That is also the thing you can sell that a data factory cannot easily buy: you live somewhere they do not, and your kitchen is a scene they do not have.

CheckWhat is expectedWhere it is documented
Camera countTwo camera views preferredHugging Face LeRobot dataset guidance
Frame rateAbout 30 fpsHugging Face LeRobot dataset guidance
ResolutionAt least 480x640, 720p preferredHugging Face LeRobot dataset guidance
What may move in frameOnly the follower arm and the manipulated objects; the leader arm must stay out of shotHugging Face LeRobot dataset guidance
Feature namingmodality.location, for example images.top, images.wrist.left, not images.laptopHugging Face LeRobot dataset guidance
Task stringA real description of the goal, roughly 25 to 50 characters, not task1 or demo2Hugging Face LeRobot dataset guidance
Episodes per taskAt least 50, with 10 episodes per locationLeRobot imitation learning tutorial
Diversity targetAround 32 environment-object pairs, 50 demos eachData Scaling Laws, ICLR 2025
Dataset formatLeRobot v2.0 or v2.1 for GR00T, v3.0 for Pi0.5, SmolVLA and ACTAY-Robots policy catalog
The format mismatch that eats a day

A LeRobot v3.0 dataset crashes the GR00T loader. It has to be converted down to v2.1 first. If a buyer says GR00T and you hand over whatever the newest LeRobot release wrote, the job bounces and you will spend a day working out why. Agree the dataset version in writing at the same time you agree the task string. The dataset rejected as v3 fix page covers the symptom, and the dataset docs cover the layout.

The other check nobody writes into a contract but everybody applies: does a policy actually train on it. A buyer who has been burned once will ask for evidence, and the cheapest evidence is a small model trained on your own data with the loss curve and a rollout video attached. ACT at roughly 80 M parameters and SmolVLA at roughly 450 M both run on a 24 GB card, which on the AY-Robots spot tier is 2 to 5 hours at 0.30 to 0.60 USD per hour, so about 1 to 3 USD per run. Three dollars to prove your dataset is trainable is the best money in this whole business.

Recording demonstrations for other people
What works about it
  • Hardware cost is low: an SO-100 is about 110 to 150 EUR in parts, an SO-101 about 130 to 170 EUR.
  • The tooling is open and stable. lerobot-record, the LeRobot dataset format and the Hugging Face Hub cost nothing to use.
  • Verification is cheap. An ACT or SmolVLA run on a 24 GB card is about 1 to 3 USD, so you can prove a dataset trains before invoicing it.
  • Scene diversity is a genuine moat for a distributed operator: your rooms, lighting and objects are not in anyone's warehouse.
  • The skill transfers: calibration, camera placement and failure diagnosis are what the on-site teleop programs hire for.
What you are signing up for
  • Roughly half your wall clock is resets and re-records at the LeRobot defaults, and rejected takes are usually unpaid.
  • There is no clearing price. No public marketplace publishes what a LeRobot episode is worth, so every deal is negotiated cold.
  • Enormous free supply: more than 72,400 LeRobot-tagged datasets on Hugging Face as of 24 August 2026 sets the floor for anything generic.
  • Buyers who want humanoid data do not want SO-100 data. Embodiment is not fungible at the top of the value ladder.
  • Privacy exposure is real. MIT Technology Review found at-home recorders who did not know how their footage would be stored or shared on.
  • The demand is venture funded and could reprice fast. MIT Technology Review put venture money into humanoids at 6.1 billion USD in 2025; that is a cycle, not a floor.

Two routes to the same recorded episode

  1. 1
    Build and calibrate the arm

    Assemble an SO-100 or SO-101 leader-follower pair. Feetech STS3215 servos run at 7.4 V; feeding them 12 V destroys them. Calibration writes a file keyed to the robot id, and you must reuse the same id when recording and evaluating.

    bash
    lerobot-teleoperate \
        --robot.type=so101_follower --robot.port=/dev/ttyACM0 --robot.id=my_follower_arm \
        --teleop.type=so101_leader --teleop.port=/dev/ttyACM1 --teleop.id=my_leader_arm
  2. 2
    Fix the camera rig and never touch it again

    Two views, 30 fps, at least 480x640: one fixed front or overhead camera, one wrist camera. Tape the tripod positions to the table. A camera moved mid-dataset is a silent quality bug you will not see until training.

    bash
    lerobot-teleoperate \
        --robot.type=so101_follower --robot.port=/dev/ttyACM0 --robot.id=my_follower_arm \
        --robot.cameras="{ front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30} }" \
        --teleop.type=so101_leader --teleop.port=/dev/ttyACM1 --teleop.id=my_leader_arm \
        --display_data=true
  3. 3
    Record against a written variation plan

    Decide the environment-object pairs before you start, not while recording. Write the task string once and keep it identical. Press r to discard a bad take immediately rather than cleaning up later.

    bash
    lerobot-record \
        --robot.type=so101_follower --robot.port=/dev/ttyACM0 --robot.id=my_follower_arm \
        --teleop.type=so101_leader --teleop.port=/dev/ttyACM1 --teleop.id=my_leader_arm \
        --dataset.repo_id=${HF_USER}/pick-cube-kitchen-01 \
        --dataset.num_episodes=50 \
        --dataset.single_task="Grab the black cube"
  4. 4
    Prove it trains before you invoice

    Train ACT on your own dataset on whatever GPU you can rent, and record a rollout video. This is the artefact that turns a folder of parquet files into a deliverable.

    bash
    lerobot-train \
      --dataset.repo_id=${HF_USER}/pick-cube-kitchen-01 \
      --policy.type=act \
      --output_dir=outputs/train/act_pick_cube \
      --job_name=act_pick_cube \
      --policy.device=cuda
  5. 5
    Handle the format conversion yourself

    If the buyer trains GR00T, convert the dataset down to LeRobot v2.1 before delivery. If they train Pi0.5, SmolVLA or ACT, v3.0 is what they want. Get this wrong and the delivery bounces.

    bash
    # confirm what you actually wrote before you ship it
    python -c "import json,os,pathlib; \
    p=pathlib.Path.home()/'.cache/huggingface/lerobot'/os.environ['HF_USER']/'pick-cube-kitchen-01/meta/info.json'; \
    print(json.load(open(p)).get('codebase_version'))"
What this route costs you

Parts, your time, and the whole of the commercial problem: finding the buyer, agreeing the spec, and absorbing rejected takes. Nothing here is hard, but all of it is on you.

Pricing a job without a market price

There is no published clearing price for a LeRobot episode, so you build one from the two numbers you control: your measured episodes per hour, and the rate you would accept for that hour doing anything else. What follows is arithmetic on your own inputs, not a market quote.

  1. 1
    Measure your real throughput for one week

    Not the theoretical ceiling. Accepted episodes divided by hours at the desk, setup and re-records included. At the LeRobot defaults the ceiling is 30 per hour; most people land well under it.

    bash
    # count the episodes actually saved
    ls -R ~/.cache/huggingface/lerobot/${HF_USER}/pick-cube-kitchen-01/videos | wc -l
    # and read the episode count the dataset itself reports
    python -c "import json,os,pathlib; \
    p=pathlib.Path.home()/'.cache/huggingface/lerobot'/os.environ['HF_USER']/'pick-cube-kitchen-01/meta/info.json'; \
    i=json.load(open(p)); print(i['total_episodes'], 'episodes,', i['total_frames'], 'frames,', i['fps'], 'fps')"
  2. 2
    Convert the spec into scene-object pairs

    A buyer asking for 500 episodes has not told you anything useful yet. Ask how many distinct environments and objects. The Data Scaling Laws recipe of 32 pairs at 50 demonstrations each is 1,600 episodes and 32 room resets, and the room resets are the expensive part.

  3. 3
    Price the setup separately from the recording

    Moving the rig, re-fixing the cameras and re-checking calibration is fixed cost per environment. Bill it as such. Price per episode only and the diverse dataset buyers actually want is the one that loses you money.

  4. 4
    Add a verification line item

    One training run and a rollout video per delivered dataset. On the 24 GB tier that is roughly 1 to 3 USD of compute, and it converts arguments about quality into a fact.

    bash
    lerobot-train --dataset.repo_id=${HF_USER}/pick-cube-kitchen-01 \
      --policy.type=act --policy.device=cuda \
      --output_dir=outputs/train/act_verify --job_name=act_verify
  5. 5
    Write the format and the reject rule into the contract

    LeRobot version, camera count and resolution, frame rate, task string, who decides a take is rejected, and whether rejected takes are paid. The version clause alone will save you a day at some point.

The AY-Robots pricing page showing what a training run costs on rented spot GPUs
The pricing page at /pricing. The compute side of the business is the only part with a published number: about 1 to 3 USD on the 24 GB tier, about 4 to 12 USD on the A100 and H100 tier.

Where this does not work

Three honest limits, none of which the vendors in this space put on their landing pages.

  • Embodiment does not transfer for free at the top of the ladder. A buyer building a humanoid wants demonstrations on that humanoid. Your SO-100 data is useful for general manipulation research and for anyone training on SO-100 class arms, and that is a smaller room than the headline funding numbers suggest.
  • Volume alone has no bid. With more than 72,400 LeRobot-tagged datasets already free on Hugging Face, another 200 episodes of a cube going into a bin has no scarcity value. What is scarce is a documented variation plan, consistent cameras, and evidence the policy learned.
  • Remote work has a physics limit. Inference must sit next to the servos for fast tasks: the control loop is 20 ms per action step for ACT and up to 485 ms for Pi0.5, and adding public-internet round trips turns a working policy into a hesitant one. Remote inference is viable for slow pick-and-place, not for fast reactive motion. See inference latency.
  • The demand curve is young. Ken Goldberg of UC Berkeley, quoted in both the Rest of World and MIT Technology Review pieces, calls the effort noble but slow: even with hundreds of people working, he says, it is going to take a long time to get enough data, and longer than people think. Plan for it taking longer than the funding cycle.
The AY-Robots operator page, Become a Robot Operator from anywhere in the world, with a photo of the SO-100 arm operators drive
The operator route at /teleoperator. Selling hours on someone else's arm removes the equipment risk and the sales problem, and removes the upside with them.

If you want to test the work before spending anything, the live arm at /live streams a physical robot with no signup, queue-based, and /try lays out the three ways to start without hardware. For the technical side of what makes a dataset trainable rather than merely large, our guide on collecting high-quality VLA training data goes deeper on the recording side, the DROID dataset write-up covers how a 50-person effort was organised, and the RoboTurk piece covers the crowdsourcing experiment that started this line of work. For the hardware, the SO-100 complete guide is the setup reference, and imitation learning and policy define the terms buyers will use in their spec. If you are still choosing which model to record for, the five trainable policies are compared side by side with their minimum episode counts.

How much can I realistically earn recording robot demonstrations?

The published anchors in 2026 are about 15 USD per hour for at-home egocentric video at micro1, reported by MIT Technology Review in April 2026, and 30 to 55 USD per hour for on-site robot teleoperation in a US facility, per a live OpenTrain AI listing verified in August 2026. Contract dataset work has no published rate. Measure your accepted episodes per hour for a week before quoting anyone, because roughly half of wall-clock time at the LeRobot defaults goes to resets and re-records.

Can I just upload datasets to Hugging Face and get paid?

No. Hugging Face listed more than 72,400 datasets tagged LeRobot on 24 August 2026 and there is no payment mechanism attached to them. Publishing openly builds a portfolio and gets you found, which is worth doing, but the revenue comes from contract work with a named buyer and an agreed spec.

How many episodes does a buyer actually need?

Diversity matters far more than count. The Data Scaling Laws paper (ICLR 2025) found generalisation follows roughly a power law in the number of environments and objects, and that past about 50 demonstrations per environment-object pair extra demonstrations add very little. Their recipe was 32 environments with a unique object each, 50 demonstrations per pair. On AY-Robots the per-model floors are 50 episodes for GR00T N1.7, GR00T N1.5, Pi0.5 and ACT, and 30 for SmolVLA.

Does an SO-100 produce data a serious lab will buy?

For SO-100 class manipulation research, yes. For a humanoid programme, no. The value ladder XDOF described to TechCrunch in June 2026 puts teleoperation on the exact deployed robot at the top, teleoperated robots collecting general data in the middle, and human egocentric video at the bottom. An SO-100 sits in the middle tier, the highest rung a single person with roughly 150 EUR of hardware can reach.

What is the fastest way to prove my dataset is worth paying for?

Train a small policy on it and show the result. ACT at about 80 M parameters and SmolVLA at about 450 M both fit a 24 GB card, which on the AY-Robots spot tier is roughly 1 to 3 USD per run. A loss curve plus a rollout video converts an argument about quality into evidence.

Which LeRobot dataset version should I deliver?

Ask the buyer which policy they are training and write it into the contract. GR00T N1.7 and GR00T N1.5 need LeRobot v2.0 or v2.1; a v3.0 dataset crashes the GR00T loader and has to be converted down. Pi0.5, SmolVLA and ACT take v3.0. This one clause prevents the single most common delivery failure.

Sell hours instead of hardware risk

If the recording work interests you but buying and maintaining an arm does not, the operator route puts you on someone else's rig. Drive a real SO-100 class arm from anywhere and record LeRobot datasets from a teleoperation session.

See the operator route

Ready for high-quality robotics data?

AY-Robots connects your robots to skilled operators worldwide.

Get Started