Post what data you are missing.Many people record. You get one Collection.
You describe the task, pick the arm and set the scope and the hourly rate. Whoever joins in records it at their own place, on their own hardware. You pay by the hour of recorded material, and nobody is paid out until you have accepted.
You can fill the whole request in without an account. The account comes at the end, and the request goes live once your credit covers the budget - reserved, not charged.
- 10 %
- Platform share, no other fee
- 1
- Collection, however many people record
- 6 - 22 €
- Hourly rates that are common today
- anonymous
- Your name appears nowhere
A model can only do as much as its data allows.
The architectures are open and interchangeable. The difference between a demo and a robot that really does the task is almost always in the recordings.
The model families people work with today are public. GR00T, Pi0.5, SmolVLA and ACT all run through the same training pipeline on this platform, and any team can download them. What a team cannot download are recordings of its own task, with its own gripper, in its own light, on its own table. That is what decides whether a policy hits the grasp or reaches past it.
Recordings like that do not happen on the side. They happen when somebody demonstrates the task a hundred times, moves the objects between runs, changes the light and deletes the failed attempts. That is work, and it comes up long before the first training run starts. Whoever only starts collecting when the model is needed waits weeks for something that would have trained in hours.
That is why a data request is an investment and not an order. You decide today which capability you will need in half a year and have it recorded from now on. The recordings stay, even when the next architecture arrives. Model weights age. Recordings of your own task do not.
- Model weights are interchangeable, recordings of your own task are not
- Collecting takes weeks, training takes hours
- The recordings survive the next change of architecture
- Collecting early costs time, collecting late costs the deadline
The recordings are the slow part. The training that follows rents its graphics cards by the minute, and needs none of your own.What training without a GPU costs
Three ways, and all three cost more than they look
This is not a criticism of the three ways. It is the calculation every team makes before picking one.
Collect it yourself
You need arms, cameras, a place the setup may stand for weeks, and above all people who record day after day. An engineer recording episodes is not writing code in that time. Doubling the volume means a second table, a second arm, a second person.
Hire a data vendor
Quote, contract, payment up front, waiting. You pay before you have seen a single recording, and you rarely learn who recorded it, under what conditions. If the gripper does not match yours, what starts is a discussion instead of a re-delivery.
Take a public dataset
Free and there immediately, but almost always slightly off: the right arm with the wrong task, the right task with a different camera layout, a format version your trainer does not read. Enough for a first attempt, rarely for what you actually want to build.
If the first of the three is what you actually want, the recording client is the page for you.How recording without Python works
The fourth way: you post, other people record
Whoever has a matching arm builds your task at their own place and records it. You pay only for deliveries you have accepted, and however many people deliver at the same time, it stays one Collection and one download.
See open requestsA data request is a posting with a pot behind it.
Two mandatory entries, a budget, a tile in the marketplace. That is all it takes to start.
You fill in a short form. Which robot arm should be visible in the recordings, and what should be happening in them. The description is a free text field and not a form tree with twelve fields for table height and grasp angle: you write the task down the way you would explain it to a person. If a photo, a short video or a document explains it faster than a paragraph, attach it. None of that is mandatory.
On top of that comes what makes the request a request: how many hours you need and what an hour should pay. The budget follows from those two, and you top it up as credit beforehand, exactly like a normal purchase in the marketplace.
The posting becomes a tile in the marketplace, right next to the finished datasets. The tile carries the fill level: how many episodes have been recorded, how many hours that adds up to, how much is missing until the pot is full, what an hour pays and how large the budget is. Whoever opens the tile sees the full description with every attachment. What is nowhere to be found: who posted the request.
Whoever wants to join builds the task at their own place. Their arm, their cameras, their table. They record it themselves, with the same client they already use for their own datasets. You provide no hardware, you set up no workstation, and you look after nobody else's machine. Your work consists of two things: describing the task understandably and reviewing the deliveries.
- Robot arm: which arm should appear in the recordings
- Description: what to do, as free text in your own words
- Attachments: images, videos, documents, all optional
- Scope and hourly rate: the budget follows from them
The request form is short on purpose
Robot arm
Picked from the arm catalogue the platform knows; custom builds go in as free text. Whoever has no matching arm never sees your request. That is not a formality: a different arm has different joints and different value ranges, and that cannot be reconciled afterwards.
Description
What the arm should do, what the scene looks like, what counts as a success and what does not. Up to 4,000 characters. Strangers build their setup from this text, so take the time here that you save on everything else.
Attachments, optional
Images, videos, documents. A photo of your setup explains more than three paragraphs. None of it is mandatory. File metadata is stripped on upload, so an author field in a PDF does not give you away.
Exclusive, if you want
One checkbox. Without it the dataset stays with the collector, who may also offer it in the marketplace. With it, that sales route is blocked for the dataset. It sits on the tile from the start, so every collector sees it before the first recording, and it cannot be flipped later, in either direction.
Your name appears nowhere
The request appears without a requester. A collector decides by the task, not by the sender.
See your whole budget before you post anything
Three sliders. Behind them is the formula that bills you later, not a marketing version of it.
You state a price per hour, and it is paid per episode: the hourly rate is scaled down to the expected episode length. That episode price is frozen when an assignment is made, so a rate you change later applies only to new assignments.
The slider covers the usual lengths. The rule is wider: between 5 seconds and 30 minutes an episode is billable, and outside that window a recording does not count at all.
common today: 6 to 22 EUR
Common today, from the pricing page: Beginner 6 to 10 euros per hour, Skilled 10 to 15, Expert 15 to 22. In a data request you set the rate yourself.
This is pure recording time. Nobody here decides how many hours a day somebody records, so we do not convert it into days.
This is how your tile looks in the marketplace
Not somebody else's request. Yours, with your numbers, exactly as a collector would see it.
Pick the sponge out of the tray and put it down on the mat, one item at a time.
You pay exclusively for deliveries that passed the automatic checks. Whatever you reject, you get back.
Posting costs nothing. Not a cent moves before the first delivery arrives.
Post a requestAn hour here is an hour of material, not an hour of presence.
There is no clock anyone could start and stop. What counts is what is inside the recording.
An hour is 3,600 seconds of accepted recorded material. The seconds are read out of the delivery itself: number of frames divided by the frame rate, both from the metadata of the LeRobot dataset, read on the server. No number from a form, no browser clock, no reported session length.
That has a consequence in your favour: an hour in which nobody records produces not a cent. Whoever joins a data request sits at their own hardware in their own home. We see neither the screen nor the arm there. A clock only the collector can start would be a self-declaration with a timestamp. So there is none.
Setup time is therefore not a separate line item either. Getting the objects, placing the cameras, calibrating the arm, doing a test run: nobody invoices that, it is priced into the hourly rate. Whoever posts a demanding task with short episodes has to raise the rate, or they will find no collectors. That balances itself out.
- You pay for material, not for presence
- The seconds come from the delivered file, not from a form
- An hour without a recording costs nothing
- Setup and rigging are included in the hourly rate, not billed separately
Fifteen people record. You download once.
Add a collector and watch the pot fill. The delivery that misses the specification stays out.
1 Collection, 4 linked datasets
20 % of 2,000 episodes
396 episodes, 6.6 hours
One download, LeRobot format, trainable without conversion
Specification: SO-101 (SO-ARM101), 30 fps, 2 camera roles
- so101_sponge_tray_01Goes into the download30 fps148 episodes, 2.4 h
- so101_sponge_tray_02Goes into the download30 fps126 episodes, 2.2 h
- so101_sponge_tray_03Goes into the download30 fps122 episodes, 2.1 h
- so101_sponge_tray_04Stays out25 fps118 episodes, 2.1 h
Specification says 30 fps
3 of 4 match, 1 stays out: different frame rate
This is where distributed collecting usually falls apart. Every upload leaves its own folder behind: its own episode numbering from zero, its own metadata file, in case of doubt a different frame rate. The tidying up lands exactly when training was supposed to start.
Everyone keeps their own dataset
Nothing is taken away from a collector: their dataset stays theirs, only linked to your request. Out of those links you get one view, one fill level, one episode count, one download.
Assembled at download, not at upload
The archive is built when you download: continuous episode numbers, shared metadata, only the camera roles you asked for. And you see what stays out before you press the button.
The specification is set at the front
Same arm, same frame rate, same camera roles. Roles, not numbers: a camera_1 at the wrist here and at the table front there still loads, still trains, and the policy drives wildly afterwards. So every collector maps their cameras to your roles before they record.
Your money moves in four steps, and you decide the last one
Click through it and see where your money sits at each point. Every step gets its own ledger entry, so it stays traceable what an amount was taken for.
Your credit
Collector's account
What you are not charged for in the first place
Too short. Too long. Discarded by the collector themselves. Aborted halfway through. Duplicate. Started without funds. Recordings like that are recorded, but not billed, and they do not need your acceptance at all.
Plans, credit and platform feesThree gates sit between a recording and your money.
One stands in front, one is done by the machine, the last one is done by you. What each of them does is here, and what it does not do as well.
The first gate stands before the first recording. Next to the description box the request carries a machine-readable specification: robot type, frame rate, camera roles, resolution, task text, expected episode length with a window. Whoever wants to join uploads a small sample batch before taking the request on. If it does not match the specification, they never start. That is not red tape, it is the condition for the deliveries to be assemblable at all afterwards: different frame rates or a different camera set cannot be reconciled after the fact.
The second gate is the machine. Every delivery runs through the same check line that every listing in the marketplace runs through today: eleven named checks, all or nothing. Whether the archive can be unpacked, whether the unpacked size is within range, whether the archive copy is unchanged, whether the metadata file is readable, whether the number of episodes matches the metadata and the number of data files, whether the frame rate is a whole number in the valid range, whether every episode has a length greater than zero, whether the total number of frames matches the sum of the episode lengths, whether the camera images are really there, whether the robot type is set and whether the total duration is plausible. If one check fails, the delivery does not pass, and the message names exactly the check that failed.
What these eleven checks do not do belongs in the same breath: they check the form, not the content. They do not read motion data and they do not look at the images. A technically flawless dataset showing the wrong thing passes them. That is exactly why there is a third gate, and exactly why it sits with you.
The third gate is you. Every delivery has a review state, you see it with recording date, duration and origin, and you decide whether it belongs in your Collection. If you reject, you give a reason, and the amount goes back. If you do not react at all, the delivery counts as accepted once the review period has run out. That is not a trap, it is the protection of the other side: a collector must not wait indefinitely for their money because somebody is not looking at their review list.
On origin: whoever delivers into a request signs the same declaration as anybody who lists a dataset in the marketplace. They confirm that they hold the rights to the recording and that it contains no personal data of third parties without consent. This confirmation is not a checkbox that disappears afterwards: it is attached to the contribution with the time, the account, the text version and the IP address. It is a declaration and not a proof, and we would rather say that ourselves. What it does achieve: it names a person who stands behind it and makes recourse possible. Compared with a delivery that comes with nothing but files and a promise, that is a difference.
- A specification and a sample batch before anyone starts
- Eleven automatic format checks per delivery, each one named
- Your acceptance per delivery, with a reason when you reject
- Authorship declaration with time, account, text version and IP
You see everything. Nobody sees you.
Whoever posts stays unnamed. Whoever collects sees only what they contributed themselves.
No company name, no logo, no contact person, no country on the tile. A competitor cannot read off the marketplace what you are working on.
That is not only the display. The public request route does not hand out the requester's identifier at all, not even as a field the page simply does not show. Whoever reads the response in the network tab does not find you there.
The other direction is fixed too. You see every delivery on your request, with recording date, duration and review state. A collector sees only their own and cannot retrieve anybody else's: it is their own dataset.
| What | Public | You as the requester | A single collector |
|---|---|---|---|
| Task description and attachments | visible | visible | visible |
| Hourly rate and budget | visible | visible | visible |
| Exclusivity checkbox | visible | visible | visible |
| Fill level as a sum | visible | visible | visible |
| Individual deliveries | not visible | all of them | only their own |
| Earnings per collector | not visible | not visible | only their own |
| Who posted the request | nobody | - | nobody |
Five things that simply do not come up on this route.
None of this is a discount. These are work steps that do not exist here.
There is no middleman who buys somewhere else themselves. Your posting stands in the marketplace, and the people who record read it directly. You do not negotiate a delivery volume, you set it. And you do not wait for a quoting process before anybody switches on a camera.
Every delivery comes with an account that uploaded it, with timestamps set by the server and not by the collector's machine, and with an authorship declaration that stays attached to the contribution. Whoever gets money has set up a payout account beforehand. That is not a certificate about the content, but it is a chain that reaches back to the person who recorded.
Publishing a request costs nothing. No setup fee, no pilot package you have to buy first in order to see whether the result fits, no project flat rate before the first episode. What is paid for are accepted deliveries, nothing else, and the platform share comes out of the hourly rate instead of being added on top.
And the fifteenth collector makes no more work for you than the first. Each of them sets up their own environment from your description. You provide nobody with hardware, you configure nothing for anybody, and you do not become the contact person for fifteen different setups. That is exactly where collecting it yourself falls apart, and exactly what is gone here.
- No contract and no negotiation with a data vendor
- Every delivery carries an account, a server timestamp and an authorship declaration
- No setup fee and no pilot package you have to buy first
- We run the format checks before a delivery reaches you
- No third-party setup to look after: the fifteenth collector is no more work than the first
Two ways, the same intention
Both lead to data that did not exist before. The difference lies in when you pay, what you see and how you get more.
| Point | Order with a data vendor | Data request on AY-Robots |
|---|---|---|
| Getting started | Request a quote, negotiate, sign a contract, often a minimum volume | Fill in a form, top up credit, the tile is live |
| Payment | Up front or by milestone, before a recording exists | By the hour of material, and paid out only after you have accepted |
| Setup cost | Setup and project flat rates before the first episode | No surcharge for posting, one fixed platform share out of the rate |
| Origin | You get files and a promise | Every delivery comes with an account, a server timestamp and an authorship declaration |
| More data | New contract, new waiting time | More collectors only means the pot fills up faster |
| Work on your side | You coordinate setups, contact people and follow-up questions | You describe and you review, the setups stand at the collectors |
| Result | Deliveries in batches, one format per batch | One Collection, one download, in LeRobot format |
| Fit | Whatever the vendor has in the catalogue or in the quote | Your arm, your task, your camera roles |
| Stopping | Remaining term and notice period | Pause the request, unspent credit stays yours and does not expire |
The left column describes the usual course of commissioning a data vendor. It is not a statement about any particular vendor.
The limits the platform actually enforces
Everything on this page in numbers, as it applies today.
- Price statement
- Euros per hour
- The hourly rate is converted into an episode price: hourly rate times expected episode length divided by 3,600, rounded to whole cents.
- Billing unit
- The episode
- You pay per episode, at the price that was frozen when the assignment was made. An hourly rate changed later applies only to new assignments.
- Platform share
- out of the rate
- On every billed episode, never on top of it. Share and payout together add up to exactly the episode price, no cent disappears in between. The current share stands at the top of this page.
- Hourly rates today
- 6 to 22 euros
- The bands that apply on the platform today: Beginner 6 to 10, Skilled 10 to 15, Expert 15 to 22. In a data request you set the rate yourself.
- Episode length
- 5 seconds to 30 minutes
- Outside that window a recording does not count as a billable episode.
- Hourly rate ceiling
- 1,000 euros per hour
- Not a business rule but a typo brake. Whoever types two zeros too many notices it right away.
- Task description
- Up to 4,000 characters
- Free text. Attachments come on top and are optional.
- Robot arms
- 49 families in the filter, 353 models in the catalogue
- From SO-100 and Koch through Franka, Universal Robots and xArm to humanoids. Custom builds can be entered as free text.
- Dataset format
- LeRobot v2.1, Parquet and MP4
- The same format the client writes. Readable by the lerobot scripts without conversion.
- Automatic checks
- Eleven, all or nothing
- From unpackable to cameras. If one of them fails, the delivery does not pass. They are format checks.
- Exclusivity
- Optional checkbox on the request
- With the checkbox, reselling the dataset is blocked. The checkbox is on the tile from the start and cannot be flipped later.
- Credit
- Does not expire
- Topped-up credit adds up and is not cut back with age. Whatever is left at the end stays yours.
Frequently asked questions
No. That is exactly what the data request is for. The arms stand at the collectors. You need an account, credit and a task that can be described.
Who actually records this
Nobody is employed here. The people who fill a request are the ones these three pages are written for - and each of them can tell you what your request looks like from the other side.
People with an arm nobody else has
The rarer the arm, the fewer recordings exist for it and the smaller the pool that can take your request. What that scarcity does to the rate, from their side.
Read moreUniversity labs and student projects
An arm between two theses stands idle most of the time. What a supervisor has to sign off before any of those hours can go into your request.
Read moreData providers and brokers
Some of what you need may already be recorded. This is the page the people selling it read - useful for judging what a finished dataset costs against a commission.
Read moreRelated on AY-Robots
Teleoperation
Remote operators drive your arm and record demonstration data around the clock.
Read moreCloud training
Fine-tune GR00T, Pi0.5, SmolVLA or ACT on a rented GPU, billed by the minute.
Read moreCloud inference
Serve a checkpoint in the cloud and watch a real arm run it in your browser.
Read moreData is the foundation, not an accessory.
A model gets better because somebody recorded the fitting data, not because it trained for longer. A request that stands in the marketplace today collects while you work on something else.
Fill it in first, sign up after. It goes live once your credit covers the budget, and that credit stays reserved until you accept a delivery.