Cloud Inference

Your policy drives real robots— without a GPU of your own.

Upload GR00T, Pi0.5, SmolVLA, or ACT as a checkpoint, serve it in the cloud, and watch the arm move live in your browser. From checkpoint to moving arm in minutes.

100+
policies ready to deploy
5
architectures supported
Minutes
from checkpoint to endpoint
Live
watch in your browser

Latency per model, honestly measured

Inference time per action chunk on the serving GPU — the same values the dashboard shows you when choosing a model.

Hugging FaceACT~20 ms
NVIDIAGR00T N1.5~98 ms
NVIDIAGR00T N1.7~152 ms
Hugging FaceSmolVLA~245 ms
Physical IntelligencePi0.5~485 ms

On top of that comes the network path between arm and GPU — the dashboard shows you both separately (p50/p99 live).

Three steps to a moving arm

01

Upload a checkpoint

Your fine-tune from training — or a checkpoint from the marketplace. LeRobot format, no repackaging.

02

GPU starts for you

The platform rents the right cloud GPU, loads the model, and lets you know as soon as the endpoint is up — in minutes, not days.

03

The arm moves, you watch

The SO-100 pulls its actions from the serving API. In your browser you see cameras, joint angles, and latency live.

Pay when the arm moves — not otherwise

Cloud inference pays off because robotics GPUs would sit idle most of the time.

No base fee

You pay GPU minutes only while your endpoint is running. Stop it, and it costs nothing more.

Billed by the minute

The price of the chosen GPU class is fixed before you start — no surprises on the bill.

Instead of a €2,000 workstation

An inference-capable GPU costs four figures and mostly sits idle. Here you rent it only for the minutes the arm is actually moving.

Set up your first endpoint

Create an account, upload a checkpoint, start the GPU — the rest is just watching.