AGI

GO-2 (Genie Operator-2)

AgiBot, China · April 2026

AGIBOT

Reasoning VLAClosed weightsAlso written Genie Operator-2, AgiBot GO-2
Parameters
not published
Not disclosed.
GPU memory
not published
weights at bf16, computed
Inference latency
not published
per action step
Weights
closed
no public checkpoint

What it is

GO-2 is AGIBOT's second-generation embodied foundation model, unveiled at the company's 2026 Partner Conference on 17 April 2026 as the successor to GO-1. AGIBOT frames the problem it targets as the semantic-actuation gap, meaning that in conventional VLAs the high-level reasoning signal and the motor command drift apart over long horizons. GO-2 answers this with an action chain-of-thought, where the model first emits a sequence of high-level action intents as a macro-plan, and with an asynchronous dual-system architecture that runs a low-frequency semantic planning module (System 2) alongside a high-frequency action following module (System 1). AGIBOT reports strong benchmark numbers including 98.5 percent average on LIBERO and 82.9 percent real-world success from simulation-only training, but publishes no weights, no code, no technical report and no parameter count, so none of it is independently reproducible.

Architecture

Backbone
Not disclosed. Described as an asynchronous dual-system design with a low-frequency semantic planning module (System 2, the general commander) and a high-frequency action following module (System 1, the agile executor).
Action head
Not disclosed. GO-2 emits an action chain-of-thought, a sequence of high-level action intents forming a macro-plan, which the action following module converts into motor commands rather than mapping language directly to actions.
Parameters
Not disclosed. AGIBOT describes GO-2 as a ViLLA (vision-language-latent-action) embodied foundation model without naming a backbone or size.
Pretraining data
Tens of thousands of hours of interaction data, not further quantified. AGIBOT separately maintains the open AgiBot World Colosseo dataset collected by more than 100 homogeneous robots, but the announcement does not state that this is the GO-2 training corpus.
Embodiments
not named in the announcement, deployed through AGIBOT's Genie Studio development platform, trained in simulation with Genie Sim 3.0

What hardware it needs

This is the section most people came for, so it is worth being precise about which numbers are measured and which are arithmetic.

Not published. No hardware or memory requirement appears in any primary source.

The authors publish no memory or latency figure for this model. Everything above is computed from the parameter count.

Reported results

Grouped by what the evaluation actually asked. Values from different suites are not comparable with each other, so each one keeps its suite and split.

Simulation

Reproducible benchmark suites. Everyone runs the same episodes, so the numbers can be compared inside a suite.

  • LIBERO average over Spatial, Object, Goal and Long
    98.5%
    success rateSelf-reported by AGIBOT, which states GO-2 ranks first across all four suites. No per-suite breakdown, seed count or evaluation protocol is published.source
  • LIBERO-Plus zero-shot under environmental disturbances
    86.6%
    success rateSelf-reported by AGIBOT. LIBERO-Plus is the perturbed variant of LIBERO, so this is the more informative robustness number.source
  • Genie Sim 3.0 sim-to-real transfer trained on simulation data only, evaluated on real hardware
    82.9%
    success rateSelf-reported by AGIBOT. Task set and robot hardware are not specified.source
  • VLABench cross-category and texture generalization
    47.4 score
    average scoreSelf-reported by AGIBOT. Score scale and comparison baselines are not given.source

Real world

No results in this category are published for this model.

Fine tuned tasks

No results in this category are published for this model.

On real hardware

The strongest real-world claim is sim-to-real: trained solely on simulation data from Genie Sim 3.0, GO-2 reached 82.9 percent success rate in real-world testing. The robot hardware used is not named.

Fine tuning it yourself

No public fine-tuning path. AGIBOT states GO-2 is integrated into its Genie Studio platform and claims about 10x improvement in training efficiency, task startup reduced to minutes, and 2 to 4 times better success rates with over 50 percent less data, but names no baseline for those comparisons and publishes no tooling.

Where it helps, where it does not

Strengths

  • Action chain-of-thought decomposes long-horizon tasks into ordered stages instead of mapping language straight to motor commands, which targets error accumulation over long horizons.
  • Asynchronous dual-system split lets planning run slowly while control runs fast, which is the standard answer to VLM inference cost in closed-loop control.
  • Reports a strong robustness number on LIBERO-Plus (86.6 percent zero-shot under disturbances), not just the near-saturated LIBERO average.
  • 82.9 percent real-world success from simulation-only training is the only pure sim-to-real number in this comparison.
  • Backed by AGIBOT's own hardware, simulator (Genie Sim 3.0) and deployment platform (Genie Studio).

Limits

  • No weights, no code, no technical report and no arXiv preprint. The only primary source is AGIBOT's own announcement page.
  • No parameter count, backbone, control frequency, latency or VRAM figure is published.
  • LIBERO, LIBERO-Plus and VLABench numbers are self-reported without per-suite breakdowns, seed counts or a published evaluation protocol, so they are not reproducible.
  • The efficiency claims (about 10x training efficiency, 2 to 4x success rate, over 50 percent less data) name no baseline.
  • The robot hardware GO-2 runs on is not named in the announcement, and neither is the sim-to-real task set.
  • LIBERO is close to saturated across the field, so a 98.5 percent average carries little discriminative information on its own.

Sources

Everything on this page was taken from these documents. Where they disagree with what you read here, they win.

Entry last checked 2026-08-11.

Try a policy on a real arm

A physical SO-100 is online and free to drive in the browser, and five of the models in this arena can be fine tuned on a dataset you record with your own.