Skip to content
RFL_GLOBAL
中文
  • WAVE 5 // SIMULATION
  • CHIRON + PYGMALION
  • CO-TRAINING

CENTAUR

SIM-AND-HUMAN CO-TRAINING

SimHum co-training framework: joint pretraining on simulation trajectories and human demonstrations, fine-tuning on small real-robot data. Simulation provides scale and diversity, human data provides physical realism. The result: policies that generalize far beyond what either data source achieves alone.

MODULE STATUS: DEVELOPMENT

Improvement Over Real-Only

SIM+H

DIVISION
ANIMA
WAVE
W5
DOMAIN
SIMULATION
WAVE 5 // ANIMA SUITE
SIMULATION — SIM-AND-HUMAN CO-TRAINING
CENTAUR // W5 // 009/079
01THE CHALLENGE

REAL DATA IS EXPENSIVE. SIM DATA IS CHEAP BUT WRONG.

Training robust robot policies demands thousands of demonstrations, but collecting real-world data is slow, expensive, and dangerous. A single manipulation task can require hundreds of hours of human teleoperation — and the resulting policy still fails on out-of-distribution scenarios it has never encountered.

Simulation offers unlimited scale and diversity, but sim-to-real transfer remains brittle. Physics engines approximate reality, rendering is imperfect, and contact dynamics diverge. Policies trained purely in simulation fail when deployed on real hardware. The field needs a framework that extracts the best of both worlds.

+40%
OVER REAL-ONLY BASELINES
62.5%
OOD SUCCESS RATE
80
REAL DEMOS REQUIRED
02THE SOLUTION

WHAT CENTAUR DELIVERS

CENTAUR implements a SimHum co-training framework that jointly pretrains on simulation trajectories and human demonstrations, then fine-tunes on a small set of real-robot data. Simulation provides the scale and diversity needed for generalization; human demonstrations provide the physical realism needed for deployment.

CAPABILITIES

  • Joint pretraining on mixed sim + human demonstration data with domain-aware weighting
  • MuJoCo simulation pipeline generating diverse manipulation trajectories at scale
  • Real-robot fine-tuning stage requiring only 80 demonstrations for robust transfer
  • Out-of-distribution generalization: 62.5% success on unseen scenarios with minimal real data
03ENGINEERING

WHY THIS IS HARD

Merging simulation and human data into a single training pipeline requires solving multiple coupled problems:

  1. 01Domain alignment: simulation observations and real-world observations occupy different distributions — naive mixing degrades both
  2. 02Action space mismatch: MuJoCo joint-space actions vs. human teleoperation Cartesian commands require unified representation
  3. 03Curriculum design: when to weight sim data vs. human data during pretraining phases for maximum transfer
  4. 04Fine-tuning stability: adapting a co-trained backbone to real hardware without catastrophic forgetting of sim-learned diversity
  5. 05Evaluation rigor: OOD benchmarks must test genuine generalization, not memorized sim scenarios

CENTAUR solves this with domain-aware data mixing, unified action representations, and a staged pretraining → fine-tuning pipeline optimized for minimal real-data requirements.

04BENCHMARKS

CO-TRAINING PERFORMANCE

Measured against real-only and sim-only baselines:

CO-TRAINING PERFORMANCE
METRICVALUE
Improvement Over Real-Only+40%
OOD Success Rate62.5%
Real Demos Required80
Sim-to-Real Improvement Factor7.1×
05BUILD STATUS

WHAT'S BUILT TODAY

3/6 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
MuJoCo Training LoopCOMPLETEStable sim trajectory generation with domain randomization
Sim Trajectory GeneratorCOMPLETEDiverse manipulation trajectories at scale in MuJoCo
Human Demo IntegrationIN PROGRESSUnified action representation for mixed data training
Real-Robot Fine-TuningIN PROGRESSStaged fine-tuning pipeline with catastrophic forgetting mitigation
Core modelsCOMPLETECo-training backbone validated on benchmark tasks
API layerIN PROGRESSREST API for training job submission and model serving
06APPLICATIONS

WHERE CENTAUR DEPLOYS

  • APP_01

    DATA-EFFICIENT TRAINING

    Train robust manipulation policies with 80 real demonstrations instead of thousands — simulation provides the missing diversity.

  • APP_02

    SIM-TO-REAL TRANSFER

    Bridge the reality gap with co-training that combines simulated scale with human physical realism for reliable hardware deployment.

  • APP_03

    SMALL-DATA ROBOTICS

    Enable new robot deployments where collecting large real datasets is impractical — startups, novel hardware, constrained environments.