Skip to content
RFL_GLOBAL
中文
  • RDT2 // 2025
  • UNIVERSAL TRANSFER
  • ANY EMBODIMENT

MORPHEUS

MANIPULATION — CROSS-EMBODIMENT VLA

Cross-embodiment Vision-Language-Action system built on the RDT2 architecture. One model works across any robot without retraining — wheeled platforms, robotic arms, quadrupeds, humanoids. Embodiment-specific action heads eliminate per-robot fine-tuning while maintaining manipulation precision.

MODULE STATUS: DEVELOPMENT

Embodiment Transfer

ANY

DIVISION
ANIMA
WAVE
W5
DOMAIN
MANIPULATION
WAVE 5 // ANIMA SUITE
CROSS-EMBODIMENT VLA
MORPHEUS // W5 // 050/079
01THE CHALLENGE

EVERY ROBOT NEEDS ITS OWN MODEL

Today's manipulation models are trained for a single robot morphology. A policy learned on a 7-DOF arm doesn't transfer to a wheeled mobile manipulator. A humanoid policy can't drive a quadruped. Every new platform requires collecting data, training a model, and validating from scratch — multiplying cost linearly with fleet diversity.

Heterogeneous fleets are the reality of deployment. Military, logistics, and industrial operations use arms, wheeled platforms, legged robots, and humanoids simultaneously. Without cross-embodiment transfer, every platform addition requires a full ML pipeline — data collection, training, validation, deployment. The scaling cost is prohibitive.

02THE SOLUTION

WHAT MORPHEUS DELIVERS

MORPHEUS is a cross-embodiment Vision-Language-Action system using the RDT2 (Robotics Diffusion Transformer 2) architecture. A shared vision-language backbone processes scene understanding, while embodiment-specific action heads translate unified representations into platform-native control signals.

CAPABILITIES

  • Shared vision-language backbone — one perception model for all embodiments
  • Embodiment-specific action heads — modular output layers per robot morphology
  • Action Head Registry — plug-and-play registration of new robot platforms
  • Zero per-robot fine-tuning — universal policy transfers directly
  • RDT2 diffusion transformer for high-fidelity action generation
  • Integrates PYGMALION + ATOMOS modules from ANIMA suite
03ENGINEERING

WHY THIS IS HARD

Cross-embodiment transfer requires solving fundamental representation mismatches:

  1. 01Action space heterogeneity: a 7-DOF arm, a differential-drive base, a 12-DOF quadruped, and a 30+ DOF humanoid have incompatible action spaces
  2. 02Embodiment-invariant representations: the shared backbone must extract manipulation intent without encoding morphology-specific biases
  3. 03Action head alignment: each embodiment-specific head must map from the same latent space to wildly different control signals
  4. 04Training data balance: preventing the model from overfitting to the most represented embodiment in the training mix
  5. 05Deployment validation: ensuring manipulation quality is maintained across all registered embodiments after each model update

MORPHEUS solves this through the RDT2 architecture — a diffusion transformer backbone that learns embodiment-invariant manipulation representations, paired with a registry of action heads that can be hot-swapped per deployment target.

04BENCHMARKS

PROOF, NOT PROMISES

Cross-embodiment transfer metrics:

PROOF, NOT PROMISES
METRICVALUE
Embodiment TransferUniversal
Per-Robot Fine-TuningZero
Supported MorphologiesMulti-embodiment
Policy CountSingle unified
Action Head SwapHot-swappable
Backbone ArchitectureRDT2 Diffusion Transformer
Perception PipelineShared ViT + LLM
Action GenerationDiffusion-based
Platform RegistrationPlug-and-play
Training EfficiencyN embodiments, 1 model
ANIMA IntegrationPYGMALION + ATOMOS
Deployment OverheadAction head only
05BUILD STATUS

WHAT'S BUILT TODAY

3/5 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
RDT2 ArchitectureCOMPLETEDiffusion transformer backbone with modular action head interface
Action Head RegistryCOMPLETEPlug-and-play registration for new embodiment action heads
Cross-Embodiment TrainingIN PROGRESSMulti-embodiment dataset curation and balanced training
Core modelsCOMPLETEShared backbone trained — action heads for 4 morphologies validated
API layerIN PROGRESSPublic REST + gRPC API — pending cross-embodiment inference routing
06APPLICATIONS

WHERE MORPHEUS DEPLOYS

  • APP_01

    HETEROGENEOUS FLEET DEPLOYMENT

    Deploy a single manipulation policy across mixed fleets — arms, wheeled platforms, quadrupeds, and humanoids operating in the same environment.

  • APP_02

    RAPID PLATFORM SCALING

    Add new robot platforms to an existing fleet by registering an action head — no data collection or retraining required.

  • APP_03

    UNIVERSAL MANIPULATION POLICY

    One policy handles pick-and-place, door opening, tool use, and object handover regardless of the executing robot's morphology.

  • APP_04

    DEFENSE & SECURITY

    Mixed autonomous fleets operating in contested environments — ground vehicles, legged platforms, and manipulator arms sharing a single intelligence layer.

  • APP_05

    WAREHOUSE AUTOMATION

    Deploy arms, mobile manipulators, and humanoids in the same facility with unified task assignment and manipulation capability.

  • APP_06

    RESEARCH PLATFORMS

    Train once, evaluate across embodiments — accelerate manipulation research by eliminating per-platform model development.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: RDT2 ARCHITECTURE

  • Robotics Diffusion Transformer 2 — diffusion-based action generation
  • Shared ViT vision encoder + language model backbone
  • Embodiment-invariant latent manipulation representations
  • Modular action head interface with standardized I/O contract

KEY INNOVATION

Decouples "what to do" from "how to move" — the shared backbone learns manipulation intent (grasp, place, push, rotate) while embodiment-specific action heads translate intent into morphology-native trajectories. New robots register an action head without touching the backbone.

ANIMA INTEGRATION

  • PYGMALION — cross-embodiment data generation and augmentation
  • ATOMOS — 1-bit quantization for edge deployment of the shared backbone

INFERENCE PIPELINE

PERCEPTION
Shared ViT encoder processes RGB input — embodiment-agnostic scene understanding
REASONING
Language model + diffusion transformer generate manipulation intent in latent space
ACTION
Embodiment-specific head decodes latent intent into native control signals — hot-swapped per robot
08PAPERS

RESEARCH PAPERS

  1. [01]RDT2: Robotics Diffusion Transformer for Cross-Embodiment Transfer
  2. [02]Cross-Embodiment Policy Learning — Pinto et al.
  3. [03]Diffusion Models for Robotic Action Generation