- RDT2 // 2025
- UNIVERSAL TRANSFER
- ANY EMBODIMENT
MORPHEUS
MANIPULATION — CROSS-EMBODIMENT VLA
Cross-embodiment Vision-Language-Action system built on the RDT2 architecture. One model works across any robot without retraining — wheeled platforms, robotic arms, quadrupeds, humanoids. Embodiment-specific action heads eliminate per-robot fine-tuning while maintaining manipulation precision.
MODULE STATUS: DEVELOPMENTEmbodiment Transfer
ANY
- DIVISION
- ANIMA
- WAVE
- W5
- DOMAIN
- MANIPULATION
- WAVE 5 // ANIMA SUITE
- CROSS-EMBODIMENT VLA
EVERY ROBOT NEEDS ITS OWN MODEL
Today's manipulation models are trained for a single robot morphology. A policy learned on a 7-DOF arm doesn't transfer to a wheeled mobile manipulator. A humanoid policy can't drive a quadruped. Every new platform requires collecting data, training a model, and validating from scratch — multiplying cost linearly with fleet diversity.
Heterogeneous fleets are the reality of deployment. Military, logistics, and industrial operations use arms, wheeled platforms, legged robots, and humanoids simultaneously. Without cross-embodiment transfer, every platform addition requires a full ML pipeline — data collection, training, validation, deployment. The scaling cost is prohibitive.
WHAT MORPHEUS DELIVERS
MORPHEUS is a cross-embodiment Vision-Language-Action system using the RDT2 (Robotics Diffusion Transformer 2) architecture. A shared vision-language backbone processes scene understanding, while embodiment-specific action heads translate unified representations into platform-native control signals.
CAPABILITIES
- Shared vision-language backbone — one perception model for all embodiments
- Embodiment-specific action heads — modular output layers per robot morphology
- Action Head Registry — plug-and-play registration of new robot platforms
- Zero per-robot fine-tuning — universal policy transfers directly
- RDT2 diffusion transformer for high-fidelity action generation
- Integrates PYGMALION + ATOMOS modules from ANIMA suite
WHY THIS IS HARD
Cross-embodiment transfer requires solving fundamental representation mismatches:
- 01Action space heterogeneity: a 7-DOF arm, a differential-drive base, a 12-DOF quadruped, and a 30+ DOF humanoid have incompatible action spaces
- 02Embodiment-invariant representations: the shared backbone must extract manipulation intent without encoding morphology-specific biases
- 03Action head alignment: each embodiment-specific head must map from the same latent space to wildly different control signals
- 04Training data balance: preventing the model from overfitting to the most represented embodiment in the training mix
- 05Deployment validation: ensuring manipulation quality is maintained across all registered embodiments after each model update
MORPHEUS solves this through the RDT2 architecture — a diffusion transformer backbone that learns embodiment-invariant manipulation representations, paired with a registry of action heads that can be hot-swapped per deployment target.
PROOF, NOT PROMISES
Cross-embodiment transfer metrics:
| METRIC | VALUE |
|---|---|
| Embodiment Transfer | Universal |
| Per-Robot Fine-Tuning | Zero |
| Supported Morphologies | Multi-embodiment |
| Policy Count | Single unified |
| Action Head Swap | Hot-swappable |
| Backbone Architecture | RDT2 Diffusion Transformer |
| Perception Pipeline | Shared ViT + LLM |
| Action Generation | Diffusion-based |
| Platform Registration | Plug-and-play |
| Training Efficiency | N embodiments, 1 model |
| ANIMA Integration | PYGMALION + ATOMOS |
| Deployment Overhead | Action head only |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| RDT2 Architecture | COMPLETE | Diffusion transformer backbone with modular action head interface |
| Action Head Registry | COMPLETE | Plug-and-play registration for new embodiment action heads |
| Cross-Embodiment Training | IN PROGRESS | Multi-embodiment dataset curation and balanced training |
| Core models | COMPLETE | Shared backbone trained — action heads for 4 morphologies validated |
| API layer | IN PROGRESS | Public REST + gRPC API — pending cross-embodiment inference routing |
WHERE MORPHEUS DEPLOYS
- APP_01
HETEROGENEOUS FLEET DEPLOYMENT
Deploy a single manipulation policy across mixed fleets — arms, wheeled platforms, quadrupeds, and humanoids operating in the same environment.
- APP_02
RAPID PLATFORM SCALING
Add new robot platforms to an existing fleet by registering an action head — no data collection or retraining required.
- APP_03
UNIVERSAL MANIPULATION POLICY
One policy handles pick-and-place, door opening, tool use, and object handover regardless of the executing robot's morphology.
- APP_04
DEFENSE & SECURITY
Mixed autonomous fleets operating in contested environments — ground vehicles, legged platforms, and manipulator arms sharing a single intelligence layer.
- APP_05
WAREHOUSE AUTOMATION
Deploy arms, mobile manipulators, and humanoids in the same facility with unified task assignment and manipulation capability.
- APP_06
RESEARCH PLATFORMS
Train once, evaluate across embodiments — accelerate manipulation research by eliminating per-platform model development.
UNDER THE HOOD
FOUNDATION: RDT2 ARCHITECTURE
- Robotics Diffusion Transformer 2 — diffusion-based action generation
- Shared ViT vision encoder + language model backbone
- Embodiment-invariant latent manipulation representations
- Modular action head interface with standardized I/O contract
KEY INNOVATION
Decouples "what to do" from "how to move" — the shared backbone learns manipulation intent (grasp, place, push, rotate) while embodiment-specific action heads translate intent into morphology-native trajectories. New robots register an action head without touching the backbone.
ANIMA INTEGRATION
- PYGMALION — cross-embodiment data generation and augmentation
- ATOMOS — 1-bit quantization for edge deployment of the shared backbone
INFERENCE PIPELINE
- PERCEPTION
- Shared ViT encoder processes RGB input — embodiment-agnostic scene understanding
- REASONING
- Language model + diffusion transformer generate manipulation intent in latent space
- ACTION
- Embodiment-specific head decodes latent intent into native control signals — hot-swapped per robot
RESEARCH PAPERS
- [01]RDT2: Robotics Diffusion Transformer for Cross-Embodiment Transfer
- [02]Cross-Embodiment Policy Learning — Pinto et al.
- [03]Diffusion Models for Robotic Action Generation