- MetricAnything // 2601.22054
- GPU-train, MLX inference
- WAVE 6
ODIN
ALL-SEEING DEPTH
ODIN delivers absolute metric depth from a single RGB image by pre-training on 20 million heterogeneous 3D data sources — LiDAR, stereo, RGB-D, and Structure-from-Motion. Unlike relative depth estimators that only rank pixels by distance, ODIN outputs real-world measurements in meters, enabling robots to judge distances with centimeter-level accuracy without any manual calibration. The system uses sparse metric prompts to adapt across sensor types, making it the foundational depth backbone for the entire ANIMA perception stack. Every downstream module — from SLAM to manipulation planning — depends on ODIN's accurate distance measurements.
MODULE STATUS: DEVELOPMENTTraining Sources
20M
- DIVISION
- ANIMA
- WAVE
- W6
- DOMAIN
- DEPTH SYSTEMS
- WAVE 6 // ANIMA SUITE
- FOUNDATION — METRIC DEPTH
DEPTH WITHOUT GROUND TRUTH
Robots need exact distances in meters, not relative ordering. Traditional depth sensors are expensive and fail in many conditions.
Every downstream module depends on accurate distance measurements.
WHAT ODIN DELIVERS
ODIN delivers absolute metric depth from a single RGB image by pre-training on 20 million heterogeneous 3D data sources. Unlike relative depth estimators, ODIN outputs real-world measurements in meters.
CAPABILITIES
- Pre-trained on 20M heterogeneous 3D sources
- Outputs absolute metric depth in meters
- Sparse metric prompts for cross-sensor adaptation
- Foundation backbone for ANIMA perception stack
WHY THIS IS HARD
Building ODIN requires solving multiple coupled problems:
- 01Unifying heterogeneous depth sources with different scales and noise
- 02Learning absolute scale without per-scene calibration
- 03Maintaining accuracy across environments
- 04Real-time inference for robotics
ODIN solves these through careful architecture design and rigorous validation.
PROOF, NOT PROMISES
Key metrics:
| METRIC | VALUE |
|---|---|
| Training Sources | 20M+ |
| Output | Metric (meters) |
| Calibration | None |
| Backend | GPU + MLX |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Pre-training Pipeline | COMPLETE | 20M multi-modal pre-training |
| Multi-Source Fusion | COMPLETE | LiDAR+stereo+RGB-D unified |
| Metric Prompting | IN PROGRESS | Cross-sensor adaptation |
| Edge Inference | IN PROGRESS | MLX optimization |
| Core models | COMPLETE | Foundation depth validated |
| API layer | IN PROGRESS | REST + streaming |
WHERE ODIN DEPLOYS
- APP_01
AUTONOMOUS NAVIGATION
Metric depth for obstacle avoidance without LiDAR.
- APP_02
SLAM BACKBONE
Foundation depth for THOR, BALDUR, HEIMDALL.
- APP_03
MANIPULATION
Distance measurements for grasp planning.
UNDER THE HOOD
FOUNDATION: METRICANYTHING
- Pre-trained on 20M heterogeneous 3D sources
- Outputs absolute metric depth in meters
- Sparse metric prompts for cross-sensor adaptation
KEY INNOVATION
ODIN delivers absolute metric depth from a single RGB image by pre-training on 20 million heterogeneous 3D data sources. Unlike relative depth estimators, ODIN outputs real-world measurements in meters.
DEPLOYMENT
- REST API
- Docker containerized
- Prometheus metrics
- Configurable backends
COMPUTE
- PRIMARY
- GPU-train, MLX inference
- EDGE
- Optimized inference
- API
- REST + streaming
PAPERS
- [01]MetricAnything (2601.22054)