- DEFM // ETH ZÜRICH 2026
- 11 ARCHITECTURES
- SELF-SUPERVISED
PETRA
THE FOUNDATION STONE
Depth images contain rich 3D geometry, but most pipelines ignore this signal. PETRA is a depth foundation model (DeFM) trained self-supervised on 60M depth images using the DINO objective. 11 variants from 3M to 307M parameters — from ultra-lightweight edge robots to research-grade distillation sources. Zero-shot transfer delivers 87–93% scene classification accuracy.
MODULE STATUS: ACTIVEZero-Shot Accuracy
60M
- DIVISION
- ANIMA
- WAVE
- W4
- DOMAIN
- FOUNDATION
- WAVE 4 // ANIMA SUITE
- DEPTH FOUNDATION MODEL
DEPTH IS IGNORED
Depth images contain rich 3D geometry, but most robotics pipelines ignore this signal. RGB foundation models dominate because they're pre-trained on billions of internet images. Depth gets relegated to post-hoc refinement.
We need a depth foundation model trained on 60 million depth images, with 11 variants scaling from 3M to 307M parameters, enabling every robot stack to extract rich geometric features from raw depth. PETRA provides that foundation.
WHAT PETRA DELIVERS
PETRA is a depth foundation model with 11 backbone architectures. Choose your backbone, fine-tune for your task, or use zero-shot features.
PIPELINE
- 0111 backbone architectures: ViT-S/B/L, EfficientNet-B0/B3, ResNet50, RegNet, ConvNeXt
- 02Self-supervised DINO training on 60M depth images (sim + real)
- 03Zero-shot scene classification: 87–93% accuracy
- 04Grasp success via linear probe: 78–85%
- 05Navigation in sim: 89–94% success
- 06Fine-tune output head: 5–10% of parameters for your task
- 07MLX Apple Silicon (M1–M5, ~12ms on M5 for ViT-S, 4× M1) + CUDA + CPU
- 08REST API: /features, /features/batch, /adapt, /classify, /segment
CAPABILITIES
- ZERO-SHOT TRANSFERScene classification 87–93% without fine-tuning→ Out-of-box features for rapid deployment
- SCALABLE BACKBONES11 architectures from 3M to 307M parameters→ Scale from edge robots to research distillation
- CROSS-PLATFORMMLX, CUDA, MPS, and CPU with REST API→ Deploy anywhere — Apple Silicon to cloud GPU
WHY THIS IS HARD
Self-supervised learning on depth requires rethinking augmentation and representation:
- 01Depth augmentation: random crops need careful handling of scale and occlusion boundaries, not just RGB-style crops
- 02Metric-aware encoding: log-compress depth to preserve scale from 1mm to 100m in 3 channels (log, linear, inverse)
- 03Scale diversity: train on sim (Isaac, MuJoCo) and real (quadrupeds, humanoids, mobile manipulators)
- 04Student-teacher DINO training with momentum updates requires careful hyperparameter tuning for depth
- 0511 architecture variants means maintaining consistent training pipelines across very different backbones
PETRA achieves this with student-teacher DINO training, log-compressed depth encoding, and 60M diverse depth images from simulators and real robots.
REAL HARDWARE PERFORMANCE
Measured across 11 DeFM backbone architectures:
| METRIC | VALUE |
|---|---|
| Zero-Shot Accuracy | 87–93% (scene classification) |
| Grasp Success (Linear Probe) | 78–85% (ViT-S to ViT-L) |
| Navigation Success | 89–94% simulated environments |
| ViT-S Latency | 12.5ms (batch 1) |
| ViT-S Throughput | 390 fps (batch 32) |
| RegNet-400MF Latency | 4.2ms (batch 1) |
| Apple M5 (MLX) | ~12ms ViT-S (4× M1) |
| Training Data | 60 million depth images |
| Model Range | 3M to 307M parameters |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | 11 DeFM architectures with production-ready weights |
| API layer | IN PROGRESS | Public API, pending dataset infrastructure |
| 11 Backbone Architectures | COMPLETE | ViT-S/B/L, EfficientNet, ResNet, RegNet, ConvNeXt |
| Log-Compressed Encoding | COMPLETE | 3-channel depth representation |
| FastAPI Server | COMPLETE | /features, /features/batch, /adapt, /classify, /segment |
| Model Loading | COMPLETE | HuggingFace integration, local caching |
| Task Adaptation | COMPLETE | Lightweight downstream heads |
| Zero-Shot Transfer | COMPLETE | Linear probing on scene/obstacle tasks |
| Docker Deployment | COMPLETE | Multi-container with Prometheus/Grafana |
| Device Support | COMPLETE | Auto (MLX > MPS > CUDA > CPU) |
| gRPC Interface | PLANNED | Architecture designed, not implemented |
WHERE PETRA DEPLOYS
- APP_01
ROBOT PERCEPTION
Drop-in depth feature backbone for vision stacks
- APP_02
NAVIGATION
Geometry-aware features for terrain traversability
- APP_03
GRASPING
Grasp point prediction from depth features
- APP_04
MOBILE ROBOTS
Low-latency depth on edge (RegNet variants)
- APP_05
RESEARCH LABS
Benchmark for depth self-supervised learning
- APP_06
SIM-TO-REAL
Pre-trained features that transfer to real depth sensors
UNDER THE HOOD
FOUNDATION: DEFM (ETH ZÜRICH 2026)
- Self-supervised DINO objective on student-teacher framework
- Log-compressed depth encoding: 3-channel (1mm–100m)
- 60 million depth images from simulators + real robots
- Augmentation: random crops + depth noise + lens distortion + temporal shifts
BACKBONE SPECIFICATIONS
- RegNet-Y-400MF
- 3.2M params / 440 dim — edge robots
- ViT-S
- 21M params / 384 dim — standard tasks
- ViT-B
- 87M params / 768 dim — high-performance
- ViT-L
- 307M params / 1024 dim — research/distillation
INTEGRATION POINTS
- Provides depth features to ABYSSOS (monocular depth)
- Feeds foundation representations to PANOPTES (camera-agnostic)
- Supplies geometric embeddings for downstream manipulation modules
RESEARCH BASIS
- [01]DeFM: Learning Foundation Representations from Depth for Robotics — ETH Zürich 2026