Skip to content
RFL_GLOBAL
中文
  • DEFM // ETH ZÜRICH 2026
  • 11 ARCHITECTURES
  • SELF-SUPERVISED

PETRA

THE FOUNDATION STONE

Depth images contain rich 3D geometry, but most pipelines ignore this signal. PETRA is a depth foundation model (DeFM) trained self-supervised on 60M depth images using the DINO objective. 11 variants from 3M to 307M parameters — from ultra-lightweight edge robots to research-grade distillation sources. Zero-shot transfer delivers 87–93% scene classification accuracy.

MODULE STATUS: ACTIVE

Zero-Shot Accuracy

60M

DIVISION
ANIMA
WAVE
W4
DOMAIN
FOUNDATION
WAVE 4 // ANIMA SUITE
DEPTH FOUNDATION MODEL
PETRA // W4 // 059/079
01THE CHALLENGE

DEPTH IS IGNORED

Depth images contain rich 3D geometry, but most robotics pipelines ignore this signal. RGB foundation models dominate because they're pre-trained on billions of internet images. Depth gets relegated to post-hoc refinement.

We need a depth foundation model trained on 60 million depth images, with 11 variants scaling from 3M to 307M parameters, enabling every robot stack to extract rich geometric features from raw depth. PETRA provides that foundation.

02THE SOLUTION

WHAT PETRA DELIVERS

PETRA is a depth foundation model with 11 backbone architectures. Choose your backbone, fine-tune for your task, or use zero-shot features.

PIPELINE

  1. 0111 backbone architectures: ViT-S/B/L, EfficientNet-B0/B3, ResNet50, RegNet, ConvNeXt
  2. 02Self-supervised DINO training on 60M depth images (sim + real)
  3. 03Zero-shot scene classification: 87–93% accuracy
  4. 04Grasp success via linear probe: 78–85%
  5. 05Navigation in sim: 89–94% success
  6. 06Fine-tune output head: 5–10% of parameters for your task
  7. 07MLX Apple Silicon (M1–M5, ~12ms on M5 for ViT-S, 4× M1) + CUDA + CPU
  8. 08REST API: /features, /features/batch, /adapt, /classify, /segment

CAPABILITIES

  • ZERO-SHOT TRANSFERScene classification 87–93% without fine-tuning→ Out-of-box features for rapid deployment
  • SCALABLE BACKBONES11 architectures from 3M to 307M parameters→ Scale from edge robots to research distillation
  • CROSS-PLATFORMMLX, CUDA, MPS, and CPU with REST API→ Deploy anywhere — Apple Silicon to cloud GPU
03ENGINEERING

WHY THIS IS HARD

Self-supervised learning on depth requires rethinking augmentation and representation:

  1. 01Depth augmentation: random crops need careful handling of scale and occlusion boundaries, not just RGB-style crops
  2. 02Metric-aware encoding: log-compress depth to preserve scale from 1mm to 100m in 3 channels (log, linear, inverse)
  3. 03Scale diversity: train on sim (Isaac, MuJoCo) and real (quadrupeds, humanoids, mobile manipulators)
  4. 04Student-teacher DINO training with momentum updates requires careful hyperparameter tuning for depth
  5. 0511 architecture variants means maintaining consistent training pipelines across very different backbones

PETRA achieves this with student-teacher DINO training, log-compressed depth encoding, and 60M diverse depth images from simulators and real robots.

04BENCHMARKS

REAL HARDWARE PERFORMANCE

Measured across 11 DeFM backbone architectures:

REAL HARDWARE PERFORMANCE
METRICVALUE
Zero-Shot Accuracy87–93% (scene classification)
Grasp Success (Linear Probe)78–85% (ViT-S to ViT-L)
Navigation Success89–94% simulated environments
ViT-S Latency12.5ms (batch 1)
ViT-S Throughput390 fps (batch 32)
RegNet-400MF Latency4.2ms (batch 1)
Apple M5 (MLX)~12ms ViT-S (4× M1)
Training Data60 million depth images
Model Range3M to 307M parameters
05BUILD STATUS

WHAT'S BUILT TODAY

9/11 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETE11 DeFM architectures with production-ready weights
API layerIN PROGRESSPublic API, pending dataset infrastructure
11 Backbone ArchitecturesCOMPLETEViT-S/B/L, EfficientNet, ResNet, RegNet, ConvNeXt
Log-Compressed EncodingCOMPLETE3-channel depth representation
FastAPI ServerCOMPLETE/features, /features/batch, /adapt, /classify, /segment
Model LoadingCOMPLETEHuggingFace integration, local caching
Task AdaptationCOMPLETELightweight downstream heads
Zero-Shot TransferCOMPLETELinear probing on scene/obstacle tasks
Docker DeploymentCOMPLETEMulti-container with Prometheus/Grafana
Device SupportCOMPLETEAuto (MLX > MPS > CUDA > CPU)
gRPC InterfacePLANNEDArchitecture designed, not implemented
06APPLICATIONS

WHERE PETRA DEPLOYS

  • APP_01

    ROBOT PERCEPTION

    Drop-in depth feature backbone for vision stacks

  • APP_02

    NAVIGATION

    Geometry-aware features for terrain traversability

  • APP_03

    GRASPING

    Grasp point prediction from depth features

  • APP_04

    MOBILE ROBOTS

    Low-latency depth on edge (RegNet variants)

  • APP_05

    RESEARCH LABS

    Benchmark for depth self-supervised learning

  • APP_06

    SIM-TO-REAL

    Pre-trained features that transfer to real depth sensors

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: DEFM (ETH ZÜRICH 2026)

  • Self-supervised DINO objective on student-teacher framework
  • Log-compressed depth encoding: 3-channel (1mm–100m)
  • 60 million depth images from simulators + real robots
  • Augmentation: random crops + depth noise + lens distortion + temporal shifts

BACKBONE SPECIFICATIONS

RegNet-Y-400MF
3.2M params / 440 dim — edge robots
ViT-S
21M params / 384 dim — standard tasks
ViT-B
87M params / 768 dim — high-performance
ViT-L
307M params / 1024 dim — research/distillation

INTEGRATION POINTS

  • Provides depth features to ABYSSOS (monocular depth)
  • Feeds foundation representations to PANOPTES (camera-agnostic)
  • Supplies geometric embeddings for downstream manipulation modules
08PAPERS

RESEARCH BASIS

  1. [01]DeFM: Learning Foundation Representations from Depth for Robotics — ETH Zürich 2026