- DEPTH ANY CAMERA // CVPR 2025
- CAMERA-AGNOSTIC
- ZERO-SHOT
PANOPTES
THE ALL-SEEING
Depth estimation models are locked to specific camera types. Train on perspective? Doesn't work on fisheye. Train on smartphone? Fails on 360 panorama. PANOPTES predicts metric depth from any camera model — perspective, fisheye, equirectangular, omnidirectional — without camera-specific training. One model for all cameras.
MODULE STATUS: ACTIVEDepth Accuracy (AbsRel)
ANY
- DIVISION
- ANIMA
- WAVE
- W4
- DOMAIN
- PERCEPTION
- WAVE 4 // ANIMA SUITE
- UNIVERSAL DEPTH ESTIMATION
ONE CAMERA, ONE MODEL
Depth estimation models are locked to specific camera types. Train on a perspective camera? Doesn't work on a fisheye. Train on a smartphone? Fails on a 360 panorama. Every new camera geometry requires retraining.
Real robots use heterogeneous sensors — action cameras, ZED stereo, phone cameras, RealSense. There is no single depth model that works everywhere. PANOPTES provides that missing universal layer.
WHAT PANOPTES DELIVERS
PANOPTES is a camera-agnostic depth estimation system. Any RGB image from any camera, automatically adapted, metric depth + uncertainty output.
PIPELINE
- 01Zero-shot depth from any camera: perspective, fisheye, equirectangular, omnidirectional
- 02Camera-adaptive embeddings that encode camera type and field-of-view dynamically
- 03ViT backbone (DINOv2) → camera adaptive layer → DPT decoder → metric depth
- 04Metric depth range: 1mm to 100m with uncertainty estimates
CAPABILITIES
- MODEL VARIANTSThree sizes: SMALL (50M), BASE (86M), LARGE (170M params)→ Scale from edge devices to data center
- DEVICE SUPPORTMLX Apple Silicon (M1–M5, ~45ms on M5, 4x M1) + CUDA + CPU→ Deploy anywhere without code changes
- REST API/depth, /depth/batch, /pointcloud, /cameras, /configure_camera→ Production-ready integration for any platform
WHY THIS IS HARD
Every camera has a different projection geometry. PANOPTES must handle them all:
- 01Perspective: pinhole model. Fisheye: non-linear radial distortion (equidistant, equisolid, stereographic)
- 02Equirectangular: spherical unwrapping for 360 cameras. Omnidirectional: multiple overlapping perspectives
- 03Camera-adaptive embeddings must encode type and FOV, then dynamically adjust feature extraction
- 04Log-compressed depth encoding preserves scale from 1mm to 100m across diverse sensor types
- 05Inference must remain fast across all camera types without per-camera fine-tuning
PANOPTES learns camera-adaptive embeddings that encode the camera type and field-of-view, dynamically adjusting feature extraction during inference. One model, any camera.
REAL HARDWARE PERFORMANCE
Measured with ViT+DPT camera-adaptive pipeline:
| METRIC | VALUE |
|---|---|
| Depth Accuracy (AbsRel) | NYU 0.098 / KITTI 0.089 / ScanNet 0.108 |
| Inference Speed (A100) | 45–180ms (512×512 to 1024×1024) |
| Apple M5 (MLX) | ~45ms at 512×512 (4x M1) |
| Throughput (BASE) | 5.5–22 fps (512×512) |
| Memory | 200–700MB (models), 500–2000MB (runtime) |
| Depth Range | 1mm to 100m (metric scale) |
| Model Variants | SMALL (50M), BASE (86M), LARGE (170M) |
| Camera Types | Perspective, Fisheye, Equirectangular, Omnidirectional |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | ViT+DPT camera-adaptive pipeline, production-ready weights |
| API layer | IN PROGRESS | Public API, pending dataset infrastructure |
| Camera Adaptive Module | COMPLETE | Type embedding, FOV projection, feature fusion |
| ViT Backbone | COMPLETE | DINOv2-based visual encoder |
| DPT Decoder | COMPLETE | Coarse-to-fine depth refinement |
| REST API | COMPLETE | /depth, /depth/batch, /pointcloud, /cameras |
| Point Cloud Generation | COMPLETE | 3D export from depth + camera params |
| Docker Deployment | COMPLETE | Full stack with Prometheus/Grafana |
| Device Support | COMPLETE | Auto (MLX > MPS > CUDA > CPU) |
| Prometheus Metrics | COMPLETE | Inference count, latency, error tracking |
| gRPC Interface | PLANNED | Not implemented; REST ready |
WHERE PANOPTES DEPLOYS
- APP_01
ROBOT PERCEPTION
Single model for multi-camera robotic platforms. No per-sensor retraining.
- APP_02
MOBILE ROBOTICS
Deploy on phones, action cameras, USB cams without retraining.
- APP_03
SURVEILLANCE
360 camera depth for anomaly detection and spatial awareness.
- APP_04
AR/VR PLATFORMS
Fast metric depth from heterogeneous phone/headset cameras.
- APP_05
EMBEDDED ROBOTICS
Lightweight depth on edge devices with small model variants.
- APP_06
AUTONOMOUS VEHICLES
Real-time depth for obstacle detection and path planning.
UNDER THE HOOD
FOUNDATION: DEPTH ANY CAMERA
- Camera-agnostic depth estimation (CVPR 2025, Guo et al.)
- Vision Transformer (ViT) backbone with DINOv2 pretraining
- Camera-adaptive feature fusion with learned type embeddings
- Dense Prediction Transformer (DPT) for coarse-to-fine depth
PANOPTES IMPLEMENTATION
- FastAPI REST server with multi-camera input schema
- Three model sizes: SMALL (50M) BASE (86M) LARGE (170M)
- Log-compressed depth encoding (1mm–100m)
- Prometheus metrics and Docker deployment
INTEGRATION POINTS
- Feeds metric depth to ABYSSOS (monocular depth baseline)
- Provides camera-agnostic features to PETRA (depth foundation model)
- Outputs point clouds for downstream 3D reconstruction (PLEROMA)
RESEARCH BASIS
- [01]Depth Any Camera: Zero-Shot Metric Depth Estimation — CVPR 2025