Skip to content
RFL_GLOBAL
中文
  • DEPTH ANY CAMERA // CVPR 2025
  • CAMERA-AGNOSTIC
  • ZERO-SHOT

PANOPTES

THE ALL-SEEING

Depth estimation models are locked to specific camera types. Train on perspective? Doesn't work on fisheye. Train on smartphone? Fails on 360 panorama. PANOPTES predicts metric depth from any camera model — perspective, fisheye, equirectangular, omnidirectional — without camera-specific training. One model for all cameras.

MODULE STATUS: ACTIVE

Depth Accuracy (AbsRel)

ANY

DIVISION
ANIMA
WAVE
W4
DOMAIN
PERCEPTION
WAVE 4 // ANIMA SUITE
UNIVERSAL DEPTH ESTIMATION
PANOPTES // W4 // 058/079
01THE CHALLENGE

ONE CAMERA, ONE MODEL

Depth estimation models are locked to specific camera types. Train on a perspective camera? Doesn't work on a fisheye. Train on a smartphone? Fails on a 360 panorama. Every new camera geometry requires retraining.

Real robots use heterogeneous sensors — action cameras, ZED stereo, phone cameras, RealSense. There is no single depth model that works everywhere. PANOPTES provides that missing universal layer.

02THE SOLUTION

WHAT PANOPTES DELIVERS

PANOPTES is a camera-agnostic depth estimation system. Any RGB image from any camera, automatically adapted, metric depth + uncertainty output.

PIPELINE

  1. 01Zero-shot depth from any camera: perspective, fisheye, equirectangular, omnidirectional
  2. 02Camera-adaptive embeddings that encode camera type and field-of-view dynamically
  3. 03ViT backbone (DINOv2) → camera adaptive layer → DPT decoder → metric depth
  4. 04Metric depth range: 1mm to 100m with uncertainty estimates

CAPABILITIES

  • MODEL VARIANTSThree sizes: SMALL (50M), BASE (86M), LARGE (170M params)→ Scale from edge devices to data center
  • DEVICE SUPPORTMLX Apple Silicon (M1–M5, ~45ms on M5, 4x M1) + CUDA + CPU→ Deploy anywhere without code changes
  • REST API/depth, /depth/batch, /pointcloud, /cameras, /configure_camera→ Production-ready integration for any platform
03ENGINEERING

WHY THIS IS HARD

Every camera has a different projection geometry. PANOPTES must handle them all:

  1. 01Perspective: pinhole model. Fisheye: non-linear radial distortion (equidistant, equisolid, stereographic)
  2. 02Equirectangular: spherical unwrapping for 360 cameras. Omnidirectional: multiple overlapping perspectives
  3. 03Camera-adaptive embeddings must encode type and FOV, then dynamically adjust feature extraction
  4. 04Log-compressed depth encoding preserves scale from 1mm to 100m across diverse sensor types
  5. 05Inference must remain fast across all camera types without per-camera fine-tuning

PANOPTES learns camera-adaptive embeddings that encode the camera type and field-of-view, dynamically adjusting feature extraction during inference. One model, any camera.

04BENCHMARKS

REAL HARDWARE PERFORMANCE

Measured with ViT+DPT camera-adaptive pipeline:

REAL HARDWARE PERFORMANCE
METRICVALUE
Depth Accuracy (AbsRel)NYU 0.098 / KITTI 0.089 / ScanNet 0.108
Inference Speed (A100)45–180ms (512×512 to 1024×1024)
Apple M5 (MLX)~45ms at 512×512 (4x M1)
Throughput (BASE)5.5–22 fps (512×512)
Memory200–700MB (models), 500–2000MB (runtime)
Depth Range1mm to 100m (metric scale)
Model VariantsSMALL (50M), BASE (86M), LARGE (170M)
Camera TypesPerspective, Fisheye, Equirectangular, Omnidirectional
05BUILD STATUS

WHAT'S BUILT TODAY

9/11 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETEViT+DPT camera-adaptive pipeline, production-ready weights
API layerIN PROGRESSPublic API, pending dataset infrastructure
Camera Adaptive ModuleCOMPLETEType embedding, FOV projection, feature fusion
ViT BackboneCOMPLETEDINOv2-based visual encoder
DPT DecoderCOMPLETECoarse-to-fine depth refinement
REST APICOMPLETE/depth, /depth/batch, /pointcloud, /cameras
Point Cloud GenerationCOMPLETE3D export from depth + camera params
Docker DeploymentCOMPLETEFull stack with Prometheus/Grafana
Device SupportCOMPLETEAuto (MLX > MPS > CUDA > CPU)
Prometheus MetricsCOMPLETEInference count, latency, error tracking
gRPC InterfacePLANNEDNot implemented; REST ready
06APPLICATIONS

WHERE PANOPTES DEPLOYS

  • APP_01

    ROBOT PERCEPTION

    Single model for multi-camera robotic platforms. No per-sensor retraining.

  • APP_02

    MOBILE ROBOTICS

    Deploy on phones, action cameras, USB cams without retraining.

  • APP_03

    SURVEILLANCE

    360 camera depth for anomaly detection and spatial awareness.

  • APP_04

    AR/VR PLATFORMS

    Fast metric depth from heterogeneous phone/headset cameras.

  • APP_05

    EMBEDDED ROBOTICS

    Lightweight depth on edge devices with small model variants.

  • APP_06

    AUTONOMOUS VEHICLES

    Real-time depth for obstacle detection and path planning.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: DEPTH ANY CAMERA

  • Camera-agnostic depth estimation (CVPR 2025, Guo et al.)
  • Vision Transformer (ViT) backbone with DINOv2 pretraining
  • Camera-adaptive feature fusion with learned type embeddings
  • Dense Prediction Transformer (DPT) for coarse-to-fine depth

PANOPTES IMPLEMENTATION

  • FastAPI REST server with multi-camera input schema
  • Three model sizes: SMALL (50M) BASE (86M) LARGE (170M)
  • Log-compressed depth encoding (1mm–100m)
  • Prometheus metrics and Docker deployment

INTEGRATION POINTS

  • Feeds metric depth to ABYSSOS (monocular depth baseline)
  • Provides camera-agnostic features to PETRA (depth foundation model)
  • Outputs point clouds for downstream 3D reconstruction (PLEROMA)
08PAPERS

RESEARCH BASIS

  1. [01]Depth Any Camera: Zero-Shot Metric Depth Estimation — CVPR 2025