- CVPR 2025 // PROMPTDA
- 4K DENSE DEPTH
- MULTI-SENSOR FUSION
ABYSSOS
THE BOTTOMLESS DEEP
Robots see in RGB. They manipulate in 3D. Every grasp, every reach, every collision check demands depth. ABYSSOS delivers production-ready metric depth from any camera and any sparse sensor — LiDAR, RealSense, ARKit — fused into one clean 4K depth map.
MODULE STATUS: ACTIVERTX 4090
4K
- DIVISION
- ANIMA
- WAVE
- W1
- DOMAIN
- PERCEPTION
- WAVE 2 // ANIMA SUITE
- METRIC DEPTH ESTIMATION
DEPTH SENSORS FAIL IN THE REAL WORLD
LiDAR is sparse. RealSense drifts in sunlight. iPhone depth is noisy at range. Stereo is slow. Every depth sensor has a failure mode that makes it unreliable for production robotics. Vision-only depth estimation exists, but generic models trained on internet images fail catastrophically on real robot hands, gripper tools, and cluttered warehouse bins.
Robotics needs a depth model that learns from sparse physical measurements and fuses multiple sensor types into one clean, metric-accurate depth map. Not approximate depth. Not relative depth. Real-world metric depth at 4K resolution, running at production speed.
WHAT ABYSSOS DELIVERS
ABYSSOS is a production-ready depth estimation service built on PromptDA (CVPR 2025). It takes an RGB image plus optional sparse depth from any sensor and outputs dense 4K metric depth in 100-300ms on GPU.
CAPABILITIES
- Dense 4K depth (3840×2160) from RGB + sparse prompts (100-1000 points)
- ViT-S/B/L model variants (25MB, 90MB, 350MB)
- Sparse fusion adapters: LiDAR, RealSense, ARKit, generic depth
- Temporal consistency for video sequences with optical flow
- Multi-device: MLX (Apple Silicon M1–M5), CUDA (NVIDIA), CPU fallback
- Batch processing, per-image metrics, calibration endpoints
WHY THIS IS HARD
PromptDA conditions depth on sparse point prompts rather than requiring dense ground truth. Perfect for robotics — but deployment is non-trivial:
- 01Sensor calibration varies per camera (focal length, principal point)
- 02Sparse depth adapters must normalize LiDAR, RealSense, ARKit to a common format
- 03Temporal consistency requires optical flow estimation (additional compute)
- 04ViT models need careful quantization for real-time throughput
- 05Production observability: metrics, health checks, batch queueing
ABYSSOS abstracts all of this. Plug in RGB + calibration, get clean depth.
REAL HARDWARE PERFORMANCE
Measured latency and memory on production devices:
| DEVICE | MODEL | RESOLUTION | LATENCY | MEMORY | USE CASE |
|---|---|---|---|---|---|
| RTX 4090 | ViT-S | 3840×2160 | 100-150ms | ~4GB | Production target |
| RTX 3080 | ViT-S | 3840×2160 | 200-300ms | ~6GB | High-end consumer |
| MacBook M5 | ViT-S | 3840×2160 | 500ms-1.2s | ~2GB | Mobile production (4× M1) |
| MacBook M3 | ViT-S | 3840×2160 | 2-5s | ~2GB | Development |
| Intel i7 | ViT-S | 1920×1080 | 10-20s | ~3GB | Validation only |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | PromptDA ViT-S/B/L trained and validated — production-ready weights |
| Sensor adapters | COMPLETE | LiDAR, RealSense, ARKit, generic |
| Calibration system | COMPLETE | Per-camera intrinsic matrix |
| Device abstraction | COMPLETE | MLX, CUDA, CPU auto-detection |
| Server framework | COMPLETE | FastAPI + gRPC dual endpoints |
| Docker build | COMPLETE | CPU, GPU, dev profiles |
| Configuration | COMPLETE | .env + TOML support |
| /depth endpoint | IN PROGRESS | PromptDA inference pipeline |
| /depth/batch | IN PROGRESS | Queue-based batch processing |
| /depth/video | IN PROGRESS | Temporal consistency engine |
| API layer | IN PROGRESS | REST + gRPC public API — pending dataset infrastructure |
| Optical flow | PLANNED | RAFT or CoTracker integration |
WHERE ABYSSOS DEPLOYS
- APP_01
MOBILE MANIPULATION
Autonomous arms and grippers needing clean metric depth for bin picking, palletizing, and assembly.
- APP_02
MOBILE ROBOTICS
Navigation and obstacle avoidance with sensor fusion from multiple depth sources.
- APP_03
PERCEPTION PIPELINES
Computer vision integrators building production systems requiring calibrated depth.
- APP_04
AR / VR HARDWARE
Depth for hand tracking, scene understanding, and spatial mapping on mobile devices.
- APP_05
AUTONOMOUS VEHICLES
Metric depth fusion from camera + LiDAR for obstacle detection and path planning.
- APP_06
WAREHOUSE AUTOMATION
High-speed depth processing for item detection and pose estimation at scale.
UNDER THE HOOD
FOUNDATION: PROMPTDA (CVPR 2025)
- Vision Transformer (ViT) backbone pre-trained on large-scale depth data
- Sparse prompt conditioning — no dense ground truth required
- Apache-2.0 licensed — safe for commercial robotics
KEY INNOVATION
Fuses sparse depth sensors (LiDAR, RealSense, ARKit) into ViT embeddings during inference — no retraining required.
DEPLOYMENT STACK
- Per-image concurrency with device queue management
- Prometheus metrics for latency, throughput, memory
- Structured JSON logging via structlog
- Sensor calibration API for multi-camera systems
MULTI-DEVICE RUNTIME
- MLX
- Apple Silicon (M1/M2/M3/M4/M5) — ~2-5s on M3, ~0.5-1.2s on M5 (4× M1)
- CUDA
- NVIDIA GPUs (RTX 3080+, A100) — 100-300ms for 4K
- CPU
- Fallback for validation — 10-20s for 4K on i7
MODEL VARIANTS
- ViT-S — 25MB — 100-150ms on RTX 4090
- ViT-B — 90MB — 150-200ms on RTX 4090
- ViT-L — 350MB — 250-350ms on RTX 4090