Skip to content
RFL_GLOBAL
中文
  • CVPR 2025 // PROMPTDA
  • 4K DENSE DEPTH
  • MULTI-SENSOR FUSION

ABYSSOS

THE BOTTOMLESS DEEP

Robots see in RGB. They manipulate in 3D. Every grasp, every reach, every collision check demands depth. ABYSSOS delivers production-ready metric depth from any camera and any sparse sensor — LiDAR, RealSense, ARKit — fused into one clean 4K depth map.

MODULE STATUS: ACTIVE

RTX 4090

4K

DIVISION
ANIMA
WAVE
W1
DOMAIN
PERCEPTION
WAVE 2 // ANIMA SUITE
METRIC DEPTH ESTIMATION
ABYSSOS // W1 // 001/079
01THE CHALLENGE

DEPTH SENSORS FAIL IN THE REAL WORLD

LiDAR is sparse. RealSense drifts in sunlight. iPhone depth is noisy at range. Stereo is slow. Every depth sensor has a failure mode that makes it unreliable for production robotics. Vision-only depth estimation exists, but generic models trained on internet images fail catastrophically on real robot hands, gripper tools, and cluttered warehouse bins.

Robotics needs a depth model that learns from sparse physical measurements and fuses multiple sensor types into one clean, metric-accurate depth map. Not approximate depth. Not relative depth. Real-world metric depth at 4K resolution, running at production speed.

02THE SOLUTION

WHAT ABYSSOS DELIVERS

ABYSSOS is a production-ready depth estimation service built on PromptDA (CVPR 2025). It takes an RGB image plus optional sparse depth from any sensor and outputs dense 4K metric depth in 100-300ms on GPU.

CAPABILITIES

  • Dense 4K depth (3840×2160) from RGB + sparse prompts (100-1000 points)
  • ViT-S/B/L model variants (25MB, 90MB, 350MB)
  • Sparse fusion adapters: LiDAR, RealSense, ARKit, generic depth
  • Temporal consistency for video sequences with optical flow
  • Multi-device: MLX (Apple Silicon M1–M5), CUDA (NVIDIA), CPU fallback
  • Batch processing, per-image metrics, calibration endpoints
03ENGINEERING

WHY THIS IS HARD

PromptDA conditions depth on sparse point prompts rather than requiring dense ground truth. Perfect for robotics — but deployment is non-trivial:

  1. 01Sensor calibration varies per camera (focal length, principal point)
  2. 02Sparse depth adapters must normalize LiDAR, RealSense, ARKit to a common format
  3. 03Temporal consistency requires optical flow estimation (additional compute)
  4. 04ViT models need careful quantization for real-time throughput
  5. 05Production observability: metrics, health checks, batch queueing

ABYSSOS abstracts all of this. Plug in RGB + calibration, get clean depth.

04BENCHMARKS

REAL HARDWARE PERFORMANCE

Measured latency and memory on production devices:

REAL HARDWARE PERFORMANCE
DEVICEMODELRESOLUTIONLATENCYMEMORYUSE CASE
RTX 4090ViT-S3840×2160100-150ms~4GBProduction target
RTX 3080ViT-S3840×2160200-300ms~6GBHigh-end consumer
MacBook M5ViT-S3840×2160500ms-1.2s~2GBMobile production (4× M1)
MacBook M3ViT-S3840×21602-5s~2GBDevelopment
Intel i7ViT-S1920×108010-20s~3GBValidation only
05BUILD STATUS

WHAT'S BUILT TODAY

7/12 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETEPromptDA ViT-S/B/L trained and validated — production-ready weights
Sensor adaptersCOMPLETELiDAR, RealSense, ARKit, generic
Calibration systemCOMPLETEPer-camera intrinsic matrix
Device abstractionCOMPLETEMLX, CUDA, CPU auto-detection
Server frameworkCOMPLETEFastAPI + gRPC dual endpoints
Docker buildCOMPLETECPU, GPU, dev profiles
ConfigurationCOMPLETE.env + TOML support
/depth endpointIN PROGRESSPromptDA inference pipeline
/depth/batchIN PROGRESSQueue-based batch processing
/depth/videoIN PROGRESSTemporal consistency engine
API layerIN PROGRESSREST + gRPC public API — pending dataset infrastructure
Optical flowPLANNEDRAFT or CoTracker integration
06APPLICATIONS

WHERE ABYSSOS DEPLOYS

  • APP_01

    MOBILE MANIPULATION

    Autonomous arms and grippers needing clean metric depth for bin picking, palletizing, and assembly.

  • APP_02

    MOBILE ROBOTICS

    Navigation and obstacle avoidance with sensor fusion from multiple depth sources.

  • APP_03

    PERCEPTION PIPELINES

    Computer vision integrators building production systems requiring calibrated depth.

  • APP_04

    AR / VR HARDWARE

    Depth for hand tracking, scene understanding, and spatial mapping on mobile devices.

  • APP_05

    AUTONOMOUS VEHICLES

    Metric depth fusion from camera + LiDAR for obstacle detection and path planning.

  • APP_06

    WAREHOUSE AUTOMATION

    High-speed depth processing for item detection and pose estimation at scale.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: PROMPTDA (CVPR 2025)

  • Vision Transformer (ViT) backbone pre-trained on large-scale depth data
  • Sparse prompt conditioning — no dense ground truth required
  • Apache-2.0 licensed — safe for commercial robotics

KEY INNOVATION

Fuses sparse depth sensors (LiDAR, RealSense, ARKit) into ViT embeddings during inference — no retraining required.

DEPLOYMENT STACK

  • Per-image concurrency with device queue management
  • Prometheus metrics for latency, throughput, memory
  • Structured JSON logging via structlog
  • Sensor calibration API for multi-camera systems

MULTI-DEVICE RUNTIME

MLX
Apple Silicon (M1/M2/M3/M4/M5) — ~2-5s on M3, ~0.5-1.2s on M5 (4× M1)
CUDA
NVIDIA GPUs (RTX 3080+, A100) — 100-300ms for 4K
CPU
Fallback for validation — 10-20s for 4K on i7

MODEL VARIANTS

  • ViT-S — 25MB — 100-150ms on RTX 4090
  • ViT-B — 90MB — 150-200ms on RTX 4090
  • ViT-L — 350MB — 250-350ms on RTX 4090
08PAPERS

RESEARCH PAPER

  1. [01]PromptDA: Prompt-Driven Depth Anything — Depth Anything Team, CVPR 2025