Skip to content
RFL_GLOBAL
中文
  • CVPR 2024 // BEST PAPER
  • 6DOF // ANY OBJECT
  • ZERO-SHOT POSE

ERGON

THE PERFECT GRASP

Grasping requires knowing exactly where an object is in 3D space. FoundationPose solves this: estimate 6DoF pose from a single reference image, no CAD model needed. Works on novel objects immediately. No retraining. CVPR 2024 Best Paper, deployed for production robotics.

MODULE STATUS: ACTIVE

Inference Speed (RTX 4090)

±2mm

DIVISION
ANIMA
WAVE
W2
DOMAIN
UNDERSTANDING
WAVE 4 // ANIMA SUITE
6DOF POSE ESTIMATION
ERGON // W2 // 019/079
01THE CHALLENGE

NOVEL OBJECTS BREAK TRADITIONAL POSE ESTIMATION

Grasping requires knowing exactly where an object is in 3D space — its position and orientation. For known objects with CAD models, pose estimation works well. But real warehouses have novel, unlabeled items. Most systems fail on objects they've never seen before.

You can't pre-register every SKU in a logistics center. You can't retrain a model for each new product. You need pose estimation that works on anything, from a single reference image, in real-time. FoundationPose (CVPR 2024 Best Paper) is the first system to deliver this.

02THE SOLUTION

WHAT ERGON DELIVERS

ERGON deploys FoundationPose — unified 6DoF pose estimation and tracking for any object. Three input modes, one output: full SE(3) pose with confidence.

PIPELINE

  1. 01CAD MODELProvide 3D model geometry → precise pose refinement with render-and-compare
  2. 02REFERENCE IMAGESingle image of object → zero-shot pose matching, no training required
  3. 03RGB-D FUSIONColor + depth sensor → high-accuracy metric pose with geometric features

CAPABILITIES

  • Full 6DoF pose output: SE(3) rotation + translation + confidence score
  • Novel object support: zero-shot from reference image, no CAD needed
  • Real-time: 40–50ms per frame on RTX 4090, ~100ms on Apple M5 (MLX)
  • Kalman filter tracking with occlusion recovery and re-detection
  • Up to 20 simultaneous tracked objects with batch processing
  • 104 tests, 70% coverage — production-grade test suite
03ENGINEERING

WHY THIS IS HARD

6DoF pose estimation requires solving three coupled problems simultaneously:

  1. 01Feature extraction: RGB features must be view-invariant yet fine-grained enough to distinguish subtle poses
  2. 02Pose hypothesis sampling: generate candidate SO(3) rotations efficiently without exhaustive search
  3. 03Iterative refinement: compare rendered hypotheses to observations (photometric + geometric losses) and refine
  4. 04Novel object generalization: match features from a single reference image to arbitrary viewpoints
  5. 05Temporal tracking: maintain pose consistency across frames with Kalman filtering and occlusion handling

FoundationPose achieves this with CNN feature extraction, depth-based hypothesis generation, and differentiable rendering for 5-iteration refinement. For novel objects, reference image matching replaces CAD models entirely.

04BENCHMARKS

REAL HARDWARE PERFORMANCE

Measured with FoundationPose on real sensor data:

REAL HARDWARE PERFORMANCE
METRICVALUE
Inference Speed (RTX 4090)40–50ms per frame (RGB-D)
RGB-Only Latency50–60ms per frame
Apple M5 (MLX)~100ms per frame (4× M1)
Translation Accuracy (known)±2–5mm
Translation Accuracy (novel)±10–20mm
Rotation Accuracy±2–5 degrees
Memory Usage2.5 GB model, 4–6 GB peak
Max Tracked Objects20 (configurable)
Refinement Iterations5 (default, up to 10)
Occlusion HandlingKalman filtering + auto re-detection
05BUILD STATUS

WHAT'S BUILT TODAY

10/13 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETEFoundationPose weights — production-ready, CVPR 2024 Best Paper
RGB-D Pose EstimationCOMPLETEIterative refinement with render-and-compare
RGB-Only Reference MatchingCOMPLETEZero-shot pose from reference image
CAD Model RegistrationCOMPLETEOBJ/STL/PLY model loading and geometry caching
Kalman Filter TrackingCOMPLETETemporal consistency with occlusion recovery
Object RegistryCOMPLETEPersistent JSON storage, auto-load on startup
FastAPI REST APICOMPLETE/estimate, /track, /register_object, /objects
Device SupportCOMPLETEAuto (MPS > CUDA > CPU) with MLX ready
Docker DeploymentCOMPLETEMulti-container with Prometheus/Grafana
Test SuiteCOMPLETE104 tests, 70% coverage, ruff linting
API layerIN PROGRESSPublic REST + gRPC API — pending dataset infrastructure
ROS2 IntegrationPLANNEDPose topic publishing (interface defined)
ZED 2i IntegrationPLANNEDDeferred until hardware arrival
06APPLICATIONS

WHERE ERGON DEPLOYS

  • APP_01

    WAREHOUSE AUTOMATION

    Pose estimation for novel objects — no CAD models needed. Handle unlabeled packages in real-time.

  • APP_02

    PICK-AND-PLACE ROBOTICS

    Fast 6DoF for grasping on conveyor lines. 40ms latency enables closed-loop control.

  • APP_03

    VISION-GUIDED ASSEMBLY

    Precise object localization for assembly tasks. ±2mm accuracy on known parts.

  • APP_04

    LOGISTICS & FULFILLMENT

    Handle arbitrary SKUs with zero-shot reference matching. No retraining per product.

  • APP_05

    MANUFACTURING BIN PICKING

    Novel part geometries handled immediately. Kalman tracking through occlusion.

  • APP_06

    RESEARCH & BENCHMARKING

    FoundationPose baseline for 6DoF pose estimation on BOP, YCB, and custom datasets.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: FOUNDATIONPOSE (CVPR 2024)

  • NVIDIA Research — CVPR 2024 Best Paper Award
  • ResNet-50 / EfficientNet backbone for feature extraction
  • Depth-encoded geometric features for pose hypothesis generation
  • Differentiable render-and-compare with 5-iteration refinement
  • Reference matching: feature similarity for novel objects without CAD

INPUT MODALITIES

RGB-D (Preferred)
Color + depth sensor fusion — most accurate, ±2–5mm translation
RGB-Only
Reference image matching — relative scale, zero-shot, no training
Depth-Only
Geometric reconstruction — works in low-light conditions

INFERENCE PIPELINE

  • Feature extraction → depth encoding → pose hypothesis sampling (SO(3))
  • Render candidate poses → photometric + geometric loss → refine (5 iterations)
  • Kalman filter for temporal tracking and occlusion recovery
  • Object registry with persistent storage and auto-load

MULTI-DEVICE RUNTIME

CUDA
NVIDIA GPUs (RTX 4090, A100) — full FoundationPose Torch, 40–50ms
MLX
Apple Silicon (M1–M5) — bootstrap inference, ~100ms on M5 (4× M1)
MPS
Apple Metal Performance Shaders — GPU-accelerated fallback
CPU
Fallback — demo/development only
08PAPERS

RESEARCH PAPER

  1. [01]FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects — Wen et al., CVPR 2024 (Best Paper)