Skip to content
RFL_GLOBAL
中文
  • 90.48% // TOP-1
  • VPR // DINOV2
  • <50MS INFERENCE

LOCI

VISUAL MEMORY FOR ROBOTS

Robots need to know where they are. LOCI provides visual place recognition — match camera frames to a stored memory of places. DINOv2 embeddings, sub-50ms inference, 90.48% top-1 accuracy. Loop closure, relocalization, fleet memory.

MODULE STATUS: ACTIVE

Top-1 Accuracy

90.5%

DIVISION
ANIMA
WAVE
W5
DOMAIN
PERCEPTION
WAVE 5 // ANIMA SUITE
VISUAL PLACE RECOGNITION
LOCI // W5 // 041/079
01THE CHALLENGE

ROBOTS LOSE THEIR PLACE

Mobile robots navigate through changing environments. Odometry drifts. GPS fails indoors. Without knowing "where am I?", mapping, planning, and coordination break down. Traditional place recognition relied on hand-crafted features — brittle, slow, and prone to failure under viewpoint or lighting changes.

You need a system that recognizes places from a single image, matches against a growing memory, and runs in real-time on robot hardware. DINOv2 and modern VPR methods deliver this — but deployment requires careful engineering.

02THE SOLUTION

WHAT LOCI DELIVERS

LOCI builds a visual memory of places. Each camera frame is encoded into a fingerprint; the system searches the memory grid for matches. Loop closure, relocalization, and fleet-wide place sharing.

PIPELINE

  1. 01Frame capture — RGB image from robot camera
  2. 02DINOv2 embedding — extract place fingerprint (768-dim vector)
  3. 03Memory search — compare against stored place slots, nearest-neighbor
  4. 04Match or store — return place ID if match > threshold, else add new slot

CAPABILITIES

  • LOOP CLOSUREDetect when robot returns to a previously visited place→ Corrects odometry drift, enables consistent mapping
  • RELOCALIZATIONRecover pose after kidnapping or sensor dropout→ Robust to failures, fast recovery
  • FLEET MEMORYShared place database across multiple robots→ One robot explores, all robots recognize
03ENGINEERING

WHY THIS IS HARD

Visual place recognition requires:

  1. 01Viewpoint invariance: same place from different angles must match
  2. 02Lighting robustness: day/night, shadows, seasonal changes
  3. 03Real-time inference: <50ms per frame on embedded hardware
  4. 04Scalable memory: thousands of places without linear search blowup
  5. 05Temporal decay: old places fade to avoid stale matches

LOCI uses DINOv2 for robust embeddings, approximate nearest-neighbor for fast search, and configurable decay for long-term operation.

04BENCHMARKS

REAL HARDWARE PERFORMANCE

Measured with DINOv2 + VPR pipeline:

REAL HARDWARE PERFORMANCE
METRICVALUE
Top-1 Accuracy90.48%
Inference Latency<50ms per frame
Embedding Dim768 (DINOv2)
Memory SlotsConfigurable (1K–100K)
Search MethodApproximate NN (FAISS/HNSW)
05BUILD STATUS

WHAT'S BUILT TODAY

4/6 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
DINOv2 backboneCOMPLETEPre-trained weights, 768-dim output
Place fingerprint pipelineCOMPLETEFrame → embedding → search
Memory grid storageCOMPLETESlot-based, configurable capacity
Temporal decayCOMPLETEConfigurable fade for old places
REST APIIN PROGRESS/embed, /search, /add_place
Fleet syncPLANNEDShared memory across robots
06APPLICATIONS

WHERE LOCI DEPLOYS

  • APP_01

    SLAM LOOP CLOSURE

    Correct odometry drift in LiDAR/visual SLAM pipelines. Detect revisited places, trigger pose graph optimization.

  • APP_02

    AUTONOMOUS NAVIGATION

    Relocalize after sensor dropout or kidnapping. Fast recovery to known map.

  • APP_03

    MULTI-ROBOT FLEETS

    Shared place memory — one robot explores, all recognize. Coordinated mapping and planning.

  • APP_04

    WAREHOUSE LOGISTICS

    Recognize aisles, docks, and landmarks. Route optimization and task assignment.

  • APP_05

    INDOOR MOBILE ROBOTS

    No GPS — visual place recognition is the primary localization cue.

  • APP_06

    RESEARCH & BENCHMARKING

    VPR baseline for Nordland, RobotCar, and custom datasets.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: DINOV2 + VPR

  • DINOv2: self-supervised vision transformer, 768-dim embeddings
  • Place recognition: cosine similarity or L2 distance in embedding space
  • Approximate nearest-neighbor: FAISS or HNSW for scalable search
  • Temporal decay: configurable slot aging for long-term operation

API ENDPOINTS

POST /embed
Encode image → 768-dim fingerprint
POST /search
Query fingerprint → top-K place IDs + scores
POST /add_place
Store fingerprint in memory grid

TECH STACK

  • PyTorch
  • DINOv2
  • FAISS / HNSW
  • FastAPI
  • Redis (optional cache)
08PAPERS

RESEARCH PAPER

  1. [01]DINOv2: Learning Robust Visual Features without Supervision — Oquab et al., ICCV 2023