- 90.48% // TOP-1
- VPR // DINOV2
- <50MS INFERENCE
LOCI
VISUAL MEMORY FOR ROBOTS
Robots need to know where they are. LOCI provides visual place recognition — match camera frames to a stored memory of places. DINOv2 embeddings, sub-50ms inference, 90.48% top-1 accuracy. Loop closure, relocalization, fleet memory.
MODULE STATUS: ACTIVETop-1 Accuracy
90.5%
- DIVISION
- ANIMA
- WAVE
- W5
- DOMAIN
- PERCEPTION
- WAVE 5 // ANIMA SUITE
- VISUAL PLACE RECOGNITION
ROBOTS LOSE THEIR PLACE
Mobile robots navigate through changing environments. Odometry drifts. GPS fails indoors. Without knowing "where am I?", mapping, planning, and coordination break down. Traditional place recognition relied on hand-crafted features — brittle, slow, and prone to failure under viewpoint or lighting changes.
You need a system that recognizes places from a single image, matches against a growing memory, and runs in real-time on robot hardware. DINOv2 and modern VPR methods deliver this — but deployment requires careful engineering.
WHAT LOCI DELIVERS
LOCI builds a visual memory of places. Each camera frame is encoded into a fingerprint; the system searches the memory grid for matches. Loop closure, relocalization, and fleet-wide place sharing.
PIPELINE
- 01Frame capture — RGB image from robot camera
- 02DINOv2 embedding — extract place fingerprint (768-dim vector)
- 03Memory search — compare against stored place slots, nearest-neighbor
- 04Match or store — return place ID if match > threshold, else add new slot
CAPABILITIES
- LOOP CLOSUREDetect when robot returns to a previously visited place→ Corrects odometry drift, enables consistent mapping
- RELOCALIZATIONRecover pose after kidnapping or sensor dropout→ Robust to failures, fast recovery
- FLEET MEMORYShared place database across multiple robots→ One robot explores, all robots recognize
WHY THIS IS HARD
Visual place recognition requires:
- 01Viewpoint invariance: same place from different angles must match
- 02Lighting robustness: day/night, shadows, seasonal changes
- 03Real-time inference: <50ms per frame on embedded hardware
- 04Scalable memory: thousands of places without linear search blowup
- 05Temporal decay: old places fade to avoid stale matches
LOCI uses DINOv2 for robust embeddings, approximate nearest-neighbor for fast search, and configurable decay for long-term operation.
REAL HARDWARE PERFORMANCE
Measured with DINOv2 + VPR pipeline:
| METRIC | VALUE |
|---|---|
| Top-1 Accuracy | 90.48% |
| Inference Latency | <50ms per frame |
| Embedding Dim | 768 (DINOv2) |
| Memory Slots | Configurable (1K–100K) |
| Search Method | Approximate NN (FAISS/HNSW) |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| DINOv2 backbone | COMPLETE | Pre-trained weights, 768-dim output |
| Place fingerprint pipeline | COMPLETE | Frame → embedding → search |
| Memory grid storage | COMPLETE | Slot-based, configurable capacity |
| Temporal decay | COMPLETE | Configurable fade for old places |
| REST API | IN PROGRESS | /embed, /search, /add_place |
| Fleet sync | PLANNED | Shared memory across robots |
WHERE LOCI DEPLOYS
- APP_01
SLAM LOOP CLOSURE
Correct odometry drift in LiDAR/visual SLAM pipelines. Detect revisited places, trigger pose graph optimization.
- APP_02
AUTONOMOUS NAVIGATION
Relocalize after sensor dropout or kidnapping. Fast recovery to known map.
- APP_03
MULTI-ROBOT FLEETS
Shared place memory — one robot explores, all recognize. Coordinated mapping and planning.
- APP_04
WAREHOUSE LOGISTICS
Recognize aisles, docks, and landmarks. Route optimization and task assignment.
- APP_05
INDOOR MOBILE ROBOTS
No GPS — visual place recognition is the primary localization cue.
- APP_06
RESEARCH & BENCHMARKING
VPR baseline for Nordland, RobotCar, and custom datasets.
UNDER THE HOOD
FOUNDATION: DINOV2 + VPR
- DINOv2: self-supervised vision transformer, 768-dim embeddings
- Place recognition: cosine similarity or L2 distance in embedding space
- Approximate nearest-neighbor: FAISS or HNSW for scalable search
- Temporal decay: configurable slot aging for long-term operation
API ENDPOINTS
- POST /embed
- Encode image → 768-dim fingerprint
- POST /search
- Query fingerprint → top-K place IDs + scores
- POST /add_place
- Store fingerprint in memory grid
TECH STACK
- PyTorch
- DINOv2
- FAISS / HNSW
- FastAPI
- Redis (optional cache)
RESEARCH PAPER
- [01]DINOv2: Learning Robust Visual Features without Supervision — Oquab et al., ICCV 2023