- RA-L 2025 // VLM-GIST
- NATURAL LANGUAGE
- OPEN-VOCAB TRACKING
LOGOS
SPEAK IT INTO EXISTENCE
"Find the red cylindrical bottle." "Track the damaged wheel." Robots should understand natural language, not rigid detection categories. LOGOS converts language descriptions into grounded detections with persistent tracking — attribute-aware, provider-flexible, production-ready.
MODULE STATUS: PRODUCTION MVPMODULE STATUS
OPEN
- DIVISION
- ANIMA
- WAVE
- W4
- DOMAIN
- UNDERSTANDING
- WAVE 1 // ANIMA SUITE
- LANGUAGE-GROUNDED TRACKING
THE GAP BETWEEN LANGUAGE AND PIXELS
Natural language is how humans think about tasks. "Find the red cylindrical bottle." "Track the damaged wheel." But most robot systems force you into rigid detection categories or hand-coded attribute rules.
VLM-powered perception exists, but it's either slow (5–10s per frame) or not grounded (gives descriptions, not pixel locations). Detectors can't understand attributes. Trackers don't use language. The gap between natural-language intention and pixel-level action is enormous.
WHAT LOGOS DELIVERS
LOGOS converts natural-language task descriptions into grounded object understanding and persistent tracking.
CAPABILITIES
- Natural language → instances: "find the red cylindrical bottle" → detections + tracked IDs
- Attribute extraction: color, shape, size, material, pose, condition
- /describe endpoint: VLM scene grounding with provider abstraction and caching
- /detect endpoint: open-vocabulary detection with prompt sanitization
- /track + /track/sequence: motion-aware IDs, appearance-based assignment
- Provider flexibility: Anthropic, OpenAI, optional MLX, or mock VLM
- Detector abstraction: Ultralytics YOLO-World or mock — auto-selection
- Apple Silicon MLX: ~200ms describe, ~80ms detect on M5 (4× M1)
WHY THIS IS HARD
Combining VLMs, detectors, and trackers in a single service requires careful orchestration:
- 01Build task-grounded prompts that guide the VLM to the right objects
- 02Run open-vocabulary detection with prompt sanitization and fallback
- 03Validate detections against VLM descriptions — not all VLM objects are detectable
- 04Maintain stable IDs using motion and appearance features across occlusion
- 05Fallback gracefully when detectors or VLMs are unavailable — mock providers
LOGOS does all of this with provider abstraction — swap VLMs and detectors without rewriting the pipeline. Session tracking with archived-track re-identification maintains persistent identity.
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | VLM + YOLO-World weights — production-ready |
| /describe (VLM grounding) | COMPLETE | Tested with Anthropic + OpenAI providers |
| /detect (open-vocab) | COMPLETE | Fallback to mock on missing detector |
| /track (single-image) | COMPLETE | Motion-aware stable IDs |
| /track/sequence (multi-frame) | COMPLETE | Per-track continuity analytics |
| Session Management | COMPLETE | Archive-based re-identification |
| Segmentation | COMPLETE | Image-refined polygons, optional SAM |
| VLM Abstraction (4 providers) | COMPLETE | Anthropic, OpenAI, MLX, mock |
| Detector Abstraction | COMPLETE | Ultralytics YOLO-World or mock |
| Docker (CPU/GPU) | COMPLETE | Production deployment |
| Benchmark Harness | COMPLETE | scripts/benchmark_perception.py |
| API layer | IN PROGRESS | Public API — pending dataset infrastructure |
WHERE LOGOS DEPLOYS
- APP_01
ROBOTIC MANIPULATION
Natural-language task descriptions into tracked objects for grasping.
- APP_02
WAREHOUSE AUTOMATION
Find and track pallets by description — no RFID or barcodes needed.
- APP_03
INSPECTION & REPAIR
Describe anomalies and track them across multiple views and angles.
- APP_04
ASSEMBLY & KITTING
"Grab the red 5mm bolt" — without predefined SKU databases.
- APP_05
HUMAN-ROBOT COLLAB
Operators describe what they want, robot understands and acts.
- APP_06
RESEARCH & DEV
Test VLM-grounded perception hypotheses without custom pipelines.
UNDER THE HOOD
FOUNDATION: VLM-GIST (RA-L 2025)
- Language-grounded structured perception for robotics
- Attribute-aware open-vocabulary detection and tracking
- Provider-abstracted VLM + detector pipeline
- Session-based persistent identity with re-identification
INFERENCE PIPELINE
- Natural language task → prompt builder (task-grounded JSON)
- VLM provider (Anthropic/OpenAI/MLX/mock) → structured scene objects
- Open-vocab detector (Ultralytics YOLO-World) → bounding boxes + masks
- Object association validator → query + attribute filtering
- Session tracker → motion-aware stable IDs + appearance matching
PROVIDER FLEXIBILITY
- Anthropic Claude
- Primary VLM — highest accuracy scene grounding
- OpenAI GPT
- Secondary VLM — fallback provider
- MLX (Apple Silicon)
- Local inference on M1–M5, ~200ms describe
- Mock VLM + Detector
- Fallback-safe — always available for testing
RESEARCH PAPER
- [01]VLM-GIST — RA-L 2025