Skip to content
RFL_GLOBAL
中文
  • RA-L 2025 // VLM-GIST
  • NATURAL LANGUAGE
  • OPEN-VOCAB TRACKING

LOGOS

SPEAK IT INTO EXISTENCE

"Find the red cylindrical bottle." "Track the damaged wheel." Robots should understand natural language, not rigid detection categories. LOGOS converts language descriptions into grounded detections with persistent tracking — attribute-aware, provider-flexible, production-ready.

MODULE STATUS: PRODUCTION MVP

MODULE STATUS

OPEN

DIVISION
ANIMA
WAVE
W4
DOMAIN
UNDERSTANDING
WAVE 1 // ANIMA SUITE
LANGUAGE-GROUNDED TRACKING
LOGOS // W4 // 042/079
01THE CHALLENGE

THE GAP BETWEEN LANGUAGE AND PIXELS

Natural language is how humans think about tasks. "Find the red cylindrical bottle." "Track the damaged wheel." But most robot systems force you into rigid detection categories or hand-coded attribute rules.

VLM-powered perception exists, but it's either slow (5–10s per frame) or not grounded (gives descriptions, not pixel locations). Detectors can't understand attributes. Trackers don't use language. The gap between natural-language intention and pixel-level action is enormous.

02THE SOLUTION

WHAT LOGOS DELIVERS

LOGOS converts natural-language task descriptions into grounded object understanding and persistent tracking.

CAPABILITIES

  • Natural language → instances: "find the red cylindrical bottle" → detections + tracked IDs
  • Attribute extraction: color, shape, size, material, pose, condition
  • /describe endpoint: VLM scene grounding with provider abstraction and caching
  • /detect endpoint: open-vocabulary detection with prompt sanitization
  • /track + /track/sequence: motion-aware IDs, appearance-based assignment
  • Provider flexibility: Anthropic, OpenAI, optional MLX, or mock VLM
  • Detector abstraction: Ultralytics YOLO-World or mock — auto-selection
  • Apple Silicon MLX: ~200ms describe, ~80ms detect on M5 (4× M1)
03ENGINEERING

WHY THIS IS HARD

Combining VLMs, detectors, and trackers in a single service requires careful orchestration:

  1. 01Build task-grounded prompts that guide the VLM to the right objects
  2. 02Run open-vocabulary detection with prompt sanitization and fallback
  3. 03Validate detections against VLM descriptions — not all VLM objects are detectable
  4. 04Maintain stable IDs using motion and appearance features across occlusion
  5. 05Fallback gracefully when detectors or VLMs are unavailable — mock providers

LOGOS does all of this with provider abstraction — swap VLMs and detectors without rewriting the pipeline. Session tracking with archived-track re-identification maintains persistent identity.

04BUILD STATUS

WHAT'S BUILT TODAY

11/12 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETEVLM + YOLO-World weights — production-ready
/describe (VLM grounding)COMPLETETested with Anthropic + OpenAI providers
/detect (open-vocab)COMPLETEFallback to mock on missing detector
/track (single-image)COMPLETEMotion-aware stable IDs
/track/sequence (multi-frame)COMPLETEPer-track continuity analytics
Session ManagementCOMPLETEArchive-based re-identification
SegmentationCOMPLETEImage-refined polygons, optional SAM
VLM Abstraction (4 providers)COMPLETEAnthropic, OpenAI, MLX, mock
Detector AbstractionCOMPLETEUltralytics YOLO-World or mock
Docker (CPU/GPU)COMPLETEProduction deployment
Benchmark HarnessCOMPLETEscripts/benchmark_perception.py
API layerIN PROGRESSPublic API — pending dataset infrastructure
05APPLICATIONS

WHERE LOGOS DEPLOYS

  • APP_01

    ROBOTIC MANIPULATION

    Natural-language task descriptions into tracked objects for grasping.

  • APP_02

    WAREHOUSE AUTOMATION

    Find and track pallets by description — no RFID or barcodes needed.

  • APP_03

    INSPECTION & REPAIR

    Describe anomalies and track them across multiple views and angles.

  • APP_04

    ASSEMBLY & KITTING

    "Grab the red 5mm bolt" — without predefined SKU databases.

  • APP_05

    HUMAN-ROBOT COLLAB

    Operators describe what they want, robot understands and acts.

  • APP_06

    RESEARCH & DEV

    Test VLM-grounded perception hypotheses without custom pipelines.

06TECHNOLOGY

UNDER THE HOOD

FOUNDATION: VLM-GIST (RA-L 2025)

  • Language-grounded structured perception for robotics
  • Attribute-aware open-vocabulary detection and tracking
  • Provider-abstracted VLM + detector pipeline
  • Session-based persistent identity with re-identification

INFERENCE PIPELINE

  • Natural language task → prompt builder (task-grounded JSON)
  • VLM provider (Anthropic/OpenAI/MLX/mock) → structured scene objects
  • Open-vocab detector (Ultralytics YOLO-World) → bounding boxes + masks
  • Object association validator → query + attribute filtering
  • Session tracker → motion-aware stable IDs + appearance matching

PROVIDER FLEXIBILITY

Anthropic Claude
Primary VLM — highest accuracy scene grounding
OpenAI GPT
Secondary VLM — fallback provider
MLX (Apple Silicon)
Local inference on M1–M5, ~200ms describe
Mock VLM + Detector
Fallback-safe — always available for testing
07PAPERS

RESEARCH PAPER

  1. [01]VLM-GIST — RA-L 2025