Skip to content
RFL_GLOBAL
中文
  • CVPR 2025 HIGHLIGHT // LLMDET
  • ZERO RETRAINING
  • PRODUCTION MVP

AZOTH

THE UNIVERSAL EYE

Vision systems today are brittle — every new object requires retraining. AZOTH detects any object by natural language description alone. "Red container," "damaged wheel," "assembled module" — no training, no vocabulary limits, no recompilation. Just describe what you need found.

MODULE STATUS: PRODUCTION

Apple Silicon MLX direct

146ms

DIVISION
ANIMA
WAVE
W2
DOMAIN
PERCEPTION
WAVE 1 // ANIMA SUITE
OPEN-VOCABULARY DETECTION
AZOTH // W2 // 005/079
01THE CHALLENGE

FIXED-VOCABULARY DETECTORS BREAK IN THE REAL WORLD

Traditional detectors lock you into a fixed vocabulary the moment they leave the lab. Factory floors need to detect hundreds of SKUs. Warehouses can't recompile their models daily. Robots hit new environments and go blind. Every new object means retraining, relabeling, redeploying.

Competitors rely on CLIP embeddings or brute-force detectors that don't understand spatial relationships. VLMs are too slow for robotics — seconds per frame. YOLO variants require vocabulary limits baked into the model. Robotics needs perception that understands language, not just labels.

02THE SOLUTION

WHAT AZOTH DELIVERS

AZOTH is the productized perception endpoint for Robot Flow Labs. It converts images and natural-language queries into precise bounding boxes — designed to feed grasping, navigation, and human-in-the-loop workflows.

CAPABILITIES

  • Open-vocabulary detection: detect arbitrary object categories from text prompts in real-time
  • Paper-aligned inference: detection, phrase grounding, and referential expression — same runtime
  • Container-first deployment: single service for local dev and production rollout
  • REST + gRPC dual interface: FastAPI for web, protobuf for systems
  • Bounded concurrency: predictable overload behavior (429/RESOURCE_EXHAUSTED on saturation)
  • Apple Silicon MLX backend: 146ms mean detection on M-series (25× faster than CPU)
03ENGINEERING

WHY THIS IS HARD

Open-vocabulary detection isn't just "CLIP + bounding boxes." It requires understanding how language maps to pixel coordinates:

  1. 01Language-to-pixel grounding: the model must learn spatial relationships, not just category similarity
  2. 02Real-time throughput: VLMs take seconds per frame — robotics needs sub-200ms on edge hardware
  3. 03Admission control: bounded concurrency prevents OOM under production load
  4. 04Multi-backend parity: paper-faithful LLMDet vs. experimental MLX — different speed/fidelity tradeoffs
  5. 05Reproducibility: pinned model revisions with manifest integrity checking for production safety

AZOTH uses LLMDet (CVPR 2025 Highlight), which learns from supervision on grounding and captions. It doesn't memorize categories — it understands how language maps to pixels.

04BENCHMARKS

LIVE MEASUREMENTS — MARCH 2026

Tested with pinned llmdet_tiny revision on standard COCO test images:

LIVE MEASUREMENTS — MARCH 2026
METRICRESULTNOTES
Apple Silicon MLX direct146ms mean59ms min, 232ms max
MLX forward pass51.6ms4.4ms image load, 1.9ms prompt
Direct inference (x86)2.5s meanCPU baseline
HTTP/ASGI transport (x86)1.8s meanREST endpoint
gRPC loopback (x86)2.0s meanProtobuf transport
Container detect (REST)5.3sInference: 4.2s
Container detect (gRPC)4.8sInference: 3.8s
Container warmup159.7sIncludes first model download
Apple M5 estimated~80ms mean4× M1 performance scaling
05BUILD STATUS

WHAT'S BUILT TODAY

9/10 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETELLMDet weights trained and validated — production-ready
FastAPI REST servicePRODUCTIONFull detection, grounding, referential expression
gRPC transportPRODUCTIONProtobuf contract, bounded concurrency
Docker CPU runtimePRODUCTIONKubernetes-ready profiles
Docker GPU runtimeCOMPLETETested and validated
Model cachingPRODUCTIONSnapshots with manifest integrity
Health/metricsPRODUCTIONPrometheus endpoints
Benchmark harnessPRODUCTIONReproducible latency testing
Apple Silicon MLXEXPERIMENTAL25× faster, fidelity gap vs. LLMDet noted
API layerIN PROGRESSPublic REST + gRPC API — pending dataset infrastructure
06APPLICATIONS

WHERE AZOTH DEPLOYS

  • APP_01

    MANUFACTURING QA

    Detect quality defects, damaged parts, and assembly steps by natural-language description — no SKU database required.

  • APP_02

    WAREHOUSE OPERATIONS

    Find pallets, boxes, and products by description without maintaining a fixed object vocabulary.

  • APP_03

    INSPECTION & REPAIR

    Identify faults and wear patterns on equipment without retraining — just describe the defect.

  • APP_04

    AGRICULTURE

    Detect crop types, disease, and ripeness stages by description across changing seasonal conditions.

  • APP_05

    ROBOTICS FLEETS

    Give robots new perception skills without recompilation or field retraining — update via text prompts.

  • APP_06

    HUMAN-IN-THE-LOOP

    Bridge the gap between human instructions and pixel-level actions for collaborative robotics.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: LLMDET (CVPR 2025 HIGHLIGHT)

  • Learns from supervision on grounding and captions — not category memorization
  • Genuinely open-vocabulary, not zero-shot approximation
  • Detection, phrase grounding, and referential expression in one unified architecture
  • Apache-2.0 licensed — safe for commercial robotics

KEY INNOVATION

Language-to-pixel grounding: LLMDet understands how natural language maps to spatial coordinates in images, enabling detection of any object described in words — no retraining, no vocabulary limits.

DEPLOYMENT STACK

  • Python 3.12, UV package manager
  • FastAPI + gRPC (protobuf) dual-mode serving
  • Pinned HuggingFace model revision for reproducibility
  • Model caching with manifest integrity checking
  • Kubernetes-ready Docker profiles (CPU + GPU)

INFERENCE BACKENDS

LLMDet (Primary)
Paper-faithful CVPR 2025 implementation — highest fidelity
AIMv2 MLX
Apple Silicon (M1–M5) — 25× faster, ~80ms on M5 (4× M1)
CPU Fallback
x86 baseline — 2.5s mean for validation and development
08PAPERS

RESEARCH PAPER

  1. [01]LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models — CVPR 2025 Highlight