- CVPR 2025 HIGHLIGHT // LLMDET
- ZERO RETRAINING
- PRODUCTION MVP
AZOTH
THE UNIVERSAL EYE
Vision systems today are brittle — every new object requires retraining. AZOTH detects any object by natural language description alone. "Red container," "damaged wheel," "assembled module" — no training, no vocabulary limits, no recompilation. Just describe what you need found.
MODULE STATUS: PRODUCTIONApple Silicon MLX direct
146ms
- DIVISION
- ANIMA
- WAVE
- W2
- DOMAIN
- PERCEPTION
- WAVE 1 // ANIMA SUITE
- OPEN-VOCABULARY DETECTION
FIXED-VOCABULARY DETECTORS BREAK IN THE REAL WORLD
Traditional detectors lock you into a fixed vocabulary the moment they leave the lab. Factory floors need to detect hundreds of SKUs. Warehouses can't recompile their models daily. Robots hit new environments and go blind. Every new object means retraining, relabeling, redeploying.
Competitors rely on CLIP embeddings or brute-force detectors that don't understand spatial relationships. VLMs are too slow for robotics — seconds per frame. YOLO variants require vocabulary limits baked into the model. Robotics needs perception that understands language, not just labels.
WHAT AZOTH DELIVERS
AZOTH is the productized perception endpoint for Robot Flow Labs. It converts images and natural-language queries into precise bounding boxes — designed to feed grasping, navigation, and human-in-the-loop workflows.
CAPABILITIES
- Open-vocabulary detection: detect arbitrary object categories from text prompts in real-time
- Paper-aligned inference: detection, phrase grounding, and referential expression — same runtime
- Container-first deployment: single service for local dev and production rollout
- REST + gRPC dual interface: FastAPI for web, protobuf for systems
- Bounded concurrency: predictable overload behavior (429/RESOURCE_EXHAUSTED on saturation)
- Apple Silicon MLX backend: 146ms mean detection on M-series (25× faster than CPU)
WHY THIS IS HARD
Open-vocabulary detection isn't just "CLIP + bounding boxes." It requires understanding how language maps to pixel coordinates:
- 01Language-to-pixel grounding: the model must learn spatial relationships, not just category similarity
- 02Real-time throughput: VLMs take seconds per frame — robotics needs sub-200ms on edge hardware
- 03Admission control: bounded concurrency prevents OOM under production load
- 04Multi-backend parity: paper-faithful LLMDet vs. experimental MLX — different speed/fidelity tradeoffs
- 05Reproducibility: pinned model revisions with manifest integrity checking for production safety
AZOTH uses LLMDet (CVPR 2025 Highlight), which learns from supervision on grounding and captions. It doesn't memorize categories — it understands how language maps to pixels.
LIVE MEASUREMENTS — MARCH 2026
Tested with pinned llmdet_tiny revision on standard COCO test images:
| METRIC | RESULT | NOTES |
|---|---|---|
| Apple Silicon MLX direct | 146ms mean | 59ms min, 232ms max |
| MLX forward pass | 51.6ms | 4.4ms image load, 1.9ms prompt |
| Direct inference (x86) | 2.5s mean | CPU baseline |
| HTTP/ASGI transport (x86) | 1.8s mean | REST endpoint |
| gRPC loopback (x86) | 2.0s mean | Protobuf transport |
| Container detect (REST) | 5.3s | Inference: 4.2s |
| Container detect (gRPC) | 4.8s | Inference: 3.8s |
| Container warmup | 159.7s | Includes first model download |
| Apple M5 estimated | ~80ms mean | 4× M1 performance scaling |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | LLMDet weights trained and validated — production-ready |
| FastAPI REST service | PRODUCTION | Full detection, grounding, referential expression |
| gRPC transport | PRODUCTION | Protobuf contract, bounded concurrency |
| Docker CPU runtime | PRODUCTION | Kubernetes-ready profiles |
| Docker GPU runtime | COMPLETE | Tested and validated |
| Model caching | PRODUCTION | Snapshots with manifest integrity |
| Health/metrics | PRODUCTION | Prometheus endpoints |
| Benchmark harness | PRODUCTION | Reproducible latency testing |
| Apple Silicon MLX | EXPERIMENTAL | 25× faster, fidelity gap vs. LLMDet noted |
| API layer | IN PROGRESS | Public REST + gRPC API — pending dataset infrastructure |
WHERE AZOTH DEPLOYS
- APP_01
MANUFACTURING QA
Detect quality defects, damaged parts, and assembly steps by natural-language description — no SKU database required.
- APP_02
WAREHOUSE OPERATIONS
Find pallets, boxes, and products by description without maintaining a fixed object vocabulary.
- APP_03
INSPECTION & REPAIR
Identify faults and wear patterns on equipment without retraining — just describe the defect.
- APP_04
AGRICULTURE
Detect crop types, disease, and ripeness stages by description across changing seasonal conditions.
- APP_05
ROBOTICS FLEETS
Give robots new perception skills without recompilation or field retraining — update via text prompts.
- APP_06
HUMAN-IN-THE-LOOP
Bridge the gap between human instructions and pixel-level actions for collaborative robotics.
UNDER THE HOOD
FOUNDATION: LLMDET (CVPR 2025 HIGHLIGHT)
- Learns from supervision on grounding and captions — not category memorization
- Genuinely open-vocabulary, not zero-shot approximation
- Detection, phrase grounding, and referential expression in one unified architecture
- Apache-2.0 licensed — safe for commercial robotics
KEY INNOVATION
Language-to-pixel grounding: LLMDet understands how natural language maps to spatial coordinates in images, enabling detection of any object described in words — no retraining, no vocabulary limits.
DEPLOYMENT STACK
- Python 3.12, UV package manager
- FastAPI + gRPC (protobuf) dual-mode serving
- Pinned HuggingFace model revision for reproducibility
- Model caching with manifest integrity checking
- Kubernetes-ready Docker profiles (CPU + GPU)
INFERENCE BACKENDS
- LLMDet (Primary)
- Paper-faithful CVPR 2025 implementation — highest fidelity
- AIMv2 MLX
- Apple Silicon (M1–M5) — 25× faster, ~80ms on M5 (4× M1)
- CPU Fallback
- x86 baseline — 2.5s mean for validation and development