Skip to content
RFL_GLOBAL
中文
  • ICLR 2025 // TRACEVLA
  • DRAW → ROBOT FOLLOWS
  • OPENVLA-7B

DAEMON

THE GUIDING SPIRIT

Training manipulation policies requires thousands of trajectories. DAEMON sidesteps this — overlay visual trace prompts on observation frames, and the robot learns from the hint. Draw the path, the robot follows. Proprioception traces are instant, calibration-perfect, and commercially licensed.

MODULE STATUS: ACTIVE

RTX 4090

100/s

DIVISION
ANIMA
WAVE
W1
DOMAIN
ACTION
WAVE 2 // ANIMA SUITE
VISUAL TRACE PROMPTING
DAEMON // W1 // 016/079
01THE CHALLENGE

VLA MODELS NEED ACTION GUIDANCE AT SCALE

Training manipulation policies requires thousands of robot trajectories. Current VLA models improve with scale but falter on novel tasks. They need action guidance — a way to hint at the solution without dense trajectory labels. You can't annotate trajectories at scale.

Humans solve this instantly: show the arm where to go. But how do you encode "trajectory" into an image-based policy? You need a visual prompt — draw on the observation frame where the gripper should move, and the policy learns from that hint. No one had productized this until TraceVLA.

02THE SOLUTION

WHAT DAEMON DELIVERS

DAEMON augments Vision Language Action models with visual trace prompting. It overlays robot end-effector trajectories as drawn paths onto observation frames, conditioning policy inference on visual guidance.

CAPABILITIES

  • Two trace generators: Proprioception (commercial-safe, MIT) and CoTracker (research-only, CC-BY-NC)
  • Proprioception path: forward kinematics → 2D projection → draw on frame (zero latency)
  • OpenVLA-7B policy inference with finetuning on custom datasets
  • Multi-device: MLX (Apple M1–M5), CUDA (RTX 4090, A100), CPU fallback
  • Real-time policy prediction: 20-100 inferences/sec on GPU, ~30+ on M5
  • LiDAR simulator (Gazebo Harmonic) + SimplerEnv benchmark integration
03ENGINEERING

WHY THIS IS HARD

Visual trace prompting is a new paradigm. It requires solving:

  1. 01Trace generation: real-time proprioceptive forward kinematics (calibration-critical) or optical flow tracking
  2. 02Policy runtime: OpenVLA inference on multiple devices — MLX bootstrap, CUDA Torch, CPU fallback
  3. 03Finetuning: dataset preprocessing, distributed training on 1-8 GPUs, custom normalization per domain
  4. 04Licensing discipline: CoTracker is CC-BY-NC (research-only); proprioception must be used for commercial
  5. 05Evaluation: SimplerEnv benchmark requires Gazebo simulation and complex task definitions

DAEMON handles all of this. Proprioception traces are instant, provably correct, and MIT-licensed. CoTracker is available for research but disabled in production.

04BENCHMARKS

REAL HARDWARE PERFORMANCE

Measured with OpenVLA-7B + Proprioception traces:

REAL HARDWARE PERFORMANCE
DEVICEPOLICYTRACELATENCYTHROUGHPUTMEMORY
RTX 4090OpenVLA-7BProprioception40-50ms100 inf/s14GB
RTX 4080OpenVLA-7BProprioception60-80ms50 inf/s12GB
A100OpenVLA-7BProprioception30-40ms100+ inf/s16GB
Apple M5 (MLX)BootstrapProprioception~50ms~30 inf/s~6GB
Apple M3 (MLX)BootstrapProprioception85ms20 inf/s6GB
05BUILD STATUS

WHAT'S BUILT TODAY

10/15 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETEOpenVLA-7B weights + trace conditioning — production-ready
Proprioception trace genCOMPLETEForward kinematics + 2D drawing, MIT licensed
Device detectionCOMPLETEMLX, CUDA, CPU auto-detection
Configuration systemCOMPLETE.env + TOML, device selection
Dataset preprocessingCOMPLETEBridge V2, Open-X-Embodiment canonicalization
Model cacheCOMPLETEHuggingFace snapshot verification
Server frameworkCOMPLETEFastAPI + gRPC scaffolding
LiDAR simulatorCOMPLETEGazebo Harmonic Docker manager
Docker buildCOMPLETECPU, GPU, dev profiles
/trace/from_proprioceptionCOMPLETEAPI ready
/policy/predictIN PROGRESSOpenVLA inference pipeline
Finetuning pipelineIN PROGRESSDistributed training + validation
API layerIN PROGRESSPublic REST + gRPC API — pending dataset infrastructure
SimplerEnv integrationIN PROGRESSTask benchmark suite
CoTracker research pathPLANNEDOptical flow + dense tracking (Phase 3)
06APPLICATIONS

WHERE DAEMON DEPLOYS

  • APP_01

    MANUFACTURING & ASSEMBLY

    Robots learning custom assembly sequences from a few video demonstrations with trace-guided policies.

  • APP_02

    WAREHOUSE AUTOMATION

    Bin picking and palletizing policies adapted to new item types via visual trace finetuning.

  • APP_03

    MOBILE MANIPULATION

    Fetch/Spot-like robots learning navigation + manipulation from human-provided trajectory traces.

  • APP_04

    FOOD & BEVERAGE

    Sorting, packing, and quality control where visual traces guide novel product handling.

  • APP_05

    ELECTRONICS ASSEMBLY

    Fine-grained manipulation (soldering, component placement) learned via trace-guided policies.

  • APP_06

    RESEARCH & EDUCATION

    TraceVLA baseline for benchmarking visual prompting approaches in manipulation tasks.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: TRACEVLA (ICLR 2025)

  • Vision Language Action (VLA) model: OpenVLA-7B base (Apache-2.0)
  • Visual trace prompting: overlay robot trajectory hints on observations
  • Key insight: traces dramatically improve policy performance without dense labels
  • Pre-trained on Google Open X-Embodiment, Bridge V2, LIBERO datasets

KEY INNOVATION

Draw the path, the robot follows. DAEMON encodes trajectory hints as visual overlays on observation frames — proprioception traces from forward kinematics are instant, calibration-perfect, and commercially licensed (MIT). No dense trajectory annotations required.

TWO TRACE PATHS

Proprioception (Commercial)
Forward kinematics → 2D projection → draw on frame. Zero latency, perfect calibration. MIT licensed.
CoTracker (Research)
Dense optical flow tracking from video. CC-BY-NC licensed — NOT for commercial use. Phase 3.

DEPLOYMENT STACK

  • FastAPI + async policy inference with trace pre-computation
  • Prometheus metrics for prediction latency, throughput, model status
  • Finetuning queue with job tracking (1-8 GPUs)
  • Gazebo Harmonic LiDAR simulator for sim-to-real validation

MULTI-DEVICE RUNTIME

CUDA
NVIDIA GPUs (RTX 4090, A100) — full OpenVLA Torch inference, 40-50ms
MLX
Apple Silicon (M1–M5) — bootstrap inference, ~50ms on M5 (4× M1)
CPU
Bootstrap fallback — slow, demo-only
08PAPERS

RESEARCH PAPER

  1. [01]TraceVLA: Visual Trace Prompting for Robot Manipulation Policies — Huang et al., ICLR 2025