Skip to content
RFL_GLOBAL
中文
  • UAV-TRACKVLA // W7
  • Python
  • ANIMA

UAV-TRACKVLA

Language-Commanded Pursuit That Doesn't Lose Its Target

Persistent surveillance UAVs lose tracked targets during occlusions, direction changes, or handoffs between operators. Re-acquisition requires human attention and precious seconds. In a dynamic pursuit — a vehicle fleeing a checkpoint, a person of interest moving through a crowd — those seconds determine whether the mission succeeds. UAV-TRACKVLA (codename: Tengu) implements UAV-Track VLA (arXiv:2604.02241), fusing a frozen SigLIP visual backbone with temporal compression and a dual-decoder head: one for grounding (where is the target) and one for flow-matching action generation (how does the UAV move to stay on it).

MODULE STATUS: PRODUCTION

Architecture

1 pass

DIVISION
GENERIC
WAVE
W7
DOMAIN
GENERAL
WAVE 7 // ANIMA SUITE
FOUNDATION — UAV-TRACKVLA
UAV-TRACKVLA // W7 // 010/012
01THE CHALLENGE

THE PROBLEM WE SOLVE

Persistent surveillance UAVs lose tracked targets during occlusions, direction changes, or handoffs between operators. Re-acquisition requires human attention and precious seconds. In a dynamic pursuit — a vehicle fleeing a checkpoint, a person of interest moving through a crowd — those seconds determine whether the mission succeeds.

A single operator can task a pursuit UAV in plain language and trust it to maintain lock through complex maneuvers — shrinking the human-machine ratio for ISR sorties and reducing cognitive load during high-tempo operations.

02THE SOLUTION

WHAT UAV-TRACKVLA DELIVERS

UAV-TRACKVLA (codename: Tengu) implements UAV-Track VLA (arXiv:2604.02241), fusing a frozen SigLIP visual backbone with temporal compression and a dual-decoder head: one for grounding (where is the target) and one for flow-matching action generation (how does the UAV move to stay on it).

CAPABILITIES

  • UAV-TRACKVLA (codename: Tengu) implements UAV-Track VLA (arXiv:2604.02241), fusing a frozen SigLIP visual backbone with temporal compression and a dual-decoder head: one for grounding (where is the target) and one for flow-matching action generation (how does the UAV move to stay on it)
  • Natural language commands update the grounding target in real-time, while the action decoder produces continuous flight commands without a separate planning module
  • The unified architecture eliminates the latency of chained detect-plan-act pipelines.
03ENGINEERING

WHY THIS IS HARD

Building UAV-TRACKVLA requires solving multiple coupled problems:

  1. 01Persistent surveillance UAVs lose tracked targets during occlusions, direction changes, or handoffs between operators
  2. 02Re-acquisition requires human attention and precious seconds
  3. 03UAV-TRACKVLA (codename: Tengu) implements UAV-Track VLA (arXiv:2604.02241), fusing a frozen SigLIP visual backbone with temporal compression and a dual-decoder head: one for grounding (where is the target) and one for flow-matching action generation (how does the UAV move to stay on it)
  4. 04Natural language commands update the grounding target in real-time, while the action decoder produces continuous flight commands without a separate planning module

UAV-TRACKVLA solves these through careful architecture design and rigorous validation.

04BENCHMARKS

PROOF, NOT PROMISES

Key performance metrics:

PROOF, NOT PROMISES
METRICVALUE
ArchitectureUnified VLA — grounding + action generation in one forward pass, no pipeline chaining
Action GenerationFlow-matching decoder produces smooth continuous flight commands — no discrete planning lag
BackboneFrozen SigLIP (no fine-tuning cost) — only temporal + decoder heads are trained
Defense AnglePersistent ISR pursuit, checkpoint escape re-acquisition, multi-operator UAV tasking
05BUILD STATUS

WHAT'S BUILT TODAY

2/4 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
ArchitectureCOMPLETEUnified VLA — grounding + action generation in one forward pass, no pipeline chaining
Action GenerationCOMPLETEFlow-matching decoder produces smooth continuous flight commands — no discrete planning lag
BackboneIN PROGRESSFrozen SigLIP (no fine-tuning cost) — only temporal + decoder heads are trained
Defense AngleIN PROGRESSPersistent ISR pursuit, checkpoint escape re-acquisition, multi-operator UAV tasking
06APPLICATIONS

WHERE UAV-TRACKVLA DEPLOYS

  • APP_01

    AUTONOMOUS SYSTEMS

    A single operator can task a pursuit UAV in plain language and trust it to maintain lock through complex maneuvers — shrinking the human-machine ratio for ISR sorties and reducing cognitive load during high-tempo operations.

  • APP_02

    RESEARCH LABS

    UAV-TRACKVLA (codename: Tengu) implements UAV-Track VLA (arXiv:2604.02241), fusing a frozen SigLIP visual backbone with temporal compression and a dual-decoder head: one for grounding (where is the target) and one for flow-matching action generation (how does the UAV move to stay on it).

  • APP_03

    EDGE COMPUTING

    Aerial tracker that understands "follow the white SUV turning left" and executes it — fusing language, vision, and flight control in one frozen backbone.