- WAVE 5 // DEVELOPMENT
- HIERARCHICAL VLA
- ZERO-SHOT
DAEDALUS
HIERARCHICAL ZERO-SHOT VLA
Three-layer hierarchical VLA (GeneralVLA-inspired): affordance segmentation identifies graspable regions, text-driven 3D planning generates approach trajectories, collision-aware grasp execution handles physical interaction. Each layer independently testable.
MODULE STATUS: DEVELOPMENTPROTEUS + ABYSSOS + CHIRON
3-LAYER
- DIVISION
- ANIMA
- WAVE
- W5
- DOMAIN
- MANIPULATION
- WAVE 5 // ANIMA SUITE
- MANIPULATION — HIERARCHICAL ZERO-SHOT VLA
MANIPULATION IS A STACK, NOT A MONOLITH
End-to-end VLAs learn everything jointly — perception, planning, execution — in a single opaque network. When they fail, you cannot tell which stage broke: did the model misidentify the object, plan an infeasible path, or miscalculate the grasp force?
Monolithic models also require hardware to debug. You can't test grasp planning without a robot, can't validate perception without a full pipeline. DAEDALUS decouples manipulation into three independently testable layers — debug each one offline, pre-hardware.
WHAT DAEDALUS DELIVERS
DAEDALUS is a three-layer hierarchical VLA: affordance segmentation → 3D trajectory planning → collision-aware grasp execution. Each layer runs independently, communicates via typed interfaces, and can be tested without hardware.
PIPELINE
- 01Affordance segmentation — identifies graspable regions, contact surfaces, and manipulation affordances from vision input
- 02Text-driven 3D planning — generates approach trajectories conditioned on natural language instructions and segmented affordances
- 03Collision-aware grasp execution — handles physical interaction with force control, collision avoidance, and adaptive grip
- 04Layer isolation — each stage independently testable with mock inputs, enabling pre-hardware validation
CAPABILITIES
- AFFORDANCE SEGMENTATIONVision-based identification of graspable regions and contact surfaces using PROTEUS→ Know what to grasp before planning how to grasp it
- KNOWLEDGE-GUIDED PLANNINGText-conditioned 3D trajectory generation via ABYSSOS depth and spatial reasoning→ Plan approach paths without trial-and-error on hardware
- COLLISION-AWARE EXECUTIONForce-controlled grasp execution with CHIRON motor primitives and collision detection→ Safe physical interaction with adaptive grip strategies
WHY THIS IS HARD
Decomposing manipulation into layers creates interface boundaries that most systems avoid:
- 01Affordance segmentation must generalize zero-shot to novel objects — no per-object training
- 02Planning layer receives segmentation masks, not raw images — information loss at the boundary must be minimized
- 03Trajectory generation must be collision-aware without full physics simulation — approximate but fast
- 04Grasp execution must handle force feedback in real-time while respecting planned trajectories
- 05Layer interfaces must be typed and versioned to allow independent testing and evolution
- 06Each layer has different latency requirements: segmentation (~100ms), planning (~200ms), execution (<10ms)
DAEDALUS solves this by defining clean typed interfaces between layers — affordance maps, trajectory specs, force profiles — enabling independent testing of each manipulation stage.
SYSTEM PERFORMANCE
Measured across the three-layer hierarchical pipeline:
| METRIC | VALUE |
|---|---|
| Testable Layers | 3 independently validated stages |
| Hardware Requirement | Offline-capable — works pre-hardware |
| Planning Mode | Knowledge-guided trajectory generation |
| Execution Transparency | Full layer-by-layer inspection |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Affordance Segmentation | COMPLETE | Zero-shot graspable region detection via PROTEUS |
| 3D Planning Layer | COMPLETE | Text-conditioned trajectory generation with ABYSSOS depth |
| Grasp Execution | IN PROGRESS | CHIRON motor primitives integration ongoing |
| Core models | COMPLETE | Segmentation and planning models production-ready |
| API layer | IN PROGRESS | Public API, pending dataset infrastructure |
WHERE DAEDALUS DEPLOYS
- APP_01
PRE-HARDWARE VALIDATION
Test and debug manipulation pipelines entirely in software before deploying to physical robots — catch failures at each layer.
- APP_02
DEBUGGABLE MANIPULATION
When a grasp fails, inspect which layer broke: bad segmentation, infeasible plan, or force miscalculation. No more black-box debugging.
- APP_03
MODULAR ROBOTICS
Swap individual layers independently — upgrade perception without retraining planning, replace execution without changing the vision stack.
UNDER THE HOOD
FOUNDATION: GENERALVLA-INSPIRED HIERARCHY
- Three-layer decomposition: perception → planning → execution
- Typed interfaces between layers for independent testing
- Zero-shot affordance detection from vision foundation models
- Text-conditioned trajectory planning with spatial reasoning
DAEDALUS IMPLEMENTATION
- Layer 1: PROTEUS-based affordance segmentation with graspability scoring
- Layer 2: ABYSSOS depth + language-conditioned 3D trajectory planner
- Layer 3: CHIRON motor primitives with force feedback and collision avoidance
- Typed inter-layer protocol: AffordanceMap → TrajectorySpec → ForceProfile
ANIMA MODULE INTEGRATION
- PROTEUS provides zero-shot visual affordance segmentation
- ABYSSOS provides monocular depth for 3D spatial reasoning
- CHIRON provides motor primitive execution and force control
LAYER SPECIFICATIONS
- Segmentation: ~100ms latency, mask output
- Planning: ~200ms latency, trajectory output
- Execution: <10ms loop, force-controlled
- Mock interfaces for offline testing
RESEARCH BASIS
- [01]GeneralVLA-inspired hierarchical manipulation — decomposed VLA with typed layer interfaces for testable robotics