Skip to content
RFL_GLOBAL
中文
  • SCENE GRAPHS // 2025
  • 22.5% OVER BASELINES
  • ZERO-SHOT TRANSFER

MIDAS

MANIPULATION — ZERO-SHOT VIA SCENE GRAPHS

Zero-shot robotic manipulation through agentic operational graphs combining VLM planning with dynamic scene graphs. Decomposes complex tasks into graph-structured plans with collision-aware trajectories and runtime failure recovery. 22.5-25% improvement over VLA and hierarchical baselines — no task-specific training required.

MODULE STATUS: DEVELOPMENT

Improvement Over VLA Baselines

ZERO

DIVISION
ANIMA
WAVE
W5
DOMAIN
MANIPULATION
WAVE 5 // ANIMA SUITE
ZERO-SHOT VIA SCENE GRAPHS
MIDAS // W5 // 045/079
01THE CHALLENGE

TRADITIONAL MANIPULATION BREAKS ON NOVEL OBJECTS

Current manipulation approaches require extensive per-task training data or brittle hand-coded planners. VLA models struggle with novel objects they haven't seen during training. Hierarchical planners decompose tasks linearly but can't reason about spatial relationships or recover from mid-execution failures.

Real-world manipulation demands zero-shot generalization to unseen objects, spatial reasoning about complex scenes, collision-aware trajectory generation, and graceful recovery when things go wrong. Linear task plans collapse when the environment doesn't match expectations.

02THE SOLUTION

WHAT MIDAS DELIVERS

MIDAS combines Vision-Language Model planning with dynamic scene graphs to create agentic operational graphs. It decomposes manipulation tasks into graph-structured plans where each node represents a sub-action, edges encode dependencies, and the graph restructures in real-time based on scene perception.

CAPABILITIES

  • VLM-driven task decomposition into directed acyclic operation graphs
  • Dynamic scene graph construction from RGB-D input — real-time object tracking and relation inference
  • Collision-aware trajectory generation using spatial graph constraints
  • Runtime failure detection and automatic graph restructuring for recovery
  • Zero-shot transfer to novel objects via semantic scene understanding
  • Integrates AZOTH + PROTEUS + ERGON + CHIRON modules from ANIMA suite
03ENGINEERING

WHY THIS IS HARD

Building zero-shot manipulation from scene graphs requires solving multiple coupled problems:

  1. 01Scene graph construction: extracting object nodes, spatial relations, and affordances from raw sensor data in real-time
  2. 02Task graph planning: VLM must generate valid directed acyclic graphs where node ordering respects physical constraints
  3. 03Trajectory collision avoidance: paths must be planned through cluttered scenes using graph-encoded spatial relationships
  4. 04Failure detection: recognizing when execution diverges from the plan requires continuous scene graph comparison
  5. 05Graph restructuring: re-planning mid-execution without starting from scratch — surgical modifications to the operational graph

MIDAS achieves 22.5-25% improvement over baselines by treating manipulation as graph optimization rather than sequential planning — enabling parallel sub-task execution, constraint propagation, and local failure recovery.

04BENCHMARKS

PROOF, NOT PROMISES

Measured against VLA and hierarchical baselines:

PROOF, NOT PROMISES
METRICVALUE
Improvement Over VLA Baselines22.5%
Improvement Over Hierarchical25%
Training RequirementZero-shot
Mobile Platform TransferDirect
Planning LatencyReal-time
Scene Graph Update Rate30 Hz
Failure RecoveryAutomatic
Object GeneralizationNovel / Unseen
Task DecompositionDAG-structured
Collision AvoidanceGraph-constrained
ANIMA IntegrationAZOTH + PROTEUS + ERGON + CHIRON
Trajectory PlanningCollision-aware
05BUILD STATUS

WHAT'S BUILT TODAY

3/6 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Scene Graph EngineCOMPLETEReal-time object graph construction from RGB-D — 30 Hz update rate
VLM PlanningCOMPLETETask decomposition into directed acyclic operation graphs
Trajectory GenerationIN PROGRESSCollision-aware path planning through cluttered scenes
Failure RecoveryIN PROGRESSRuntime graph restructuring and re-planning
Core modelsCOMPLETEVLM + scene graph + trajectory models trained and validated
API layerIN PROGRESSPublic REST + gRPC API — pending integration testing
06APPLICATIONS

WHERE MIDAS DEPLOYS

  • APP_01

    SEARCH & RESCUE

    Manipulate debris, open doors, and clear obstacles in disaster environments with zero prior training on specific objects.

  • APP_02

    EOD / HAZMAT

    Handle unknown explosive ordnance and hazardous materials through scene-graph-driven manipulation without object-specific models.

  • APP_03

    UNSTRUCTURED ENVIRONMENTS

    Operate in cluttered, dynamic spaces where pre-programmed manipulation sequences fail — warehouses, field operations, disaster zones.

  • APP_04

    LOGISTICS AUTOMATION

    Pick and place novel items in warehouse settings without per-SKU training or fixture-based solutions.

  • APP_05

    FIELD ROBOTICS

    Autonomous manipulation in agricultural, construction, and mining environments with unpredictable object configurations.

  • APP_06

    HUMAN-ROBOT COLLABORATION

    Safe co-manipulation with humans through real-time scene understanding and collision-aware trajectory adaptation.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: AGENTIC OPERATIONAL GRAPHS

  • VLM-driven task decomposition into directed acyclic graphs (DAGs)
  • Dynamic scene graph construction — nodes (objects), edges (spatial relations)
  • Graph-constrained trajectory optimization with collision avoidance
  • Runtime failure detection via scene graph state comparison

KEY INNOVATION

Treats manipulation as graph optimization rather than sequential planning. The operational graph enables parallel sub-task execution, constraint propagation across nodes, and surgical local modifications for failure recovery — without replanning the entire task.

ANIMA INTEGRATION

  • AZOTH — foundational VLM reasoning for task graph generation
  • PROTEUS — adaptive embodiment interface for cross-platform execution
  • ERGON — action primitive library for graph node execution
  • CHIRON — runtime monitoring and failure detection

PROCESSING PIPELINE

PERCEPTION
RGB-D → scene graph at 30 Hz — object detection, pose estimation, relation inference
PLANNING
VLM generates operational DAG — validated against scene graph constraints
EXECUTION
Graph nodes execute as collision-aware trajectories — parallel where dependencies allow
08PAPERS

RESEARCH PAPERS

  1. [01]Scene Graphs for Robot Manipulation — Johnson et al.
  2. [02]Vision-Language Models for Robotic Planning — Driess et al.
  3. [03]Agentic Task Decomposition with Directed Acyclic Graphs