Skip to content
RFL_GLOBAL
中文
  • MANIP-RAAP // W7
  • Python
  • ANIMA

MANIP-RAAP

The Robot That Learned to Grip by Watching

Logistics robots, explosive ordnance disposal systems, and field maintenance platforms fail on novel objects — configurations they have never seen in training. Hardcoded grasp libraries cover known items; the real world is full of improvised, damaged, or unfamiliar objects where brittle rule-based grasping breaks down at the worst moment. MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment.

MODULE STATUS: PRODUCTION

Task Coverage

PROD

DIVISION
GENERIC
WAVE
W7
DOMAIN
GENERAL
WAVE 7 // ANIMA SUITE
FOUNDATION — MANIP-RAAP
MANIP-RAAP // W7 // 004/012
01THE CHALLENGE

THE PROBLEM WE SOLVE

Logistics robots, explosive ordnance disposal systems, and field maintenance platforms fail on novel objects — configurations they have never seen in training. Hardcoded grasp libraries cover known items; the real world is full of improvised, damaged, or unfamiliar objects where brittle rule-based grasping breaks down at the worst moment.

A field robot encountering an unknown piece of equipment or debris can predict viable grasp points on first encounter, using the same demonstrated knowledge base that a human technician would consult — dramatically expanding operational flexibility.

02THE SOLUTION

WHAT MANIP-RAAP DELIVERS

MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment. Given a target image, the module retrieves visually similar demonstrations from a reference bank, aligns their affordance annotations across the image pair, and transfers predicted contact points and action directions to the query.

CAPABILITIES

  • MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment
  • Given a target image, the module retrieves visually similar demonstrations from a reference bank, aligns their affordance annotations across the image pair, and transfers predicted contact points and action directions to the query
  • The result is robust affordance prediction that generalizes to novel objects without retraining — purely through retrieval.
03ENGINEERING

WHY THIS IS HARD

Building MANIP-RAAP requires solving multiple coupled problems:

  1. 01Logistics robots, explosive ordnance disposal systems, and field maintenance platforms fail on novel objects — configurations they have never seen in training
  2. 02Hardcoded grasp libraries cover known items; the real world is full of improvised, damaged, or unfamiliar objects where brittle rule-based grasping breaks down at the worst moment.
  3. 03MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment
  4. 04Given a target image, the module retrieves visually similar demonstrations from a reference bank, aligns their affordance annotations across the image pair, and transfers predicted contact points and action directions to the query

MANIP-RAAP solves these through careful architecture design and rigorous validation.

04BENCHMARKS

PROOF, NOT PROMISES

Key performance metrics:

PROOF, NOT PROMISES
METRICVALUE
Task CoverageContact point estimation + action direction prediction — both outputs in one inference call
GeneralizationNovel object handling via retrieval — zero retraining required for new object classes
InferenceSingle RGB image input — no depth, no CAD model, no object identity required
Defense AngleEOD robot grasping of improvised devices, logistics handling of unfamiliar cargo, field maintenance robotics
05BUILD STATUS

WHAT'S BUILT TODAY

2/4 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Task CoverageCOMPLETEContact point estimation + action direction prediction — both outputs in one inference call
GeneralizationCOMPLETENovel object handling via retrieval — zero retraining required for new object classes
InferenceIN PROGRESSSingle RGB image input — no depth, no CAD model, no object identity required
Defense AngleIN PROGRESSEOD robot grasping of improvised devices, logistics handling of unfamiliar cargo, field maintenance robotics
06APPLICATIONS

WHERE MANIP-RAAP DEPLOYS

  • APP_01

    AUTONOMOUS SYSTEMS

    A field robot encountering an unknown piece of equipment or debris can predict viable grasp points on first encounter, using the same demonstrated knowledge base that a human technician would consult — dramatically expanding operational flexibility.

  • APP_02

    RESEARCH LABS

    MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment.

  • APP_03

    EDGE COMPUTING

    Retrieval-augmented affordance prediction that tells a robot exactly where to grab and which way to push — from a single RGB image.