- MANIP-RAAP // W7
- Python
- ANIMA
MANIP-RAAP
The Robot That Learned to Grip by Watching
Logistics robots, explosive ordnance disposal systems, and field maintenance platforms fail on novel objects — configurations they have never seen in training. Hardcoded grasp libraries cover known items; the real world is full of improvised, damaged, or unfamiliar objects where brittle rule-based grasping breaks down at the worst moment. MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment.
MODULE STATUS: PRODUCTIONTask Coverage
PROD
- DIVISION
- GENERIC
- WAVE
- W7
- DOMAIN
- GENERAL
- WAVE 7 // ANIMA SUITE
- FOUNDATION — MANIP-RAAP
THE PROBLEM WE SOLVE
Logistics robots, explosive ordnance disposal systems, and field maintenance platforms fail on novel objects — configurations they have never seen in training. Hardcoded grasp libraries cover known items; the real world is full of improvised, damaged, or unfamiliar objects where brittle rule-based grasping breaks down at the worst moment.
A field robot encountering an unknown piece of equipment or debris can predict viable grasp points on first encounter, using the same demonstrated knowledge base that a human technician would consult — dramatically expanding operational flexibility.
WHAT MANIP-RAAP DELIVERS
MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment. Given a target image, the module retrieves visually similar demonstrations from a reference bank, aligns their affordance annotations across the image pair, and transfers predicted contact points and action directions to the query.
CAPABILITIES
- MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment
- Given a target image, the module retrieves visually similar demonstrations from a reference bank, aligns their affordance annotations across the image pair, and transfers predicted contact points and action directions to the query
- The result is robust affordance prediction that generalizes to novel objects without retraining — purely through retrieval.
WHY THIS IS HARD
Building MANIP-RAAP requires solving multiple coupled problems:
- 01Logistics robots, explosive ordnance disposal systems, and field maintenance platforms fail on novel objects — configurations they have never seen in training
- 02Hardcoded grasp libraries cover known items; the real world is full of improvised, damaged, or unfamiliar objects where brittle rule-based grasping breaks down at the worst moment.
- 03MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment
- 04Given a target image, the module retrieves visually similar demonstrations from a reference bank, aligns their affordance annotations across the image pair, and transfers predicted contact points and action directions to the query
MANIP-RAAP solves these through careful architecture design and rigorous validation.
PROOF, NOT PROMISES
Key performance metrics:
| METRIC | VALUE |
|---|---|
| Task Coverage | Contact point estimation + action direction prediction — both outputs in one inference call |
| Generalization | Novel object handling via retrieval — zero retraining required for new object classes |
| Inference | Single RGB image input — no depth, no CAD model, no object identity required |
| Defense Angle | EOD robot grasping of improvised devices, logistics handling of unfamiliar cargo, field maintenance robotics |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Task Coverage | COMPLETE | Contact point estimation + action direction prediction — both outputs in one inference call |
| Generalization | COMPLETE | Novel object handling via retrieval — zero retraining required for new object classes |
| Inference | IN PROGRESS | Single RGB image input — no depth, no CAD model, no object identity required |
| Defense Angle | IN PROGRESS | EOD robot grasping of improvised devices, logistics handling of unfamiliar cargo, field maintenance robotics |
WHERE MANIP-RAAP DEPLOYS
- APP_01
AUTONOMOUS SYSTEMS
A field robot encountering an unknown piece of equipment or debris can predict viable grasp points on first encounter, using the same demonstrated knowledge base that a human technician would consult — dramatically expanding operational flexibility.
- APP_02
RESEARCH LABS
MANIP-RAAP (codename: Benzaiten) implements RAAP (ICRA 2026, arXiv:2603.29419), combining a retrieval-augmented inference loop with cross-image action alignment.
- APP_03
EDGE COMPUTING
Retrieval-augmented affordance prediction that tells a robot exactly where to grab and which way to push — from a single RGB image.