- GS-GHOST // W7
- Python
- ANIMA
GS-GHOST
Reconstruct the Grasp Before the Robot Reaches
Training manipulation policies for bomb disposal robots, EOD systems, or surgical field robots requires enormous libraries of demonstrated grasps. Motion capture is expensive, instrumented gloves are impractical in operational gear, and depth sensors fail in adverse lighting. Most collected video is unusable because there is no 3D ground truth. GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates.
MODULE STATUS: PRODUCTIONModality
Mono RGB
- DIVISION
- GENERIC
- WAVE
- W7
- DOMAIN
- GENERAL
- WAVE 7 // ANIMA SUITE
- FOUNDATION — GS-GHOST
THE PROBLEM WE SOLVE
Training manipulation policies for bomb disposal robots, EOD systems, or surgical field robots requires enormous libraries of demonstrated grasps. Motion capture is expensive, instrumented gloves are impractical in operational gear, and depth sensors fail in adverse lighting. Most collected video is unusable because there is no 3D ground truth.
Defense and logistics operators can harvest 3D manipulation demonstrations from existing helmet cam and body cam footage — turning a library of operational video into a training dataset for dexterous robotic systems without a single depth sensor.
WHAT GS-GHOST DELIVERS
GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates.
CAPABILITIES
- GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates
- The Gaussian Splatting representation enables novel view rendering for data augmentation, while batch inference mode processes large demonstration libraries without per-video tuning.
WHY THIS IS HARD
Building GS-GHOST requires solving multiple coupled problems:
- 01Training manipulation policies for bomb disposal robots, EOD systems, or surgical field robots requires enormous libraries of demonstrated grasps
- 02Motion capture is expensive, instrumented gloves are impractical in operational gear, and depth sensors fail in adverse lighting
- 03GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates
- 04The Gaussian Splatting representation enables novel view rendering for data augmentation, while batch inference mode processes large demonstration libraries without per-video tuning.
GS-GHOST solves these through careful architecture design and rigorous validation.
PROOF, NOT PROMISES
Key performance metrics:
| METRIC | VALUE |
|---|---|
| Modality | Monocular RGB only — no depth, no markers, no object templates |
| Batch Mode | Parallel inference across demonstration libraries *(estimated, from batch_inference capability)* |
| Output | 3DGS representation — renderable from arbitrary viewpoints for augmentation pipelines |
| Defense Angle | EOD and logistics robot training from operational body-cam footage; field demo data collection |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Modality | COMPLETE | Monocular RGB only — no depth, no markers, no object templates |
| Batch Mode | COMPLETE | Parallel inference across demonstration libraries *(estimated, from batch_inference capability)* |
| Output | IN PROGRESS | 3DGS representation — renderable from arbitrary viewpoints for augmentation pipelines |
| Defense Angle | IN PROGRESS | EOD and logistics robot training from operational body-cam footage; field demo data collection |
WHERE GS-GHOST DEPLOYS
- APP_01
AUTONOMOUS SYSTEMS
Defense and logistics operators can harvest 3D manipulation demonstrations from existing helmet cam and body cam footage — turning a library of operational video into a training dataset for dexterous robotic systems without a single depth sensor.
- APP_02
RESEARCH LABS
GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates.
- APP_03
EDGE COMPUTING
Monocular RGB video reconstructs hand-object interactions in real-time via Gaussian Splatting — unlocking imitation learning from raw video at scale.