Skip to content
RFL_GLOBAL
中文
  • GS-GHOST // W7
  • Python
  • ANIMA

GS-GHOST

Reconstruct the Grasp Before the Robot Reaches

Training manipulation policies for bomb disposal robots, EOD systems, or surgical field robots requires enormous libraries of demonstrated grasps. Motion capture is expensive, instrumented gloves are impractical in operational gear, and depth sensors fail in adverse lighting. Most collected video is unusable because there is no 3D ground truth. GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates.

MODULE STATUS: PRODUCTION

Modality

Mono RGB

DIVISION
GENERIC
WAVE
W7
DOMAIN
GENERAL
WAVE 7 // ANIMA SUITE
FOUNDATION — GS-GHOST
GS-GHOST // W7 // 002/012
01THE CHALLENGE

THE PROBLEM WE SOLVE

Training manipulation policies for bomb disposal robots, EOD systems, or surgical field robots requires enormous libraries of demonstrated grasps. Motion capture is expensive, instrumented gloves are impractical in operational gear, and depth sensors fail in adverse lighting. Most collected video is unusable because there is no 3D ground truth.

Defense and logistics operators can harvest 3D manipulation demonstrations from existing helmet cam and body cam footage — turning a library of operational video into a training dataset for dexterous robotic systems without a single depth sensor.

02THE SOLUTION

WHAT GS-GHOST DELIVERS

GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates.

CAPABILITIES

  • GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates
  • The Gaussian Splatting representation enables novel view rendering for data augmentation, while batch inference mode processes large demonstration libraries without per-video tuning.
03ENGINEERING

WHY THIS IS HARD

Building GS-GHOST requires solving multiple coupled problems:

  1. 01Training manipulation policies for bomb disposal robots, EOD systems, or surgical field robots requires enormous libraries of demonstrated grasps
  2. 02Motion capture is expensive, instrumented gloves are impractical in operational gear, and depth sensors fail in adverse lighting
  3. 03GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates
  4. 04The Gaussian Splatting representation enables novel view rendering for data augmentation, while batch inference mode processes large demonstration libraries without per-video tuning.

GS-GHOST solves these through careful architecture design and rigorous validation.

04BENCHMARKS

PROOF, NOT PROMISES

Key performance metrics:

PROOF, NOT PROMISES
METRICVALUE
ModalityMonocular RGB only — no depth, no markers, no object templates
Batch ModeParallel inference across demonstration libraries *(estimated, from batch_inference capability)*
Output3DGS representation — renderable from arbitrary viewpoints for augmentation pipelines
Defense AngleEOD and logistics robot training from operational body-cam footage; field demo data collection
05BUILD STATUS

WHAT'S BUILT TODAY

2/4 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
ModalityCOMPLETEMonocular RGB only — no depth, no markers, no object templates
Batch ModeCOMPLETEParallel inference across demonstration libraries *(estimated, from batch_inference capability)*
OutputIN PROGRESS3DGS representation — renderable from arbitrary viewpoints for augmentation pipelines
Defense AngleIN PROGRESSEOD and logistics robot training from operational body-cam footage; field demo data collection
06APPLICATIONS

WHERE GS-GHOST DEPLOYS

  • APP_01

    AUTONOMOUS SYSTEMS

    Defense and logistics operators can harvest 3D manipulation demonstrations from existing helmet cam and body cam footage — turning a library of operational video into a training dataset for dexterous robotic systems without a single depth sensor.

  • APP_02

    RESEARCH LABS

    GS-GHOST (codename: Fujin) implements GHOST (arXiv:2603.18912), a category-agnostic hand-object interaction reconstruction pipeline that takes ordinary monocular RGB video and outputs photorealistic 3D Gaussian reconstructions of both the hand and the held object — without category priors or object templates.

  • APP_03

    EDGE COMPUTING

    Monocular RGB video reconstructs hand-object interactions in real-time via Gaussian Splatting — unlocking imitation learning from raw video at scale.