- FUSIONSENSE // ICRA 2025
- CROSS-MODAL
- TACTILE FUSION
TACTIS
THE KNOWING TOUCH
Vision alone cannot reconstruct surfaces that cameras can't see: inside grippers, under occluding fingers, in shiny materials. TACTIS fuses RGB-D vision and tactile sensor data into a learned cross-modal representation for real-time 3D reconstruction during grasping and manipulation. Built on FusionSense (ICRA 2025).
MODULE STATUS: ACTIVEVISION + TOUCH // FUSED
SUB-MM
- DIVISION
- ANIMA
- WAVE
- W3
- DOMAIN
- FOUNDATION
- WAVE 3 // ANIMA SUITE
- VISION-TOUCH FUSION
VISION CAN'T FEEL
Vision alone cannot reconstruct surfaces that cameras can't see: inside grippers, under occluding fingers, in shiny materials that blind RGB-D sensors. Tactile sensors reveal geometry that vision misses, but require contact — slow and potentially dangerous.
Combining vision and touch enables fast, safe, detailed reconstruction. But learning a fused representation requires paired training data from manipulation (expensive) and new architectures (vision features ≠ touch features). TACTIS solves this.
WHAT TACTIS DELIVERS
TACTIS is a containerized vision-touch fusion reconstruction service built on FusionSense (ICRA 2025). Fuse vision + touch, reconstruct occluded surfaces, quantify materials.
PIPELINE
- 01Fuses vision and touch — learned shared 3D representation from RGB-D + tactile sensors
- 02Reconstructs occluded regions — predicts hidden geometry from touch feedback
- 03Quantifies material properties — infers stiffness, friction, and deformability
- 04Flexible outputs — meshes, point clouds, or signed distance fields
CAPABILITIES
- CROSS-MODAL LEARNINGContrastive learning on grasping and manipulation trajectories→ End-to-end fusion from raw multi-modal sensor data
- DEVICE FLEXIBILITYMLX Apple Silicon (M1–M5, ~50ms on M5, 4× M1) + CUDA + CPU→ Run anywhere: laptop, workstation, robot
- REAL-TIME APISREST + gRPC APIs for real-time multi-modal fusion in control loops→ Sub-50ms latency for online decision-making
WHY THIS IS HARD
Vision and touch operate at different scales and modalities — fusing them is fundamentally hard:
- 01RGB images have millions of pixels; tactile arrays have 10s–100s of sensors — different scales entirely
- 02Cross-modal embedding: learning a shared latent space where visual and tactile observations align
- 03Contrastive learning on manipulation trajectories requires careful data collection and alignment
- 04Real-time inference in robotic control loops — no batch processing, sub-50ms required
- 05Supporting diverse tactile hardware: GelSight, capacitive arrays, force-torque sensors
FusionSense trains this embedding end-to-end using contrastive learning on manipulation trajectories. TACTIS makes it deployable with proper device abstraction and real-time APIs.
REAL HARDWARE PERFORMANCE
Measured with FusionSense cross-modal pipeline:
| METRIC | VALUE |
|---|---|
| Fusion Latency | ~50ms (M5 MLX), ~15ms (RTX 4090) |
| Apple M5 (MLX) | ~50ms per fusion step (4× M1) |
| Surface Resolution | Sub-millimeter from touch |
| Output Formats | Mesh, point cloud, SDF |
| Material Properties | Stiffness, friction, deformability |
| Tactile Sensors | GelSight, Xela uSkin, Weiss, ATI |
| Vision Input | RGB-D camera |
| Learning Method | Cross-modal contrastive |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | FusionSense cross-modal pipeline, production-ready |
| API layer | IN PROGRESS | Public API, pending dataset infrastructure |
| REST API | COMPLETE | FastAPI with multi-modal input schema |
| gRPC Service | COMPLETE | Protobuf vision-touch protocol |
| Vision Backbone | COMPLETE | RGB-D feature extraction |
| Touch Encoder | COMPLETE | Tactile array embedding |
| Cross-Modal Fusion | COMPLETE | Learned shared representation |
| Surface Reconstruction | COMPLETE | Mesh, point cloud, SDF generation |
| Material Inference | COMPLETE | Stiffness and friction prediction |
| Quality Gates | COMPLETE | ruff, pytest, mypy passing |
WHERE TACTIS DEPLOYS
- APP_01
DEXTEROUS MANIPULATION
Complex grasping and in-hand reorientation with real-time surface reconstruction for adaptive grip control.
- APP_02
PROSTHETICS & REHABILITATION
Detailed hand-object feedback enabling natural interaction and precise force modulation.
- APP_03
FOOD & FRAGILE GOODS
Grip force precisely controlled for delicate items — bruise-free produce handling and glass manipulation.
- APP_04
ASSEMBLY AUTOMATION
High-precision contact detection during assembly — sub-millimeter alignment from tactile feedback.
- APP_05
HAPTIC TELEOPERATION
Detailed surface models for remote operation — operators feel what the robot touches.
- APP_06
SURGICAL ROBOTICS
Tactile feedback for delicate procedures — tissue stiffness and deformation mapping in real time.
UNDER THE HOOD
FOUNDATION: FUSIONSENSE (ICRA 2025)
- Vision-touch fusion for manipulation-aware 3D reconstruction
- Cross-modal contrastive learning on grasping tasks
- Apache-2.0 licensed reference implementation
- End-to-end learned shared latent space
TACTIS IMPLEMENTATION
- FastAPI REST with RGB-D + tactile sensor input schema
- gRPC for real-time multi-modal fusion
- Device abstraction: MLX → CUDA → CPU
- Vision backbone: CNN or ViT for depth/RGB
- Touch encoder: dense tactile array processing
INTEGRATION POINTS
- Consumes manipulator state from NEXUS (semantic understanding)
- Provides surface geometry to downstream grasp planning
- Feeds reconstruction confidence to HARMONIA (sensor weighting)
TACTILE HARDWARE
- GelSight-style optical tactile sensors
- Capacitive arrays (Xela uSkin, Weiss Robotics)
- Force-torque sensors (ATI, JR3)
- Custom modalities via sensor calibration protocol
RESEARCH BASIS
- [01]FusionSense — vision-touch fusion for manipulation 3D reconstruction (ICRA 2025)