- TACTIS MODULE
- ANYTOUCH2
- CROSS-SENSOR
HAPTOS
TACTILE REPRESENTATION LEARNING
AnyTouch2 approach for general tactile representation learning. Learns optical tactile embeddings generalizing across sensor types. Static property understanding (texture/hardness), dynamic action-aware perception (slip detection), force-aware reasoning (grasp pressure). One embedding space for all touch.
MODULE STATUS: DEVELOPMENTCross-Sensor Transfer
TOUCH
- DIVISION
- ANIMA
- WAVE
- W5
- DOMAIN
- HARDWARE & SENSORS
- WAVE 5 // ANIMA SUITE
- PERCEPTION — TACTILE REPRESENTATION LEARNING
TACTILE SENSORS SPEAK DIFFERENT LANGUAGES
Every optical tactile sensor — GelSight, DIGIT, GelSlim, Soft-bubble — produces different image patterns for the same physical contact. Models trained on one sensor fail catastrophically on another. This fragmentation means every new robot hand, every new sensor revision, requires retraining from scratch. Tactile perception is stuck in a sensor-specific silo.
Manipulation requires understanding texture, hardness, slip, and force simultaneously. Current approaches handle these as separate tasks with separate models. A robot needs a unified tactile representation that works across sensors, understands static properties, detects dynamic events, and reasons about applied forces — all from the same embedding.
WHAT HAPTOS DELIVERS
HAPTOS learns general-purpose optical tactile embeddings that generalize across sensor types. Built on the AnyTouch2 framework, it provides static property understanding, dynamic action-aware perception, and force-aware reasoning — all from a single learned representation.
CAPABILITIES
- Cross-sensor generalization — embeddings trained on GelSight transfer to DIGIT, GelSlim, and novel sensors
- Static property understanding — texture classification, hardness estimation, surface roughness from contact images
- Dynamic action-aware perception — real-time slip detection, contact event classification, manipulation phase recognition
- Force-aware reasoning — grasp pressure estimation, contact force distribution, load prediction from optical deformation
- Unified embedding space — 256-dim vectors encoding contact geometry, material properties, and force state
- Zero-shot sensor adaptation — new sensor types require only a lightweight calibration pass, not full retraining
WHY THIS IS HARD
Building cross-sensor tactile representations requires solving deeply intertwined perceptual problems:
- 01Sensor domain gap: optical tactile sensors have vastly different illumination patterns, gel geometries, and camera configurations — bridging these requires learning invariant contact features
- 02Static vs. dynamic: texture and hardness are spatial properties from single frames, while slip and contact events are temporal — the representation must encode both without interference
- 03Force estimation: mapping optical deformation patterns to physical force vectors requires precise calibration and nonlinear gel mechanics modeling
- 04Scale ambiguity: the same embedding must distinguish microscopic surface textures and macroscopic contact geometry across different sensor resolutions
- 05Data scarcity: collecting paired tactile data across multiple sensors on identical objects is extremely labor-intensive — self-supervised and contrastive learning are essential
HAPTOS resolves this through a multi-task contrastive learning framework that aligns cross-sensor embeddings while preserving task-specific information through dedicated projection heads for static, dynamic, and force modalities.
PROOF, NOT PROMISES
Target performance metrics for production deployment:
| METRIC | VALUE |
|---|---|
| Cross-Sensor Transfer | >85% accuracy (zero-shot) |
| Texture Classification | 20+ material classes |
| Slip Detection Latency | <15ms from onset |
| Force Estimation Error | <0.3N (normal force) |
| Embedding Dimension | 256-d unified space |
| Sensor Types Supported | GelSight, DIGIT, GelSlim, Soft-bubble |
| Inference Rate | 100+ Hz per sensor |
| Contact Resolution | Sub-millimeter spatial |
| Training Data | 500K+ contact samples |
| Adaptation (new sensor) | <1000 calibration frames |
| Dynamic Event Classes | Slip, stick, roll, lift-off |
| ANIMA Integration | TACTIS |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | Tactile encoder backbone trained on multi-sensor dataset |
| Tactile Embedding Network | COMPLETE | Contrastive learning with cross-sensor alignment loss |
| Cross-Sensor Adapter | IN PROGRESS | Lightweight domain bridges for GelSight↔DIGIT transfer |
| Slip Detection Module | IN PROGRESS | Temporal convolution on embedding sequences — tuning thresholds |
| Force Reasoning | PLANNED | Gel deformation → force vector mapping — architecture designed |
| API layer | IN PROGRESS | Streaming inference endpoint for real-time tactile data |
| Texture Classifier | COMPLETE | 20-class material recognition from static contact |
| Hardness Estimator | COMPLETE | Shore A scale prediction from indentation depth |
| Evaluation Suite | PLANNED | Cross-sensor benchmark with standardized test objects |
WHERE HAPTOS DEPLOYS
- APP_01
DEXTEROUS MANIPULATION
Real-time tactile feedback for multi-finger grasping — slip detection prevents drops, force reasoning prevents crushing.
- APP_02
QUALITY INSPECTION
Automated surface defect detection through tactile scanning — texture anomalies invisible to cameras become obvious to touch.
- APP_03
MATERIAL SORTING
Classify materials by texture, hardness, and compliance — sorting recyclables, fabrics, or food items by touch alone.
- APP_04
SURGICAL ROBOTICS
Force-aware tissue manipulation — distinguish tissue types by tactile response, maintain safe contact forces.
- APP_05
PROSTHETICS
Restore tactile feedback to prosthetic hands — cross-sensor embeddings adapt to any integrated sensor type.
- APP_06
HUMAN-ROBOT HANDOVER
Detect grip transitions and slip events during object handovers — safe, natural physical interaction.
UNDER THE HOOD
FOUNDATION: ANYTOUCH2 FRAMEWORK
- Multi-task contrastive learning with cross-sensor alignment objective
- Vision transformer backbone with tactile-specific patch tokenization
- Sensor-agnostic contact feature extraction from optical tactile images
- Projection heads for static (texture/hardness), dynamic (slip/event), and force modalities
KEY INNOVATION
HAPTOS decouples sensor-specific appearance from sensor-invariant contact physics through a two-stage architecture: a sensor adapter normalizes raw optical images into a canonical representation, then a shared tactile encoder extracts unified embeddings. This enables zero-shot transfer to new sensors while preserving fine-grained contact information.
PERCEPTION PIPELINE
- Raw tactile image → sensor adapter → canonical contact representation
- Contact representation → tactile encoder → 256-d unified embedding
- Embedding → task-specific heads: texture, slip, force, material
- Streaming inference at 100+ Hz for real-time manipulation control
SUPPORTED SENSORS
- GelSight
- High-resolution gel-based sensor — photometric stereo contact imaging
- DIGIT
- Compact optical tactile sensor — designed for multi-finger integration
- GelSlim
- Thin-profile gel sensor — low-profile for parallel-jaw grippers
RESEARCH PAPERS
- [01]AnyTouch2: General Tactile Representation Learning (2024)
- [02]Optical Tactile Sensing for Robotic Manipulation — Survey
- [03]Contrastive Learning for Cross-Modal Sensor Transfer