- ECMR 2025 // FAB-NAV
- SEMANTIC GOALS
- BEHAVIOR TREES
HERMES
THE INTELLIGENT NAVIGATOR
Most robot navigation uses geometry: build a map, avoid obstacles. This fails in human environments. HERMES combines SmolVLM perception, SmolLM2 reasoning, and composable behavior trees for semantic navigation — "Go to the kitchen", not grid coordinates. Understands hazards, adapts strategies in real-time.
MODULE STATUS: ACTIVEPerception Latency (SmolVLM)
50Hz
- DIVISION
- ANIMA
- WAVE
- W3
- DOMAIN
- ACTION
- WAVE 4 // ANIMA SUITE
- AI-POWERED NAVIGATION
GEOMETRY IS NOT UNDERSTANDING
Most robot navigation uses geometry: build a map, compute shortest path, avoid obstacles detected by LiDAR. This fails in unstructured human environments. A robot doesn't know if a door is locked, if a wet floor is slippery, if a person is about to move.
Humans navigate by understanding: we recognize objects, assess hazards, make semantic decisions. A wet floor deserves a different strategy than a moving person. A closed door requires finding an alternate route. HERMES brings this understanding to autonomous robots.
WHAT HERMES DELIVERS
A three-stage pipeline: perceive the world, reason about it, act through composable behaviors.
PIPELINE
- 01PERCEPTIONSmolVLM (256M): RGB-D → objects, hazards, scene description at 2-5 fps
- 02REASONINGSmolLM2 (1.7B): scene understanding → optimal action selection with confidence
- 03EXECUTIONBehavior Trees: composable navigation behaviors at 10-50 Hz tick rate
CAPABILITIES
- Semantic goal navigation: "Go to the kitchen" (not grid coordinates)
- Hazard-aware planning: wet floor, moving person, closed door → different strategies
- Semantic costmap layer: hazard costs, preferred routes overlaid on Nav2-style maps
- WebSocket API for real-time navigation state streaming
- ROS2 integration: RGB-D, odometry, cmd_vel, semantic costmap topics
- Device support: MLX (M1–M5, ~150ms on M5), MPS, CUDA, CPU
WHY THIS IS HARD
Semantic navigation requires solving three coupled problems simultaneously:
- 01Perception at scale: extract objects, hazards, and traversability from RGB-D at real-time rates (2–5 fps)
- 02Reasoning with context: use an LLM to select actions based on goals and scene, not just geometry
- 03Behavior trees for adaptation: compose reusable navigation behaviors that adapt without full replan
- 04Model orchestration: coordinate VLM + LLM + BT framework across async sensor inputs
- 05Safety guarantees: emergency stop, safety radius, velocity limits within the BT execution loop
HERMES achieves this with SmolVLM for lightweight perception, SmolLM2 for reasoning, and composable behavior tree nodes — all integrated with ROS2 topics for real robot deployment.
REAL HARDWARE PERFORMANCE
Measured with SmolVLM + SmolLM2 on real sensor data:
| METRIC | VALUE |
|---|---|
| Perception Latency (SmolVLM) | 200–400ms |
| Reasoning Latency (SmolLM2) | 150–300ms |
| Behavior Tree Tick Rate | 10–50 Hz (configurable) |
| BT Tick Latency | 10–50ms |
| Perception Throughput | 2–5 frames/second |
| Apple M5 (MLX) | ~150ms perception (4× M1) |
| VLM Memory | ~1.2 GB |
| LLM Memory | ~3.5 GB |
| Max Linear Velocity | 0.5 m/s (configurable) |
| Safety Radius | 0.3 m (configurable) |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | SmolVLM + SmolLM2 weights — production-ready |
| PerceptionModule (SmolVLM) | COMPLETE | Object detection, scene description, hazard ID |
| ReasoningModule (SmolLM2) | COMPLETE | Action selection, path planning, confidence |
| Behavior Tree Framework | COMPLETE | Nodes, composites, tick execution |
| SemanticCostmapLayer | COMPLETE | Hazard encoding, preferred routes |
| FABNavigator Orchestrator | COMPLETE | Perception + reasoning + execution integration |
| FastAPI Server | COMPLETE | /navigate, /observe, /plan, /semantic_map |
| WebSocket API | COMPLETE | Real-time navigation state streaming |
| ROS2 Integration | COMPLETE | Topics: RGB-D, odom, cmd_vel, costmap |
| Docker Deployment | COMPLETE | Multi-container with Prometheus/Grafana |
| API layer | IN PROGRESS | Public API — pending dataset infrastructure |
| Nav2 Integration | PLANNED | Nav2 stack compatibility design |
WHERE HERMES DEPLOYS
- APP_01
MOBILE ROBOTICS
Semantic indoor navigation for delivery, patrolling, and exploration.
- APP_02
SERVICE ROBOTS
Understanding human environments — adapting behavior to context.
- APP_03
WAREHOUSE AUTOMATION
Navigate around dynamic obstacles and hazards with semantic awareness.
- APP_04
HOSPITALITY ROBOTS
Greeting visitors, avoiding collisions in crowded spaces.
- APP_05
RESEARCH PLATFORMS
Benchmark for semantic navigation and behavior tree execution.
- APP_06
DROP-IN NAVIGATION
Intelligent navigation layer for any ROS2 mobile base.
UNDER THE HOOD
FOUNDATION: FAB-NAV (ECMR 2025)
- Foundation-Model-Based Action Selection for Behavior Trees
- Three-stage pipeline: Perception → Reasoning → Execution
- SmolVLM (256M) + SmolLM2 (1.7B) for lightweight semantic intelligence
- Composable behavior tree nodes for real-time adaptive navigation
- Semantic costmap layer with hazard costs and preferred routes
MODEL STACK
- SmolVLM (Perception)
- 256M params · 200–400ms
- SmolLM2 (Reasoning)
- 1.7B params · 150–300ms
- BT Framework (Execution)
- ~0 · 10–50ms
ROS2 INTEGRATION
- Subscriptions: /camera/rgb, /camera/depth, /odom, /tf
- Publications: /cmd_vel, /hermes/semantic_costmap, /hermes/navigation_state
- Services: hermes/navigate, hermes/observe, hermes/stop
RESEARCH PAPER
- [01]Foundation-Model-Based Action Selection for Behavior Trees in Navigation — ECMR 2025