Skip to content
RFL_GLOBAL
中文
  • ECMR 2025 // FAB-NAV
  • SEMANTIC GOALS
  • BEHAVIOR TREES

HERMES

THE INTELLIGENT NAVIGATOR

Most robot navigation uses geometry: build a map, avoid obstacles. This fails in human environments. HERMES combines SmolVLM perception, SmolLM2 reasoning, and composable behavior trees for semantic navigation — "Go to the kitchen", not grid coordinates. Understands hazards, adapts strategies in real-time.

MODULE STATUS: ACTIVE

Perception Latency (SmolVLM)

50Hz

DIVISION
ANIMA
WAVE
W3
DOMAIN
ACTION
WAVE 4 // ANIMA SUITE
AI-POWERED NAVIGATION
HERMES // W3 // 035/079
01THE CHALLENGE

GEOMETRY IS NOT UNDERSTANDING

Most robot navigation uses geometry: build a map, compute shortest path, avoid obstacles detected by LiDAR. This fails in unstructured human environments. A robot doesn't know if a door is locked, if a wet floor is slippery, if a person is about to move.

Humans navigate by understanding: we recognize objects, assess hazards, make semantic decisions. A wet floor deserves a different strategy than a moving person. A closed door requires finding an alternate route. HERMES brings this understanding to autonomous robots.

02THE SOLUTION

WHAT HERMES DELIVERS

A three-stage pipeline: perceive the world, reason about it, act through composable behaviors.

PIPELINE

  1. 01PERCEPTIONSmolVLM (256M): RGB-D → objects, hazards, scene description at 2-5 fps
  2. 02REASONINGSmolLM2 (1.7B): scene understanding → optimal action selection with confidence
  3. 03EXECUTIONBehavior Trees: composable navigation behaviors at 10-50 Hz tick rate

CAPABILITIES

  • Semantic goal navigation: "Go to the kitchen" (not grid coordinates)
  • Hazard-aware planning: wet floor, moving person, closed door → different strategies
  • Semantic costmap layer: hazard costs, preferred routes overlaid on Nav2-style maps
  • WebSocket API for real-time navigation state streaming
  • ROS2 integration: RGB-D, odometry, cmd_vel, semantic costmap topics
  • Device support: MLX (M1–M5, ~150ms on M5), MPS, CUDA, CPU
03ENGINEERING

WHY THIS IS HARD

Semantic navigation requires solving three coupled problems simultaneously:

  1. 01Perception at scale: extract objects, hazards, and traversability from RGB-D at real-time rates (2–5 fps)
  2. 02Reasoning with context: use an LLM to select actions based on goals and scene, not just geometry
  3. 03Behavior trees for adaptation: compose reusable navigation behaviors that adapt without full replan
  4. 04Model orchestration: coordinate VLM + LLM + BT framework across async sensor inputs
  5. 05Safety guarantees: emergency stop, safety radius, velocity limits within the BT execution loop

HERMES achieves this with SmolVLM for lightweight perception, SmolLM2 for reasoning, and composable behavior tree nodes — all integrated with ROS2 topics for real robot deployment.

04BENCHMARKS

REAL HARDWARE PERFORMANCE

Measured with SmolVLM + SmolLM2 on real sensor data:

REAL HARDWARE PERFORMANCE
METRICVALUE
Perception Latency (SmolVLM)200–400ms
Reasoning Latency (SmolLM2)150–300ms
Behavior Tree Tick Rate10–50 Hz (configurable)
BT Tick Latency10–50ms
Perception Throughput2–5 frames/second
Apple M5 (MLX)~150ms perception (4× M1)
VLM Memory~1.2 GB
LLM Memory~3.5 GB
Max Linear Velocity0.5 m/s (configurable)
Safety Radius0.3 m (configurable)
05BUILD STATUS

WHAT'S BUILT TODAY

10/12 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETESmolVLM + SmolLM2 weights — production-ready
PerceptionModule (SmolVLM)COMPLETEObject detection, scene description, hazard ID
ReasoningModule (SmolLM2)COMPLETEAction selection, path planning, confidence
Behavior Tree FrameworkCOMPLETENodes, composites, tick execution
SemanticCostmapLayerCOMPLETEHazard encoding, preferred routes
FABNavigator OrchestratorCOMPLETEPerception + reasoning + execution integration
FastAPI ServerCOMPLETE/navigate, /observe, /plan, /semantic_map
WebSocket APICOMPLETEReal-time navigation state streaming
ROS2 IntegrationCOMPLETETopics: RGB-D, odom, cmd_vel, costmap
Docker DeploymentCOMPLETEMulti-container with Prometheus/Grafana
API layerIN PROGRESSPublic API — pending dataset infrastructure
Nav2 IntegrationPLANNEDNav2 stack compatibility design
06APPLICATIONS

WHERE HERMES DEPLOYS

  • APP_01

    MOBILE ROBOTICS

    Semantic indoor navigation for delivery, patrolling, and exploration.

  • APP_02

    SERVICE ROBOTS

    Understanding human environments — adapting behavior to context.

  • APP_03

    WAREHOUSE AUTOMATION

    Navigate around dynamic obstacles and hazards with semantic awareness.

  • APP_04

    HOSPITALITY ROBOTS

    Greeting visitors, avoiding collisions in crowded spaces.

  • APP_05

    RESEARCH PLATFORMS

    Benchmark for semantic navigation and behavior tree execution.

  • APP_06

    DROP-IN NAVIGATION

    Intelligent navigation layer for any ROS2 mobile base.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: FAB-NAV (ECMR 2025)

  • Foundation-Model-Based Action Selection for Behavior Trees
  • Three-stage pipeline: Perception → Reasoning → Execution
  • SmolVLM (256M) + SmolLM2 (1.7B) for lightweight semantic intelligence
  • Composable behavior tree nodes for real-time adaptive navigation
  • Semantic costmap layer with hazard costs and preferred routes

MODEL STACK

SmolVLM (Perception)
256M params · 200–400ms
SmolLM2 (Reasoning)
1.7B params · 150–300ms
BT Framework (Execution)
~0 · 10–50ms

ROS2 INTEGRATION

  • Subscriptions: /camera/rgb, /camera/depth, /odom, /tf
  • Publications: /cmd_vel, /hermes/semantic_costmap, /hermes/navigation_state
  • Services: hermes/navigate, hermes/observe, hermes/stop
08PAPERS

RESEARCH PAPER

  1. [01]Foundation-Model-Based Action Selection for Behavior Trees in Navigation — ECMR 2025