Skip to content
RFL_GLOBAL
中文
  • SMOLVLA // HUGGINGFACE
  • 450M PARAMS
  • 15× COMPRESSION

PYGMALION

FROM CLAY TO CREATION

Robotic manipulation at scale requires models small enough to run locally but capable enough for complex tasks. PYGMALION is a 450M parameter SmolVLA that predicts robot actions from images and language instructions. 15x smaller than 7B baseline models. ~45ms inference. Runs on Apple Silicon and CUDA. From clay to creation — instruction in, action out.

MODULE STATUS: ACTIVE

Model Size

450M

DIVISION
ANIMA
WAVE
W4
DOMAIN
ACTION
WAVE 4 // ANIMA SUITE
COMPACT VLA MODEL
PYGMALION // W4 // 063/079
01THE CHALLENGE

TOO BIG TO DEPLOY

Robotic manipulation at scale requires models small enough to run locally on robots but capable enough to handle complex tasks. Existing VLAs are 7B+ parameters, requiring expensive GPUs. There's no middle ground.

Most manipulation stacks force a choice: pay for cloud inference with latency and cost, or deploy models that can't understand complex scenes. PYGMALION breaks this trade-off with a 450M model that matches 7B performance.

02THE SOLUTION

WHAT PYGMALION DELIVERS

PYGMALION is a 450M parameter SmolVLA that predicts robot actions from images and natural language. Image + instruction in, 7D action trajectory out.

PIPELINE

  1. 01450M parameters: 15x smaller than 7B baseline VLA models
  2. 02Input: RGB image (224×224) + text instruction
  3. 03Output: 7D robot actions (6D pose + gripper) over 10 timesteps
  4. 04~45ms single inference latency via REST API

CAPABILITIES

  • EDGE INFERENCESmolViT vision encoder (384 features) + SmolLM2 language model (576 features)→ MLX Apple Silicon (M1–M5, ~45ms on M5, 4x M1) + CUDA + CPU
  • REAL-TIME STREAMINGWebSocket streaming for real-time action prediction + REST + batch→ Per-image and batch inference with configurable max batch
  • AUTO DEVICEDevice detection with MPS/CUDA/CPU auto-selection→ First-class support for Apple Silicon, NVIDIA GPUs, and CPU fallback
03ENGINEERING

WHY THIS IS HARD

Vision-language-action modeling requires three encoder-decoder pipelines working in lockstep:

  1. 01Vision encoder must extract spatial features without destroying fine-grained manipulation details
  2. 02Language encoder must ground instructions in visual context — not just embed text, but align with image
  3. 03Action decoder must predict robot joint trajectories in a learned action space (7D normalized)
  4. 04Parameter-efficient fusion: small vision + small language + lightweight adapter = 450M total
  5. 05Comparable performance to 7B models with 15x fewer parameters requires careful architecture design

SmolVLA solves this with parameter-efficient fusion. SmolViT (384-dim) + SmolLM2 (576-dim) + lightweight fusion adapter. Production-ready on Apple Silicon and CUDA.

04BENCHMARKS

REAL HARDWARE PERFORMANCE

Measured with SmolVLA 450M pipeline:

REAL HARDWARE PERFORMANCE
METRICVALUE
Model Size450M parameters
Baseline Comparison7B+ (15x smaller)
Action Dimension7-DOF (6D pose + gripper)
Prediction Horizon10 timesteps
Single Inference~45ms (REST API)
Apple M5 (MLX)~45ms (4x M1)
Vision EncoderSmolViT (384 features)
Language EncoderSmolLM2 (576 features)
Max Batch Size16 (configurable)
05BUILD STATUS

WHAT'S BUILT TODAY

9/12 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Core modelsCOMPLETESmolVLA 450M architecture with production weights
API layerIN PROGRESSPublic API, pending dataset infrastructure
Model LoadingCOMPLETEHuggingFace integration, local caching
REST APICOMPLETE/health, /predict, /batch_predict, /info
WebSocket APICOMPLETEReal-time streaming via ws://
Device DetectionCOMPLETEAuto (MPS > CUDA > CPU)
Apple Silicon (MPS)COMPLETEFirst-class MLX-ready packaging
Prometheus MetricsCOMPLETERequest latency, inference time, throughput
Docker / ComposeCOMPLETEMulti-container with health checks
Unit TestsCOMPLETETest suite with pytest coverage
gRPC InterfacePLANNEDArchitecture planned, not implemented
ROS2 IntegrationPLANNEDTopic-based VLA action subscription
06APPLICATIONS

WHERE PYGMALION DEPLOYS

  • APP_01

    ROBOT ARM MANUFACTURERS

    Deploy manipulation without cloud dependency. On-device VLA inference for pick-and-place, assembly, and general manipulation tasks.

  • APP_02

    LOGISTICS & WAREHOUSE

    Edge inference for pick-and-place operations. Language-guided sorting and palletizing without cloud round-trips.

  • APP_03

    RESEARCH LABS

    Baseline VLA for imitation learning experiments. Small enough to iterate fast, large enough to produce real results.

  • APP_04

    MOBILE MANIPULATORS

    On-board model for collaborative tasks. Real-time action prediction on mobile platforms with limited compute.

  • APP_05

    EDGE AI DEVELOPERS

    Proof-of-concept for small models at scale. Demonstrate VLA capabilities on consumer hardware.

  • APP_06

    ASSEMBLY AUTOMATION

    Language-guided robotic assembly steps. Instruction-driven manipulation for flexible manufacturing lines.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: SMOLVLA (HUGGINGFACE LEROBOT)

  • SmolViT vision encoder: 384-dimensional patch embeddings
  • SmolLM2 language model: 576-dimensional token embeddings
  • Fusion adapter: cross-modal attention + action decoder
  • Action space: 7D normalized trajectory (joint control)

PYGMALION IMPLEMENTATION

  • FastAPI REST + WebSocket for real-time action streaming
  • Per-image and batch inference with configurable max batch
  • Device detection with MPS/CUDA/CPU auto-selection
  • Prometheus metrics and Docker deployment

INTEGRATION POINTS

  • Feeds action trajectories to CHIRON (control loops)
  • Provides manipulation plans to ERGON (pose estimation)
  • Consumes scene understanding from LOGOS (language grounding)
08PAPERS

RESEARCH BASIS

  1. [01]SmolVLA: A Small Vision-Language-Action Model — HuggingFace LeRobot, Apache-2.0