- SMOLVLA // HUGGINGFACE
- 450M PARAMS
- 15× COMPRESSION
PYGMALION
FROM CLAY TO CREATION
Robotic manipulation at scale requires models small enough to run locally but capable enough for complex tasks. PYGMALION is a 450M parameter SmolVLA that predicts robot actions from images and language instructions. 15x smaller than 7B baseline models. ~45ms inference. Runs on Apple Silicon and CUDA. From clay to creation — instruction in, action out.
MODULE STATUS: ACTIVEModel Size
450M
- DIVISION
- ANIMA
- WAVE
- W4
- DOMAIN
- ACTION
- WAVE 4 // ANIMA SUITE
- COMPACT VLA MODEL
TOO BIG TO DEPLOY
Robotic manipulation at scale requires models small enough to run locally on robots but capable enough to handle complex tasks. Existing VLAs are 7B+ parameters, requiring expensive GPUs. There's no middle ground.
Most manipulation stacks force a choice: pay for cloud inference with latency and cost, or deploy models that can't understand complex scenes. PYGMALION breaks this trade-off with a 450M model that matches 7B performance.
WHAT PYGMALION DELIVERS
PYGMALION is a 450M parameter SmolVLA that predicts robot actions from images and natural language. Image + instruction in, 7D action trajectory out.
PIPELINE
- 01450M parameters: 15x smaller than 7B baseline VLA models
- 02Input: RGB image (224×224) + text instruction
- 03Output: 7D robot actions (6D pose + gripper) over 10 timesteps
- 04~45ms single inference latency via REST API
CAPABILITIES
- EDGE INFERENCESmolViT vision encoder (384 features) + SmolLM2 language model (576 features)→ MLX Apple Silicon (M1–M5, ~45ms on M5, 4x M1) + CUDA + CPU
- REAL-TIME STREAMINGWebSocket streaming for real-time action prediction + REST + batch→ Per-image and batch inference with configurable max batch
- AUTO DEVICEDevice detection with MPS/CUDA/CPU auto-selection→ First-class support for Apple Silicon, NVIDIA GPUs, and CPU fallback
WHY THIS IS HARD
Vision-language-action modeling requires three encoder-decoder pipelines working in lockstep:
- 01Vision encoder must extract spatial features without destroying fine-grained manipulation details
- 02Language encoder must ground instructions in visual context — not just embed text, but align with image
- 03Action decoder must predict robot joint trajectories in a learned action space (7D normalized)
- 04Parameter-efficient fusion: small vision + small language + lightweight adapter = 450M total
- 05Comparable performance to 7B models with 15x fewer parameters requires careful architecture design
SmolVLA solves this with parameter-efficient fusion. SmolViT (384-dim) + SmolLM2 (576-dim) + lightweight fusion adapter. Production-ready on Apple Silicon and CUDA.
REAL HARDWARE PERFORMANCE
Measured with SmolVLA 450M pipeline:
| METRIC | VALUE |
|---|---|
| Model Size | 450M parameters |
| Baseline Comparison | 7B+ (15x smaller) |
| Action Dimension | 7-DOF (6D pose + gripper) |
| Prediction Horizon | 10 timesteps |
| Single Inference | ~45ms (REST API) |
| Apple M5 (MLX) | ~45ms (4x M1) |
| Vision Encoder | SmolViT (384 features) |
| Language Encoder | SmolLM2 (576 features) |
| Max Batch Size | 16 (configurable) |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Core models | COMPLETE | SmolVLA 450M architecture with production weights |
| API layer | IN PROGRESS | Public API, pending dataset infrastructure |
| Model Loading | COMPLETE | HuggingFace integration, local caching |
| REST API | COMPLETE | /health, /predict, /batch_predict, /info |
| WebSocket API | COMPLETE | Real-time streaming via ws:// |
| Device Detection | COMPLETE | Auto (MPS > CUDA > CPU) |
| Apple Silicon (MPS) | COMPLETE | First-class MLX-ready packaging |
| Prometheus Metrics | COMPLETE | Request latency, inference time, throughput |
| Docker / Compose | COMPLETE | Multi-container with health checks |
| Unit Tests | COMPLETE | Test suite with pytest coverage |
| gRPC Interface | PLANNED | Architecture planned, not implemented |
| ROS2 Integration | PLANNED | Topic-based VLA action subscription |
WHERE PYGMALION DEPLOYS
- APP_01
ROBOT ARM MANUFACTURERS
Deploy manipulation without cloud dependency. On-device VLA inference for pick-and-place, assembly, and general manipulation tasks.
- APP_02
LOGISTICS & WAREHOUSE
Edge inference for pick-and-place operations. Language-guided sorting and palletizing without cloud round-trips.
- APP_03
RESEARCH LABS
Baseline VLA for imitation learning experiments. Small enough to iterate fast, large enough to produce real results.
- APP_04
MOBILE MANIPULATORS
On-board model for collaborative tasks. Real-time action prediction on mobile platforms with limited compute.
- APP_05
EDGE AI DEVELOPERS
Proof-of-concept for small models at scale. Demonstrate VLA capabilities on consumer hardware.
- APP_06
ASSEMBLY AUTOMATION
Language-guided robotic assembly steps. Instruction-driven manipulation for flexible manufacturing lines.
UNDER THE HOOD
FOUNDATION: SMOLVLA (HUGGINGFACE LEROBOT)
- SmolViT vision encoder: 384-dimensional patch embeddings
- SmolLM2 language model: 576-dimensional token embeddings
- Fusion adapter: cross-modal attention + action decoder
- Action space: 7D normalized trajectory (joint control)
PYGMALION IMPLEMENTATION
- FastAPI REST + WebSocket for real-time action streaming
- Per-image and batch inference with configurable max batch
- Device detection with MPS/CUDA/CPU auto-selection
- Prometheus metrics and Docker deployment
INTEGRATION POINTS
- Feeds action trajectories to CHIRON (control loops)
- Provides manipulation plans to ERGON (pose estimation)
- Consumes scene understanding from LOGOS (language grounding)
RESEARCH BASIS
- [01]SmolVLA: A Small Vision-Language-Action Model — HuggingFace LeRobot, Apache-2.0