- WAVE 5 // WORLD PREDICTION
- DAEMON + PYGMALION
- ZERO-SHOT GENERALIZATION
SIBYL
THE ORACLE
Rather than directly mapping observations to actions, SIBYL first imagines what the world will look like after the action, then generates actions achieving that predicted outcome. Reproducing WoG (World of Generation): predict future world states in latent space to condition action generation. The oracle sees the future — the policy follows.
MODULE STATUS: DEVELOPMENTZero-Shot Generalization
PRED
- DIVISION
- ANIMA
- WAVE
- W5
- DOMAIN
- MANIPULATION
- WAVE 5 // ANIMA SUITE
- MANIPULATION — WORLD PREDICTION FOR ACTION
DIRECT OBS→ACTION MAPPING FAILS ON NOVEL SCENES
Current manipulation policies map observations directly to actions. This works in training distributions but fails catastrophically on novel objects, unseen arrangements, and unexpected scene configurations. The policy has no internal model of what success looks like — it just repeats learned motions.
Humans don't work this way. Before reaching for an object, you imagine the outcome — the cup in your hand, the lid on the box. This mental simulation guides your actions. Without world prediction, robots are blind pattern matchers. They need an oracle that sees the future before acting.
WHAT SIBYL DELIVERS
SIBYL predicts future world states in a learned latent space, then conditions action generation on those predictions. Instead of obs→action, it runs obs→predicted_future→action. The predicted outcome acts as a goal signal that guides manipulation in novel environments.
CAPABILITIES
- Latent world prediction: encode current state, predict future latent state after action
- Action conditioning: generate actions that achieve the predicted outcome
- Zero-shot generalization to novel objects and scene configurations
- Outcome-driven policy: actions guided by what success looks like, not memorized trajectories
- Extensive training pipeline with configurable world model architectures
- Investor demo export with visualization of predicted vs actual outcomes
WHY THIS IS HARD
World prediction for action conditioning is a frontier research problem requiring:
- 01Latent space design: the world model must compress high-dimensional observations into a space where prediction is tractable and meaningful for action
- 02Temporal consistency: predicted futures must be coherent over multi-step horizons without drift or mode collapse
- 03Action conditioning: bridging the gap between abstract latent predictions and concrete motor commands
- 04Training stability: jointly learning world models and action generators without one collapsing the other
- 05Generalization: the latent space must capture task-relevant structure, not pixel-level memorization
SIBYL solves this with a two-stage architecture: predict the outcome in latent space, then generate actions conditioned on that prediction. The oracle sees — the policy acts.
KEY PERFORMANCE METRICS
Measured across manipulation benchmarks:
| METRIC | VALUE | DETAIL |
|---|---|---|
| Zero-Shot Generalization | ENABLED | Novel objects/scenes via outcome prediction — no retraining required |
| Latent Prediction Accuracy | HIGH | World model faithfully predicts future latent states over multi-step horizons |
| Outcome-Driven Actions | ACTIVE | Policy conditioned on predicted futures, not memorized trajectories |
| Novel Scene Handling | ROBUST | Generalizes to unseen object arrangements and workspace configurations |
WHAT'S BUILT TODAY
| COMPONENT | STATUS | NOTES |
|---|---|---|
| Latent World Prediction | COMPLETE | Encode current state → predict future latent state |
| Action Conditioning | COMPLETE | Generate actions achieving predicted outcomes |
| Core models | COMPLETE | World model + action generator — trained and validated |
| Demo Export | COMPLETE | Investor demo with predicted vs actual visualization |
| Training Pipeline | IN PROGRESS | Extensive pipeline with configurable world model architectures |
| API layer | IN PROGRESS | REST + gRPC inference endpoints — pending infrastructure |
WHERE SIBYL DEPLOYS
- APP_01
NOVEL OBJECT MANIPULATION
Grasping and manipulating objects never seen during training — the oracle predicts what success looks like for any object.
- APP_02
PREDICTIVE ACTION PLANNING
Multi-step manipulation sequences guided by predicted future states, enabling complex assembly and rearrangement tasks.
- APP_03
GENERALIZATION AT SCALE
Deploy once, handle anything. Zero-shot transfer to new scenes, new objects, and new workspace configurations without retraining.
UNDER THE HOOD
FOUNDATION: WORLD OF GENERATION (WoG)
- Latent world model: compress observations into a predictive latent space
- Future state prediction: forecast what the world looks like after action execution
- Action conditioning: generate motor commands that achieve predicted outcomes
- Two-stage inference: predict → act (not direct obs → action mapping)
KEY INNOVATION
SIBYL decouples "what should happen" from "how to make it happen." The world model predicts the outcome; the action generator produces the motor commands. This separation enables zero-shot generalization — predict the right future for any scene, then generate actions to reach it.
ANIMA MODULE DEPENDENCIES
- DAEMON
- Visual trace prompting provides trajectory guidance for action conditioning
- PYGMALION
- Embodied foundation model supplies the observation encoding backbone
DEPLOYMENT STACK
- World model inference with configurable latent dimensions
- Action generator with predicted-outcome conditioning
- Training pipeline with multi-GPU support
- Visualization toolkit: predicted vs actual outcome comparison