Skip to content
RFL_GLOBAL
中文
  • WAVE 5 // WORLD PREDICTION
  • DAEMON + PYGMALION
  • ZERO-SHOT GENERALIZATION

SIBYL

THE ORACLE

Rather than directly mapping observations to actions, SIBYL first imagines what the world will look like after the action, then generates actions achieving that predicted outcome. Reproducing WoG (World of Generation): predict future world states in latent space to condition action generation. The oracle sees the future — the policy follows.

MODULE STATUS: DEVELOPMENT

Zero-Shot Generalization

PRED

DIVISION
ANIMA
WAVE
W5
DOMAIN
MANIPULATION
WAVE 5 // ANIMA SUITE
MANIPULATION — WORLD PREDICTION FOR ACTION
SIBYL // W5 // 066/079
01THE CHALLENGE

DIRECT OBS→ACTION MAPPING FAILS ON NOVEL SCENES

Current manipulation policies map observations directly to actions. This works in training distributions but fails catastrophically on novel objects, unseen arrangements, and unexpected scene configurations. The policy has no internal model of what success looks like — it just repeats learned motions.

Humans don't work this way. Before reaching for an object, you imagine the outcome — the cup in your hand, the lid on the box. This mental simulation guides your actions. Without world prediction, robots are blind pattern matchers. They need an oracle that sees the future before acting.

02THE SOLUTION

WHAT SIBYL DELIVERS

SIBYL predicts future world states in a learned latent space, then conditions action generation on those predictions. Instead of obs→action, it runs obs→predicted_future→action. The predicted outcome acts as a goal signal that guides manipulation in novel environments.

CAPABILITIES

  • Latent world prediction: encode current state, predict future latent state after action
  • Action conditioning: generate actions that achieve the predicted outcome
  • Zero-shot generalization to novel objects and scene configurations
  • Outcome-driven policy: actions guided by what success looks like, not memorized trajectories
  • Extensive training pipeline with configurable world model architectures
  • Investor demo export with visualization of predicted vs actual outcomes
03ENGINEERING

WHY THIS IS HARD

World prediction for action conditioning is a frontier research problem requiring:

  1. 01Latent space design: the world model must compress high-dimensional observations into a space where prediction is tractable and meaningful for action
  2. 02Temporal consistency: predicted futures must be coherent over multi-step horizons without drift or mode collapse
  3. 03Action conditioning: bridging the gap between abstract latent predictions and concrete motor commands
  4. 04Training stability: jointly learning world models and action generators without one collapsing the other
  5. 05Generalization: the latent space must capture task-relevant structure, not pixel-level memorization

SIBYL solves this with a two-stage architecture: predict the outcome in latent space, then generate actions conditioned on that prediction. The oracle sees — the policy acts.

04BENCHMARKS

KEY PERFORMANCE METRICS

Measured across manipulation benchmarks:

KEY PERFORMANCE METRICS
METRICVALUEDETAIL
Zero-Shot GeneralizationENABLEDNovel objects/scenes via outcome prediction — no retraining required
Latent Prediction AccuracyHIGHWorld model faithfully predicts future latent states over multi-step horizons
Outcome-Driven ActionsACTIVEPolicy conditioned on predicted futures, not memorized trajectories
Novel Scene HandlingROBUSTGeneralizes to unseen object arrangements and workspace configurations
05BUILD STATUS

WHAT'S BUILT TODAY

4/6 COMPONENTS COMPLETE
WHAT'S BUILT TODAY
COMPONENTSTATUSNOTES
Latent World PredictionCOMPLETEEncode current state → predict future latent state
Action ConditioningCOMPLETEGenerate actions achieving predicted outcomes
Core modelsCOMPLETEWorld model + action generator — trained and validated
Demo ExportCOMPLETEInvestor demo with predicted vs actual visualization
Training PipelineIN PROGRESSExtensive pipeline with configurable world model architectures
API layerIN PROGRESSREST + gRPC inference endpoints — pending infrastructure
06APPLICATIONS

WHERE SIBYL DEPLOYS

  • APP_01

    NOVEL OBJECT MANIPULATION

    Grasping and manipulating objects never seen during training — the oracle predicts what success looks like for any object.

  • APP_02

    PREDICTIVE ACTION PLANNING

    Multi-step manipulation sequences guided by predicted future states, enabling complex assembly and rearrangement tasks.

  • APP_03

    GENERALIZATION AT SCALE

    Deploy once, handle anything. Zero-shot transfer to new scenes, new objects, and new workspace configurations without retraining.

07TECHNOLOGY

UNDER THE HOOD

FOUNDATION: WORLD OF GENERATION (WoG)

  • Latent world model: compress observations into a predictive latent space
  • Future state prediction: forecast what the world looks like after action execution
  • Action conditioning: generate motor commands that achieve predicted outcomes
  • Two-stage inference: predict → act (not direct obs → action mapping)

KEY INNOVATION

SIBYL decouples "what should happen" from "how to make it happen." The world model predicts the outcome; the action generator produces the motor commands. This separation enables zero-shot generalization — predict the right future for any scene, then generate actions to reach it.

ANIMA MODULE DEPENDENCIES

DAEMON
Visual trace prompting provides trajectory guidance for action conditioning
PYGMALION
Embodied foundation model supplies the observation encoding backbone

DEPLOYMENT STACK

  • World model inference with configurable latent dimensions
  • Action generator with predicted-outcome conditioning
  • Training pipeline with multi-GPU support
  • Visualization toolkit: predicted vs actual outcome comparison