ANIMA CATALOGUE // 196 UNITS
MODULES
Every capability in the ANIMA stack, one page each. Filter by division, domain and wave, or search by name, codename or task.
- MODULES
- 91
- DOMAIN
- 16
- ANIMA
- 79
- GENERIC
- 12
INDEX // 91 RECORDS
91 OF 91 MODULES
- ANIMAW1ABYSSOSTHE BOTTOMLESS DEEP4KMetric depth from any camera — the depth foundation layer for robotic vision systems.PERCEPTION
- ANIMAW6AEGIRTHE OCEANMULTISurvey of multi-modal sensing approaches for robotic perception.SURVEYS
- ANIMAW5AGORAMULTI-ROBOT COORDINATIONFLEETRoboOS-NeXT brain-cerebellum architecture for multi-robot fleet coordination.SIMULATION
- ANIMAW2ATOMOSTHE INDIVISIBLE UNIT3.35×1-bit quantized VLA — 3.35x compression with ternary {-1, 0, 1} weights. Runs on anything.FOUNDATION
- ANIMAW2AZOTHTHE UNIVERSAL EYE146msDetect any object by name, no retraining — open-vocabulary detection on demand.PERCEPTION
- ANIMAW6BALDURTHE DYNAMIC GUARDDYNSLAM that works when the world moves — separates static from dynamic in real-time.SLAM & 3D
- ANIMAW6BESTLATHE BRIDGE+9%Depth as prior knowledge boosts detection — +9% mAP, +7% mAR for small objects.SLAM & 3D
- ANIMAW6BRAGITHE POETVLAComprehensive survey of Vision-Language-Action models for robotics.SURVEYS
- ANIMAW5CENTAURSIM-AND-HUMAN CO-TRAININGSIM+HSimHum co-training framework — joint pretraining on simulation trajectories and human demonstrations, fine-tuning on small real-robot data.SIMULATION
- ANIMAW2CHIRONTHE BRIDGE BETWEEN MIND AND BODY1kHzCRISP-compliant ROS2 controllers for VLA deployment — where AI policy meets real-time robot control.ACTION
- ANIMAW2CHRONOSTIME REVEALS ALL DEPTH60FPSTemporal depth estimation — video frames become consistent depth maps over time.PERCEPTION
- ANIMAW5CHRYSALISTHE METAMORPHOSIS0-TRAINTransfer dynamics priors from world models into frozen VLAs without retraining.MANIPULATION
- ANIMAW5COLOSSEUMROBOTIC COMPETITION SYSTEMSARENAStandardized multi-agent robotic competition and evaluation framework.SIMULATION
- ANIMAW5CORNUCOPIASYNTHETIC VLA DATA630KLarge-scale synthetic VLA data generation: 630K+ trajectories, zero-shot sim-to-real transfer.SIMULATION
- ANIMAW5DAEDALUSHIERARCHICAL ZERO-SHOT VLA3-LVLThree-layer hierarchical VLA: affordance segmentation, 3D planning, collision-aware grasp execution.MANIPULATION
- ANIMAW1DAEMONTHE GUIDING SPIRIT100/sVisual trace prompting for robotic policy learning — draw the path, the robot follows.ACTION
- ANIMAW6EIRTHE DEFENDER>95%Attack + defense for LiDAR SLAM — detect spoofing with >95% rate.SPECIALIZED
- ANIMAW5ELYSIUMLLM-GENERATED ENVIRONMENTS5.1K+Generate simulation environments at scale using LLMs — natural language to full 3D worlds with 5,140+ assets.SIMULATION
- ANIMAW2ERGONTHE PERFECT GRASP±2mmFoundationPose 6DoF pose estimation — CVPR 2024 Best Paper. Works on any object. No retraining.UNDERSTANDING
- ANIMAW6FENRIRTHE WOLF3DCamera-LiDAR late fusion for 3D detection with domain generalization.DETECTION
- ANIMAW6FJALARTHE ALL-WEATHER4D3D detection from 4D radar only — works in rain, fog, snow, dust.DETECTION
- ANIMAW5FORGEEDGE VLA SCALING LAWSEDGEHardware co-design scaling laws for VLA edge deployment.HARDWARE & SENSORS
- ANIMAW6FORSETITHE CALIBRATORAUTOAutomatic LiDAR-camera calibration — no checkerboards, no targets, field-deployable.DEPTH SYSTEMS
- ANIMAW6FREKITHE TRACKERANY-PTTrack any 3D point across multi-view cameras via feed-forward transformer.DETECTION
- ANIMAW6FREYATHE FINDERTEXTText-prompted 3D instance segmentation for indoor scenes.SCENE INTELLIGENCE
- ANIMAW6FRIGGTHE EFFICIENT14M14M params, <100ms on M3 Max — 85-89% fewer than DPT.DEPTH SYSTEMS
- ANIMAW3GENESISTHE ORIGIN OF FORM87.8%RoboSplat Gaussian Splatting — 87.8% task success from a single demonstration. 100x data efficiency.UNDERSTANDING
- ANIMAW6GERITHE URBAN EYECITYOpen-vocabulary urban point cloud segmentation at city scale.SCENE INTELLIGENCE
- ANIMAW3GNOMONKNOW WHERE YOU STAND<2%Legged robots lose 8–15% vertical accuracy on stairs. GNOMON cuts that to under 2%.FOUNDATION
- ANIMAW6GRIDTHE ACCELERATOR237FPS237 FPS on RTX 4090, 161 FPS Jetson — 3.83M params, 25x reduction.DEPTH SYSTEMS
- ANIMAW5HAPTOSTACTILE REPRESENTATION LEARNINGTOUCHGeneral tactile representation learning — cross-sensor optical tactile embeddings for static, dynamic, and force-aware perception.HARDWARE & SENSORS
- ANIMAW3HARMONIAPERFECT BALANCE5MODAdaptively fuse sensors in real-time as environmental conditions change.ACTION
- ANIMAW6HEIMDALLTHE BRIDGE KEEPER6DOFLiDAR + Camera + IMU fusion — Gaussian Splatting maps with 6-DoF localization.SLAM & 3D
- ANIMAW6HELTHE UNIVERSALUNICross-domain depth that works everywhere — indoor, outdoor, aerial, underwater.DEPTH SYSTEMS
- ANIMAW3HERMESTHE INTELLIGENT NAVIGATOR50HzFoundation models + behavior trees for ROS2 navigation — navigate with understanding, not just geometry.ACTION
- ANIMAW6HERMODTHE NAVIGATORCROSSOne navigation model for wheeled, legged, aerial, and vehicular robots.PLANNING
- ANIMAW6HROARRTHE DEXTEROUSDEXVLM task planning + dexterous grasp detection for complex manipulation.SCENE INTELLIGENCE
- ANIMAW6HUGINNTHE THOUGHT RAVENVLMVisual manipulation plans from language — trajectories, waypoints, grasps.PLANNING
- ANIMAW6IDUNTHE ETERNAL24/7Gaussian SLAM for 24/7 service robots — maps that evolve over months.DYNAMIC MAPPING
- ANIMAW3KAIROSTHE DECISIVE MOMENT10μsSee in near-darkness and capture microsecond motion with event camera processing.PERCEPTION
- ANIMAW5LOCIVISUAL MEMORY FOR ROBOTS90.5%Visual place recognition — 90.48% accuracy, <50ms inference. DINOv2 embeddings.PERCEPTION
- ANIMAW4LOGOSSPEAK IT INTO EXISTENCEOPENDescribe what you want, the robot finds and follows it — attribute-aware open-vocabulary tracking.UNDERSTANDING
- ANIMAW6LOKITHE TRICKSTEROPENOpen-vocabulary detection matching GDINO 1.5 accuracy but faster.DETECTION
- ANIMAW6MAGNITHE MIGHTYHYBRIDDepth-aware hybrid fusion for accurate 3D bounding boxes.DETECTION
- ANIMAW5MIDASTHE GOLDEN TOUCHZEROZero-shot robotic manipulation via agentic operational graphs — VLM planning meets dynamic scene graphs.MANIPULATION
- ANIMAW6MIMIRTHE ORACLE3SECPredict future 3D occupancy — where obstacles WILL be in 3 seconds.SCENE INTELLIGENCE
- ANIMAW5MNEMOSYNEVLA MEMORY BENCHMARKCERTStandardized memory evaluation harness for robotic policies — certify before you deploy.SIMULATION
- ANIMAW6MODITHE ANIMATORLIVEReal-time Gaussian Splatting for dynamic scenes — motion graph decomposition.DYNAMIC MAPPING
- ANIMAW1MONADEVERY OBJECT, SOVEREIGN30+Click once, track forever — interactive segmentation with persistent object identity.UNDERSTANDING
- ANIMAW5MORPHEUSTHE SHAPE-SHIFTERANYCross-embodiment VLA — one model, any robot. RDT2 architecture with embodiment-specific action heads.MANIPULATION
- ANIMAW6MUNINNTHE MEMORY RAVENSSLSelf-supervised point cloud foundation model — pre-trained 3D representations without labels.SLAM & 3D
- ANIMAW6NANNATHE ARCHITECTINTEGMulti-module integration patterns for ANIMA stack composition.SURVEYS
- ANIMAW3NEXUSTHE CONNECTING POINTOPENBuild searchable maps where every object is labeled and semantically understood.UNDERSTANDING
- ANIMAW6NJORDTHE PATH SEERIF/THENWhat-if reasoning — predict world states conditioned on planned trajectories.PLANNING
- ANIMAW6NOTTTHE NIGHT SEER$150Thermal-only SLAM for GPS-denied night operations — $150-200 camera.DYNAMIC MAPPING
- ANIMAW6ODINALL-SEEING DEPTH20MAbsolute metric depth from single RGB — 20M 3D sources, zero calibration.DEPTH SYSTEMS
- ANIMAW1OSIRISWE SEE THE UNSEEN30HzDetect people through walls using WiFi + radar + camera fusion — in GPS-denied spaces.FOUNDATION
- ANIMAW4PANOPTESTHE ALL-SEEINGANYCamera-agnostic depth estimation — single model works with any camera. Fisheye, 360, phone, stereo.PERCEPTION
- ANIMAW4PETRATHE FOUNDATION STONE60MDeFM depth foundation model — 3M to 307M parameters. The bedrock depth layer for all stacks.FOUNDATION
- ANIMAW2PLEROMATHE FULLNESS OF GEOMETRY1PASSUniversal feed-forward 3D reconstruction — any image to accurate 3D in one pass.FOUNDATION
- ANIMAW5PRISMMULTI-SENSOR GAUSSIAN SLAM3FUSEMulti-sensor Gaussian SLAM fusing LiDAR, camera, and IMU for real-time photorealistic scene mapping.HARDWARE & SENSORS
- ANIMAW2PROTEUSTHE SHAPE-SHIFTERSAM2Promptable segmentation and tracking for any shape — SAM2-powered, works on anything.UNDERSTANDING
- ANIMAW4PYGMALIONFROM CLAY TO CREATION450M450M parameter VLA small enough for edge, powerful enough for manipulation — 15x smaller than OpenVLA.ACTION
- ANIMAW6RANTHE INSPECTORDEFECTText-conditioned anomaly detection — describe defect, detect it.SPECIALIZED
- ANIMAW6SAGATHE WATCHFULMETADetects when perception degrades — fog, blur, rain, noise — and flags it.DETECTION
- ANIMAW5SIBYLTHE ORACLEPREDPredict future world states in latent space to condition action generation.MANIPULATION
- ANIMAW6SIFTHE SHARPENERULTRAUltra-dense depth by fusing stereo with sparse LiDAR via diffusion distillation.DEPTH SYSTEMS
- ANIMAW6SKADITHE HUNTRESSLANG3DSurvey of language models integrated with 3D scene representations.SURVEYS
- ANIMAW6SKULDTHE FATE WEAVERSATNavigation costmaps from satellite imagery via natural language.PLANNING
- ANIMAW6SOLTHE SUN+33%Synthesize thermal images from RGB via VLM — 33% improvement over baselines.SPECIALIZED
- ANIMAW6SURTTHE COMPRESSOR190×190x storage reduction for 4D Gaussian Splatting scenes.DYNAMIC MAPPING
- ANIMAW3SYNTHESISMANY EYES, ONE MAP5+Multiple robots build one shared world model through collaborative SLAM.ACTION
- ANIMAW3TACTISTHE KNOWING TOUCHSUB-MMFuse vision and touch sensors for sub-millimeter surface mapping and reconstruction.FOUNDATION
- ANIMAW6THORTHE PATHFINDER30FPSReal-time monocular SLAM — 30 FPS RTX 4090, 20 FPS Jetson Orin.SLAM & 3D
- ANIMAW5TITANWORLD-MODEL VLAWMWorld model conditioned VLA: predict dynamics, condition policy, collect rollouts, human correction.MANIPULATION
- ANIMAW6TYRTHE WARRIORMANIPSurvey of VLM for robotic manipulation — task decomposition to motion gen.SURVEYS
- ANIMAW6URDTHE RECONSTRUCTOR1PASSMetric 3D reconstruction from multi-view in a single feed-forward pass — no COLMAP.SLAM & 3D
- ANIMAW6VALITHE PREDICTORPRE-XPredict consequences of actions before execution via VLM.SCENE INTELLIGENCE
- ANIMAW6VIDARTHE SIGHT6DOF6-DoF object pose from single RGB via Gaussian Splatting.SCENE INTELLIGENCE
- GENERICW7CALIB-PROJFUSIONAlign Your Sensors Once. Trust Them Forever.PROD25.3M-parameter camera-LiDAR calibration engine that self-corrects after hard landings, vibration, or thermal drift — no checkerboard required.GENERAL
- GENERICW7GS-GHOSTReconstruct the Grasp Before the Robot ReachesDEVMonocular RGB video reconstructs hand-object interactions in real-time via Gaussian Splatting — unlocking imitation learning from raw video at scale.GENERAL
- GENERICW7GS-OVIEReconstruct Any Scene From a Single PhotographDEVOne image is all it takes — monocular novel view synthesis that generalizes to environments it has never seen before.GENERAL
- GENERICW7MANIP-RAAPThe Robot That Learned to Grip by WatchingPRODRetrieval-augmented affordance prediction that tells a robot exactly where to grab and which way to push — from a single RGB image.GENERAL
- GENERICW7MANIP-ULTRADEXTwo Hands, Any Object, No RehearsalPRODUniversal bimanual dexterous grasping policy trained entirely on synthetic data — deploys to dual-arm robotic systems without a single real-world demonstration.GENERAL
- GENERICW7SIM-CARLAAIROne Simulator to Rule Air and Ground Together5Merge drone and ground vehicle simulation in one Unreal Engine process — 18.6x faster coordinate transforms, 63-topic ROS2 bridge, zero config friction.GENERAL
- GENERICW7SLAM-COKOA Robot Swarm That Shares One Map, Even After Comms DropPRODMulti-agent Gaussian Splatting SLAM with loop closure and pose graph fusion — the distributed mapping backbone for autonomous robot teams.GENERAL
- GENERICW7SLAM-GS3LAMSee the World in 3D and Know What Everything Is3Semantic Gaussian Splatting SLAM — simultaneously builds a photorealistic map and labels every surface by object class, in real time.GENERAL
- GENERICW7SLAM-MIPSLAMCrisp Maps at Any Scale, Without the Aliasing LiesPRODAlias-free Gaussian SLAM that renders accurate maps at multiple resolutions — because a map that looks sharp but is geometrically wrong is a safety failure, not a metric.GENERAL
- GENERICW7UAV-TRACKVLALanguage-Commanded Pursuit That Doesn't Lose Its TargetDEVAerial tracker that understands "follow the white SUV turning left" and executes it — fusing language, vision, and flight control in one frozen backbone.GENERAL
- GENERICW7VIS-FORESTSIMPerception That Survives the TreelinePRODSynthetic forest perception benchmark that trains vehicles to navigate unstructured terrain where GPS fails and road markings don't exist.GENERAL
- GENERICW7VIS-OCCANY3D Occupancy Prediction That Works Where It Has Never BeenPRODCVPR 2026 occupancy model that predicts full 3D voxel occupancy in unseen environments — zero-shot, no domain adaptation required.GENERAL