Matrix-Game 3.5 is an open-source real-time streaming interactive world model developed by Riemann Dynamics at Kunlun Tech, leveraging Patch Memory and Warped PRoPE to enable long-term memory and geometric consistency modeling at the 3D spatial level. The model supports real-time generation at 720P/20FPS on a single GPU, with minute-level scene retroactive recall capabilities and responsive keyboard/mouse interaction. Transitioning from game engines toward general-purpose physical world simulation, it provides open infrastructure for robot training, autonomous driving simulation, and embodied intelligence.

Patch Memory Long-Term Memory: Divides historical frames into 3D spatial patches, reconstructs them via spatial position retrieval, solving issues of object reappearances and scene drift. Object re-representation score surpasses the previous industry ceiling of 0.6.
PRoPE + Warped RoPE Geometric Encoding: Integrates camera projection matrix into Transformer’s spatiotemporal positional encoding, enabling precise perception of camera rotation, translation, and projection relationships—achieving accurate view control and geometric consistency.
Dynamic-Static Memory Decoupling: Static scenes are maintained by Patch Memory; dynamic entities use lightweight Reference Tokens to preserve identity consistency, preventing moving objects from polluting memory and causing ghosting artifacts.
Real-Time Streaming Generation: Achieves single-GPU 720P/20FPS real-time generation by distilling sampling steps down to 3, combined with chunked causal inference, KV Cache, INT8 quantization, and VAE structured pruning.
Automated Data Pipeline: Constructs three data systems—Unreal-Gen (Unreal Engine), AAA game pipelines, and internet video pipelines—producing over 5 million high-quality video clips and more than 10,000 hours of training data, covering 1,200+ game environments.
NPC Interaction & Open-World Generalization: Supports text-driven world generation, character motion control, and multi-agent collaboration.
Unified Geometric-Aware Memory Framework: PRoPE encodes camera poses as relative positional codes without introducing additional learnable parameters; Warped RoPE recalculates historical patch coordinates under current viewpoint, enabling coordinate alignment during memory reuse.
Progressive Distillation: Two-stage distillation transforms bidirectional DiT into a causal generator—first using Perceptual Flow Matching to obtain high-quality low-step causal initialization, then training via Self-Rollout DMD (Distribution Matching Distillation) along autoregressive trajectories, combined with conditional curriculum distillation for classifier-free guidance, camera control, and memory-conditioned generation.
Lightweight Plug-and-Play Architecture: Core interaction components (PRoPE, Patch Memory, dynamic-static decoupling) introduce no additional parameters, implemented modularly to adapt to various video foundation models.

Follow WeChat and reply “open source” to join the AI Open Source Project Community
Download Model: Obtain weights and inference code via GitHub (https://github.com/Riemann-Dynamics/Matrix-Game-3.5) or HuggingFace.
Deployment Environment: Prepare a single CUDA-enabled GPU, install dependencies, and load the model.
Input Control: Navigate with WASD keys, control perspective via mouse, or generate interactive worlds through text prompts or camera trajectory inputs.
Real-Time Interaction: The model streams subsequent frames at 20 FPS, supporting continuous exploration for minutes without losing scene consistency.
Long-Term Temporal Consistency: First open-source solution to systematically address the minute-level memory bottleneck in world models—scene elements remain fully consistent upon repeated revisits.
Real-Time Interactivity: Evolved from offline video generator to playable living world, supporting real-time keyboard/mouse control with latency reduced to playable levels.
Open Ecosystem: Core architecture fully open-sourced; available on GitHub and HuggingFace, cited as benchmark by Xie Saining's team (Solaris), NVIDIA, Adobe, and others.
Zero-Parameter Incremental Design: No parameter inflation, preserves original video distribution and native open content generation capabilities of base models.
| Dimension | Matrix-Game 3.5 | Google Genie 3 |
|---|---|---|
| Open-Source Status | Core architecture fully open-source | Proprietary |
| Resolution | 720P | 720P |
| Real-Time Frame Rate | 20 FPS (single GPU) | Real-time |
| Interaction Duration | Minute-level | Several minutes |
| Memory Mechanism | Patch Memory (3D spatial block retrieval) | Details not disclosed |
| Control Method | Keyboard/mouse/camera trajectory/text prompt | Navigational commands + modifiable world events |
| Parameter Increment | Nearly zero added parameters | Not disclosed |
| Data Pipeline | Self-developed three automated pipelines (UE5/AAA/internet) | Based on Google Street View and similar datasets |
| Ecosystem Citations | Cited by Xie Saining’s team, NVIDIA, Adobe | Industry benchmark |
AI Game Engine: Replaces traditional game engines to generate interactive open worlds in real time, reducing 3A game development costs.
Robot Virtual Training: Serves as a virtual sandbox for humanoid robots and robotic arms, enabling low-cost, large-scale scenario experimentation and long-term task planning.
Autonomous Driving Simulation: Builds digital twins of traffic scenarios governed by real physical rules, used for end-to-end driving policy training.
Embodied Intelligence Research: Provides an interactive physical world simulation environment for embodied agents, supporting joint action-state training.
XR/Metaverse Content Generation: Generates immersive virtual spaces in real time, enabling user-driven exploration and content creation.
Disclaimer: Contains third-party opinions, does not constitute financial advice
Wan-Streamer v0.2 – A Multimodal Understanding and Generation Model Launched by Alibaba Tongyi
4 days ago
StaffDeck – Open-source Enterprise-grade Digital Employee Platforms by Membrane Intelligence and Others
4 days ago
Nemotron 3 Embed – NVIDIA's Open-Source Text Embedding Model Series
4 days ago
Inkling – A multimodal foundational model launched by Thinking Machines Lab
5 days ago
Backup – K12 Primary and Secondary School Teachers' AI Lesson Planning Platform Launched by 360
5 days ago
X2.0 – The World's First Real-Time Interactive Video Generation Model Launches by Xmax AI
6 days ago
Xiaomi-Robotics-U0 – Xiaomi's Unified Embodied Synthetic Model
6 days ago






