MiniCPM-Robot is a lightweight, open-source embodied AI Vision-Language-Action (VLA) model series developed by Moonshot Intelligence, comprising the 1.5B-parameter MiniCPM-RobotManip general-purpose manipulation model and the 0.9B-parameter MiniCPM-RobotTrack tracking and navigation model, complemented by the PhyAI inference framework. The models leverage Visual Token Compression and Streaming Context Memory techniques, enabling long-range fine-grained robotic arm operations and indoor/outdoor target tracking for robotic dogs. They support on-device local execution under weak network conditions with inference latency as low as 120ms, placing them among the top tier in mainstream benchmarks.

MiniCPM-RobotManip: Controls robotic arms to perform long-haul task execution (e.g., making sandwiches), supports multi-camera input, and features context memory capability to retain task progress and historical actions.
MiniCPM-RobotTrack: Enables robotic dogs to continuously track designated targets based on natural language instructions, supporting single-target tracking, multi-person interference tracking, and fuzzy target tracking, all operable locally under weak or offline network conditions.

Follow WeChat and reply "open source" to join the AI Open Source Project Community
Clone the GitHub repository and install dependencies.
Prepare hardware: robotic arm (RobotManip) or Unitree Go2 Edu robotic dog (RobotTrack).
Minimal Parameters, Top-Tier Performance: With only 1.5B parameters, it matches larger models like π0.5 in LIBERO, Calvin, and RoboTwin2 benchmarks; achieves ~53 points on RMBench for context memory, far exceeding π0.5’s ~10 points.
Inference Latency Halved: Single-frame decision latency on H100 is ~120ms—approximately half that of π0.5 (~234ms).
Truly On-Device Execution: Robotic dogs achieve stable tracking at over 5Hz using only local compute power, continuing operation even when disconnected from the network.
| Dimension | MiniCPM-RobotManip | π0.5 |
|---|---|---|
| Parameter Count | 1.5B | 3B |
| Open Source | ✓ | ✓ |
| LIBERO | 97.5 | 96.9 |
| Calvin | 4.1 | 4.1 |
| RMBench | 53.5 | ~10 |
| Single-Frame Latency | ~120ms | ~234ms |
| Context Memory | Streaming reuse of historical context | Primarily frame-based judgment |
| On-Device Deployment | Supports local execution | Typically requires cloud computing |
Disclaimer: Contains third-party opinions, does not constitute financial advice
Wan-Streamer v0.2 – A Multimodal Understanding and Generation Model Launched by Alibaba Tongyi
4 days ago
StaffDeck – Open-source Enterprise-grade Digital Employee Platforms by Membrane Intelligence and Others
4 days ago
Nemotron 3 Embed – NVIDIA's Open-Source Text Embedding Model Series
4 days ago
Inkling – A multimodal foundational model launched by Thinking Machines Lab
5 days ago
Backup – K12 Primary and Secondary School Teachers' AI Lesson Planning Platform Launched by 360
5 days ago
X2.0 – The World's First Real-Time Interactive Video Generation Model Launches by Xmax AI
6 days ago
Xiaomi-Robotics-U0 – Xiaomi's Unified Embodied Synthetic Model
6 days ago






