Miles implements MXFP8 and token-wise NVFP4 low-precision reinforcement learning schemes on the Blackwell architecture

Miles implements MXFP8 and token-wise NVFP4 low-precision reinforcement learning schemes on the Blackwell architecture

2026-07-29 18:07View Original

It has been reported that the Miles team implemented two native low-precision reinforcement learning schemes on the Blackwell architecture: end-to-end MXFP8 and token-wise NVFP4 tailored for MoE expert weights. The team conducted ablation experiments on 8x B200 using Qwen3-30B-A3B, with results showing that the original reward curves under BF16 closely align with those under five low-precision configurations, indicating that these low-precision settings did not significantly alter reward performance under the current experimental setup. Meanwhile, both MXFP8 and NVFP4 configurations reduced inference latency, demonstrating substantial efficiency optimization potential in reinforcement learning training and inference workflows on the Blackwell architecture.

Disclaimer: Contains third-party opinions, does not constitute financial advice

Recommended Reading

NVIDIA fell 5% on Monday, during which retail investors bought $108 million, marking the weakest accumulation since early 2025.

2 hours ago
NVIDIA fell 5% on Monday, during which retail investors bought $108 million, marking the weakest accumulation since early 2025.

SK hynix reports strong chip demand in Q2, with AI and memory stocks rebounding pre-market on Wednesday

10 hours ago
SK hynix reports strong chip demand in Q2, with AI and memory stocks rebounding pre-market on Wednesday

South Korean President Yoon Suk-yeol hosts a dinner in San Francisco for executives from NVIDIA, Broadcom, Microsoft, and South Korean technology firms

2 days ago
South Korean President Yoon Suk-yeol hosts a dinner in San Francisco for executives from NVIDIA, Broadcom, Microsoft, and South Korean technology firms

OpenRouter July 26, 2026 LLM Token Usage Rankings: MiMo V2.5, DeepSeek V4 Flash, Hy3 Rank in Top Three

2 days ago
OpenRouter July 26, 2026 LLM Token Usage Rankings: MiMo V2.5, DeepSeek V4 Flash, Hy3 Rank in Top Three

OpenAI and Microsoft Update Their Partnership, Gaining the Right to Deploy Products and Services Across All Cloud Platforms

4 days ago
OpenAI and Microsoft Update Their Partnership, Gaining the Right to Deploy Products and Services Across All Cloud Platforms

Chip and AI-related stocks declined on Monday as the market awaits tech giants' earnings reports and the Federal Reserve's interest rate decision.

4 days ago
Chip and AI-related stocks declined on Monday as the market awaits tech giants' earnings reports and the Federal Reserve's interest rate decision.

Citigroup says tech giants' earnings reports will test AI capital expenditure returns, as constraints around power and permitting begin to show

4 days ago
Citigroup says tech giants' earnings reports will test AI capital expenditure returns, as constraints around power and permitting begin to show