PyTorch: Low-Precision Grouped GEMM Push and Compiler Fixes

A five-PR stack moves scaled grouped matrix multiply onto the C-U-BLAS-L-T backend with new FP8 and block-scaling formats, while Dynamo, Inductor and distributed checkpointing land correctness and speed fixes.

Duration: PT2M30S

Episode overview

This episode is a short developer briefing from PyTorch.

It explains recent repository work in plain language.

  • Show: PyTorch
  • Published: 2026-10-02T13:00:24Z
  • Audio duration: PT2M30S

Transcript excerpt

This excerpt keeps the crawler page concise. Listen to the episode or use the RSS feed for the full update.

Good morning, it's Friday, October 2nd, 2026. You're listening to PyTorch Daily.

The lead signal is scale for low-precision inference: a coordinated push to give grouped quantized matmuls a full high-performance backend, paired with compiler and checkpoint reliability work developers will feel directly.

First, quantized compute on C-U-D-A. Pull request 199393 builds a new C-U-BLAS-L-T backend for scaled grouped M-M, starting with tensor-wise and group-wise FP8 scaling that the existing backend didn't support. Four follow-ups stack on it: pull request 199394 adds N-V-F-P-4, 199395 adds Hopper-style block scaling,…

Second, the compile stack gets leaner and more correct. Pull request 199417 removes Python 3.10-only handling from Dynamo now that 2.15 requires 3.11, which reduces maintenance risk. For correctness, pull request 199402 fixes device copies for materialized tensor views that could hand a C-P-U buffer to a C-U-D-A…

Third, scale-out and hardware parity. Distributed checkpoint reads get two to three times faster in pull request 199419 by loading items in parallel straight into their targets. On R-O-C-m, pull requests 199396, 199443 and 199425 fix ill-conditioned S-V-D accuracy,…

W…

Nearby episodes from PyTorch

  1. Block Sharding for MoE, Collective Transparency, and Inductor Fixes
  2. GEMM Fusion and Portable CUDA Graphs
  3. Precompile Groundwork and Compiler Reliability Fixes
  4. Compiler Correctness and Precompile Hardening
  5. Weekly Recap - Precompile Scale, Mac Reliability, and Compiler Fixes
  6. Precompile for Serving, Faster First Compile, MPS Correctness
  7. CI Pod Cleanup and Compiler Correctness
  8. Dynamo Fidelity, MPS Scale, and CI Moves