PyTorch: Weekly Recap - FSDP Contracts, MPS Speedups, and Compiler Fidelity

FSDP2 all-gather redesign debate led 50 pull request activities and 30 commits, alongside major MPS speedups and linalg fixes. The weekly pattern was correctness at the edges across distributed training, Apple silicon, and compiled paths.

Duration: PT3M30S

Episode overview

This episode is a short developer briefing from PyTorch.

It explains recent repository work in plain language.

  • Show: PyTorch
  • Published: 2026-10-05T09:05:31Z
  • Audio duration: PT3M30S

Transcript excerpt

This excerpt keeps the crawler page concise. Listen to the episode or use the RSS feed for the full update.

Hello and welcome to your PyTorch weekly recap for September 28th through October 5th.

50 pull request activity items, 30 additional commits this week.

Our lead insight: this was a correctness-at-the-edges week. Across distributed training, Apple silicon, and the compiler, teams chased empty inputs, rare numeric cases, and threading and memory edge conditions rather than new features.

First, distributed training and large-scale performance. The biggest discussion is the future all-gather contract for F S D P 2, laid out in three draft alternatives, numbers 199720, 199721 and 199726. They explore who owns communication buffers and how custom backends plug in, which will shape extensions and memory…

Second, Apple M P S maturity, both speed and math. Number 199687 migrates upsample backward to Metal, reporting roughly twenty times faster gradients on large image upsampling. On correctness, number 199086 fixes Jacobi S V D convergence for small correlated columns, while numbers 199694 and 199696 fix linear least…

Third, compiler fidelity and developer workflow. For Dynamo, number 199667 compiles constant source at trace time to match plain Python errors, numbers 199675 and 199685 add…

Nearby episodes from PyTorch

  1. Compiler Parity Push
  2. Distributed Hang Fixes and Compiler Hardening
  3. Low-Precision Grouped GEMM Push and Compiler Fixes
  4. Block Sharding for MoE, Collective Transparency, and Inductor Fixes
  5. GEMM Fusion and Portable CUDA Graphs
  6. Precompile Groundwork and Compiler Reliability Fixes
  7. Compiler Correctness and Precompile Hardening
  8. Weekly Recap - Precompile Scale, Mac Reliability, and Compiler Fixes