PyTorch: All-Gather Showdown and Stack-Aware Merges
Fully sharded training is weighing three competing designs for its next all-gather contract, while merge tooling gains full support for GitHub-native stacks. Apple-silicon speedups and empty-input crash fixes round out a reliability-focused day.
Duration: PT2M31S
Episode overview
This episode is a short developer briefing from PyTorch.
It explains recent repository work in plain language.
- Show: PyTorch
- Published: 2026-10-05T13:00:40Z
- Audio duration: PT2M31S
Transcript excerpt
This excerpt keeps the crawler page concise. Listen to the episode or use the RSS feed for the full update.
Good morning, it's Monday, October 5th, 2026, and this is your PyTorch briefing.
The through-line today is contracts under stress: how large-scale sharding moves data, how stacked pull requests land, and how edge cases crash kernels.
First, fully sharded data parallel is at a design fork. Three draft proposals, PRs 199720, 199721 and 199726, lay out competing options for the all-gather contract, centered on a new all gather layout with prepare, copy-in and finalize stages, backend-owned outputs, and who owns parameter buffers. Two smaller fixes…
Second, the merge bot is learning native stacks. A six-part series, including PRs 199714 for merging, 199716 for rebasing, and 199715 for reverting, lets the bot land, rebase, and roll back an entire GitHub-native stack together, with validation and branch refresh built in. Paired with PR 199734, which posts…
Finally, speed and edge-case hardening. On Apple silicon, PR 199687 replaces the upsample backward graph with a single metal kernel, reporting over twenty times faster on large upscales, while PR 199689 swaps a storage resize blit for a byte-copy kernel. On CUDA, PR 199710 extends tiled transpose copies to batched…
What's next: watch…
Nearby episodes from PyTorch
- Weekly Recap - FSDP Contracts, MPS Speedups, and Compiler Fidelity
- Compiler Parity Push
- Distributed Hang Fixes and Compiler Hardening
- Low-Precision Grouped GEMM Push and Compiler Fixes
- Block Sharding for MoE, Collective Transparency, and Inductor Fixes
- GEMM Fusion and Portable CUDA Graphs
- Precompile Groundwork and Compiler Reliability Fixes
- Compiler Correctness and Precompile Hardening