PyTorch: Native Kernel Push and Distributed Backend Parity
Today's activity centers on a major new "native AOT" kernel effort from slayton58, a four-part MPS reduction rewrite, and a cluster of c10d changes bringing the in-tree NCCL2 backend up to parity with stock NCCL — all while Dynamo continues a steady stream of correctness and overhead fixes.
Duration: PT2M50S
Episode overview
This episode is a short developer briefing from PyTorch.
It explains recent repository work in plain language.
- Show: PyTorch
- Published: 2026-07-25T13:00:33Z
- Audio duration: PT2M50S
Transcript excerpt
This excerpt keeps the crawler page concise. Listen to the episode or use the RSS feed for the full update.
Good day, and welcome to PyTorch, your developer briefing for July 25th, 2026.
The signal today is investment in low-level kernel infrastructure and closing feature gaps between backends, alongside continued Dynamo hardening.
First theme: native ahead-of-time kernels. Slayton58 landed an eleven-part stack, PRs 191044 through 191049, building shared CuTeDSL machinery and Triton toolchains, then using them for real ops — bmm outer product, scatter add, and index add. This is a new code path that pre-compiles reduction and pointwise kernels…
Second theme: backend parity in distributed. D4l3k's five-part stack, PRs 191064 through 191073, brings the in-tree NCCL2 backend up to speed with stock NCCL — adding a NaN-check hook, configurable communicator options, collective timing, sequence numbers, and a renamed gather-single collective. Related, PR 191034…
Third theme: MPS reductions. Isalia20's four-part stack, PRs 191097 through 191100, replaces MPSGraph-based reduction kernels with hand-written Metal kernels for inner-dim, strided, batched, and argmax paths, finishing by migrating min and max off MPSGraph entirely. This targets shapes — skinny matrices, strided…
On Dynamo: PR 191024…
Nearby episodes from PyTorch
- Correctness Hardening Across Inductor and AOTI
- The NVGEMM Epilogue Fusion Marathon
- Distributed Correctness Push
- CUDA Graphs Get a Lifecycle, Distributed Gets Cleaned Up
- One Engineer, Many Kernels
- Weekly Recap - Cleanup, Correctness, and a Random Number Overhaul
- Cleanup Sweeps and a New RNG Foundation
- Hardening Memory and Lifetime Management Across GPU Backends