PyTorch: Compile Regions Get Finer Control

Today's changes center on giving compiled and distributed code more granular, region-scoped control—over CUDA graph capture, precompile serving, and pipeline communication—while a parallel thread of correctness fixes targets large-tensor overflow bugs across CUDA, MPS, and XPU backends.

Duration: PT3M4S

Episode overview

This episode is a short developer briefing from PyTorch.

It explains recent repository work in plain language.

  • Show: PyTorch
  • Published: 2026-09-18T13:00:04Z
  • Audio duration: PT3M4S

Transcript excerpt

This excerpt keeps the crawler page concise. Listen to the episode or use the RSS feed for the full update.

Good day. It's September 18th, and here's what mattered in PyTorch today.

The dominant theme: making compiled regions behave independently instead of inheriting settings from whatever wraps them. Desertfire's three-part stack, PRs 197438, 197439, and 197440, lets a nested compile region opt in or out of CUDA graph capture on its own terms, and teaches the scheduler to look inside invoke…

Pipeline parallelism got the same simplification treatment. Sanketpurandare's large stack around PR 197488 unifies manual and traced stage communication initialization, which previously used different, inconsistent protocols depending on frontend and metadata mode.

The second theme is dimension-size correctness. Pablo Garay's fix in PR 197470 addresses sort and topk silently producing wrong results, not just errors, for dimensions past INT_MAX on CUDA and ROCm. Ivan Bogatyy's MPS fix, commit 779a26b, closes a related gap: concatenation outputs over two-to-the-31 elements could…

Worth flagging two reverts: PR 196149's Triton device-properties helper and PR 196937's stream synchronization change were both rolled back this cycle, tied to the same internal diff. If you're touching Triton metadata or…

On…

Nearby episodes from PyTorch

  1. Test Infrastructure Gets a Deep Clean
  2. Sharding, Fake Tensors, and the Great Test Cleanup
  3. Correctness Fixes and the AOT Compile Hardening Push
  4. Compile-On-One-Rank Gets Serious About Device Portability
  5. Weekly Recap - Compiler Correctness and MPS Backend Consolidation
  6. Apple Silicon Hardening and Dynamo's Global State Cleanup
  7. Hardening the Compile-and-Ship Pipeline
  8. Distributed Training Takes Control of Its Own Timing