pytorch

mirror of https://github.com/pytorch/pytorch.git synced 2025-10-22 22:25:10 +08:00

Author	SHA1	Message	Date
Edward Z. Yang	f00a1b0349	Fix profiler stack trace names	2025-08-06 21:20:14 -07:00
Edward Yang	bc67bce2e5	Working setup with runnable PyTorch on Codex. Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 132668d46021090fe3ef197fb25ba762ce42667c Pull-Request: https://github.com/pytorch/pytorch/pull/159968	2025-08-06 14:56:40 -07:00
Zhengxu Chen	79eca4677b	[precompile] Skip serializing unnecesssary objects for guards. (#158926 ) Summary: The following type of objects don't need to be serialized for precompile: 1. PyCapsule because we don't guard on C binding objects in meaningful ways. 2. Code object because we only id matching on these but id matches will always be dropped for precompile. 3. Nested function objects since we also ban CLOSURE_MATCH. Test Plan: buck run mode/opt test/dynamo:test_dynamo -- -k test_skipped_objects Rollback Plan: Differential Revision: D78816888 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158926 Approved by: https://github.com/jamesjwu	2025-08-06 15:00:28 +00:00
PyTorch MergeBot	2855688a1d	Revert "Replace C array with std::array in formatSockAddr (#159812 )" This reverts commit e7feedf6a9bb346ad205796aa4084c8dcfb18072. Reverted https://github.com/pytorch/pytorch/pull/159812 on behalf of https://github.com/malfet due to Looks like it broke distribtued tests, see `2231c3ca3a/1` ([comment](https://github.com/pytorch/pytorch/pull/159812#issuecomment-3160513656))	2025-08-06 14:55:48 +00:00
Nikita Shulga	2231c3ca3a	[CI][CD] Fix `install_nvshem` function (#159907 ) When one builds CD docker, all CUDA dependencies must be installed into `/usr/local/cuda/` folder Test plan: Looks at the binary build logs, for example [here](https://github.com/pytorch/pytorch/actions/runs/16768141521/job/47477380147?pr=159907): ``` 2025-08-06T05:58:00.7347471Z -- NVSHMEM_HOME set to: '' 2025-08-06T05:58:00.7348378Z -- NVSHMEM wheel installed at: '' 2025-08-06T05:58:00.7392528Z -- NVSHMEM_HOST_LIB: '/usr/local/cuda/lib64/libnvshmem_host.so' 2025-08-06T05:58:00.7393251Z -- NVSHMEM_DEVICE_LIB: '/usr/local/cuda/lib64/libnvshmem_device.a' 2025-08-06T05:58:00.7393792Z -- NVSHMEM_INCLUDE_DIR: '/usr/local/cuda/include' 2025-08-06T05:58:00.7394252Z -- NVSHMEM found, building with NVSHMEM support ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159907 Approved by: https://github.com/Skylion007, https://github.com/ngimel	2025-08-06 14:44:37 +00:00
can-gaa-hou	c03a734ba1	[OpenReg] Disable automatic inclusion of data files (#159845 ) # Background After I built torch_openreg, I noticed that the wheel package contained the stub.c file under the csrc directory, which was not used in the runtime. # Motivation This PR aims to remove the stub.c file and any unused file when running torch_openreg. Changes: - Setting include_package_data keyword to false in the setup function Pull Request resolved: https://github.com/pytorch/pytorch/pull/159845 Approved by: https://github.com/albanD	2025-08-06 10:35:13 +00:00
Benji Beck	98316e5896	[WOQ] Add CUDA kernel for _weight_int8pack_mm (#159325 ) Summary This issue proposes implementing a CUDA kernel for aten._weight_int8pack_mm, a weight-only quantized (WOQ) linear operation that is currently only supported on CPU. On CUDA, the fallback path uses an unfused .mul().sum() pattern in quantization.py, which is less efficient for inference. https://github.com/pytorch/pytorch/issues/158849 Motivation A fused GPU kernel for aten._weight_int8pack_mm would: - Eliminate reliance on the .mul().sum() fallback in quantization.py - Improve performance for quantized inference on CUDA - Extend Inductor’s GPU quantization support across more workloads Implementation - Implement a Triton kernel for: ``` out[b, n] = sum_k(x[b, k] * w[n, k]) * scale[n] where: x: [B, K] float32 w: [N, K] int8 scale: [N] float32 out: [B, N] float32 ``` - Integrate the kernel with register_woq_mm_ops() in torch/_inductor/quantized_lowerings.py - Route it conditionally in quantization.py where GPU currently falls back to .mul().sum() - Add unit tests comparing results to the reference fallback path Test Plan: ``` buck2 run 'fbcode//mode/opt' :linalg test_linalg.TestLinalgCUDA.test__int8_mm_m_64_k_64_n_64_compile_True_slice_True_cuda ``` Log: P1882799769 ``` buck2 test 'fbcode//mode/opt' caffe2/test:linalg ``` https://www.internalfb.com/intern/testinfra/testconsole/testrun/6755399722424741/ Benchmark Results: ``` [Shape B=256, K=1024, N=512] CPU and CUDA outputs match Max abs diff: 2.59e-04, max rel diff: 0.75 CPU: 144.14 ms, CUDA: 303.67 µs Speedup: ×474.6 [Shape B=512, K=2048, N=1024] CPU and CUDA outputs match Max abs diff: 5.49e-04, max rel diff: 0.15 CPU: 1173.27 ms, CUDA: 2.40 ms Speedup: ×488.5 ``` Rollback Plan: Differential Revision: D79042656 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159325 Approved by: https://github.com/danielvegamyhre, https://github.com/jerryzh168	2025-08-06 10:28:08 +00:00
angelayi	23cf241039	[aoti][mps] Initialize mps kernels first (#159753 ) In some cases we have mps kernels which are reused across higher-order-op subgraphs and the toplevel code. However, currently we initialize the variable for the mps kernel the first time we use it, which runs into an issue if we run into the mps kernel within a subgraph since the kernel will only be initialized within the subgraph scope. For instance: ``` if ... auto mps_lib_0_func = ... mps_lib_0_func->run() // since we already used mps_lib_0 once, we don't re-initialize it mps_lib_0_func->run() // error, mps_lib_0_func not initialized ``` So the solution we took here is to initialize all the kernels at the beginning: ``` const std::shared_ptr<at::native::mps::MetalKernelFunction> get_mps_lib_0() { static const auto func = mps_lib_0.getKernelFunction("generated_kernel"); return func; } AOTIMetalKernelFunctionHandle get_mps_lib_0_handle() { static const auto handle = AOTIMetalKernelFunctionHandle(get_mps_lib_0().get()); return handle; } ... if ... get_mps_lib_0()->run() get_mps_lib_0()->run() // success ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159753 Approved by: https://github.com/malfet ghstack dependencies: #159456, #159695	2025-08-06 07:54:29 +00:00
Will Constable	e7feedf6a9	Replace C array with std::array in formatSockAddr (#159812 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159812 Approved by: https://github.com/Skylion007	2025-08-06 07:44:29 +00:00
Will Constable	dad2a05bec	[DTensor] Set up DTensorContinuousTestBase (#159885 ) Also migrate `test_common_rules.py` since it was a short file `python test/distributed/tensor/test_common_rules.py` Before: Ran 10 tests in 91.516s After: Ran 10 tests in 5.604s Pull Request resolved: https://github.com/pytorch/pytorch/pull/159885 Approved by: https://github.com/ezyang	2025-08-06 07:40:31 +00:00
Colin L Reliability Rice	0495cab545	Wire in pt2_triton_builds (#159897 ) Summary: This allows us to start seeing the failure rate on these models (and potentially alert on it). Test Plan: ``` FORCE_LOG_TRITON_BUILDS_TO_PROD=1 TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 buck2 run @//mode/opt :compile 2>&1 \| tee out ``` P1889607054 Waiting for scuba table to generate, but manual logging show it should show up at https://fburl.com/scuba/pt2_triton_builds_inc_archive/7852kt8h soon. Rollback Plan: Reviewed By: masnesral Differential Revision: D79308333 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159897 Approved by: https://github.com/masnesral	2025-08-06 07:39:51 +00:00
Mengtian Xu	abfe403981	[AIDIR] Internal util function to insert MLHub debugging insight for dynamic shape (#159391 ) Summary: This feature is Meta internal only Add a util function to put dynamic shape-related suggestion to MLHubDebugInsightService, which will then be surfaced to users in the MLHub . The rollout will be controlled by JK. Test Plan: MAST job aps-omnifmv3_dev_baseline_test-a34fdccf21 {F1980593060} * If you're not able to see the insight, please add yourself to this gk 'mlhub_debugging_insights_dev_visibility' * The URL link should route to a new Job Inspector page that will provide details and straight forward instructions of how to config the ds. The page is currently still in development so here we use the general PT2 compile JI page. * Test fails because of the export checks. I'll export after addressing all the comments from reviewers. Rollback Plan: Reviewed By: pianpwk Differential Revision: D78526522 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159391 Approved by: https://github.com/jingsh	2025-08-06 07:39:39 +00:00
Jane Xu	1690c0c3a0	[Reland] Migrate ScalarType to headeronly (#159911 ) The non ghstack version of #159416, to make sure we don't get reverted again Pull Request resolved: https://github.com/pytorch/pytorch/pull/159911 Approved by: https://github.com/mikaylagawarecki	2025-08-06 07:36:37 +00:00
Aidyn-A	e9d27aa8fd	[CUDA 13] CMake/Dependencies: no need to call find_package(CUB) (#159854 ) CUB library is the part of CCCL of the CUDA Toolkit 13. If CUDA Found, CUB is found as well. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159854 Approved by: https://github.com/eqy	2025-08-06 06:03:58 +00:00
PyTorch MergeBot	2457e62c90	Revert "Set PYTHONHOME for inductor subprocesses using torch (#159382 )" This reverts commit fe8984a9f43bde10d1956abe7cb40710ed7ceed2. Reverted https://github.com/pytorch/pytorch/pull/159382 on behalf of https://github.com/malfet due to Broke MacOS testing see `d0fccbc99c/1` ([comment](https://github.com/pytorch/pytorch/pull/159382#issuecomment-3157455367))	2025-08-06 05:30:20 +00:00
Nikita Shulga	d0fccbc99c	[CI] Delete sm86 tests from pull (#159903 ) And delete sm89+cuda12.4 builds from periodic (as sm86+legacy driver should be enough) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159903 Approved by: https://github.com/huydhn	2025-08-06 05:16:55 +00:00
PyTorch UpdateBot	3461988a4b	[audio hash update] update the pinned audio hash (#159823 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159823 Approved by: https://github.com/pytorchbot	2025-08-06 05:02:35 +00:00
Will Constable	9764981116	Pass fw/bw compilers to aot_export_joint_with_descriptors (#159814 ) Allow overriding nop compilers with real ones when using this flow. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159814 Approved by: https://github.com/fmassa	2025-08-06 04:50:56 +00:00
Michael Lazos	704594eb23	[Dynamo] make HOPs hashable (#159910 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159910 Approved by: https://github.com/yf225	2025-08-06 04:02:17 +00:00
eqy	bfc27cf468	[Distributed] Fix `@parametrize` on unordered iterable in distributed test (#159793 ) seems to fix https://github.com/pytorch/pytorch/issues/145807 sets aren't ordered so `@parametrize` can cause two processes to spawn with different settings originally debugged thanks to @k-artem, see https://github.com/pytorch/pytorch/issues/145807#issuecomment-2971009451 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159793 Approved by: https://github.com/Skylion007, https://github.com/wconstab	2025-08-06 03:51:42 +00:00
bobrenjc93	311f74089a	remove print (#159917 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159917 Approved by: https://github.com/laithsakka	2025-08-06 03:48:23 +00:00
Tianhao Huang	14c7358c64	Enable fr_trace to read local traces from multiple hosts. (#159490 ) Summary: For training jobs particularly from GenAI, NCCL trace dumps are generated in the format of `<hostname>.pci3_rank_<rank>`. For multi-node training jobs, the hostname varies across traces. The current prefix matching logic can't handle this case. Test Plan: Create a local folder `dumps` and several empty files: `host0.pci3_rank_0`, `host0.pci3_rank_1`, `host1.pci3_rank_0`, `host1.pci3_rank_1` inside it. Then run ``` buck2 run fbcode//caffe2/fb/flight_recorder:fr_trace -- trace_dir dumps ``` Before this diff, fr_trace cannot locate any trace files, giving the following assertion error: ``` AssertionError: no files loaded from /home/tianhaoh/dumps with prefix pci3_rank_ ``` After this diff, fr_trace is able to locate the trace files, resulting in the exceptions like ``` dump = pickle.load(infile) ^^^^^^^^^^^^^^^^^^^ EOFError: Ran out of input ``` (since the trace files are fake and empty). Rollback Plan: Differential Revision: D79224727 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159490 Approved by: https://github.com/fduwjj	2025-08-06 03:15:34 +00:00
Dave Lei	8ce81bcee1	[Torch Package] Make get names of OrderedImporters support fallback to importers (#155743 ) Summary: OrderedImporters is supposed to be an importer which tries out every single importer in self._importers. However the get_name API does not follow this behavior and only uses the get_name from the basic Importer class. This change is to update the OrderedImporters get_name API so that it tries the get_name API of every single importers. Differential Revision: D76463252 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155743 Approved by: https://github.com/jcwchen, https://github.com/jingsh	2025-08-06 02:26:10 +00:00
Yu, Guangye	4604f0482c	Add UT for torch.accelerator memory-related API (#155200 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155200 Approved by: https://github.com/albanD ghstack dependencies: #138222, #152932	2025-08-06 02:22:18 +00:00
Yu, Guangye	15f1173e5d	Add unified memory APIs for torch.accelerator (#152932 ) # Motivation The following API will be put under torch.accelerator - empty_cache - max_memory_allocated - max_memory_reserved - memory_allocated - memory_reserved - memory_stats - reset_accumulated_memory_stats - reset_peak_memory_stats Pull Request resolved: https://github.com/pytorch/pytorch/pull/152932 Approved by: https://github.com/albanD ghstack dependencies: #138222	2025-08-06 02:22:18 +00:00
henrylhtsang	e16c48ae97	[BE] Fix type hint in AOTIRunnerUtil (#159577 ) Not sure why it was labelled as list in the first place. In test_aot_inductor.py, I scanned a few use cases and they are tuple as well. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159577 Approved by: https://github.com/Skylion007	2025-08-06 01:20:45 +00:00
Yu, Guangye	f7a66da5f9	Add DeviceAllocator as the base device allocator (#138222 ) # Motivation In line with [RFC] [A device-agnostic Python device memory related API design for stream-based accelerators](https://github.com/pytorch/pytorch/issues/134978), some memory-related APIs are widely used in popular repositories, such as HuggingFace [so many if-else conditional code](https://github.com/search?q=repo%3Ahuggingface%2Faccelerate%20torch.cuda.empty_cache&type=code). We would like to introduce a generic API set under torch.accelerator namespace to generalize these user cases. <div align="center"> <table> <tr> <td> Device-specific memory APIs torch.xxx.foo</td> <td> Device-agnostic memory APIs torch.accelerator.foo</td> </tr> <tr> <td> ```python torch.xxx.empty_cache ``` </td> <td> ```python torch.accelerator.empty_cache ``` </td> </tr> <tr> <td> ```python torch.xxx.reset_peak_memory_stats ``` </td> <td> ```python torch.accelerator.reset_peak_memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.reset_accumulated_memory_stats ``` </td> <td> ```python torch.accelerator.reset_accumulated_memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_stats ``` </td> <td> ```python torch.accelerator.memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_allocated ``` </td> <td> ```python torch.accelerator.memory_allocated ``` </td> </tr> <tr> <td> ```python torch.xxx.max_memory_allocated ``` </td> <td> ```python torch.accelerator.max_memory_allocated ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_reserved ``` </td> <td> ```python torch.accelerator.memory_reserved ``` </td> </tr> <tr> <td> ```python torch.xxx.max_memory_reserved ``` </td> <td> ```python torch.accelerator.max_memory_reserved ``` </td> </tr> </table> </div> # Solution This design follows a similar pattern to `HostAllocator`. We're introducing a base class `DeviceAllocator`, from which `CUDAAllocator` and `XPUAllocator` will inherit. This allows us to provide a unified call path like: `torch.accelerator.empty_cache()` -> `GetDeviceAllocator(allocator)->empty_cache()`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/138222 Approved by: https://github.com/albanD, https://github.com/Camyll	2025-08-06 00:40:29 +00:00
Animesh Jain	3eb3da9b4b	[dynamo][guards] Skip ID_MATCH guard on self.__class__.__closure__ (#159888 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159888 Approved by: https://github.com/williamwen42	2025-08-06 00:36:43 +00:00
Jane Xu	3ddfd46bd2	Cut a version of TORCH_ERROR_CODE_CHECK in headeronly from AOTI (#159604 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159604 Approved by: https://github.com/albanD, https://github.com/desertfire	2025-08-06 00:29:56 +00:00
Zhengxu Chen	6a82da392e	[export] Fix generated schema for C++20/23 (#159871 ) Summary: Fixing the issue from https://github.com/pytorch/pytorch/issues/159838 Test Plan: buck run caffe2/:export_update_schema -- --prefix /data/users/$USER/fbsource/fbcode/caffe2/ Rollback Plan: Differential Revision: D79647167 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159871 Approved by: https://github.com/malfet	2025-08-06 00:23:05 +00:00
Simon Fan	22bedc429f	Extract some HOP utils to be importable (#159705 ) Useful helper function for stage 1 export -> manual partitioner -> stage 2 compile users Pull Request resolved: https://github.com/pytorch/pytorch/pull/159705 Approved by: https://github.com/zou3519 ghstack dependencies: #159134	2025-08-05 23:59:47 +00:00
Huy Do	49abc0e3f8	[Take 2] Setup TorchBench in Docker (#159300 ) Fix and reland https://github.com/pytorch/pytorch/pull/158613, I keep `checkout_install_torchbench` in `.ci/pytorch/macos-test.sh` script because it's still used there, and there is no Docker. ### Testing MacOS perf nightly run https://github.com/pytorch/pytorch/actions/runs/16580798470 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159300 Approved by: https://github.com/ZainRizvi	2025-08-05 23:47:42 +00:00
Xu Han	1052604acd	fix logging setup issue for Windows.. (#159887 ) When we setup logging config as guide: https://docs.pytorch.org/docs/stable/logging.html Such as: TORCH_LOGS="+schedule,+inductor,+output_code" On Linux, it shows as: ```cmd declare -x SSH_TTY="/dev/pts/0" declare -x TERM="xterm" declare -x TORCH_LOGS="+schedule,+inductor,+output_code" declare -x USER="xu" ``` On Windows, it shows as: ```cmd TORCHINDUCTOR_WINDOWS_TESTS=1 TORCH_LOGS="+schedule,+inductor,+output_code" UCRTVersion=10.0.22000.0 ``` For Linux, it shows quotes by default, And Windows is not shows quotes. Besides that, Windows would auto assemble quotes when env var processing. On Linux, we will get variable: "+schedule,+inductor,+output_code" On Windows, we will get variable: '"+schedule,+inductor,+output_code"' So, we need remove the outer quotes for Windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159887 Approved by: https://github.com/angelayi	2025-08-05 23:44:38 +00:00
Alex Malyshev	fe8984a9f4	Set PYTHONHOME for inductor subprocesses using torch (#159382 ) Summary: This is needed for subprocesses that are trying to call back into torch functionality, i.e. anything that's also setting `PYTHONPATH`. There are more `sys.executable` subprocesses in torch/ but it seems like they're fine. Test Plan: Local inference runs. Reviewed By: aorenste Differential Revision: D79124705 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159382 Approved by: https://github.com/aorenste	2025-08-05 23:32:48 +00:00
angelayi	74a754aae9	Add meta kernel for sdpa_math_for_mps (#159695 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159695 Approved by: https://github.com/malfet ghstack dependencies: #159456	2025-08-05 22:27:06 +00:00
angelayi	b1ec088113	[mps] Turn on inductor dynamic shapes tests (#159456 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159456 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-08-05 22:27:06 +00:00
angelayi	fb35a9ea4a	[export] Improve error messages (#159881 ) Originally, if the PT2 errored when loading, we would try to load using the old loader to fit BC issues. However this hides the error messages for if an up-to-date PT2 is erroring when loading due to some other reason. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159881 Approved by: https://github.com/yushangdi	2025-08-05 22:26:48 +00:00
Sandeep Narendranath Karjala	8034b2a732	[inductor] Add TLParse artifact for logging runtime of collective and compute ops (#159730 ) Summary: - debug.py: Added log_runtime_estimates() function to dump runtime estimation data as structured tlparse artifacts in JSON format - test_structured_trace.py: Added comprehensive test coverage with testing compute and collective ops Pull Request resolved: https://github.com/pytorch/pytorch/pull/159730 Approved by: https://github.com/yushangdi ghstack dependencies: #159190	2025-08-05 22:06:32 +00:00
anwang	64cc6f06b1	[Inductor] Revert minimal changes to avoid internal test failures (#159809 ) The diff/PR https://github.com/pytorch/pytorch/pull/159211 caused a bunch of test failures for graph compiler(T232684410). But I couldn't figure out a forward fix so far. So with this diff/PR, I'm proposing to revert the minimal changes to resolve the test failures. I'll continue the debugging, and re-land the reverted changes once we find out a forward fix. Differential Revision: [D79221721](https://our.internmc.facebook.com/intern/diff/D79221721/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159809 Approved by: https://github.com/blaine-rister, https://github.com/eellison	2025-08-05 22:05:26 +00:00
PyTorch MergeBot	410812763b	Revert "[Inductor][Triton] Support TMA before strict 3.4 cutoff (#159777 )" This reverts commit bbc0df1094b5a4dcd2cce83f8402127b07913231. Reverted https://github.com/pytorch/pytorch/pull/159777 on behalf of https://github.com/izaitsevfb due to breaking inductor test on ROCm ([comment](https://github.com/pytorch/pytorch/pull/159777#issuecomment-3156770098))	2025-08-05 22:00:24 +00:00
Michael Lazos	bdb07a2bc5	[Cutlass] Allow offsets to be passed as arguments to kernel (#159761 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159761 Approved by: https://github.com/henrylhtsang ghstack dependencies: #159760	2025-08-05 21:59:07 +00:00
Simon Fan	8085edc8f9	[autograd] torch._C._set_view_replay_enabled state leaking into other tests (#159840 ) This was causing view_fns to pop up in tests that ran after `TestAutograd.test_view_replay_enabled` where it isn't used as a context manager. It is unclear to me why we would want `_force_original_view_tracking` to mutate global state on __init__ rather than on __enter__, that could be an alternative fix. FIXES https://github.com/pytorch/pytorch/issues/156306 https://github.com/pytorch/pytorch/issues/156289 https://github.com/pytorch/pytorch/issues/156265 https://github.com/pytorch/pytorch/issues/156209 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159840 Approved by: https://github.com/albanD	2025-08-05 21:57:49 +00:00
Nikita Shulga	882d50c5bf	[C10] Add `Scalar::isUnsigned()` method (#159877 ) That returns true if Scalar hold unsigned integral value With the implications of `Tag::HAS_u` semantic. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159877 Approved by: https://github.com/Skylion007, https://github.com/ezyang	2025-08-05 21:43:21 +00:00
Catherine Lee	b52a4d0821	[ez][CI] Remove some unused docker images (#159171 ) Removes unused docker images from the docker build workflow Then removes unused definitions in build.sh The only one I left is the vllm one because I'm pretty sure it's going to be used in the future I assume everything not mentioned is old and we forgot to remove them Pull Request resolved: https://github.com/pytorch/pytorch/pull/159171 Approved by: https://github.com/yangw-dev	2025-08-05 21:31:53 +00:00
Nikita Shulga	a45a840926	[CI] Disable check-labels and check_mergeability (#159900 ) See https://github.com/pytorch/pytorch/issues/159825 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159900 Approved by: https://github.com/clee2000	2025-08-05 21:16:12 +00:00
Nikita Shulga	9b953bb3fb	[BE] Update TensorPipe pin (#159834 ) No functional changes, just: - Update C++ standard to C++17 - Update `cmake` min version to 3.18 - Update `libuv` dependency to 1.51 (to move its cmake min version to 3.10) - Replace boost optional implementation with `std::optional` wrapper - Make it compilable with gcc-14.x plus by including `cstddef` in few headers - Avoid using deprecated enums for MacOS builds Pull Request resolved: https://github.com/pytorch/pytorch/pull/159834 Approved by: https://github.com/Skylion007	2025-08-05 20:45:09 +00:00
eellison	eb25a95a6e	Fix inductor memory estimation when a single buf has multiple mutations. Add runtime verification of mem tracking (#159569 ) With fsdp, we sometimes have multiple, non-overlapping views of a single buffer which are all mutated. Previously we considered the original buffer as an allocation, and make the mutated buffer the deallocation. With multiple mutations of the same buffer, we need to consider the original buffer as deallocated only when all of its aliases die (and avoid double counting the input buffer size). See comment inline: ``` When an operation mutates a buffer in-place, the scheduler creates a new buffer name to track the "before" and "after" states, even though they share the same memory. The mutated buffer represents a rename with zero allocation and deallocation cost. During dependency tracking, we transfer dependencies from the mutated name back to the original buffer, ensuring the original memory is only freed when all aliases are done. This handles cases where a buffer has multiple non-overlapping aliases - rather than trying to assign free costs to individual aliases, we forward all alias dependencies to the original buffer. Consider: buf0 = op0() buf1 = mutation_op_(buf0) del buf0 ... op(buf1) del buf1 The only memory events are the creation prior to op0, and the deletion following buf1. ``` As @IvanKobzarev 's logs in https://github.com/pytorch/pytorch/pull/158361/files#diff-e173a1d52aff49959c9f6d17ecc09946d8a616fc5909df884e62a15e1ebd1d41R1776-R1807 show, it can a bit of a pain to pinpoint which part of our memory calculation is incorrect. This pr also adds a runtime verifier `config.test_configs.track_memory_lifecycle` which tracks buffer allocation and deallocation, and errors if their lifetime does not match our expectations. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159569 Approved by: https://github.com/IvanKobzarev	2025-08-05 19:58:11 +00:00
eqy	9884d0351e	[CUDA] Decrease launch bounds of CTCLoss backward for blackwell (#159522 ) Otherwise we see `CUDA error: too many resources requested for launch` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159522 Approved by: https://github.com/janeyx99	2025-08-05 19:26:25 +00:00
Eli Uriegas	d7c83972d5	tools: Add mode to find python automatically (#159820 ) Add support for automatically finding Python interpreters in manylinux environments to our wheel building script. Scaffolding for sequential builds Signed-off-by: Eli Uriegas <eliuriegas@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159820 Approved by: https://github.com/malfet	2025-08-05 19:19:22 +00:00
Nikita Shulga	e06b110f73	[Testing] Add MPS to NATIVE_DEVICES (#153835 ) This would allow me to enable more opinfo tests against MPS device eventually and supposed to be a very simple test, but actually required minor adjustments to lots of test files, namely: - Introduce `all_mps_types_and` that is very similar to `all_types_and`, but skips `float64` - Decorate lots of tests with `@dtypesIfMPS(*all_mps_types())` - Skip `test_from_dlpack_noncontinguous` as it currently crashes (need to be fixed) - Add lots of `expectedFailureIfMPS` - Delete all `@onlyNativeDeviceTypesAnd("mps")` <sarcasm> I love how well documented this variable are </sarcasm> Pull Request resolved: https://github.com/pytorch/pytorch/pull/153835 Approved by: https://github.com/Skylion007	2025-08-05 18:57:35 +00:00
Zheng, Zhaoqiong	0ba09a6d34	fix link for tutorial of inductor on windows (#159853 ) fix link issue from https://docs.pytorch.org/tutorials/prototype/inductor_windows.html to https://docs.pytorch.org/tutorials/unstable/inductor_windows.html due to structure change with pr https://github.com/pytorch/tutorials/pull/3489 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159853 Approved by: https://github.com/sekyondaMeta Co-authored-by: sekyondaMeta <127536312+sekyondaMeta@users.noreply.github.com> Co-authored-by: Zesheng Zong <zesheng.zong@outlook.com>	2025-08-05 18:37:47 +00:00
Luca Wehrstedt	aeb5321b63	Allow controlling PG backend and options via init_device_mesh (#159371 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159371 Approved by: https://github.com/wconstab, https://github.com/fduwjj, https://github.com/wanchaol	2025-08-05 12:44:14 +00:00
Ruben Rodriguez Buchillon	625108ede2	[inductor] consolidate common GEMM triton param retrieval (#159383 ) \# Why - Make loop iteration simpler - Have a common spot where to make modifications that affect all the GEMM Triton templates, avoiding missed spots \# What - pull out commong logic of taking the BaseConfig objects and turning them into kwargs to feed into maybe_append_choice for Triton GEMM templates Differential Revision: [D79186962](https://our.internmc.facebook.com/intern/diff/D79186962) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159383 Approved by: https://github.com/jansel	2025-08-05 11:42:25 +00:00
Edward Z. Yang	09e5a93fcb	Improve graph output alias with subclass error message (#159619 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159619 Approved by: https://github.com/albanD	2025-08-05 06:47:31 +00:00
Yu, Guangye	908c5cc4c0	Generalize torch._C._set_allocator_settings to be generic (#156175 ) # Motivation This PR moves the implementation of `torch.cuda.memory._set_allocator_settings` to `torch._C._accelerator_setAllocatorSettings`. Since the original API was intended as a temporary/internal utility, I am not exposing the new function as a public API. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156175 Approved by: https://github.com/albanD ghstack dependencies: #159629, #150312, #156165	2025-08-05 04:08:42 +00:00
Yu, Guangye	c1145852a5	Deprecate overleap functions in CUDAAllocatorConfig, use AcceleratorAllocatorConfig instead (#156165 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156165 Approved by: https://github.com/albanD ghstack dependencies: #159629, #150312	2025-08-05 04:08:42 +00:00
Yu, Guangye	ae1a706444	Refactor CUDAAllocatorConfig to reuse AcceleratorAllocatorConfig (#150312 ) # Motivation Refactor `CUDAAllocatorConfig` to reuse `AcceleratorAllocatorConfig` and `ConfigTokenizer`. We would deprecate those option that overleap with `AcceleratorAllocatorConfig` in the following PR and keep them only for BC. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150312 Approved by: https://github.com/albanD ghstack dependencies: #159629	2025-08-05 04:08:04 +00:00
Yu, Guangye	56d19a5ced	Fix AllocatorConfig potential SIO issue (#159629 ) # Motivation As @ScottTodd identified in this [comment](https://github.com/pytorch/pytorch/pull/150312#issuecomment-3141524874), using STL containers like `std::string` and `std::unordered_set` at static init time can cause static initialization order issues. This PR is based on and modified from his original PR: https://github.com/pytorch/pytorch/pull/159607. I’m stacking this PR here to help facilitate the landing and validation process. Co-authored-by: @ScottTodd Pull Request resolved: https://github.com/pytorch/pytorch/pull/159629 Approved by: https://github.com/ScottTodd, https://github.com/albanD	2025-08-05 04:07:51 +00:00
Lucas Kabela	b6c53383fe	[Dynamo][Better Engineering] Type annotation for `torch/_dynamo/output_graph.py` (#159602 ) As part of better engineering effort, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to `torch/_dynamo/output_graph.py` Running ``` mypy torch/_dynamo/output_graph.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Annotated \| Lines Total \| % lines covered \| Funcs Annotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 2163 \| 4792 \| 45.14% \| 121 \| 268 \| 45.15% \| \| This PR \| 4818 \| 4818 \| 100.00% \| 268 \| 268 \| 100.00% \| \| Delta \| +2655 \| +26 \| +54.84% \| +147 \| 0 \| +54.85% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/159602 Approved by: https://github.com/Skylion007	2025-08-05 03:50:54 +00:00
Divyansh Khanna	4fd5fabee9	skip XPU for dataloader CPU only unit test (#159811 ) Fixes [#159802](https://github.com/pytorch/pytorch/issues/159802) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159811 Approved by: https://github.com/izaitsevfb	2025-08-05 03:44:01 +00:00
Nick Riasanovsky	bbc0df1094	[Inductor][Triton] Support TMA before strict 3.4 cutoff (#159777 ) Summary: Inductor's 3.4 Triton release is the most common used variant of Triton, but if someone is working with an alternative version of Triton this may not match. This moves the version check from 3.4 Triton to any variant that has support for the TMA APIs. Test Plan: Relying on CI. Should be a NFC. Rollback Plan: Reviewed By: davidberard98 Differential Revision: D79378792 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159777 Approved by: https://github.com/davidberard98	2025-08-05 03:29:13 +00:00
Mark Harfouche	33ec6e3e9a	Remove pin on libuv from instructions (#159504 ) This package doesn't exist at conda-forge and causes some confusion for users. see https://anaconda.org/conda-forge/libuv/files?version=1.39.0 libuv is quite stable, so the newer versions should be fine. we build with them anyway at conda-forge. see: https://github.com/conda-forge/libuv-feedstock/issues/80 Hopefully this can help future users. Fixes https://github.com/conda-forge/libuv-feedstock/issues/80 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159504 Approved by: https://github.com/seemethere	2025-08-05 03:18:42 +00:00
CaoE	efc4b460b3	Add cascade sum support for Inductor CPP backend (#156296 ) Fixes #154703 Add cascade summation support for Inductor CPP backend to improve precision for large size summation. Currently, Inductor CPP directly do reduction for sum. As shown in #154703, when the size of the sum is large and the number of parallel is small, direct reduction will cause an intolerable precision loss: ``` extern "C" void kernel(float* in_out_ptr0, const float* in_ptr0) { auto out_ptr0 = in_out_ptr0; { { float tmp_acc0 = 0; at::vec::Vectorized<float> tmp_acc0_vec = at::vec::Vectorized<float>(0); for(int64_t x0=static_cast<int64_t>(0L); x0<static_cast<int64_t>(3000000000L); x0+=static_cast<int64_t>(16L)) { { if(C10_LIKELY(x0 >= static_cast<int64_t>(0) && x0 < static_cast<int64_t>(3000000000L))) { auto tmp0 = at::vec::Vectorized<float>::loadu(in_ptr0 + static_cast<int64_t>(x0), static_cast<int64_t>(16)); tmp_acc0_vec = tmp_acc0_vec + tmp0; } } } tmp_acc0 = tmp_acc0 + at::vec::vec_reduce_all<float, 1>([](at::vec::Vectorized<float>& x, at::vec::Vectorized<float>& y) { return x + y; }, tmp_acc0_vec); out_ptr0[static_cast<int64_t>(0L)] = static_cast<float>(tmp_acc0); } } { { { auto tmp0 = out_ptr0[static_cast<int64_t>(0L)]; auto tmp1 = static_cast<float>(3000000000.0); auto tmp2 = tmp0 / tmp1; in_out_ptr0[static_cast<int64_t>(0L)] = tmp2; } } } } ``` After adding cascade sum support: ``` extern "C" void kernel(float* in_out_ptr0, const float* in_ptr0) { auto out_ptr0 = in_out_ptr0; { { float tmp_acc0 = 0; at::vec::Vectorized<float> tmp_acc0_vec = at::vec::Vectorized<float>(0); at::vec::Vectorized<float> masked_tmp_acc0_vec = at::vec::Vectorized<float>(0); CascadeSumHelper<float, 65536> scalar_cascade_helper0(static_cast<int64_t>(3000000000L)); CascadeSumHelper<at::vec::Vectorized<float>, 65536> cascade_helper0(static_cast<int64_t>(187500000L)); CascadeSumHelper<at::vec::Vectorized<float>, 65536> masked_cascade_helper0(static_cast<int64_t>(0L)); for(int64_t x0=static_cast<int64_t>(0L); x0<static_cast<int64_t>(3000000000L); x0+=static_cast<int64_t>(16L)) { { if(C10_LIKELY(x0 >= static_cast<int64_t>(0) && x0 < static_cast<int64_t>(3000000000L))) { auto tmp0 = at::vec::Vectorized<float>::loadu(in_ptr0 + static_cast<int64_t>(x0), static_cast<int64_t>(16)); tmp_acc0_vec = cascade_sum_combine(tmp0, &cascade_helper0); } } } tmp_acc0 = cascade_sum_final(&scalar_cascade_helper0); tmp_acc0_vec = cascade_sum_final(&cascade_helper0); masked_tmp_acc0_vec = cascade_sum_final(&masked_cascade_helper0); tmp_acc0 = tmp_acc0 + at::vec::vec_reduce_all<float, 1>([](at::vec::Vectorized<float>& x, at::vec::Vectorized<float>& y) { return x + y; }, tmp_acc0_vec + masked_tmp_acc0_vec); out_ptr0[static_cast<int64_t>(0L)] = static_cast<float>(tmp_acc0); } } { { { auto tmp0 = out_ptr0[static_cast<int64_t>(0L)]; auto tmp1 = static_cast<float>(3000000000.0); auto tmp2 = tmp0 / tmp1; in_out_ptr0[static_cast<int64_t>(0L)] = tmp2; } } } } ``` This will inevitably reduce performance when cascade sum is turned on. For the case shown in #154703: performance reduced by ~3%. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156296 Approved by: https://github.com/leslie-fang-intel, https://github.com/jansel	2025-08-05 02:54:32 +00:00
Nikita Shulga	1ca8388442	[BE][MPS] Remove unused size12 variable (#159832 ) Fixes following compilation warning ``` /Users/nshulga/git/pytorch/pytorch/aten/src/ATen/native/mps/kernels/Pooling.metal:433:8: warning: unused variable 'size12' [-Wunused-variable] auto size12 = input_sizes[1] * input_sizes[2]; ^ 1 warning generated. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159832 Approved by: https://github.com/dcci	2025-08-05 02:32:06 +00:00
dolpm	b69497351d	[nativert] force resize to zero. (#159683 ) Summary: this was quite a miserable bug. there are a few kernels that don't explicitly resize outputs to zero, which led to some weird UB. Rollback Plan: Differential Revision: D79476454 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159683 Approved by: https://github.com/SherlockNoMad, https://github.com/henryoier	2025-08-05 02:25:31 +00:00
Will Constable	482f069c41	[C10D] fix slow init due to repeated dns resolution failure (#159596 ) It can be be very slow to repeatedly hit DNS resolution failure, but its very helpful to have DNS names in logs by default. So we try to use DNS but if we hit a transient failure we just disable it for the remainder of the job, logging IP addresses instead. Fixes #159007 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159596 Approved by: https://github.com/d4l3k	2025-08-05 02:15:26 +00:00
Benjamin Hottell	85d931f29e	Use uppercase OR when checking for system XNNPACK (#159527 ) This PR fixes `cmake/Dependencies.cmake` to work when compiling with `USE_SYSTEM_XNNPACK=ON` by changing a lowercase `or` to an uppercase `OR`. --- For a personal project, I was building pytorch with a customized build of XNNPACK. When trying to do so I encountered the following error: ``` CMake Error at cmake/Dependencies.cmake:566 (if): if given arguments: "NOT" "XNNPACK_LIBRARY" "or" "NOT" "microkernels-prod_LIBRARY" Unknown arguments specified Call Stack (most recent call first): CMakeLists.txt:868 (include) ``` Upon making the change in this PR (changing `or` to `OR`), the process continued as expected. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159527 Approved by: https://github.com/janeyx99	2025-08-05 02:10:53 +00:00
Jack Taylor	8a2f53c523	Recursively sync fbgemm submodules before build (#159477 ) ROCm inductor benchmark builds failing fbgemm build stage https://ossci-raw-job-status.s3.amazonaws.com/log/46800456622 ``` 2025-07-27T08:00:32.3443858Z /var/lib/jenkins/pytorch/fbgemm/src/RowWiseSparseAdagradFused.cc:389:18: error: no matching function for call to ‘asmjit::v1_17::x86::Vec::Vec(uint32_t)’ 2025-07-27T08:00:32.3444080Z 389 \| x86::Xmm partial_sum_xmm(partial_sum_vreg.id()); ``` It looks like asmjit fails to build, this seems to be due to submodules of fbgemm not being updated after checking out to new commit. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159477 Approved by: https://github.com/pruthvistony, https://github.com/eqy	2025-08-05 02:00:54 +00:00
Kurt Mohler	b59b61a099	Add `avg_pool3d` backward pass for MPS (#159089 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159089 Approved by: https://github.com/malfet	2025-08-05 01:55:38 +00:00
Cui, Yifeng	57ab39f7e4	Update torch-xpu-ops commit pin (#159621 ) Update the torch-xpu-ops commit to [intel/torch-xpu-ops@1f7a57](`1f7a57f507`) includes: - Add Template Parameter to the function `gpu_kernel` for Controlling Broadcasting Vectorization - Add optional NaN checks to XCCL - Fix NllLossForwardReduce2DKernelFunctor accuracy - Extend the existing communication logging to include the reduction operation for collective calls - [Reland] Install xpu codegen header to torch/include Pull Request resolved: https://github.com/pytorch/pytorch/pull/159621 Approved by: https://github.com/EikanWang	2025-08-05 01:46:15 +00:00
Michael Lazos	182975e01a	[Dynamo] Enable torch function dispatch on HOPs (#159708 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159708 Approved by: https://github.com/zou3519, https://github.com/XilunWu ghstack dependencies: #159707	2025-08-05 01:43:22 +00:00
Michael Lazos	9f8cfe7476	[Dynamo] Fix arg ordering in tf modes (#159707 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159707 Approved by: https://github.com/zou3519	2025-08-05 01:43:21 +00:00
Oguz Ulgen	e273ff028a	Fix failing test (#159800 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159800 Approved by: https://github.com/aorenste	2025-08-05 00:28:51 +00:00
David Berard	5e0fc2c9a9	[AOTI] don't allow int32 indices if {non-inf, > int32_max} upper bound is provided (#159433 ) Motivation / Context: (what I _think_ is happening here) In "eager"/just-in-time PT2 usage, dynamo/inductor will guard on whether indices fit in int32 or not. So it's generally safe in Inductor code to rely on the example values for symbolic ints in order to determine whether indices fit in int32, because the indices will be guarded on anyway; and if the inputs ever increase to `>int32_max`, dynamo will cause a recompilation. But with AOTI, those int32 guards aren't respected; so if the example input is `< int32_max` but can be `> int32_max` during future execution, then the future execution might fail / IMA. Solution space Export allows users to specify which dimension are dynamic, and to provide ranges of valid sizes. One solution idea is to always respect the upper bound of the dynamic shape range when doing AOTI; if the index's range includes values `>int32_max`, then don't use the hint and assume that this index doesn't fit in int32. However, the problem with this is that many users may specify dynamism without specifying a range of values - the upper bound of the range will be set to the default of `inf`. Such use cases could potentially experience a perf regression if we implemented the idea above. To prevent any such regressions, this implementation will rely solely on the specified range only if the upper bound of the range isn't inf. In other words, we'll ignore the hints/example values for AOTI (and rely only on the specified range) only if the upper bound of the range isn't inf - if users explicitly specify a range that extends past int32, we can be fairly sure that they actually do need values `>int32_max`. If we continue to see correctness issues even with this implementation, we could consider more aggressively relying on the ranges. Differential Revision: [D79220301](https://our.internmc.facebook.com/intern/diff/D79220301) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159433 Approved by: https://github.com/jingsh, https://github.com/ColinPeppler	2025-08-05 00:17:09 +00:00
Shangdi Yu	bc4b04e058	DeviceCopy should have the same layout as input (#159615 ) Summary: Fix https://github.com/pytorch/pytorch/issues/159612 - Fix the meta implementation of `nan_to_num`, it should preserve the stride of the input - The DeviceCopy IR node should always preserve the input's layout, so we don't end up with a contiguous call during device copy Test Plan: ``` buck2 run @mode/dev-nosan fbcode//caffe2/test/inductor:test_aot_inductor -- -r test_d2h_copy ``` Rollback Plan: Differential Revision: D79411407 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159615 Approved by: https://github.com/eellison	2025-08-04 23:56:58 +00:00
David Berard	6b414f56a4	Revert "[inductor] add lowering for repeat_interleave.Tensor with output size specified (#147160 ) (#158462 )" (#159798 ) This reverts commit 305a03727672de42870f956ddf4ad9fa424443e1. Reason: causes device-side assertion failures when running with this repro (a minimized version of a failure seen in a real model) ``` import torch def ri(inp, repeats, output_size): return torch.repeat_interleave(inp, repeats, output_size=output_size) inp = torch.arange(0, 4, device="cuda").reshape(-1, 1) x = torch.tensor([1, 2, 3, 4], device="cuda") ri_c = torch.compile(ri) print(ri(inp, x, 10)) print(ri_c(inp, x, 10)) ``` which leads to errors like ``` /tmp/torchinductor_dberard/3h/c3hlb22fpptebupstsuhl6kexa6z3upgbnyxln7c24gfcr5747iu.py:30: unknown: block: [0,0,0], thread: [10,0,0] Assertion `index out of bounds: 0 <= tmp5 < 4` failed. ``` Differential Revision: [D79591561](https://our.internmc.facebook.com/intern/diff/D79591561) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159798 Approved by: https://github.com/danzimm	2025-08-04 23:39:20 +00:00
PyTorch MergeBot	fb8f32ef52	Revert "[mps] Turn on inductor dynamic shapes tests (#159456 )" This reverts commit 19f1f9960db7f29f2110a7f49f06a1a23c651ecf. Reverted https://github.com/pytorch/pytorch/pull/159456 on behalf of https://github.com/davidberard98 due to Sorry - this causes a merge conflict with https://github.com/pytorch/pytorch/pull/159798, which I'm trying to land with co-dev to resolve a sev ([comment](https://github.com/pytorch/pytorch/pull/159456#issuecomment-3152751821))	2025-08-04 23:11:05 +00:00
Michael Lazos	7ba996bbaa	[Cutlass] Fix wrapper code generation breakage (#159760 ) Fixes issues introduced by https://github.com/pytorch/pytorch/pull/159355 The issue got past OSS CI because the H100 tag wasn't added, not sure how to prevent these kinds of issues in the future, perhaps we should run H100 on Inductor PRs? Pull Request resolved: https://github.com/pytorch/pytorch/pull/159760 Approved by: https://github.com/angelayi	2025-08-04 23:03:03 +00:00
henrylhtsang	ddbdcdc710	[cutlass backend][test] Expand FP8 tests to FP16 (#159538 ) Differential Revision: [D79317343](https://our.internmc.facebook.com/intern/diff/D79317343/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159538 Approved by: https://github.com/mlazos	2025-08-04 23:01:55 +00:00
angelayi	19f1f9960d	[mps] Turn on inductor dynamic shapes tests (#159456 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159456 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-08-04 22:44:31 +00:00
yewentao256	fd6655a0f5	Feature: Implement support for `cudnn_batch_norm_out` kernel to replace the autogen approach. (#123020 ) Fixes #115611 Autogen kernel may cause redundant copy, so we develop the kernel to improve efficiency. Test Case: ```c++ #include <torch/torch.h> #include <iostream> #include <ATen/ATen.h> #include <ATen/cuda/CUDAContext.h> int main() { auto input = torch::rand({2, 3, 4, 4}, torch::device(torch::kCUDA)); auto weight = torch::randn({3}, torch::device(torch::kCUDA)); auto bias = torch::randn({3}, torch::device(torch::kCUDA)); auto running_mean = torch::zeros({3}, torch::device(torch::kCUDA)); auto running_var = torch::ones({3}, torch::device(torch::kCUDA)); bool training = true; double exponential_average_factor = 0.1; double epsilon = 1e-5; auto output = torch::empty_like(input); auto save_mean = torch::empty({3}, torch::device(torch::kCUDA)); auto save_var = torch::empty({3}, torch::device(torch::kCUDA)); auto reserve = torch::empty({0}, torch::device(torch::kCUDA)); // empty place-holder at::native::cudnn_batch_norm_out(input, weight, bias, running_mean, running_var, training, exponential_average_factor, epsilon, output, save_mean, save_var, reserve); auto outputs = at::native::cudnn_batch_norm(input, weight, bias, running_mean, running_var, training, exponential_average_factor, epsilon); bool is_close_output = torch::allclose(output, std::get<0>(outputs)); bool is_close_save_mean = torch::allclose(save_mean, std::get<1>(outputs)); bool is_close_save_var = torch::allclose(save_var, std::get<2>(outputs)); bool is_close_reserve = torch::allclose(reserve, std::get<3>(outputs)); std::cout << "Is output close: " << is_close_output << std::endl; std::cout << "Is save_mean close: " << is_close_save_mean << std::endl; std::cout << "Is save_var close: " << is_close_save_var << std::endl; std::cout << "Is reserve close: " << is_close_reserve << std::endl; return 0; } ``` Please CC @albanD Pull Request resolved: https://github.com/pytorch/pytorch/pull/123020 Approved by: https://github.com/andrewor14, https://github.com/eqy, https://github.com/albanD	2025-08-04 22:40:33 +00:00
Lucas Kabela	a7f3bdf550	[Dynamo][Better Engineering] Type coverage for `torch/_dynamo/utils.py` (#159580 ) As part of better engineering effort, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to `torch/_dynamo/utils.py` Running ``` mypy torch/_dynamo/utils.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Annotated \| Lines Total \| % lines covered \| Funcs Annotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 2163 \| 4792 \| 45.14% \| 121 \| 268 \| 45.15% \| \| This PR \| 4818 \| 4818 \| 100.00% \| 268 \| 268 \| 100.00% \| \| Delta \| +2655 \| +26 \| +54.84% \| +147 \| 0 \| +54.85% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/159580 Approved by: https://github.com/williamwen42	2025-08-04 21:51:53 +00:00
Xu Han	510e8b4ae0	[inductor] use writable temp file on windows (#159738 ) Use `WritableTempFile` on Windows, reference to: https://github.com/pytorch/pytorch/pull/159342 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159738 Approved by: https://github.com/angelayi, https://github.com/Skylion007	2025-08-04 21:51:02 +00:00
PyTorch MergeBot	83ba3f1101	Revert "[inductor] allocate non-blocking copy destinations in pinned memory (#155121 ) (#158758 )" This reverts commit 6085bf7565fec0d2ed26e8590001f09c05adbbe4. Reverted https://github.com/pytorch/pytorch/pull/158758 on behalf of https://github.com/davidberard98 due to I need to revert #158462 (it causes device-side asserts), and this PR causes a merge conflict in the test file. Sorry about that! ([comment](https://github.com/pytorch/pytorch/pull/158758#issuecomment-3152490371))	2025-08-04 21:47:11 +00:00
PyTorch MergeBot	1fad16aacb	Revert "[inductor] move all cpu scalars using pinned memory for graph partition (#155360 ) (#158983 )" This reverts commit 444e2381d07a14cb501c00d11f9e63a3f1d2c86e. Reverted https://github.com/pytorch/pytorch/pull/158983 on behalf of https://github.com/davidberard98 due to I need to revert #158462 (it causes device-side asserts), and this PR causes a merge conflict in the test file. Sorry about that! ([comment](https://github.com/pytorch/pytorch/pull/158758#issuecomment-3152490371))	2025-08-04 21:47:11 +00:00
Markus Hoehnerbach	444e2381d0	[inductor] move all cpu scalars using pinned memory for graph partition (#155360 ) (#158983 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158983 Approved by: https://github.com/eellison ghstack dependencies: #158758	2025-08-04 21:42:05 +00:00
Markus Hoehnerbach	6085bf7565	[inductor] allocate non-blocking copy destinations in pinned memory (#155121 ) (#158758 ) Fixes #155121 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158758 Approved by: https://github.com/EikanWang, https://github.com/eellison	2025-08-04 21:22:11 +00:00
Natalia Gimelshein	8201dbf4bc	check driver to be >=12.4 to use fabric handles (#159697 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/159697 Approved by: https://github.com/malfet	2025-08-04 21:05:39 +00:00
atalman	26d045bb60	Linux py 3.14 wheel builds (#157559 ) Related to https://github.com/pytorch/pytorch/issues/156856 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157559 Approved by: https://github.com/malfet, https://github.com/albanD	2025-08-04 20:55:19 +00:00
PyTorch MergeBot	356ac3103a	Revert "Stop parsing command line arguments every time common_utils is imported. (#156703 )" This reverts commit 310f901a71e53688866b14bb2f2b4c8eef9979b3. Reverted https://github.com/pytorch/pytorch/pull/156703 on behalf of https://github.com/izaitsevfb due to breaking tests internally with `assert common_utils.SEED is not None` ([comment](https://github.com/pytorch/pytorch/pull/156703#issuecomment-3152337518))	2025-08-04 20:37:39 +00:00
Kurt Mohler	d4109a0f99	[MPS] Add max_unpool1d/2d/3d (#159789 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159789 Approved by: https://github.com/malfet	2025-08-04 20:00:59 +00:00
Arsh Zahed	7ea789ccfb	Revert #156868 : Bring back symint check for sharding propagation cache (#159671 ) Fixes #159601 Unfortunately #156868 introduced a couple regressions (see #159590 and #159601). This reverts the commit while I am working on a permanent fix. This means the `in_compiled_autograd_initial_trace` global flag will be removed and the `_are_we_tracing()` will instead be replaced with the symint preprocessing step during sharding prop post init. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159671 Approved by: https://github.com/xmfan	2025-08-04 19:58:48 +00:00
PyTorch MergeBot	7e8197e34d	Revert "Migrate ScalarType to headeronly (#159416 )" This reverts commit 1371a98b0e727f8a8916dd473b6dd0cff78c0449. Reverted https://github.com/pytorch/pytorch/pull/159416 on behalf of https://github.com/izaitsevfb due to breaking internal builds, see D79452481 ([comment](https://github.com/pytorch/pytorch/pull/159416#issuecomment-3152138508))	2025-08-04 19:55:09 +00:00
Benjamin Glass	50eac811a6	[typing] Constrain OrderedSet generic to be Hashable (#159684 ) Ran across this typing bug while creating an OrderedSet from a type I didn't realize wasn't hashable, which failed at runtime. With this constraint, typing would've failed pre-runtime. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159684 Approved by: https://github.com/Skylion007	2025-08-04 18:08:01 +00:00
ILCSFNO	4e0f179d0b	Update the signature and test of torch.hamming_window() (#152682 ) Fixes #146590 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152682 Approved by: https://github.com/albanD	2025-08-04 17:50:42 +00:00
Tan Hoang	36e59d9b12	[c10d][nvshmem] fix missing override compilation error for nvshmem symmetric code (#159557 ) Summary: Fix error when compiling nvshmem code section `NVSHMEMSymmetricMemory.cu` with BUCK ``` fbcode/caffe2/torch/csrc/distributed/c10d/symm_mem/NVSHMEMSymmetricMemory.cu:154:20: error: 'get_buffer' overrides a member function but is not marked 'override' [-Werror,-Winconsistent-missing-override] 154 \| virtual at::Tensor get_buffer(int \| ^ fbcode/caffe2/torch/csrc/distributed/c10d/symm_mem/SymmetricMemory.hpp:56:20: note: overridden virtual function is here 56 \| virtual at::Tensor get_buffer(int rank, c10::IntArrayRef sizes, c10::ScalarType dtype, int64_t storage_offset) = 0; ``` Test Plan: Build test + CI Rollback Plan: Differential Revision: D78813586 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159557 Approved by: https://github.com/kwen2501	2025-08-04 17:46:30 +00:00
angelayi	fc340d0ca3	[export] Allow comparing device w/o index with device w/ index (#159665 ) In the case where we have expected device "cuda" and given device "cuda:0" I think we should succeed? Pull Request resolved: https://github.com/pytorch/pytorch/pull/159665 Approved by: https://github.com/yushangdi	2025-08-04 17:00:07 +00:00
Animesh Jain	53e47af0f7	[dynamo][guards] Read the attr name from GetAttrGuardAccessor (#159754 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159754 Approved by: https://github.com/jansel ghstack dependencies: #159752	2025-08-04 16:51:27 +00:00
Animesh Jain	66ad881fc7	[dynamo][guards][refactor] Simplify type extraction from GuardManager (#159752 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159752 Approved by: https://github.com/jansel	2025-08-04 16:51:27 +00:00
amdfaa	1d3eef27ac	[ROCm CI] Migrate to MI325 Capacity (#159649 ) Migrate mi300s to gfx942. Related to https://github.com/pytorch/pytorch/pull/159059 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159649 Approved by: https://github.com/huydhn	2025-08-04 16:48:12 +00:00
Xu Han	dd95900cec	[AOTI] normalize_path_separator file path for Windows. (#159726 ) `normalize_path_separator` file path for Windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159726 Approved by: https://github.com/angelayi, https://github.com/jansel	2025-08-04 15:57:19 +00:00
yuchengliu1	1cdd665526	fix test_verbose_logs_dynamic_shapes with MSVC (#159573 ) Operator `typeid` have different outputs in different compiler. There is a good example in [cppreference](https://www.en.cppreference.com/w/cpp/language/typeid.html). Pull Request resolved: https://github.com/pytorch/pytorch/pull/159573 Approved by: https://github.com/angelayi, https://github.com/jansel	2025-08-04 15:56:53 +00:00
Tan Hoang	7cb2dcd2dd	[c10d][nvshmem] modify is_nvshmem_available runtime check to work with static-linked library (#159558 ) (#159561 ) Summary: Currently this function rely on the logic that we load `libnvshmem_device.a` statically and load `libnvshmem_host.so` at runtime. For loading `libnvshmem.a` (the combine 2 thing together) statically this will fail. Add a section to check if the symbol from host API exist at runtime to check if nvshmem is loaded statically Test Plan: CI + sample run Rollback Plan: Differential Revision: D79177525 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159561 Approved by: https://github.com/kwen2501	2025-08-04 15:40:29 +00:00
Aleksei Nikiforov	e5a81aa7ba	Fix conversion of values in libtorch agnostic tests (#155115 ) Due to different byteorder, when copying data, it has to be put into last bytes to ensure that int32_t converted to int64_t keeps same value. Same has to be done when it's converted back. This change fixes test TestLibtorchAgnosticCPU::test_my_ones_like_cpu from cpp_extensions/libtorch_agnostic_extension/test/test_libtorch_agnostic.py on s390x. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155115 Approved by: https://github.com/huydhn	2025-08-04 13:40:22 +00:00
Andrey Talman	3e2aa4b0e3	Update pin to include Python 3.14 support (#159725 ) Update Triton Pin to top of rel/3.4 branch : https://github.com/triton-lang/triton/tree/rel/3.4 . This is the same as release/3.4.x branch but also includes Python 3.14 support This should unblock enablement of Python 3.14 support in this PR: https://github.com/pytorch/pytorch/pull/157559 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159725 Approved by: https://github.com/davidberard98	2025-08-04 13:30:12 +00:00
Aleksei Nikiforov	6646461764	S390X: fix detection of magic number placeholder in inductor (#157784 ) This change fixes multiple tests in test/inductor/test_aot_inductor_arrayref.py such as test_cond_with_parameters_cpu_with_stack_allocation, test_issue_140766_cpu_with_stack_allocation, test_model_modified_weights_cpu_with_stack_allocation, test_nested_tensor_from_jagged_cpu_with_stack_allocation. Enable tests in test/inductor/test_aot_inductor_arrayref.py This change is split off from https://github.com/pytorch/pytorch/pull/150116 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157784 Approved by: https://github.com/huydhn	2025-08-04 12:42:31 +00:00
PyTorch UpdateBot	f74da2a136	[xla hash update] update the pinned xla hash (#159758 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned xla hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159758 Approved by: https://github.com/pytorchbot	2025-08-04 11:21:45 +00:00
eqy	d35b27dde5	[CUDA] Add some more missing `@serialTest` decorators (#159672 ) Seems to fix #159663 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159672 Approved by: https://github.com/Skylion007	2025-08-04 07:44:35 +00:00
anwang	a9dc1566d4	[MTIA Aten Backend] Migrate arange.start_out (#159540 ) Differential Revision: [D79317519](https://our.internmc.facebook.com/intern/diff/D79317519/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159540 Approved by: https://github.com/malfet, https://github.com/nautsimon	2025-08-04 07:38:05 +00:00
Jiang, Yanbing	33a1996714	Fix perf downgrad by reverting template use in use_mkldnn_matmul (#159024 ) This PR is to fix the performance downgrad by reverting template use in `use_mkldnn_matmul` in #157520 . Fix https://github.com/pytorch/pytorch/issues/159031 and https://github.com/pytorch/pytorch/issues/159551. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159024 Approved by: https://github.com/mingfeima	2025-08-04 05:49:46 +00:00
Animesh Jain	ee62177c19	[dynamo] Be consistent with storing func source for UserMethodVariable (#159696 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159696 Approved by: https://github.com/jansel ghstack dependencies: #159534	2025-08-04 05:12:44 +00:00
Animesh Jain	64cbaa876c	[dynamo][guards] Make class members go through obj.__class__.__dict__ (#159534 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159534 Approved by: https://github.com/jansel	2025-08-04 05:12:44 +00:00
Animesh Jain	4516c59f5f	[dynamo][source] Add special source for __code__ and __closure__ (#159722 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159722 Approved by: https://github.com/jansel	2025-08-04 05:02:05 +00:00
PyTorch UpdateBot	8bc843a9ec	[vllm hash update] update the pinned vllm hash (#159610 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned vllm hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159610 Approved by: https://github.com/pytorchbot	2025-08-04 04:06:09 +00:00
Jason Ansel	e39a62c70d	Fix warnings in triton_helpers.py (#159719 ) ``` /home/jansel/pytorch/torch/_inductor/runtime/triton_helpers.py:152: UserWarning: Logical operators 'and' and 'or' are deprecated for non-scalar tensors; please use '&' or '\|' instead equal \|= a_isnan and b_isnan ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159719 Approved by: https://github.com/Skylion007	2025-08-04 03:21:09 +00:00
Laith Sakka	978e3a9142	refresh expected results (#159727 ) Just regular update due to recent <10% changes CI is stable. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159727 Approved by: https://github.com/anijain2305	2025-08-03 22:47:50 +00:00
Nikita Shulga	e2a5c42e7e	[BE][MPS] Build metal kernels of MacOS-14+ (#159733 ) Which makes `#if __METAL_VERSION__ >= 310` guards for `bfloat` use support unnecessary. Rename `kernels_bfloat.metallib` into `kernels_basic` and remove custom build/selection logic. Part of https://github.com/pytorch/pytorch/issues/159275 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159733 Approved by: https://github.com/dcci ghstack dependencies: #159731, #159732	2025-08-03 20:53:58 +00:00
Nikita Shulga	5116c49b52	[BE] Remove macos-13 guard from bench_mps_ops (#159732 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159732 Approved by: https://github.com/dcci ghstack dependencies: #159731	2025-08-03 20:53:58 +00:00
Nikita Shulga	fecdebe385	[CI][MPS] Fix compile benchmark correctness (#159731 ) By passing `fullgraph=True` attribute and increasing cache size limit to 2**16 Otherwise, compiler might decide not to fall back to eager to avoid recompilations Pull Request resolved: https://github.com/pytorch/pytorch/pull/159731 Approved by: https://github.com/dcci	2025-08-03 20:53:50 +00:00
Nikita Shulga	e136a9175b	[BE] Fix dev warning in `Dependencies.cmake` (#159702 ) Namely ``` CMake Warning (dev) in cmake/Dependencies.cmake: A logical block opening on the line /Users/nshulga/git/pytorch/pytorch/cmake/Dependencies.cmake:261 (if) closes on the line /Users/nshulga/git/pytorch/pytorch/cmake/Dependencies.cmake:263 (endif) with mis-matching arguments. ``` Introduced by https://github.com/pytorch/pytorch/pull/143846 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159702 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-08-03 18:45:07 +00:00
Francisco Massa	9a680e14b7	[bucketing] Reduce CPU overhead for reduce_scatter_merge_fn_to_trace (#159723 ) The previous implementation was creating `n_gpu * n_tensors` intermediate tensors, which was adding a lot of CPU overhead, specially given that inductor was generating a number of individual tensor copy kernels for `torch.cat` . This PR changes the implementation so that only `n_tensors` are created, making the CPU overhead proportional to the number of tensors being bucketed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159723 Approved by: https://github.com/IvanKobzarev	2025-08-03 09:16:55 +00:00
PyTorch MergeBot	805a102beb	Revert "[dynamo][guards] Make class members go through obj.__class__.__dict__ (#159534 )" This reverts commit 1616777cd2a3170ff76afa3e7860b0969420c445. Reverted https://github.com/pytorch/pytorch/pull/159534 on behalf of https://github.com/malfet due to Broke some inductor test and lint among other things, see `9c18901bfd/1` ([comment](https://github.com/pytorch/pytorch/pull/159534#issuecomment-3146983186))	2025-08-03 04:58:32 +00:00
PyTorch MergeBot	6e8d705a22	Revert "[dynamo] Be consistent with storing func source for UserMethodVariable (#159696 )" This reverts commit be71000ff5292293d1976f313218e2df4d5046d3. Reverted https://github.com/pytorch/pytorch/pull/159696 on behalf of https://github.com/malfet due to Broke some inductor test and lint among other things, see `9c18901bfd/1` ([comment](https://github.com/pytorch/pytorch/pull/159534#issuecomment-3146983186))	2025-08-03 04:58:32 +00:00
anwang	9c18901bfd	[MTIA Aten Backend] Migrate all.out (#159539 ) Differential Revision: [D79317033](https://our.internmc.facebook.com/intern/diff/D79317033/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159539 Approved by: https://github.com/malfet ghstack dependencies: #159098	2025-08-03 02:08:35 +00:00
Oguz Ulgen	a29ed5e1ac	Add torch compile force disable caches alias (#158072 ) Bunch of people keep thinking current alias only disables inductor cache because it has the name inductor in it. lets globalize the name Pull Request resolved: https://github.com/pytorch/pytorch/pull/158072 Approved by: https://github.com/ezyang	2025-08-02 23:23:17 +00:00
Francisco Massa	d2792f51b2	[bucketing] Use max of input/output size for bucketing (#159717 ) The output of a reduce_scatter is n_gpu times smaller than its input, while the output of an all_gather is n_gpu times larger than its input. This means that in the current heuristic for bucketing reduce_scatter, we would need to use a bucket size which is n_gpu times larger than the bucket for all_gather, making it gpu-dependent and less intuitive. This PRs propose to use instead the max between the input and output sizes, so that one can use the same bucket_size value for both passes Pull Request resolved: https://github.com/pytorch/pytorch/pull/159717 Approved by: https://github.com/wconstab	2025-08-02 22:42:22 +00:00
Animesh Jain	be71000ff5	[dynamo] Be consistent with storing func source for UserMethodVariable (#159696 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159696 Approved by: https://github.com/jansel ghstack dependencies: #159186, #159534	2025-08-02 21:40:38 +00:00
Aaron Orenstein	3f86076775	gc before warming up benchmarking (#159670 ) #158649 turned off automatic GCs during cudagraph recording. This is causing a small uptick in some internal benchmark numbers because of memory the benchmark is leaving around before the benchmark starts - so GC before warming up the model. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159670 Approved by: https://github.com/oulgen	2025-08-02 19:37:24 +00:00
Animesh Jain	1616777cd2	[dynamo][guards] Make class members go through obj.__class__.__dict__ (#159534 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159534 Approved by: https://github.com/jansel ghstack dependencies: #159186	2025-08-02 18:04:35 +00:00
rajeshvshiyal	38895c0ac2	Update RuntimeError message in is_nonzero(input) method from bool to Boolean (#159712 ) RuntimeError message updated in is_nonzero(input) method from bool to Boolean. Case 1: t = torch.tensor([]) torch.is_nonzero(t) Case 2: t = torch.tensor([1,2]) torch.is_nonzero(t) Existing Error message in documentation: for case 1: RuntimeError: bool value of Tensor with no values is ambiguous for case 2: RuntimeError: bool value of Tensor with more than one value is ambiguous Proposed Error message in documentation: for case 1: RuntimeError: Boolean value of Tensor with no values is ambiguous for case 2: RuntimeError: Boolean value of Tensor with more than one value is ambiguous Fixes #159710 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159712 Approved by: https://github.com/malfet	2025-08-02 17:23:45 +00:00
Anthony Barbier	310f901a71	Stop parsing command line arguments every time common_utils is imported. (#156703 ) Last PR in the series to re-submit https://github.com/pytorch/pytorch/pull/134592 as smaller PRs: https://github.com/pytorch/pytorch/pull/154612 https://github.com/pytorch/pytorch/pull/154628 https://github.com/pytorch/pytorch/pull/154715 https://github.com/pytorch/pytorch/pull/154716 https://github.com/pytorch/pytorch/pull/154725 https://github.com/pytorch/pytorch/pull/154728 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156703 Approved by: https://github.com/clee2000	2025-08-02 16:38:54 +00:00
Nichols A. Romero	e11b1cd97e	[ROCm] fix nightly wheel due to rocBLAS environment variable (#159570 ) Fixes #159070 The TunableOp failure is due to missing rocBLAS files in our manywheels packaging. This bug has been present since June 7-8 time frame. It was caused by a typo in the rocBLAS environment variable that stores the list of files. It was introduced in this PR: https://github.com/pytorch/pytorch/pull/155388 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159570 Approved by: https://github.com/malfet	2025-08-02 06:54:43 +00:00
Wenyuan Chi	b599d91738	Log autotune choices and benchmark result to scuba/chrome trace (#159496 ) Summary: Report the kernel choices and benchmark data to better understand how kernels are selected and the performance gap between the best kernel (likely a CUDA kernel) and Triton kernels. Example Event: mm_template_autotuning Column: autotune_choices ```json { "num_choices": 52, "num_triton_choices": 19, "best_kernel": "cutlass_f6c25cf2", "best_kernel_desc": "cutlass3x_sm90_tensorop_gemm_f16_f16_f32_void_f16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8", "best_time": 0.6283040046691895, "best_triton_pos": 26, "best_triton_time": 0.6832960247993469, "best_triton_kernel": "triton_mm_17", "best_triton_kernel_desc": "ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4, num_consumer_groups=0, num_buffers_warp_spec=0" } ``` Test Plan: ``` TORCHINDUCTOR_MAX_AUTOTUNE_REPORT_CHOICES_STATS =1 buck2 run //scripts/wychi:test_autotune_mm 2>&1 > /tmp/mylog.txt ``` Rollback Plan: Differential Revision: D79235037 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159496 Approved by: https://github.com/masnesral	2025-08-02 05:34:17 +00:00
Xiao, Wang	fd6a6658c3	Enable _int_mm on Intel GPU (#157769 ) # Moativation This PR is used to enable _int_mm on Intel GPU. And _int_mm is used by int8 quantization on torchao. # Model Test Result: We run meta-llama/Llama-3.1-8B-Instruct on Intel GPU and A100 using torchao int8-dynamic-quantization. The model configs as below: Precision : torch.bfloat16 quantization configuration : Int8DynamicActivationInt8WeightConfig dataset : wikitext Result: The perplexity values for Intel GPU and A100 are 9.582953453063965 and 9.57755184173584, respectively. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157769 Approved by: https://github.com/EikanWang, https://github.com/desertfire	2025-08-02 05:16:01 +00:00
PyTorch UpdateBot	04973496a8	[audio hash update] update the pinned audio hash (#159611 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159611 Approved by: https://github.com/pytorchbot	2025-08-02 05:15:47 +00:00
Sam Larsen	1548b011ea	Fix rand_like decomposition to preserve strides (#159294 ) Summary: Like https://github.com/pytorch/pytorch/pull/158898, the rand_like variants are not preserving strides. Followed the pattern established in https://github.com/pytorch/pytorch/pull/158898. Test Plan: New unit test (fails before this PR; but fixed after) Differential Revision: [D79472604](https://our.internmc.facebook.com/intern/diff/D79472604) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159294 Approved by: https://github.com/eellison	2025-08-02 03:54:41 +00:00
angelayi	e57a92734d	[export] Fix nn_module_stack of assert_tensor_metadata nodes (#159625 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/159625 Approved by: https://github.com/yushangdi	2025-08-02 02:52:42 +00:00
Dylan Maloy	79ff3b320b	Back out "[ez] get rid of unused var" (#159677 ) Summary: turns out i added this to reduce the frequency we'd call try_update_max_size_at_index when a new maximum is found before the replan is called. oops. Test Plan: backout Rollback Plan: Differential Revision: D79474114 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159677 Approved by: https://github.com/georgiaphillips	2025-08-02 01:50:16 +00:00
nandesuka	426f249f20	Fix launch grid calculation (#159497 ) Summary: The launch grid calculation code is using a python trick to achieve CeilDiv() through negative integer division with FloorDiv(). This is language dependent behaviour that doesn't apply to all languages. In the FXIR backend we negate this behaviour and replace the experssion with CeilDiv() operation so the computation is correct regardless of language used. Not directly directly changing the orginal computation as it leads to a performance degredation. Test Plan: CI Rollback Plan: Differential Revision: D79275534 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159497 Approved by: https://github.com/blaine-rister	2025-08-02 01:12:58 +00:00
Edward Z. Yang	d33a484763	Use boxed_nop_preserve_node_meta for aot_export_joint_with_descriptors (#159545 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159545 Approved by: https://github.com/xmfan, https://github.com/wconstab ghstack dependencies: #159336, #159337	2025-08-02 00:33:41 +00:00
Natalia Gimelshein	a81ffbc5f5	improve shape checks for grouped_mm (#159666 ) Check that contraction dimension matches between tensors if it's known, and do device-side checks for correct offsets Pull Request resolved: https://github.com/pytorch/pytorch/pull/159666 Approved by: https://github.com/danielvegamyhre, https://github.com/eqy	2025-08-02 00:12:25 +00:00
Huy Do	465fe4d9f7	Enable sample nightly PT2 benchmark on B200 (#158011 ) Per the discussion with @nWEIdia, this resumes the work on https://github.com/pytorch/pytorch/pull/157870 to enable PT2 benchmark on B200 ### Testing https://github.com/pytorch/pytorch/actions/runs/16615101382 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158011 Approved by: https://github.com/nWEIdia, https://github.com/atalman	2025-08-01 23:47:44 +00:00
Natalia Gimelshein	9477af1063	fix compilation on cuda < 12.3 (#159657 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/159657 Approved by: https://github.com/kwen2501	2025-08-01 23:40:55 +00:00
Lucas Kabela	dcc36e38bb	[Graph Breaks] Remove unsupported Additional Info field (#159658 ) Race condition when landing PR#158800 caused us to add this field when it is deprecated, so remove it Pull Request resolved: https://github.com/pytorch/pytorch/pull/159658 Approved by: https://github.com/williamwen42	2025-08-01 23:25:50 +00:00
Zain Rizvi	efd78584a8	[EZ] Add linux-aarch64.yml workflow to the viable/strict blocking set (#159668 ) Since it's required to be run on every PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/159668 Approved by: https://github.com/malfet	2025-08-01 23:19:08 +00:00
Oguz Ulgen	135762ea20	Unpin helion (#159579 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159579 Approved by: https://github.com/jansel	2025-08-01 23:08:06 +00:00
Sherlock Huang	e2ee9cfaa2	[NativeRT] Turn on enableStaticCPUKernels by default (#159422 ) Summary: As title. Test Plan: Need to manual test on production models. Rollback Plan: Differential Revision: D78747742 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159422 Approved by: https://github.com/dolpm	2025-08-01 22:27:07 +00:00
Andy Lugo	06d28de17a	Update CK Kernel generation and update ck submodule (#157964 ) changes required to reduce the number of ck kernels generated. This change depends on https://github.com/ROCm/composable_kernel/pull/2480 to be merged first. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157964 Approved by: https://github.com/842974287	2025-08-01 22:24:27 +00:00
anwang	df9720b8b5	[MTIA Aten Backend] Migrate all foreach ops (#159098 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate all foreach operators to in-tree, including: - _foreach_abs - _foreach_abs_ - _foreach_add.List - _foreach_add_.List - _foreach_add_.Scalar - _foreach_add_.Tensor - _foreach_addcmul.Scalar - _foreach_addcmul_.Scalar - _foreach_copy - _foreach_copy_ - _foreach_mul.List - _foreach_mul_.List - _foreach_mul_.Scalar - _foreach_mul.Tensor - _foreach_mul_.Tensor - _foreach_norm.Scalar - _foreach_sqrt_ Differential Revision: [D78913847](https://our.internmc.facebook.com/intern/diff/D78913847/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159098 Approved by: https://github.com/malfet	2025-08-01 22:10:12 +00:00
Sandeep Narendranath Karjala	85e74d5ace	[inductor] Add logging for distributed collective ops for multi‑rank diagnostics (#159190 ) This change introduces structured logging of the collective communication schedule, enabling downstream tools (e.g. TLParse) to ingest and analyze per‑rank collective‐order information for multi‑rank jobs. - Iterates over scheduler.nodes, filters for _CollectiveKernel nodes - Extracts each op’s python_kernel_name - Emits a structured JSON payload under the inductor_collective_schedule artifact name - Dumps the full schedule list to collective_schedule.json via the PyTorch trace‑structured artifact - Added comprehensive unit tests for collective schedule tracing: Created test_collective_schedule_empty() and test_collective_schedule_real() tests to verify structured trace logging works correctly for both empty collective schedules and real collective operations (like all_reduce and wait_tensor from _c10d_functional ops). Pull Request resolved: https://github.com/pytorch/pytorch/pull/159190 Approved by: https://github.com/yushangdi, https://github.com/xmfan	2025-08-01 21:51:42 +00:00
Sheng Fu	0450f05658	Output tensor meta data for FX graph node (#159311 ) FX graph segment in CompiledFxGraph does not include tensor meta data, for example, tensor shape, tensor stride, tensor data type, tensor device. AI system co-design team requested to include these information in FX graph segment so they can use FX graph segment to project the performance on different hardware. This DIFF is to modify the Graph::Node::format_node to include tensor meta data. Before this DIFF, the triton kernel FX graph segment looks like the following: ``` # %mm : Tensor "f32[4, 4][4, 1]cuda:0" = PlaceHolder[target=mm] # %arg2_1 : Tensor "f32[4, 4][4, 1]cuda:0" = PlaceHolder[target=arg2_1] # %sin : Tensor "f32[4, 4][4, 1]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.sin.default](args = (%mm,), kwargs = {}) # %permute_1 : [num_users=1] = call_function[target=torch.ops.aten.permute.default](args = (%sin, [1, 0]), kwargs = {}) # %mul : [num_users=1] = call_function[target=torch.ops.aten.mul.Tensor](args = (%arg2_1, 1111), kwargs = {}) # %add : [num_users=1] = call_function[target=torch.ops.aten.add.Tensor](args = (%permute_1, %mul), kwargs = {}) # %cos : cuda:0"[num_users=1] = call_function[target=torch.ops.aten.cos.default](args = (%add,), kwargs = {}) # return %cos After this DIFF: # %mm : Tensor "f32[4, 4][4, 1]cuda:0" = PlaceHolder[target=mm] # %arg2_1 : Tensor "f32[4, 4][4, 1]cuda:0" = PlaceHolder[target=arg2_1] # %sin : Tensor "f32[4, 4][4, 1]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.sin.default](args = (%mm,), kwargs = {}) # %permute_1 : Tensor "f32[4, 4][1, 4]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.permute.default](args = (%sin, [1, 0]), kwargs = {}) # %mul : Tensor "f32[4, 4][4, 1]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.mul.Tensor](args = (%arg2_1, 1111), kwargs = {}) # %add : Tensor "f32[4, 4][1, 4]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.add.Tensor](args = (%permute_1, %mul), kwargs = {}) # %cos : Tensor "f32[4, 4][1, 4]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.cos.default](args = (%add,), kwargs = {}) # return %cos ``` If format_node can not be changed, I can copy the code to caffe2/torch/_inductor/utils.py. Differential Revision: D77973076 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159311 Approved by: https://github.com/angelayi	2025-08-01 21:40:29 +00:00
zeshengzong	595a65f5c2	[dynamo] Replace unimplemented with unimplemented_v2 in `torch/_dynamo/variables/script_object.py` (#159343 ) Fixes part of #147913 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159343 Approved by: https://github.com/williamwen42 Co-authored-by: William Wen <william.wen42@gmail.com>	2025-08-01 21:30:41 +00:00
tiandeyu-cs	8c6c2e40eb	Edit a test case to detect potential bugs in all-gathering noncontiguous inputs in the Gloo backend (#159542 ) As suggested in the pull request #158903 by @H-huang, this pull request edits a test case to detect potential bugs in all-gathering noncontiguous inputs in the Gloo backend. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159542 Approved by: https://github.com/d4l3k, https://github.com/H-Huang	2025-08-01 21:20:25 +00:00
henrylhtsang	32840d19f9	[cutlass backend] skip stream k if shape is dynamic (#159442 ) Differential Revision: [D79229210](https://our.internmc.facebook.com/intern/diff/D79229210/) Motivation is workspace size is hard to determine, and varies for different shape. What I observed is sometimes the shape got smaller, but the workspace can increase. So it is hard to upper bound it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159442 Approved by: https://github.com/ColinPeppler	2025-08-01 20:42:24 +00:00
Xuehai Pan	2040f00112	[BE][Easy] respect `os.environ` in subprocess calls in tools/nightly.py (#159572 ) Respect parent shell's envvars, such as `UV_INDEX_STRATEGY`, `http{,s}_proxy`, etc. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159572 Approved by: https://github.com/Skylion007	2025-08-01 20:40:31 +00:00
Lucas Kabela	c137f9da0b	[Dynamo][Better Engineering] Add type coverage to dynamo/compiled_autograd.py (#159518 ) As part of better engineering effort, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to `torch/_dynamo/compiled_autograd.py` Running ``` mypy torch/_dynamo/compiled_autograd.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Annotated \| Lines Total \| % lines covered \| Funcs Annotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 425 \| 1553 \| 27.37% \| 17 \| 62 \| 27.42% \| \| This PR \| 1623 \| 1623 \| 100.00% \| 62 \| 62 \| 100.00% \| \| Delta \| +1198\| +0 \| +72.63% \| +45 \| 0 \| +72.58% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/159518 Approved by: https://github.com/xmfan	2025-08-01 20:24:58 +00:00
Howard Huang	5e8b95605f	[PP] Support OVERLAP_F_B computation type (#158978 ) Some changes to validation code and visualizer to support a new computation type that will be used in DualPipeV (see https://github.com/pytorch/pytorch/pull/159591) The IR looks like: ``` [0F0, 0F1, 0F2, 0F3, 0F4, 0F5, 0F6, 7F0, 7I0, 7W0, 7F1, 7I1, 7W1, 7F2, 7I2, 7W2, 7F3, (0F7;7B3)OVERLAP_F_B, (7F4;0B0)OVERLAP_F_B, (0F8;7B4)OVERLAP_F_B, (7F5;0B1)OVERLAP_F_B, (0F9;7B5)OVERLAP_F_B, (7F6;0B2)OVERLAP_F_B, 7B6, (7F7;0B3)OVERLAP_F_B, 7B7, (7F8;0B4)OVERLAP_F_B, 7B8, (7F9;0B5)OVERLAP_F_B, 7B9, 0I6, 0W6, 0I7, 0W7, 0I8, 0W8, 0I9, 0W9] [1F0, 1F1, 1F2, 1F3, 1F4, 6F0, 1F5, 6F1, 6I0, 6W0, 6F2, 6I1, 6W1, 6F3, (1F6;6B2)OVERLAP_F_B, (6F4;1B0)OVERLAP_F_B, (1F7;6B3)OVERLAP_F_B, (6F5;1B1)OVERLAP_F_B, (1F8;6B4)OVERLAP_F_B, (6F6;1B2)OVERLAP_F_B, (1F9;6B5)OVERLAP_F_B, (6F7;1B3)OVERLAP_F_B, 6B6, (6F8;1B4)OVERLAP_F_B, 6B7, (6F9;1B5)OVERLAP_F_B, 6B8, 1B6, 6I9, 1I7, 6W9, 1I8, 1W7, 1I9, 1W8, 1W9] [2F0, 2F1, 2F2, 5F0, 2F3, 5F1, 2F4, 5F2, 5I0, 5W0, 5F3, (2F5;5B1)OVERLAP_F_B, (5F4;2B0)OVERLAP_F_B, (2F6;5B2)OVERLAP_F_B, (5F5;2B1)OVERLAP_F_B, (2F7;5B3)OVERLAP_F_B, (5F6;2B2)OVERLAP_F_B, (2F8;5B4)OVERLAP_F_B, (5F7;2B3)OVERLAP_F_B, (2F9;5B5)OVERLAP_F_B, (5F8;2B4)OVERLAP_F_B, 5B6, (5F9;2B5)OVERLAP_F_B, 5B7, 2B6, 5B8, 2I7, 5I9, 2I8, 2W7, 2I9, 5W9, 2W8, 2W9] [3F0, 4F0, 3F1, 4F1, 3F2, 4F2, 3F3, 4F3, 3F4, 4B0, (4F4;3B0)OVERLAP_F_B, (3F5;4B1)OVERLAP_F_B, (4F5;3B1)OVERLAP_F_B, (3F6;4B2)OVERLAP_F_B, (4F6;3B2)OVERLAP_F_B, (3F7;4B3)OVERLAP_F_B, (4F7;3B3)OVERLAP_F_B, (3F8;4B4)OVERLAP_F_B, (4F8;3B4)OVERLAP_F_B, (3F9;4B5)OVERLAP_F_B, (4F9;3B5)OVERLAP_F_B, 4B6, 3B6, 4B7, 3B7, 4I8, 3I8, 4I9, 3I9, 4W8, 3W8, 4W9, 3W9] ``` In this PR, the schedule execution will just treat the OVERLAP_F_B as two separate operations of F and B (so there is no actual overlap). The next step is to allow users to create a custom function to plug in what this operation does. `814629043a/torch/distributed/pipelining/schedules.py (L1205-L1216)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158978 Approved by: https://github.com/wconstab	2025-08-01 20:22:30 +00:00
Jane Xu	8ea86a6e31	Actually test STD_TORCH_CHECK, add testfile to CMake (#159603 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159603 Approved by: https://github.com/Skylion007, https://github.com/albanD	2025-08-01 19:53:41 +00:00
PyTorch MergeBot	acad808545	Revert "[inductor] consolidate common GEMM triton param retrieval (#159383 )" This reverts commit e7cc42df58a86bee05944f6e80c535aa1d099443. Reverted https://github.com/pytorch/pytorch/pull/159383 on behalf of https://github.com/jataylo due to sorry but rocm CI is broken due to this PR ([comment](https://github.com/pytorch/pytorch/pull/159383#issuecomment-3145604831))	2025-08-01 19:49:21 +00:00
PyTorch MergeBot	c687446374	Revert "Fix rand_like decomposition to preserve strides (#159294 )" This reverts commit 2c46922ce4b33c39b1c48c302604805510a3f889. Reverted https://github.com/pytorch/pytorch/pull/159294 on behalf of https://github.com/yangw-dev due to breaking internal test ([comment](https://github.com/pytorch/pytorch/pull/159294#issuecomment-3145541845))	2025-08-01 19:19:51 +00:00
Will Constable	dd22ba09b4	[C10D] Document barrier interaction with device_id (#159389 ) Addresses #159262 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159389 Approved by: https://github.com/malfet, https://github.com/H-Huang, https://github.com/kwen2501, https://github.com/fduwjj	2025-08-01 18:12:21 +00:00
Yu, Guangye	c0e0126399	Remove unused input parameter in ExpandableSegment (#159356 ) # Motivation While refactoring the caching allocator, I noticed that the `ExpandableSegment` constructor on CUDA had an unused parameter. This change removes that unused argument to avoid potential confusion. # Additional Context I noticed that `ExpandableSegment` is defined in cpp file, so it should be safe to make this change. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159356 Approved by: https://github.com/ngimel, https://github.com/albanD ghstack dependencies: #159159	2025-08-01 17:47:51 +00:00
Ivan Zaitsev	e4b123b5e4	Revert direct updates (#159654 ) reverts: ``` commit 5711a8f06948eeee56ed5f53f171fa519f78491c (tag: trunk/5711a8f06948eeee56ed5f53f171fa519f78491c, origin/main, main) Author: Jovian Anthony Jaison <38627145+jovianjaison@users.noreply.github.com> Date: Fri Aug 1 09:32:52 2025 -0700 Update test_utils.py commit b4b71d011ed07a41c2086ff0dec2988a63662877 (tag: trunk/b4b71d011ed07a41c2086ff0dec2988a63662877) Author: Jovian Anthony Jaison <38627145+jovianjaison@users.noreply.github.com> Date: Fri Aug 1 09:27:54 2025 -0700 Update utils.py commit 52376b9b6fbf9fe24f5d82038dc520f0c64b6f8d (tag: trunk/52376b9b6fbf9fe24f5d82038dc520f0c64b6f8d) Author: Jovian Anthony Jaison <38627145+jovianjaison@users.noreply.github.com> Date: Fri Aug 1 09:26:05 2025 -0700 ``` (commits pushed directly to main by mistake) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159654 Approved by: https://github.com/atalman	2025-08-01 16:54:51 +00:00
Jovian Anthony Jaison	5711a8f069	Update test_utils.py	2025-08-01 09:32:52 -07:00
Jovian Anthony Jaison	b4b71d011e	Update utils.py	2025-08-01 09:27:54 -07:00
Jovian Anthony Jaison	52376b9b6f	Update convert_frame.py	2025-08-01 09:26:05 -07:00
Jane Xu	1371a98b0e	Migrate ScalarType to headeronly (#159416 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159416 Approved by: https://github.com/albanD ghstack dependencies: #159415, #159411	2025-08-01 16:07:01 +00:00
linmin	2a286cbdf4	Allow register_buffer with Tensor-like object (#159455 ) As torch allows extending the tensor with `__torch_function__`, it would be desirable to allow registering it as a buffer. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159455 Approved by: https://github.com/mikaylagawarecki	2025-08-01 15:31:38 +00:00
Scott Todd	7c37b8e1e0	[ROCm][Windows] Switch __builtin_clz ifdef from WIN32 to MSC_VER. (#159273 ) PyTorch with ROCm on Windows is built with clang-cl and not MSVC. This code path is specific to the MSVC compiler so it should be checking for MSC_VER, not just WIN32. The change here is similar to https://github.com/pytorch/pytorch/pull/146606. This fixes downstream build errors using clang-cl like https://github.com/ROCm/TheRock/actions/runs/16569646709/job/46858176812 (patched and tested downstream at https://github.com/ROCm/TheRock/pull/1140): ``` [7099/7147] Building CXX object functorch\CMakeFiles\functorch.dir\csrc\dim\dim.cpp.obj FAILED: functorch/CMakeFiles/functorch.dir/csrc/dim/dim.cpp.obj C:\home\runner\_work\_tool\Python\3.11.9\x64\Lib\site-packages\_rocm_sdk_devel\lib\llvm\bin\clang-cl.exe /nologo -TP -DEXPORT_AOTI_FUNCTIONS -DFUNCTORCH_BUILD_MAIN_LIB -DMINIZ_DISABLE_ZIP_READER_CRC32_CHECKS -DNOMINMAX -DONNXIFI_ENABLE_EXT=1 -DONNX_ML=1 -DONNX_NAMESPACE=onnx_torch -DROCM_ON_WINDOWS -DROCM_USE_FLOAT16 -DROCM_VERSION=70000 -DTORCH_API_INCLUDE_EXTENSION_H -DTORCH_EXTENSION_NAME=_C -DTORCH_HIP_VERSION=700 -DUSE_EXTERNAL_MZCRC -DUSE_MIMALLOC -DUSE_PROF_API=1 -DWIN32_LEAN_AND_MEAN -D_CRT_SECURE_NO_DEPRECATE=1 -D_UCRT_LEGACY_INFINITY -D__HIP_PLATFORM_AMD__ -D__HIP_PLATFORM_AMD__=1 -Dfunctorch_EXPORTS -IB:\src\torch\build\aten\src -IB:\src\torch\aten\src -IB:\src\torch\build -IB:\src\torch -IB:\src\torch\nlohmann -IB:\src\torch\moodycamel -IB:\src\torch\third_party\mimalloc\include -IB:\src\torch\functorch -IB:\src\torch\torch\csrc\api -IB:\src\torch\torch\csrc\api\include -IB:\src\torch\c10\.. -IB:\src\torch\c10\hip\..\.. -IB:\src\torch\torch\.. -IB:\src\torch\torch\..\aten\src -IB:\src\torch\torch\..\aten\src\TH -IB:\src\torch\build\caffe2\aten\src -IB:\src\torch\build\third_party -IB:\src\torch\build\third_party\onnx -IB:\src\torch\torch\..\third_party\valgrind-headers -IB:\src\torch\torch\..\third_party\gloo -IB:\src\torch\torch\..\third_party\onnx -IB:\src\torch\torch\..\third_party\flatbuffers\include -IB:\src\torch\torch\..\third_party\kineto\libkineto\include -IB:\src\torch\torch\..\third_party\cpp-httplib -IB:\src\torch\torch\..\third_party\nlohmann\include -IB:\src\torch\torch\csrc -IB:\src\torch\torch\lib -IB:\src\torch\torch\standalone -IB:\src\torch\torch\lib\libshm_windows -imsvcC:\home\runner\_work\_tool\Python\3.11.9\x64\Lib\site-packages\_rocm_sdk_devel\include -imsvcB:\src\torch\third_party\protobuf\src -imsvcB:\src\torch\third_party\XNNPACK\include -imsvcB:\src\torch\third_party\ittapi\include -imsvcB:\src\torch\cmake\..\third_party\eigen -imsvcB:\src\torch\third_party\ideep\mkl-dnn\include\oneapi\dnnl -imsvcB:\src\torch\third_party\ideep\include -imsvcB:\src\torch\INTERFACE -imsvcB:\src\torch\third_party\nlohmann\include -imsvcB:\src\torch\third_party\concurrentqueue -imsvcC:\home\runner\_work\_tool\Python\3.11.9\x64\Lib\site-packages\_rocm_sdk_devel\include\hiprand -imsvcC:\home\runner\_work\_tool\Python\3.11.9\x64\Lib\site-packages\_rocm_sdk_devel\include\rocrand -imsvcB:\src\torch\cmake\..\third_party\pybind11\include -imsvcC:\home\runner\_work\_tool\Python\3.11.9\x64\include /DWIN32 /D_WINDOWS /EHsc /Zc:__cplusplus /bigobj /FS /utf-8 -DUSE_PTHREADPOOL -DNDEBUG -DUSE_FBGEMM -DUSE_XNNPACK -DSYMBOLICATE_MOBILE_DEBUG_HANDLE /wd4624 /wd4068 /wd4067 /wd4267 /wd4661 /wd4717 /wd4244 /wd4804 /wd4273 /O2 /Ob2 /DNDEBUG /bigobj -DNDEBUG -std:c++17 -MD -Z7 -Wmissing-prototypes -Werror=missing-prototypes /permissive- /d2implyavx512upperregs- /EHsc /bigobj -fms-runtime-lib=dll -D__HIP_PLATFORM_AMD__=1 -DCUDA_HAS_FP16=1 -DUSE_ROCM -D__HIP_NO_HALF_OPERATORS__=1 -D__HIP_NO_HALF_CONVERSIONS__=1 -DTORCH_HIP_VERSION=700 -Wno-shift-count-negative -Wno-shift-count-overflow -Wno-duplicate-decl-specifier -DCAFFE2_USE_MIOPEN -DTHRUST_DEVICE_SYSTEM=THRUST_DEVICE_SYSTEM_HIP -std=c++17 -DHIPBLAS_V2 -DHIP_ENABLE_WARP_SYNC_BUILTINS -fms-extensions -Wno-ignored-attributes /showIncludes /Fofunctorch\CMakeFiles\functorch.dir\csrc\dim\dim.cpp.obj /Fdfunctorch\CMakeFiles\functorch.dir\ -c -- B:\src\torch\functorch\csrc\dim\dim.cpp clang-cl: warning: unknown argument ignored in clang-cl: '-std=c++17' [-Wunknown-argument] clang-cl: warning: argument unused during compilation: '/d2implyavx512upperregs-' [-Wunused-command-line-argument] In file included from B:\src\torch\functorch\csrc\dim\dim.cpp:36: B:\src\torch\functorch\csrc\dim\arena.h(14,21): error: functions that differ only in their return type cannot be overloaded 14 \| inline unsigned int __builtin_clz(unsigned int x) { \| ~~~~~~~~~~~~ ^ C:\home\runner\_work\_tool\Python\3.11.9\x64\Lib\site-packages\_rocm_sdk_devel\lib\llvm\lib\clang\20\include\ia32intrin.h(60,15): note: '__builtin_clz' is a builtin with type 'int (unsigned int) noexcept' 60 \| return 31 - __builtin_clz((unsigned int)__A); \| ^ 1 error generated. [7100/7147] Building CXX object caffe2\torch\CMakeFiles\torch_python.dir\csrc\utils\tensor_list.cpp.obj ``` > [!NOTE] > I haven't been able to reproduce those errors locally, but we have CI jobs that consistently fail when building for Python 3.11 but not 3.12 or 3.13. I'm not sure what is different between those builds, but the code fix seems correct. There are a few other variations on fixes to this floating around, such as: * `a97a957af0/lz4.c (L34-L43)` (checking with `__has_builtin`) * `c98c55ec7e/lj92.c (L31-L46)` (the same code as here, but with `_MSC_VER`) * `2760e5a2bb/def.h (L23-L25)` (using `__lzcnt` instead of a custom implementation) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159273 Approved by: https://github.com/Skylion007, https://github.com/m-gallus	2025-08-01 15:21:26 +00:00
Raphael Reme	ee2649219c	Fix max_width computation in _tensor_str._Formatter (#126859 ) Previous version of `torch._tensor_str._Formatter` was not using `PRINT_OPTS.sci_mode` for the `max_width` computation but was using it for the formatting of values leading to a weird discrepancy. Now, the code first checks if it should be in sci_mode, then compute `max_width` Here is an example to test the behavior: ```python A = torch.tensor([10, 1e-1, 1e-2]) B = torch.tensor([10, 1e-1, 1e-1]) print("================= Default =================") print(A, f"Formatter max_width: {torch._tensor_str._Formatter(A).max_width}") print(B, f"Formatter max_width: {torch._tensor_str._Formatter(B).max_width}") print("================= sci_mode=False =================") with torch._tensor_str.printoptions(sci_mode=False): print(A, f"Formatter max_width: {torch._tensor_str._Formatter(A).max_width}") print(B, f"Formatter max_width: {torch._tensor_str._Formatter(B).max_width}") print("================= sci_mode=True =================") with torch._tensor_str.printoptions(sci_mode=True): print(A, f"Formatter max_width: {torch._tensor_str._Formatter(A).max_width}") print(B, f"Formatter max_width: {torch._tensor_str._Formatter(B).max_width}") ``` In the current version this prints: ``` ================= Default ================= tensor([1.0000e+01, 1.0000e-01, 1.0000e-02]) Formatter max_width: 10 tensor([10.0000, 0.1000, 0.1000]) Formatter max_width: 7 ================= sci_mode=False ================= tensor([ 10.0000, 0.1000, 0.0100]) Formatter max_width: 10 tensor([10.0000, 0.1000, 0.1000]) Formatter max_width: 7 ================= sci_mode=True ================= tensor([1.0000e+01, 1.0000e-01, 1.0000e-02]) Formatter max_width: 10 tensor([1.0000e+01, 1.0000e-01, 1.0000e-01]) Formatter max_width: 7 ``` On can see that in `sci_mode=False`, the values of A are prefixed with unneeded 0 and does not have the same `max_width` as B (It keeps the `max_width` from `sci_mode = None`) Also in `sci_mode = True`, for B, the `max_width` is 7 but each value takes 10 chars... (But it is fine as the code that uses `max_width` do not rely much on it, but still, this is missleading) After this commit, this will print ``` ================= Default ================= tensor([1.0000e+01, 1.0000e-01, 1.0000e-02]) Formatter max_width: 10 tensor([10.0000, 0.1000, 0.1000]) Formatter max_width: 7 ================= sci_mode=False ================= tensor([10.0000, 0.1000, 0.0100]) Formatter max_width: 7 tensor([10.0000, 0.1000, 0.1000]) Formatter max_width: 7 ================= sci_mode=True ================= tensor([1.0000e+01, 1.0000e-01, 1.0000e-02]) Formatter max_width: 10 tensor([1.0000e+01, 1.0000e-01, 1.0000e-01]) Formatter max_width: 10 ``` This also allows to align A with B for `sci_mode=False`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/126859 Approved by: https://github.com/malfet	2025-08-01 15:05:41 +00:00
Howard Huang	b0b3e6e48b	[PP] Refactor test_schedule_multiproc (#158780 ) This refactors the pipelining schedule tests since a lot of them have the same repeated code of: 1. Create pipelined model and reference model 2. Run reference model and pipelined model 3. compare gradients So this refactors those parts above into helper methods and reduces ~300 LOC. Also adds a better gradient check to resolve flakiness (fixes https://github.com/pytorch/pytorch/issues/154408). Pull Request resolved: https://github.com/pytorch/pytorch/pull/158780 Approved by: https://github.com/wconstab	2025-08-01 15:02:18 +00:00
Xilun Wu	3967dbedf4	[ContextParallel][FlexAttention] Prototype of supporting FlexAttention in Context Parallel (#158692 ) Summary This PR adds an all-gather based FlexAttention and uses TorchFunctionMode to dispatch `FlexAttentionHOP.__call__` to it. This PR makes the following changes: - add a user-facing API `create_cp_block_mask` for creating CP-specific `BlockMask` which masks over the attention result of Q shard and KV global. - add `_ContextParallelGlobalVars` to store all necessary global vars that CP FlexAttention requires. `torch_function_mode` is critical to maintain singleton mode to avoid dynamo recompilations. - add a dispatch path for `FlexAttentionForwardHOP.__call__` (TorchFunctionMode dispatch won't work correctly without this line) What's not in this PR: - QKV load balancing - Test on other masking besides `causal_mask`. - Support on small attention (i.e. qkv size is smaller than 128) because the block mask rewrite function requires `Q_BLOCK_SIZE == KV_BLOCK_SIZE == 128`. Test `pytest test/distributed/tensor/test_attention.py -s -k test_ring_flex_attention` Followup 1. create an issue to reproduce the error in `create_fw_bw_graph()` when trying to call `create_block_mask` to re-write `block_mask` in `FlexAttentionHOP` dispatch in `TorchFunctionMode`. 2. Merge `_ContextParallelGlobalVars` and `_cp_options`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158692 Approved by: https://github.com/drisspg	2025-08-01 06:49:01 +00:00
gaoyufeng	4396b15aa7	remove co_lnotab in favor of co_linetable (#159227 ) Fixes #158833 DeprecationWarning: remove co_lnotab in favor of co_linetable Pull Request resolved: https://github.com/pytorch/pytorch/pull/159227 Approved by: https://github.com/ezyang	2025-08-01 06:34:38 +00:00
zpcore	bb6766053b	fix strategy hashing arg mismatch (#159506 ) Reland https://github.com/pytorch/pytorch/pull/159289. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159506 Approved by: https://github.com/XilunWu	2025-08-01 05:42:40 +00:00
tiandeyu-cs	a4fc051c9a	Fix a bug of distributed 'gather' with noncontiguous tensors on the NCCL backend. (#159549 ) Fixes #159548 * Throw an error message when the input tensors for the distributed `gather` are noncontiguous. This behaviour is consistent with the distributed `all_gather`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159549 Approved by: https://github.com/d4l3k	2025-08-01 03:26:06 +00:00
PyTorch MergeBot	5cc6a0abc1	Revert "Refactor CUDAAllocatorConfig to reuse AcceleratorAllocatorConfig (#150312 )" This reverts commit dfacf11f66d6512396382bdf5088f0ba9de00406. Reverted https://github.com/pytorch/pytorch/pull/150312 on behalf of https://github.com/guangyey due to Static initialization order issue impact the downstream repo ([comment](https://github.com/pytorch/pytorch/pull/150312#issuecomment-3142035444))	2025-08-01 03:24:54 +00:00
PyTorch MergeBot	90f13f3b2a	Revert "Deprecate overleap functions in CUDAAllocatorConfig, use AcceleratorAllocatorConfig instead (#156165 )" This reverts commit 1fc010a9d8ea95bb74e54b31d17eba56ef16c27c. Reverted https://github.com/pytorch/pytorch/pull/156165 on behalf of https://github.com/guangyey due to Static initialization order issue impact the downstream repo ([comment](https://github.com/pytorch/pytorch/pull/150312#issuecomment-3142035444))	2025-08-01 03:24:54 +00:00
PyTorch MergeBot	cb9b74872b	Revert "Generalize torch._C._set_allocator_settings to be generic (#156175 )" This reverts commit d3ce45012ed42cd1e13d5048b046b781f0feabe0. Reverted https://github.com/pytorch/pytorch/pull/156175 on behalf of https://github.com/guangyey due to Static initialization order issue impact the downstream repo ([comment](https://github.com/pytorch/pytorch/pull/150312#issuecomment-3142035444))	2025-08-01 03:24:54 +00:00
Catherine Lee	c964204829	[CI] Disable executorch jobs (#159595 ) The current executorch pin needs to be updated The next time the docker image gets rebuilt, the executorch docker build is going to fail like https://github.com/pytorch/pytorch/actions/runs/16626853655/job/47137807966 The failure is that the pin uses a version of the nightly that has been removed from the nightly index ``` #62 72.30 ERROR: Could not find a version that satisfies the requirement torch==2.8.0.dev20250601 (from versions: 1.11.0, 1.12.0, 1.12.1, 1.13.0, 1.13.1, 2.0.0, 2.0.1, 2.1.0, 2.1.1, 2.1.2, 2.2.0, 2.2.1, 2.2.2, 2.3.0, 2.3.1, 2.4.0, 2.4.1, 2.5.0, 2.5.1, 2.6.0, 2.7.0, 2.7.1, 2.8.0.dev20250602+cpu, 2.8.0.dev20250603+cpu, 2.8.0.dev20250604+cpu, 2.8.0.dev20250605+cpu, 2.8.0.dev20250606+cpu, 2.8.0.dev20250607+cpu, 2.8.0.dev20250608+cpu, 2.8.0.dev20250609+cpu, 2.8.0.dev20250610+cpu, 2.8.0.dev20250611+cpu, 2.8.0.dev20250612+cpu, 2.8.0.dev20250613+cpu, 2.8.0.dev20250614+cpu, 2.8.0.dev20250615+cpu, 2.8.0.dev20250616+cpu, 2.8.0.dev20250617+cpu, 2.8.0.dev20250618+cpu, 2.8.0.dev20250619+cpu, 2.8.0.dev20250620+cpu, 2.8.0.dev20250621+cpu, 2.8.0.dev20250622+cpu, 2.8.0.dev20250623+cpu, 2.8.0.dev20250624+cpu, 2.8.0.dev20250625+cpu, 2.8.0.dev20250626+cpu, 2.8.0.dev20250627+cpu, 2.9.0.dev20250628+cpu, 2.9.0.dev20250629+cpu, 2.9.0.dev20250630+cpu, 2.9.0.dev20250701+cpu, 2.9.0.dev20250702+cpu, 2.9.0.dev20250703+cpu, 2.9.0.dev20250704+cpu, 2.9.0.dev20250705+cpu, 2.9.0.dev20250706+cpu, 2.9.0.dev20250707+cpu, 2.9.0.dev20250708+cpu, 2.9.0.dev20250709+cpu, 2.9.0.dev20250710+cpu, 2.9.0.dev20250711+cpu, 2.9.0.dev20250712+cpu, 2.9.0.dev20250713+cpu, 2.9.0.dev20250714+cpu, 2.9.0.dev20250715+cpu, 2.9.0.dev20250716+cpu, 2.9.0.dev20250717+cpu, 2.9.0.dev20250718+cpu, 2.9.0.dev20250719+cpu, 2.9.0.dev20250720+cpu, 2.9.0.dev20250722+cpu, 2.9.0.dev20250723+cpu, 2.9.0.dev20250724+cpu, 2.9.0.dev20250725+cpu, 2.9.0.dev20250726+cpu, 2.9.0.dev20250727+cpu, 2.9.0.dev20250728+cpu, 2.9.0.dev20250729+cpu, 2.9.0.dev20250730+cpu, 2.9.0.dev20250731+cpu) #62 72.30 ERROR: No matching distribution found for torch==2.8.0.dev20250601 ``` The executorch hash update currently fails due to https://github.com/pytorch/pytorch/actions/runs/16636773244/job/47079169392 ``` 2025-07-31T01:56:57.0249165Z + echo 'expecting triton to not be installed, but it is' 2025-07-31T01:56:57.0249614Z expecting triton to not be installed, but it is 2025-07-31T01:56:57.0249969Z + exit 1 2025-07-31T01:58:27.6764352Z ##[error]Final attempt failed. Child_process exited with error code 1 ``` I believe the cause is https://github.com/pytorch/executorch/pull/11653 where the nightly pytorch is installed from our index, but then requirements-examples installs timm from pypi, which reinstalls pytorch, except its the release build for cuda from pypi? Which then causes triton to be installed. I don't know what the intended behavior is so I'm disabling the executorch docker build, executorch build, and the nightly hash update, and apparently the test was already disabled because it was failing Pull Request resolved: https://github.com/pytorch/pytorch/pull/159595 Approved by: https://github.com/malfet	2025-08-01 02:18:03 +00:00
Tugsbayasgalan (Tugsuu) Manlaibaatar	2ac45c2752	Fix autocast context manager when there is exception (#159565 ) Summary: When exception occurs inside context manager, we need to either return False OR properly propagage exceptions via __exit__(exc_type, exc_val). But previously while tracing, we don't actually run the exit node so we end up swallowing the exception in a very weird way as outlined in https://github.com/pytorch/pytorch/issues/153202. This PR fixes it Test Plan: new test case Rollback Plan: Differential Revision: D79348382 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159565 Approved by: https://github.com/zou3519, https://github.com/yushangdi	2025-08-01 02:12:24 +00:00
Xia, Weiwen	83e2ea8135	[CPU] fix _weight_int8pack_mm with large output shape (#158341 ) Summary `_weight_int8pack_mm` on CPU may cause segmentation fault if output shape is large (i.e., M * N is large). It's because the kernel compute output buffer address by ```c++ auto* C_ptr = C_data + mb_start * N + nb_start; ``` where both `mb_start` and `N` are `int` and when they are large their product may overflow. The solution is simple: declare these variables as `int64_t` so that the product won't overflow. Test plan ``` pytest -sv test/test_linalg.py -k test__int8_mm_large_shape ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158341 Approved by: https://github.com/mingfeima, https://github.com/drisspg	2025-08-01 01:55:48 +00:00
Raman Kumar	d994027a41	[Doc fix] fix spelling of enough (#159587 ) fixes typo in word `enought` to correct `enough` at 3 places in these files ``` aten/src/ATen/native/cuda/AdaptiveAveragePooling.cu aten/src/ATen/native/cuda/CuFFTPlanCache.h ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159587 Approved by: https://github.com/ezyang	2025-08-01 01:50:57 +00:00
PyTorch MergeBot	cb4f41e125	Revert "[dynamo] [guard] Add caching for inside torch.compile.disable function to avoid unnecessary recompilation. (#157566 )" This reverts commit 8e07c9870d07c5a318ab21bb16b3fa27576851e6. Reverted https://github.com/pytorch/pytorch/pull/157566 on behalf of https://github.com/yangw-dev due to failed an odd internal test, please reach out to metamate to fix it, D79112610 ([comment](https://github.com/pytorch/pytorch/pull/157566#issuecomment-3141840110))	2025-08-01 01:27:45 +00:00
rzou	690fc9cf88	[merge_rules] add some expected failure and skips (#159581 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159581 Approved by: https://github.com/anijain2305	2025-08-01 01:18:40 +00:00
henrylhtsang	eb853e222b	[cutlass upgrade] Ignore unused-but-set-variable for AsyncMM.cu (#159578 ) Fixes inductor-perf-nightly-h100. This was caused by cutlass upgrade https://github.com/pytorch/pytorch/pull/158854. I missed it in https://github.com/pytorch/pytorch/pull/159276 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159578 Approved by: https://github.com/Skylion007	2025-08-01 00:10:59 +00:00
Sam Larsen	06395276e4	Remove dynamo_timed from the CachingAutotuner.coordinate_descent_tuning() hot path. (#159588 ) Summary: When coordinate_descent_tuning==True, CachingAutotuner.coordinate_descent_tuning() is called for every call of CachingAutotuner.run() (at least for Triton templates), but immediately returns the launcher. Move the dynamo_timed call after the check for triton template so we don't incur the context manager overhead on every call. Fixes https://github.com/pytorch/pytorch/issues/159525 Test Plan: Used the repro in https://github.com/pytorch/pytorch/issues/159525 to make sure the overhead goes away. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159588 Approved by: https://github.com/eellison	2025-07-31 23:33:10 +00:00
Rob Timpe	8becf646ef	[dynamo] Make filter handle None as filter function (#159500 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159500 Approved by: https://github.com/guilhermeleobas, https://github.com/zou3519 ghstack dependencies: #158774, #159102	2025-07-31 23:28:57 +00:00
Rob Timpe	fa68216ca1	[itertools] Implement itertools.cycle with a polyfill (#159102 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159102 Approved by: https://github.com/guilhermeleobas, https://github.com/zou3519 ghstack dependencies: #158774	2025-07-31 23:28:57 +00:00
angelayi	25ef3d315d	[aoti][mps] Dynamic reductions (#159355 ) Dynamic kernel: ```cpp [[max_total_threads_per_threadgroup(1024)]] kernel void generated_kernel( device float* out_ptr0, constant float* in_ptr0, constant long& r0_numel, uint2 thread_pos [[thread_position_in_grid]], uint2 group_pos [[thread_position_in_threadgroup]] ) { auto xindex = thread_pos.x; auto r0_index = thread_pos.y; int x0 = xindex; threadgroup float tmp_acc_0[32]; float tmp_acc_1 = 0; for(auto r0_1_cnt = 0; r0_1_cnt < static_cast<int>(metal::floor(static_cast<float>(0.99902343750000000 + 0.00097656250000000000r0_numel))); ++r0_1_cnt) { int r0_1 = 1024 r0_1_cnt + r0_index; if (r0_1 >= r0_numel) break; auto tmp0 = in_ptr0[x0 + 5r0_1]; tmp_acc_1 += tmp0; } auto tmp1 = c10:🤘:threadgroup_sum(tmp_acc_0, tmp_acc_1, r0_index 1, metal::min(static_cast<decltype(1024+r0_numel)>(1024), static_cast<decltype(1024+r0_numel)>(r0_numel))); if (r0_index == 0) out_ptr0[x0] = static_cast<float>(tmp1); } void AOTInductorModel::run_impl(...) { ... auto arg0_1_size = arg0_1.sizes(); int64_t s77 = arg0_1_size[0]; inputs.clear(); [[maybe_unused]] auto& kernels = static_cast<AOTInductorModelKernels&>(this->kernels_.get()); static constexpr int64_t int_array_0[] = {5LL, }; static constexpr int64_t int_array_1[] = {1LL, }; AtenTensorHandle buf0_handle; AOTI_TORCH_ERROR_CODE_CHECK(aoti_torch_empty_strided(1, int_array_0, int_array_1, cached_torch_dtype_float32, cached_torch_device_type_mps, this->device_idx_, &buf0_handle)); RAIIAtenTensorHandle buf0(buf0_handle); auto mps_lib_0_func = mps_lib_0.getKernelFunction("generated_kernel"); auto mps_lib_0_func_handle = AOTIMetalKernelFunctionHandle(mps_lib_0_func.get()); mps_lib_0_func->runCommandBlock([&] { mps_lib_0_func->startEncoding(); aoti_torch_mps_set_arg_tensor(mps_lib_0_func_handle, 0, buf0); aoti_torch_mps_set_arg_tensor(mps_lib_0_func_handle, 1, arg0_1); aoti_torch_mps_set_arg_int(mps_lib_0_func_handle, 2, s77); mps_lib_0_func->dispatch({static_cast<uint64_t>(5LL), static_cast<uint64_t>(std::min(static_cast<int64_t>(1024LL), static_cast<int64_t>(s77)))}, {static_cast<uint64_t>(1), static_cast<uint64_t>(std::min(static_cast<int64_t>(1024LL), static_cast<int64_t>(s77)))}); }); arg0_1.reset(); output_handles[0] = buf0.release(); } // AOTInductorModel::run_impl ``` Static kernel: ```cpp kernel void generated_kernel( device float out_ptr0, constant float* in_ptr0, uint xindex [[thread_position_in_grid]] ) { int x0 = xindex; auto tmp0 = in_ptr0[x0]; auto tmp1 = in_ptr0[5 + x0]; auto tmp3 = in_ptr0[10 + x0]; auto tmp5 = in_ptr0[15 + x0]; auto tmp2 = tmp0 + tmp1; auto tmp4 = tmp2 + tmp3; auto tmp6 = tmp4 + tmp5; out_ptr0[x0] = static_cast<float>(tmp6); } void AOTInductorModel::run_impl(...) { ... static constexpr int64_t int_array_0[] = {5LL, }; static constexpr int64_t int_array_1[] = {1LL, }; AtenTensorHandle buf0_handle; AOTI_TORCH_ERROR_CODE_CHECK(aoti_torch_empty_strided(1, int_array_0, int_array_1, cached_torch_dtype_float32, cached_torch_device_type_mps, this->device_idx_, &buf0_handle)); RAIIAtenTensorHandle buf0(buf0_handle); auto mps_lib_0_func = mps_lib_0.getKernelFunction("generated_kernel"); auto mps_lib_0_func_handle = AOTIMetalKernelFunctionHandle(mps_lib_0_func.get()); mps_lib_0_func->runCommandBlock([&] { mps_lib_0_func->startEncoding(); aoti_torch_mps_set_arg_tensor(mps_lib_0_func_handle, 0, buf0); aoti_torch_mps_set_arg_tensor(mps_lib_0_func_handle, 1, arg0_1); mps_lib_0_func->dispatch({static_cast<uint64_t>(5LL)}); }); arg0_1.reset(); output_handles[0] = buf0.release(); } // AOTInductorModel::run_impl ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159355 Approved by: https://github.com/malfet	2025-07-31 23:15:02 +00:00
Xu Han	7e00f2ec9d	[AOTI] add zero size consts asm handler (#159225 ) Add `get_zero_consts_asm_code` to handle zero size consts to object. This function is used to handle zero consts situation. Because cpp standard does not allow zero size array: https://stackoverflow.com/questions/9722632/what-happens-if-i-define-a-0-size-array-in-c-c 1. On Windows, MSVC will report error C2466: https://learn.microsoft.com/en-us/cpp/error-messages/compiler-errors-1/compiler-error-c2466?view=msvc-170 So, we can use assmbely compiler to handle this situation. 2. On Windows, why not use Win32 asm to handle all path? Because ml64 only supports up to align `16`, it is not aligned to pytorch's `64`. Reference: https://learn.microsoft.com/en-us/cpp/assembler/masm/ml-and-ml64-command-line-reference?view=msvc-170 ``` Packs structures on the specified byte boundary. The alignment can be 1, 2, 4, 8, or 16. ``` 3. It function can handle zero size case on both Windows and Linux, as that: A. On Linux, we added `-pedantic` to disable zero size array on C++ compiler. `8e07c9870d/torch/_inductor/cpp_builder.py (L580)` B. On Windows, msvc is not support zero size array by default. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159225 Approved by: https://github.com/desertfire	2025-07-31 22:46:33 +00:00
PyTorch MergeBot	490cb3f1a4	Revert "[inductor] Add logging for distributed collective ops for multi‑rank diagnostics (#159190 )" This reverts commit bb62e1f769ef51e2ec149d7256c135d09425aaa0. Reverted https://github.com/pytorch/pytorch/pull/159190 on behalf of https://github.com/clee2000 due to broke [GH job link](https://github.com/pytorch/pytorch/actions/runs/16658705097/job/47150840171) [HUD commit link](`bb62e1f769`) on mac ([comment](https://github.com/pytorch/pytorch/pull/159190#issuecomment-3141513921))	2025-07-31 22:22:13 +00:00
Jane Xu	b95cf5c91d	Move complex to headeronly (#159411 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159411 Approved by: https://github.com/albanD ghstack dependencies: #159415	2025-07-31 22:05:43 +00:00
Jane Xu	5e2ef2a465	Move Float8 variations to headeronly (#159415 ) This PR is a big copy pasta from `c10/util/Float8*` -> `torch/headeronly/util/` which is why we are breaking PR sanity :C (sorry @albanD!). Why is it not a clean copy paste? - For BC reasons, we have to keep the old c10 file around so that OSS devs relying on those files can still get the same APIs - Because we reexpose APIs that are headeronly through torch::headeronly, so there is an extra chunk of code in the new torch::headeronly files to do that. Outside of the copy paste, I: - changed the tests to call torch::headeronly instead of c10 - updated header_only_apis.txt - added `// NOLINTNEXTLINE(bugprone-narrowing-conversions,cppcoreguidelines-narrowing-conversions)` to pass lint (which was previously skipped for -inl.h files) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159415 Approved by: https://github.com/albanD	2025-07-31 22:05:43 +00:00
zpcore	9f753f8c0d	[DTensor] Improve `sort` strategy (#159189 ) - Sort strategy now supports sharding on non sorted dim. ~~- Fix histc xfail.~~ - ~~Previously `python test/distributed/tensor/test_dtensor_ops.py TestDTensorOpsCPU.test_dtensor_op_db_histc_cpu_float32` will fail with `PYTORCH_OPINFO_SAMPLE_INPUT_INDEX=18`. However, if we run `PYTORCH_OPINFO_SAMPLE_INPUT_INDEX=18 python test/distributed/tensor/test_dtensor_ops.py TestDTensorOpsCPU.test_dtensor_op_db_histc_cpu_float32`, the test will pass. This kind of error is due to DTensor reuses the strategy schema hashing. It turns out that not only the strategy, the result correctness also depends on `static_argnum` or the op will reuse the previous args from hashed schema and output wrong results. I updated the document also.~~ (fixed in https://github.com/pytorch/pytorch/pull/159289) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159189 Approved by: https://github.com/XilunWu	2025-07-31 21:52:42 +00:00
Jane Xu	db437690d1	Add myself as a reviewer for when someone touches headeronly or stable (#159583 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159583 Approved by: https://github.com/mikaylagawarecki	2025-07-31 21:30:05 +00:00
Simon Fan	669009bcd1	[inductor] respect layout tags for ops with registered lowerings (#159134 ) scaled_grouped_mm's kernel only supports column-major on the second operand. I -think- this is just for efficiency reasons. But inductor treats that buffer as flexible and may tweak the strides to be row-major instead, as seen in the issue. ~Tagging the op as "needs_fixed_stride_order"/"needs_exact_strides" does not work. Inductor only considers those tags for ops that don't have registered lowering (not sure if this is intended). scaled_grouped_mm does have a lowering, so we never check its tags.~ From discussion below, the op tags are expected to work. FIXES https://github.com/pytorch/pytorch/issues/159097 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159134 Approved by: https://github.com/eellison	2025-07-31 21:29:40 +00:00
Svetlana Karslioglu	e4e2701429	Add the RunLLM widget to the website (#152055 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/152055 Approved by: https://github.com/albanD	2025-07-31 20:53:53 +00:00
Rob Timpe	64cc649275	[itertools] Fix accumulate (#158774 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158774 Approved by: https://github.com/guilhermeleobas, https://github.com/zou3519	2025-07-31 20:32:02 +00:00
PyTorch MergeBot	b1fb552974	Revert "Fix ep deepcopy when there is python builitin name (#159478 )" This reverts commit de7376537f2a11783169fee2b3bc276d266898bf. Reverted https://github.com/pytorch/pytorch/pull/159478 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/159478#issuecomment-3141228423))	2025-07-31 20:20:53 +00:00
Sandeep Narendranath Karjala	bb62e1f769	[inductor] Add logging for distributed collective ops for multi‑rank diagnostics (#159190 ) This change introduces structured logging of the collective communication schedule, enabling downstream tools (e.g. TLParse) to ingest and analyze per‑rank collective‐order information for multi‑rank jobs. - Iterates over scheduler.nodes, filters for _CollectiveKernel nodes - Extracts each op’s python_kernel_name - Emits a structured JSON payload under the inductor_collective_schedule artifact name - Dumps the full schedule list to collective_schedule.json via the PyTorch trace‑structured artifact - Added comprehensive unit tests for collective schedule tracing: Created test_collective_schedule_empty() and test_collective_schedule_real() tests to verify structured trace logging works correctly for both empty collective schedules and real collective operations (like all_reduce and wait_tensor from _c10d_functional ops). Pull Request resolved: https://github.com/pytorch/pytorch/pull/159190 Approved by: https://github.com/yushangdi, https://github.com/xmfan	2025-07-31 19:58:07 +00:00
Dylan Maloy	327e2ca580	[ez] get rid of unused var (#159571 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D79320299 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159571 Approved by: https://github.com/houseroad, https://github.com/georgiaphillips	2025-07-31 19:11:57 +00:00
Neil Tenenholtz	1ebcba4e1b	Fix typo in link to torch memory_viz tool (#159214 ) Fixes a small typo in the torch_cuda_memory docs Pull Request resolved: https://github.com/pytorch/pytorch/pull/159214 Approved by: https://github.com/yewentao256, https://github.com/HDCharles, https://github.com/Skylion007	2025-07-31 18:50:54 +00:00
Divyansh Khanna	5f7eae697d	Deprecate DataLoader pin_memory_device param (#158323 ) Build on top of https://github.com/pytorch/pytorch/pull/146821 - Moves enabling pin_memory back inside `_BaseDataLoaderIter` - This is required for `StatefulDataloader` which leveraged `_BaseDataLoaderIter` directly and not the `Dataloader` class init - Add a simple test for CPU only env where setting `pin_memory=True` is a no-op. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158323 Approved by: https://github.com/ramanishsingh Co-authored-by: zeshengzong <zesheng.zong@outlook.com>	2025-07-31 18:42:07 +00:00
Sherlock Huang	c1722db0f7	[NativeRT] Make VariadicOpConverter and FuseListUnpackConverter for cpu nodes only (#159519 ) Summary: VariadicOpConverter and FuseListUnpackConverter would introduce ops that only have CPU kernels. Currently, the graph passes are ran if static_dispatch is enabled. As we plan to enable static_dispatch by default, this diff add the additional check for the graph pass to only work on the node that has all the inputs/outputs on CPU. Test Plan: CI Rollback Plan: Differential Revision: D79295640 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159519 Approved by: https://github.com/dolpm, https://github.com/henryoier	2025-07-31 18:17:21 +00:00
PyTorch MergeBot	8a233d6000	Revert "[ContextParallel][FlexAttention] Prototype of supporting FlexAttention in Context Parallel (#158692 )" This reverts commit 07fad04181321d18963b71e9566d44f86a25c9f7. Reverted https://github.com/pytorch/pytorch/pull/158692 on behalf of https://github.com/yangw-dev due to failed some internal testapf.metrics.tests.generate_graph_def_test.GenerateGraphDefTest: test_aps_generate_inference_graph_def_with_justknobs1) AssertionError: Expected 'check' to be called once. Called 3 times., please fix the internal test and reland it ([comment](https://github.com/pytorch/pytorch/pull/158692#issuecomment-3140873894))	2025-07-31 18:00:30 +00:00
Aleksandar Samardžić	bf3ebd7ad4	Fix grouped MM load along K when TMA loads are not used (#159485 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159485 Approved by: https://github.com/ngimel	2025-07-31 17:58:02 +00:00
PyTorch MergeBot	c07bb277a0	Revert "fix strategy hashing arg mismatch (#159506 )" This reverts commit 3a556762002ec0027b2120a7e6675182c0e50dbd. Reverted https://github.com/pytorch/pytorch/pull/159506 on behalf of https://github.com/yangw-dev due to failed the internal tests test_get_bwd_hook (torch.equal(output * 2, input_tensor.grad)) ([comment](https://github.com/pytorch/pytorch/pull/159506#issuecomment-3140858905))	2025-07-31 17:54:29 +00:00
Markus Hoehnerbach	f89c28cc6b	[inductor] add lowering for repeat_interleave.Tensor with output size specified (#147160 ) (#158462 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158462 Approved by: https://github.com/eellison	2025-07-31 17:00:32 +00:00
Pian Pawakapan	8fedcfa59a	[export] _ccode for PythonMod (#158851 ) Summary: Adds ccode impl to PythonMod Test Plan: test_export Rollback Plan: Differential Revision: D76463347 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158851 Approved by: https://github.com/kalpit-meta-1	2025-07-31 16:46:51 +00:00
henrylhtsang	6662a76f59	[cutlass backend] Fix EVT tests post buf name change (#159541 ) Differential Revision: [D79317791](https://our.internmc.facebook.com/intern/diff/D79317791/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159541 Approved by: https://github.com/mlazos	2025-07-31 16:39:49 +00:00
eqy	05aade1b6d	[CUDA] Add `serialTest` decorator to `largeTensorTest` in `test_cuda.py` (#159271 ) Hopefully helps with disabled tests due to OOM such as https://github.com/pytorch/pytorch/issues/159069 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159271 Approved by: https://github.com/Skylion007, https://github.com/ngimel	2025-07-31 16:27:16 +00:00
Nikita Shulga	f946b25865	[MPS] Speedup `argmax`/`argmin` (#159524 ) By using efficient `threadgroup_arg[max\|min]` primitives. - Fixed bug in `simd_argmax` when result of the `simd_ballot` were prematurely cast to `ushort` and adjusted unit test - Fixed nan handling in compiled argmax, but can't reliably test it as MPS(eager) implementaiton of argmax is buggy Now according to `bench_mps_ops.py` `max(x, dim=0)` is reliably faster than eager implementaiton: ``` [--------------------------------------------------------------------------------------------- --------------------------------------------------------------------------------------------] \| eager-512x512 \| compile-512x512 \| eager-1024x1024 \| compile-1024x1024 \| eager-2048x2048 \| compile-2048x2048 \| eager-4096x4096 \| compile-4096x4096 1 threads: ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- max (torch.float16) \| 285.8 \| 272.2 \| 422.3 \| 354.5 \| 721.6 \| 683.5 \| 2224.0 \| 1979.1 max (torch.float32) \| 300.2 \| 267.0 \| 389.6 \| 342.5 \| 769.4 \| 682.6 \| 2995.7 \| 2609.8 max (torch.int32) \| 299.6 \| 275.4 \| 390.0 \| 361.7 \| 758.7 \| 686.1 \| 3103.4 \| 2646.5 max (torch.int64) \| 297.5 \| 275.5 \| 417.0 \| 382.1 \| 856.1 \| 722.6 \| 5467.7 \| 3156.8 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159524 Approved by: https://github.com/Skylion007, https://github.com/dcci ghstack dependencies: #158990	2025-07-31 16:18:32 +00:00
Bin Bao	d2e02585b8	[AOTI] Explicitly delete wait_tensor returned tensor (#159502 ) Summary: In the Python wrapper codegen, the returned tensor from wait_tensor is not assigned or used anywhere, because wait_tensor always returns its input, see more discussion in https://github.com/pytorch/pytorch/issues/126773. Similarly, we should just immediately delete the returned tensor handle from aoti_torch_cpu__c10d_functional_wait_tensor in the cpp wrapper codegen, otherwise it may cause tensor's lifetime expansion and even cause OOM in some cases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159502 Approved by: https://github.com/yushangdi, https://github.com/jingsh ghstack dependencies: #159476, #159487	2025-07-31 15:33:36 +00:00
Bin Bao	3dd7ebf418	[BE] Fix buf name mismatch in test_c10d_functional_native.py (#159487 ) Summary: test_c10d_functional_native.py uses hard-coded buf names to check the generated code string. This is fragile given that Inductor can update its buffer naming implementation freely. Thus this PR uses name regex matching to find buffer names at the run time. This will solve issues like https://github.com/pytorch/pytorch/issues/147754. Currently we do name matching based on empty_strided_ calls. We can expand it later if needed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159487 Approved by: https://github.com/yushangdi ghstack dependencies: #159476	2025-07-31 15:33:36 +00:00
Bin Bao	8273ee0646	[BE] Fix global config leak in test_c10d_functional_native.py (#159476 ) Summary: test_c10d_functional_native.py tests torch._inductor.config.cpp_wrapper as True and False. Currently torch._inductor.config.cpp_wrapper is set globally which can cause a problem when running the whole test file. This PR changes it to use patch context. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159476 Approved by: https://github.com/yushangdi	2025-07-31 15:33:36 +00:00
Jane Xu	c57382a493	Move BFloat16.h to headeronly (#159412 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159412 Approved by: https://github.com/desertfire	2025-07-31 15:29:17 +00:00
Ruben Rodriguez Buchillon	e7cc42df58	[inductor] consolidate common GEMM triton param retrieval (#159383 ) \# Why - Make loop iteration simpler - Have a common spot where to make modifications that affect all the GEMM Triton templates, avoiding missed spots \# What - pull out commong logic of taking the BaseConfig objects and turning them into kwargs to feed into maybe_append_choice for Triton GEMM templates Differential Revision: [D79186962](https://our.internmc.facebook.com/intern/diff/D79186962) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159383 Approved by: https://github.com/jansel	2025-07-31 13:05:04 +00:00
cyy	72c69e731f	set MSVC debug information only on debug builds (#159533 ) Fixes: https://github.com/pytorch/pytorch/issues/159515 To reduce the binary size increment in release builds by removing debug information. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159533 Approved by: https://github.com/atalman	2025-07-31 12:57:33 +00:00
Tom Ritchford	78b9dea754	[inductor] Fix set_linter's handling of f-strings for Python 3.12 and up (fix #159056 ) (#159252 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159252 Approved by: https://github.com/Skylion007	2025-07-31 12:56:09 +00:00
LifengWang	838924436e	update the baseline for nightly max_autotune tests (#154973 ) Hi @desertfire, according to the latest test [results](https://github.com/pytorch/pytorch/actions/runs/15385952839) from the inductor nightly for max_autotune tests, we plan to update the baseline data: In the latest nightly test, two models require baseline updates: - vision_maskrcnn: This model shows improved graph breaks, so I’ve updated the baseline accordingly. - detectron2_fcos_r_50_fpn: This model has a different number of graph breaks. However, since its accuracy result still shows fail_accuracy, so I skipped the graph break check for this model. ``` vision_maskrcnn IMPROVED: graph_breaks=29, expected=30 Improvement: 1 models have fixed dynamo graph breaks: vision_maskrcnn ``` ``` detectron2_fcos_r_50_fpn XFAIL detectron2_fcos_r_50_fpn FAIL: graph_breaks=24, expected=22 Error: 1 models have new dynamo graph breaks: detectron2_fcos_r_50_fpn ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154973 Approved by: https://github.com/desertfire	2025-07-31 11:38:55 +00:00
xinan.lin	2ffb510942	[Break XPU][Indutor UT] Fix failures introduced by community. (#159463 ) Fixes #159000, Fixes #159335, Fixes #159334, Fixes #159332, Fixes #159331, Fixes #159330 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159463 Approved by: https://github.com/jansel	2025-07-31 08:37:41 +00:00
Michael Lazos	20b5f694f8	[Dynamo] Make frozen dataclasses hashable (#159529 ) Fixes https://github.com/pytorch/pytorch/issues/159424 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159529 Approved by: https://github.com/oulgen ghstack dependencies: #159513	2025-07-31 07:03:01 +00:00
Michael Lazos	447e300d55	[Dynamo] Frozen dataclass attr access test (#159513 ) Verifies https://github.com/pytorch/pytorch/issues/159424, but perhaps the issue is not fixed yet. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159513 Approved by: https://github.com/oulgen	2025-07-31 07:03:01 +00:00
Pian Pawakapan	5b2ad9279c	[draft export] logging (#159004 ) Summary: adds logging for draft export Test Plan: loggercli stage actualize-stage TorchDraftExportUsageLoggerConfig Rollback Plan: Differential Revision: D78308105 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159004 Approved by: https://github.com/angelayi	2025-07-31 05:52:13 +00:00
Georgia Phillips	78d7f0cdec	disable execution frame cleanup (#159531 ) Summary: Want to disable execution frame cleanup until fix in D78621408 is merged Test Plan: CI Rollback Plan: Differential Revision: D79306602 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159531 Approved by: https://github.com/SherlockNoMad	2025-07-31 05:02:36 +00:00
Xu Han	d5c719ec3c	[inductor] fix open temp file failed on Windows. (#159342 ) Fix open temp file failed on Windows. Error message: <img width="1181" height="239" alt="image" src="https://github.com/user-attachments/assets/e4a6f438-cb06-44c6-959b-0a6a49d2f44f" /> Here two option to fix this issue: https://stackoverflow.com/questions/66744497/python-tempfile-namedtemporaryfile-cant-use-generated-tempfile 1. `tempfile.NamedTemporaryFile` must setup `delete=False` on Windows 2. Use `WritableTempFile` to handle this case on Windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159342 Approved by: https://github.com/jansel	2025-07-31 04:58:02 +00:00
yewentao256	c44efc3755	[Refactor] Fix Compile Warning: `possibly dangling reference to a temporary` (#159517 ) ```bash DEBUG pytorch/torch/csrc/dynamo/compiled_autograd.h:1388:25: warning: possibly dangling reference to a temporary [-Wdangling-reference] DEBUG 1388 \| for (const at::IValue& elt : lst) { DEBUG \| ^~~ DEBUG pytorch/torch/csrc/dynamo/compiled_autograd.h:1388:1: note: the temporary was destroyed at the end of the full expression ‘__for_begin .c10::impl::ListIterator<c10::IValue, __gnu_cxx::__normal_iterator<c10::IValue, std::vector<c10::IValue> > >::operator().c10::impl::ListElementReference<c10::IValue, __gnu_cxx::__normal_iterator<c10::IValue*, std::vector<c10::IValue> > >::operator std::conditional_t<true, const c10::IValue&, c10::IValue>()’ DEBUG 1388 \| for (const at::IValue& elt : lst) { DEBUG \| ^ ``` This PR fixes this warning Pull Request resolved: https://github.com/pytorch/pytorch/pull/159517 Approved by: https://github.com/xmfan	2025-07-31 04:49:43 +00:00
Boyuan Feng	6b9473469f	[Graph Partition] add log for graph partition reasons and #partitions (#159425 ) Previously, we log `skipping cudagraphs due to [xxx reasons]` when there are cudagraph-unsafe ops. With graph partition, we will split off these ops and cudagraph remaining parts. But the log message is also skipped. In this PR, we add logs for graph partition reasons and the number of partitions to better understand the workload. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159425 Approved by: https://github.com/eellison	2025-07-31 04:21:06 +00:00
Natalia Gimelshein	7a4167a164	support fabric handles with symmetric memory (#159319 ) enable fabric handles for symmetric memory Enables handle exchange via CU_MEM_HANDLE_TYPE_FABRIC on the systems that support it. This is needed to enable symmetric memory on NVLS72 systems. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159319 Approved by: https://github.com/malfet, https://github.com/kwen2501	2025-07-31 04:16:20 +00:00
PyTorch UpdateBot	8e67a6ae89	[vllm hash update] update the pinned vllm hash (#159320 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned vllm hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159320 Approved by: https://github.com/pytorchbot	2025-07-31 04:08:14 +00:00
Animesh Jain	c68ad1bd6a	[dynamo][guards] Always record user.stack for informative tlparse guards (#159526 ) Before <img width="1146" height="280" alt="image" src="https://github.com/user-attachments/assets/4ddb11b2-dec8-4010-a28d-63b3cd4a7929" /> After <img width="1248" height="248" alt="image" src="https://github.com/user-attachments/assets/8aafc5be-92cd-4468-bb8f-ad966de8c717" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159526 Approved by: https://github.com/Lucaskabela	2025-07-31 03:18:33 +00:00
PyTorch MergeBot	3e5e094615	Revert "Fix large_tensor_test skipping cpu (#158617 )" This reverts commit debc0591b888f211bfe846bdc7cfa0626a5f6f6a. Reverted https://github.com/pytorch/pytorch/pull/158617 on behalf of https://github.com/ZainRizvi due to Sorry but this seems to be breaking trunk. See [GH job link](https://github.com/pytorch/pytorch/actions/runs/16631113381/job/47062415099) [HUD commit link](`debc0591b8`) ([comment](https://github.com/pytorch/pytorch/pull/158617#issuecomment-3138387762))	2025-07-31 02:57:22 +00:00
clr	c65efc8ea1	torch.compile: Record a pt2_compile_event for combo kernels (#159306 ) This is off by default, but some jobs have it on. Having this show up in perfetto and be globally queryable would be useful to see how expensive this is. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159306 Approved by: https://github.com/masnesral	2025-07-31 02:51:38 +00:00
Animesh Jain	a9049413e2	[dynamo] Turn on recursive dict tag optimization (#159186 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159186 Approved by: https://github.com/jansel	2025-07-31 02:36:37 +00:00
ILCSFNO	d7a5ec9355	Fix the Doc of `padding` in `avg_poolnd` (#159142 ) Fixes #159141 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159142 Approved by: https://github.com/mikaylagawarecki	2025-07-31 02:02:48 +00:00
Sam Larsen	2c46922ce4	Fix rand_like decomposition to preserve strides (#159294 ) Summary: Like https://github.com/pytorch/pytorch/pull/158898, the rand_like variants are not preserving strides. Followed the pattern established in https://github.com/pytorch/pytorch/pull/158898. Test Plan: New unit test (fails before this PR; but fixed after) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159294 Approved by: https://github.com/eellison	2025-07-31 01:36:50 +00:00
wengshiy	668d414ae7	[CPU] Fix bias dtype issue for FP8 qlinear (#159125 ) Fixes `RuntimeError: self and mat2 must have the same dtype, but got BFloat16 and Float` With bf16 autocast, bias converted into BFloat16, but fp8_qlinear_onednn_ref not support bf16 bias. In this pr, convert bias into bf16 on fp8_qlinear_onednn_ref. Add this case into ut and reproduce: `python test/test_quantization.py -k test_qlinear_fp8` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159125 Approved by: https://github.com/Xia-Weiwen, https://github.com/cyyever, https://github.com/CaoE	2025-07-31 01:26:45 +00:00
Nick Riasanovsky	4541509237	[Triton] [Inductor] Fix an incorrect descriptor (#159407 ) Summary: Fixes a clear template typo where `a_desc_ptr` was passed instead of `b_desc_ptr` to define `b_desc`. Test Plan: Found by inspection. Rollback Plan: Reviewed By: NoamPaz Differential Revision: D79178538 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159407 Approved by: https://github.com/NikhilAPatel	2025-07-31 00:34:19 +00:00
eellison	6c7f88c2c9	Check addmm dtypes (#159509 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159509 Approved by: https://github.com/eqy	2025-07-31 00:15:46 +00:00
Chris Thi	c400c8e2e0	[ROCm] Add FP8 rowwise support to _scaled_grouped_mm + Submodule update (#159075 ) Summary: In this PR we integrate the [FBGEMM AMD FP8 rowwise scaling grouped GEMM kernel](https://github.com/pytorch/FBGEMM/tree/main/fbgemm_gpu/experimental/gen_ai/src/quantize/ck_extensions/fp8_rowwise_grouped) to add support for the `_scaled_grouped_mm` API on AMD. `_scaled_grouped_mm` is [currently supported on Nvidia](`9faef3d17c/aten/src/ATen/native/cuda/Blas.cpp (L1614)`), this PR aims to bring parity to AMD. Related: [[RFC]: PyTorch Low-Precision GEMMs Public API](https://github.com/pytorch/pytorch/issues/157950#top) #157950. The kernel is developed using the Composable Kernel framework. Only MI300X is currently supported. In the near future we plan to add support for MI350X as well. For data types we support FP8 e3m4. The kernel support will be gated with the `USE_FBGEMM_GENAI` flag. We hope to enable this by default for relevant AMD builds. Note we also update submodule `third_party/fbgemm` to 0adf62831 for the required updates from fbgemm. Test Plan: Hipify & build ``` python tools/amd_build/build_amd.py USE_FBGEMM_GENAI=1 python setup.py develop ``` Unit tests ``` python test/test_matmul_cuda.py -- TestFP8MatmulCUDA Ran 488 tests in 32.969s OK (skipped=454) ``` Performance Sample \| G \| M \| N \| K \| Runtime Ms \| GB/S \| TFLOPS \| \| -- \| -- \| -- \| -- \| -- \| -- \| -- \| \| 128 \| 1 \| 2048 \| 5120 \| 0.37\| 3590 \| 7.17 \| \| 128 \| 64 \| 2048 \| 5120 \| 0.51\| 2792 \| 338.34 \| \| 128 \| 128 \| 2048 \| 5120 \| 0.66\| 2272 \| 522.72 \| \| 128 \| 1 \| 5120 \| 1024 \| 0.21\| 3224 \| 6.43 \| \| 128 \| 64 \| 5120 \| 1024 \| 0.29\| 2590 \| 291.40 \| \| 128 \| 128 \| 5120 \| 1024 \| 0.40\| 2165 \| 434.76 \| \| 128 \| 1 \| 4096 \| 4096 \| 0.69\| 3126 \| 6.25 \| \| 128 \| 64 \| 4096 \| 4096 \| 0.85\| 2655 \| 324.66 \| \| 128 \| 128 \| 4096 \| 4096 \| 1.10\| 2142 \| 501.40 \| \| 128 \| 1 \| 8192 \| 8192 \| 2.45\| 3508 \| 7.01 \| \| 128 \| 64 \| 8192 \| 8192 \| 3.27\| 2692 \| 336.74 \| \| 128 \| 128 \| 8192 \| 8192 \| 4.04\| 2224 \| 543.76 \| \| 16 \| 1 \| 2048 \| 5120 \| 0.04\| 3928 \| 7.85 \| \| 16 \| 64 \| 2048 \| 5120 \| 0.05\| 3295 \| 399.29 \| \| 16 \| 128 \| 2048 \| 5120 \| 0.07\| 2558 \| 588.69 \| \| 16 \| 1 \| 5120 \| 1024 \| 0.03\| 3119 \| 6.23 \| \| 16 \| 64 \| 5120 \| 1024 \| 0.03\| 2849 \| 320.62 \| \| 16 \| 128 \| 5120 \| 1024 \| 0.05\| 2013 \| 404.11 \| \| 16 \| 1 \| 4096 \| 4096 \| 0.06\| 4512 \| 9.02 \| \| 16 \| 64 \| 4096 \| 4096 \| 0.09\| 3124 \| 381.95 \| \| 16 \| 128 \| 4096 \| 4096 \| 0.13\| 2340 \| 547.67 \| \| 16 \| 1 \| 8192 \| 8192 \| 0.32\| 3374 \| 6.75 \| \| 16 \| 64 \| 8192 \| 8192 \| 0.42\| 2593 \| 324.28 \| \| 16 \| 128 \| 8192 \| 8192 \| 0.53\| 2120 \| 518.36 \| - Using ROCm 6.4.1 - Collected through `triton.testing.do_bench_cudagraph` Binary size with gfx942 arch Before: 116103856 Jul 23 14:12 build/lib/libtorch_hip.so After: 118860960 Jul 23 14:29 build/lib/libtorch_hip.so The difference is 2757104 bytes (~2.6 MiB). Reviewers: @drisspg @ngimel @jwfromm @jeffdaily Pull Request resolved: https://github.com/pytorch/pytorch/pull/159075 Approved by: https://github.com/drisspg	2025-07-30 23:53:58 +00:00
Eddie Yan	25c3a7e317	[CUDA][CUDA Graphs] Move cuda graphs test to subprocess to avoid polluting mempool tests (#159305 ) Otherwise mempool test will fail as the previous graph capture failed but doesn't have its state in the caching allocator fully cleaned up. See also #159301 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159305 Approved by: https://github.com/eellison, https://github.com/BoyuanFeng, https://github.com/naromero77amd	2025-07-30 23:31:38 +00:00
Tugsbayasgalan (Tugsuu) Manlaibaatar	de7376537f	Fix ep deepcopy when there is python builitin name (#159478 ) Summary: title Test Plan: CI Rollback Plan: Differential Revision: D79261007 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159478 Approved by: https://github.com/pianpwk	2025-07-30 23:14:31 +00:00
Shangdi Yu	fd2c64e286	Fix duplicated sources in inductor provenance tracking (#159484 ) Summary: The `replace_hook` is called once for each user of the replaced node. This fix avoids adding duplicated node sources. This also means that if there are two nested pass like: ``` with GraphTransformObserver(gm, "outer"): with GraphTransformObserver(gm, "inner"): ..... ``` We'll only see the outer pass's pass name recorded for the replaced node in the "from_node" node meta. I think this is fine. In practice, the outer pass usually contains a more meaningful name, e.g. `decompose_auto_functionalized`, and the inner pass name is just a default pass name like `pattern_matcher`. Test Plan: ``` buck2 run @mode/dev-nosan fbcode//caffe2/test:fx -- -r test_graph_transform_observer_replace ``` Rollback Plan: Differential Revision: D79203058 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159484 Approved by: https://github.com/angelayi	2025-07-30 23:03:11 +00:00
Lucas Kabela	2b1ae29960	[Dynamo][Better Engineering] Add typing annotations to guard and source (#158397 ) (#159491 ) Summary: X-link: https://github.com/pytorch/executorch/pull/12986 As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to a critical set of files for dynamo, `source.py` and the base `_guards.py` Running ``` mypy torch/_dynamo/source.py torch/_guards.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 1227 \| 2208 \| 55.57% \| 207 \| 362 \| 57.18% \| \| This PR \| 2217 \| 2217 \| 100.00% \| 362 \| 362 \| 100.00% \| \| Delta \| +990 \| +9 \| +44.43% \| +155 \| 0 \| +42.82% \| cc jgong5 mingfeima XiaobingSuper sanchitintel ashokei jingxu10 jerryzh168 voznesenskym penguinwu EikanWang Guobing-Chen zhuhaozhe blzheng wenzhe-nrv jiayisunx ipiszy chenyang78 kadeng muchulee8 amjames chauhang aakhundov coconutruben Test Plan: Imported from GitHub, without a `Test Plan:` line. Rollback Plan: Reviewed By: JacobSzwejbka, yangw-dev Differential Revision: D79199389 Pulled By: Lucaskabela Pull Request resolved: https://github.com/pytorch/pytorch/pull/159491 Approved by: https://github.com/anijain2305, https://github.com/yangw-dev	2025-07-30 22:57:50 +00:00
Nikita Shulga	1293405c8d	[MPS] Add `simd_[arg][max\|min]` (#158990 ) And add eager tests for those. Re-implement `threadgroup_[max\|min]` using those function as they are significantly faster (though much slower than eager, due to the arg part) than before, which could be verified by running the following script ```python import itertools import timeit import torch from torch.utils.benchmark import Compare, Measurement, Timer def bench_unary_op(func, x, label) -> Measurement: sync_cmd = "torch.mps.synchronize()" if "mps" in str(x.device) else "" t = Timer( stmt=f"f(x);{sync_cmd}", globals={"f": func, "x": x}, language="python", timer=timeit.default_timer, sub_label=f"{func.__name__} ({str(x.dtype)})", description=label, env=torch.__version__, ) return t.blocked_autorange() def bench_reduction( reduction_func, device: str = "mps", dtype: torch.dtype = torch.float32 ) -> list[Measurement]: rc = [] # Bench 2D with reduction over dim=0 def f(t): return reduction_func(t, dim=0)[0] f.__name__ = reduction_func.__name__ f_c = torch.compile(f, dynamic=False, fullgraph=True) for size in (512, 1024, 2048, 4096): x = torch.testing.make_tensor(size, size, device=device, dtype=dtype) rc_c, rc_e = f(x), f_c(x) rc_c, rc_e = (rc_c[0], rc_e[0]) if isinstance(rc_c, tuple) else (rc_c, rc_e) rc.append(bench_unary_op(f, x, f"eager-{size}x{size}")) rc.append(bench_unary_op(f_c, x, f"compile-{size}x{size}")) return rc def main() -> None: #dtypes = [torch.float16, torch.float32, torch.bfloat16, torch.int32, torch.int64] dtypes = [torch.float32, torch.int32, torch.int64] # Profile reduction ops rc = [] for op, dtype in itertools.product([torch.max], dtypes): rc.extend(bench_reduction(op, dtype=dtype)) Compare(rc).print() if __name__ == "__main__": torch._dynamo.config.cache_size_limit = 2**16 main() ``` Produces the following table before ``` [--------------------------------------------------------------------------------------------- --------------------------------------------------------------------------------------------] \| eager-512x512 \| compile-512x512 \| eager-1024x1024 \| compile-1024x1024 \| eager-2048x2048 \| compile-2048x2048 \| eager-4096x4096 \| compile-4096x4096 1 threads: ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- max (torch.float32) \| 297.3 \| 531.6 \| 394.1 \| 2550.5 \| 773.0 \| 4904.7 \| 3647.2 \| 9682.0 max (torch.int32) \| 297.8 \| 359.2 \| 387.7 \| 1179.4 \| 768.2 \| 2175.0 \| 3677.1 \| 4495.9 max (torch.int64) \| 278.7 \| 541.4 \| 410.2 \| 2873.3 \| 858.9 \| 5620.4 \| 6107.2 \| 11176.1 Times are in microseconds (us). ``` And after ``` [--------------------------------------------------------------------------------------------- --------------------------------------------------------------------------------------------] \| eager-512x512 \| compile-512x512 \| eager-1024x1024 \| compile-1024x1024 \| eager-2048x2048 \| compile-2048x2048 \| eager-4096x4096 \| compile-4096x4096 1 threads: ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- max (torch.float32) \| 307.9 \| 265.3 \| 401.0 \| 340.8 \| 766.5 \| 661.9 \| 3463.5 \| 2829.5 max (torch.int32) \| 293.5 \| 263.1 \| 405.0 \| 338.8 \| 761.4 \| 672.5 \| 3050.0 \| 2688.6 max (torch.int64) \| 308.2 \| 255.7 \| 417.4 \| 341.4 \| 877.0 \| 695.0 \| 5812.2 \| 5762.2 ``` `argmax`/`argmin` are much tricker due to the nan-handling logic that need to be added there. Also fixes `torch.max/min` compilation for half-precision types, added regression types for it. This PR also introduces a bunch of helper functions, such as `simd_broadcast` that works for int64 and `c10:🤘:pair` template, which are used by `simd_argmax` to return both value and index Pull Request resolved: https://github.com/pytorch/pytorch/pull/158990 Approved by: https://github.com/dcci, https://github.com/Skylion007	2025-07-30 21:57:25 +00:00
William Wen	3a65ff84b6	[dynamo, easy] add comment on skipping sys.monitoring frames (#159493 ) Add a comment so we know why we're doing this code (followup to https://github.com/pytorch/pytorch/pull/159369) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159493 Approved by: https://github.com/azahed98, https://github.com/Lucaskabela, https://github.com/zou3519, https://github.com/jingsh ghstack dependencies: #159369	2025-07-30 21:54:38 +00:00
tiandeyu-cs	acf13a9b75	Fix a bug of distributed 'gather' with uncontiguous tensors on the Gloo backend (#158903 ) Fixes #158902 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158903 Approved by: https://github.com/H-Huang	2025-07-30 21:44:29 +00:00
zpcore	3a55676200	fix strategy hashing arg mismatch (#159506 ) Reland https://github.com/pytorch/pytorch/pull/159289. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159506 Approved by: https://github.com/XilunWu	2025-07-30 21:37:13 +00:00
Sam Larsen	af39144a93	Don't use torch.backends.cuda.matmul.allow_tf32 in inductor cache key (#159480 ) Summary: According to https://github.com/pytorch/pytorch/pull/158209, the API is deprecated and we should be using torch.backends.cuda.matmul.fp32_precision instead. Fixes https://github.com/pytorch/pytorch/issues/159440 Test Plan: CI Pull Request resolved: https://github.com/pytorch/pytorch/pull/159480 Approved by: https://github.com/xmfan, https://github.com/oulgen	2025-07-30 21:29:38 +00:00
Aidyn-A	25343b343e	[ATen][CUDA][cuFFT] Guard against deprecated error codes (#159466 ) This PR adds a guard based on CUDA version, per latest cuFFT [documentation](https://docs.nvidia.com/cuda/cufft/index.html#return-value-cufftresult): >The following error codes are deprecated and will be removed in a future release: `CUFFT_INCOMPLETE_PARAMETER_LIST`, `CUFFT_PARSE_ERROR`, `CUFFT_LICENSE_ERROR`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159466 Approved by: https://github.com/albanD, https://github.com/eqy, https://github.com/Skylion007	2025-07-30 21:10:32 +00:00
Xilun Wu	07fad04181	[ContextParallel][FlexAttention] Prototype of supporting FlexAttention in Context Parallel (#158692 ) Summary This PR adds an all-gather based FlexAttention and uses TorchFunctionMode to dispatch `FlexAttentionHOP.__call__` to it. This PR makes the following changes: - add a user-facing API `create_cp_block_mask` for creating CP-specific `BlockMask` which masks over the attention result of Q shard and KV global. - add `_ContextParallelGlobalVars` to store all necessary global vars that CP FlexAttention requires. `torch_function_mode` is critical to maintain singleton mode to avoid dynamo recompilations. - add a dispatch path for `FlexAttentionForwardHOP.__call__` (TorchFunctionMode dispatch won't work correctly without this line) What's not in this PR: - QKV load balancing - Test on other masking besides `causal_mask`. - Support on small attention (i.e. qkv size is smaller than 128) because the block mask rewrite function requires `Q_BLOCK_SIZE == KV_BLOCK_SIZE == 128`. Test `pytest test/distributed/tensor/test_attention.py -s -k test_ring_flex_attention` Followup 1. create an issue to reproduce the error in `create_fw_bw_graph()` when trying to call `create_block_mask` to re-write `block_mask` in `FlexAttentionHOP` dispatch in `TorchFunctionMode`. 2. Merge `_ContextParallelGlobalVars` and `_cp_options`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158692 Approved by: https://github.com/drisspg	2025-07-30 21:01:53 +00:00
PyTorch MergeBot	7ac70ac4cd	Revert "Fix rand_like decomposition to preserve strides (#159294 )" This reverts commit a3a51282dbabe0220c2c3947a89f7d2ecc514d33. Reverted https://github.com/pytorch/pytorch/pull/159294 on behalf of https://github.com/yangw-dev due to failed internal build Failed to load config ([comment](https://github.com/pytorch/pytorch/pull/159294#issuecomment-3137796767))	2025-07-30 20:59:19 +00:00
drisspg	e221a1c853	[Code Motion]Restructure flex attention kernel into flex subdirectory (#159437 ) Mostly code motion, updating relative paths, moving some imports that had to be lazy before to top level scope now that we are free from the curse. This will make it easier to add newer templates and provide some organization Pull Request resolved: https://github.com/pytorch/pytorch/pull/159437 Approved by: https://github.com/Chillee, https://github.com/BoyuanFeng, https://github.com/eellison, https://github.com/Skylion007	2025-07-30 20:12:35 +00:00
Junjie Wang (PyTorch)	4defea1e2c	[c10d] Fix setGroupName and setGroupDesc in `group_split` and `merge_remote_group` (#159429 ) Summary: We found that we don't really set group_name inside group_split correctly, because we are setting group_name to `deviceTypeToBackend_` which is set after `setBackend`. Same thing as group_desc. I added more unit tests for it. We need to setGroupName correctly, otherwise, this will break DeviceMesh use case when split_group is used in DeviceMesh Also ncclx needs to be aware of that its Option is a subclass of BackendOption Test Plan: CI Rollback Plan: Differential Revision: D79201132 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159429 Approved by: https://github.com/xunnanxu	2025-07-30 19:55:55 +00:00
saienduri	53d68b95de	[ROCm CI] Migrate to MI325 Capacity. (#159059 ) This PR moves PyTorch CI capacity from mi300 to a new, larger mi325 cluster. Both of these GPUs are the same architecture gfx942 and our testing plans don't change within an architecture, so we pool them under the same label `linux.rocm.gpu.gfx942.<#gpus>` with this PR as well to reduce overhead and confusion. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159059 Approved by: https://github.com/jithunnair-amd, https://github.com/atalman Co-authored-by: deedongala <deekshitha.dongala@amd.com>	2025-07-30 19:47:59 +00:00
Howard Huang	f74842d57f	[PP] Fix zero bubble schedules for eval() (#159475 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159475 Approved by: https://github.com/tianyu-l, https://github.com/Skylion007	2025-07-30 19:46:10 +00:00
Simon Fan	644fee2610	Fix TestAutogradFallback flaky tests under Dynamo: migrate to lib._destroy() (#159443 ) under dynamo, the libraries couldn't properly be cleared unless we manually did `gc.collect()`, but that's slow. it also worked if we just used the _destroy() method to tear down FIXES #159398 #159349 #159254 #159237 #159153 #159114 #159040 #158910 #158841 #158763 #158735 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159443 Approved by: https://github.com/zou3519, https://github.com/Skylion007	2025-07-30 19:30:55 +00:00
Jane Xu	7821fbc560	[BE] Clarify comment to not revert when command has been edited (#159495 ) This is mostly a nit. I was a bit confused when I saw <img width="1032" height="183" alt="image" src="https://github.com/user-attachments/assets/7a18f167-78c1-4c33-ba6f-3588914c642e" /> in https://github.com/pytorch/pytorch/pull/159172 So I decided I should clean up this message a bit. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159495 Approved by: https://github.com/yangw-dev, https://github.com/clee2000, https://github.com/ZainRizvi, https://github.com/malfet	2025-07-30 19:23:33 +00:00
Justin Chu	73ee323380	[ONNX] RMS Norm (#159377 ) - Implement rms norm using onnx RMSNormalization-23 - Use the correct eps for float32 `eaadd1282c/aten/src/ATen/native/cuda/layer_norm_kernel.cu (L1844-L1866)` <img width="743" height="107" alt="image" src="https://github.com/user-attachments/assets/a6fd45aa-01d9-4667-924d-3012232cfcde" /> - Created facility to run tests with the reference runtime by extending ONNXProgram and assert_onnx_program. Fix https://github.com/pytorch/pytorch/issues/159257 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159377 Approved by: https://github.com/titaiwangms	2025-07-30 18:55:47 +00:00
Justin Chu	176c6446f8	Update CODEOWNERS for ONNX (#159390 ) Update CODEOWNERS for ONNX to reflect current maintainers. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159390 Approved by: https://github.com/titaiwangms, https://github.com/malfet	2025-07-30 18:54:25 +00:00
drisspg	debc0591b8	Fix large_tensor_test skipping cpu (#158617 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158617 Approved by: https://github.com/BoyuanFeng	2025-07-30 18:48:07 +00:00
Nikita Shulga	0df78f0c11	Remove `/d2implyavx512upperregs-` flag (#159431 ) And reopen https://github.com/pytorch/pytorch/issues/145702 As this flag is not documented anywhere, slows down sccache accelerated build and per https://developercommunity.visualstudio.com/t/Invalid-code-gen-when-using-AVX2-and-SSE/10527298#T-N10562579 it does not workaround a compiler bug, but rather disables some optimizations of AVX512 instructions which are being invoked in AVX2 codepath Fixes https://github.com/pytorch/pytorch/issues/159082 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159431 Approved by: https://github.com/clee2000	2025-07-30 18:47:03 +00:00
Guilherme Leobas	d0e8a0ec4c	Add CPython test for heapq (#159370 ) Not used directly but used internally by `collections.Counter` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159370 Approved by: https://github.com/zou3519, https://github.com/Skylion007	2025-07-30 18:43:06 +00:00
Aaron Gokaslan	22492848b6	[BE]: Update CUTLASS submodule to 4.1.0 (#158854 ) Update the CUTLASS submodule to the latest version with new supported architectures and new features we can use. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158854 Approved by: https://github.com/henrylhtsang	2025-07-30 17:44:38 +00:00
Rohit Manav	5c14315b05	fixed typo error (#159451 ) Fixes #159375 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159451 Approved by: https://github.com/albanD	2025-07-30 17:41:30 +00:00
PaliC	1b99c1859c	[BE] Make PyObjectSlot use a global PyInterpreter and remove (#158427 ) This PR is a bit more involved but effectively works to drastically simplify PyObjectSlot and PyInterpreter. 1) For PyObjectSlot we now use a global pyinterpreter since there only is one. From here we change all of the call sites to rely on this assumption. 2) We also remove the "tags" of the PyInterpreter by deprecating `PyInterpreterStatus`. For the reviewer, sadly it seems like `functorch/csrc/dim/dim.cpp` needed to get linted, so there is an unreadable amount of changes there. Fortunately, the only actual change in the file is as follows which just removes `getPyInterpreter()` from the `check_pyobj` call. ``` mpy::handle handle_from_tensor(Arena& A, TensorRef t) { - // fast case: tensor is live in python - std::optional<PyObject> mb_obj = - t->unsafeGetTensorImpl()->pyobj_slot()->check_pyobj(getPyInterpreter(), /ignore_hermetic_tls=/false); - if (mb_obj.has_value() && !t->unsafeGetTensorImpl()->pyobj_slot()->owns_pyobj()) { - return mb_obj; - } - return A.autorelease(mpy::object::checked_steal(THPVariable_Wrap(t))); -} -} + // fast case: tensor is live in python + std::optional<PyObject> mb_obj = + t->unsafeGetTensorImpl()->pyobj_slot()->check_pyobj( + /ignore_hermetic_tls=/false); + if (mb_obj.has_value() && + !t->unsafeGetTensorImpl()->pyobj_slot()->owns_pyobj()) { + return mb_obj; + } + return A.autorelease(mpy::object::checked_steal(THPVariable_Wrap(t))); +} ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158427 Approved by: https://github.com/albanD	2025-07-30 17:29:43 +00:00
Boyuan Feng	435edbcb5d	[Graph Partition] add graph partition doc (#159450 ) This pr adds doc for graph partition. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159450 Approved by: https://github.com/eellison	2025-07-30 17:01:10 +00:00
PyTorch MergeBot	6c6e11c206	Revert "Fix max_width computation in _tensor_str._Formatter (#126859 )" This reverts commit 1465757959dd7e63715b7621650896eca977aefa. Reverted https://github.com/pytorch/pytorch/pull/126859 on behalf of https://github.com/yangw-dev due to broke trunk with test distributed/test_c10d_functional_native.py::CompileTest::test_inductor_all_reduce_single - RuntimeError: Expected to find buf7 = empty but did not find it ([comment](https://github.com/pytorch/pytorch/pull/126859#issuecomment-3137137030))	2025-07-30 16:56:32 +00:00
Denghui Dong	a775c8e73e	[Profiler] Fix lost C call events problem in Python 3.12.0-3.12.4 (#155446 ) Hi team, Please help review this patch. This PR https://github.com/pytorch/pytorch/pull/150370 tried to fix the "Empty C Call Queue" problem on Python 3.12. It added C calls for each starting Python event with a callable. I found the root cause is not that we cannot get C function frames by `PyFrame_GetBack` when PythonTracer is filling start frames, but the c call event loss problem bug on Python 3.12.0-3.12.4. And that problem was fixed by `257c413cd1` on 3.12.5. So I think the https://github.com/pytorch/pytorch/pull/150370 cannot fix the problem, this patch reverts the change of it. There are solutions to fix the problem correctly, such as we can add a new monitoring callback to compensate call events of methods with C function or we can override the callback registered by `PyEval_SetProfile`. These solutions may make the code hard to maintain. ~~Since upgrading the micro version of Python is not difficult for users, we can just ignore C functions and suggest user upgrade.~~ Pull Request resolved: https://github.com/pytorch/pytorch/pull/155446 Approved by: https://github.com/sraikund16	2025-07-30 16:35:51 +00:00
Arsh Zahed	24d07b3a67	[inductor] Fix mm decomposition evaluating symints (#158998 ) Fixes #154111 Resolves an issue during compilation with dynamic shapes where `torch._inductor.decomposition.mm` evaluates the SymInt expression for the input tensor due to a for loop, and thus the output tensor is not dynamically shaped. This issue is limited to (Mx1)x(1xN) small matrix multiplications, and creates an explicit error with tensor subclasses such as DTensor. The proposed fix replaces the loop with a simple product instead. Benchmark currently running https://hud.pytorch.org/benchmark/compilers Pull Request resolved: https://github.com/pytorch/pytorch/pull/158998 Approved by: https://github.com/jansel, https://github.com/BoyuanFeng	2025-07-30 16:34:15 +00:00
James Wu	90fd06be71	Various bugfixes for running NanoGPT training (#159166 ) Fix various small bugs with running nanogpt on torchbenchmark in OSS under python 3.10. After these changes, the following now succeeds: ``` tlp python benchmarks/dynamo/torchbench.py --only nanogpt --performance --training --backend inductor --caching-precompile --warm-start-latency ``` Cold start: https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmp12LuZ5/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 Warm start (we are invesigating the recompile): https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmpT5YTB2/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159166 Approved by: https://github.com/zhxchen17	2025-07-30 16:30:22 +00:00
Meet Vadakkanchery	002f18807e	[DCP] Improve error handling for process based async checkpointing (#159374 ) Summary: ### PR Context - Kill background process only when PG init fails or there is an explicit `TERMINATE` signal from main process. - When a checkpoint fails to save, log and return the error but continue the serving loop. Test Plan: CI Rollback Plan: Differential Revision: D79177410 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159374 Approved by: https://github.com/sibuachu	2025-07-30 16:25:28 +00:00
Jane Xu	259e79e3ff	Move Half to headeronly (#159172 ) Essence of this copypasta: - combine Half-inl.h and Half.h in c10/util -> torch/headeronly/util/Half.h - Add NOLINTNEXTLINE's to the portions of Half-inl.h that were previously in the ignore list of clangtidy - Re-expose all APIs in namespaces and through includes of the original files. Ideally, we would have the APIs in torch::headeronly and reexpose them in c10, but that runs into BC issues (see D78997465) so for now we are keeping the APIs in c10 but reexposing them in torch::headeronly. - Change test cases in test_aoti_abi_check to test torch::headeronly::Half vs c10::Half (they're the same thing but we eventually want all the tests for headeronly APIs to only import from headeronly). Pull Request resolved: https://github.com/pytorch/pytorch/pull/159172 Approved by: https://github.com/albanD, https://github.com/desertfire	2025-07-30 16:11:58 +00:00
Aidyn-A	ee343ce60c	[RPC][TensorPipe] Fix import torch if compiled without TensorPipe (#159461 ) This is a follow up on the PR #154382, as the issue still persists: ``` File "/opt/pytorch/pytorch/torch/distributed/rpc/__init__.py", line 81, in <module> from . import api, backend_registry, functions File "/opt/pytorch/pytorch/torch/distributed/rpc/api.py", line 35, in <module> from .constants import DEFAULT_SHUTDOWN_TIMEOUT, UNSET_RPC_TIMEOUT File "/opt/pytorch/pytorch/torch/distributed/rpc/constants.py", line 3, in <module> from torch._C._distributed_rpc import ( ImportError: cannot import name '_DEFAULT_NUM_WORKER_THREADS' from 'torch._C._distributed_rpc' (unknown location) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159461 Approved by: https://github.com/lw	2025-07-30 16:04:02 +00:00
Avik Chaudhuri	ea5369113a	unflatten closure (#159418 ) Summary: Sometimes the call history recorded in a `nn_module_stack` does not have the stack property, where each FQN is a prefix of the next FQN. This can cause errors during `unflatten`. Instead of erroring we now drop entries from such a `nn_module_stack` to restore the stack property. This effectively leads to less unflattening: the last FQN in the call history before the stack property was broken keeps the entire flat subgraph of its call. Test Plan: added test, updated another Rollback Plan: Differential Revision: D79204669 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159418 Approved by: https://github.com/angelayi	2025-07-30 15:42:18 +00:00
Jane Xu	b268f22ab2	Move Float4 to headeronly (#159414 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159414 Approved by: https://github.com/desertfire	2025-07-30 15:34:01 +00:00
Animesh Jain	52a52d1b78	[dynamo][guards] Skip no tensor aliasing guard on inbuilt nn module buffers (#159453 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159453 Approved by: https://github.com/jansel	2025-07-30 15:31:07 +00:00
PyTorch MergeBot	eaadd1282c	Revert "Move Half to headeronly (#159172 )" This reverts commit 6d0f4566e2b6e05369d8bb6c0d0e83a0eee982aa. Reverted https://github.com/pytorch/pytorch/pull/159172 on behalf of https://github.com/clee2000 due to broke lint [GH job link](https://github.com/pytorch/pytorch/actions/runs/16613893793/job/47002486679) [HUD commit link](`6d0f4566e2`). Note to self: why isn't Dr. CI updating ([comment](https://github.com/pytorch/pytorch/pull/159172#issuecomment-3136769493))	2025-07-30 15:10:26 +00:00
Raphael Reme	1465757959	Fix max_width computation in _tensor_str._Formatter (#126859 ) Previous version of `torch._tensor_str._Formatter` was not using `PRINT_OPTS.sci_mode` for the `max_width` computation but was using it for the formatting of values leading to a weird discrepancy. Now, the code first checks if it should be in sci_mode, then compute `max_width` Here is an example to test the behavior: ```python A = torch.tensor([10, 1e-1, 1e-2]) B = torch.tensor([10, 1e-1, 1e-1]) print("================= Default =================") print(A, f"Formatter max_width: {torch._tensor_str._Formatter(A).max_width}") print(B, f"Formatter max_width: {torch._tensor_str._Formatter(B).max_width}") print("================= sci_mode=False =================") with torch._tensor_str.printoptions(sci_mode=False): print(A, f"Formatter max_width: {torch._tensor_str._Formatter(A).max_width}") print(B, f"Formatter max_width: {torch._tensor_str._Formatter(B).max_width}") print("================= sci_mode=True =================") with torch._tensor_str.printoptions(sci_mode=True): print(A, f"Formatter max_width: {torch._tensor_str._Formatter(A).max_width}") print(B, f"Formatter max_width: {torch._tensor_str._Formatter(B).max_width}") ``` In the current version this prints: ``` ================= Default ================= tensor([1.0000e+01, 1.0000e-01, 1.0000e-02]) Formatter max_width: 10 tensor([10.0000, 0.1000, 0.1000]) Formatter max_width: 7 ================= sci_mode=False ================= tensor([ 10.0000, 0.1000, 0.0100]) Formatter max_width: 10 tensor([10.0000, 0.1000, 0.1000]) Formatter max_width: 7 ================= sci_mode=True ================= tensor([1.0000e+01, 1.0000e-01, 1.0000e-02]) Formatter max_width: 10 tensor([1.0000e+01, 1.0000e-01, 1.0000e-01]) Formatter max_width: 7 ``` On can see that in `sci_mode=False`, the values of A are prefixed with unneeded 0 and does not have the same `max_width` as B (It keeps the `max_width` from `sci_mode = None`) Also in `sci_mode = True`, for B, the `max_width` is 7 but each value takes 10 chars... (But it is fine as the code that uses `max_width` do not rely much on it, but still, this is missleading) After this commit, this will print ``` ================= Default ================= tensor([1.0000e+01, 1.0000e-01, 1.0000e-02]) Formatter max_width: 10 tensor([10.0000, 0.1000, 0.1000]) Formatter max_width: 7 ================= sci_mode=False ================= tensor([10.0000, 0.1000, 0.0100]) Formatter max_width: 7 tensor([10.0000, 0.1000, 0.1000]) Formatter max_width: 7 ================= sci_mode=True ================= tensor([1.0000e+01, 1.0000e-01, 1.0000e-02]) Formatter max_width: 10 tensor([1.0000e+01, 1.0000e-01, 1.0000e-01]) Formatter max_width: 10 ``` This also allows to align A with B for `sci_mode=False`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/126859 Approved by: https://github.com/malfet	2025-07-30 14:01:00 +00:00
Ke Wen	17b9c618dd	[a2av] not returning out tensor from ops (#159435 ) torch.compile of `all_to_all_vdev_2d` hits the following error: ``` torch._dynamo.exc.BackendCompilerFailed: backend='aot_eager' raised: RuntimeError: Found a custom (non-ATen) operator whose output has alias annotations: symm_mem::all_to_all_vdev_2d(Tensor input, Tensor(a!) out, Tensor in_splits, Tensor(a!) out_splits_offsets, str group_name, int? major_align=None) -> Tensor(a!). We only support functionalizing operators whose outputs do not have alias annotations (e.g. 'Tensor(a)' is a Tensor with an alias annotation whereas 'Tensor' is a Tensor without. The '(a)' is the alias annotation). The alias annotation specifies that the output Tensor shares storage with an input that has the same annotation. Please check if (1) the output needs to be an output (if not, don't return it), (2) if the output doesn't share storage with any inputs, then delete the alias annotation. (3) if the output indeed shares storage with an input, then add a .clone() before returning it to prevent storage sharing and then delete the alias annotation. Otherwise, please file an issue on GitHub. ``` This PR selects option (1). Pull Request resolved: https://github.com/pytorch/pytorch/pull/159435 Approved by: https://github.com/ngimel, https://github.com/xmfan	2025-07-30 08:30:25 +00:00
Yu, Guangye	d3ce45012e	Generalize torch._C._set_allocator_settings to be generic (#156175 ) # Motivation This PR moves the implementation of `torch.cuda.memory._set_allocator_settings` to `torch._C._accelerator_setAllocatorSettings`. Since the original API was intended as a temporary/internal utility, I am not exposing the new function as a public API. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156175 Approved by: https://github.com/albanD ghstack dependencies: #149601, #157908, #150312, #156165	2025-07-30 06:37:15 +00:00
Yu, Guangye	1fc010a9d8	Deprecate overleap functions in CUDAAllocatorConfig, use AcceleratorAllocatorConfig instead (#156165 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156165 Approved by: https://github.com/albanD ghstack dependencies: #149601, #157908, #150312	2025-07-30 06:37:15 +00:00
Yu, Guangye	dfacf11f66	Refactor CUDAAllocatorConfig to reuse AcceleratorAllocatorConfig (#150312 ) # Motivation Refactor `CUDAAllocatorConfig` to reuse `AcceleratorAllocatorConfig` and `ConfigTokenizer`. We would deprecate those option that overleap with `AcceleratorAllocatorConfig` in the following PR and keep them only for BC. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150312 Approved by: https://github.com/albanD ghstack dependencies: #149601, #157908	2025-07-30 06:37:06 +00:00
Yu, Guangye	c8cf811995	Enable AcceleratorAllocatorConfig key check (#157908 ) # Motivation Add a mechanism to ensure raise the key if the key is unrecognized in allocator config. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157908 Approved by: https://github.com/albanD ghstack dependencies: #149601	2025-07-30 06:36:56 +00:00
Yu, Guangye	914b1a3873	Introduce AcceleratorAllocatorConfig as the common class (#149601 ) # Motivation This PR aims to generalize `AllocatorConfig` to be device-agnostic. Introduce the class `AcceleratorAllocatorConfig` to clarify its scope as a configuration manager for accelerator backends (e.g., CUDA, XPU). The another name `AllocatorConfig` is now reserved for a potential future base class that can unify configuration handling for both CPU and accelerator allocators, should similar requirements arise for the CPU path. # Design Rule ## Overall This class configures memory allocation for both device and host memory. A single `AcceleratorAllocatorConfig` instance is shared across all accelerator backends, such as CUDA and XPU, under the assumption that relevant environment variables apply uniformly to all accelerators. Device-specific configuration extensions are supported via hooks (see `registerDeviceConfigParserHook`). Introduce a new class `ConfigTokenizer` to help process the env variable config key-value pair ## Naming Convention: - Public API names in `AcceleratorAllocatorConfig` should be device-generic. - Members prefixed with `pinned_` are specific to the host/pinned allocator. - Environment variable names should be generic across backends. - Comma-separated key-value pairs in the format: `key:value`. Use square brackets `[]` for list values Example: `key1:123, key2:[val1,val2]` ## Environment Variables: - The default environment variable for configuration is `PYTORCH_ALLOC_CONF`. - For backward compatibility, `PYTORCH_CUDA_ALLOC_CONF` and `PYTORCH_HIP_ALLOC_CONF` are also supported with lower priority. Differential Revision: [D79011786](https://our.internmc.facebook.com/intern/diff/D79011786) Pull Request resolved: https://github.com/pytorch/pytorch/pull/149601 Approved by: https://github.com/albanD	2025-07-30 06:36:46 +00:00
Animesh Jain	7eb5fdb358	[dynamo][guards] Recursive dict tag optimization (#159183 ) Design doc here - https://docs.google.com/document/d/1W29DrWID5miGWlZXspsQVN5U0zydE3kjZpziOXrhuaY/edit?tab=t.0#bookmark=id.sba04iw9sp68 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159183 Approved by: https://github.com/jansel	2025-07-30 06:01:32 +00:00
Sheng Fu	f1fb57d854	Add user annotation for FX graph cache key (#159318 ) Summary: AI system co-design team requested to add user annotation for FX graph cache key in PyTorch Kineto trace and Execution trace. With this annotation, they can know the FX graph to which the kernels belong. Test Plan: buck2 run mode/opt caffe2/test:test_profiler_cuda -- profiler.test_execution_trace.TestExecutionTraceCUDA Rollback Plan: Differential Revision: D79019069 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159318 Approved by: https://github.com/sraikund16, https://github.com/jansel	2025-07-30 05:52:50 +00:00
Jane Xu	6d0f4566e2	Move Half to headeronly (#159172 ) Essence of this copypasta: - combine Half-inl.h and Half.h in c10/util -> torch/headeronly/util/Half.h - Add NOLINTNEXTLINE's to the portions of Half-inl.h that were previously in the ignore list of clangtidy - Re-expose all APIs in namespaces and through includes of the original files. Ideally, we would have the APIs in torch::headeronly and reexpose them in c10, but that runs into BC issues (see D78997465) so for now we are keeping the APIs in c10 but reexposing them in torch::headeronly. - Change test cases in test_aoti_abi_check to test torch::headeronly::Half vs c10::Half (they're the same thing but we eventually want all the tests for headeronly APIs to only import from headeronly). Pull Request resolved: https://github.com/pytorch/pytorch/pull/159172 Approved by: https://github.com/albanD, https://github.com/desertfire	2025-07-30 05:02:13 +00:00
PyTorch UpdateBot	e785c087c5	[audio hash update] update the pinned audio hash (#159321 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159321 Approved by: https://github.com/pytorchbot	2025-07-30 04:35:01 +00:00
Svetlana Karslioglu	d214901133	Add a title to distributed._dist2.md (#159385 ) Sphinx likes titles and complains about them when they are not there. So adding a title to address this Wartning in the build: ``` WARNING: toctree contains reference to document 'distributed._dist2' that doesn't have a title: no link will be generated ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159385 Approved by: https://github.com/d4l3k	2025-07-30 04:09:41 +00:00
Jane Xu	96ac64d00c	Migrate easy q(u)int/bits stuff to torch/headeronly (#159302 ) Straightup copy pasta. Keeps APIs in c10 and reexposes them to torch::headeronly. It is arguable that we should just get rid of some of these unused dtypes but that is outside the scope of this PR, which is meant to build up to ScalarType moving to headeronly. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159302 Approved by: https://github.com/malfet, https://github.com/albanD	2025-07-30 03:41:27 +00:00
Colin Peppler	46d34d6766	(should_fold) gso to guard_or_false when checking folding whether to 3d bmm into 2d mm (#159184 ) Switch from guard_size_oblivious to guard_or_false if you encounter a DDE, this would then avoid folding this 3d bmm into a mm. `806d9e3fe7/torch/_decomp/decompositions.py (L4506-L4512)` ## DDE ``` File "/data/users/colinpeppler/pytorch/torch/_decomp/decompositions.py", line 4506, in matmul elif should_fold(tensor1, tensor2, is_out): File "/data/users/colinpeppler/pytorch/torch/_decomp/decompositions.py", line 4472, in should_fold if guard_size_oblivious(t1.numel() == 0): torch.fx.experimental.symbolic_shapes.GuardOnDataDependentSymNode: Could not guard on data-dependent expression Eq(12((u0//2)), 0) (unhinted: Eq(12((u0//2)), 0)). (Size-like symbols: none) Caused by: (_decomp/decompositions.py:4472 in should_fold) ``` ``` File "/data/users/colinpeppler/pytorch/torch/_decomp/decompositions.py", line 4506, in matmul elif should_fold(tensor1, tensor2, is_out): File "/data/users/colinpeppler/pytorch/torch/_decomp/decompositions.py", line 4483, in should_fold return all( torch.fx.experimental.symbolic_shapes.GuardOnDataDependentSymNode: Could not guard on data-dependent expression Eq(3((u0//2)), 3) (unhinted: Eq(3((u0//2)), 3)). (Size-like symbols: none) Caused by: (_decomp/decompositions.py:4483 in should_fold) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159184 Approved by: https://github.com/ezyang ghstack dependencies: #158894	2025-07-30 03:12:14 +00:00
clr	880249adbc	dynamo: handle AttributeErrors from nn_module when infer_paramaters throws. (#158501 ) This only handles AttributeError, but in general, any exception coming from here is a user exception. let me know if we prefer to catch all exceptions, and then reraise them as observed exceptions. ``` File "/packages/aps.ads.gmp/launcher_with_publish#link-tree/torch/_dynamo/symbolic_convert.py", line 2200, in CALL_FUNCTION self.call_function(fn, args, {}) File "/packages/aps.ads.gmp/launcher_with_publish#link-tree/torch/_dynamo/symbolic_convert.py", line 1210, in call_function self.push(fn.call_function(self, args, kwargs)) # type: ignore[arg-type] File "/packages/aps.ads.gmp/launcher_with_publish#link-tree/torch/_dynamo/variables/lazy.py", line 201, in realize_and_forward return getattr(self.realize(), name)(args, kwargs) File "/packages/aps.ads.gmp/launcher_with_publish#link-tree/torch/_dynamo/variables/nn_module.py", line 472, in call_function initialize_lazy_module(tx, mod, args, kwargs) File "/packages/aps.ads.gmp/launcher_with_publish#link-tree/torch/_dynamo/variables/nn_module.py", line 104, in initialize_lazy_module mod._infer_parameters(mod, fake_args, fake_kwargs) File "/packages/aps.ads.gmp/launcher_with_publish#link-tree/torch/nn/modules/lazy.py", line 261, in _infer_parameters module.initialize_parameters(args, *kwargs) ..., File "/packages/aps.ads.gmp/launcher_with_publish#link-tree/torch/nn/modules/module.py", line 1962, in __getattr__ raise AttributeError( torch._dynamo.exc.InternalTorchDynamoError: AttributeError: '...' object has no attribute '...' ``` Note that we crash with a sligthly different exception trace in the other test I added. Let me know if we want this to not throw directly to the end user. ``` ====================================================================== ERROR: test_lazy_module_bad_params (__main__.NNModuleTests.test_lazy_module_bad_params) ---------------------------------------------------------------------- Traceback (most recent call last): File "/data/users/clr/pytorch/torch/testing/_internal/common_utils.py", line 3223, in wrapper method(args, *kwargs) ~~~~~~^^^^^^^^^^^^^^^^^ File "/data/users/clr/pytorch/test/dynamo/test_modules.py", line 1683, in test_lazy_module_bad_params exp_res = opt_m(x, y) File "/data/users/clr/pytorch/torch/_dynamo/eval_frame.py", line 411, in __call__ return super().__call__(args, *kwargs) ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^ File "/data/users/clr/pytorch/torch/nn/modules/module.py", line 1773, in _wrapped_call_impl return self._call_impl(args, *kwargs) ~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^ File "/data/users/clr/pytorch/torch/nn/modules/module.py", line 1784, in _call_impl return forward_call(args, *kwargs) File "/data/users/clr/pytorch/torch/_dynamo/eval_frame.py", line 473, in _call_lazy_check self._orig_mod._infer_parameters(self._orig_mod, args, kwargs) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/data/users/clr/pytorch/torch/nn/modules/lazy.py", line 261, in _infer_parameters module.initialize_parameters(args, **kwargs) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^ File "/data/users/clr/pytorch/test/dynamo/test_modules.py", line 711, in initialize_parameters self.foo += 1 ^^^^^^^^ File "/data/users/clr/pytorch/torch/nn/modules/module.py", line 1962, in __getattr__ raise AttributeError( f"'{type(self).__name__}' object has no attribute '{name}'" ) AttributeError: 'LazyModuleBadInferParams' object has no attribute 'foo' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158501 Approved by: https://github.com/williamwen42, https://github.com/jansel	2025-07-30 02:41:41 +00:00
Xu Han	846ada4973	[AOTI] disable crashed AOTI UTs on Windows. (#159427 ) disable crashed AOTI UTs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159427 Approved by: https://github.com/angelayi	2025-07-30 02:23:27 +00:00
Yu, Guangye	badd0618e4	Remove unused paramter on CUDA AllocParams (#159159 ) # Motivation While refactoring the caching allocator, I noticed that the `AllocParams` constructor on CUDA had an unused parameter. This change removes that unused argument to avoid potential confusion. # Additional Context I noticed that `AllocParams` is defined in cpp file, so it should be safe to make this change. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159159 Approved by: https://github.com/cyyever, https://github.com/albanD	2025-07-30 02:05:25 +00:00
PaliC	a753a72b14	[BE] Modify PyObjectSlot the assume only a single interpreter is in use (#158407 ) This PR makes some less risky changes to PyObjectSlot as there is a lot of stuff we do not need since there is only one interpreter. Specifically `check_interpreter` and `has_pyobj_nonhermetic` are removed Pull Request resolved: https://github.com/pytorch/pytorch/pull/158407 Approved by: https://github.com/albanD ghstack dependencies: #158290, #158291	2025-07-30 01:36:03 +00:00
PaliC	b57d1ef110	[BE] Remove __reduce_deploy__ (#158291 ) This PR removes the integration point torch.fx had with torch::deploy (and another minor change). Note: This PR has some broken mypy errors, but I believe those should have been in the code base beforehand, and should be fixed in a separate PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/158291 Approved by: https://github.com/albanD ghstack dependencies: #158290	2025-07-30 01:36:03 +00:00
PaliC	dd7c996d5c	[BE] Remove torch deploy \| remove torch deploy specific files (#158290 ) This PR removes specific files found in pytorch which are only used for torch::deploy. This is mostly testing code and a debugger. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158290 Approved by: https://github.com/albanD	2025-07-30 01:36:03 +00:00
Kurt Mohler	70d2e9ba45	[MPS] Avoid outputing zeros from `exponential_` for MPS (#159386 ) Fixes #159103 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159386 Approved by: https://github.com/malfet	2025-07-30 00:20:31 +00:00
eqy	62f98dbb44	[CUDA][Convolution] Add `tf32_on_and_off` decorator to `test_deconv_freezing_cuda` (#159280 ) Blackwell seems to select TF32 kernels for this case Pull Request resolved: https://github.com/pytorch/pytorch/pull/159280 Approved by: https://github.com/zou3519, https://github.com/jingsh, https://github.com/Skylion007	2025-07-29 23:44:10 +00:00
PyTorch MergeBot	e288c258f7	Revert "Remove tensorexpr tests (#158928 )" This reverts commit d742a2896c571a535003d5928fe80397325575a5. Reverted https://github.com/pytorch/pytorch/pull/158928 on behalf of https://github.com/yangw-dev due to this breaks bunch of internal dependency since some tests are still using the deleted test files from this pr, the internal reviewer please help fix this using codev ([comment](https://github.com/pytorch/pytorch/pull/158928#issuecomment-3134378616))	2025-07-29 23:32:07 +00:00
William Wen	df58db8831	[dynamo, docs] add recompilation, observability, reporting issues docs (#159062 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159062 Approved by: https://github.com/svekars, https://github.com/zou3519, https://github.com/anijain2305	2025-07-29 23:23:51 +00:00
Nikita Shulga	15bb81ea4f	[2/N][CI] Remove MacOS-13 workarounds from tests (#159304 ) Part of https://github.com/pytorch/pytorch/issues/159275 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159304 Approved by: https://github.com/dcci, https://github.com/cyyever ghstack dependencies: #159277, #159278	2025-07-29 23:12:13 +00:00
saienduri	8d37073bac	[ROCm] Update jit_utils.cpp trait modification based on HIP version. (#159292 ) The mi355 ci regression and hiprtc kernel compilation is failing due to duplicate definitions of traits leading to errors like `error: redefinition of 'integral_constant'`. This seems to be the culprit: https://github.com/pytorch/pytorch/pull/158868. Checking if using hip version instead of rocm version for the check would help with resolution here as rocm version and hip version aren't synced. ROCm 7.0 Alpha build used in CI is still on HIP 6.5. Confirmed that this patch works here: https://github.com/pytorch/pytorch/actions/runs/16579227179?pr=159292 Also, this PR increases the frequency of this MI355 CI to twice a day so we can catch and identify regressions easier if they happen for now. Jeff is on vacation, so Jithun asked me to reach out to y'all. Please help stamp and approve, so we can resolve the recent MI355 CI regression/timeout (https://github.com/pytorch/pytorch/actions/workflows/rocm-mi355.yml) :) @huydhn @malfet @atalman @seemethere Pull Request resolved: https://github.com/pytorch/pytorch/pull/159292 Approved by: https://github.com/malfet	2025-07-29 22:45:27 +00:00
AaronWang04	dc286aef61	Fused RMSNorm Housekeeping (#159317 ) Small PR to address comments that were made from the original fused rmsnorm PR that were not landed Changes: - Warning message when input.dtype doesn't match weight.dtype - Ensure default epsilon value is correct Comments: https://github.com/pytorch/pytorch/pull/153666#discussion_r2114735005 https://github.com/pytorch/pytorch/pull/153666#discussion_r2223518064 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159317 Approved by: https://github.com/ngimel, https://github.com/Skylion007, https://github.com/eqy	2025-07-29 22:39:18 +00:00
Oguz Ulgen	b4619f0272	Pin Helion to 0.0.10 in PyTorch CI (#159420 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159420 Approved by: https://github.com/aorenste, https://github.com/malfet	2025-07-29 22:06:50 +00:00
William Wen	477c2273e1	[dynamo] better way to skip tracing sys.monitoring callables (#159369 ) Better approach to https://github.com/pytorch/pytorch/pull/158171, according to https://github.com/python/cpython/issues/137178#issuecomment-3131617493. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159369 Approved by: https://github.com/Skylion007	2025-07-29 21:54:58 +00:00
Will Constable	2176d481c1	[DTensor] dispatch to sharding prop over decomps (#159324 ) Fixes #159110 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159324 Approved by: https://github.com/ezyang	2025-07-29 21:28:36 +00:00
Guilherme Leobas	b97274e8ac	[iter] Raise TypeError if iter arg cannot be iterable (#158410 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158410 Approved by: https://github.com/XuehaiPan, https://github.com/zou3519 ghstack dependencies: #156371, #156416, #156460	2025-07-29 21:24:21 +00:00
Guilherme Leobas	f9be65cea4	[iter] Wrap iter(..) call in a ObjectIteratorVariable (#156460 ) This object keeps track when the iterator is exhausted (raise Stopiteration). Pull Request resolved: https://github.com/pytorch/pytorch/pull/156460 Approved by: https://github.com/zou3519 ghstack dependencies: #156371, #156416	2025-07-29 21:24:20 +00:00
Guilherme Leobas	4e3e3dc0a7	[iter] support `iter(callable, sentinel)` (#156416 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156416 Approved by: https://github.com/XuehaiPan, https://github.com/zou3519 ghstack dependencies: #156371	2025-07-29 21:24:20 +00:00
Guilherme Leobas	fcf59df2b6	[iter] Add support for sequence protocol in `iter(..)` (#156371 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156371 Approved by: https://github.com/zou3519	2025-07-29 21:24:20 +00:00
Nick Riasanovsky	1bcb2f41e0	[BE] Eliminate workspace info in templates with new API (#159055 ) Summary: Moves the workspace info calculations to the old TMA API. Test Plan: NFC Rollback Plan: Differential Revision: D78904434 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159055 Approved by: https://github.com/NikhilAPatel	2025-07-29 21:22:36 +00:00
Zhengxu Chen	8460131087	[nativert] Add OSS version of ModelRunner (#159268 ) Summary: Implement a ModelRunner from scratch with the minimum features for OSS only Test Plan: test_export -r NativeRT Rollback Plan: Differential Revision: D78979812 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159268 Approved by: https://github.com/dolpm	2025-07-29 21:08:14 +00:00
PyTorch MergeBot	c0c24b61ff	Revert "Partitioner: Fix to align partition node order with original graph (#157892 )" This reverts commit 2d1e92307d3e67622f4fe8058d62e44fe4fa2f4e. Reverted https://github.com/pytorch/pytorch/pull/157892 on behalf of https://github.com/yangw-dev due to fails internal tests : [executorch/backends/xnnpack/partition/xnnpack_partitioner.py:101:24] Incompatible parameter type [6]: In call `Partition.__init__`, for argument `nodes`, expected `Optional[Iterable[Tuple[Node, Optional[int]]]]` but got `dict_keys[Node, str]`. ([comment](https://github.com/pytorch/pytorch/pull/157892#issuecomment-3134004881))	2025-07-29 20:41:45 +00:00
Sahan Paliskara	4fac43b21f	[BE] Move _freeze.py to torch/fb/utils (#159307 ) Summary: We are trying to deprecate torch deploy externally. However a bunch of legacy stuff still uses it. This PR allows the legacy tests to still run if neccessary Test Plan: It's a targets change so CI should suffice Rollback Plan: Differential Revision: D78910653 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159307 Approved by: https://github.com/albanD	2025-07-29 20:07:17 +00:00
Aaron Orenstein	b794e77b7b	Disable cudagraph GCs by default (#158649 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158649 Approved by: https://github.com/eellison ghstack dependencies: #158193	2025-07-29 19:56:11 +00:00
PyTorch MergeBot	d987a6f7f0	Revert "[Dynamo][Better Engineering] Add typing annotations to guard and source (#158397 )" This reverts commit abcb24f4de11f8fedf2c2c9ff53b6092ef42306d. Reverted https://github.com/pytorch/pytorch/pull/158397 on behalf of https://github.com/yangw-dev due to Suggested to fix failing internal signals on D78911890 ([comment](https://github.com/pytorch/pytorch/pull/158397#issuecomment-3133823766))	2025-07-29 19:49:40 +00:00
PyTorch MergeBot	5d93127c87	Revert "[HOP, map] Rework of map autograd to the new interface (#153343 )" This reverts commit 24b1f10ca13d682430725c511812e43a35fcd6a6. Reverted https://github.com/pytorch/pytorch/pull/153343 on behalf of https://github.com/yangw-dev due to a older pr this pr dependes on needed to revert, rebase it after it's in ([comment](https://github.com/pytorch/pytorch/pull/153343#issuecomment-3133816812))	2025-07-29 19:46:42 +00:00
Sam Larsen	a3a51282db	Fix rand_like decomposition to preserve strides (#159294 ) Summary: Like https://github.com/pytorch/pytorch/pull/158898, the rand_like variants are not preserving strides. Followed the pattern established in https://github.com/pytorch/pytorch/pull/158898. Test Plan: New unit test (fails before this PR; but fixed after) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159294 Approved by: https://github.com/eellison	2025-07-29 19:26:20 +00:00
PyTorch MergeBot	e557b3d5e5	Revert "[inductor] Fix mm decomposition evaluating symints (#158998 )" This reverts commit 52e180c3799a7638ee668b1291a711865ab8cfec. Reverted https://github.com/pytorch/pytorch/pull/158998 on behalf of https://github.com/yangw-dev due to it broke trunk with pr_time_benchmark test ([comment](https://github.com/pytorch/pytorch/pull/158998#issuecomment-3133696775))	2025-07-29 19:04:11 +00:00
eellison	f3a9e99036	Fix inductor cuda sort nan behavior (#159308 ) Fix for https://github.com/pytorch/pytorch/issues/152423 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159308 Approved by: https://github.com/isuruf	2025-07-29 19:02:45 +00:00
Animesh Jain	f7d6e9f500	[dynamo][guards] More small guard optimizations (#159345 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159345 Approved by: https://github.com/williamwen42 ghstack dependencies: #159288	2025-07-29 18:36:49 +00:00
Animesh Jain	e43e09e6c1	[dynamo][guards] Use lambda guards for object aliasing to improve object aliasing guards (#159288 ) # Note - On Lambda guarding of object aliasing # We previously installed object‑aliasing guards as relational guards, # but that undermined the recursive‑dict guard optimization: placing the # aliasing guard at a leaf prevented the parent dict node from # qualifying as a recursive‑dict guard root. Because aliasing guards are # rare, we now emit them as epilogue guards via a small Python lambda. # This repeats the access in Python—adding a bit of work—but the # overhead is outweighed by the gains from enabling recursive‑dict guard # optimization. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159288 Approved by: https://github.com/StrongerXi	2025-07-29 18:36:49 +00:00
milliechen	2004f8aa10	FXConverter handling of generic output in inductor fallback kernel (#159002 ) (#159297 ) Summary: A fallback kernel's output may be a non-list/tuple but a `MultiOutput` with empty indices. Allow the `FXConverter` to handle such case. Test Plan: Modified the fxir test for fallbacks, then ran `buck2 test mode/dev-nosan caffe2/test/inductor:fxir_backend -- test_fallback`. Before this diff the modified test would fail with ``` File "/re_cwd/buck-out/v2/gen/fbcode/e2105f7329ead90a/caffe2/test/inductor/__fxir_backend__/fxir_backend#link-tree/torch/_inductor/codegen/wrapper_fxir.py", line 341, in generate line.codegen_fx(self)(line) File "/re_cwd/buck-out/v2/gen/fbcode/e2105f7329ead90a/caffe2/test/inductor/__fxir_backend__/fxir_backend#link-tree/torch/_inductor/codegen/wrapper_fxir.py", line 489, in _generate_multi_output inds = line.indices[0][1:] torch._dynamo.exc.BackendCompilerFailed: backend='inductor' raised: IndexError: list index out of range ``` (Full error paste in P1878839403) With this diff the error is no longer present. Rollback Plan: Differential Revision: [D79126619](https://our.internmc.facebook.com/intern/diff/D79126619) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159297 Approved by: https://github.com/blaine-rister	2025-07-29 18:29:01 +00:00
Edward Z. Yang	31b3b38e3a	Ensure export joint with descriptors + compile works (#159337 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159337 Approved by: https://github.com/wconstab ghstack dependencies: #159336	2025-07-29 17:43:52 +00:00
Edward Z. Yang	2f0db0444e	Track previous MetricsContext edits for ease of debugging. (#159336 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159336 Approved by: https://github.com/wconstab	2025-07-29 17:43:52 +00:00
PaliC	6162e650b0	[BE] remove torch deploy - conditionals (#158288 ) This PR is part of the work to deprecate torch::deploy in OSS. Effectively it does 3 things to get started. 1. Remove test_deploy_interaction as we no longer need to worry about this 2. Remove all torch._running_with_deploy checks and use the False path always (surfaced 1) 3. Remove `USE_DEPLOY` and switch to the default path always Note: MyPy does fail on a bunch of things here as a bunch of older files are touched. It may be better to fix these things on a separate PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/158288 Approved by: https://github.com/albanD	2025-07-29 17:40:49 +00:00
Lucas Kabela	5d89634ca8	Graph break with error message (#158800 ) Fixes #157452 Test with ``` python test/dynamo/test_repros.py ReproTests.test_nn_parameter_ctor_graph_breaks ``` ### Release Notes Change to nn.Parameter Constructor Behavior in Dynamo Semantic change introduced in the nn.Parameter constructor; previously, if the constructor lacked a clean source, the system would attempt to infer arguments to construct a clone and lift this synthetic proxy in the computation graph. This approach had many potential edge cases and was difficult to reason about. The new behavior defaults to graph breaking when the nn.Parameter constructor does not have a clean source. Users are now suggested to manually move the constructor out of the graph in such cases. This change improves clarity and reduces complexity in graph construction and debugging. Users can escape hatch to old semantics with `torch.dynamo.config.graph_break_on_nn_param_ctor=False` if this cannot be done. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158800 Approved by: https://github.com/anijain2305	2025-07-29 17:34:49 +00:00
Arsh Zahed	52e180c379	[inductor] Fix mm decomposition evaluating symints (#158998 ) Fixes #154111 Resolves an issue during compilation with dynamic shapes where `torch._inductor.decomposition.mm` evaluates the SymInt expression for the input tensor due to a for loop, and thus the output tensor is not dynamically shaped. This issue is limited to (Mx1)x(1xN) small matrix multiplications, and creates an explicit error with tensor subclasses such as DTensor. The proposed fix replaces the loop with a simple product instead. Benchmark currently running https://hud.pytorch.org/benchmark/compilers Pull Request resolved: https://github.com/pytorch/pytorch/pull/158998 Approved by: https://github.com/jansel, https://github.com/BoyuanFeng	2025-07-29 17:29:38 +00:00
anwang	c55e72bea1	[Re-land][Inductor] Support native Inductor as backend for MTIA (#159211 ) The previous [diff/PR] (https://github.com/pytorch/pytorch/pull/158526) was reverted due to this docstring lint error: <img width="1736" height="722" alt="image" src="https://github.com/user-attachments/assets/216b1720-4002-48da-b5f3-32b5d48aaa54" /> I didn't add the docstring cause I thought I'm not supposed to add docstring for an EXISTING function. So this diff/PR is an exactly copy of the previous one, except for adding the docstring. ------------- This diff/PR includes the changes to support native Inductor integration for MTIA. The goal is to support `torch.compile(backend="inductor")` for MTIA. Inductor should generate code(triton kernel + python wrapper code) similar to CUDA. And the triton kernels can be launched eagerly. The changes include: - Add MTIA device interfaces used by Dynamo and Inductor, including APIs on device, stream, event, etc. - Add required torch.mtia APIs, like is_bf16_supported, memory_allocated, set_stream_by_id, etc. - MTIA specific codegen logic, for example, loading MTIA dynamic_library. - Other necessary changes to integrate with Inductor codegn, following other devices like CUDA, XPU. - Integrate with the [empty_strided_mtia](https://www.internalfb.com/code/fbsource/[0d017d3a4a1bdff7253f9c66a9f38e77bd62166b]/fbcode/caffe2/aten/src/ATen/native/mtia/EmptyTensor.cpp?lines=49%2C63%2C71%2C74%2C78) API that we’ve added for the new MTIA ATen backend. - A change in Inductor runtime to avoid re-initialize MTIADriver. - BUCK changes to include ATen-mtia in Inductor, and to use -USE_MTIA preprocessor flag. - Update `test_mnist_e2e.py` to cover native Inductor as backend, using the `--use_native_inductor` flag. - Add a personal script(`scripts/anwang/run_native_inductor_script.py`) for testing purpose. Note: - This approach(option 3) aims to provide a pytorch native approach of Inductor integration for MTIA, minimizing the onboarding overhead. The downside of this approach is that it doesn't leverage MTIA specific graph optimization, and is limited to eagerly launch overhead. - MTIA will support another approach(option 2) to provide best performance, based on WrapperFxCodegen. We should be able to reuse the fundamental changes of this diff for option 2, like the device interfaces, steam/event APIs, etc, especially as WrapperFxCodegen inherits PythonWrapperCodegen. Internal: References: - [post for context](https://fb.workplace.com/groups/mtiasw/permalink/1718377262384606/) - [Inductor integration discussion(option 1/2/3)](https://docs.google.com/document/d/1p6363OXtVIRv1hPoaKlRSK3j-iir3QIbDd5bjyqCNig/edit?tab=t.0#heading=h.7s4ns6wcnhmb) - [Project design doc(option 3)](https://docs.google.com/document/d/1jXUmhgoV9WvkMf-bcY3Od_kK9K_RDOdgHdt1LoQ5Tc4/edit?tab=t.0#heading=h.y43gwdqlv46w) - [early prototying diff](https://www.internalfb.com/diff/D75110196) - [MPS integration PR](https://github.com/pytorch/pytorch/pull/153959) - [empty_strided_xpu PR](https://github.com/pytorch/pytorch/pull/126678) Differential Revision: [D79040806](https://our.internmc.facebook.com/intern/diff/D79040806/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159211 Approved by: https://github.com/eellison, https://github.com/blaine-rister, https://github.com/jansel	2025-07-29 17:03:24 +00:00
Sherlock Huang	750348b579	[NativeRT] Clean up use of TargetDevice in KernelFactory (#159298 ) Summary: Remove use of targetDevice in KernelFactory. AOTI would infer device when creating AOTIDelegateExecutor. Test Plan: CI Rollback Plan: Reviewed By: dolpm Differential Revision: D79007317 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159298 Approved by: https://github.com/dolpm	2025-07-29 16:24:33 +00:00
Kurt Mohler	52b9af163c	Add `avg_pool3d` for MPS (#158877 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158877 Approved by: https://github.com/malfet	2025-07-29 15:22:22 +00:00
James Wu	f4bfac11c7	[Precompile] [easy] API For Editable PrecompileCacheArtifacts (#158586 ) This adds an option for backend precompile artifacts to be editable, i.e. to not serialize them right away, but instead be able to apply a Callable edit_fn to them. This allows us to support editing the precompile artifact with more updated autotune results at a later time in the next PR. The goal flow here is: - User runs AOTAutograd -> Inductor -> Triton - User saves to AOTAutogradCache the normal results - User runs autotuning - User calls serialize(), it takes the new autotuning results at runtime and saves only the necessary triton kernels. This PR just implements the API for editing the cache artifacts. The next PR actually adds the autotuning saving support. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158586 Approved by: https://github.com/zhxchen17	2025-07-29 14:53:21 +00:00
Howard Huang	8d00833fdb	[PP] Fix eval step under no_grad() (#159293 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159293 Approved by: https://github.com/tianyu-l, https://github.com/wconstab	2025-07-29 14:42:33 +00:00
Justin Chu	de529ef002	[ONNX] onnx.md to simplify deprecated entities (#159312 ) Simplify documentation of deprecated entities and remove the auto-generated page for JitScalarType Pull Request resolved: https://github.com/pytorch/pytorch/pull/159312 Approved by: https://github.com/titaiwangms	2025-07-29 14:24:17 +00:00
PyTorch MergeBot	61aa2ae20f	Revert "[CPU] fix _weight_int8pack_mm with large output shape (#158341 )" This reverts commit e469414b59ceeaae2860e36708de8852b9892776. Reverted https://github.com/pytorch/pytorch/pull/158341 on behalf of https://github.com/albanD due to Breaks slowtest ([comment](https://github.com/pytorch/pytorch/pull/158341#issuecomment-3132641530))	2025-07-29 13:56:20 +00:00
Mark Harfouche	9d32aa9789	Help fix numpy detection in cross compiled layouts (#137084 ) We had trouble at conda-forge getting numpy to get detected on aarch64 due to our splayed layout and cross compilation needs. see: * https://github.com/conda-forge/pytorch-cpu-feedstock/pull/256 * https://github.com/conda-forge/pytorch-cpu-feedstock/issues/266 * https://github.com/conda-forge/pytorch-cpu-feedstock/pull/267 This is my attempt at making an "upstreamable patch" that tries to follow your structure. It could introduce a new environment variable `Python_NumPy_INCLUDE_DIR` if you want, but CMake doesn't use it as an environment variable, so I feel like that would be weird. Pull Request resolved: https://github.com/pytorch/pytorch/pull/137084 Approved by: https://github.com/atalman	2025-07-29 12:08:56 +00:00
Francisco Massa	5cf77a0ea2	Fix redistribution costs for slice_scatter (#159223 ) We were previously assuming that the `input_strategy == src_strategy`, which is not true in all cases. This should fix this. On the side, I also realized that for `slice_scatter` some DTensorSpecs don't have TensorMeta, e.g., https://github.com/pytorch/pytorch/blob/main/torch/distributed/tensor/_ops/_tensor_ops.py#L524 It would be good to fix it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159223 Approved by: https://github.com/ezyang, https://github.com/wconstab	2025-07-29 12:00:39 +00:00
Xuehai Pan	efcf87654e	[CI] update flake8 and mypy lint dependencies (#158720 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158720 Approved by: https://github.com/Skylion007	2025-07-29 08:05:56 +00:00
Laith Sakka	2523e58781	unbacked handling for view_copy (#159244 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159244 Approved by: https://github.com/bobrenjc93	2025-07-29 07:10:46 +00:00
Jane Xu	222fa451a2	Move some of vec into headeronly in preparation for Half.h (#158976 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158976 Approved by: https://github.com/albanD, https://github.com/desertfire	2025-07-29 05:43:53 +00:00
Jason Ansel	6de24135e5	Fix flaky test_inductor_multiple_specializations (#159264 ) Summary: This test was using do_bench, so it was flaky performance is non-deterministic. Test Plan: buck test 'fbcode//mode/opt' fbcode//caffe2/test/inductor:compile_subprocess -- --exact 'caffe2/test/inductor:compile_subprocess - test_inductor_multiple_specializations_cuda (caffe2.test.inductor.test_compile_subprocess.GPUTests)' --run-disabled Rollback Plan: Differential Revision: D79098692 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159264 Approved by: https://github.com/jingsh	2025-07-29 05:16:55 +00:00
henrylhtsang	27ae72036d	[cutlass] Prep for cutlass upgrade by ignoring Wunused-but-set-variable (#159276 ) Differential Revision: [D79106238](https://our.internmc.facebook.com/intern/diff/D79106238/) This is in prep for cutlass upgrade. More context: https://github.com/NVIDIA/cutlass/issues/2487 Tested in https://github.com/pytorch/pytorch/pull/159115 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159276 Approved by: https://github.com/adamomainz, https://github.com/njriasan, https://github.com/Skylion007	2025-07-29 04:40:24 +00:00
Sherlock Huang	e924df23a6	[NativeRT] Strengthen matcher check for StaticDispatch kernel (#159187 ) Summary: Strength matcher for StaticDispatch kernels: all input, output tensor must be on CPU, all Device-typed attribute must be CPU. Previously, we only check output tensor on CPU. This will miss catching the case where we do DeviceToHost aten._to_copy. Prepare for turning on static dispatch kernel by default. Test Plan: I should add some test before land. Rollback Plan: Differential Revision: D78747600 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159187 Approved by: https://github.com/dolpm	2025-07-29 04:03:49 +00:00
fduwjj	67e68e0785	[c10d] Cleanup split_group logic using the newly built splitGroup (#158488 ) with https://github.com/pytorch/pytorch/pull/157716 merged we want to further clean up the code on the python side for `split_group` API. We do need to keep some old global book keeping for bc. The rest of logic is now all in cpp. Regarding the change brought in https://github.com/pytorch/pytorch/pull/152175, we did clean up in https://github.com/pytorch/pytorch/pull/158790 (including internal changes) so that we can safely remove it. Differential Revision: [D78777152](https://our.internmc.facebook.com/intern/diff/D78777152) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158488 Approved by: https://github.com/d4l3k	2025-07-29 03:27:11 +00:00
Xuehai Pan	775788f93b	[BE][PYFMT] migrate PYFMT for `test/[i-z]*/` to `ruff format` (#144556 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/144556 Approved by: https://github.com/ezyang	2025-07-29 03:26:09 +00:00
Mu-Chu Lee	19ce1beb05	[AOTInductor] Add test for enabling CUDACachingAllocator for AOTInductor's Weight (#159279 ) Summary: Add test for enabling CUDACachingAllocator for AOTInductor's Weight. Implementation TBD Test Plan: N/A, commit is adding a test. Rollback Plan: Differential Revision: D79107507 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159279 Approved by: https://github.com/desertfire, https://github.com/jingsh	2025-07-29 02:52:10 +00:00
Guilherme Leobas	a91ddea61f	Add CPython tests for `collections` module (#158950 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158950 Approved by: https://github.com/zou3519	2025-07-29 02:24:27 +00:00
William Wen	ffccb90ff4	[dynamo, docs] add fullgraph=False docs (#159050 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159050 Approved by: https://github.com/svekars, https://github.com/anijain2305 ghstack dependencies: #157985, #158055, #158531	2025-07-29 01:53:47 +00:00
William Wen	f916f34739	[dynamo, docs] non-strict programming model docs (#158531 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158531 Approved by: https://github.com/AlannaBurke, https://github.com/mlazos, https://github.com/anijain2305 ghstack dependencies: #157985, #158055 Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-07-29 01:53:47 +00:00
William Wen	c32994ce4b	[docs, dynamo] add fullgraph=True, common graph breaks docs (#158055 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158055 Approved by: https://github.com/AlannaBurke, https://github.com/anijain2305 ghstack dependencies: #157985 Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-07-29 01:53:41 +00:00
William Wen	433e43cbec	[dynamo, docs] programming model dynamo core concepts (#157985 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157985 Approved by: https://github.com/svekars, https://github.com/anijain2305	2025-07-29 01:53:34 +00:00
Xia, Weiwen	e469414b59	[CPU] fix _weight_int8pack_mm with large output shape (#158341 ) Summary `_weight_int8pack_mm` on CPU may cause segmentation fault if output shape is large (i.e., M * N is large). It's because the kernel compute output buffer address by ```c++ auto* C_ptr = C_data + mb_start * N + nb_start; ``` where both `mb_start` and `N` are `int` and when they are large their product may overflow. The solution is simple: declare these variables as `int64_t` so that the product won't overflow. Test plan ``` pytest -sv test/test_linalg.py -k test__int8_mm_large_shape ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158341 Approved by: https://github.com/mingfeima, https://github.com/drisspg	2025-07-29 01:14:50 +00:00
rzou	657e5e9aa6	All custom operators go through Inductor's graph.call_function (#159174 ) Fixes #158892 All custom operators should go through the graph.call_function path. The other fallback path is for aten/prim operations that don't have support for things (like torch.float8_e8m0fn). Test Plan: - new tests Pull Request resolved: https://github.com/pytorch/pytorch/pull/159174 Approved by: https://github.com/eellison	2025-07-29 00:31:57 +00:00
Nikita Shulga	f02b783aae	[1/N] Remove MacOS-13 MPS testing (#159278 ) Starts addressing https://github.com/pytorch/pytorch/issues/159275 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159278 Approved by: https://github.com/dcci ghstack dependencies: #159277	2025-07-28 23:52:47 +00:00
Xu Han	8ad96a563c	[inductor] normalize path of the code. (#159255 ) Error stack: <img width="1361" height="345" alt="image" src="https://github.com/user-attachments/assets/50fb2baa-34fd-4a48-a3e7-76e3185391d4" /> After fix: <img width="1103" height="398" alt="image" src="https://github.com/user-attachments/assets/ece5a9ba-a085-46fe-b061-0c2ebda3a2df" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159255 Approved by: https://github.com/desertfire	2025-07-28 23:42:11 +00:00
PyTorch MergeBot	59e261bbd8	Revert "[CI] update flake8 and mypy lint dependencies (#158720 )" This reverts commit f5130bf339f12ccf5c6296130c47685bdc4858e4. Reverted https://github.com/pytorch/pytorch/pull/158720 on behalf of https://github.com/yangw-dev due to this pr failed internally when build torchgen due to rror: fail: Unknown PyPI project: pyyaml, it seems like this is caused by change PyYAML into pyyaml, please fix it ([comment](https://github.com/pytorch/pytorch/pull/158720#issuecomment-3129995414))	2025-07-28 22:02:10 +00:00
Catherine Lee	08ea8fccaf	[ez][docker] Remove some unused vars and scripts (#158680 ) `CUDNN_VERSION` isn't used in any Dockerfiles, it's picked automatically based on the cuda version in `install_cuda.sh` `install_cudnn.sh` isn't used anywhere, cudnn installation happens in `install_cuda.sh` I didn't find any mentions of `GRADLE_VERSION` or `TENSORRT_VERSION` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158680 Approved by: https://github.com/janeyx99, https://github.com/atalman, https://github.com/malfet	2025-07-28 21:44:47 +00:00
atalman	41754539be	Add 3.14 triton wheel build (#159261 ) Related to https://github.com/pytorch/pytorch/issues/156856 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159261 Approved by: https://github.com/malfet, https://github.com/albanD	2025-07-28 20:34:16 +00:00
Nikita Shulga	716d52779f	[BE] Delete non-existing labels (#159277 ) As no such runners has been online for last 2+ month Pull Request resolved: https://github.com/pytorch/pytorch/pull/159277 Approved by: https://github.com/clee2000	2025-07-28 20:28:57 +00:00
Michael Lazos	3bf41f26c8	[cutlass] rename EVT args within kernels for code caching (#159243 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159243 Approved by: https://github.com/henrylhtsang	2025-07-28 19:01:40 +00:00
Eddie Yan	19aa8eb4f5	[TF32][Flex Attention] Turn off TF32 for reference computation in `test_flex_decoding` (#158979 ) Seems to avoid threshold (fudge factor) twiddling games as this causes the checks to go down the "very small ref error" path instead. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158979 Approved by: https://github.com/drisspg, https://github.com/BoyuanFeng, https://github.com/nWEIdia	2025-07-28 18:38:23 +00:00
Animesh Jain	8c0c5c58c7	[benchmarks] Set model name early to keep warmup and main model same (#159231 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159231 Approved by: https://github.com/williamwen42 ghstack dependencies: #159209	2025-07-28 18:18:16 +00:00
Xiaochang Wu	2d1e92307d	Partitioner: Fix to align partition node order with original graph (#157892 ) Fixes #157891 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157892 Approved by: https://github.com/ezyang	2025-07-28 17:36:29 +00:00
Lucca Bertoncini	399c89e15c	fix torch/distributed contributing doc (#158934 ) both pointers are pointing to a page of empty github issues. I'm moving this to point to all issues tagged with `pt_distributed_rampup` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158934 Approved by: https://github.com/d4l3k	2025-07-28 17:01:05 +00:00
PyTorch MergeBot	14d67eec05	Revert "[dynamo][fsdp] Consistent behavior of int attributes (#157262 )" This reverts commit 9b4d938f04c95cebe0fbd96974f64c935567e039. Reverted https://github.com/pytorch/pytorch/pull/157262 on behalf of https://github.com/ZainRizvi due to This was reverted internally. Somehow this PR didn't get reverted alongside it. See D78772867. To validate your fixes internally, you can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/157262#issuecomment-3128148475))	2025-07-28 16:58:27 +00:00
Benson Ma	9ad7dd54f9	[fbgemm_gpu] Upgrade KernelLauncher kernelLaunchCheck to print help string (#158896 ) Summary: - Upgrade KernelLauncher kernelLaunchCheck to print help string, following D78440016 Test Plan: ``` buck test 'fbcode//mode/opt' fbcode//deeplearning/fbgemm/fbgemm_gpu/test/utils:kernel_launcher ``` Rollback Plan: Differential break Revision: D78572009 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158896 Approved by: https://github.com/atalman	2025-07-28 16:11:13 +00:00
Deepak Seshadri	387db86ef1	Name Inductor's Subproc pool threads. (#158815 ) Differential hack Revision: D78710371 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158815 Approved by: https://github.com/d4l3k	2025-07-28 16:08:08 +00:00
dolpm	e5a1d839c5	[nativert] ensure planner once flag is class-local, not static. (#159116 ) Summary: att - otherwise only one global planner will be made even though we need it to be per-model if models are colocated. Differential hack Revision: D78939141 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159116 Approved by: https://github.com/SherlockNoMad	2025-07-28 16:06:21 +00:00
zhxchen17	c06164a9c5	[nativert][ez] Remove unused dist collectives ops. (#159220 ) Removing dependency to c10d/ in ExecutionFrame.h. We don't need c10d::Work in the frame. Differential Revision: [D79041618](https://our.internmc.facebook.com/intern/diff/D79041618/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159220 Approved by: https://github.com/SherlockNoMad, https://github.com/dolpm	2025-07-28 16:03:14 +00:00
Karim Abou Zeid	c7586d4ed3	typo (#156560 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/156560 Approved by: https://github.com/albanD, https://github.com/Skylion007	2025-07-28 15:40:06 +00:00
thenumberouscode	8e07c9870d	[dynamo] [guard] Add caching for inside torch.compile.disable function to avoid unnecessary recompilation. (#157566 ) inside torch.compile.disable function always triggers recompilation. because a user inside function decorated with torch._dynamo.disable would be used as an argument in the resume_in_xx function. In the current implementation, it will always be a new object, resulting in the ID_MATCH guard always failing and triggering recompilation. Fixes https://github.com/pytorch/pytorch/issues/157399 @xmfan Pull Request resolved: https://github.com/pytorch/pytorch/pull/157566 Approved by: https://github.com/mlazos, https://github.com/anijain2305	2025-07-28 12:44:22 +00:00
PyTorch UpdateBot	a76147c9e0	[xla hash update] update the pinned xla hash (#158223 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned xla hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158223 Approved by: https://github.com/pytorchbot	2025-07-28 11:19:05 +00:00
pzzp	f3913ea641	[CUDA] fix nansum in non-JIT build (#158633 ) This change fix crash of ``` import torch a = torch.tensor([[1, 2]], dtype=torch.complex32).to('cuda') b = torch.nansum(a, dim=0) print(b) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158633 Approved by: https://github.com/ngimel	2025-07-28 08:11:32 +00:00
Sherlock Huang	1abff80fae	Reland D78841818 (#159216 ) Summary: Relanding D78841818 with fixes Test Plan: Tested all failing tests buck build --config fbcode.use_link_groups=true --flagfile fbcode//mode/dev-nosan fbcode//sigmoid/core/executor/memory/test:layout_planner_tests buck test 'fbcode//mode/opt' fbcode//sigmoid/inference/test:test_passes Rollback Plan: Reviewed By: hl475 Differential Revision: D79038615 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159216 Approved by: https://github.com/dolpm	2025-07-28 07:39:35 +00:00
zeshengzong	799303f655	Fix atleast_{1,2,3}d() with no arguments description (#156042 ) Fixes #130667 ## Test Result ### Before ![image](https://github.com/user-attachments/assets/7e3a6764-872a-4573-8bec-e7219f920a15) ![image](https://github.com/user-attachments/assets/194be00c-9a29-44cf-b6bc-4d261a12d04e) ![image](https://github.com/user-attachments/assets/21cd6a4f-0793-44e3-9073-7b8b801f997c) ### After ![image](https://github.com/user-attachments/assets/fdbaa2ff-f13c-4fa9-bf52-0810faa698bd) ![image](https://github.com/user-attachments/assets/0374b474-4c6b-4b7d-abea-70e3df0c0a06) ![image](https://github.com/user-attachments/assets/9f9dc188-60e2-4c0f-9e23-36a39310008c) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156042 Approved by: https://github.com/zou3519	2025-07-28 06:25:23 +00:00
PyTorch MergeBot	d26ab281d2	Revert "Setup TorchBench in Docker (#158613 )" This reverts commit d72ebefe3fa7d3ee0e9c9b399f5c07611e790664. Reverted https://github.com/pytorch/pytorch/pull/158613 on behalf of https://github.com/XuehaiPan due to checkout_install_torchbench function is removed but still referenced in trunk ([comment](https://github.com/pytorch/pytorch/pull/158613#issuecomment-3125695250))	2025-07-28 06:19:00 +00:00
PyTorch MergeBot	1cffb217ef	Revert "[Profiler] Fix lost C call events problem in Python 3.12.0-3.12.4 (#155446 )" This reverts commit e88f804a2eecf967dbbf95c5643248352626dafd. Reverted https://github.com/pytorch/pytorch/pull/155446 on behalf of https://github.com/XuehaiPan due to Breaks Windows wheels ([comment](https://github.com/pytorch/pytorch/pull/155446#issuecomment-3125566269))	2025-07-28 05:29:37 +00:00
PyTorch UpdateBot	c8342b7231	[vllm hash update] update the pinned vllm hash (#159235 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned vllm hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159235 Approved by: https://github.com/pytorchbot	2025-07-28 04:16:31 +00:00
Animesh Jain	f63673626d	[dynamo][guards] Skip guards on constant func.__defaults__ elements (#159209 ) Func.__defaults__ is a tuple. Therefore, we can skip guards on immutable elements. Mutable elements are still guarded. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159209 Approved by: https://github.com/jansel	2025-07-27 22:46:17 +00:00
Sampath Victor	37638c303e	Addressing some linter errors (#158670 ) Summary: Addressing the linter errors reported in the changed files. Test Plan: ``` buck test mode/opt deeplearning/fbgemm:QuantUtilsTest ``` https://www.internalfb.com/intern/testinfra/testrun/11821949118528688 ``` buck test mode/opt caffe2/torch/fb/model_transform/splitting/tests:split_dispatcher_test ``` https://www.internalfb.com/intern/testinfra/testrun/7881299627525465 Rollback Plan: Differential Revision: D78352311 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158670 Approved by: https://github.com/excelle08, https://github.com/cyyever, https://github.com/digantdesai	2025-07-27 21:55:50 +00:00
Max Podkorytov	ee2edf3d37	[ROCm][CK][Inductor] enable gfx950 for max autotune with CK (#159195 ) + update inductor config for new gfx arch + fixes in codegen for conv2d and ck-tile matmul + use appropriate fp8 dtypes + test cleanup Pull Request resolved: https://github.com/pytorch/pytorch/pull/159195 Approved by: https://github.com/chenyang78	2025-07-27 20:47:13 +00:00
Chinmay Shrivastava	51eb41a57e	Enable dynamic shapes for foreach operations by default (#158985 ) ## Summary This PR changes the default value of `combo_kernel_foreach_dynamic_shapes` from `False` to `True` in `torch/_inductor/config.py`. ## Context The `combo_kernel_foreach_dynamic_shapes` configuration was introduced in PR #134477 (August 2024) to support dynamic shapes for foreach and combo kernels. It was initially disabled by default as a conservative approach to avoid disrupting production workflows. ## Why This Change? After several months of the feature being available and stable, it's time to enable it by default. This improves the user experience for developers using `torch.compile(dynamic=True)` with foreach operations. ### Current behavior: - Users must manually discover and enable `combo_kernel_foreach_dynamic_shapes` - Without this flag, foreach operations may fail with dynamic shapes - This creates friction and confusion ### With this change: - Foreach operations work seamlessly with dynamic compilation - No manual configuration needed - Better "it just works" experience ## Testing Extensive testing was performed with PyTorch 2.5.0+ and 2.7.1: - ✅ Various tensor sizes (8, 16, 32, 64, 128) - ✅ Multiple tensors in operations (tested up to 20) - ✅ Nested foreach operations - ✅ Mixed operations (foreach + standard operations) - ✅ Both CPU and CUDA devices - ✅ Symbolic shapes with dynamic compilation ## Impact Assessment - Performance: No impact - this only affects compilation behavior - Backward Compatibility: Fully maintained - users can still set to `False` - Risk: Minimal - feature has been stable since August 2024 ## References - Original implementation: PR #134477 by @qchip - This completes the feature rollout by making it available by default Pull Request resolved: https://github.com/pytorch/pytorch/pull/158985 Approved by: https://github.com/jansel, https://github.com/mlazos	2025-07-27 19:56:07 +00:00
Howard Huang	ede6186c86	[PP] Allow intermediate nodes in ZB to have multiple grads (#159084 ) Fixes a ZB regression (https://github.com/pytorch/torchtitan/actions/runs/16478292562/job/46585646792) Previously we only allowed an intermediate node to have 1 gradient. Recently a torchtitan ZB test started failing and I tracked to back to FusedRMSNorm grad_fn having two values `(grad, None)` (see https://github.com/pytorch/pytorch/pull/153666) and it started breaking our ZB tests. This PR allows `stage_backward_weight` intermediate nodes to have multiple grads (it sums them together or if the grad value is None, then ignores it). Here is an example where the backward would have two grad values (gI1, gI2): ```python class Func(torch.autograd.Function): @staticmethod def forward(ctx, x): return x, 2 @staticmethod def backward(ctx, gI1, gI2): assert gI2 is None return gI1 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159084 Approved by: https://github.com/tianyu-l	2025-07-27 19:16:51 +00:00
Nikita Shulga	6d071bd65d	Remove numpy dependency from onnx (#159177 ) One should not expect numpy to be there during onnx import Forward fix for : https://github.com/pytorch/pytorch/pull/157734 Added regression test to `test_without_numpy` function Test plan: Run `python -c "import sys;sys.path.insert(0, 'fake_numpy');import torch; import torch.onnx"` with/without this fix Pull Request resolved: https://github.com/pytorch/pytorch/pull/159177 Approved by: https://github.com/atalman, https://github.com/justinchuby, https://github.com/titaiwangms, https://github.com/cyyever, https://github.com/Skylion007, https://github.com/andrewboldi	2025-07-27 13:23:03 +00:00
cyy	d742a2896c	Remove tensorexpr tests (#158928 ) The tests are not maintained. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158928 Approved by: https://github.com/albanD, https://github.com/malfet	2025-07-27 07:13:27 +00:00
Xu Han	11d6559a58	[inductor] disable failed UTs of test_misc.py (#159210 ) Disable failed UTs. <img width="1195" height="118" alt="image" src="https://github.com/user-attachments/assets/da0933fb-3c4c-44c9-ba85-45971f03405f" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159210 Approved by: https://github.com/jansel Co-authored-by: Jason Ansel <jansel@jansel.net>	2025-07-27 05:41:44 +00:00
PyTorch UpdateBot	e7667e5702	[vllm hash update] update the pinned vllm hash (#159217 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned vllm hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159217 Approved by: https://github.com/pytorchbot	2025-07-27 04:16:35 +00:00
cyy	f6c89c1ef3	Detach tensor before clone in SGD optimiser and other code (#159204 ) Reverse the pattern of tensor clone followed by detach in SGD and other code. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159204 Approved by: https://github.com/Skylion007	2025-07-27 03:31:12 +00:00
Huy Do	d72ebefe3f	Setup TorchBench in Docker (#158613 ) Signed-off-by: Huy Do <huydhn@gmail.com>	2025-07-26 12:56:03 -07:00
Anatoly Myachev	46b925681c	[inductor] Update `to(tl.int8).to(tl.uint8)` workaround from #94717 to handle entire range of `torch.uint8` (#158567 ) https://github.com/pytorch/pytorch/pull/94717/files#r2210265070 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158567 Approved by: https://github.com/ngimel, https://github.com/jansel	2025-07-26 19:11:37 +00:00
PyTorch MergeBot	fe0ff12dab	Revert "[Inductor] Support native Inductor as backend for MTIA (#158526 )" This reverts commit cd68559d0451185f8521912c23e77b83d76b87cf. Reverted https://github.com/pytorch/pytorch/pull/158526 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/158526#issuecomment-3122186057))	2025-07-26 17:58:00 +00:00
Francisco Massa	7dafab6a93	Fix SDPA sharding when `return_debug_mask` is False (#159205 ) If `return_debug_mask` is False (which is the default value for SDPA), the attention tensor returned is an empty tensor (which has 0 dimensions). This means that the shardings for the batch and CP case are that are passed can yield invalid dimensions. This PR fixes it for `scaled_dot_product_flash_attention_strategy`. Note that `scaled_dot_product_cudnn_attention_strategy` doen't have this issue Pull Request resolved: https://github.com/pytorch/pytorch/pull/159205 Approved by: https://github.com/wconstab	2025-07-26 17:41:42 +00:00
Xuehai Pan	f5130bf339	[CI] update flake8 and mypy lint dependencies (#158720 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158720 Approved by: https://github.com/Skylion007	2025-07-26 17:12:29 +00:00
PyTorch MergeBot	f62772f365	Revert "Remove tensorexpr tests (#158928 )" This reverts commit 517eebc1dd4ae6430a95818b16c5f8b4b10fd1bc. Reverted https://github.com/pytorch/pytorch/pull/158928 on behalf of https://github.com/ZainRizvi due to Sorry but this breaks trunk test_jit_fuser_te.py::TestNNCOpInfoCPU::test_nnc_correctness_frac_cpu_bfloat16 [GH job link](https://github.com/pytorch/pytorch/actions/runs/16534544469/job/46768022799) [HUD commit link](`517eebc1dd`) ([comment](https://github.com/pytorch/pytorch/pull/158928#issuecomment-3122158944))	2025-07-26 17:01:54 +00:00
Xu Han	e2b2685f84	[inductor] enable compiled autograd on CPU windows - v2 (#159185 ) The first version: https://github.com/pytorch/pytorch/pull/158432 compiled autograd on windows is disabled in PR #144707 because cuda windows cannot compile this code. However these code can be compiled on CPU. This PR enable these code on CPU windows. But the first version changed ifdef block logical, and caused torch audio build fail: https://github.com/pytorch/audio/issues/3992 Here is the version two, which keep the original logical. # Local test torch audio build pass: <img width="874" height="1043" alt="image" src="https://github.com/user-attachments/assets/9657be86-04f7-4c66-b8c6-802ec2a7c5c8" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159185 Approved by: https://github.com/xmfan	2025-07-26 16:21:28 +00:00
PyTorch MergeBot	3db8623dcb	Revert "[NativeRT] Apply Device placement once when loading the graph (#158996 )" This reverts commit 28ee8be5bfeebb2e44daace6551462b52557e451. Reverted https://github.com/pytorch/pytorch/pull/158996 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/158996#issuecomment-3121540050))	2025-07-26 09:05:26 +00:00
anwang	cd68559d04	[Inductor] Support native Inductor as backend for MTIA (#158526 ) This diff/PR includes the changes to support native Inductor integration for MTIA. The goal is to support `torch.compile(backend="inductor")` for MTIA. Inductor should generate code(triton kernel + python wrapper code) similar to CUDA. And the triton kernels can be launched eagerly. The changes include: - Add MTIA device interfaces used by Dynamo and Inductor, including APIs on device, stream, event, etc. - Add required torch.mtia APIs, like is_bf16_supported, memory_allocated, set_stream_by_id, etc. - MTIA specific codegen logic, for example, loading MTIA dynamic_library. - Other necessary changes to integrate with Inductor codegn, following other devices like CUDA, XPU. - Integrate with the [empty_strided_mtia](https://www.internalfb.com/code/fbsource/[0d017d3a4a1bdff7253f9c66a9f38e77bd62166b]/fbcode/caffe2/aten/src/ATen/native/mtia/EmptyTensor.cpp?lines=49%2C63%2C71%2C74%2C78) API that we’ve added for the new MTIA ATen backend. - A change in Inductor runtime to avoid re-initialize MTIADriver. - BUCK changes to include ATen-mtia in Inductor, and to use -USE_MTIA preprocessor flag. - Update `test_mnist_e2e.py` to cover native Inductor as backend, using the `--use_native_inductor` flag. - Add a personal script(`scripts/anwang/run_native_inductor_script.py`) for testing purpose. Note: - This approach(option 3) aims to provide a pytorch native approach of Inductor integration for MTIA, minimizing the onboarding overhead. The downside of this approach is that it doesn't leverage MTIA specific graph optimization, and is limited to eagerly launch overhead. - MTIA will support another approach(option 2) to provide best performance, based on WrapperFxCodegen. We should be able to reuse the fundamental changes of this diff for option 2, like the device interfaces, steam/event APIs, etc, especially as WrapperFxCodegen inherits PythonWrapperCodegen. Internal: References: - [post for context](https://fb.workplace.com/groups/mtiasw/permalink/1718377262384606/) - [Inductor integration discussion(option 1/2/3)](https://docs.google.com/document/d/1p6363OXtVIRv1hPoaKlRSK3j-iir3QIbDd5bjyqCNig/edit?tab=t.0#heading=h.7s4ns6wcnhmb) - [Project design doc(option 3)](https://docs.google.com/document/d/1jXUmhgoV9WvkMf-bcY3Od_kK9K_RDOdgHdt1LoQ5Tc4/edit?tab=t.0#heading=h.y43gwdqlv46w) - [early prototying diff](https://www.internalfb.com/diff/D75110196) - [MPS integration PR](https://github.com/pytorch/pytorch/pull/153959) - [empty_strided_xpu PR](https://github.com/pytorch/pytorch/pull/126678) Differential Revision: [D78458745](https://our.internmc.facebook.com/intern/diff/D78458745/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158526 Approved by: https://github.com/blaine-rister, https://github.com/jansel, https://github.com/eellison	2025-07-26 08:16:34 +00:00
PyTorch UpdateBot	62a49d929b	[vllm hash update] update the pinned vllm hash (#159198 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned vllm hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159198 Approved by: https://github.com/pytorchbot	2025-07-26 04:44:38 +00:00
Laith Sakka	c6b479bc09	remove guard_or_x from allowlist_for_publicAPI (#159181 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159181 Approved by: https://github.com/albanD	2025-07-26 01:22:17 +00:00
cyy	517eebc1dd	Remove tensorexpr tests (#158928 ) The tests are not maintained. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158928 Approved by: https://github.com/albanD, https://github.com/malfet	2025-07-26 01:21:01 +00:00
zpcore	7f266020de	add softmax_backward_strategy missing field (#159167 ) Add input_specs in softmax_backward_strategy, as is needed by AutoParallel. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159167 Approved by: https://github.com/XilunWu	2025-07-26 00:53:53 +00:00
Meet Patel	e06798191b	Split out C++ code from fused adagrad PR (#159008 ) The original fused Adagrad pull request was: PR#153038 This PR contains only the c++ code of that original PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159008 Approved by: https://github.com/janeyx99	2025-07-26 00:36:59 +00:00
eqy	c89fa88acb	[conv][cuDNN][64-bit indexing] reduce memory usage of depthwise conv 64-bit indexing test (#158981 ) Use half instead for reduced memory usage Pull Request resolved: https://github.com/pytorch/pytorch/pull/158981 Approved by: https://github.com/soulitzer, https://github.com/Skylion007	2025-07-25 23:58:45 +00:00
gaoyufeng	f5cf05c983	Throw invalid_argument instead of RuntimeError when parameters exceed… (#158267 ) Throw invalid_argument instead of RuntimeError when parameters exceed limits (for torch.int32 dtype) Fixes #157707 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158267 Approved by: https://github.com/albanD	2025-07-25 23:49:46 +00:00
NikhilAPatel	21a95bdf7c	[Inductor] [Triton] Enabling TMA for flex-attention for supported device types (#157822 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157822 Approved by: https://github.com/drisspg ghstack dependencies: #159123	2025-07-25 23:45:26 +00:00
Colin Peppler	fb029accb7	(is_non_overlapping_and_dense) gso to guard_or_false in when checking length 1 (#158894 ) Switch from `guard_size_oblivious` to `guard_or_false` if you encounter a DDE, this would then fallback to computing elementwise strides. `2dccff7dcf/torch/_prims/__init__.py (L1919-L1923)` We think it's safe because Laith tested whether this fallback would fail any tests. It did not. https://github.com/pytorch/pytorch/pull/158157 ## Data-dependent exceptions (DDE) ``` File "/data/users/colinpeppler/pytorch/torch/_decomp/decompositions.py", line 2139, in _to_copy x_tensor = torch._prims.convert_element_type(x_tensor, dtype) ... File "/data/users/colinpeppler/pytorch/torch/_prims/__init__.py", line 1920, in _convert_element_type_meta if torch._prims_common.is_non_overlapping_and_dense(a): File "/data/users/colinpeppler/pytorch/torch/_prims_common/__init__.py", line 494, in is_non_overlapping_and_dense if guard_size_oblivious(length == 1): GuardOnDataDependentSymNode: Could not guard on data-dependent expression Eq(u0 - 4, 1) (unhinted: Eq(u0 - 4, 1)). (Size-like symbols: u0) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158894 Approved by: https://github.com/pianpwk, https://github.com/laithsakka	2025-07-25 23:43:38 +00:00
drisspg	26f4dd5160	Scaled MM Fix NVfp4 (#159170 ) Fixes mm on B200: Before: ```Shell def _addmm_nvfp4_dispatch( a: NVFP4Tensor, b: NVFP4Tensor, aten_op, bias: Optional[torch.Tensor] = None ) -> torch.Tensor: """ Core implementation shared between nvfp4_mm, nvfp4_addmm, and nvfp4_linear. The only difference is whether bias is None or not. """ assert a._data.is_contiguous() assert b._data.t().is_contiguous() assert a._block_size == 16, f"NVFP4 requires block_size=16, got {a._block_size}" assert b._block_size == 16, f"NVFP4 requires block_size=16, got {b._block_size}" M, K = a.shape[0], a.shape[1] N = b.shape[1] # Swizzle Dizzle if a._is_swizzled_scales: a_scale_blocked = a._scale_e4m3 # Already swizzled else: a_scale = a._scale_e4m3.view(M, K // a._block_size) a_scale_blocked = to_blocked(a_scale) if b._is_swizzled_scales: b_scale_blocked = b._scale_e4m3 # Already swizzled else: b_scale = b._scale_e4m3.view(N, K // b._block_size) b_scale_blocked = to_blocked(b_scale) # Merge double quant scales into 1 scale for Scale_In^D if a._per_tensor_scale is not None: assert b._per_tensor_scale is not None scale_result = a._per_tensor_scale * b._per_tensor_scale else: assert b._per_tensor_scale is None and a._per_tensor_scale is None scale_result = None # THIS IS A WORKAROUND: # RuntimeError: CUDA error: CUBLAS_STATUS_INVALID_VALUE when calling # When we have per-tensor scaling, we need to apply it before bias # since bias is not quantized should_add_bias_separately = (scale_result is not None) and (bias is not None) # should_add_bias_separately = bias is not None > result = torch._scaled_mm( a._data.view(torch.float4_e2m1fn_x2), b._data.view(torch.float4_e2m1fn_x2), a_scale_blocked.view(torch.float8_e4m3fn), b_scale_blocked.view(torch.float8_e4m3fn), bias=None if should_add_bias_separately else bias, out_dtype=a._orig_dtype, # scale_result=scale_result, # Not supported yet ) E RuntimeError: Invalid scaling configuration. E - For TensorWise scaling, a and b should be float8, scales should be float and singletons. E - For RowWise scaling, a and b should be float8, scales should be float, scale_a should be (200, 1) and scale_b should be (1, 256), and both should be contiguous. E - For BlockWise 1x128 scaling, a and b should be float8, scales should be float, scale_a should be (200, 1) and scale_b should be (1, 256), and both should be outer-dim-major. E - For BlockWise 128x128 scaling, a and b should be float8, scales should be float, scale_a should be (2, 1) and scale_b should be (1, 2), and both should be near-inner-dim-major (with 16-byte aligned strides). E - For Blockwise 1x32 scaling, a and b should be float8, scales should be float8_e8m0fnu, scale_a should have 1024 elements and scale_b should have 1024 elements, and both should be contiguous. E - For Blockwise 1x16 scaling, a and b should be float4 (packed 2x), scales should be float8_e4m3fn, scale_a should have 3072 elements and scale_b should have 3072 elements, and both should be contiguous. E Got a.dtype()=Float4_e2m1fn_x2, scale_a.dtype()=Float8_e4m3fn, scale_a.size()=[256, 12], scale_a.stride()=[12, 1], b.dtype()=Float4_e2m1fn_x2, scale_b.dtype()=Float8_e4m3fn, scale_b.size()=[256, 12] and scale_b.stride()=[12, 1] ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159170 Approved by: https://github.com/ngimel	2025-07-25 23:34:03 +00:00
Menglu Yu	b9e3eb64a7	[Optimus] Support decompose mm with dynamic shapes (#158821 ) Summary: The current implementation will not do the decompose for GEMM with dynamic shapes, thus we add one more option for users to enable this feature Test Plan: ### how to enable Step 1: Set decompose_mem_bound_mm = false Step 2: Add the decompose_mm_pass pattern to the post_grad_fusion_options json config example: "post_grad_fusion_options": { "decompose_mm_pass": { "min_first_dimension_decomposition": 10240, -> default value "max_other_dimention_decomposition": 32, -> default value "skip_dynamic_shape_dim_check": true, -> default is false } }, yaml config example ``` post_grad_fusion_options: decompose_mm_pass: skip_dynamic_shape_dim_check: true ``` Note that all these hyper-parameters can be set by the users, if nothing gives, a default value will be used ### unit test ``` buck2 test @mode/dev-nosan //caffe2/test/inductor:decompose_mem_bound_mm -- test_dynamic_shape_decompose_addmm ``` Buck UI: https://www.internalfb.com/buck2/a98eb4b3-da1d-4450-9e49-472ba98b2267 Test UI: https://www.internalfb.com/intern/testinfra/testrun/6473924745731095 Network: Up: 86KiB Down: 1.3MiB (reSessionID-96cf35cc-5189-4372-8f25-1fc6a52a3963) Executing actions. Remaining 0/3 1.4s exec time total Command: test. Finished 2 local Time elapsed: 2:00.6s Tests finished: Pass 3. Fail 0. Fatal 0. Skip 0. Build failure 0 ### E2E before: aps-DPA_new_v0_amd_20250716-e7927755df after: aps-DPA_new_v0_amd_20250716_optimus-f2175fc9fb tlparse: https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/aps-DPA_new_v0_amd_20250716_optimus-f2175fc9fb/attempt_0/version_0/rank_0/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 ### qps and NE {F1980635506} {F1980635505} - 12.5% qps improvement with NE neutral ### trace analysis baseline:https://www.internalfb.com/intern/perfdoctor/trace_view?filepath=tree%2Ftraces%2Fdynocli%2Faps-DPA_new_v0_amd_20250716-e7927755df%2F0%2Frank-1.Jul_22_22_28_01.4592.pt.trace.json.gz&bucket=aps_traces {F1980633952} proposal:https://www.internalfb.com/intern/perfdoctor/trace_view?filepath=tree%2Ftraces%2Fdynocli%2Faps-DPA_new_v0_amd_20250716_optimus-f2175fc9fb%2F0%2Frank-1.Jul_24_14_37_59.4576.pt.trace.json.gz&bucket=aps_traces {F1980633966} ``` unsqueeze_default: "bf16[32s54, 8, 1][8, 1, 1]cuda:0" = torch.ops.aten.unsqueeze.default(constant_pad_nd_default_2, 2) unsqueeze_default_1: "bf16[1, 8, 8][64, 8, 1]cuda:0" = torch.ops.aten.unsqueeze.default(constant_pad_nd_default_3, 0); constant_pad_nd_default_3 = None mul_tensor: "bf16[32s54, 8, 8][64, 8, 1]cuda:0" = torch.ops.aten.mul.Tensor(unsqueeze_default, unsqueeze_default_1); unsqueeze_default = unsqueeze_default_1 = None ``` ### what have been decomposed P1880443593 Rollback Plan: Differential Revision: D78716034 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158821 Approved by: https://github.com/Yuzhen11	2025-07-25 23:19:53 +00:00
PyDevC	69cc99525c	[nn]: updated type alias for padddingmode in module/conv.py (#158843 ) Fixes #152280 Changed type of `padding_mode` from `str` to `Literal["zeros", "reflect", "replicate", "circular"]` cc @Skylion007 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158843 Approved by: https://github.com/mikaylagawarecki	2025-07-25 23:05:02 +00:00
Edward Z. Yang	72af19dadf	Add aot_autograd.fx_utils (#159005 ) See docblock for details. The API here has been validated by use in autoparallel but I'm always open to suggestions for tweaks. One particular choice I made is to make most of the functions return dicts by default; this isn't strictly necessary for inputs but it is very convenient for outputs as the output desc lives on the output node, not the argument that feeds into the node. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159005 Approved by: https://github.com/wconstab	2025-07-25 22:52:33 +00:00
IvanKobzarev	8aebf01287	[bucketing] Rewrite all_gather, reduce_scatter passes via tracing merge_fn (#158663 ) Rewriting bucketing of all_gather and reduce_scatter with defining of "merge graph" via torch function. `all_gather_merge_fn_to_trace` `reduce_scatter_merge_fn_to_trace` (Instead of creating nodes and doing FakeTensor prop manually) This allows to experiment with merge function. Used foreach_copy_ in merging function for all_gather - added lowering for inductor for `foreach_copy_` Adding topological sort after bucketing passes (comment in post_grad.py): ``` # Fx collectives bucketing passes require topological sort for the cases: # when bucketed collectives have users before the last collective in the bucket # AND when inputs of bucketed collective have ancestors after the first collective in the bucket. # # In this case we can not manually pick the place for bucketed collective insertion. # But we are guaranteed by the bucketing (independent collectives in the bucket), # that it is possible to reorder nodes to satisfy all ordering requirements. # # --- before bucketing --- # in0 = ... # wait_ag0 = ag(in0) # user0(wait_ag0) # ... # pre_in1 = ... # in1 = transform(pre_in1) # wait_ag1 = ag(in1) # user1(wait_ag1) # # --- after bucketing --- # # in0 = ... # user(wait_ag0) <--- wait_ag0 is defined only after bucketed collective. # # pre_in1 = ... # in1 = transform(pre_in1) # ag_bucket(in0+in1) # wait_bucket # wait_ag0 = wait_bucket[0] # wait_ag1 = wait_bucket[1] # user1(wait_ag1) ```` Correctness of the passes verified by loss curve for llama3 8b for simple_fsdp and for autoparallel: <img width="1364" height="495" alt="Screenshot 2025-07-22 at 14 27 28" src="https://github.com/user-attachments/assets/67b2cabb-3206-450b-b529-e23c24292fc6" /> <img width="1355" height="509" alt="Screenshot 2025-07-22 at 14 27 56" src="https://github.com/user-attachments/assets/4d0e6b25-2eb1-47b2-8d68-dcec185239c4" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158663 Approved by: https://github.com/wconstab	2025-07-25 22:49:51 +00:00
Yu Guo	bc5dbbbb78	support scalar tensor for functional all_gather (#149913 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/149913 Approved by: https://github.com/H-Huang ghstack dependencies: #149912	2025-07-25 22:38:08 +00:00
Mikayla Gawarecki	36cf8f1ed8	[BE] Use .md instead of .rst for nn.aliases doc (#158666 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158666 Approved by: https://github.com/janeyx99 ghstack dependencies: #158491, #158654	2025-07-25 22:03:55 +00:00
Mikayla Gawarecki	1e79872f2e	[BE] More torch.nn docs coverage test (except for torch.nn.parallel) (#158654 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158654 Approved by: https://github.com/janeyx99 ghstack dependencies: #158491	2025-07-25 22:03:55 +00:00
Mikayla Gawarecki	9e8f27cc79	[BE] Make torch.nn.modules.* satisfy the docs coverage test (#158491 ) Options to address the "undocumented python objects": 1. Reference the functions in the .rst via the torch.nn.modules namespace. Note that this changes the generated doc filenames / locations for most of these functions! 2. [Not an option] Monkeypatch `__module__` for these objects (broke several tests in CI due to `inspect.findsource` failing after this change) 3. Update the .rst files to also document the torch.nn.modules forms of these functions, duplicating docs. #### [this is the docs page added](https://docs-preview.pytorch.org/pytorch/pytorch/158491/nn.aliases.html) This PR takes option 3 by adding an rst page nn.aliases that documents the aliases in nested namespaces, removing all the torch.nn.modules.* entries from the coverage skiplist except - NLLLoss2d (deprecated) - Container (deprecated) - CrossMapLRN2d (what is this?) - NonDynamicallyQuantizableLinear This mostly required adding docstrings to `forward`, `extra_repr` and `reset_parameters`. Since forward arguments are already part of the module docstrings I just added a very basic docstring. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158491 Approved by: https://github.com/janeyx99	2025-07-25 22:03:55 +00:00
Mikayla Gawarecki	e65ab9a868	Enable generating generic c_shim that doesn't bypass dispatcher (#158974 ) Adds `c_shim_aten.{h/cpp}` and use this for `fill_` This is the generated `c_shim_aten.cpp` for reference ```cpp // WARNING: THIS FILE IS AUTOGENERATED BY torchgen. DO NOT MODIFY BY HAND. // See `7e86a7c015/torchgen/gen.py (L2424-L2436)` for details // This file corresponds to the aten_shimified_ops list in torchgen/aoti/fallback_ops.py #include <torch/csrc/inductor/aoti_torch/generated/c_shim_aten.h> #include <torch/csrc/inductor/aoti_torch/utils.h> #ifndef AT_PER_OPERATOR_HEADERS #include <ATen/Functions.h> #include <ATen/CompositeExplicitAutogradFunctions.h> #include <ATen/CompositeExplicitAutogradNonFunctionalFunctions.h> #include <ATen/CompositeImplicitAutogradFunctions.h> #else #include <ATen/ops/fill.h> #endif // AT_PER_OPERATOR_HEADERS using namespace torch::aot_inductor; AOTITorchError aoti_torch_aten_fill__Scalar(AtenTensorHandle self, double value) { AOTI_TORCH_CONVERT_EXCEPTION_TO_ERROR_CODE({ at::fill_( *tensor_handle_to_tensor_pointer(self), value ); ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158974 Approved by: https://github.com/albanD, https://github.com/janeyx99	2025-07-25 21:59:14 +00:00
Pian Pawakapan	bfe6765d6b	[export] assert fix in serdes (#159060 ) Summary: catch asserts on True Test Plan: T232064560 Rollback Plan: Differential Revision: D78907485 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159060 Approved by: https://github.com/yiming0416	2025-07-25 21:46:20 +00:00
Denghui Dong	e88f804a2e	[Profiler] Fix lost C call events problem in Python 3.12.0-3.12.4 (#155446 ) Hi team, Please help review this patch. This PR https://github.com/pytorch/pytorch/pull/150370 tried to fix the "Empty C Call Queue" problem on Python 3.12. It added C calls for each starting Python event with a callable. I found the root cause is not that we cannot get C function frames by `PyFrame_GetBack` when PythonTracer is filling start frames, but the c call event loss problem bug on Python 3.12.0-3.12.4. And that problem was fixed by `257c413cd1` on 3.12.5. So I think the https://github.com/pytorch/pytorch/pull/150370 cannot fix the problem, this patch reverts the change of it. There are solutions to fix the problem correctly, such as we can add a new monitoring callback to compensate call events of methods with C function or we can override the callback registered by `PyEval_SetProfile`. These solutions may make the code hard to maintain. ~~Since upgrading the micro version of Python is not difficult for users, we can just ignore C functions and suggest user upgrade.~~ Pull Request resolved: https://github.com/pytorch/pytorch/pull/155446 Approved by: https://github.com/sraikund16	2025-07-25 21:44:57 +00:00
raghavhrishi	7ef3c3357d	NUMA binding integration with elastic agent and torchrun (#149334 ) Implements #148689 Pull Request resolved: https://github.com/pytorch/pytorch/pull/149334 Approved by: https://github.com/d4l3k Co-authored-by: Paul de Supinski <pdesupinski@gmail.com>	2025-07-25 21:19:49 +00:00
Thomas Bohnstingl	24b1f10ca1	[HOP, map] Rework of map autograd to the new interface (#153343 ) This PR reworks the current autograd implementation of map to the new interface. @pytorchbot label "topic: not user facing" Pull Request resolved: https://github.com/pytorch/pytorch/pull/153343 Approved by: https://github.com/ydwu4	2025-07-25 21:17:06 +00:00
Yidi Wu	0006dd5c43	[test][torchbind] don't allow set torchbind attr at runtime (#158608 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158608 Approved by: https://github.com/zou3519 ghstack dependencies: #158583, #158606, #158607	2025-07-25 20:55:41 +00:00
Yidi Wu	0f31e9a656	[torchbind] fix fakifying a staitc tensor returns dynamic accidentally (#158607 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158607 Approved by: https://github.com/zou3519 ghstack dependencies: #158583, #158606	2025-07-25 20:55:41 +00:00
Yidi Wu	0427e439aa	[test][torchbind] turn on inductor backend for compile torchbind tests (#158606 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158606 Approved by: https://github.com/zou3519 ghstack dependencies: #158583	2025-07-25 20:55:41 +00:00
Yidi Wu	4aa69ae336	[torchbind] support register_autocast for torchbind custom op (#158583 ) Fix https://github.com/pytorch/pytorch/issues/158414 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158583 Approved by: https://github.com/zou3519	2025-07-25 20:55:41 +00:00
dolpm	14c314b30d	[nativert] make per-node benchmark work with memory planning (#159117 ) Summary: this will use-after-free otherwise Rollback Plan: Differential Revision: D78934104 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159117 Approved by: https://github.com/SherlockNoMad	2025-07-25 20:46:17 +00:00
Pian Pawakapan	0b01e11416	[ez][export] add sym_sum to verified ops (#159111 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159111 Approved by: https://github.com/angelayi	2025-07-25 20:42:42 +00:00
NikhilAPatel	806d9e3fe7	[Inductor][TMA] Split config-gated and pure compatibility logic for TMA template eligibility checks (#159123 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159123 Approved by: https://github.com/drisspg	2025-07-25 20:35:49 +00:00
Yu Guo	d90ce83027	add a util function _make_all_gather_out_tensor to reduce code duplication (#149912 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/149912 Approved by: https://github.com/H-Huang	2025-07-25 20:29:01 +00:00
Xu Han	dfcb07bdfa	[Inductor] disable windows failed UTs temporary. (#159163 ) Disable windows failed UTs temporary. <img width="1238" height="107" alt="image" src="https://github.com/user-attachments/assets/c8a40408-a793-4016-99bb-19c1bb09860a" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159163 Approved by: https://github.com/desertfire	2025-07-25 20:25:36 +00:00
Sam Larsen	fa0355c18d	Fix full_like decomposition to preserve strides (#158898 ) Summary: See original PR at: https://github.com/pytorch/pytorch/pull/144765, which landed internally but was reverted due to test failures. Addressing reviewer comments and trying again. Rollback Plan: Differential hack Revision: D78783627 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158898 Approved by: https://github.com/eellison	2025-07-25 20:21:36 +00:00
Sherlock Huang	28ee8be5bf	[NativeRT] Apply Device placement once when loading the graph (#158996 ) Summary: Placement is leaked to too many classes! In this diff, we consolidate all placement lookup into one place: Graph::ApplyDevicePlacement. After applying placement, the in-memory graph, tensorMeta, weightMeta would already have the re-mapped device. The subsequence weight loading, sample input loading, target device inference would look up the re-mapped device from graph's tensorMeta. graph's tensorMeta becomes the only ground truth! Test Plan: Need to add some tests before landing. This is a big change. Rollback Plan: Differential Revision: D78841818 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158996 Approved by: https://github.com/henryoier	2025-07-25 20:11:35 +00:00
Yidi Wu	ed472257d1	[associative_scan] stop manually set example inputs in dynamo (#159065 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159065 Approved by: https://github.com/zou3519 ghstack dependencies: #159063, #159064	2025-07-25 20:08:08 +00:00
Yidi Wu	57eea56a9a	[scan] stop manually set example inputs in dynamo (#159064 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159064 Approved by: https://github.com/zou3519 ghstack dependencies: #159063	2025-07-25 20:08:08 +00:00
Yidi Wu	dd681f7f59	[while_loop] stop manually setting example inputs in dynamo (#159063 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159063 Approved by: https://github.com/zou3519	2025-07-25 20:08:08 +00:00
Howard Huang	0d4d3e8a89	[TCPStore] Allow ping to be retried (#159165 ) On client setup we retry connections with server: `f8fafdc7a6/torch/csrc/distributed/c10d/TCPStore.cpp (L313-L350)` I noticed `ping()` raises `TORCH_INTERNAL_ASSERT` AKA a runtime error rather than a `DistNetworkError`. So updating that so it can be retried as well. We have seen this pop up internally: - https://fb.workplace.com/groups/319878845696681/permalink/1478849733132914/ - https://fb.workplace.com/groups/319878845696681/permalink/1479368959747658/ Pull Request resolved: https://github.com/pytorch/pytorch/pull/159165 Approved by: https://github.com/d4l3k	2025-07-25 20:03:00 +00:00
cz2h	ee4c5c7cd2	Add torchcheck for replication_pad3d_backward (#151986 ) Fixes #142833 Add check on channel dimension, logic same to the CUDA implementation `78bbb468c6/aten/src/ATen/native/cuda/ReplicationPadding.cu (L347)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/151986 Approved by: https://github.com/mikaylagawarecki	2025-07-25 19:48:51 +00:00
Sameer	51cd6697cd	Fix: Use memory_order_relaxed instead of memory_order_relaxed (#159105 ) Addresses #159074 by using `memory_order_release` instead of `memory_order_relaxed` here: `9c10760662/c10/core/DeviceType.cpp (L161)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/159105 Approved by: https://github.com/colesbury	2025-07-25 19:39:04 +00:00
Xu Han	ba949c54a7	[inductor] fix test_save_graph_repro on Windows. (#159148 ) The issue is caused by Windows path separator work as escape character. Fixed by `normalize_path_separator` in torch front end codegen. Error message: <img width="855" height="542" alt="image" src="https://github.com/user-attachments/assets/ad08b521-05e6-4c93-9507-ad19c68ac7b5" /> Fixed: <img width="855" height="312" alt="image" src="https://github.com/user-attachments/assets/4a0a142a-2dbe-4226-a4cb-8eacfab2c3fc" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159148 Approved by: https://github.com/desertfire	2025-07-25 19:11:08 +00:00
morrison-turnansky	2a528e80ce	Add more type hints for _inductor/ir.py (#159049 ) Fixes #146167 Incremental step to add type hints for _inductor/ir.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/159049 Approved by: https://github.com/Skylion007	2025-07-25 18:56:30 +00:00
Edward Z. Yang	56c45f863b	Add aot_export_joint_with_descriptors and aot_compile_joint_with_descriptors (#158715 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158715 Approved by: https://github.com/fmassa, https://github.com/wconstab, https://github.com/xmfan ghstack dependencies: #158624, #158708, #158734	2025-07-25 18:49:00 +00:00
albanD	d30f89b9b8	Add host protoc script back (#159157 ) Following comment from https://github.com/pytorch/pytorch/pull/158475#issuecomment-3116518904 Also this is a fake issue as protoc is dead anyways: https://github.com/pytorch/pytorch/issues/159156 Also also, macos cross compilation is not something that is tested :/ But I guess we're ok with that given how niche it it... Pull Request resolved: https://github.com/pytorch/pytorch/pull/159157 Approved by: https://github.com/janeyx99	2025-07-25 18:44:20 +00:00
PyTorch MergeBot	3fb78501f0	Revert "enable compiled autograd on CPU windows (#158432 )" This reverts commit a369350065493109d1abfbb994695777ab11bcf4. Reverted https://github.com/pytorch/pytorch/pull/158432 on behalf of https://github.com/atalman due to Broke audio cuda windows builds see: https://github.com/pytorch/audio/issues/3992 ([comment](https://github.com/pytorch/pytorch/pull/158432#issuecomment-3119912177))	2025-07-25 18:29:16 +00:00
angelayi	8a0508335f	[export] Fix public bindings (#159109 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/159109 Approved by: https://github.com/jbschlosser	2025-07-25 18:18:52 +00:00
rajeshvshiyal	4c0d5ad4be	Fix docstring for clip_grads_with_norm_ to reflect clamping behavior (#158200 ) Fix docstring for clip_grads_with_norm_ to reflect clamping behavior This PR updates the docstring for torch.nn.utils.clip_grads_with_norm_ to accurately reflect the implementation behavior. The current documentation suggests that gradients are always scaled by: grad = grad * (max_norm / (total_norm + eps)) However, the actual implementation clamps the scale coefficient to a maximum of 1.0, ensuring gradients are only scaled down, not up. This PR corrects the formula and adds a clarifying note to avoid confusion for users. Updated the formula in the docstring to: grad = grad * min(max_norm / (total_norm + eps), 1.0) Added a note explaining the rationale for clamping (to prevent gradient amplification). Ensured consistency with the behavior of clip_grad_norm_. Fixes #151554 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158200 Approved by: https://github.com/mikaylagawarecki	2025-07-25 18:07:41 +00:00
Joel Schlosser	316c188a5e	Remove torch.functional entries from the doc ignore list (#158581 ) Options to address the "undocumented python objects": 1. Reference the functions in the .rst via the `torch.functional` namespace. Note that this changes the generated doc filenames / locations for most of these functions! 2. Document these functions by referencing them from the `torch.` namespace instead, in line with common usage. This would also require setting the `__module__` for these functions and moving entries from `torch.functional`'s `__all__` -> `torch`'s `__all__`, which is BC-breaking. 3. Update the .rst files to also document the `torch.functional` forms of these functions, duplicating docs. This PR takes option (3) above and: * Removes all 20 `torch.functional` entries from the doc ignore list * Removes `torch.functional.align_tensors()` entirely, since we don't want to document it. * This is technically BC-breaking, although the previous impl simply errored out. This change could be moved to a separate isolated PR for safety. * Introduces `torch.aliases.md` as a hidden page for the `torch.functional` aliases to the `torch` analogue functions Pull Request resolved: https://github.com/pytorch/pytorch/pull/158581 Approved by: https://github.com/janeyx99	2025-07-25 17:19:01 +00:00
Edward Z. Yang	191eca0bf0	Use simple_wraps instead of functools.wraps in AOTAutograd (#158734 ) Wrapping is load bearing for things that introspect argument signatures, but use of functools.wraps to do this is undesirable as this overrides the name/module of the wrapping function, which is bad for tracking down exactly what code is actually being run at runtime. simple_wraps is like wraps but it doesn't override the name information, so you still get an appropriate printout. To see the stack of all functions wrapping each other, there is now a helper fn_stack. I also make some assertions tighter in the descriptor PR. These didn't catch any bugs but I figure might as well. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158734 Approved by: https://github.com/wconstab ghstack dependencies: #158624, #158708	2025-07-25 17:08:54 +00:00
Sheng Fu	74f64d3c84	Add inputs and outputs in Triton Kernel FX Graph segment (#158174 ) Summary: Add inputs and outputs in Triton Kernel FX Graph segment The FX graph segment in Triton kernel does not include the input tensors and return tensors, for example Python code: ``` @torchdynamo.optimize("inductor") def fn(a, b, c): x = torch.nn.functional.linear(a, b) x = x.sin() x = x.t() + c * 2 return x ``` ``` # %sin : "f32[4, 4][4, 1]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.sin.default](args = (%mm,), kwargs = {}) # %permute_1 : "f32[4, 4][1, 4]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.permute.default](args = (%sin, [1, 0]), kwargs = {}) # %mul : "f32[4, 4][4, 1]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.mul.Tensor](args = (%arg2_1, 2), kwargs = {}) # %add : "f32[4, 4][1, 4]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.add.Tensor](args = (%permute_1, %mul), kwargs = {}) ``` The fix is to add the input and output tensors into FX graph segment ``` # %mm : Tensor "f32[4, 4][4, 1]cuda:0" = PlaceHolder[target=mm] # %arg2_1 : Tensor "f32[4, 4][4, 1]cuda:0" = PlaceHolder[target=arg2_1] # %sin : "f32[4, 4][4, 1]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.sin.default](args = (%mm,), kwargs = {}) # %permute_1 : "f32[4, 4][1, 4]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.permute.default](args = (%sin, [1, 0]), kwargs = {}) # %mul : "f32[4, 4][4, 1]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.mul.Tensor](args = (%arg2_1, 2), kwargs = {}) # %add : "f32[4, 4][1, 4]cuda:0"[num_users=1] = call_function[target=torch.ops.aten.add.Tensor](args = (%permute_1, %mul), kwargs = {}) # return %add ``` Differential Revision: D78131358 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158174 Approved by: https://github.com/jansel	2025-07-25 17:01:17 +00:00
PyTorch MergeBot	f8fafdc7a6	Revert "[BE] remove torch deploy - conditionals (#158288 )" This reverts commit ab26d4fbeb5bc4b4e6ef1c37fbec9fab6e5a9edd. Reverted https://github.com/pytorch/pytorch/pull/158288 on behalf of https://github.com/ZainRizvi due to Reverting as per offline discussion to fix internal breaks. @PaliC will reland this as a codev diff. Instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158288#issuecomment-3119037960))	2025-07-25 16:09:39 +00:00
PyTorch MergeBot	c8316d0e79	Revert "[BE] Remove torch deploy \| remove torch deploy specific files (#158290 )" This reverts commit 6ed2cb6ccd00e64f67fd414d42dff54393140c8f. Reverted https://github.com/pytorch/pytorch/pull/158290 on behalf of https://github.com/ZainRizvi due to Reverting as per offline discussion to fix internal breaks. @PaliC will reland this as a codev diff. Instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158288#issuecomment-3119037960))	2025-07-25 16:09:39 +00:00
PyTorch MergeBot	a9f6770edd	Revert "[BE] Remove __reduce_deploy__ (#158291 )" This reverts commit 9c68c4d08f4c4da49f0086b80e382f0cdd518f60. Reverted https://github.com/pytorch/pytorch/pull/158291 on behalf of https://github.com/ZainRizvi due to Reverting as per offline discussion to fix internal breaks. @PaliC will reland this as a codev diff. Instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158288#issuecomment-3119037960))	2025-07-25 16:09:39 +00:00
PyTorch MergeBot	5620e617c9	Revert "[BE] Modify PyObjectSlot the assume only a single interpreter is in use (#158407 )" This reverts commit 255c0545e7eac2ec6d00a41a3fc9d6d8201f8f39. Reverted https://github.com/pytorch/pytorch/pull/158407 on behalf of https://github.com/ZainRizvi due to Reverting as per offline discussion to fix internal breaks. @PaliC will reland this as a codev diff. Instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158288#issuecomment-3119037960))	2025-07-25 16:09:39 +00:00
Huy Do	ee84ba42ea	[Experiment] Run PT2 benchmark twice a day (#159162 ) Running every 4 hours seems too many, lower it to twice a day. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159162 Approved by: https://github.com/desertfire, https://github.com/eellison	2025-07-25 15:58:29 +00:00
Catherine Lee	561193e5f2	[CI][testing] Use 3 processes for testing on sm89 and sm90 jobs (#158691 ) 3 procs were used for sm86, but we switched to sm89 and the check failed so it switched back to 2 sm90 is H100, but idk what unittests we have running there, but I assume they also have a lot of memory They use larger runners, which have more GPU memory, so its usually ok. I think it's ~22GB -> 10GB per proc if 2, 6GB per proc if 3 (cuda context maybe 1GB) I've applied skips to the ones that OOMed Time decreases from ~2.7hr per test job -> ~2hr Pull Request resolved: https://github.com/pytorch/pytorch/pull/158691 Approved by: https://github.com/huydhn	2025-07-25 15:26:29 +00:00
PyTorch MergeBot	9535995bbc	Revert "Remove tensorexpr tests (#158928 )" This reverts commit a0bc865123dba047aa1507e281bf2462780cf271. Reverted https://github.com/pytorch/pytorch/pull/158928 on behalf of https://github.com/clee2000 due to broke cpp static runtime test? [GH job link](https://github.com/pytorch/pytorch/actions/runs/16517697273/job/46715871457) [HUD commit link](`a0bc865123`) ([comment](https://github.com/pytorch/pytorch/pull/158928#issuecomment-3118554478))	2025-07-25 15:22:51 +00:00
William Wen	6fcb2b4413	[dynamo] unimplemented -> unimplemented_v2 for user_defined.py (#156652 ) For https://github.com/pytorch/pytorch/issues/147913 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156652 Approved by: https://github.com/zou3519 Co-authored-by: Sidharth <ssubbarao8@meta.com>	2025-07-25 15:04:17 +00:00
Edward Z. Yang	204eb4da5e	Add expanded_def option for FX printing, render descriptor, update tests (#158708 ) ---- - First, we add a new expanded_def to FX, which will expand the definitions of variables into multiple lines, one per variable definition. This makes extremely long args/return lists much more readable. - Next, we extend this mechanism to also print out descriptors on placeholders and return values, as comments, if available. This is how we will test descriptors. - We update tlparse for AOTAutograd to use this format. - We update expect tests to use this format and update their formats, so you can inspect what it can look at. There may be other tests I should update, open to suggestions. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158708 Approved by: https://github.com/wconstab ghstack dependencies: #158624	2025-07-25 13:22:32 +00:00
Edward Z. Yang	bf311141d6	Track descriptors for all inputs/outputs of AOTAutograd traced graph (#158624 ) One of the recurring challenges of working with FX graphs produced by AOTAutograd is that there is a very intricate input/output calling convention that is essentially impossible to understand without actually reverse engineering the AOTAutograd code. It is so bad that there is a bit of logic for stashing indices of relevant arguments/outputs in TracingContext so Inductor can figure out what the correct arguments are. This PR introduces the necessary scaffolding to keep track of "descriptors" of every input/output to a (joint) FX graph produced by AOTAutograd. First read through descriptors.py to get a sense for what is available: for inputs, you can figure out if you have a plain input, tangent, parameter, or something more exotic like one of the fields of a subclass or view base. For outputs, you can determine if you have a plain output or grad, or something more exotic like the contents of a mutated input or an intermediate base of several views that were returned. There are two distinct parts of this patch: AOTInput tracking, and AOTOutput tracking. AOTInput tracking. The way this works is that AOTAutograd starts of with some Tensor `flat_args` that are the inputs to the graph being traced, and then updates these arguments as it modifies the input calling convention. Anywhere these `args` are passed around, we now add a news argument `args_descs` which is updated in synchrony with args. Add a new arg? Add a new AOTInput to `args_descs`. AOTOutput tracking. Originally, I wanted to also add an `outs_descs` analogous to `args_descs` tracking output metadata. However, it is often difficult to compute what the output will be until you're actually tracing the function for real (and are able to peek at the real outputs). So we only compute `outs_desc` when we actually trace. To do this, we change the calling convention of the function we trace to return not just outputs, but a tuple of `outs` and `outs_descs`. Before we bottom out at the `make_fx` invocation, we save `outs_descs` to a nonlocal and bottom out. To actually make use of this information in a useful way, see the next PR. Potentially the two PRs could be combined together but I think it's actually clearer for them to be separate. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158624 Approved by: https://github.com/xmfan	2025-07-25 13:22:32 +00:00
thenumberouscode	92e93bb580	[inductor][cpu] Stop lowering div to reciprocal multiplication to preserve precision when the divisor is a scalar and device is on cpu (#158231 ) ## Fixes https://github.com/pytorch/pytorch/issues/157959 ## mini repro from issue ```c++ import torch from torch import nn class Foo(nn.Module): def __init__( self, use_parameter: bool ) -> None: super().__init__() self.b = 101 if use_parameter: self.b = nn.Parameter(torch.Tensor([self.b]), requires_grad=False) def forward(self, x: torch.Tensor) -> torch.Tensor: # return x + self.b # return x - self.b return x / self.b # return x * self.b torch.manual_seed(42) x = torch.rand((5, 5)) expected = Foo(False)(x) models = [ Foo(False), Foo(True), torch.compile(Foo(False), fullgraph=True), torch.compile(Foo(True), fullgraph=True), ] for m in models: print((m(x) - expected).sum()) ``` all outputs equal zero except the result of torch.compile(Foo(False), fullgraph=True) ## summary: when divisor is a scalar, inductor will lower div to mul the scalar's reciprocal. this could lead precision lost in c++ kernel. but not in triton kernel ## why: Generated C++ kernel; thanks to @xmfan for supplying the code. ```c++ #include <torch/csrc/inductor/cpp_prefix.h> extern "C" void kernel(const float* in_ptr0, float* out_ptr0) { { for(int64_t x0=static_cast<int64_t>(0L); x0<static_cast<int64_t>(25L); x0+=static_cast<int64_t>(16L)) { { if(C10_LIKELY(x0 >= static_cast<int64_t>(0) && x0 < static_cast<int64_t>(16L))) { auto tmp0 = at::vec::Vectorized<float>::loadu(in_ptr0 + static_cast<int64_t>(x0), static_cast<int64_t>(16)); auto tmp1 = static_cast<float>(0.009900990099009901); auto tmp2 = at::vec::Vectorized<float>(tmp1); auto tmp3 = tmp0 * tmp2; tmp3.store(out_ptr0 + static_cast<int64_t>(x0)); } if(C10_UNLIKELY(x0 >= static_cast<int64_t>(16L) && x0 < static_cast<int64_t>(25L))) { auto tmp0 = at::vec::Vectorized<float>::loadu(in_ptr0 + static_cast<int64_t>(x0), static_cast<int64_t>(9L)); auto tmp1 = static_cast<float>(0.009900990099009901); auto tmp2 = at::vec::Vectorized<float>(tmp1); auto tmp3 = tmp0 * tmp2; tmp3.store(out_ptr0 + static_cast<int64_t>(x0), static_cast<int64_t>(9L)); } } } } } ``` The float type in C typically has 6 to 7 significant digits, while the double type has 15 to 16 significant digits. ```c++ #include <iostream> #include <iomanip> int main() { auto tmp1 = static_cast<float>(0.009900990099009901); auto tmp2 = static_cast<double>(0.009900990099009901); std::cout << std::setprecision(20) << "tmp1 = " << tmp1 << std::endl; std::cout << std::setprecision(20) << "tmp2 = " << tmp2 << std::endl; return 0; } ``` the ouput is ```bash tmp1 = 0.0099009899422526359558 tmp2 = 0.0099009900990099011103 ``` `auto tmp1 = static_cast<float>(0.009900990099009901);` This will cause tmp1 to become 0.0099009, resulting in a loss of precision, so the final result will not match the expected value. I also found that the bug occurred at that position `86d8af6a6c/torch/_inductor/lowering.py (L6238)` The commit states that the precision lost is expected in cuda implementation. original commit `03439d4c1c` cuda implementation `0636c11811/aten/src/ATen/native/cuda/BinaryDivTrueKernel.cu (L36-L38)` What is interesting is that the Triton kernel works correctly due to the precision of float type in python. ```python def triton_poi_fused_div_0(in_ptr0, out_ptr0, xnumel, XBLOCK : tl.constexpr): xnumel = 25 xoffset = tl.program_id(0) * XBLOCK xindex = xoffset + tl.arange(0, XBLOCK)[:] xmask = xindex < xnumel x0 = xindex tmp0 = tl.load(in_ptr0 + (x0), xmask) tmp1 = 0.009900990099009901 tmp2 = tmp0 * tmp1 tl.store(out_ptr0 + (x0), tmp2, xmask) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158231 Approved by: https://github.com/eellison	2025-07-25 08:57:17 +00:00
cyy	a0bc865123	Remove tensorexpr tests (#158928 ) The tests are not maintained. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158928 Approved by: https://github.com/albanD	2025-07-25 08:37:51 +00:00
Laith Sakka	aaa384b2d4	move view_meta to fake impl (#158406 ) Python dispatcher is not always enabled in fake tensors and have to be called explicitly. While it should be, it requires some work to get all tests working. I have been running in several issues where I add to add enable_python_dispatcher ex XLA, Helom ..etc to avoid issues related to that for the view specifically i moved it to fake tensor impl. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158406 Approved by: https://github.com/bobrenjc93	2025-07-25 08:21:27 +00:00
Jeff Daily	0fd5f1c294	[ROCm][CI] upgrade wheels to 6.4.2 patch release (#158886 ) Upgrade wheels to ROCm 6.4.2 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158886 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-25 08:11:41 +00:00
Xu Han	e38a2b3d0f	[inductor] add missing ignore_errors parameter for Windows. (#159025 ) The origin code comemnts: ```python # Let's not fail if we can't clean up the temp dir. Also note that for # Windows, we can't delete the loaded modules because the module binaries # are open. ``` But we are missing the `ignore_errors` parameter for Windows. I help to add it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159025 Approved by: https://github.com/jansel	2025-07-25 07:58:22 +00:00
Robert Burke	ae183d6092	Aten vector default constructors set to 0, add fnmadd and fnmsub (#158508 ) cc jgong5 mingfeima XiaobingSuper sanchitintel ashokei jingxu10 jerryzh168 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158508 Approved by: https://github.com/swolchok	2025-07-25 06:55:37 +00:00
Animesh Jain	659f8fb115	[dynamo][guards] Add some relational guard helpers (#159077 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159077 Approved by: https://github.com/jansel ghstack dependencies: #158995	2025-07-25 06:28:10 +00:00
Animesh Jain	05a748d287	[dynamo][guards] Expand is_immutable_object to have None (#158995 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158995 Approved by: https://github.com/Lucaskabela, https://github.com/jansel	2025-07-25 06:12:05 +00:00
Han, Chao1	02ca965560	Device agnostic for DCP (#158337 ) Enable device-agnostic implementation of DCP-related functionality, allowing the new DCP features to be supported on XPU as well. use_cuda_non_blocking_copy to use_non_blocking_copy because non-blocking copy is supported by most GPUs and is not exclusive to CUDA devices. Test plan: test cases have not yet been updated to be fully device agnostic; this will be addressed in future work. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158337 Approved by: https://github.com/guangyey, https://github.com/EikanWang, https://github.com/Saiteja64 Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-07-25 05:24:09 +00:00
Dylan Maloy	511d987378	only call re-plan if historic max's were updated. (#159016 ) Summary: wasteful. only update the plan if a new maximum has been found. Test Plan: ci Rollback Plan: Reviewed By: SherlockNoMad Differential Revision: D78859344 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159016 Approved by: https://github.com/SherlockNoMad	2025-07-25 05:07:30 +00:00
zeshengzong	9685fc36d4	Add missing optional for tensor ops (#159028 ) ## Test Result <img width="872" height="340" alt="image" src="https://github.com/user-attachments/assets/20c3f1a2-0160-4ea3-b9f3-14630b4ec06d" /> <img width="906" height="429" alt="image" src="https://github.com/user-attachments/assets/68f8d8da-0570-4ae8-8e45-573b2c64cae5" /> <img width="906" height="429" alt="image" src="https://github.com/user-attachments/assets/42d133f6-94eb-4a38-8b4b-5586f52bff88" /> <img width="878" height="285" alt="image" src="https://github.com/user-attachments/assets/d3ad8950-81fa-4c4c-a5b5-621b0d9df99b" /> <img width="889" height="430" alt="image" src="https://github.com/user-attachments/assets/9aabeaff-bb8f-4990-b253-1bb053e72aca" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/159028 Approved by: https://github.com/Skylion007	2025-07-25 04:36:55 +00:00
PyTorch UpdateBot	9e5cfd3ee5	[audio hash update] update the pinned audio hash (#159108 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159108 Approved by: https://github.com/pytorchbot	2025-07-25 04:35:21 +00:00
Nikita Shulga	cdf8e9ec1a	[MPS] Add support for unsigned types (#159094 ) As both Metal and MPS support uint16/uint32 and uint64 Test plan: `python3 -c "import torch;print(torch.randint(55, 66, (16,), device='mps', dtype=torch.uint16)[10:])"` Fixes https://github.com/pytorch/pytorch/issues/159076 Pull Request resolved: https://github.com/pytorch/pytorch/pull/159094 Approved by: https://github.com/Skylion007, https://github.com/dcci	2025-07-25 04:31:42 +00:00
PyTorch UpdateBot	bcf34d24eb	[vllm hash update] update the pinned vllm hash (#159107 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned vllm hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159107 Approved by: https://github.com/pytorchbot	2025-07-25 04:03:39 +00:00
Jeff Daily	9b29166f57	[ROCm] add flag torch.backends.miopen.immediate (#158951 ) The MIOpen integration has changed over the years. In the past, the MIOpen default for benchmark was True and if it were set to False it would use MIOpen Immediate Mode. But with #145294 the MIOpen benchmark default changed to False and to activate immediate mode you would set the deterministic flag to True. This has proved too restrictive because benchmark and deterministic flags are independent from immediate mode. Thus, immediate mode needs its own flag. Though MIOpen still masquerades behind torch.backends.cudnn and its flags, it seemed inappropriate to add an miopen-exclusive flag to the set of cudnn flags. This PR adds the first miopen-only flag to control its immediate mode. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158951 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-25 04:01:51 +00:00
Jeff Daily	1fced0c7d5	[ROCm] enable hipblaslt on gfx908 for ROCm >= 6.3 (#159092 ) Fixes #159030. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159092 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-25 03:54:30 +00:00
Nichols A. Romero	16c0ccd669	[ROCm][CI] upgrade to 6.4.2 patch release (#158887 ) Upgrade to ROCm 6.4.2. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158887 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-25 03:45:44 +00:00
Xuehai Pan	f5e2de928b	[BE] fix remaining flake8 v7 warnings (#159044 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159044 Approved by: https://github.com/Skylion007 ghstack dependencies: #159043	2025-07-25 02:56:34 +00:00
Xuehai Pan	f903bc475c	[BE] add noqa for flake8 rule B036: found `except BaseException` without re-raising (#159043 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/159043 Approved by: https://github.com/Skylion007	2025-07-25 02:56:34 +00:00
FFFrog	4261e26a8b	[OpenReg] move fallback tests into test_openreg.py (#158441 ) ---- - move fallback tests into test_operneg - remove the test_cpp_extensions_open_device_registration.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/158441 Approved by: https://github.com/albanD ghstack dependencies: #158415, #158440	2025-07-25 02:39:41 +00:00
FFFrog	b635359e4c	[OpenReg] add pyproject.toml for openreg (#158440 ) As the title stated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158440 Approved by: https://github.com/albanD ghstack dependencies: #158415	2025-07-25 02:39:41 +00:00
FFFrog	f1a1aa9490	[OpenReg] Improve README.md and optimize some codes for OpenReg (#158415 ) ---- - add description for DSO dependencies - remove unnecessary code Pull Request resolved: https://github.com/pytorch/pytorch/pull/158415 Approved by: https://github.com/albanD	2025-07-25 02:39:41 +00:00
FFFrog	6fc0ad22f0	Using the latest torch.library.register_fake API instead of torch.library.impl_abstract (#158839 ) As the title stated. `torch.library.impl_abstract` have beed deprecated in PyTorch2.4, so change to use the new API. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158839 Approved by: https://github.com/jingsh, https://github.com/zou3519 ghstack dependencies: #158838	2025-07-25 02:37:30 +00:00
FFFrog	c60d382870	Add tests for torch.ops.load_library (#158838 ) According to this [comment](https://github.com/pytorch/pytorch/pull/157524#issuecomment-3097899129), adding a related test to keep BC. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158838 Approved by: https://github.com/zou3519	2025-07-25 02:37:30 +00:00
Chuan Jiang	64cb349b81	Extract a method that filters frames in the captured stack trace (#158266 ) Summary: The subclass can override the filtering logic to customize which frames to keep or drop. Test Plan: ``` buck run caffe2/test:test_export -- -r test_stack_trace buck2 run 'fbcode//mode/dev-nosan' fbcode//caffe2/test:others -- -r test_constant_random buck2 run 'fbcode//mode/dev-nosan' fbcode//caffe2/test:test_export -- -r test_custom_obj_list_out buck2 run 'fbcode//mode/dev-nosan' fbcode//caffe2/test:fx -- -r class_member_back_compat ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158266 Approved by: https://github.com/ezyang, https://github.com/yushangdi	2025-07-25 02:22:03 +00:00
PyTorch MergeBot	a53db90e21	Revert "[inductor] consolidate common GEMM triton param retrieval (#158015 )" This reverts commit 9faef3d17c2e422d5d62f62b266155e2deb52c40. Reverted https://github.com/pytorch/pytorch/pull/158015 on behalf of https://github.com/henrylhtsang due to breaking tests ([comment](https://github.com/pytorch/pytorch/pull/158015#issuecomment-3115384824))	2025-07-25 00:16:50 +00:00
Ke Wen	9c10760662	[SymmMem] Use host/nvshmem_api.h for backward compat (#159061 ) Resolves #159045 `nvshmem_host.h` was introduced in 3.3.9. Use `host/nvshmem_api.h` and `host/nvshmemx_api.h` for prior versions. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159061 Approved by: https://github.com/ngimel, https://github.com/fduwjj, https://github.com/fegin	2025-07-24 22:56:26 +00:00
PyTorch MergeBot	8d2a1d6e18	Revert "Graph break with error message (#158800 )" This reverts commit cae4746952afbb6d26ecf7599cb7c6c449c69ef4. Reverted https://github.com/pytorch/pytorch/pull/158800 on behalf of https://github.com/clee2000 due to broke some tests on main inductor/test_distributed_patterns.py::DistributedPatternTests::test_nn_param_return4 [GH job link](https://github.com/pytorch/pytorch/actions/runs/16507837934/job/46685704688) [HUD commit link](`cae4746952`), note to self: bad TD, but also dynamo/test_repros failed but didn't get skipped by TD so maybe a landrace, or I just blaming the wrong commit entirely.. ([comment](https://github.com/pytorch/pytorch/pull/158800#issuecomment-3115224608))	2025-07-24 22:45:58 +00:00
PyTorch MergeBot	751285cb22	Revert "Move some of vec into headeronly in preparation for Half.h (#158976 )" This reverts commit 5564f2ca2e0836d75c4ee45899b1b981582c3e2d. Reverted https://github.com/pytorch/pytorch/pull/158976 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally. See D78924504 for details. To validate your fixes internally, you can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158976#issuecomment-3115198443))	2025-07-24 22:31:49 +00:00
Lucas Kabela	efc810c7d0	[Bugfix] Fix circular import between export and dynamo from tensor fn map (#158931 ) Fixes #158120 The issue was caused by populating a builtin tensor fn map at import time; if torch.export.export was called before any dynamo imports with the `meta` device, this map would not be populated, and so would populate on import time which would try to call `torch.disable`, which would not yet be initialized Fix is to populate this map lazily ``` python test/dynamo/imports_non_circular_repro.py TestImports.test_circular_import_with_export_meta ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158931 Approved by: https://github.com/StrongerXi, https://github.com/mlazos, https://github.com/anijain2305	2025-07-24 22:24:57 +00:00
Xu Han	abb0bf45df	[AOTI] skip crashed case on Windows temporary. (#158929 ) skip crashed case on Windows temporary. This case will crashed application: <img width="1053" height="275" alt="image" src="https://github.com/user-attachments/assets/3225e9c8-cbe7-4998-86da-f20fbb12ead2" /> Quick analysis: <img width="1400" height="261" alt="image" src="https://github.com/user-attachments/assets/9c21fefc-9ed8-40f2-84c5-edde2004777c" /> 1. It is crashed on OpenMP. 2. stack is dameged, need consider how to debug. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158929 Approved by: https://github.com/desertfire	2025-07-24 22:08:19 +00:00
PyTorch MergeBot	b533f12120	Revert "[Profiler] Fix lost C call events problem in Python 3.12.0-3.12.4 (#155446 )" This reverts commit da94023b0205bf98c3da366f2f86e0a443f4db17. Reverted https://github.com/pytorch/pytorch/pull/155446 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally. @sraikund16 can you please help validate the fix? (See D78845227 for details). You can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/155446#issuecomment-3115072504))	2025-07-24 21:46:00 +00:00
Aaron Orenstein	e20736bf1d	Dont't GC as often when collecting cudagraphs (#158193 ) TL;DR: Cuts vLLM cudagraph collection from 80s -> 24s Stop garbage collecting by default on every cudagraph recording. The old behavior can be re-enabled by setting `TORCH_CUDAGRAPH_GC=1` or the config `force_cudagraph_gc`. We were previously garbage collecting at the beginning of each cudagraph capture. vLLM collects 5427 graphs and most of those garbage collections weren't actually collecting any memory (CPU or GPU). This changes it to not collect more than every 10s so if we're capturing in a loop we don't burn all our cycles looking for garbage. (These number have a lot of variance from run to run but give the correct general scale) ``` \| calls \| total \| synchronize \| gcs \| collect \| empty cache \| sys freed \| cuda freed \| -------+-------+-------+-------------+------+---------+-------------+-----------+------------+ before \| 5427 \| 78s \| 1.48s \| 5427 \| 53.22s \| 1.21s \| 145855 \| 1539309568 \| -------+-------+-------+-------------+------+---------+-------------+-----------+------------+ after \| 5427 \| 24s \| 0s \| 3 \| 1.53s \| 0.84s \| 592 \| 1539309568 \| -------+-------+-------+-------------+------+---------+-------------+-----------+------------+ ``` total - this is the total time reported by vLLM's "Graph capturing finished" log. The rest of these are measured in torch.cuda.graphs.graph.__enter__(): calls - number of times torch.cuda.graphs.graph.__enter__ was called synchronize - this is the duration taken by the cuda.synchronize call gcs - number of times gc.collect was called collect - this is the duration taken by the gc.collect call empty cache - this is the duration taken by the torch.cuda.empty_cache call sys freed - the number of bytes reported freed by gc.collect cuda freed - the number of bytes reported freed by torch.cuda.memory_reserved So it seems like the heavy lifting is done by torch.cuda.empty_cache() which is fairly quick. Cudagraph results from the TorchInductor Performance DashBoard (this is from the original version using the GC clock so the real results will be slightly better than this): <img width="1494" height="382" alt="image" src="https://github.com/user-attachments/assets/69b705ef-47ce-4b6e-9733-1ec941cad93d" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158193 Approved by: https://github.com/ngimel	2025-07-24 21:37:11 +00:00
Lucas Kabela	cae4746952	Graph break with error message (#158800 ) Fixes #157452 Test with ``` python test/dynamo/test_repros.py ReproTests.test_nn_parameter_ctor_graph_breaks ``` ### Release Notes Change to nn.Parameter Constructor Behavior in Dynamo Semantic change introduced in the nn.Parameter constructor; previously, if the constructor lacked a clean source, the system would attempt to infer arguments to construct a clone and lift this synthetic proxy in the computation graph. This approach had many potential edge cases and was difficult to reason about. The new behavior defaults to graph breaking when the nn.Parameter constructor does not have a clean source. Users are now suggested to manually move the constructor out of the graph in such cases. This change improves clarity and reduces complexity in graph construction and debugging. Users can escape hatch to old semantics with `torch.dynamo.config.graph_break_on_nn_param_ctor=False` if this cannot be done. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158800 Approved by: https://github.com/anijain2305	2025-07-24 21:05:17 +00:00
Jithun Nair	4a13d4d7d0	[ROCm] Update jit_utils.cpp for compatibility with ROCm7.0 (#158868 ) Resolves error when running tests such as `test_nn.py::TestNN::test_L1Loss_no_reduce_complex_cuda` etc. on ROCm7.0: ``` /tmp/comgr-4cd8ad/input/CompileSourceU53Ndb:1016:7: error: no template named 'is_floating_point'; did you mean '__hip_internal::is_floating_point'? 1016 \| is_floating_point<_Tp>::value, \| ^~~~~~~~~~~~~~~~~ \| __hip_internal::is_floating_point /tmp/comgr-4cd8ad/include/hiprtc_runtime.h:1481:31: note: '__hip_internal::is_floating_point' declared here 1481 \| template<typename _Tp> struct is_floating_point : public false_type {}; \| ^ /tmp/comgr-4cd8ad/input/CompileSourceU53Ndb:1017:16: error: too few template arguments for class template '__libcpp_complex_overload_traits' 1017 \| typename __libcpp_complex_overload_traits<_Tp>::_ComplexType \| ^ /tmp/comgr-4cd8ad/input/CompileSourceU53Ndb:850:10: note: template is declared here 847 \| template <class _Tp, bool = is_integral<_Tp>::value, \| ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 848 \| bool = is_floating_point<_Tp>::value \| ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 849 \| > \| ~ 850 \| struct __libcpp_complex_overload_traits {}; \| ^ fatal error: too many errors emitted, stopping now [-ferror-limit=] 20 errors generated when compiling for gfx90a. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158868 Approved by: https://github.com/jeffdaily	2025-07-24 21:00:37 +00:00
Ti-Tai Wang	da35562bba	[ONNX] Filter out torchscript sentences (#158850 ) Fixes #157300 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158850 Approved by: https://github.com/justinchuby, https://github.com/svekars	2025-07-24 20:59:06 +00:00
Guilherme Leobas	de85ee73ae	Update context in `unimplemented_v2` when exception bubbles up to the interpreter (#158924 ) Before: ``` .Observed exception Explanation: Dynamo found no exception handler at the top-level compiled function when encountering an exception. Exception will propagate outside the compiled region. Hint: Dynamo has detected that tracing the code will result in an error when running in eager. Please double check that your code doesn't contain a similar error when actually running eager/uncompiled. Hint: It may be possible to write Dynamo tracing rules for this code. Please report an issue to PyTorch if you encounter this graph break often and it is causing performance issues. Developer debug context: ``` After: ``` Observed exception Explanation: Dynamo found no exception handler at the top-level compiled function when encountering an exception. Exception will propagate outside the compiled region. Hint: Dynamo has detected that tracing the code will result in an error when running in eager. Please double check that your code doesn't contain a similar error when actually running eager/uncompiled. Hint: It may be possible to write Dynamo tracing rules for this code. Please report an issue to PyTorch if you encounter this graph break often and it is causing performance issues. Developer debug context: raised exception TypeError([ConstantVariable(str: "unhashable type: <class 'torch._dynamo.variables.dicts.SetVariable'>")]) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158924 Approved by: https://github.com/williamwen42, https://github.com/zou3519	2025-07-24 20:50:22 +00:00
eqy	8573a2beda	[CUDA] Fix missing `__syncthreads` in MultiMarginLoss backward (#158994 ) Turns out issue in #158921 is detectable with a simple unit test and adding the missing sync fixes it Pull Request resolved: https://github.com/pytorch/pytorch/pull/158994 Approved by: https://github.com/malfet, https://github.com/Skylion007 Co-authored-by: Nikita Shulga <2453524+malfet@users.noreply.github.com>	2025-07-24 20:47:29 +00:00
PyTorch MergeBot	13398dab79	Revert "Remove tensorexpr tests (#158928 )" This reverts commit a3f9f79f591102afa93145bb67dc7e34df44f9a4. Reverted https://github.com/pytorch/pytorch/pull/158928 on behalf of https://github.com/clee2000 due to Theres still some references to the things removed in this PR in test.sh, the jobs on this PR are failing because of that but log classifier is probably pointing to a wrong line, should be an easy fix tho ([comment](https://github.com/pytorch/pytorch/pull/158928#issuecomment-3114873706))	2025-07-24 20:45:30 +00:00
Jane Xu	5564f2ca2e	Move some of vec into headeronly in preparation for Half.h (#158976 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158976 Approved by: https://github.com/albanD, https://github.com/desertfire	2025-07-24 20:32:33 +00:00
Sidharth	f3edcac23a	[dynamo] Added back weblink generation (#159011 ) Added back weblink generation for v2.9 development Note: It is fine to bring the weblink generation back since v2.9 isn't released for a while Pull Request resolved: https://github.com/pytorch/pytorch/pull/159011 Approved by: https://github.com/williamwen42	2025-07-24 20:27:11 +00:00
zhxchen17	90c241dedd	[precompile] Support user defined function calls from bytecode. (#158947 ) Previously precompile was implemented under the assumption that dynamo always inlines the user code and generate resume functions when a graph break is hit. In cases like nanogpt training, there exists nontrivial amount of code causing dynamo to fail the speculation and stop inlining certain type of user function. This results in more code objects to be tracked by CompilePackage. Since these new code objects are user defined, we need to also serialize the location of these code so that we can load the precompile entries to the these code objects in another process. With this fix, we are able to run nanogpt inference+training with precompile under torchbench. Differential Revision: [D78691422](https://our.internmc.facebook.com/intern/diff/D78691422/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158947 Approved by: https://github.com/jamesjwu	2025-07-24 20:10:57 +00:00
Luca Wehrstedt	5ab0eb28f7	Support DeepSeek-style blockwise scaling scaled-mm for fp8 on Hopper+ (#158037 ) cuBLAS added support for them in CUDA 12.9. It's rather easy to call into them, the hardest thing is allowing the lhs and rhs operands to have different scaling types, as that changes the whole callstack. The scaling format is still detected from the sizes of the scale tensors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158037 Approved by: https://github.com/eqy, https://github.com/drisspg	2025-07-24 20:10:51 +00:00
Laith Sakka	0b2ef76e85	DDE-Free select with unbacked index. (#157605 ) When select has data dependent input, we cant tell if the actual index shall be index+size or index. to avoid throwing dde, we allocate a new unbacked symbol to represent the storage offset of the output view and we compute its value dynamically at runtime when inductor is lowered. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157605 Approved by: https://github.com/ColinPeppler	2025-07-24 20:08:05 +00:00
Ruben Rodriguez Buchillon	9faef3d17c	[inductor] consolidate common GEMM triton param retrieval (#158015 ) \# Why - Make loop iteration simpler - Have a common spot where to make modifications that affect all the GEMM Triton templates, avoiding missed spots \# What - pull out commong logic of taking the BaseConfig objects and turning them into kwargs to feed into maybe_append_choice for Triton GEMM templates Differential Revision: [D78081314](https://our.internmc.facebook.com/intern/diff/D78081314) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158015 Approved by: https://github.com/PaulZhang12, https://github.com/jansel	2025-07-24 19:17:48 +00:00
namgyu-youn	aeaa20083f	[profiler] update CUDA runtime kernel identification logic (#157890 ) Update CUDA kernel detection to exclude memory API calls References: - https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__MEMORY.html - https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__EXECUTION.html Pull Request resolved: https://github.com/pytorch/pytorch/pull/157890 Approved by: https://github.com/sraikund16	2025-07-24 19:14:08 +00:00
zpcore	5be7e187ba	Support `sort` and `scatter_add` strategy (#159022 ) Add `sort`, `scatter_add` strategy. I am reusing the strategy for `scatter` related ops for a quick support. The strategy can be potential improved after we fix index related strategies. Minor fix: fix `replicate_op_strategy` to support output multiple tensors, which is required by aten.sort. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159022 Approved by: https://github.com/XilunWu, https://github.com/wconstab	2025-07-24 18:33:18 +00:00
Nikita Shulga	347a97da66	[MPS] Enable dlpack integration (#158888 ) Though testing is a lie and dependent on https://github.com/pytorch/pytorch/pull/153835 Fixes https://github.com/pytorch/pytorch/issues/153789 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158888 Approved by: https://github.com/albanD ghstack dependencies: #158874	2025-07-24 18:05:41 +00:00
Conan Truong	78aa3bd6b6	Added Emscripten __assert_fail declaration to Macros.h (#158580 ) Summary: __assert_fail is declared slightly differently in the Emscripten stdlib. This may cause errors when compiling with Emscripten. Test Plan: N/A Rollback Plan: Differential Revision: D78500790 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158580 Approved by: https://github.com/JacobSzwejbka	2025-07-24 17:10:29 +00:00
Jeff Daily	ee97dbf2e7	[ROCm][CI] update HIP patch for 6.4.1, again (#159001 ) Another fix for hipGraph capture of MIOpen OCL kernels. Pull Request resolved: https://github.com/pytorch/pytorch/pull/159001 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-24 16:36:19 +00:00
zeshengzong	b7d41729e0	Add `zerotensor` design description in code (#158837 ) Fix `TODO: add a note explaining the design decisions` by adding design description in https://github.com/pytorch/pytorch/issues/69687 to codebase. Make it easier to get and read by other developers. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158837 Approved by: https://github.com/soulitzer	2025-07-24 16:35:42 +00:00
Lucas Kabela	abcb24f4de	[Dynamo][Better Engineering] Add typing annotations to guard and source (#158397 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to a critical set of files for dynamo, `source.py` and the base `_guards.py` Running ``` mypy torch/_dynamo/source.py torch/_guards.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 1227 \| 2208 \| 55.57% \| 207 \| 362 \| 57.18% \| \| This PR \| 2217 \| 2217 \| 100.00% \| 362 \| 362 \| 100.00% \| \| Delta \| +990 \| +9 \| +44.43% \| +155 \| 0 \| +42.82% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158397 Approved by: https://github.com/anijain2305	2025-07-24 15:55:18 +00:00
fduwjj	fd48681b6a	[DeviceMesh][ez] Make the logic within flatten simpler (#158999 ) While looking at the code of device mesh I find that this logic can be simplified. Also the naming needs to be correct. Because this mesh is not "flattened" yet, so we can just call it flatten. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158999 Approved by: https://github.com/wz337, https://github.com/wconstab ghstack dependencies: #158900	2025-07-24 15:40:13 +00:00
cyy	a3f9f79f59	Remove tensorexpr tests (#158928 ) The tests are not maintained. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158928 Approved by: https://github.com/albanD	2025-07-24 15:38:36 +00:00
Ke Wen	2fc0b1605e	[a2av] Make test input more random (#157029 ) Stack from [ghstack](https://github.com/ezyang/ghstack) (oldest at bottom): Use torch.randn to fill input buffer. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157029 Approved by: https://github.com/fegin, https://github.com/ngimel ghstack dependencies: #158234, #158235, #156743, #156881, #157026	2025-07-24 15:35:12 +00:00
PyTorch MergeBot	11ea3736dd	Revert "[CI][testing] Use 3 processes for testing on sm89 and sm90 jobs (#158691 )" This reverts commit 0c0fcb53ff5ee1eb5f0d1f535ed3726d01f8abb5. Reverted https://github.com/pytorch/pytorch/pull/158691 on behalf of https://github.com/ZainRizvi due to Sorry but these are causing jobs to fail with out of memory errors on trunk ([comment](https://github.com/pytorch/pytorch/pull/158691#issuecomment-3113922186))	2025-07-24 15:31:53 +00:00
Ke Wen	43d4ff6851	[a2av] Test dispatch-then-combine (#157026 ) Stack from [ghstack](https://github.com/ezyang/ghstack) (oldest at bottom): Putting both the dispatch API and combine API in battlefield, one following the other, i.e. ``` all_to_all_vdev_2d(inp, out, inp_splits, out_splits_offsets, ...) all_to_all_vdev_2d_offset( input=out, out=combine_out, in_splits_offsets=out_splits_offsets, out_splits_offsets=combine_out_splits_offsets ) ``` Here the `out_splits_offsets` from dispatch perfectly serves as the `in_splits_offsets` argument for combine. Then we assert that the output of combine is exactly the same as the original input to shuffle, and combine's output splits are exactly the same as the original input splits. It works! Pull Request resolved: https://github.com/pytorch/pytorch/pull/157026 Approved by: https://github.com/Skylion007, https://github.com/ngimel ghstack dependencies: #158234, #158235, #156743, #156881	2025-07-24 15:21:02 +00:00
Ke Wen	83957d1c03	[a2av] Add token combine operator (#156881 ) Added `all_to_all_vdev_2d_offset`, which: Perform a 2D AllToAllv operation, with input split and offset information provided on device. The input offsets need not to be exact prefix sum of the input splits, i.e. paddings are allowed between the splitted chunks. The paddings, however, will not be transferred to peer ranks. In Mixure of Experts models, this operation can be used to combine tokens processed by experts on remote ranks. This operation can be viewed as an "reverse" operation to the `all_to_all_vdev_2d` operation (which shuffles tokens to experts). The change may seem a bit dense, sorry. But it is mainly two changes: 1. templating existing device functions (to use provided input offset or calculate it) 2. generalizing variable names, e.g. npes, ne --> minor_size, major_size, so that I can use the same alltoall function for matrix of (nranks, ne) as well as matrix of (ne, nranks). Pull Request resolved: https://github.com/pytorch/pytorch/pull/156881 Approved by: https://github.com/ngimel ghstack dependencies: #158234, #158235, #156743	2025-07-24 15:08:04 +00:00
Pian Pawakapan	48fe4ff247	[export] set enable_gqa in export flash->math decomp (#158604 ) Differential Revision: D78524147 For `scaled_dot_product_attention(..., enable_gqa=True)`: - the Math backend passes the flag through, performing the extra [KV broadcast](`6e07d6a0ff/aten/src/ATen/native/transformers/attention.cpp (L902)`) if set to True - the Flash backend has no flag, and relies on correct indexing in the C++ kernel - Export used to default to Math for `enable_gqa=True`, but https://github.com/pytorch/pytorch/pull/157893 landed and enabled Flash. At the same time, there's an export-only [decomp](`6e07d6a0ff/torch/_decomp/decompositions.py (L4968)`) redirecting flash -> math, calling with `enable_gqa` unset, because that info isn't available. This led to https://fb.workplace.com/groups/1028545332188949/posts/1264609398582540 crashing, calling the Math non-GQA variant, with GQA inputs. This assumes GQA for seqlen mismatches in the export decomp, setting `enable_gqa = <q seqlen> != <kv seqlen>`, relying on prior backend checks to raise on invalid input shapes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158604 Approved by: https://github.com/angelayi, https://github.com/drisspg	2025-07-24 14:46:13 +00:00
James Wu	f55c5d085e	[Precompile] Various small bugfixes, add CachingPrecompile to torchbench (#158847 ) This PR addresses a few small bugfixes needed to make NanoGPT inference work, and also adds a new `--caching-precompile` argument to torchbench. With `--caching-precompile`, after every benchmark we save precompile artifacts to DynamoCache, allowing us to test caching precompile on all existing benchmarks. The following bugfixes are in this PR to make all of this work: - Fix global variables being pruned with DUPLICATE_INPUT guards. DUPLICATE_INPUT guards have additional vars from the second input, which we track with additional_local_vars, but we never tracked additional global variables. This fixes the issue. (See torch/_dynamo/guards.py changes) - Return None from PRecompileContext.serialize() if no new dynamo compiles occurred. There's no reason to save artifacts (i.e. autotuning artifacts, etc) if no dynamo_compile occurred, so we return None early. We may later want to support editing existing dynamo artifacts as a TODO, but that's upcoming. - log `dynamo_start` on CompilePackage.load: This is only needed so that tlparse doesn't ignore TORCH_TRACE logs generated when caching precompile hits. If there are no actual compiles, we never log a "dynamo_start" entry, which makes internal tlparse ignore the TORCH_TRACE file. ## Test Plan After this PR, the following now works: ``` TORCH_LOGS=dynamo tlp python benchmarks/dynamo/torchbench.py --only nanogpt --performance --inference --backend inductor --caching-precompile --warm-start-latency ``` tlparse result (internal): Cold Start (6 seconds): https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmpAWe0zD/dedicated_log_torch_trace_vk9nkp4m.log/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 Warm Start (~1 s): https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmpAWe0zD/dedicated_log_torch_trace_5l4iwrpm.log/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 The 1 second of warm start here can be improved: the costs here are mostly in starting up workers and triton and initializing CUDA, a lot of which should not be included in the compile time cost in real world scenarios where these are already loaded before training begins. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158847 Approved by: https://github.com/zhxchen17	2025-07-24 14:09:54 +00:00
Nick Westlake	a3025e17b2	Fix inductor non-stable argsort/sort test (#146622 ) - Prevent the inductor test for argsort/sort from wrongly failing when the argsort/sort output with stable=False differs from pytorch but is still a valid argsort output. - Add functionality to allow alternative assert_equal functions in inductor tests for future cases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/146622 Approved by: https://github.com/eellison Co-authored-by: George Wigley <georgewi@graphcore.ai>	2025-07-24 14:02:12 +00:00
atalman	afd6eb0d49	[docker release] Remove build layer as not used (#158988 ) [docker release] Remove build layer as not used in any of the : https://hud.pytorch.org/hud/pytorch/pytorch/nightly/1?per_page=50&name_filter=Build%20Official Pull Request resolved: https://github.com/pytorch/pytorch/pull/158988 Approved by: https://github.com/oulgen, https://github.com/malfet	2025-07-24 12:22:55 +00:00
IvanKobzarev	3ced1079a4	[inductor] Fix collectives_reordering overwrite real_dep with fake_dep with the same name (#158960 ) Differential Revision: [D78839734](https://our.internmc.facebook.com/intern/diff/D78839734) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158960 Approved by: https://github.com/wconstab	2025-07-24 11:08:58 +00:00
Hameer Abbasi	3e954d3943	better testing for subclasses + compile (#158742 ) Fixes #114398 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158742 Approved by: https://github.com/ezyang	2025-07-24 10:28:44 +00:00
Sherlock Huang	fb067de550	[NativeRT] Remove device_ member from OpKernel base class (#158944 ) Summary: In general, device_ is not very useful in OpKernel. Remove it to avoid misuse. Also, the meaning of `device_` is also ambiguous in the OpKernel. For StaticDispatch kernels, we always call cpu kernel. For C10Kernel, we rely on input tensor's device and dispatcher to determine which device to run on. For ops involves multiple device, e.g. aten._to_copy(device), the meaning of device is ill-defined. Test Plan: CI Rollback Plan: Reviewed By: henryoier, dolpm, kqfu, zhxchen17 Differential Revision: D78704840 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158944 Approved by: https://github.com/dolpm	2025-07-24 09:21:37 +00:00
Wei (Will) Feng	693197eed6	[doc] remove FSDP1 developer note (#158991 ) this resolve pytorch doc audit - we remove fsdp1 doc and promote fsdp2 https://docs.pytorch.org/tutorials/intermediate/FSDP_tutorial.html Pull Request resolved: https://github.com/pytorch/pytorch/pull/158991 Approved by: https://github.com/svekars, https://github.com/mori360 ghstack dependencies: #158989	2025-07-24 08:21:54 +00:00
cyy	65c1109ca2	Remove CUDA 11 CMake code (#156795 ) CUDA 11 is no longer supported. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156795 Approved by: https://github.com/atalman, https://github.com/malfet	2025-07-24 08:00:41 +00:00
Ke Wen	70fb5bb6fb	[CI] Add smoke test for NVSHMEM availability (#158938 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158938 Approved by: https://github.com/huydhn, https://github.com/atalman	2025-07-24 06:34:21 +00:00
zero000064	30bb7636da	removed zero dim cpu logic from fake_tensor.py (#147501 ) Fixes #144748 In #144748, the inconsistency between the eager mode and the inductor mode is reported as a bug. The root cause is fake_tenosr.py's find-common-device method, `0b0da81021/torch/_subclasses/fake_tensor.py (L833)`, takes zero dim cpu tensor into account but the device check in adaption.h doesn't. This fix is to add a list for some ops to bypass zero-dim-cpu-tensor check to align with the eager mode. Pull Request resolved: https://github.com/pytorch/pytorch/pull/147501 Approved by: https://github.com/ezyang	2025-07-24 06:19:46 +00:00
Wei (Will) Feng	68349118b5	[doc] add weifengpy to torch distributed pocs (#158989 ) <img width="415" height="355" alt="Screenshot 2025-07-23 at 16 02 12" src="https://github.com/user-attachments/assets/35b6bb45-d5ed-4d74-8369-e8e66aaa2618" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158989 Approved by: https://github.com/mori360	2025-07-24 04:42:33 +00:00
PyTorch UpdateBot	e09d80c545	[vllm hash update] update the pinned vllm hash (#158997 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned vllm hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158997 Approved by: https://github.com/pytorchbot	2025-07-24 04:04:17 +00:00
Nikita Shulga	07df6ba7f5	[BE] Remove unused `test_python_gloo_with_tls` (#158964 ) This was last modified in 2021 and has not been invokved at least since 2.0 release Pull Request resolved: https://github.com/pytorch/pytorch/pull/158964 Approved by: https://github.com/Camyll, https://github.com/atalman ghstack dependencies: #158961, #158962, #158963	2025-07-24 02:34:27 +00:00
Nikita Shulga	d61153a300	Delete mobile merge rule (#158963 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158963 Approved by: https://github.com/atalman ghstack dependencies: #158961, #158962	2025-07-24 02:34:27 +00:00
Nikita Shulga	da9e120e3f	[BE] Remove unused `build-android` action (#158962 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158962 Approved by: https://github.com/Camyll, https://github.com/atalman ghstack dependencies: #158961	2025-07-24 02:34:27 +00:00
Nikita Shulga	611b61e758	[BE] Remove android build rules (#158961 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158961 Approved by: https://github.com/Camyll, https://github.com/atalman	2025-07-24 02:34:27 +00:00
cyy	d352c28dd1	[2/N] Remove FindPackageHandleStandardArgs.cmake (#156559 ) Following #157188, this PR removes FindPackageHandleStandardArgs.cmake Pull Request resolved: https://github.com/pytorch/pytorch/pull/156559 Approved by: https://github.com/albanD	2025-07-24 02:34:10 +00:00
Catherine Lee	0c0fcb53ff	[CI][testing] Use 3 processes for testing on sm89 and sm90 jobs (#158691 ) 3 procs were used for sm86, but we switched to sm89 and the check failed so it switched back to 2 sm90 is H100, but idk what unittests we have running there, but I assume they also have a lot of memory They use larger runners, which have more GPU memory, so its usually ok. I think it's ~22GB -> 10GB per proc if 2, 6GB per proc if 3 (cuda context maybe 1GB) I've applied skips to the ones that OOMed Time decreases from ~2.7hr per test job -> ~2hr Pull Request resolved: https://github.com/pytorch/pytorch/pull/158691 Approved by: https://github.com/huydhn	2025-07-24 01:51:28 +00:00
Teja	febf3c475e	fix forced loglevel in pytorch oss code (#158820 ) Differential Revision: [D78715806](https://our.internmc.facebook.com/intern/diff/D78715806/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158820 Approved by: https://github.com/Skylion007, https://github.com/pradeepfn	2025-07-24 00:40:28 +00:00
Aditya Tewari	7001d6fbc9	Skip slow tests for aarch64-inductor-benchmarks (#158842 ) This PR suggests adding some models to `cpu_skip_list` which are currently being run in TIMM and Torchbench. The suggested models takes a long time which leads to the benchmark runs being `timeout`. [benchmark runs for aarch64](https://github.com/pytorch/pytorch/actions/workflows/inductor-perf-test-nightly-aarch64.yml) • The issue stems from unoptimized groupwise convolution (BF16 /F16 dtype) kernels for aarch64 platforms , which significantly slow down execution leading to the timeout. Action: • An optimized BF16 groupwise convolution kernel is currently being developed in oneDNN, targeted for release in Q4 2025. To maintain dashboard consistency and signal clarity, I’ve skipped the affected tests in: * timm benchmarks * torchbench benchmarks As suggested, skip is applied at the CPU - arch level, explicitly branching for aarch64 and adding models which needs to be skipped. This keeps the logic clean, but: • An alternative considered was increasing shard counts for aarch64 runners, but given the known performance bottleneck, skipping avoids wasted compute cycles. Suggestions around this will be appreciated. Benchmark does not timeout after the suggested change: https://github.com/pytorch/pytorch/actions/runs/16447200138 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158842 Approved by: https://github.com/malfet	2025-07-24 00:21:38 +00:00
Bin Bao	0118931e27	[Inductor] Fix a user-defined Triton kernel bool param codegen issue (#158845 ) Summary: Fixes https://github.com/pytorch/pytorch/issues/158778. When handling a boolean type parameter to a user-defined Triton kernel, we need to treat it differently from integer. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158845 Approved by: https://github.com/davidberard98, https://github.com/eellison	2025-07-24 00:19:27 +00:00
atalman	ebb032a202	[docker release] Fix push nightly tag (#158984 ) This is a typo. I see that this step is not executing in nightly builds: https://github.com/pytorch/pytorch/actions/runs/16464544564/job/46538759844 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158984 Approved by: https://github.com/oulgen	2025-07-23 23:39:49 +00:00
Ke Wen	60ac3414eb	[a2av] Split in_out_splits into in_splits and out_splits_offsets (#156743 ) So that it would be easier if user would like to feed `out_splits_offsets` as input to a combining a2av (coming next). An example is in #157029. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156743 Approved by: https://github.com/ngimel ghstack dependencies: #158234, #158235	2025-07-23 23:34:48 +00:00
PyTorch MergeBot	d34cee4cf3	Revert "[Torch Native] Add test for packaging weight (#158750 )" This reverts commit 85ee2fb8c5c57b513526b0cc968ba13012167572. Reverted https://github.com/pytorch/pytorch/pull/158750 on behalf of https://github.com/ZainRizvi due to Sorry but this is failing on trunk: inductor/test_aot_inductor_package.py::TestAOTInductorPackageCpp_cuda::test_compile_with_exporter_weights [GH job link](https://github.com/pytorch/pytorch/actions/runs/16478978095/job/46590552109) [HUD commit link](`85ee2fb8c5`) ([comment](https://github.com/pytorch/pytorch/pull/158750#issuecomment-3111188266))	2025-07-23 23:24:55 +00:00
Anshul Sinha	5cdb3d896e	[FSDP][Replicate] added replicate function that uses FSDP instead of DDP (#158207 ) Summary Users would like to use Replicate with TP. Currently, the replicate function uses DDP, which has not been maintained resulting in a lack of integration options. Since users can use FSDP with TP, we will make the replicate function use FSDP so that users can use replicate with FSDP. To that end I have created a replicate function that uses FSDP instead of DDP. One blocker that I ran into is that the replicate function has a contract which assigns a module "replicate" attribute in registry. This would mean that fully_shards is_composable requirement would not be satisfied making it impossible to apply fully_shard to a replicate module. The solution to this was to copy the fully_shard function and state and modify it for replicate. In the future, it should be explored making the replicate_state inherit from FSDP_state to get rid of code duplicity. I have attached below the profile tracing of a replicated Net Module. https://interncache-all.fbcdn.net/manifold/perfetto-artifacts/tree/ui/index.html#!/?url=https://interncache-all.fbcdn.net/manifold/perfetto_internal_traces/tree/shared_trace/anshulsi_270fcc36-194a-42f5-9841-cace984c2132_devgpu263.prn2.facebook.com_1792146.1753232748025155780.pt.trace.json Test Case 1. pytest test/distributed/_composable/test_replicate_with_fsdp.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/158207 Approved by: https://github.com/weifengpy Co-authored-by: Anshul Sinha <50644008+sinhaanshul@users.noreply.github.com>	2025-07-23 22:53:06 +00:00
Guilherme Leobas	0204099762	Raise exception in Dynamo if op fails in the interpreter (#158661 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158661 Approved by: https://github.com/williamwen42 ghstack dependencies: #158660	2025-07-23 22:31:51 +00:00
Guilherme Leobas	b67f97c166	Correctly handle `OP_CONTAINS` (#158660 ) CPython can fallback to `__iter__` if object doesn't implement `__contains__` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158660 Approved by: https://github.com/zou3519	2025-07-23 22:31:51 +00:00
Mikayla Gawarecki	7f649ed4f8	Add basic torch.hash_tensor op (#154149 ) Added `torch.hash_tensor` reduction function with a `mode` argument that defaults to reduction with xor. - The hash is always uint64. - Integers will be casted to uint64 before performing the xor_sum reduction - Floats will be upcasted to double and then bitcasted to uint64 before performing the xor_sum reduction Pull Request resolved: https://github.com/pytorch/pytorch/pull/154149 Approved by: https://github.com/albanD	2025-07-23 22:28:03 +00:00
Max Ren	86df3ff1f1	fix xnnpack build on mac (#158881 ) Summary: Fix a bug for not getting the correct sources Test Plan: CI on my mac: ``` buck2 build @//fbobjc/mode/profile --show-full-output //xplat/executorch/examples/portable/executor_runner:executor_runner_opt File changed: fbsource//xplat/caffe2/third_party/xnnpack.buck.bzl Buck UI: https://www.internalfb.com/buck2/67b59179-4de8-462a-9202-0b9c34a35aef Network: Up: 2.4MiB Down: 1.3KiB (reSessionID-f687a7cd-5961-4851-bc67-b07043baa52a) Loading targets. Remaining 0/1 504 targets declared Analyzing targets. Remaining 0/42 1960 actions, 2424 artifacts declared Executing actions. Remaining 0/975 37.2s exec time total Command: build. Finished 40 local Time elapsed: 7.7s BUILD SUCCEEDED fbsource//xplat/executorch/examples/portable/executor_runner:executor_runner_opt /Users/maxren/fbsource/buck-out/v2/gen/fbsource/267ffdee31edf15e/xplat/executorch/examples/portable/executor_runner/__executor_runner_opt__/executor_runner_opt ``` Rollback Plan: Reviewed By: swolchok Differential Revision: D78771697 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158881 Approved by: https://github.com/digantdesai	2025-07-23 22:06:27 +00:00
fduwjj	82f8e04f27	Update distributed maintainers (#158900 ) I maintain couple components of distributed like devicemesh, c10d and PGNCCL, gloo, etc. Can I be marked not as emeritus? Thanks! Pull Request resolved: https://github.com/pytorch/pytorch/pull/158900 Approved by: https://github.com/albanD	2025-07-23 21:53:27 +00:00
saienduri	5619bf9971	Enable MI355X PyTorch CI testing. (#158889 ) This PR consists of all the changes required to enable PyTorch ROCm CI on MI355X nodes. - Rework aotriton cmake configuration to rely on `HIP_VERSION` instead of `ROCM_VERSION` as aotriton depnds on hip. Hip loosely track the rocm major version, but the two are not actually synchronized as observed in the ROCm 7 alpha build. - Bump composable-kernel submodule to [df6023e305f389bbf7249b0c4414e649f3ad6598](`df6023e305`) for mi350 compatibility. - Extend the change docker permissions step to the MI355x runners as well. This step is included to apply the required permission change to the test folder for a successful upload of artifacts in k8s docker. - Create new rocm-mi355 workflow to trigger core PyTorch tests on a nightly basis at 2:30 am PST. - Successfully tested running the test suites listed in rocm-mi355.yml on MI355 runners by temporarily hacking rocm-mi300.yml: `ca7d5fae11 (rocm-mi300)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158889 Approved by: https://github.com/jeffdaily	2025-07-23 21:50:31 +00:00
zpcore	d8425e9c75	[1/N] support of replication fallback strategy (#158046 ) #### 1. Provide a default fallback strategy that can apply to arbitrary operator with output in type of single tensor. We can call register_op_strategy to register using the `fallback_op_strategy`: - For op without List[Tensor] as input, call: ``` register_op_strategy(op_overload)(replicate_op_strategy) ``` - For op contains List[Tensor] as input, call: ``` register_op_strategy(op_overload, schema_info=RuntimeSchemaInfo(needs_pytree=True))(replicate_op_strategy) ``` The strategy will force all input and output to be replicated with the corresponding redistribute_cost. #### 2. Add a test function as a necessary condition for strategy function. ``` detect_exists_identical_opspec(*args, op, mesh, strategy_function) ``` This function detects if identical strategies will be produced given the sample `args`. It will iterate all combinations of placements for each arg and produce the output strategy from the registered `strategy_function`. #### 3. Provide a context manger `op_strategy_context` to easily register/unregister strategies for testing. E.g., ``` with op_strategy_context(test_op.default, replicate_op_strategy): ... ``` #### 4. Fix a bug that TupleStrategy never get flatten as expected: `9df0176408/torch/distributed/tensor/_op_schema.py (L286)` Basically we need to 1) register_pytree_node for TupleStrategy, 2) propagate the schema_info to `strategy_schema` after `strategy_schema = _wrap_with_op_strategy(op_schema)`. This is the first implementation. Plan to add support to enable sharding on the batch dim as the output strategy next. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158046 Approved by: https://github.com/wanchaol, https://github.com/wconstab	2025-07-23 21:14:20 +00:00
fduwjj	633d5faf3f	[DeviceMesh] Enable slicing a submesh with warnings (#158899 ) We don't create new PGs when doing slicing in DeviceMesh so it is relatively safe to relax the requirement of one can only do slicing from root mesh. But this does come with caveat when it is asymmetric, for example, only some have the sliced out submesh, for example. So aside from removing the requirement we also add a warning here. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158899 Approved by: https://github.com/wz337	2025-07-23 21:13:41 +00:00
Sidharth	4d5d56a30e	[dynamo] lintrunner for gb_registry adds/updates (#158460 ) This PR adds automation to adding/updating the JSON registry through the lintrunner. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158460 Approved by: https://github.com/williamwen42	2025-07-23 21:02:54 +00:00
Xuehai Pan	64e8d7d66b	[BE] bump test dependency `z3-solver` to drop using deprecated `pkg_resources` (#158905 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158905 Approved by: https://github.com/albanD, https://github.com/ezyang ghstack dependencies: #158904	2025-07-23 21:01:02 +00:00
Xuehai Pan	b935ad17d5	[BE][Easy] add missing Python 3.14 PyPI classifier (#158904 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158904 Approved by: https://github.com/albanD	2025-07-23 21:01:02 +00:00
henrylhtsang	f7f550649f	[cutlass backend] Change default inst level mm config number (#158901 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158901 Approved by: https://github.com/ColinPeppler, https://github.com/jingsh, https://github.com/Skylion007	2025-07-23 20:53:22 +00:00
PaliC	255c0545e7	[BE] Modify PyObjectSlot the assume only a single interpreter is in use (#158407 ) This PR makes some less risky changes to PyObjectSlot as there is a lot of stuff we do not need since there is only one interpreter. Specifically `check_interpreter` and `has_pyobj_nonhermetic` are removed Pull Request resolved: https://github.com/pytorch/pytorch/pull/158407 Approved by: https://github.com/albanD ghstack dependencies: #158288, #158290, #158291	2025-07-23 20:27:28 +00:00
PaliC	9c68c4d08f	[BE] Remove __reduce_deploy__ (#158291 ) This PR removes the integration point torch.fx had with torch::deploy (and another minor change). Note: This PR has some broken mypy errors, but I believe those should have been in the code base beforehand, and should be fixed in a separate PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/158291 Approved by: https://github.com/albanD ghstack dependencies: #158288, #158290	2025-07-23 20:27:28 +00:00
PaliC	6ed2cb6ccd	[BE] Remove torch deploy \| remove torch deploy specific files (#158290 ) This PR removes specific files found in pytorch which are only used for torch::deploy. This is mostly testing code and a debugger. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158290 Approved by: https://github.com/albanD ghstack dependencies: #158288	2025-07-23 20:27:28 +00:00
PaliC	ab26d4fbeb	[BE] remove torch deploy - conditionals (#158288 ) This PR is part of the work to deprecate torch::deploy in OSS. Effectively it does 3 things to get started. 1. Remove test_deploy_interaction as we no longer need to worry about this 2. Remove all torch._running_with_deploy checks and use the False path always (surfaced 1) 3. Remove `USE_DEPLOY` and switch to the default path always Note: MyPy does fail on a bunch of things here as a bunch of older files are touched. It may be better to fix these things on a separate PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/158288 Approved by: https://github.com/albanD	2025-07-23 20:27:28 +00:00
Denghui Dong	da94023b02	[Profiler] Fix lost C call events problem in Python 3.12.0-3.12.4 (#155446 ) Hi team, Please help review this patch. This PR https://github.com/pytorch/pytorch/pull/150370 tried to fix the "Empty C Call Queue" problem on Python 3.12. It added C calls for each starting Python event with a callable. I found the root cause is not that we cannot get C function frames by `PyFrame_GetBack` when PythonTracer is filling start frames, but the c call event loss problem bug on Python 3.12.0-3.12.4. And that problem was fixed by `257c413cd1` on 3.12.5. So I think the https://github.com/pytorch/pytorch/pull/150370 cannot fix the problem, this patch reverts the change of it. There are solutions to fix the problem correctly, such as we can add a new monitoring callback to compensate call events of methods with C function or we can override the callback registered by `PyEval_SetProfile`. These solutions may make the code hard to maintain. ~~Since upgrading the micro version of Python is not difficult for users, we can just ignore C functions and suggest user upgrade.~~ Pull Request resolved: https://github.com/pytorch/pytorch/pull/155446 Approved by: https://github.com/sraikund16, https://github.com/cyyever	2025-07-23 20:03:52 +00:00
Nichols A. Romero	c996aff6ed	[ROCm] UT verifies a runtime error is raised if tensor.item() is captured in a cudagraph (#158878 ) Unit test for this PR: https://github.com/pytorch/pytorch/pull/158165 This unit test verifies that a runtime error is raised when tensor.item() operation is captured in a cudagraph. Equally valid for ROCm and CUDA. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158878 Approved by: https://github.com/jeffdaily, https://github.com/ngimel	2025-07-23 20:01:50 +00:00
drisspg	691736ae07	Add kernel options to flex docs (#158875 ) Fixes https://github.com/pytorch/pytorch/issues/158741 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158875 Approved by: https://github.com/BoyuanFeng, https://github.com/albanD	2025-07-23 19:05:19 +00:00
PaulZhang12	fe8f556006	Fix Triton GEMM templates with k=1 (#158650 ) Thanks to @davidberard98 for much of the analysis here. For GEMMs of K=1, the hints, `tl.multiple_of` and `tl.max_contiguous` apply completely, as the indices to the loads are only dependent on `offs_m` and `offs_n`. For shapes like `(97x1), (1x97)`, this results in misaligned address errors, due to the fact that for all BLOCK_M and BLOCK_N sizes, the last tile is not a contiguous load. With K > 1 case, the hint is not as strict given the dependency on the k indices for the load as well. In the K=1 case, only `offs_m` and `offs_n` are used and broadcasted to the index shape. One can say these hints are "wrong", but in various cases in the hints being wrong, such as with the shape `9999x4, 4x9999`, there is a substantial performance improvement with the hint. For nice shapes with K=1, where M, N are a multiple 8 to where these hints are fine and there is no misaligned address, there is no performance regression observed on H100: <img width="547" height="402" alt="Screenshot 2025-07-18 at 5 05 47 PM" src="https://github.com/user-attachments/assets/fee2bbaa-784c-422e-bb8c-43c6c2607ad2" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158650 Approved by: https://github.com/davidberard98	2025-07-23 18:45:51 +00:00
Shangdi Yu	85ee2fb8c5	[Torch Native] Add test for packaging weight (#158750 ) Add test that require weights to be packaged for torch native For now, we need `package_weights_in_so=True` for compile standalone. The constants are in a `.o` file and will be added as a source to the CMakeLists.txt of the model. After we added weight deduping, we should be able to let this config be False. ``` python test/inductor/test_aot_inductor_package.py -k test_compile_with_exporter_weights ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158750 Approved by: https://github.com/desertfire	2025-07-23 18:36:10 +00:00
Mikayla Gawarecki	fef236da69	Add zero_() and empty_like(t) to torch/csrc/stable/ops.h (#158866 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158866 Approved by: https://github.com/janeyx99	2025-07-23 18:31:05 +00:00
PyTorch MergeBot	76be282e3a	Revert "[Precompile] Various small bugfixes, add CachingPrecompile to torchbench (#158847 )" This reverts commit d898d0d437bfdc0719e6c69d5005606c5e64fca8. Reverted https://github.com/pytorch/pytorch/pull/158847 on behalf of https://github.com/jithunnair-amd due to Broke ROCm CI jobs on MI200 and MI300 ([comment](https://github.com/pytorch/pytorch/pull/158847#issuecomment-3109664713))	2025-07-23 18:25:46 +00:00
PaulZhang12	9905ed616a	[Inductor] Expose decomposeK knobs as envvars (#158745 ) Fix up decomposeK autotuning, by removing condition to return more than `k_splits_limit` and setting default to 10 instead of 5. Allow `k_splits_limit` to be configurable to the user via `TORCHINDUCTOR_NUM_DECOMPOSE_K_SPLITS` and also allow user to configure threshold in which to use decompose_k via `TORCHINDUCTOR_DECOMPOSE_K_THRESHOLD` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158745 Approved by: https://github.com/eellison	2025-07-23 18:23:44 +00:00
PyTorch MergeBot	30b0ad5c68	Revert "Fix decorators skipping NCCL tests (#158846 )" This reverts commit 57024913c409764f129d6a7792625f5b05462e31. Reverted https://github.com/pytorch/pytorch/pull/158846 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking trunk. See distributed/_composable/fsdp/test_fully_shard_logging.py::LoggingTests::test_fsdp_logging [GH job link](https://github.com/pytorch/pytorch/actions/runs/16472103496/job/46564570609) [HUD commit link](`57024913c4`) ([comment](https://github.com/pytorch/pytorch/pull/158846#issuecomment-3109553414))	2025-07-23 17:47:35 +00:00
PyTorch MergeBot	41b6cdaf76	Revert "Fix Triton GEMM templates with k=1 (#158650 )" This reverts commit 9df0f565972a8a034fd77d65aff2c53e6e9856d1. Reverted https://github.com/pytorch/pytorch/pull/158650 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally, see D78805560 for details. To validate your fixes internally, you can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158650#issuecomment-3109538827))	2025-07-23 17:42:10 +00:00
Animesh Jain	1b456c580d	[dynamo][guards] Add type info of the guarded value in guard managers (#158765 ) tlparse looks like this <img width="1165" height="226" alt="image" src="https://github.com/user-attachments/assets/04c4e6b1-34a3-4d9d-8304-6eb6d9a94980" /> This will aid in reading guards. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158765 Approved by: https://github.com/Lucaskabela, https://github.com/StrongerXi	2025-07-23 16:59:15 +00:00
Xu Han	5e386eec94	[AOTI] enable aot inductor on Windows (#158915 ) With many PRs landed, we can run the first aot inductor example on Windows. <img width="640" height="427" alt="image" src="https://github.com/user-attachments/assets/131db159-ce17-4857-a3d5-a4b03638f01d" /> Let's remove the Windows check on `AotCodeCompiler`. CC: @angelayi , @desertfire , @jansel Pull Request resolved: https://github.com/pytorch/pytorch/pull/158915 Approved by: https://github.com/desertfire	2025-07-23 16:29:15 +00:00
iremyux	00da8e63eb	CI for Windows Arm64 (#148753 ) This pull request adds a new CI workflow for Windows Arm64, named win-arm64-build-test.yml. It can be triggered on any pull request by including the ciflow/win-arm64 tag. Pull Request resolved: https://github.com/pytorch/pytorch/pull/148753 Approved by: https://github.com/malfet	2025-07-23 16:12:20 +00:00
Guilherme Leobas	576253c476	[math] Trace `float.fromhex` (#156976 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156976 Approved by: https://github.com/zou3519 ghstack dependencies: #156975, #156977	2025-07-23 16:12:08 +00:00
Guilherme Leobas	f5314f89c8	[struct] Add `struct.pack` and `struct.unpack` polyfills (#156977 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156977 Approved by: https://github.com/XuehaiPan, https://github.com/jansel ghstack dependencies: #156975	2025-07-23 16:12:08 +00:00
Guilherme Leobas	671e22a951	[math] Raise exception in Dynamo if constant fold call fail (#156975 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156975 Approved by: https://github.com/zou3519	2025-07-23 16:12:08 +00:00
Mwiza Kunda	d3d9bc1c31	[inductor] Allow backends to register their own custom config object (#158254 ) An out of tree backend can have its own configuration options that the user can enable to control inductor compilation. These config options need to be taken into account when calculating the key that is used to determine cache miss / hits. This PR allows out of tree backends to specify a custom config module that has the same type as `torch._inductor.config` that can be used to control codegen (in addition to the default config), and will be used when creating the cache key. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158254 Approved by: https://github.com/eellison	2025-07-23 15:56:06 +00:00
angelayi	7d296d5c19	[aoti][mps] Enable more tests (#158703 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158703 Approved by: https://github.com/malfet, https://github.com/desertfire ghstack dependencies: #158349, #158350, #158351	2025-07-23 15:38:56 +00:00
Zhengxu Chen	2a60b8fc97	[export][ez] Fix packaging (#158855 ) Summary: as title, seems ytpo Test Plan: CI Rollback Plan: Differential Revision: D78758466 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158855 Approved by: https://github.com/henryoier	2025-07-23 15:36:14 +00:00
James Wu	d898d0d437	[Precompile] Various small bugfixes, add CachingPrecompile to torchbench (#158847 ) This PR addresses a few small bugfixes needed to make NanoGPT inference work, and also adds a new `--caching-precompile` argument to torchbench. With `--caching-precompile`, after every benchmark we save precompile artifacts to DynamoCache, allowing us to test caching precompile on all existing benchmarks. The following bugfixes are in this PR to make all of this work: - Fix global variables being pruned with DUPLICATE_INPUT guards. DUPLICATE_INPUT guards have additional vars from the second input, which we track with additional_local_vars, but we never tracked additional global variables. This fixes the issue. (See torch/_dynamo/guards.py changes) - Return None from PRecompileContext.serialize() if no new dynamo compiles occurred. There's no reason to save artifacts (i.e. autotuning artifacts, etc) if no dynamo_compile occurred, so we return None early. We may later want to support editing existing dynamo artifacts as a TODO, but that's upcoming. - log `dynamo_start` on CompilePackage.load: This is only needed so that tlparse doesn't ignore TORCH_TRACE logs generated when caching precompile hits. If there are no actual compiles, we never log a "dynamo_start" entry, which makes internal tlparse ignore the TORCH_TRACE file. ## Test Plan After this PR, the following now works: ``` TORCH_LOGS=dynamo tlp python benchmarks/dynamo/torchbench.py --only nanogpt --performance --inference --backend inductor --caching-precompile --warm-start-latency ``` tlparse result (internal): Cold Start (6 seconds): https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmpAWe0zD/dedicated_log_torch_trace_vk9nkp4m.log/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 Warm Start (~1 s): https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmpAWe0zD/dedicated_log_torch_trace_5l4iwrpm.log/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 The 1 second of warm start here can be improved: the costs here are mostly in starting up workers and triton and initializing CUDA, a lot of which should not be included in the compile time cost in real world scenarios where these are already loaded before training begins. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158847 Approved by: https://github.com/zhxchen17	2025-07-23 15:06:54 +00:00
Nikita Shulga	5998cd4eaa	[MPS] Speedup torch.full for 1-byte types (#158874 ) By using [`fillBuffer:range:value:`](https://developer.apple.com/documentation/metal/mtlblitcommandencoder/fillbuffer:range:value:?language=objc) rather than MPSGraph op, which should be faster and also does not have INT_MAX limit Which in turn fixes `test_index_put_accumulate_large_tensor_mps` test Pull Request resolved: https://github.com/pytorch/pytorch/pull/158874 Approved by: https://github.com/dcci	2025-07-23 14:00:40 +00:00
Alexander Grund	57024913c4	Fix decorators skipping NCCL tests (#158846 ) Avoid failures caused by tests exiting via sys.exit instead of `unittest.skip` In particular it will not try to start the test (causing forks into subprocess) just to stop them (killing the subprocess) which is done in the test setup Using `unittest.skip` decorators avoids the starting of the test in the first place. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158846 Approved by: https://github.com/Skylion007	2025-07-23 13:31:21 +00:00
yuchengliu1	ee72338f0c	[Inductor] MSVC use pointer when generating temporary array pointer (#158913 ) MSVC cannot implicitly convert a const iterator to a const pointer. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158913 Approved by: https://github.com/desertfire Co-authored-by: Xu Han <xu.han@outlook.com>	2025-07-23 13:19:11 +00:00
Han, Xu	c665594c1e	[AOTI] fix extract file failed on Windows. (#158702 ) Changes: 1. rename zip index filename, and keep it out of normalize path. 2. normalize output path for extract file. Extract files successful: <img width="683" height="247" alt="image" src="https://github.com/user-attachments/assets/72dff7b9-5ec0-4523-a6ee-7768b37bbe63" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158702 Approved by: https://github.com/angelayi	2025-07-23 08:00:14 +00:00
Ruben Rodriguez Buchillon	255a04baf1	[pt2 event logging] send autotuning data for strides and hinted shapes (#158852 ) Summary: # Why capture relevant data for offline lookup table generation # What report the hinted sizes not just the symbolic sizes Test Plan: ``` buck2 run mode/opt scripts/coconutruben/torchmm:experiment 2>&1 \| tee /tmp/epx040 ``` This only validates that this change does not break anything, as the schema is not on scuba yet (not actualized) Rollback Plan: Reviewed By: stashuk-olek Differential Revision: D77837548 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158852 Approved by: https://github.com/jingsh	2025-07-23 06:44:27 +00:00
Yang Wang	1d302eaee8	[vllm] add vllm test base docker image (#158755 ) # description Add base docker image for vllm. It seems like we use the base docker image for both pytorch build, and tests. Configure a base image for vllm against pytorch CI. # Others Added readme regarding how the base docker images are used, and how to add one, this also explain what is the right file to modify Pull Request resolved: https://github.com/pytorch/pytorch/pull/158755 Approved by: https://github.com/seemethere, https://github.com/huydhn	2025-07-23 05:42:44 +00:00
Colin Peppler	a6b7bea244	[inductor] support linear & layer_norm unbacked (#155267 ) ### What - Use `statically_known_true` over `guard_size_oblivious` in cases where we're checking an optimization path. Otherwise, it will DDE and we can't take the safe/slower path. - For broadcast checks, use `fallback=False` if we encounter a DDE. Typically, unbackeds would be ≥2 and that falls inline with size-oblivious reasoning (i.e. when `size_oblivious=True`). ### Example DDE ``` torch._inductor.exc.InductorError: LoweringException: GuardOnDataDependentSymNode: Could not guard on data-dependent expression Eq((u0//387), 1) (unhinted: Eq((u0//387), 1)). (Size-like symbols: u0) Caused by: (_inductor/lowering.py:488 in broadcast_symbolic_shapes) ``` ``` torch._inductor.exc.InductorError: LoweringException: GuardOnDataDependentSymNode: Could not guard on data-dependent expression Eq((u0//387), 1) (unhinted: Eq((u0//387), 1)). (Size-like symbols: u0) Caused by: (_inductor/ir.py:2797 in create) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155267 Approved by: https://github.com/eellison	2025-07-23 05:42:01 +00:00
PyTorch UpdateBot	be72bcf828	[vllm hash update] update the pinned vllm hash (#158806 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned vllm hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158806 Approved by: https://github.com/pytorchbot	2025-07-23 04:41:53 +00:00
PyTorch UpdateBot	f80f97d192	[audio hash update] update the pinned audio hash (#158807 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158807 Approved by: https://github.com/pytorchbot	2025-07-23 04:39:50 +00:00
anwang	42a69f7c2b	[MTIA Aten Backend] Migrate addmm.out / baddbmm.out / bmm.out (#158749 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate addmm.out / baddbmm.out / bmm.out to in-tree. Differential Revision: [D78578483](https://our.internmc.facebook.com/intern/diff/D78578483/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158749 Approved by: https://github.com/albanD, https://github.com/nautsimon ghstack dependencies: #158748	2025-07-23 03:45:28 +00:00
anwang	b87471e66f	[MTIA Aten Backend] Migrate addcdiv.out / addcmul.out / eq.Tensor_out / eq.Scalar_out (#158748 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate addcdiv.out / addcmul.out / eq.Tensor_out / eq.Scalar_out to in-tree. Differential Revision: [D78568103](https://our.internmc.facebook.com/intern/diff/D78568103/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158748 Approved by: https://github.com/albanD, https://github.com/nautsimon	2025-07-23 03:45:20 +00:00
Xu Han	f10e4430e2	[AOTI] normalize path and process model files. (#158705 ) Continued to https://github.com/pytorch/pytorch/pull/158702 , split `zip_filename_str` and real file path. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158705 Approved by: https://github.com/desertfire	2025-07-23 02:58:21 +00:00
Xu Han	2dccff7dcf	[inductor] pass_fds not supported on Windows, skip them on Windows. (#158830 ) <img width="1366" height="806" alt="image" src="https://github.com/user-attachments/assets/ddf3d27a-36da-47ce-9ba9-00c43805bb06" /> Almost UTs are failed on `AssertionError: pass_fds not supported on Windows.`, let's skip them on Windows. TODO: I will also debug and confirm `pass_fds` on Windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158830 Approved by: https://github.com/jansel	2025-07-23 02:24:35 +00:00
Pian Pawakapan	dec0d3101c	[export] fix unbacked range deserialization (#158681 ) Fixes https://github.com/pytorch/pytorch/issues/151809, by reading shape assertion nodes into ShapeEnv, and deferring instantiation of node example values, to be done node-by-node. Differential Revision: D78588406 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158681 Approved by: https://github.com/ydwu4, https://github.com/avikchaudhuri	2025-07-23 02:13:11 +00:00
PaulZhang12	9df0f56597	Fix Triton GEMM templates with k=1 (#158650 ) Thanks to @davidberard98 for much of the analysis here. For GEMMs of K=1, the hints, `tl.multiple_of` and `tl.max_contiguous` apply completely, as the indices to the loads are only dependent on `offs_m` and `offs_n`. For shapes like `(97x1), (1x97)`, this results in misaligned address errors, due to the fact that for all BLOCK_M and BLOCK_N sizes, the last tile is not a contiguous load. With K > 1 case, the hint is not as strict given the dependency on the k indices for the load as well. In the K=1 case, only `offs_m` and `offs_n` are used and broadcasted to the index shape. One can say these hints are "wrong", but in various cases in the hints being wrong, such as with the shape `9999x4, 4x9999`, there is a substantial performance improvement with the hint. For nice shapes with K=1, where M, N are a multiple 8 to where these hints are fine and there is no misaligned address, there is no performance regression observed on H100: <img width="547" height="402" alt="Screenshot 2025-07-18 at 5 05 47 PM" src="https://github.com/user-attachments/assets/fee2bbaa-784c-422e-bb8c-43c6c2607ad2" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158650 Approved by: https://github.com/davidberard98	2025-07-23 02:05:57 +00:00
albanD	91602a9254	Cleanup old caffe2 scripts (#158475 ) Testing on this one is grep based: if there were no reference to that script I can find, I deleted. We can easily add any of these back if needed! Pull Request resolved: https://github.com/pytorch/pytorch/pull/158475 Approved by: https://github.com/seemethere, https://github.com/huydhn, https://github.com/cyyever	2025-07-23 01:21:31 +00:00
angelayi	cc372ad557	[aoti][mps] Improve tabbing in cpp generation (#158351 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158351 Approved by: https://github.com/desertfire, https://github.com/malfet ghstack dependencies: #158349, #158350	2025-07-23 00:54:53 +00:00
angelayi	84058d1179	[aoti][mps] Fix cpu kernel generation (#158350 ) In the case where we have both mps and cpu code which can be inductor compiled, we need to case on the device -- this requires the device field to be correctly passed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158350 Approved by: https://github.com/malfet ghstack dependencies: #158349	2025-07-23 00:54:53 +00:00
angelayi	096dc35d77	[aoti][mps] Fix update constants buffer (#158349 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158349 Approved by: https://github.com/malfet	2025-07-23 00:54:52 +00:00
rzou	56d07d0bde	Add merge_rules category for Dynamo; add guilhermeleobas (#158620 ) Adds guilhermeleobas to merge_rules for Dynamo and functorch. Guilherme has done good work on both of these subsystems and I am tired of him approving my PRs and me not being able to merge them. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158620 Approved by: https://github.com/anijain2305	2025-07-23 00:44:27 +00:00
Pian Pawakapan	39b54b78d7	[export] runtime asserts for while HOP subgraphs (#158467 ) Differential Revision: D78431075 For #158366 - Calls runtime asserts pass for HOP subgraphs (in reenter_make_fx) - For while_loop only (can be expanded), clones input tensors for subgraph tracing, so unbacked memos (item, nonzero, etc.) aren't reused Pull Request resolved: https://github.com/pytorch/pytorch/pull/158467 Approved by: https://github.com/ydwu4	2025-07-23 00:34:18 +00:00
Nichols A. Romero	3703dabe42	[ROCm] delete un-needed workaround for tensor.item() (#158486 ) Deleting unused workaround per discussion here: https://github.com/pytorch/pytorch/pull/158165#discussion_r2207968880 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158486 Approved by: https://github.com/jeffdaily, https://github.com/houseroad	2025-07-23 00:31:57 +00:00
albanD	d3f9107d68	Remove top limit for cpython version and fix lint appropriately. (#158853 ) As per title. Sorry for the churn in the main commit. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158853 Approved by: https://github.com/seemethere, https://github.com/Skylion007, https://github.com/jingsh, https://github.com/malfet, https://github.com/ZainRizvi	2025-07-22 23:59:00 +00:00
Benjamin Glass	cab96b5879	[tests] Reduce sizes of unnecessarily large tensors to reduce OOM flakes (#158456 ) Downsizes several tensors that were massively oversized to test the problem at hand, to reduce test flaking. Fixes #126867 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158456 Approved by: https://github.com/desertfire	2025-07-22 23:41:48 +00:00
Xinya Zhang	6100ed457c	[ROCm] Improve Type Safety of C10_WARP_SIZE (#158271 ) # Background The `C10_WARP_SIZE`, although always be `32` on CUDA platform, varies across different AMD GPUs. Therefore, to correctly refer this value, the host code must be a variable instead of a literal defined by macro, or a `constexpr int`. This PR may cause more compiler errors for third party code on AMD GPU, which is intentional. Having a fixed `C10_WARP_SIZE` value on host code for AMD GPU only defers compile time error to runtime. This PR is recommended to be included as part of Release Notes to describe an API change for whoever uses this macro. Users are recommended to use `C10_WARP_SIZE` directly, which adapts for various scenarios, or define a macro to use `C10_WARP_SIZE`. Assignment of this macro to symbols shared by host/device code causes problems on ROCM platform. (See the fix at `aten/src/ATen/native/cuda/layer_norm_kernel.cu` for a concrete example) # Behaviors * If compiling with HIPCC (i.e `defined(__HIPCC__)`): + Define `C10_WARP_SIZE` to be non-`constexpr` `at::cuda::warp_size()` for host-compilation pass (as compared to `static constexpr int C10_WARP_SIZE = 1;` set in 04bd7e6850e8efec77994963ffee87549555b9c3) + Define `C10_WARP_SIZE` to be a function returning `constexpr int` `64` for `__GFX9__`, and `32` otherwise, for device-compilation pass - `__GFX8__` is also 64 but we do not support any GFX8 GPU. * If not compiling with HIPCC: + Define `C10_WARP_SIZE` to be non-constexpr `at::cuda::warp_size()` # `constexpr` variant for host code For host-compilation cases where a `constexpr` value is needed for warp size (eg. launch bounds), use `C10_WARP_SIZE_STATIC`, which is defined as `64`. This macro follows the pre 04bd7e6850e8efec77994963ffee87549555b9c3 behavior of `C10_WARP_SIZE` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158271 Approved by: https://github.com/jeffdaily Co-authored-by: Jithun Nair <37884920+jithunnair-amd@users.noreply.github.com>	2025-07-22 23:19:38 +00:00
PyTorch MergeBot	badfebf29e	Revert "[Inductor] Expose decomposeK knobs as envvars (#158745 )" This reverts commit eac777c4f46b381106f2f2b78fe05b506f8c558c. Reverted https://github.com/pytorch/pytorch/pull/158745 on behalf of https://github.com/jeffdaily due to sorry but rocm CI is broken due to this PR ([comment](https://github.com/pytorch/pytorch/pull/158745#issuecomment-3105071170))	2025-07-22 23:04:16 +00:00
Yahaya Suleiman	fc5a404eb1	[gtest][listing] fixing caffe2:verify_api_visibility - main (#158229 ) Summary: Remove the custom main from this test file Test Plan: https://www.internalfb.com/intern/testinfra/testrun/9570149303161031 Rollback Plan: Reviewed By: patskovn Differential Revision: D78015676 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158229 Approved by: https://github.com/Skylion007	2025-07-22 22:45:28 +00:00
AaronWang04	04a393507b	Fused RMSNorm implementation (#153666 ) Relevant #72643 Benchmarked versus unfused torch implementation and torch.compile implementation. Around 9x speedup vs unfused implementation on cuda and slightly faster vs inductor compile on 5090. ```py import torch import torch.nn as nn class RMSNorm(nn.Module): def __init__(self, dim, eps=1e-5): super().__init__() self.eps = eps self.scale = nn.Parameter(torch.ones(dim)) def forward(self, x): norm_x = x.norm(2, dim=-1, keepdim=True) rms_x = norm_x * torch.rsqrt(torch.tensor(x.shape[-1], dtype=x.dtype)) x_normed = x / (rms_x + self.eps) return self.scale * x_normed def benchmark_rmsnorm_cuda(input_shape, normalized_dim, num_iterations=100, warmup_iterations=10, dtype=torch.float16): rms_norm_layer = torch.nn.RMSNorm(normalized_dim, device='cuda', dtype=dtype) input_data = torch.randn(input_shape, device='cuda', dtype=dtype) for _ in range(warmup_iterations): _ = rms_norm_layer(input_data) torch.cuda.synchronize() start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() for _ in range(num_iterations): _ = rms_norm_layer(input_data) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = elapsed_time_ms / num_iterations print(f"--- RMSNorm CUDA Benchmark ---") print(f"Input Shape: {input_shape}") print(f"Normalized Dimension: {normalized_dim}") print(f"Benchmark Iterations: {num_iterations}") print(f"--- Fused Implementation ---") print(f"Average Time per Iteration: {avg_time_ms:.4f} ms") print(f"Total Time for {num_iterations} Iterations: {elapsed_time_ms:.3f} ms") compiled_rms_norm = torch.compile(RMSNorm(dim=normalized_dim)).cuda() for _ in range(warmup_iterations): _ = compiled_rms_norm(input_data) torch.cuda.synchronize() start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() for _ in range(num_iterations): _ = compiled_rms_norm(input_data) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = elapsed_time_ms / num_iterations print(f"--- TorchCompile Implementation ---") print(f"Average Time per Iteration: {avg_time_ms:.4f} ms") print(f"Total Time for {num_iterations} Iterations: {elapsed_time_ms:.3f} ms") print("-" * 50) if __name__ == '__main__': parameter_sets = [ {'batch_size': 16, 'sequence_length': 256, 'hidden_features': 512, 'dtype': torch.float16}, {'batch_size': 32, 'sequence_length': 512, 'hidden_features': 768, 'dtype': torch.float16}, {'batch_size': 64, 'sequence_length': 1024, 'hidden_features': 1024, 'dtype': torch.float16}, {'batch_size': 32, 'sequence_length': 512, 'hidden_features': 768, 'dtype': torch.float32}, {'batch_size': 8, 'sequence_length': 2048, 'hidden_features': 2048, 'dtype': torch.float16}, ] num_benchmark_iterations = 200 num_warmup_iterations = 20 for params in parameter_sets: batch_size = params['batch_size'] sequence_length = params['sequence_length'] hidden_features = params['hidden_features'] data_type = params.get('dtype', torch.float16) shape = (batch_size, sequence_length, hidden_features) norm_dim_to_normalize = hidden_features print(f"Benchmarking with: BS={batch_size}, SeqLen={sequence_length}, Hidden={hidden_features}, DType={data_type}") benchmark_rmsnorm_cuda(input_shape=shape, normalized_dim=norm_dim_to_normalize, num_iterations=num_benchmark_iterations, warmup_iterations=num_warmup_iterations, dtype=data_type) ``` Here are the triton compile tests ran on a 5090 (comparing this branch vs main) ```py import torch import torch.nn as nn from torch._inductor.utils import run_and_get_code, run_fw_bw_and_get_code torch.manual_seed(0) device = torch.device("cuda") for batch in range(0, 9): for i in range(9, 16): normalized_shape_arg = (2batch, 2i) input_tensor = torch.randn(2batch, 2i, device=device, requires_grad=True) weight_tensor = torch.randn(2batch, 2i,device=device, requires_grad=True) model = torch.nn.functional.rms_norm compiled_model = torch.compile(model) loss = torch.randn_like(input_tensor) num_iter = 5 for j in range(num_iter): output = compiled_model(input_tensor, normalized_shape_arg, weight_tensor) output.backward(loss) start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() num_iter = 10 for j in range(num_iter): output = compiled_model(input_tensor, normalized_shape_arg, weight_tensor) output.backward(loss) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = round(elapsed_time_ms / num_iter, 5) print(2batch, 2i, avg_time_ms) ``` main ``` 32 512 0.1812 32 1024 0.19021 32 2048 0.18871 32 4096 0.17019 32 8192 0.21944 32 16384 0.38871 32 32768 0.83282 64 512 0.14705 64 1024 0.13987 64 2048 0.14111 64 4096 0.21699 64 8192 0.43141 64 16384 0.90652 64 32768 2.18573 128 512 0.19361 128 1024 0.1963 128 2048 0.20122 128 4096 0.38888 128 8192 0.93795 128 16384 2.23437 128 32768 5.50079 256 512 0.16722 256 1024 0.22856 256 2048 0.39421 256 4096 0.96621 256 8192 2.48746 256 16384 5.53571 256 32768 11.97932 ``` current branch ``` 32 512 0.16328 32 1024 0.18104 32 2048 0.15508 32 4096 0.14356 32 8192 0.20111 32 16384 0.45974 32 32768 0.94799 64 512 0.16874 64 1024 0.18701 64 2048 0.16107 64 4096 0.20152 64 8192 0.46568 64 16384 0.96599 64 32768 2.21661 128 512 0.14982 128 1024 0.15565 128 2048 0.22241 128 4096 0.46128 128 8192 0.88883 128 16384 2.3097 128 32768 5.84448 256 512 0.14346 256 1024 0.2007 256 2048 0.45927 256 4096 0.87876 256 8192 2.10571 256 16384 5.73948 256 32768 12.98581 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153666 Approved by: https://github.com/ngimel, https://github.com/albanD	2025-07-22 22:25:44 +00:00
Xu Han	a626dc8f16	[AOTI] windows package load dev (#158671 ) changes: 1. add extract file fail handler for Windows develop. 2. normalize more file paths. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158671 Approved by: https://github.com/angelayi, https://github.com/desertfire	2025-07-22 21:35:57 +00:00
Panagiotis Kourdis	fd47401536	[doc] Updates to distributed.md for XCCL backend (#155834 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155834 Approved by: https://github.com/guangyey, https://github.com/AlannaBurke, https://github.com/d4l3k Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-07-22 21:01:43 +00:00
Sandeep Narendranath Karjala	e44e05f7ae	[dynamo] Move skipIf decorator to class level in test_fx_graph_runnable (#157594 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157594 Approved by: https://github.com/xmfan ghstack dependencies: #157162	2025-07-22 20:41:49 +00:00
Nikita Shulga	ddd74d10fc	More fixes to `MakeTensor::computeStorageSize()` (#158813 ) Followup after https://github.com/pytorch/pytorch/pull/158690 that fixessimilar logic if `strides` are not explicitly specified Expanded testing to cover both cases Pull Request resolved: https://github.com/pytorch/pytorch/pull/158813 Approved by: https://github.com/ZainRizvi, https://github.com/Skylion007, https://github.com/albanD ghstack dependencies: #158690	2025-07-22 20:36:12 +00:00
Xinya Zhang	823e223893	[ROCm] logsumexp on ROCm needs scaling back to natural base. (#156903 ) Fixes #156012 This is a temporary solution that makes context parallelism working before logsumexp behavior changes landed in AOTriton. After discussion we are not going to release AOTriton 0.10.1 to fix this due to * Even if the interface is not changed, changing the behavior of returned logsumexp tensor should still be considered as an ABI break. Such changes do not fall into the "ABI compatible" category and should be postponed to next release. * AOTriton 0.11 is scheduled to be released before end of July, which is less than five weeks Pull Request resolved: https://github.com/pytorch/pytorch/pull/156903 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-22 20:32:34 +00:00
fduwjj	6499420e45	[DeviceMesh] Make the repr shorter when debug ENV not set (#158822 ) Users want a shorter repr so this PR is trying to address that when TORCH_DISTRIBUTED_DEBUG is not set to DETAIL. Feedback and discussion is welcomed. Somehow I found that torch.set_printoptions is global, so I am hesitated to use it. Now the print is like <img width="435" height="79" alt="image" src="https://github.com/user-attachments/assets/8f173287-7138-4fbe-a4a3-8483523b21e4" /> or <img width="485" height="104" alt="image" src="https://github.com/user-attachments/assets/21e34db9-56b5-47e2-9767-750d6105a273" /> or <img width="675" height="97" alt="image" src="https://github.com/user-attachments/assets/53aa763e-7edd-4622-9cdb-37e2af8ec11f" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158822 Approved by: https://github.com/wz337, https://github.com/wconstab, https://github.com/xmfan	2025-07-22 20:31:44 +00:00
Goswami, Subrata	e17538022a	Making input dynamically adjust. (#157324 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/157324 Approved by: https://github.com/Skylion007, https://github.com/d4l3k	2025-07-22 20:14:05 +00:00
Goswami, Subrata	37ded2ac90	Using torch.accelerator in comm_mode_features_example.py and visualize_sharding_example.py (#157317 ) Continuation of https://github.com/pytorch/pytorch/pull/153213 . @guangyey @kwen2501 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157317 Approved by: https://github.com/guangyey, https://github.com/EikanWang, https://github.com/d4l3k Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-07-22 19:58:48 +00:00
Justin Chu	767791943d	[ONNX] Set default opset to 20 (#158802 ) Bump default opset to 20, which is a newer opset and the max torchscript exporter supports. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158802 Approved by: https://github.com/titaiwangms	2025-07-22 19:55:05 +00:00
Nichols A. Romero	c917c63282	[ROCm][tunableop] UT tolerance increase for matmul_small_brute_force_tunableop at FP16 (#158788 ) TunableOp will sometimes find a less precise solution due to the small input vectors used in this UT. Bumping op tolerance to eliminate flakiness. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158788 Approved by: https://github.com/jeffdaily	2025-07-22 19:45:35 +00:00
Zain Rizvi	659bfbf443	Revert "We do support 3.14" (#158856 ) Reverting to fix lint This reverts commit 2a249f1967d29626fe6ac6a07f28440348d1cc93. An emergency fix since the change needed to fix this is a little more complex than expected (see https://github.com/pytorch/pytorch/pull/158853 for reference) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158856 Approved by: https://github.com/Camyll, https://github.com/atalman	2025-07-22 19:40:53 +00:00
Electron4444	832ab990c9	Use init_device_mesh API for select tests where possible (#158675 ) This addresses reviews made for: #158538 #108749 It interchanged all the specific DevideMesh constructor calls with the API provided by the test cases, to improve abstraction Pull Request resolved: https://github.com/pytorch/pytorch/pull/158675 Approved by: https://github.com/wconstab	2025-07-22 19:28:42 +00:00
Jay Tsou	56df025d51	Add caching for `_rename_without_collisions` (#158594 ) Fixes #158357 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158594 Approved by: https://github.com/pianpwk	2025-07-22 19:19:13 +00:00
Eddie Yan	55ff4f85e9	[FP8][CUTLASS] xFail `honor_sm_carveout` on `sm100` (#152378 ) CUTLASS only supports SM carveout via green contexts on `sm100` Pull Request resolved: https://github.com/pytorch/pytorch/pull/152378 Approved by: https://github.com/Skylion007, https://github.com/albanD, https://github.com/nWEIdia	2025-07-22 18:39:50 +00:00
William Wen	7d2ceaff21	[dynamo] skip tracing functions registered in sys.monitoring (#158171 ) Fixes https://github.com/pytorch/pytorch/issues/158164 This was fixed by applying `skip_code_recursive` to any function registered to `sys.monitoring` (via `PyThreadState_GET()->interp->monitoring_callables`). This check is done whenever we attempt to set the eval frame callback from Python. Microbenchmark: `benchmarks/dynamo/microbenchmarks/overheads.py`: BEFORE: ``` requires_grad=False eager 7.1us (warmup=0.0s) compiled 24.6us (warmup=10.0s) requires_grad=True eager 8.9us (warmup=0.0s) compiled 57.8us (warmup=0.1s) inference_mode() eager 6.5us (warmup=0.0s) compiled 23.4us (warmup=0.1s) ``` AFTER: ``` requires_grad=False eager 7.0us (warmup=0.0s) compiled 23.2us (warmup=15.2s) requires_grad=True eager 9.0us (warmup=0.0s) compiled 55.1us (warmup=0.1s) inference_mode() eager 6.4us (warmup=0.0s) compiled 22.2us (warmup=0.1s) ``` Followup thought: how do we let users know that a frame is skipped because the code object is a callable registered to sys.monitoring? (or any other reason?) Differential Revision: [D78530528](https://our.internmc.facebook.com/intern/diff/D78530528) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158171 Approved by: https://github.com/jansel	2025-07-22 18:02:30 +00:00
albanD	2a249f1967	We do support 3.14 This has been added a bit back.	2025-07-22 10:40:18 -07:00
Yidi Wu	52c294008e	[hop] allow non fake inputs when check input alias and mutation (#158798 ) https://github.com/pytorch/pytorch/pull/154193 gets reverted due to a test failure. The root cause being that: an executorch pass turns int inputs into a scalar tensor in cond's subgraph. The pass have been around on the critical path of executorch since two years ago. Changing it would be difficult. So we just allow non-fake inputs for check input mutation and aliasing, which shoudn't affect the correctness of the analysis. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158798 Approved by: https://github.com/pianpwk	2025-07-22 17:22:37 +00:00
Alexander Novikov	0971637c11	Fix torch.tensor warning in ONNX symbolic_opset10 export (#158835 ) Fix PyTorch tensor copying warning in ONNX export ## Problem PyTorch ONNX exporter was generating a warning about incorrect tensor copying method: ``` UserWarning: To copy construct from a tensor, it is recommended to use sourceTensor.clone().detach() or sourceTensor.clone().detach().requires_grad_(True), rather than torch.tensor(sourceTensor). ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158835 Approved by: https://github.com/justinchuby	2025-07-22 16:32:49 +00:00
PyTorch MergeBot	7d6f340238	Revert "[AOTI] Add more default options to compile_standalone (#158560 )" This reverts commit a991e285ae35159680b0ad4be24669906a6fa256. Reverted https://github.com/pytorch/pytorch/pull/158560 on behalf of https://github.com/jeffdaily due to broke rocm CI, no test signal was available from rocm ciflow/trunk, need to add ciflow/rocm to reland ([comment](https://github.com/pytorch/pytorch/pull/158560#issuecomment-3103633964))	2025-07-22 16:20:17 +00:00
Benjamin Glass	4060f30042	[AOTI] Convert C-struct zip handling to RAII container (#158687 ) Attempts to fix a memory leak reported in #158614 by wrapping manually managed MiniZ C-structs in an RAII container. I have been unable to reproduce the reported leak, but this seems like the most likely candidate. Fixes #158614 (hopefully) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158687 Approved by: https://github.com/desertfire	2025-07-22 16:01:51 +00:00
PyTorch MergeBot	9a28e23d97	Revert "removed zero dim cpu logic from fake_tensor.py (#147501 )" This reverts commit 9e0473b56621162bd85e94943a516be4727e5651. Reverted https://github.com/pytorch/pytorch/pull/147501 on behalf of https://github.com/ZainRizvi due to Seems to have broken ROCm. See inductor/test_aot_inductor_package.py::TestAOTInductorPackageCpp_cuda::test_compile_standalone_cos [GH job link](https://github.com/pytorch/pytorch/actions/runs/16428359564/job/46426243808) [HUD commit link](`a991e285ae`) ([comment](https://github.com/pytorch/pytorch/pull/147501#issuecomment-3103494041))	2025-07-22 15:45:34 +00:00
Nikita Shulga	d0c00d9a69	[MPS] Do not crash if tensor dim > INT_MAX (#158824 ) Looks like all MPS operations will crash if one of tensor dimentions are greater than `231-1` Change it into a structured exception, by checking tensor size before attempting to create MPS Tensor Add regression test for it. Before this change running following will abort with exception ``` % python3 -c "import torch; torch.randint(0, 10, (231,), dtype=torch.uint8, device='mps')" /AppleInternal/Library/BuildRoots/1c8f7852-1ca9-11f0-b28b-226177e5bb69/Library/Caches/com.apple.xbs/Sources/MetalPerformanceShaders/MPSCore/Types/MPSNDArray.mm:829: failed assertion `[MPSNDArray initWithDevice:descriptor:isTextureBacked:] Error: NDArray dimension length > INT_MAX' zsh: abort python3 -c· ``` Skip the test on MacOS-13, as it crashes somewhere deep in MPSGraph framework with ``` /AppleInternal/Library/BuildRoots/c651a45f-806e-11ed-a221-7ef33c48bc85/Library/Caches/com.apple.xbs/Sources/MetalPerformanceShaders/MPSCore/Types/MPSNDArray.mm:724: failed assertion `[MPSTemporaryNDArray initWithDevice:descriptor:] Error: total bytes of NDArray > 2**32' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158824 Approved by: https://github.com/dcci ghstack dependencies: #158690, #158823	2025-07-22 15:12:26 +00:00
IvanKobzarev	371ffaf415	[bucketing] Support case of several pgs in graph (#158632 ) Main changes: - bucketing collectives only from the same process_group by group_name - Support of groups like [0,2,4,6], [0,1,3,5] using `rank_idx_dict` for in pass operations for slice idxs etc. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158632 Approved by: https://github.com/wconstab	2025-07-22 14:50:39 +00:00
James Wu	1b772de397	Still run TritonBundler with BundledAOTAutogradCache, save autotune results (#158048 ) When running BundledAOTAutogradCache with precompile, we still need to run triton bundling so that the precompiled CompiledFxGraph has triton cuda kernels. We also pre save the autotune results in the precompile artifact. It would be even better to pre trim the cuda kernels on save and apply them, which we can work on later. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158048 Approved by: https://github.com/zhxchen17	2025-07-22 14:12:21 +00:00
Nikita Shulga	8e99714204	[EZ][BE][MPS] Remove unused `ndArrayFromTensor` (#158823 ) And `printTensorNDArray`, both of which according to https://github.com/search?type=code&q=ndArrayFromTensor+org%3Apytorch are not used anywhere Pull Request resolved: https://github.com/pytorch/pytorch/pull/158823 Approved by: https://github.com/dcci ghstack dependencies: #158690	2025-07-22 14:06:42 +00:00
Animesh Jain	9b4d938f04	[dynamo][fsdp] Consistent behavior of int attributes (#157262 ) Reimpl of https://github.com/pytorch/pytorch/pull/150954 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157262 Approved by: https://github.com/bdhirsh	2025-07-22 11:26:54 +00:00
PyTorch MergeBot	0142d5f4e2	Revert "Remove is_arvr_mode() from xnnpack.buck.bzl (#158682 )" This reverts commit f09a484b8164aaadd57a79354f0ccf47733f365e. Reverted https://github.com/pytorch/pytorch/pull/158682 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/158682#issuecomment-3101648365))	2025-07-22 08:33:08 +00:00
Nichols A. Romero	91b69deeb0	[ROCm][CI] update fbgemm_gpu hash used by inductor tests (#158602 ) fbgemm_gpu build started failing with asmjit errors. Moving to latest tip of fbgemm for inductor tests resolves the build failures. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158602 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-22 08:04:59 +00:00
Shangdi Yu	392fa75411	Change from import trace to import config (#158796 ) Summary: for this particular instance, we're doing from torch._inductor.config import trace ...trace.provenance_tracking... but for all other call sites, we're doing from torch._inductor import config ... config.trace.provenance_tracking.... Test Plan: CI Rollback Plan: Differential Revision: D78699876 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158796 Approved by: https://github.com/c00w	2025-07-22 06:10:38 +00:00
Junjie Wang (PyTorch)	3a67bf9c62	[PGNCCLx] Bring split and merge for PGNCCLx (#158790 ) Summary: We added group split in D78300794 and remote_group_merge in D78450094. We first want to upstream this change to PGNCCLx as well so that NCCLx can use this new API and we can continue our c10d clean up in https://github.com/pytorch/pytorch/pull/158488. Test Plan: CI ``` buck test -c hpc_comms.use_ncclx=stable comms/ncclx/pg/tests:test_c10d_ncclx -- test_group_split_and_merge ``` Rollback Plan: Differential Revision: D78521060 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158790 Approved by: https://github.com/d4l3k	2025-07-22 06:05:00 +00:00
henrylhtsang	d984143a74	[ci][cutlass backend] Add ci for cutlass backend tests (#156626 ) redo of https://github.com/pytorch/pytorch/pull/156136 Differential Revision: [D77327309](https://our.internmc.facebook.com/intern/diff/D77327309) I want to try land the full version first. If the ci is taking too long, we can revert back to only testing for a few names. ``` -k 'test_max_autotune_cutlass_backend_regular_mm and not test_max_autotune_cutlass_backend_regular_mm_streamk' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156626 Approved by: https://github.com/huydhn, https://github.com/mlazos	2025-07-22 05:18:13 +00:00
Shangdi Yu	21c97bd565	[reland] Transfer "stack_trace" in post_grad passes (#158752 ) Summary: We transfer stack trace in post_grad passes. We shouldn't add "stack_trace" to _COPY_META_FIELDS because _COPY_META_FIELDS is used in proxy.py where stack_trace is explicitly set. Since the stack_trace is being used by more and more debugging tools, we should also start testing it more rigorously. This PR start by adding a first test for testing that stack trace is preserved through post_grad_passes. Test Plan: ``` buck run mode/dev-nosan fbcode//caffe2/test/inductor:provenance_tracing -- -r test_pattern_matcher_transfer_meta buck run mode/dev-nosan fbcode//caffe2/test/inductor:auto_functionalize -- --rcaffe2/test/inductor:auto_functionalize_old ``` Rollback Plan: Differential Revision: D78669729 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158752 Approved by: https://github.com/jingsh	2025-07-22 03:49:13 +00:00
Boyuan Feng	a155f742ad	[benchmark] allow default mode for compile (#158792 ) Allow default mode for compile when users cannot run "max-autotune-no-cudagraphs" due to compilation time. Overall, "default" mode is slower than "[max-autotune-no-cudagraphs](https://github.com/pytorch/pytorch/pull/158536)" depending on input shapes. <img width="3564" height="2368" alt="CrossEntropyBackward_bench" src="https://github.com/user-attachments/assets/5d25c0e4-6714-42bb-a544-b7ef9cbc1b17" /> <img width="3564" height="2368" alt="CrossEntropyForward_bench" src="https://github.com/user-attachments/assets/40e0bbf9-657f-48f2-ac0c-1f0fd6a0ac1d" /> <img width="3564" height="2368" alt="LayerNormBackward_bench" src="https://github.com/user-attachments/assets/db582bb2-d8d4-414a-9de7-b9af061ad0cd" /> <img width="3564" height="2368" alt="LayerNormForward_bench" src="https://github.com/user-attachments/assets/2ce18bd8-73fc-434a-820f-46aa9ad9ddce" /> <img width="3564" height="2368" alt="RMSNormBackward_bench" src="https://github.com/user-attachments/assets/f4cb5f4b-93d3-4d96-973f-37643912325a" /> <img width="3564" height="2368" alt="RMSNormForward_bench" src="https://github.com/user-attachments/assets/231c5805-b156-4587-9c5f-504a33b60883" /> <img width="3564" height="2368" alt="SoftmaxBackward_bench" src="https://github.com/user-attachments/assets/f651c578-813b-4a8e-bffc-b5b34bd879fc" /> <img width="3564" height="2368" alt="SoftmaxForward_bench" src="https://github.com/user-attachments/assets/bfdcc043-4370-4355-af84-9f463426b21a" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158792 Approved by: https://github.com/zou3519	2025-07-22 03:07:22 +00:00
cyy	3639d29ea1	Fix warnings of unused-variable (#158627 ) Fixes ``` /var/lib/jenkins/workspace/test/cpp/tensorexpr/test_kernel.cpp:42:22: error: unused variable 'verification_pattern' [-Werror,-Wunused-variable] ``` and also extra semicolons. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158627 Approved by: https://github.com/albanD	2025-07-22 02:49:06 +00:00
FFFrog	aee8a2e985	Remove duplicated installation for python dependencies. (#158339 ) As the title stated. The `Common` Section have installed the python dependencies `1b389025ba/README.md (L247)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158339 Approved by: https://github.com/ezyang	2025-07-22 02:39:28 +00:00
PaulZhang12	eac777c4f4	[Inductor] Expose decomposeK knobs as envvars (#158745 ) Fix up decomposeK autotuning, by removing condition to return more than `k_splits_limit` and setting default to 10 instead of 5. Allow `k_splits_limit` to be configurable to the user via `TORCHINDUCTOR_NUM_DECOMPOSE_K_SPLITS` and also allow user to configure threshold in which to use decompose_k via `TORCHINDUCTOR_DECOMPOSE_K_THRESHOLD` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158745 Approved by: https://github.com/eellison	2025-07-22 01:59:51 +00:00
Xu Han	1a6b21c59f	[AOTI] fix load_pt2 split wrong model name on Windows (#158711 ) fix load_pt2 split wrong model name on Windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158711 Approved by: https://github.com/jansel	2025-07-22 01:54:44 +00:00
Nikita Shulga	abe0c9538a	[BE] Fix extra-semi warnings (#158730 ) And prevent new ones from appearing by removing `-Wno-error=extra-semi` (not sure what was thereason behind adding the warning but not erroring on on it when building with -Werror introduced by https://github.com/pytorch/pytorch/pull/140236 ) 300+ violations of that rule were fixed by running `sed -i -e "s/});/})/" /` against `torch/nativert` Other 3p deps that needs updates: - TensorPipe - LLVM - FBGEMM Pull Request resolved: https://github.com/pytorch/pytorch/pull/158730 Approved by: https://github.com/Skylion007	2025-07-22 01:05:03 +00:00
PyTorch MergeBot	95b658427d	Revert "Add DeviceAllocator as the base device allocator (#138222 )" This reverts commit 1179e333237b02ed8fe2ba10cb9a23adf98d7d7a. Reverted https://github.com/pytorch/pytorch/pull/138222 on behalf of https://github.com/ZainRizvi due to Very sorry but this is still breaking internally. @albanD would you be able to help get this past the finish line? D78496124 has more details on the failure and the workaround might be to do something like what's in D78684669. To validate the fixes internally, you can follow the instructions here to ghimport the changes: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/138222#issuecomment-3100195370))	2025-07-22 01:01:41 +00:00
PyTorch MergeBot	6341311333	Revert "Add unified memory APIs for torch.accelerator (#152932 )" This reverts commit 2ad5c25cfc603c3656e6699d6137419dbb009495. Reverted https://github.com/pytorch/pytorch/pull/152932 on behalf of https://github.com/ZainRizvi due to Very sorry but this is still breaking internally. @albanD would you be able to help get this past the finish line? D78496124 has more details on the failure and the workaround might be to do something like what's in D78684669. To validate the fixes internally, you can follow the instructions here to ghimport the changes: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/138222#issuecomment-3100195370))	2025-07-22 01:01:41 +00:00
Han, Xu	350d6af52c	[AOTI] add windows support for get_cpp_compile_command (#158732 ) add windows support for `get_cpp_compile_command`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158732 Approved by: https://github.com/desertfire	2025-07-22 00:23:10 +00:00
PyTorch MergeBot	9281625a9b	Revert "Setup TorchBench in Docker (#158613 )" This reverts commit cab28330f8c49cdb66d6a299755dc09c87c14a9d. Reverted https://github.com/pytorch/pytorch/pull/158613 on behalf of https://github.com/ZainRizvi due to Seems to have broken trunk. See [GH job link](https://github.com/pytorch/pytorch/actions/runs/16429779764/job/46430634676) [HUD commit link](`b3c868d603`) ([comment](https://github.com/pytorch/pytorch/pull/158613#issuecomment-3100023071))	2025-07-22 00:12:49 +00:00
Huamin Li	2c37acfd89	[AOTI][CPU] Consider bias=None case for fbgemm_linear_fp16_weight (#158535 ) Test Plan: Rollback Plan: Differential Revision: D78458214 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158535 Approved by: https://github.com/houseroad, https://github.com/henryoier, https://github.com/jingsh	2025-07-21 23:42:44 +00:00
Raymond Li	08540b13c6	Use cuda error code instead of error text in get_cuda_error_help (#158688 ) Use cudaError_t and switch through the enum to prevent impact by upstream changes in wording Pull Request resolved: https://github.com/pytorch/pytorch/pull/158688 Approved by: https://github.com/q10, https://github.com/aorenste	2025-07-21 23:34:50 +00:00
zpcore	187c2deb40	Fix clamp(min/max) strategy (#158619 ) Part of plan https://github.com/pytorch/pytorch/issues/157495. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158619 Approved by: https://github.com/wanchaol	2025-07-21 23:26:08 +00:00
Catherine Lee	67be2f27e1	[CI][lintrunner] Only run on non deleted changed files (#158794 ) My PR was failing lint because I removed a file, and then lintrunner would try to run on the deleted file and error, so this changes how the changed files are retrieved to only retrieve changed files that have not been removed. I don't think this is possible through `gh pr view`, so instead it uses `gh api` Testing: https://github.com/pytorch/pytorch/pull/158795 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158794 Approved by: https://github.com/seemethere	2025-07-21 23:22:37 +00:00
henrylhtsang	d293022c47	[cutass backend] memorize parts of cache key to reduce general overhead (#158311 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158311 Approved by: https://github.com/ColinPeppler ghstack dependencies: #156781	2025-07-21 23:21:12 +00:00
PyTorch MergeBot	ee5a434f8c	Revert "[BE] remove torch deploy - conditionals (#158288 )" This reverts commit 1a4268b8113d5160d71225bab980f03c2318a0a4. Reverted https://github.com/pytorch/pytorch/pull/158288 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally, see D78496147 for details. To validate your fixes internally, you can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158288#issuecomment-3099826158))	2025-07-21 23:17:39 +00:00
PyTorch MergeBot	4c18e85300	Revert "[BE] Remove torch deploy \| remove torch deploy specific files (#158290 )" This reverts commit a6de309ca15cda6b2792fc74e82814dc8d2f9dd9. Reverted https://github.com/pytorch/pytorch/pull/158290 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally, see D78496147 for details. To validate your fixes internally, you can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158288#issuecomment-3099826158))	2025-07-21 23:17:39 +00:00
PyTorch MergeBot	920f26c761	Revert "[BE] Remove __reduce_deploy__ (#158291 )" This reverts commit 0b9fb91f17edfbc51ae36584dcb8350b2d8bb23b. Reverted https://github.com/pytorch/pytorch/pull/158291 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally, see D78496147 for details. To validate your fixes internally, you can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158288#issuecomment-3099826158))	2025-07-21 23:17:38 +00:00
PyTorch MergeBot	99cc3633f6	Revert "[BE] Modify PyObjectSlot the assume only a single interpreter is in use (#158407 )" This reverts commit d9426a81d2ab54f809a3b32a6ab2e606073fe66f. Reverted https://github.com/pytorch/pytorch/pull/158407 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally, see D78496147 for details. To validate your fixes internally, you can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158288#issuecomment-3099826158))	2025-07-21 23:17:38 +00:00
PyTorch MergeBot	15a50dcf1c	Revert "[BE] Make PyObjectSlot use a global PyInterpreter and remove (#158427 )" This reverts commit eb7365072315be2bc4259114e25e269801441748. Reverted https://github.com/pytorch/pytorch/pull/158427 on behalf of https://github.com/ZainRizvi due to Reverting this as part of reverting the stack for https://github.com/pytorch/pytorch/pull/158288 ([comment](https://github.com/pytorch/pytorch/pull/158427#issuecomment-3099815367))	2025-07-21 23:14:57 +00:00
Pian Pawakapan	1227ed6674	[dynamic shapes] fix _maybe_evaluate_static axioms bug (#158672 ) Summary: couldn't get a minimal repro, but xref for size change during dict iteration error: https://fb.workplace.com/groups/1075192433118967/posts/1709439696360901 Test Plan: - Rollback Plan: Differential Revision: D78047846 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158672 Approved by: https://github.com/bobrenjc93	2025-07-21 23:14:19 +00:00
Svetlana Karslioglu	2bb684304d	Fix the typos in the right nav by pulling the latest theme (#158746 ) This will fix broken links in the right nav. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158746 Approved by: https://github.com/malfet	2025-07-21 22:51:07 +00:00
Ketan Ambati	f09a484b81	Remove is_arvr_mode() from xnnpack.buck.bzl (#158682 ) Summary: Changes * Deleted function import from build definition utilities * Removed `load("//tools/build_defs:fbsource_utils.bzl", "is_arvr_mode")` * Replaced is_arvr_mode() function calls with direct references to configuration flags * Changed from `is_arvr_mode()` to `"ovr_config//build_mode:arvr_mode"` * Changed conditional expressions to Buck `select()` statements Test Plan: Check if CI passes Rollback Plan: Differential Revision: D78520947 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158682 Approved by: https://github.com/malfet	2025-07-21 22:49:26 +00:00
PyTorch MergeBot	feaa02f9ad	Revert "[build] pin `setuptools>=77` to enable PEP 639 (#158104 )" This reverts commit a78fb63dbdf98a1db219095293de1a11005e0390. Reverted https://github.com/pytorch/pytorch/pull/158104 on behalf of https://github.com/malfet due to It still breaks inductor-perf-nightly, see https://github.com/pytorch/pytorch/actions/runs/16425364208/job/46417088208, I'm going to dismiss all previous reviews ([comment](https://github.com/pytorch/pytorch/pull/158104#issuecomment-3099706457))	2025-07-21 22:46:53 +00:00
Yang Wang	b3c868d603	[vllm]Add vllm.txt for pinned commit (#158754 ) It seems the nightly.yml won't auto-generate txt file when it does not existed, so added the file with latest merged commit from vllm: [vllm commit](https://github.com/vllm-project/vllm/commits/main) Error: https://github.com/pytorch/pytorch/actions/runs/16405915719/job/46351847504 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158754 Approved by: https://github.com/huydhn	2025-07-21 22:41:07 +00:00
Huy Do	cab28330f8	Setup TorchBench in Docker (#158613 ) This reduces the time spending to setup TorchBench in A100/H100 by another half an hour ### Testing * H100 benchmark https://github.com/pytorch/pytorch/actions/runs/16396172453. Once this done, I will review the results on [HUD](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Fri%2C%2011%20Jul%202025%2023%3A01%3A24%20GMT&stopTime=Fri%2C%2018%20Jul%202025%2023%3A01%3A24%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cuda%20(h100)&lBranch=gh/huydhn/6/head&lCommit=14a38c719b29a19f518239b5edb084838ac5d2fb&rBranch=main&rCommit=0a99b026d6bd0f67dc2c0a20fe3228ddc4144854) to confirm that all models are there * A100 benchmark https://github.com/pytorch/pytorch/actions/runs/16396173932 Signed-off-by: Huy Do <huydhn@gmail.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158613 Approved by: https://github.com/janeyx99	2025-07-21 22:34:08 +00:00
Tristan Rice	4366610f5a	[c10d] block_current_stream: correctness fixes (#158757 ) This fixes a number of issues that were present in https://github.com/pytorch/pytorch/pull/156883 as pointed out by @ngimel Test plan: Expanded tests to cover use after free behavior + non-default stream ``` pytest test/distributed/test_c10d_pypg.py -v -k block_current_stream ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158757 Approved by: https://github.com/ngimel	2025-07-21 22:23:44 +00:00
codingwithsurya	dd0adc9386	[SymmMem] Add NVSHMEM broadcast support into Triton (#158514 ) Adds broadcast collective operation for distributing data from root PE to all other PEs in NVSHMEM Triton kernels. Tests: `python test/distributed/test_nvshmem_triton.py -k test_triton_broadcast` <details> <summary> Quick debug print for sanity check </summary> ```markdown ============================================================ [Rank 0] Starting broadcast test with world_size=2 ============================================================ [Rank 0] Configuration: - nelems: 4 - dtype: torch.int64, element_size: 8 bytes - nelems_bytes: 32 ============================================================ [Rank 1] Starting broadcast test with world_size=2 ============================================================ [Rank 1] Configuration: - nelems: 4 - dtype: torch.int64, element_size: 8 bytes - nelems_bytes: 32 [Rank 1] Non-root source data: [-1, -1, -1, -1] [Rank 0] Root source data: [100, 101, 102, 103] [Rank 1] Initial destination: [-999, -999, -999, -999] [Rank 0] Initial destination: [-999, -999, -999, -999] [Rank 0] Executing broadcast operation... [Rank 1] Executing broadcast operation... [Rank 0] Broadcast operation completed /data/users/suryasub/pytorch/torch/distributed/distributed_c10d.py:4809: UserWarning: No device id is provided via `init_process_group` or `barrier `. Using the current device set by the user. warnings.warn( # warn only once [Rank 1] Broadcast operation completed /data/users/suryasub/pytorch/torch/distributed/distributed_c10d.py:4809: UserWarning: No device id is provided via `init_process_group` or `barrier `. Using the current device set by the user. warnings.warn( # warn only once [Rank 1] Results after broadcast: [Rank 0] Results after broadcast: [Rank 1] Destination buffer: [100, 101, 102, 103] [Rank 1] Expected: [100, 101, 102, 103] [Rank 0] Destination buffer: [100, 101, 102, 103] [Rank 0] Expected: [100, 101, 102, 103] [Rank 1] Match: ✓ [Rank 0] Match: ✓ [Rank 1] ============================================================ [Rank 1] Broadcast test PASSED ✓ [Rank 1] Summary: Root PE 0 broadcasted [100, 101, 102, 103] to all PEs [Rank 1] ============================================================ [Rank 0] ============================================================ [Rank 0] Broadcast test PASSED ✓ [Rank 0] Summary: Root PE 0 broadcasted [100, 101, 102, 103] to all PEs [Rank 0] ============================================================ ``` </details> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158514 Approved by: https://github.com/fduwjj, https://github.com/mandroid6 ghstack dependencies: #158511, #158512, #158513	2025-07-21 22:23:26 +00:00
PyTorch MergeBot	734826d88e	Revert "[AOTI] windows package load dev (#158671 )" This reverts commit d42c40976727fed4c9908d4194f26917d0a3da66. Reverted https://github.com/pytorch/pytorch/pull/158671 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally. @angelayi can you please help them validate the fixes internally? You can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158671#issuecomment-3099570374))	2025-07-21 22:20:46 +00:00
PyTorch MergeBot	5a56e6a72b	Revert "[AOTI] fix extract file failed on Windows. (#158702 )" This reverts commit 7cc1a9546c135f8e7635e0d38aa2bba797f8907d. Reverted https://github.com/pytorch/pytorch/pull/158702 on behalf of https://github.com/ZainRizvi due to Sorry but I had to revert this PR in order to revert https://github.com/pytorch/pytorch/pull/158671 ([comment](https://github.com/pytorch/pytorch/pull/158702#issuecomment-3099556215))	2025-07-21 22:18:19 +00:00
PyTorch MergeBot	e8af168ee0	Revert "[AOTI] normalize path and process model files. (#158705 )" This reverts commit ff0da08f4bc5ee135b495926cd58a36a1c0e1a5b. Reverted https://github.com/pytorch/pytorch/pull/158705 on behalf of https://github.com/ZainRizvi due to Sorry but I had to revert this PR in order to revert https://github.com/pytorch/pytorch/pull/158671 ([comment](https://github.com/pytorch/pytorch/pull/158705#issuecomment-3099532516))	2025-07-21 22:16:03 +00:00
PyTorch MergeBot	97d7dc197f	Revert "[AOTI] Convert C-struct zip handling to RAII container (#158687 )" This reverts commit 8ed5e1844c77d952bcea89ca7d0225d876fec4e8. Reverted https://github.com/pytorch/pytorch/pull/158687 on behalf of https://github.com/ZainRizvi due to Sorry but I had to revert this PR in order to revert https://github.com/pytorch/pytorch/pull/158671 ([comment](https://github.com/pytorch/pytorch/pull/158687#issuecomment-3099515618))	2025-07-21 22:13:26 +00:00
Lucas Kabela	9498d95b9c	[Dynamo][BetterEngineering] Type trace_rules.py (#158679 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to a core file, `trace_rules.py` Running ``` mypy torch/_dynamo/trace_rules.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 2564 \| 3997 \| 64.15% \| 34 \| 53 \| 64.15% \| \| This PR \| 4022 \| 4022 \| 100.00% \| 53 \| 53 \| 100.00% \| \| Delta \| +1458 \| +25 \| +35.85% \| +19 \| 0 \| +35.85% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158679 Approved by: https://github.com/williamwen42	2025-07-21 22:12:59 +00:00
Jeff Daily	0e46f54286	[ROCm][CI] update HIP patch for 6.4.1 (#158651 ) patch is intended to fix hipGraph capture for some miopen kernels Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/158651 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-21 22:09:36 +00:00
zeshengzong	216ba6e5f2	Fix `MaskedTensor` to device ignored mask (#151205 ) Fixes #147140 ## Changes - Add `to` implementation in `MaskedTensor` to support move `mask` to target device ## Test Result ```python In [1]: import torch ...: from torch.masked import as_masked_tensor ...: data = torch.tensor([1,2,3]) ...: mask = torch.tensor([True,False,True]) ...: mt = as_masked_tensor(data, mask).to('cuda') ...: mt.get_data().device, mt.get_mask().device /home/zong/code/pytorch/torch/masked/maskedtensor/core.py:247: UserWarning: The PyTorch API of MaskedTensors is in prototype stage and will change in the near future. Please open a Github issue for features requests and see our documentation on the torch.masked module for further information about the project. return MaskedTensor(data, mask) /home/zong/code/pytorch/torch/masked/maskedtensor/_ops_refs.py:354: UserWarning: The PyTorch API of MaskedTensors is in prototype stage and will change in the near future. Please open a Github issue for features requests and see our documentation on the torch.masked module for further information about the project. return MaskedTensor(new_data, _maybe_get_mask(args[0])) Out[1]: (device(type='cuda', index=0), device(type='cuda', index=0)) In [2]: mt.sum(dim=0) /home/zong/code/pytorch/torch/masked/maskedtensor/core.py:247: UserWarning: The PyTorch API of MaskedTensors is in prototype stage and will change in the near future. Please open a Github issue for features requests and see our documentation on the torch.masked module for further information about the project. return MaskedTensor(data, mask) Out[2]: MaskedTensor(4, True) ``` ```bash pytest test/test_maskedtensor.py -vv ``` ![image](https://github.com/user-attachments/assets/640b809c-b4f0-4aca-a09e-04049017a745) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151205 Approved by: https://github.com/ezyang	2025-07-21 21:44:49 +00:00
dependabot[bot]	c774180e59	Bump requests from 2.32.2 to 2.32.4 in /tools/build/bazel (#158006 ) Bumps [requests](https://github.com/psf/requests) from 2.32.2 to 2.32.4. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/psf/requests/releases">requests's releases</a>.</em></p> <blockquote> <h2>v2.32.4</h2> <h2>2.32.4 (2025-06-10)</h2> <p><strong>Security</strong></p> <ul> <li>CVE-2024-47081 Fixed an issue where a maliciously crafted URL and trusted environment will retrieve credentials for the wrong hostname/machine from a netrc file. (<a href="https://redirect.github.com/psf/requests/issues/6965">#6965</a>)</li> </ul> <p><strong>Improvements</strong></p> <ul> <li>Numerous documentation improvements</li> </ul> <p><strong>Deprecations</strong></p> <ul> <li>Added support for pypy 3.11 for Linux and macOS. (<a href="https://redirect.github.com/psf/requests/issues/6926">#6926</a>)</li> <li>Dropped support for pypy 3.9 following its end of support. (<a href="https://redirect.github.com/psf/requests/issues/6926">#6926</a>)</li> </ul> <h2>v2.32.3</h2> <h2>2.32.3 (2024-05-29)</h2> <p><strong>Bugfixes</strong></p> <ul> <li>Fixed bug breaking the ability to specify custom SSLContexts in sub-classes of HTTPAdapter. (<a href="https://redirect.github.com/psf/requests/issues/6716">#6716</a>)</li> <li>Fixed issue where Requests started failing to run on Python versions compiled without the <code>ssl</code> module. (<a href="https://redirect.github.com/psf/requests/issues/6724">#6724</a>)</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/psf/requests/blob/main/HISTORY.md">requests's changelog</a>.</em></p> <blockquote> <h2>2.32.4 (2025-06-10)</h2> <p><strong>Security</strong></p> <ul> <li>CVE-2024-47081 Fixed an issue where a maliciously crafted URL and trusted environment will retrieve credentials for the wrong hostname/machine from a netrc file.</li> </ul> <p><strong>Improvements</strong></p> <ul> <li>Numerous documentation improvements</li> </ul> <p><strong>Deprecations</strong></p> <ul> <li>Added support for pypy 3.11 for Linux and macOS.</li> <li>Dropped support for pypy 3.9 following its end of support.</li> </ul> <h2>2.32.3 (2024-05-29)</h2> <p><strong>Bugfixes</strong></p> <ul> <li>Fixed bug breaking the ability to specify custom SSLContexts in sub-classes of HTTPAdapter. (<a href="https://redirect.github.com/psf/requests/issues/6716">#6716</a>)</li> <li>Fixed issue where Requests started failing to run on Python versions compiled without the <code>ssl</code> module. (<a href="https://redirect.github.com/psf/requests/issues/6724">#6724</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="`021dc729f0`"><code>021dc72</code></a> Polish up release tooling for last manual release</li> <li><a href="`821770e822`"><code>821770e</code></a> Bump version and add release notes for v2.32.4</li> <li><a href="`59f8aa2adf`"><code>59f8aa2</code></a> Add netrc file search information to authentication documentation (<a href="https://redirect.github.com/psf/requests/issues/6876">#6876</a>)</li> <li><a href="`5b4b64c346`"><code>5b4b64c</code></a> Add more tests to prevent regression of CVE 2024 47081</li> <li><a href="`7bc45877a8`"><code>7bc4587</code></a> Add new test to check netrc auth leak (<a href="https://redirect.github.com/psf/requests/issues/6962">#6962</a>)</li> <li><a href="`96ba401c12`"><code>96ba401</code></a> Only use hostname to do netrc lookup instead of netloc</li> <li><a href="`7341690e84`"><code>7341690</code></a> Merge pull request <a href="https://redirect.github.com/psf/requests/issues/6951">#6951</a> from tswast/patch-1</li> <li><a href="`6716d7c9f2`"><code>6716d7c</code></a> remove links</li> <li><a href="`a7e1c745dc`"><code>a7e1c74</code></a> Update docs/conf.py</li> <li><a href="`c799b8167a`"><code>c799b81</code></a> docs: fix dead links to kenreitz.org</li> <li>Additional commits viewable in <a href="https://github.com/psf/requests/compare/v2.32.2...v2.32.4">compare view</a></li> </ul> </details> <br /> [![Dependabot compatibility score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=requests&package-manager=pip&previous-version=2.32.2&new-version=2.32.4)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot merge` will merge this PR after your CI passes on it - `@dependabot squash and merge` will squash and merge this PR after your CI passes on it - `@dependabot cancel merge` will cancel a previously requested merge and block automerging - `@dependabot reopen` will reopen this PR if it is closed - `@dependabot close` will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) You can disable automated security fix PRs for this repo from the [Security Alerts page](https://github.com/pytorch/pytorch/network/alerts). </details> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158006 Approved by: https://github.com/Skylion007 Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2025-07-21 21:35:38 +00:00
Bin Bao	a991e285ae	[AOTI] Add more default options to compile_standalone (#158560 ) Summary: When compiling for standalone, make embed_kernel_binary and emit_multi_arch_kernel default to True, and add a default name for model_name_for_generated_files to make the generated cpp project easier to understand. Also improved the weights object file naming to be more readable. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158560 Approved by: https://github.com/yushangdi	2025-07-21 21:16:48 +00:00
zero000064	9e0473b566	removed zero dim cpu logic from fake_tensor.py (#147501 ) Fixes #144748 In #144748, the inconsistency between the eager mode and the inductor mode is reported as a bug. The root cause is fake_tenosr.py's find-common-device method, `0b0da81021/torch/_subclasses/fake_tensor.py (L833)`, takes zero dim cpu tensor into account but the device check in adaption.h doesn't. This fix is to add a list for some ops to bypass zero-dim-cpu-tensor check to align with the eager mode. Pull Request resolved: https://github.com/pytorch/pytorch/pull/147501 Approved by: https://github.com/ezyang	2025-07-21 21:11:10 +00:00
Howard Huang	5e17932c22	[DCP] Add support for ShardedTensor to PgTransport (#158573 ) Add support for ShardedTensors in when PGTransport is used for send/recv checkpoints Test is pulled from https://github.com/pytorch/pytorch/pull/157963 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158573 Approved by: https://github.com/meetv18	2025-07-21 21:04:23 +00:00
Xuan Zhang	6b0526a2c4	ban fusion of large amount of reads (#158667 ) This is an reland attempt of https://github.com/pytorch/pytorch/pull/157563, but insteading of introducing the `realize_acc_reads_size_threshold` config and setting to a default value, we set it to `None` for now to unblock an internal use case. Will deep dive into the issue and harden the logic in later PRs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158667 Approved by: https://github.com/yf225	2025-07-21 21:00:40 +00:00
PyTorch MergeBot	bc379aebe2	Revert "Still run TritonBundler with BundledAOTAutogradCache, save autotune results (#158048 )" This reverts commit 8e57cdb746b4ab28865fdf01532f87b0d21700e9. Reverted https://github.com/pytorch/pytorch/pull/158048 on behalf of https://github.com/jeffdaily due to rocm failures due to unit test introduced in this PR, but no pre-merge signal available ([comment](https://github.com/pytorch/pytorch/pull/158048#issuecomment-3098746624))	2025-07-21 20:45:21 +00:00
Ruben Rodriguez Buchillon	b1a0c34dd3	[pt2 event logging] add configurable prefix (#157678 ) Summary: # Why make experiments easier to find # What - dynamo config to provide a prefix - use the prefix when sending data to scuba through the self.id_ field Test Plan: ``` # code edited to set the prefix as `coconutruben-02` buck2 run mode/opt scripts/coconutruben/torchmm:experiment 2>&1 \| tee /tmp/epx040 ``` on scuba ``` \| autotune_dtypes \| autotune_offset \| autotune_shape \| autotune_strides \| event \| run_id \| \| -----\| -----\| -----\| -----\| -----\| ----- \| \| "torch.float16, torch.float16" \| "0, 0" \| "4096x3008, 3008x2048" \| "[3008, 1], [2048, 1]" \| "mm_template_autotuning" \| "coconutruben-02-e6bdccc5-6dcf-4d68-9a04-b34f2c6d94fd" \| \| "torch.float16, torch.float16" \| "0, 0" \| "4096x3008, 3008x2048" \| "[3008, 1], [2048, 1]" \| "mm_template_autotuning" \| "coconutruben-02-14165153-5842-4eaa-9e6c-3b0cbc016375" \| ``` Rollback Plan: Differential Revision: D77837550 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157678 Approved by: https://github.com/stashuk-olek	2025-07-21 20:41:03 +00:00
Eli Uriegas	851e953f68	ci: Only run lint jobs on relevant files (#158773 ) Conditionally run lint jobs on relevant files, this is mainly targetd at clangtidy since it takes a long time but also includes mypy since that's an additional 4 minutes of runtime that we can save. Signed-off-by: Eli Uriegas <eliuriegas@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158773 Approved by: https://github.com/malfet	2025-07-21 20:21:34 +00:00
zeshengzong	b66f429827	Fix `torch.randint`, `torch.mul` param missing description (#158731 ) Wrong separator cause param description truncated. - Change separator of param and its description - Remove quote make `torch.dtype` display as reference to the class ## Test Result ### Before <img width="1092" height="784" alt="image" src="https://github.com/user-attachments/assets/e8d96b26-07e9-40ff-9392-fa6665d4bbe4" /> <img width="1111" height="457" alt="image" src="https://github.com/user-attachments/assets/a3c2e333-f861-4aeb-b4fb-05c8d880ae81" /> ### After <img width="897" height="820" alt="image" src="https://github.com/user-attachments/assets/d1b5cefa-717a-4223-84b0-4346b7eecf44" /> <img width="872" height="409" alt="image" src="https://github.com/user-attachments/assets/96223c37-cd9d-4656-9e55-032d09cbe5c1" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158731 Approved by: https://github.com/ngimel	2025-07-21 20:17:27 +00:00
Lucas Kabela	ea5b06ed5b	[Dynamo][BetterEngineering] Type side_effects.py (#158605 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to a core file, `side_effects.py` Running ``` mypy torch/_dynamo/side_effects.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 365 \| 1166 \| 31.30% \| 16 \| 51 \| 31.37% \| \| This PR \| 1185 \| 1185 \| 100.00% \| 51 \| 51 \| 100.00% \| \| Delta \| +820 \| +19 \| +68.70% \| +35 \| 0 \| +68.63% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158605 Approved by: https://github.com/StrongerXi	2025-07-21 19:34:14 +00:00
Natalia Gimelshein	25fbf09d5f	Use more fine-grained locks in sym mem kernels (#158523 ) Summary: Use only acq in the beginning of the kernel, and only release in the end Test Plan: Existing tests Rollback Plan: Differential Revision: D78458020 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158523 Approved by: https://github.com/drisspg, https://github.com/kwen2501	2025-07-21 19:23:47 +00:00
Benjamin Glass	22920c9138	Grab bag of (mostly) typing improvements (#158075 ) Collects some scattershot improvements made while attempting to enable training for AOTInductor. Non-typing changes are: 1. Swapping a few custom searches for the output node in an FX graph for calling `graph.output_node()`. 2. Removing two unused parameters from `torch.export._unlift._unlift`. 3. Switching handles to constants in `cpp_wrapper_cpu` to use C++ references for memory efficiency. 4. Cleaning out unused, unexported imports from `torch/export/__init__.py`, and adding one missing export to `__all__`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158075 Approved by: https://github.com/Skylion007	2025-07-21 19:17:01 +00:00
codingwithsurya	ad2dec1997	[SymmMem] Add NVSHMEM alltoall support into Triton (#158513 ) Implements collective alltoall operation for NVSHMEM Triton kernels. Enables data exchange where each PE sends unique data to every other PE in the team. Tests: `python test/distributed/test_nvshmem_triton.py -k test_triton_alltoall` <details> <summary>Quick debug print for sanity check</summary> ```markdown ============================================================ [Rank 0] Starting alltoall test with world_size=2 ============================================================ [Rank 0] Configuration: - nelems_per_pe: 2 - dtype: torch.int64, element_size: 8 bytes - nelems_bytes: 16 /dvs/p4/build/sw/rel/gpgpu/toolkit/r12.8/main_nvshmem/src/modules/transport/ibrc/ibrc.cpp:1653: NULL value get_device_list failed /dvs/p4/build/sw/rel/gpgpu/toolkit/r12.8/main_nvshmem/src/modules/transport/ibrc/ibrc.cpp:1653: NULL value get_device_list failed [Rank 0] Preparing source data: [Rank 1] Preparing source data: - Data for PE 0: [0, 0] (indices 0-1) - Data for PE 1: [1, 1] (indices 2-3) [Rank 0] Complete source buffer: [0, 0, 1, 1] - Data for PE 0: [100, 100] (indices 0-1) - Data for PE 1: [101, 101] (indices 2-3) [Rank 1] Complete source buffer: [100, 100, 101, 101] [Rank 1] Initial destination buffer: [-1, -1, -1, -1] [Rank 0] Initial destination buffer: [-1, -1, -1, -1] /data/users/suryasub/pytorch/torch/distributed/distributed_c10d.py:4809: UserWarning: No device id is provided via `init_process_group` or `barrier `. Using the current device set by the user. warnings.warn( # warn only once /data/users/suryasub/pytorch/torch/distributed/distributed_c10d.py:4809: UserWarning: No device id is provided via `init_process_group` or `barrier `. Using the current device set by the user. warnings.warn( # warn only once [rank0]:[W716 15:30:06.215666766 ProcessGroupNCCL.cpp:5064] [PG ID 0 PG GUID 0 Rank 0] using GPU 0 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can specify device_id in init_process_group() to force use of a particular device. [rank1]:[W716 15:30:06.215752786 ProcessGroupNCCL.cpp:5064] [PG ID 0 PG GUID 0 Rank 1] using GPU 1 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can specify device_id in init_process_group() to force use of a particular device. NCCL version 2.27.5+cuda12.4 [Rank 1] Executing alltoall operation... [Rank 0] Executing alltoall operation... [Rank 1] alltoall operation completed /data/users/suryasub/pytorch/torch/distributed/distributed_c10d.py:4809: UserWarning: No device id is provided via `init_process_group` or `barrier `. Using the current device set by the user. warnings.warn( # warn only once [Rank 0] alltoall operation completed /data/users/suryasub/pytorch/torch/distributed/distributed_c10d.py:4809: UserWarning: No device id is provided via `init_process_group` or `barrier `. Using the current device set by the user. warnings.warn( # warn only once [Rank 0] Results after alltoall: [Rank 1] Results after alltoall:[Rank 0] Destination buffer: [0, 0, 100, 100] [Rank 0] Verifying results: - From PE 0 (indices 0-1): Expected: [0, 0] Actual: [0, 0] [Rank 1] Destination buffer: [1, 1, 101, 101] [Rank 1] Verifying results: - From PE 0 (indices 0-1): Expected: [1, 1] Actual: [1, 1] Match: ✓ Match: ✓ - From PE 1 (indices 2-3): Expected: [100, 100] - From PE 1 (indices 2-3): Expected: [101, 101] Actual: [100, 100] Actual: [101, 101] Match: ✓ Match: ✓ [Rank 0] ============================================================ [Rank 0] Summary: ALL TESTS PASSED ✓ [Rank 0] Data flow explanation: - Each rank sends 2 elements to every other rank [Rank 1] ============================================================ [Rank 1] Summary: ALL TESTS PASSED ✓ - Rank 0 sent: [0, 0, 1, 1] [Rank 1] Data flow explanation: - Each rank sends 2 elements to every other rank - Rank 0 received: [0, 0, 100, 100] - My data for PE 0 (0) went to PE 0's buffer - I received PE 0's data for me (0) - My data for PE 1 (1) went to PE 1's buffer - Rank 1 sent: [100, 100, 101, 101] - I received PE 1's data for me (100) [Rank 0] ============================================================ - Rank 1 received: [1, 1, 101, 101] - My data for PE 0 (100) went to PE 0's buffer - I received PE 0's data for me (1) - My data for PE 1 (101) went to PE 1's buffer - I received PE 1's data for me (101) [Rank 1] ============================================================ ``` </details> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158513 Approved by: https://github.com/fduwjj, https://github.com/mandroid6 ghstack dependencies: #158511, #158512	2025-07-21 19:14:47 +00:00
henrylhtsang	662dd7db5b	[cutlass backend] cache maybe_append_choices (#156781 ) This PR attempts to cache: * codegen for cutlass backend for the same kernel. Even if runtime params are different. From some profiling, most of the time spent is on render. So we only target to cache that part for now. The output of render is `code`, and we are able to cache that easily. Also, I have to cache size_args, since it depends on `kernel.get_dynamic_shape_args()`, which depends on the state of self when we call render. make_key is doing most of the work here: We are hashing on input node layouts, output node layout and op.configuration_name() (this is what hash(op) would do anyway). Pull Request resolved: https://github.com/pytorch/pytorch/pull/156781 Approved by: https://github.com/ColinPeppler	2025-07-21 19:02:39 +00:00
PyTorch MergeBot	72db0a98a3	Revert "[DTensor] Assert DTensorSpec has valid placements (#158133 )" This reverts commit 1839e8d04b81ee6eda0cff6fbfc218a7a600f6f7. Reverted https://github.com/pytorch/pytorch/pull/158133 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking internally. See D78496151 for details. To validate your fixes internally, you can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts ([comment](https://github.com/pytorch/pytorch/pull/158133#issuecomment-3097994857))	2025-07-21 18:54:07 +00:00
Benjamin Glass	8ed5e1844c	[AOTI] Convert C-struct zip handling to RAII container (#158687 ) Attempts to fix a memory leak reported in #158614 by wrapping manually managed MiniZ C-structs in an RAII container. I have been unable to reproduce the reported leak, but this seems like the most likely candidate. Fixes #158614 (hopefully) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158687 Approved by: https://github.com/desertfire	2025-07-21 18:53:14 +00:00
Menglu Yu	393fecb2cc	[Optimus][Unit test] clean up the unit test (#158696 ) Summary: We should only patch the specific pattern(s) for each unit test. Test Plan: ``` buck2 test 'fbcode//mode/dev-nosan' fbcode//caffe2/test/inductor:group_batch_fusion ``` Buck UI: https://www.internalfb.com/buck2/f8d37674-91c4-4244-90fa-f24fc3f91e4b Test UI: https://www.internalfb.com/intern/testinfra/testrun/2533275088644915 Network: Up: 100KiB Down: 233KiB (reSessionID-92039f44-bc6f-4e78-87b1-93bca1bd1c66) Analyzing targets. Remaining 0/296 Executing actions. Remaining 0/20196 5.8s exec time total Command: test. Finished 2 local, 2 cache (50% hit) 4.6s exec time cached (79%) Time elapsed: 3:55.1s Tests finished: Pass 13. Fail 0. Fatal 0. Skip 0. Build failure 0 Rollback Plan: Differential Revision: D78598127 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158696 Approved by: https://github.com/Skylion007, https://github.com/masnesral	2025-07-21 18:05:09 +00:00
Sam Larsen	9285b8245c	[BE][testing] fix test_cat_max_autotune_triton (#158589 ) Summary: This test often fails internally -- looks like it's because autotuning sometimes chooses not to do the epilog tuning. Turning off `benchmark_epilogue_fusion` seems to fix. Test Plan: `buck test '@fbcode//mode/opt' fbcode//caffe2/test/inductor:max_autotune -- --exact 'caffe2/test/inductor:max_autotune - test_cat_max_autotune_triton (caffe2.test.inductor.test_max_autotune.TestMaxAutotune)' --run-disabled` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158589 Approved by: https://github.com/eellison	2025-07-21 18:02:18 +00:00
Xuehai Pan	637e75433c	[BE] always use `uv pip` if possible in `pip_init.py` for `lintrunner init` (#157199 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157199 Approved by: https://github.com/ezyang, https://github.com/ZainRizvi	2025-07-21 17:56:05 +00:00
Xuehai Pan	a78fb63dbd	[build] pin `setuptools>=77` to enable PEP 639 (#158104 ) For reference here is the link PEP 639: [peps.python.org/pep-0639](https://peps.python.org/pep-0639/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158104 Approved by: https://github.com/rgommers, https://github.com/Skylion007, https://github.com/atalman	2025-07-21 17:46:40 +00:00
FFFrog	7205458b85	[Easy] Show some clear error when torch.ops.load_library fails. (#157524 ) Background: ```Shell torch 2.5.1+cpu torchvision 0.20.1 ``` ```Python import torch import torchvision Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torchvision/__init__.py", line 10, in <module> from torchvision import _meta_registrations, datasets, io, models, ops, transforms, utils # usort:skip File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torchvision/_meta_registrations.py", line 164, in <module> def meta_nms(dets, scores, iou_threshold): File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torch/library.py", line 795, in register use_lib._register_fake(op_name, func, _stacklevel=stacklevel + 1) File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torch/library.py", line 184, in _register_fake handle = entry.fake_impl.register(func_to_register, source) File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torch/_library/fake_impl.py", line 31, in register if torch._C._dispatch_has_kernel_for_dispatch_key(self.qualname, "Meta"): RuntimeError: operator torchvision::nms does not exist ``` Cause: ``` torchvision's .so file lacks some symbol definitions, because these symbols come from CUDA, but the current environment does not have CUDA and GPU. The above error message is very confusing. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157524 Approved by: https://github.com/ezyang	2025-07-21 17:32:31 +00:00
PyTorch MergeBot	35f1b4ad9e	Revert "Fused RMSNorm implementation (#153666 )" This reverts commit 15ef4f28df0a14e9f0d55a57a4e2db415a303be7. Reverted https://github.com/pytorch/pytorch/pull/153666 on behalf of https://github.com/ZainRizvi due to Sorry but this is breaking tests internally. @albanD can you please help land this change?You can follow the instructions here: https://fburl.com/fixing-ghfirst-reverts. See D78599667 for more info ([comment](https://github.com/pytorch/pytorch/pull/153666#issuecomment-3097690935))	2025-07-21 17:31:42 +00:00
Yu, Guangye	cbe1cb7018	[CMake] Move xpu flag to xpu.cmake (#158542 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158542 Approved by: https://github.com/gujinghui, https://github.com/ezyang	2025-07-21 17:19:59 +00:00
Xu Han	9894d43b6c	[AOTI] explicit aoti wrapper functions for Windows. (#158713 ) On Windows, we need to explicit declaration for export APIs. Because the package loader call these API via GetProcAddress. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158713 Approved by: https://github.com/desertfire	2025-07-21 15:59:44 +00:00
Zain Rizvi	f168cf49a8	[BE] Always use python 3.9 for pre-push hook's lintrunner (#158693 ) A follow up to https://github.com/pytorch/pytorch/pull/158389 Sets up the pre-push lintrunner to always use python 3.9 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158693 Approved by: https://github.com/atalman	2025-07-21 15:19:27 +00:00
PyTorch MergeBot	393377d215	Revert "[CI] update flake8 and mypy lint dependencies (#158720 )" This reverts commit a527e816935957a164d74dd7c5069310b2857695. Reverted https://github.com/pytorch/pytorch/pull/158720 on behalf of https://github.com/malfet due to This broke lint, see `8e57cdb746/1` ([comment](https://github.com/pytorch/pytorch/pull/158720#issuecomment-3096893256))	2025-07-21 13:58:50 +00:00
James Wu	8e57cdb746	Still run TritonBundler with BundledAOTAutogradCache, save autotune results (#158048 ) When running BundledAOTAutogradCache with precompile, we still need to run triton bundling so that the precompiled CompiledFxGraph has triton cuda kernels. We also pre save the autotune results in the precompile artifact. It would be even better to pre trim the cuda kernels on save and apply them, which we can work on later. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158048 Approved by: https://github.com/zhxchen17	2025-07-21 13:35:46 +00:00
Edward Z. Yang	d5a29fc58a	De-abstract premature generalization with InductorWrapper (#158528 ) See docblock on InductorWrapper for the distinction. This will matter on a later refactor PR where I will change the signature for one of these but not the other. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158528 Approved by: https://github.com/jamesjwu ghstack dependencies: #158449	2025-07-21 13:27:07 +00:00
Edward Z. Yang	979fae761c	Rename modules in AOTAutograd (#158449 ) Fixes https://github.com/pytorch/pytorch/issues/158382 ``` renamed: torch/_functorch/_aot_autograd/dispatch_and_compile_graph.py -> torch/_functorch/_aot_autograd/graph_capture.py renamed: torch/_functorch/_aot_autograd/traced_function_transforms.py -> torch/_functorch/_aot_autograd/graph_capture_wrappers.py renamed: torch/_functorch/_aot_autograd/jit_compile_runtime_wrappers.py -> torch/_functorch/_aot_autograd/graph_compile.py ``` Everything else is ONLY import changes. I did not rename any functions even if we probably should have. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158449 Approved by: https://github.com/jamesjwu	2025-07-21 13:27:07 +00:00
Sun, Jiayi	1eb6b2089f	[Inductor] Set the default value of min_chunk_size to 512 (#150762 ) Change the default value of min_chunk_size from 4096 to 512 to allow more for loops to be parallelized. I tested the Inductor benchmark with this PR on CPU, and saw ~10% improvement in torchbench geomean speedup, and no change in huggingface/timm_models. There are about 15 torchbench models with different degrees of performance improvement, among which functorch_dp_cifar10, opacus_cifar10, hf_Reformer, and pyhpc_turbulent_kinetic_energy have more than 50% performance improvement. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150762 Approved by: https://github.com/leslie-fang-intel, https://github.com/jansel	2025-07-21 12:46:05 +00:00
codingwithsurya	bbc32d680f	[SymmMem] Add NVSHMEM sync_all support into Triton (#158512 ) Adds `sync_all()` function for local store visibility synchronization in NVSHMEM Triton kernels. Provides memory ordering for local operations without remote completion guarantees. Tests: `python test/distributed/test_nvshmem_triton.py -k test_triton_sync` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158512 Approved by: https://github.com/fduwjj ghstack dependencies: #158511	2025-07-21 10:27:59 +00:00
Xuehai Pan	a527e81693	[CI] update flake8 and mypy lint dependencies (#158720 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158720 Approved by: https://github.com/Skylion007	2025-07-21 09:24:29 +00:00
Nikita Shulga	1c6328a588	[EZ][BE] Fix compilation warning in Pooling.metal (#158729 ) This one ``` Compiling /Users/malfet/git/pytorch/pytorch/aten/src/ATen/native/mps/kernels/Pooling.metal to Pooling_30.air /Users/malfet/git/pytorch/pytorch/aten/src/ATen/native/mps/kernels/Pooling.metal:172:1: warning: non-void function does not return a value in all control paths [-Wreturn-type] } ^ 1 warning generated. ``` Although functionally one is not supposed to hit this codepath ever, it's not not to throw warning Pull Request resolved: https://github.com/pytorch/pytorch/pull/158729 Approved by: https://github.com/Skylion007	2025-07-21 04:34:14 +00:00
codingwithsurya	70b4a8880b	[SymmMem] Add NVSHMEM barrier_all, my_pe, n_pes support into Triton (#158511 ) Adds device-side barrier synchronization and PE identification functions for NVSHMEM Triton integration. Includes `barrier_all()` for collective synchronization and `my_pe()`/`n_pes()` for PE identification within kernels. We are launching with cooperative grid launch (for all the PRs in this stack) because the `nvshmemx_collective_launch` function must be used to launch kernels on the GPU when the kernels use NVSHMEM synchronization or collective APIs, and `nvshmemx_collective_launch` essentially boils down to a CUDA cooperative group launch. Tests: `python test/distributed/test_nvshmem_triton.py -k test_triton_barrier` Also tested that if you remove the barrier, you get an assertion error/race conditions. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158511 Approved by: https://github.com/fduwjj	2025-07-21 02:37:33 +00:00
PyTorch MergeBot	5e1232871b	Revert "[build] pin `setuptools>=77` to enable PEP 639 (#158104 )" This reverts commit a4ec381302f8acd279033707b182bed30ffd2091. Reverted https://github.com/pytorch/pytorch/pull/158104 on behalf of https://github.com/malfet due to This break inductor-perf-nighly-macos by failing to build torchvision, see https://github.com/pytorch/pytorch/issues/158728 ([comment](https://github.com/pytorch/pytorch/pull/158104#issuecomment-3095048940))	2025-07-21 02:24:11 +00:00
Xu Han	ff0da08f4b	[AOTI] normalize path and process model files. (#158705 ) Continued to https://github.com/pytorch/pytorch/pull/158702 , split `zip_filename_str` and real file path. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158705 Approved by: https://github.com/desertfire	2025-07-21 01:08:59 +00:00
Nikita Shulga	2cdafab0bd	[BE] Raise ValueError from `torch.cat` meta func (#158249 ) Followup after https://github.com/pytorch/pytorch/pull/155460 From [Python documentation](https://docs.python.org/3/library/exceptions.html#ValueError): > Raised when an operation or function receives an argument that has the right type but an inappropriate value, and the situation is not described by a more precise exception such as IndexError. Raise [`TypeError`](https://docs.python.org/3/library/exceptions.html#TypeError) when input-output types are incompatible with each other > Raised when an operation or function is applied to an object of inappropriate type. The associated value is a string giving details about the type mismatch. > This exception may be raised by user code to indicate that an attempted operation on an object is not supported, and is not meant to be. If an object is meant to support a given operation but has not yet provided an implementation, [NotImplementedError](https://docs.python.org/3/library/exceptions.html#NotImplementedError) is the proper exception to raise. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158249 Approved by: https://github.com/jbschlosser, https://github.com/Skylion007, https://github.com/albanD	2025-07-20 23:49:18 +00:00
Ankita George	4b02bd76d3	DCP safetensors test fix (#158685 ) https://github.com/pytorch/pytorch/pull/158069 removed the consolidated output path argument without updating the test. Reported by a user here https://github.com/pytorch/pytorch/pull/156705#issuecomment-3090748034. Adding back the logic from the original PR https://github.com/pytorch/pytorch/pull/158069 and fixing the test. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158685 Approved by: https://github.com/teja-rao	2025-07-20 22:52:54 +00:00
Mwiza Kunda	2e038793ef	[inductor][templates] Finalize all registered hooks (#157270 ) This refactor ensures all registered template hooks have been finalised before accessing the code object of the template. In `simd.SimdScheduling.codegen_template` the template hooks are finalised manually with `template.finalize_hook(hook_name)` calls, so it is the responsibility of the caller to finalise all the template hooks. This PR adds: - `RenderPartial.finalize_remaining` a function that can be called at the end to finalise the remaining active hooks after a selection of hooks have been finalised manually. - A test with a custom template implementation that registers custom hooks that the scheduler needs to finalise. This test should fail if the scheduler does not finalise the registered custom hook. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157270 Approved by: https://github.com/eellison	2025-07-20 22:07:32 +00:00
Tugsbayasgalan (Tugsuu) Manlaibaatar	5e149a6482	Add deprecation warning (#158203 ) Summary: export_for_training exist because we couldn't migrate internal usages of export to the final IR. Now that we have completed the migration, we should deprecate and delete this API. Test Plan: CI Rollback Plan: Differential Revision: D78240836 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158203 Approved by: https://github.com/JacobSzwejbka	2025-07-20 17:02:01 +00:00
Andrey Talman	badf002014	[Reland] Add warning about removed sm50 and sm60 arches (#158700 ) Related to https://github.com/pytorch/pytorch/issues/157517 Detect when users are executing torch build with cuda 12.8/12.9 and running on Maxwell or Pascal architectures. We would like to include reference to the issue: https://github.com/pytorch/pytorch/issues/157517 as well as ask people to install CUDA 12.6 builds if they are running on sm50 or sm60 architectures. Test: ``` >>> torch.cuda.get_arch_list() ['sm_70', 'sm_75', 'sm_80', 'sm_86', 'sm_90', 'sm_100', 'sm_120', 'compute_120'] >>> torch.cuda.init() /home/atalman/.conda/envs/py312/lib/python3.12/site-packages/torch/cuda/__init__.py:263: UserWarning: Found <GPU Name> which is of cuda capability 5.0. PyTorch no longer supports this GPU because it is too old. The minimum cuda capability supported by this library is 7.0. warnings.warn( /home/atalman/.conda/envs/py312/lib/python3.12/site-packages/torch/cuda/__init__.py:268: UserWarning: Support for Maxwell and Pascal architectures is removed for CUDA 12.8+ builds. Please see https://github.com/pytorch/pytorch/issues/157517 Please install CUDA 12.6 builds if you require Maxwell or Pascal support. ``` Please note I reverted original PR https://github.com/pytorch/pytorch/pull/158301 because it broke internal users. This is a reland, added added check for non empty torch.cuda.get_arch_list() Pull Request resolved: https://github.com/pytorch/pytorch/pull/158700 Approved by: https://github.com/huydhn, https://github.com/Skylion007, https://github.com/eqy	2025-07-20 14:57:46 +00:00
Natalia Gimelshein	4869f71170	don't set CUDA_MODULE_LOADING (#158712 ) If needed, it'll be set in `_C._cuda_init()`. setenv is not threadsafe, so this can cause segfaults due to getenv/setenv races. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158712 Approved by: https://github.com/eqy	2025-07-20 01:36:26 +00:00
Yukio Siraichi	b4abf41425	Raise `BufferError` for DLPack buffer-related errors. (#150691 ) This PR addresses the Array API documentation for [`__dlpack__`][1] and [`from_dlpack`][2] by making some buffer-related errors `BufferError` instead of `RuntimeError`, e.g. incompatible dtype, strides, or device. [1]: https://data-apis.org/array-api/latest/API_specification/generated/array_api.array.__dlpack__.html [2]: https://data-apis.org/array-api/latest/API_specification/generated/array_api.from_dlpack.html#from-dlpack Pull Request resolved: https://github.com/pytorch/pytorch/pull/150691 Approved by: https://github.com/Skylion007, https://github.com/albanD ghstack dependencies: #150216, #150217, #150218	2025-07-20 00:46:21 +00:00
Yukio Siraichi	a10f15718d	[DLPack] Add support for missing keyword-arguments. (#150218 ) This PR introduces the rest of the keyword-arguments added in DLPack version 2023.12: `dl_device` and `copy`. In summary, we handle these arguments in the C++ implementation of `to_dlpack(...)` at _torch/csrc/Module.cpp_, by calling the `maybeCopyTensor` function at _aten/src/ATen/DLConvertor.cpp_. It also introduces the following changes: - Add a new Python API `torchDeviceToDLDevice()`, which is simply a refactoring of the `getDLDevice()` function at _aten/src/ATen/DLConvertor.cpp_. - Add both keyword-arguments to the `from_dlpack()` function at _torch/utils/dlpack.py_ and to the `Tensor.__dlpack__()` dunder method. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150218 Approved by: https://github.com/albanD ghstack dependencies: #150216, #150217	2025-07-20 00:46:20 +00:00
Yukio Siraichi	1d526fe78f	Fix DLPack stream logic. (#150217 ) This PR fixes the logic for dealing with CUDA and ROCm streams whenever we are trying to create a DLPack capsule from a tensor. In summary, this PR: - Uses the legacy default stream if `tensor.__dlpack__(stream=None)` is called for a CUDA tensor. - Errors if `tensor.__dlpack__(stream=2)` is called for a CUDA tensor: PyTorch doesn't support the per-thread default stream. - Errors if `tensor.__dlpack__(stream=stream)`, where `stream` is 1 or 2, is called for a CUDA tensor using ROCm. For more details, see [the documentation][1]. [1]: https://data-apis.org/array-api/latest/API_specification/generated/array_api.array.__dlpack__.html Pull Request resolved: https://github.com/pytorch/pytorch/pull/150217 Approved by: https://github.com/msaroufim, https://github.com/albanD ghstack dependencies: #150216	2025-07-20 00:46:20 +00:00
Yukio Siraichi	b64f338da4	[DLPack] add NumPy exchange tests. (#150216 ) This PR resolves an old TODO that requested NumPy DLPack exchange tests once version 1.22 was required. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150216 Approved by: https://github.com/msaroufim, https://github.com/albanD	2025-07-20 00:46:20 +00:00
dolpm	a1cfe7f1df	[nativert] benchmark util (#158678 ) Differential Revision: D78514241 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158678 Approved by: https://github.com/SherlockNoMad, https://github.com/georgiaphillips	2025-07-20 00:28:09 +00:00
Huy Do	d36afac83b	Build domain libraries for all workflows with TorchBench config (#158601 ) They are expensive GPU runners and should not spend time building packages Signed-off-by: Huy Do <huydhn@gmail.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158601 Approved by: https://github.com/ZainRizvi	2025-07-19 21:51:39 +00:00
Xu Han	7cc1a9546c	[AOTI] fix extract file failed on Windows. (#158702 ) Changes: 1. rename zip index name, and keep it out of normalize path. 2. normalize output path for extract file. Extract files successful: <img width="683" height="247" alt="image" src="https://github.com/user-attachments/assets/72dff7b9-5ec0-4523-a6ee-7768b37bbe63" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158702 Approved by: https://github.com/angelayi	2025-07-19 08:58:42 +00:00
Jane Xu	7cc5d03dfc	Document the rest of the specific optimizer module APIs (#158669 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158669 Approved by: https://github.com/albanD ghstack dependencies: #158483	2025-07-19 07:27:15 +00:00
Jane Xu	f73594164a	[BE] document Adadelta and Adagrad APIs properly (#158483 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158483 Approved by: https://github.com/albanD	2025-07-19 07:27:15 +00:00
AaronWang04	a9f84021fb	[CI] Fixes CI for CUDA Version > 12.9 (#157385 ) Compute capabilities older than volta (inclusive) is no longer supported in CUDA Version > 12.9 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157385 Approved by: https://github.com/eqy	2025-07-19 06:51:57 +00:00
Boyuan Feng	22d82222c6	GenAI Layer Benchmark (#158536 ) This PR adds GenAI layer benchmark. It compares pytorch eager, pytorch compiler, liger, and quack. It covers all kernels supported by [quack](https://github.com/Dao-AILab/quack?tab=readme-ov-file#kernels-) (CrossEntropy Fwd/Bwd, Softmax Fwd/Bwd, RMSNorm Fwd/Bwd, LayerNorm Fwd) and LayerNormBwd. ## Motivations - Many OSS users asked how to properly benchmark torch.compile generated kernels. One common error is to compile a kernel/layer for one shape (e.g., batch size=1) and benchmark for another shape (e.g., batch size = 1024), which leads to bad performance. This provides an simple & clear example for proper benchmark. - We recently added GenAI model benchmark (based on [vLLM](https://hud.pytorch.org/benchmark/llms?repoName=vllm-project%2Fvllm)). But it's usually hard to optimize models directly due to complexity. Layer benchmarks are easier to reason and optimize. ## Key Settings - Avoid reusing a kernel specializing on 1 shape for benchmark on another shape. ```python torch._dynamo.config.automatic_dynamic_shapes = False # Needed since changing args to function causes recompiles torch._dynamo.config.recompile_limit = 1000000 ``` - For forward, people may mark batch size as dynamic to avoid runtime recompilation. We respect the setting in this kernel-level benchmark. ``` torch._dynamo.mark_dynamic(x, 0) ``` GPU: H100 (devvm006.dkl0) Results: [P1874246170](https://www.internalfb.com/phabricator/paste/view/P1874246170) Note: for numerical accuracy, we use the default tolerance of torch.testing.assert_close (i.e., for `torch.bfloat16`, use rtol `1.6e-2` and atol `1e-5`). It shows numerical issues for some backends and kernels. Next step is to add roofline analysis, add to ci for checking regression, cover more GenAI Kernels, and include GenAI Layers for common fusion patterns. <img width="3564" height="2368" alt="CrossEntropyBackward_bench" src="https://github.com/user-attachments/assets/7aa77ad1-83eb-41ea-a27d-50fd5b1dd6be" /> <img width="3564" height="2368" alt="CrossEntropyForward_bench" src="https://github.com/user-attachments/assets/a26ec028-3791-4a41-a12a-05e10f60e9aa" /> <img width="3564" height="2368" alt="LayerNormBackward_bench" src="https://github.com/user-attachments/assets/cc6673ed-c148-4dd2-a729-5f02e717ab3e" /> <img width="3564" height="2368" alt="LayerNormForward_bench" src="https://github.com/user-attachments/assets/f71f9f9d-7b45-4ce7-89d0-e9bce727efae" /> <img width="3564" height="2368" alt="RMSNormBackward_bench" src="https://github.com/user-attachments/assets/e012821a-b7e6-4e83-a24c-c97fa8cd37b5" /> <img width="3564" height="2368" alt="RMSNormForward_bench" src="https://github.com/user-attachments/assets/2d52ee1e-9a8c-4bd1-a180-97b93f07171d" /> <img width="3564" height="2368" alt="SoftmaxBackward_bench" src="https://github.com/user-attachments/assets/02aad056-3ce1-4b40-8cfe-adae81fd017a" /> <img width="3564" height="2368" alt="SoftmaxForward_bench" src="https://github.com/user-attachments/assets/779f6b0d-a102-4164-8300-86fff0329ddf" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158536 Approved by: https://github.com/yf225, https://github.com/eellison	2025-07-19 05:41:01 +00:00
Nikita Shulga	5cde34473c	Fix `MakeTensor::computeStorageSize()` (#158690 ) For tensor with non-zero offset, it must be multiplied by element size Add regression test by creating Tensor in array of 6 elements with offset 3, which before the fix crashed with ``` C++ exception with description "setStorage: sizes [3, 3], strides [0, 1], storage offset 3, and itemsize 4 requiring a storage size of 24 are out of bounds for storage of size 15 Exception raised from checkInBoundsForStorage at /Users/nshulga/git/pytorch/pytorch/aten/src/ATen/native/Resize.h:123 (most recent call first): frame #0: c10::Error::Error(c10::SourceLocation, std::__1::basic_string<char, std::__1::char_traits<char>, std::__1::allocator<char>>) + 56 (0x104a9cd44 in libc10.dylib) frame #1: c10::detail::torchCheckFail(char const, char const, unsigned int, std::__1::basic_string<char, std::__1::char_traits<char>, std::__1::allocator<char>> const&) + 120 (0x104a9a05c in libc10.dylib) frame #2: void at::native::checkInBoundsForStorage<long long>(c10::ArrayRef<long long>, c10::ArrayRef<long long>, long long, caffe2::TypeMeta const&, c10::Storage const&) + 656 (0x111dbd314 in libtorch_cpu.dylib) frame #3: void at::native::setStrided<long long>(at::Tensor const&, c10::ArrayRef<long long>, c10::ArrayRef<long long>, long long) + 152 (0x111dcd22c in libtorch_cpu.dylib) frame #4: at::native::as_strided_tensorimpl(at::Tensor const&, c10::ArrayRef<long long>, c10::ArrayRef<long long>, std::__1::optional<long long>) + 312 (0x111dccf98 in libtorch_cpu.dylib) frame #5: c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::ArrayRef<c10::SymInt>, std::__1::optional<c10::SymInt>), &at::(anonymous namespace)::(anonymous namespace)::wrapper_CPU__as_strided(at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::ArrayRef<c10::SymInt>, std::__1::optional<c10::SymInt>)>, at::Tensor, c10::guts::typelist::typelist<at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::ArrayRef<c10::SymInt>, std::__1::optional<c10::SymInt>>>, at::Tensor (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::ArrayRef<c10::SymInt>, std::__1::optional<c10::SymInt>)>::call(c10::OperatorKernel, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::ArrayRef<c10::SymInt>, std::__1::optional<c10::SymInt>) + 104 (0x1129a1e94 in libtorch_cpu.dylib) frame #6: at::_ops::as_strided::call(at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::ArrayRef<c10::SymInt>, std::__1::optional<c10::SymInt>) + 476 (0x112200ad0 in libtorch_cpu.dylib) frame #7: at::Tensor::as_strided(c10::ArrayRef<long long>, c10::ArrayRef<long long>, std::__1::optional<long long>) const + 236 (0x1115db098 in libtorch_cpu.dylib) frame #8: at::native::expand(at::Tensor const&, c10::ArrayRef<long long>, bool) + 348 (0x111dcc0d4 in libtorch_cpu.dylib) frame #9: c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool), &torch::ADInplaceOrView::(anonymous namespace)::expand(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool)>, at::Tensor, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool>>, at::Tensor (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool)>::call(c10::OperatorKernel, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool) + 116 (0x1157ac410 in libtorch_cpu.dylib) frame #10: c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool), &torch::autograd::VariableType::(anonymous namespace)::expand(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool)>, at::Tensor, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool>>, at::Tensor (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool)>::call(c10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool) + 992 (0x114e8b010 in libtorch_cpu.dylib) frame #11: at::_ops::expand::call(at::Tensor const&, c10::ArrayRef<c10::SymInt>, bool) + 316 (0x112743c90 in libtorch_cpu.dylib) frame #12: at::expand_size(at::Tensor const&, c10::ArrayRef<long long>) + 164 (0x1047d82b4 in basic) frame #13: BasicTest_TestForBlobResizeCPU_Test::TestBody() + 284 (0x1047d8048 in basic) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158690 Approved by: https://github.com/angelayi	2025-07-19 05:21:33 +00:00
Luca Wehrstedt	fac0be7b9c	[async-TP] Turn asserts back into silent skips (#158572 ) https://github.com/pytorch/pytorch/pull/149946 modified some checks that verify whether async-TP is "applicable" to a given collective operation in a graph. Before, the pattern-mathcing+replacement would just be skipped, but now these are asserts that fail and raise. This is causing concrete issues in some graphs where 2-dimensional device meshes are being used (e.g., TP + CP) but only one dimension has symm-mem enabled. See #158569. This PR is turning these asserts back into harmless early-exits. Note that this only needed to be done for reduce-scatters, as it was already the case for all-gathers. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158572 Approved by: https://github.com/danielvegamyhre, https://github.com/atalman	2025-07-19 04:54:38 +00:00
Laith Sakka	64dabb2cf5	only fail regressions>10% on pr_time benchmarks (#158577 ) Moving to a new framework, maintaitning the pr_time benchmark test right now is hard and often breaking. 1. only fail PRs >10% regressions. 2. post monitor with pr_time benchmarks dashboard (oncall), and update expected results (frequently or on big changes) (supposed to already be doing https://www.internalfb.com/unidash/dashboard/pt2_diff_time_metrics) 3. setting up some one detections detectors warnings that would be triggered at regressions and notify internally post land https://www.internalfb.com/monitoring/detector/1140915271179237 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158577 Approved by: https://github.com/xmfan, https://github.com/janeyx99	2025-07-19 04:35:31 +00:00
Tristan Rice	ab557421a4	[cca] [c10d] Refactor CUDAEventCache into separate files (#158616 ) Summary: Refactored CUDAEventCache from ProcessGroupNCCL.hpp/.cpp into dedicated header and implementation files for better code organization and maintainability. Split out CUDAEventCache into: - New header file: CUDAEventCache.hpp - New implementation file: CUDAEventCache.cpp - Updated build_variables.bzl to include the new file This change improves code maintainability, readability, and follows better code organization practices. --- > Generated by [Confucius Code Assist (CCA)](https://www.internalfb.com/wiki/Confucius/Analect/Shared_Analects/Confucius_Code_Assist_(CCA)/) [Session](https://www.internalfb.com/confucius?session_id=61b9029a-636b-11f0-9d9a-f1bcc55be1ce&tab=Chat), [Trace](https://www.internalfb.com/confucius?session_id=61b9029a-636b-11f0-9d9a-f1bcc55be1ce&tab=Trace) Test Plan: Verified build with: ``` buck build //caffe2/test/distributed:c10d ``` --- > Generated by [Confucius Code Assist (CCA)](https://www.internalfb.com/wiki/Confucius/Analect/Shared_Analects/Confucius_Code_Assist_(CCA)/) [Session](https://www.internalfb.com/confucius?session_id=61b9029a-636b-11f0-9d9a-f1bcc55be1ce&tab=Chat), [Trace](https://www.internalfb.com/confucius?session_id=61b9029a-636b-11f0-9d9a-f1bcc55be1ce&tab=Trace) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158616 Approved by: https://github.com/fduwjj	2025-07-19 02:51:28 +00:00
Laith Sakka	90b082e207	enable_caching_generated_triton_templates=True by default (#158592 ) Got some risk, but good to catch issues if there is any, easy to revert single flag flip. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158592 Approved by: https://github.com/eellison	2025-07-19 02:19:34 +00:00
Huy Do	a741094159	Build domain libraries on the build job (#158600 ) By setting the name of the domain libraries to build via `BUILD_ADDITIONAL_PACKAGES` environment variable, the build job will build them and make them available as artifacts in the same way as the PyTorch CI wheel. To ensure that this doesn't break CI, the test job will still build them as usual if the wheels are not there. Building dependencies like FBGEMM on the test job is bad, especially for GPU jobs, because it leave the GPU resource idle Fixes https://github.com/pytorch/pytorch/issues/152024 Signed-off-by: Huy Do <huydhn@gmail.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158600 Approved by: https://github.com/yangw-dev ghstack dependencies: #158598, #158599	2025-07-19 02:03:50 +00:00
Huy Do	2955acaed6	Clean up some unused build env variables (#158599 ) * Parameter build-with-debug isn't needed, it isn't even passed into Docker. Debug build is detected via the build environment name * AWS_DEFAULT_REGION is a leftover from ARC and isn't used anywhere in .ci/pytorch nor .github Signed-off-by: Huy Do <huydhn@gmail.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158599 Approved by: https://github.com/cyyever, https://github.com/ZainRizvi ghstack dependencies: #158598	2025-07-19 01:59:00 +00:00
Ryan Guo	2c16eb9f3d	[dynamo] Support more basic output types for `nonstrict_trace` (#157969 ) Fixes #157397 and improves the user-facing error message for remaining unsupported cases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157969 Approved by: https://github.com/zou3519	2025-07-19 00:59:54 +00:00
PyTorch MergeBot	c2c88846a9	Revert "[Easy] Show some clear error when torch.ops.load_library fails. (#157524 )" This reverts commit 555f3562541992b66a550eca8e8740884b1247f8. Reverted https://github.com/pytorch/pytorch/pull/157524 on behalf of https://github.com/wdvr due to reverting for now to reopen the discussion ([comment](https://github.com/pytorch/pytorch/pull/157524#issuecomment-3091317252))	2025-07-19 00:45:31 +00:00
PyTorch MergeBot	5b40f6581e	Revert "Add warning about removed sm50 and sm60 arches (#158301 )" This reverts commit fb731fe371cb1b5bf95de84b19c213590526acb2. Reverted https://github.com/pytorch/pytorch/pull/158301 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/158301#issuecomment-3091307023))	2025-07-19 00:32:04 +00:00
Xu Han	d42c409767	[AOTI] windows package load dev (#158671 ) changes: 1. add extract file fail handler for Windows develop. 2. normalize more file paths. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158671 Approved by: https://github.com/angelayi	2025-07-19 00:06:40 +00:00
Will Constable	a3aacd6cb2	[DTensor] fix copy_ strategy (#158538 ) The previous strategy directly used 'self' input strategy for 'src' input. The fixed strategy correctly maps the self dim to src dim so that it works even if the src input is broadcast. E.g. for this program, broadcasting will occur on dims 0,1,3 of self. ``` self = torch.ones((2,3,4,5)) src = torch.ones((4,1)) self.copy_(src) ``` These are the correct sharding combinations: \| self \| src \| \|-------\|------\| \| Shard(0) \| Replicate() \| \| Shard(1) \| Replicate() \| \| Shard(2) \| Shard(0) \| \| Shard(3) \| Shard(1) \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158538 Approved by: https://github.com/zpcore, https://github.com/XilunWu, https://github.com/wanchaol ghstack dependencies: #158490	2025-07-18 23:44:43 +00:00
Will Constable	36bddcd18c	[DTensor] Fix default_strategy and rename for clarity (#158490 ) Fixes several bugs in the original. - foremost, fixes a serious bug where we returned incorrect strategies by mixing input_specs that were frozen from select_strategy.strategies[0] with output_specs that varied across select_strategy.strategies[0..N] (e.g. we could create a nonsense strategy like input:Shard(0) output(Replicate) for an op like clone - fixes the redistribute costs: they should not actually be 0, they should be the cost of redistributing our single input from another strategy to the current strategy, in our list of output strategies - adds a note, wondering if we should have just literally returned the input strategy instead of creating this new object - Currently, using default_strategy is incorrect becuase it maps 'self' tensor's strategies directly onto 'src' tensor without accounting for the fact that copy_ supports broadcasting a smaller rank tensor into a larger one. Separates out copy_ op from default strategy, adds missing test case, but does not fix the underlying issue with copy_, leaves that for future PR Renames to `propagate_single_input_strategy` since that's more descriptive Pull Request resolved: https://github.com/pytorch/pytorch/pull/158490 Approved by: https://github.com/wanchaol, https://github.com/XilunWu	2025-07-18 23:44:42 +00:00
AaronWang04	15ef4f28df	Fused RMSNorm implementation (#153666 ) Relevant #72643 Benchmarked versus unfused torch implementation and torch.compile implementation. Around 9x speedup vs unfused implementation on cuda and slightly faster vs inductor compile on 5090. ```py import torch import torch.nn as nn class RMSNorm(nn.Module): def __init__(self, dim, eps=1e-5): super().__init__() self.eps = eps self.scale = nn.Parameter(torch.ones(dim)) def forward(self, x): norm_x = x.norm(2, dim=-1, keepdim=True) rms_x = norm_x * torch.rsqrt(torch.tensor(x.shape[-1], dtype=x.dtype)) x_normed = x / (rms_x + self.eps) return self.scale * x_normed def benchmark_rmsnorm_cuda(input_shape, normalized_dim, num_iterations=100, warmup_iterations=10, dtype=torch.float16): rms_norm_layer = torch.nn.RMSNorm(normalized_dim, device='cuda', dtype=dtype) input_data = torch.randn(input_shape, device='cuda', dtype=dtype) for _ in range(warmup_iterations): _ = rms_norm_layer(input_data) torch.cuda.synchronize() start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() for _ in range(num_iterations): _ = rms_norm_layer(input_data) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = elapsed_time_ms / num_iterations print(f"--- RMSNorm CUDA Benchmark ---") print(f"Input Shape: {input_shape}") print(f"Normalized Dimension: {normalized_dim}") print(f"Benchmark Iterations: {num_iterations}") print(f"--- Fused Implementation ---") print(f"Average Time per Iteration: {avg_time_ms:.4f} ms") print(f"Total Time for {num_iterations} Iterations: {elapsed_time_ms:.3f} ms") compiled_rms_norm = torch.compile(RMSNorm(dim=normalized_dim)).cuda() for _ in range(warmup_iterations): _ = compiled_rms_norm(input_data) torch.cuda.synchronize() start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() for _ in range(num_iterations): _ = compiled_rms_norm(input_data) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = elapsed_time_ms / num_iterations print(f"--- TorchCompile Implementation ---") print(f"Average Time per Iteration: {avg_time_ms:.4f} ms") print(f"Total Time for {num_iterations} Iterations: {elapsed_time_ms:.3f} ms") print("-" * 50) if __name__ == '__main__': parameter_sets = [ {'batch_size': 16, 'sequence_length': 256, 'hidden_features': 512, 'dtype': torch.float16}, {'batch_size': 32, 'sequence_length': 512, 'hidden_features': 768, 'dtype': torch.float16}, {'batch_size': 64, 'sequence_length': 1024, 'hidden_features': 1024, 'dtype': torch.float16}, {'batch_size': 32, 'sequence_length': 512, 'hidden_features': 768, 'dtype': torch.float32}, {'batch_size': 8, 'sequence_length': 2048, 'hidden_features': 2048, 'dtype': torch.float16}, ] num_benchmark_iterations = 200 num_warmup_iterations = 20 for params in parameter_sets: batch_size = params['batch_size'] sequence_length = params['sequence_length'] hidden_features = params['hidden_features'] data_type = params.get('dtype', torch.float16) shape = (batch_size, sequence_length, hidden_features) norm_dim_to_normalize = hidden_features print(f"Benchmarking with: BS={batch_size}, SeqLen={sequence_length}, Hidden={hidden_features}, DType={data_type}") benchmark_rmsnorm_cuda(input_shape=shape, normalized_dim=norm_dim_to_normalize, num_iterations=num_benchmark_iterations, warmup_iterations=num_warmup_iterations, dtype=data_type) ``` Here are the triton compile tests ran on a 5090 (comparing this branch vs main) ```py import torch import torch.nn as nn from torch._inductor.utils import run_and_get_code, run_fw_bw_and_get_code torch.manual_seed(0) device = torch.device("cuda") for batch in range(0, 9): for i in range(9, 16): normalized_shape_arg = (2batch, 2i) input_tensor = torch.randn(2batch, 2i, device=device, requires_grad=True) weight_tensor = torch.randn(2batch, 2i,device=device, requires_grad=True) model = torch.nn.functional.rms_norm compiled_model = torch.compile(model) loss = torch.randn_like(input_tensor) num_iter = 5 for j in range(num_iter): output = compiled_model(input_tensor, normalized_shape_arg, weight_tensor) output.backward(loss) start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() num_iter = 10 for j in range(num_iter): output = compiled_model(input_tensor, normalized_shape_arg, weight_tensor) output.backward(loss) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = round(elapsed_time_ms / num_iter, 5) print(2batch, 2i, avg_time_ms) ``` main ``` 32 512 0.1812 32 1024 0.19021 32 2048 0.18871 32 4096 0.17019 32 8192 0.21944 32 16384 0.38871 32 32768 0.83282 64 512 0.14705 64 1024 0.13987 64 2048 0.14111 64 4096 0.21699 64 8192 0.43141 64 16384 0.90652 64 32768 2.18573 128 512 0.19361 128 1024 0.1963 128 2048 0.20122 128 4096 0.38888 128 8192 0.93795 128 16384 2.23437 128 32768 5.50079 256 512 0.16722 256 1024 0.22856 256 2048 0.39421 256 4096 0.96621 256 8192 2.48746 256 16384 5.53571 256 32768 11.97932 ``` current branch ``` 32 512 0.16328 32 1024 0.18104 32 2048 0.15508 32 4096 0.14356 32 8192 0.20111 32 16384 0.45974 32 32768 0.94799 64 512 0.16874 64 1024 0.18701 64 2048 0.16107 64 4096 0.20152 64 8192 0.46568 64 16384 0.96599 64 32768 2.21661 128 512 0.14982 128 1024 0.15565 128 2048 0.22241 128 4096 0.46128 128 8192 0.88883 128 16384 2.3097 128 32768 5.84448 256 512 0.14346 256 1024 0.2007 256 2048 0.45927 256 4096 0.87876 256 8192 2.10571 256 16384 5.73948 256 32768 12.98581 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153666 Approved by: https://github.com/ngimel, https://github.com/eqy, https://github.com/albanD	2025-07-18 23:24:21 +00:00
Grace Cheng	60b9b06a53	[caffe2] Fix Missing override in get_buffer of NCCLSymmetricMemory (#158597 ) Summary: Fix the error that occurs in the devarm environment when compiling with Clang: ``` caffe2/torch/csrc/distributed/c10d/symm_mem/NCCLSymmetricMemory.cu:97:20: error: 'get_buffer' overrides a member function but is not marked 'override' [-Werror,-Winconsistent-missing-override] 97 \| virtual at::Tensor get_buffer(int \| ^ caffe2/torch/csrc/distributed/c10d/symm_mem/SymmetricMemory.hpp:56:20: note: overridden virtual function is here 56 \| virtual at::Tensor get_buffer(int rank, c10::IntArrayRef sizes, c10::ScalarType dtype, int64_t storage_offset) = 0; \| ^ 1 error generated. ``` Test Plan: See D78520305 Rollback Plan: Differential Revision: D78517953 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158597 Approved by: https://github.com/janeyx99	2025-07-18 23:12:29 +00:00
fduwjj	a835dbc096	[c10d][ez] Fix error message to reflect the correct API name (#158668 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158668 Approved by: https://github.com/VieEeEw	2025-07-18 23:10:47 +00:00
Yang Wang	f76f4abf3f	Track monitor (#156907 ) Tracking gpu mem allocation, we were tracking the gpu bandwidth memory, the mem allocation is the one reflect wether the gpu is oom or not, upcoming ui fix. UI fix: https://github.com/pytorch/test-infra/pull/6878/files Pull Request resolved: https://github.com/pytorch/pytorch/pull/156907 Approved by: https://github.com/huydhn	2025-07-18 22:54:13 +00:00
Yang Wang	be483a5481	setup pinned commit for vllm in pytorch ci (#158591 ) Set up pinned commit for vllm in nightly Pull Request resolved: https://github.com/pytorch/pytorch/pull/158591 Approved by: https://github.com/seemethere, https://github.com/huydhn	2025-07-18 22:30:20 +00:00
Huamin Li	bc7b1f5252	[AOTI] Use libstdc++ only for fbcode cpu case (#158659 ) Differential Revision: D78567218 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158659 Approved by: https://github.com/kflu, https://github.com/zoranzhao	2025-07-18 22:27:10 +00:00
Simon Fan	07c4c2a792	[dynamo][be] hide warnings without invalidating warnings cache (#158520 ) I feel uneasy about touching `__warningregistry__` since it is undocumented and private surface. The only public API hook that doesn't increment warnings version seems to be https://docs.python.org/3/library/warnings.html#warnings.showwarning. So we could wack a mole all the warnings muters in compile to just not display warnings, and we wouldn't invalidate warnings cache. This PR adds it for torch/_dynamo, and I didn't find any warnings versioning mutation from torch/_inductor. There is a behavior change if someone calls a compiled graph with simplefilter("error"): ```python # e.g. test/dynamo_expected_failures/TestAutogradFallback.test_no_autograd_kernel_inplace_mode_nothing with warnings.catch_warnings(): warnings.simplefilter("error") # turns all warnings into errors compiled_fn() # will throw if any of the muted warnings fire ``` FIXES https://github.com/pytorch/pytorch/issues/128427 A note for the future: The warnings module doesn't offer a thread safe way of using it. Even regular filters have this problem, directly editing `__warningregistry__` would be very bad, and this PR would mute all threads. Someone will need to build a thread safe warnings interface. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158520 Approved by: https://github.com/anijain2305, https://github.com/zou3519	2025-07-18 22:02:31 +00:00
Michael Lazos	89850bbc07	[Dynamo] Use proper sources for constructing dataclass defaults (#157993 ) Partially fixes https://github.com/pytorch/pytorch/issues/154009 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157993 Approved by: https://github.com/williamwen42, https://github.com/anijain2305	2025-07-18 21:51:40 +00:00
PyTorch MergeBot	3bb729df97	Revert "Fix test consolidate hf safetensors (#157386 )" This reverts commit fa1c20ae9285f7994a73d2d06025065f96b67a57. Reverted https://github.com/pytorch/pytorch/pull/157386 on behalf of https://github.com/jithunnair-amd due to Need to revert this so we can revert PR 156705, which introduced errors on ROCm CI. These errors were not seen on CUDA CI because CUDA CI docker images do not have safetensors installed and the test silently passes ([comment](https://github.com/pytorch/pytorch/pull/157386#issuecomment-3090706074))	2025-07-18 21:00:12 +00:00
PyTorch MergeBot	e3351b3ddf	Revert "[DCP][HF] [ez]Change where sharded tensors are saved (#158069 )" This reverts commit 627ba411366bcc15019c49756d3f22fd3914bd50. Reverted https://github.com/pytorch/pytorch/pull/158069 on behalf of https://github.com/jithunnair-amd due to Didn't remove reference to `consolidated_output_path` in test_hf_safetensor_e2e.py; CUDA runs do not surface issue because safetensors is not installed and the test silently passes ([comment](https://github.com/pytorch/pytorch/pull/158069#issuecomment-3090692336))	2025-07-18 20:54:19 +00:00
Huy Do	1ab1ab38a0	Use linux.12xlarge.memory to build for H100/sm_90 (#158598 ) Use a bigger runner here because CUDA_ARCH 9.0 is only built for H100 or newer GPUs, so it doesn't benefit much from existing compiler cache from trunk. Also use a memory-intensive runner here because memory is usually the bottleneck Signed-off-by: Huy Do <huydhn@gmail.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158598 Approved by: https://github.com/ZainRizvi, https://github.com/malfet	2025-07-18 20:31:56 +00:00
Colin L Reliability Rice	8b2a650572	pt2_remote_cache: Log sample for failures, and log the explicit reason we're faling. (#156874 ) Summary: This allows us to start alerting on cache failures, based on scuba data Test Plan: Added new tests explicitly for the Remote Cache API. Note that we have existing tests for memcache, but not for manifold AFAICT. There are two potential wrinkles. One we're adding a new field (and everything uses ScubaData AFAICT, so this should just work). The other one is the implicit api contract that if the sample is None, then it will be ignored (and not crash). I believe the second one is implemented correctly (and tested). The first one is a little more nebulous, but I think won't cause any breakages. Also manually ran a compile and made sure it didn't break - P1851504490 as well as forcing it to break and checking we didn't screw up the exception handling - P1851504243 Rollback Plan: Differential Revision: D77054339 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156874 Approved by: https://github.com/oulgen, https://github.com/masnesral	2025-07-18 20:28:27 +00:00
Apostolos Kokolis	ec0b538961	[inductor] Make times and repeat parameters command line args (#158590 ) Summary: Small change to make the `times` and `repeat` variables controllable as command line args. Test Plan: Execute: ``` buck2 run <run params> <path>:inductor_benchmark -- --times=1 --repeat=1 ``` Only runs once, and without passing the args it runs with default values of 10. Rollback Plan: Reviewed By: malfet Differential Revision: D78458680 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158590 Approved by: https://github.com/FindHao, https://github.com/malfet	2025-07-18 20:07:55 +00:00
Xu Han	599f94e7b9	[AOTI] add Windows file ext to package loader. (#158578 ) Add `object` and `extension` file type for Windows Pull Request resolved: https://github.com/pytorch/pytorch/pull/158578 Approved by: https://github.com/angelayi	2025-07-18 19:57:12 +00:00
Sam Larsen	04ac258cf6	[BE][testing] Fix test_cudacodecache.py (#158259 ) Summary: According to internal test failures, looks like we're missing a check for cuda: https://fburl.com/testinfra/eznzkyha Test Plan:c`buck test` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158259 Approved by: https://github.com/exclamaforte, https://github.com/BoyuanFeng	2025-07-18 19:56:13 +00:00
Zain Rizvi	1b5fdb23b9	[BE] Add pre-push hook for lintrunner to the PyTorch repo (#158389 ) Adds a pre-commit hook (technically a pre-push hook) to the PyTorch repo. This is currently an opt-in feature, which one can opt into by running `python scripts/setup_hooks.py` locally. ### Features - Run Lintrunner Before Push: Before every `git push`, automatically runs lintrunner on your changes. - Really need to skip the checks? Run `git push --no-verify` - Consistent, Isolated, Lintrunner Environment: During pre-push, Lintrunner runs in it's own virtual en environment that contain all lintrunner dependencies in a consistent, isolated environment. No more lintrunner failures because you created a new .venv. (Did you know you needed to run `lintrunner init` every time you make a new .venv?) - Dependencies Automatically Updated: If .lintrunner.toml is updated, this will automatically re-run `lintrunner init` to ensure you install the latest dependencies specified ### Installation - Run `python scripts/setup_hooks.py`. Now every `git push` will first run lintrunner. ### Additional details - The lintrunner used by the pre-push hook runs in a special per-repo virtual environment managed by the commit-hook tool located under `$USER/.cache/pre-commit` - Does not affect your regularly used lintrunner - Manual invocations of lintrunner will continue to depend on your local environment instead of the special pre-push one. If there's enough interest, we could explore consolidating them. - Does not run `lintrunner -a` for you. - You still need to manually run that (can be changed later though!) - Have staged/unstaged changes? No worries - This runs `git stash` before running the pre-commit hooks and pops back your changes afterwards, so only the changes actaully being pushed will be tested ### Downsides - No streaming UI updates - While you still get the same output from lintrunner that you're used to, the commit-hook framework doesn't show any output while lintrunner is actually running. Instead, it shows the entire output after linter has completed execution, which could be a few minutes (especially if it has to run `lintrunner init` first) - `uv` installation is required to run the setup script. The setup script will ask users to install uv if it's not available. - This is required to be able to install the pre-commit package in a safe way that's available no matter what .venv you are running in. ### Opting out - Disable hook for a single push: Run `git push --no-verify` - Disable hook permanently: If something goes wrong and you need to wipe your setup: - Delete the `$USER/.cache/pre-commit` folder and the `.git/hooks/pre-push` file in your local repo. - You can now rerun `python scripts/setup_hooks.py` to setup your git push hook again if you want. ### Potential Future Changes Things that could be done to make this even better if folks like these ideas: - Automatic setup - Our `CONTRIBUTING.md` file tells devs to run `make setup-env`. That could be a good entry point to hook the installation into - Fix the console output streaming - Make every lintrunner invocation (including manual ones) use the same repo-specific venv that the commit-hook uses. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158389 Approved by: https://github.com/seemethere	2025-07-18 19:55:35 +00:00
dsashidh	75e2628782	Add lower bounds for fsspec and networkx dependencies (#158565 ) Fixes #156587 This sets lower bounds for fsspec and networkx in both setup.py and requirements,txt. - fsspec>= 0.8.5 (released December 15, 2020) - netowrkx>= 2.5.1 (released April 3, 2021) These are the first stable versions released after Python 3.9 came out on October 5, 2020. Since Python 3.8 is no longer maintained, setting these minimums helps ensure PyTorch won't be installed alongside unexpectedly old versions of these packages. Tested with these versions locally to make sure they don't break anything. Adding CI for lower-bound testing could be a follow up later if need. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158565 Approved by: https://github.com/janeyx99	2025-07-18 19:42:09 +00:00
Svetlana Karslioglu	79e49efadd	Pull latest Sphinx theme (#158595 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158595 Approved by: https://github.com/albanD	2025-07-18 18:46:47 +00:00
Sam Larsen	b87e50db5e	[BE][testing] Fix internal test failures in test/dynamo/test_unspec (#158485 ) Summary: These tests failing internally because the number of underlying calls to the rng differ by virtue of various library initializations that get sucked in with an internal build. Test Plan: ``` buck test '@fbcode//mode/opt' fbcode//caffe2/test/dynamo:test_dynamo -- --exact 'caffe2/test/dynamo:test_dynamo - test_unspec.py::UnspecTests::test_random_object' --run-disabled buck test '@fbcode//mode/opt' fbcode//caffe2/test/dynamo:test_dynamo -- --exact 'caffe2/test/dynamo:test_dynamo - test_unspec.py::UnspecTests::test_random_values_with_graph_break' --run-disabled buck test '@fbcode//mode/opt' fbcode//caffe2/test/dynamo:test_dynamo -- --exact 'caffe2/test/dynamo:test_dynamo - test_unspec.py::UnspecTests::test_feed_random_values_into_graph_only' --run-disabled ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158485 Approved by: https://github.com/williamwen42	2025-07-18 18:41:03 +00:00
Lucas Kabela	656885b614	[Dynamo][Better Engineering] Type devices, resume_execution and testing utils (#158593 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to a set of utilities in dynamo, `device_interface.py`, `resume_execution.py`, `tensor_version_ops.py`, `test_case.py`, and `test_minifier_common.py` Running ``` mypy torch/_dynamo/device_interface.py torch/_dynamo/resume_execution.py torch/_dynamo/tensor_version_op.py torch/_dynamo/test_case.py torch/_dynamo/test_minifier_common.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 976 \| 1672 \| 58.37% \| 76 \| 112 \| 67.86% \| \| This PR \| 1719 \| 1719 \| 100.00% \| 112 \| 112 \| 100.00% \| \| Delta \| +743 \| +47 \| +41.63% \| +36 \| 0 \| +32.14% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158593 Approved by: https://github.com/mlazos	2025-07-18 18:22:06 +00:00
Lucas Kabela	6e07d6a0ff	[Dynamo][Better Engineering] Add typing support for _dynamo/repro and debug_utils (#158504 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to an important set of utilities in dynamo, `repro/` and the base `debug_utils.py` Running ``` mypy torch/_dynamo/repro/ torch/_dynamo/debug_utils.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 905 \| 3268 \| 27.69% \| 22 \| 81 \| 27.16% \| \| This PR \| 3368 \| 3368 \| 100.00% \| 81 \| 81 \| 100.00% \| \| Delta \| +2463 \| +100 \| +72.31% \| +59 \| 0 \| +72.84% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158504 Approved by: https://github.com/mlazos	2025-07-18 18:15:55 +00:00
yuchengliu1	b4358c5e87	[inductor] Explicitly link c10 in inductor. (#158622 ) MSVC have error "unresolved external symbol" when compiling inductor. Explicitly link c10 in inductor. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158622 Approved by: https://github.com/desertfire Co-authored-by: Xu Han <xu.han@outlook.com>	2025-07-18 18:00:50 +00:00
PyTorch MergeBot	86675af3f0	Revert "[ROCm][CI] update fbgemm_gpu hash used by inductor tests (#158602 )" This reverts commit 9308261a2afb69d807ea06508bb8582b066d9ccd. Reverted https://github.com/pytorch/pytorch/pull/158602 on behalf of https://github.com/ZainRizvi due to The lint job failure was hiding a real lint failure. See here for more details: [GH job link](https://github.com/pytorch/pytorch/actions/runs/16375911199/job/46275682191) [HUD commit link](`6f73e06796`) ([comment](https://github.com/pytorch/pytorch/pull/158602#issuecomment-3090209891))	2025-07-18 17:46:11 +00:00
Deepak Seshadri	725cdb218e	Name threads in caffe2/torch/distributed/checkpoint AsyncCheckpointExecutor (#158612 ) Differential Revision: D78493333 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158612 Approved by: https://github.com/d4l3k	2025-07-18 17:33:12 +00:00
Xu Han	8c3f84908b	[aot] fix greater_than_max build fail on Windows. (#158479 ) Error snapshot: <img width="937" height="110" alt="image" src="https://github.com/user-attachments/assets/10195f84-83c4-42db-af3c-76f875a6a983" /> Reason: `std::numeric_limits::max` is confilct to windef.h:`max(a, b)` Fix code: <img width="488" height="269" alt="image" src="https://github.com/user-attachments/assets/3328c37b-7c89-435e-944c-4ca7c9b6c5b6" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158479 Approved by: https://github.com/desertfire	2025-07-18 17:18:10 +00:00
Guilherme Leobas	6f73e06796	[iter] exhaust `ListIterator` when `unpack_var_sequence` is called (#156370 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156370 Approved by: https://github.com/zou3519 ghstack dependencies: #156369	2025-07-18 16:48:27 +00:00
Guilherme Leobas	acffd1a297	[iter] Update some of the tests to not call pickle (#156369 ) Some tests in test_iter only fail because of pickle. I'm skipping the pickle section as Dynamo doesn't support it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156369 Approved by: https://github.com/zou3519	2025-07-18 16:48:27 +00:00
PyTorch MergeBot	bf4aa78279	Revert "[DTensor] Fix default_strategy and rename for clarity (#158490 )" This reverts commit d8b084312b54e97bdbaf6a178fe2fc628a23243b. Reverted https://github.com/pytorch/pytorch/pull/158490 on behalf of https://github.com/clee2000 due to broke lint? [GH job link](https://github.com/pytorch/pytorch/actions/runs/16361950974/job/46231492581) [HUD commit link](`d8b084312b`) ([comment](https://github.com/pytorch/pytorch/pull/158490#issuecomment-3090042448))	2025-07-18 16:45:32 +00:00
PyTorch MergeBot	50f33a6fca	Revert "[DTensor] fix copy_ strategy (#158538 )" This reverts commit 7b05bdd925f0f4b49e68662f9761fabaa27f2faf. Reverted https://github.com/pytorch/pytorch/pull/158538 on behalf of https://github.com/clee2000 due to broke lint? [GH job link](https://github.com/pytorch/pytorch/actions/runs/16361950974/job/46231492581) [HUD commit link](`d8b084312b`) ([comment](https://github.com/pytorch/pytorch/pull/158490#issuecomment-3090042448))	2025-07-18 16:45:32 +00:00
Xu Han	35df895d05	[AOTI] package loader normalize path separator (#158630 ) Add `normalize_path_separator` to handle Windows path simplify. This solution is working well on `torch/_inductor/cpp_builder.py`: `a00cd8cf25/torch/_inductor/cpp_builder.py (L406-L409)` Let's copy it to package loader. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158630 Approved by: https://github.com/angelayi	2025-07-18 15:55:24 +00:00
Zain Rizvi	193b29ee0c	[BE][EZ] Minor doc fixes (#158574 ) [BE] Minor doc fixes	2025-07-18 10:34:55 -05:00
Zhengxu Chen	036eb1f65d	[precompile] Filter out ID_MATCH family of guards with caching_precompile. (#158368 ) Summary: For case like caching_precompile, we almost always want to drop ID_MATCH-type guards since they will block serialization. This diff add this behavior when this global flag is toggled on so that ID_MATCH guards are excluded from compilation and serialization. Test Plan: test_dynamo -- -k test_id_match_with_config Rollback Plan: Differential Revision: D78363609 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158368 Approved by: https://github.com/jamesjwu	2025-07-18 14:47:11 +00:00
Jane Xu	e882c761dd	Add STD_TORCH_CHECK to headeronly (#158377 ) Differential Revision: [D78366519](https://our.internmc.facebook.com/intern/diff/D78366519/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158377 Approved by: https://github.com/albanD	2025-07-18 14:35:20 +00:00
bobrenjc93	0eae6b68f4	Unify torch.tensor and torch.ops.aten.scalar_tensor behavior (#158537 ) Fixes #158376 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158537 Approved by: https://github.com/atalman	2025-07-18 14:05:52 +00:00
Xuehai Pan	a4ec381302	[build] pin `setuptools>=77` to enable PEP 639 (#158104 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158104 Approved by: https://github.com/rgommers, https://github.com/Skylion007, https://github.com/atalman	2025-07-18 11:49:54 +00:00
Aidyn-A	27af877f84	[ATen][CUDA][SDPA] Flash Attention: Refactor sm version checks (#158558 ) The architecture version checks are unnecessary fine-grained in PyTorch. Considering the fact that PyTorch's Flash Attention works on all `sm_80+` machines, it makes more sense to just check for lower bound. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158558 Approved by: https://github.com/eqy	2025-07-18 09:59:41 +00:00
Will Constable	7b05bdd925	[DTensor] fix copy_ strategy (#158538 ) The previous strategy directly used 'self' input strategy for 'src' input. The fixed strategy correctly maps the self dim to src dim so that it works even if the src input is broadcast. E.g. for this program, broadcasting will occur on dims 0,1,3 of self. ``` self = torch.ones((2,3,4,5)) src = torch.ones((4,1)) self.copy_(src) ``` These are the correct sharding combinations: \| self \| src \| \|-------\|------\| \| Shard(0) \| Replicate() \| \| Shard(1) \| Replicate() \| \| Shard(2) \| Shard(0) \| \| Shard(3) \| Shard(1) \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158538 Approved by: https://github.com/zpcore, https://github.com/XilunWu, https://github.com/wanchaol ghstack dependencies: #158495, #158490	2025-07-18 09:59:37 +00:00
Aleksei Nikiforov	ead80f3202	Fix s390x CI: ensure that all python dependencies are installed when … (#158552 ) …building pytorch for tests on s390x Pull Request resolved: https://github.com/pytorch/pytorch/pull/158552 Approved by: https://github.com/huydhn	2025-07-18 09:13:41 +00:00
PyTorch MergeBot	32aade9d8d	Revert "Support DeepSeek-style blockwise scaling scaled-mm for fp8 on Hopper+ (#158037 )" This reverts commit 39ac189808c61588f3594dbc2fc1d69bb6194c47. Reverted https://github.com/pytorch/pytorch/pull/158037 on behalf of https://github.com/jithunnair-amd due to Ignored ROCm failures while ROCm was unstable, but HUD clearly shows this PR introduced failures on trunk ([comment](https://github.com/pytorch/pytorch/pull/158037#issuecomment-3087982975))	2025-07-18 07:47:46 +00:00
PyTorch MergeBot	be896d6b41	Revert "Forward-fix unused variables warning/error (#158549 )" This reverts commit eeda1a75ace75ce8a6763050fb91d236a6d3287b. Reverted https://github.com/pytorch/pytorch/pull/158549 on behalf of https://github.com/jithunnair-amd due to Sorry, need to revert this first, so we can revert PR 158037, which broke ROCm CI ([comment](https://github.com/pytorch/pytorch/pull/158549#issuecomment-3087942475))	2025-07-18 07:44:14 +00:00
Yidi Wu	a3396a9b85	[hop] set capture_scalar_outputs=True by default for compiled hops (#158480 ) We want to do it for two reasons: 1. It's tedious for users to manually turn on capture_scalar_outputs=True when compiling map and scan with inductor, where we decomposing them into while_loop and use the idx tensor.item() to select a slice of output buffer and write into it. This pr turns on the flag by default. 2. a graph break caused by capture_scalar_outputs=False would cause the hop to fail, and we should turn it on by default so that the error message is more meaningful. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158480 Approved by: https://github.com/zou3519	2025-07-18 07:16:50 +00:00
Yidi Wu	fda3f3b2ec	[while_loop] fix constant tensor used as carried inputs (#158381 ) Address second part of #158366, where torch.tensor(0), is treated as a constant tensor and its .item() gets specailized to 0 which causes a silent specialization. The fix is to unspecialize the constant carries and make them non-constant. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158381 Approved by: https://github.com/zou3519	2025-07-18 07:08:11 +00:00
drisspg	a00cd8cf25	Add a way to disable compile for debugging flex-attention (#158534 ) Finally got around to doing this, this flag lets us do: ```Python #!/usr/bin/env python3 """ FlexAttention Debug: Using breakpoints and unwrap """ import torch import torch.nn.attention.flex_attention as fa unwrap = torch._C._functorch.get_unwrapped def score_mod(score, batch, head, q_idx, kv_idx): # Set breakpoint here to debug breakpoint() # In debugger, unwrap to see actual tensor values: # >>> actual_score = unwrap(unwrap(unwrap(unwrap(score)))) # >>> actual_batch = unwrap(batch) # >>> actual_head = unwrap(head) # >>> actual_q_idx = unwrap(q_idx) # >>> actual_kv_idx = unwrap(kv_idx) # >>> print(actual_score) # >>> print(f"q_idx: {actual_q_idx}, kv_idx: {actual_kv_idx}") return torch.where(q_idx >= kv_idx, score, torch.tensor(float('-inf'))) def main(): # Enable debug mode fa._FLEX_ATTENTION_DISABLE_COMPILE_DEBUG = True # Small example B, H, S, D = 1, 2, 4, 8 q = torch.randn(B, H, S, D) k = torch.randn(B, H, S, D) v = torch.randn(B, H, S, D) # Run - will hit breakpoint output = fa.flex_attention(q, k, v, score_mod=score_mod) # Disable debug mode fa._FLEX_ATTENTION_DISABLE_COMPILE_DEBUG = False if __name__ == "__main__": main() ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158534 Approved by: https://github.com/Chillee, https://github.com/zou3519	2025-07-18 05:33:45 +00:00
PaliC	eb73650723	[BE] Make PyObjectSlot use a global PyInterpreter and remove (#158427 ) This PR is a bit more involved but effectively works to drastically simplify PyObjectSlot and PyInterpreter. 1) For PyObjectSlot we now use a global pyinterpreter since there only is one. From here we change all of the call sites to rely on this assumption. 2) We also remove the "tags" of the PyInterpreter by deprecating `PyInterpreterStatus`. For the reviewer, sadly it seems like `functorch/csrc/dim/dim.cpp` needed to get linted, so there is an unreadable amount of changes there. Fortunately, the only actual change in the file is as follows which just removes `getPyInterpreter()` from the `check_pyobj` call. ``` mpy::handle handle_from_tensor(Arena& A, TensorRef t) { - // fast case: tensor is live in python - std::optional<PyObject> mb_obj = - t->unsafeGetTensorImpl()->pyobj_slot()->check_pyobj(getPyInterpreter(), /ignore_hermetic_tls=/false); - if (mb_obj.has_value() && !t->unsafeGetTensorImpl()->pyobj_slot()->owns_pyobj()) { - return mb_obj; - } - return A.autorelease(mpy::object::checked_steal(THPVariable_Wrap(t))); -} -} + // fast case: tensor is live in python + std::optional<PyObject> mb_obj = + t->unsafeGetTensorImpl()->pyobj_slot()->check_pyobj( + /ignore_hermetic_tls=/false); + if (mb_obj.has_value() && + !t->unsafeGetTensorImpl()->pyobj_slot()->owns_pyobj()) { + return mb_obj; + } + return A.autorelease(mpy::object::checked_steal(THPVariable_Wrap(t))); +} ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158427 Approved by: https://github.com/albanD	2025-07-18 05:23:00 +00:00
Jeff Daily	9308261a2a	[ROCm][CI] update fbgemm_gpu hash used by inductor tests (#158602 ) fbgemm_gpu build started failing with asmjit errors. Moving to latest tip of fbgemm for inductor tests resolves the build failures. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158602 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-18 05:02:31 +00:00
PyTorch MergeBot	9a7c2f1f64	Revert "Add torch compile force disable caches alias (#158072 )" This reverts commit 2ecf083b7247f265a03ec296ba9d7b795f035118. Reverted https://github.com/pytorch/pytorch/pull/158072 on behalf of https://github.com/jeffdaily due to fails on rocm, signal ignored while rocm was unstable ([comment](https://github.com/pytorch/pytorch/pull/158072#issuecomment-3086740829))	2025-07-18 04:58:24 +00:00
Will Constable	d8b084312b	[DTensor] Fix default_strategy and rename for clarity (#158490 ) Fixes several bugs in the original. - foremost, fixes a serious bug where we returned incorrect strategies by mixing input_specs that were frozen from select_strategy.strategies[0] with output_specs that varied across select_strategy.strategies[0..N] (e.g. we could create a nonsense strategy like input:Shard(0) output(Replicate) for an op like clone - fixes the redistribute costs: they should not actually be 0, they should be the cost of redistributing our single input from another strategy to the current strategy, in our list of output strategies - adds a note, wondering if we should have just literally returned the input strategy instead of creating this new object - Currently, using default_strategy is incorrect becuase it maps 'self' tensor's strategies directly onto 'src' tensor without accounting for the fact that copy_ supports broadcasting a smaller rank tensor into a larger one. Separates out copy_ op from default strategy, adds missing test case, but does not fix the underlying issue with copy_, leaves that for future PR Renames to `propagate_single_input_strategy` since that's more descriptive Pull Request resolved: https://github.com/pytorch/pytorch/pull/158490 Approved by: https://github.com/wanchaol, https://github.com/XilunWu ghstack dependencies: #158495	2025-07-18 04:09:32 +00:00
Shangdi Yu	1e86fa2e5b	Add stack trace to Inductor IR nodes if `inductor.config.trace.provenance_tracing=True` (#158576 ) Summary: - Split `create_mapping` to `create_mapping_pre_post_grad_nodes` and ` create_node_mapping_kernel_to_post_grad` - Store a mapping from pre_grad graph node names to stack traces in `_inductor_pre_grad_node_stack_trace` - Add `stack_traces` member to ir.Node and add it to the string representation of ir.Node - When we create an IR node, if `inductor.config.trace.provenance_tracing=True`, we populate `stack_traces` from `origins`. The nodes in `origins` are post_grad graph nodes. If a node has `node.stack_trace`, we store the stack_trace directly. This is particularly important for backward graph nodes because they don't have a mapping to pre-grad graph nodes. If a node doesn't have `.stack_trace ` (such as `linear`-> `addmm` nodes), we use the stack trace of the pre_grad graph nodes that it maps to. - A post grad graph node might not have stack trace if it correspond to multiple pre grad graph nodes, e.g. [GroupLinearFusion](`a00442421a/torch/_inductor/fx_passes/group_batch_fusion.py (L299)`) Example: ``` scheduling ExternKernelOut( python_kernel_name='extern_kernels.mm', name=buf0, layout=FixedLayout('cuda:0', torch.float32, size=[8, 16], stride=[16, 1]), inputs=[InputBuffer(name='arg2_1', layout=FixedLayout('cuda:0', torch.float32, size=[8, 10], stride=[10, 1])), ReinterpretView( StorageBox( ConstantBuffer(name='fc1_weight', layout=FixedLayout('cuda:0', torch.float32, size=[16, 10], stride=[10, 1])) ), FixedLayout('cuda:0', torch.float32, size=[10, 16], stride=[1, 10]), origins=OrderedSet([mm_default_1]), stack_traces = {, File "/data/users/shangdiy/fbsource/buck-out/v2/gen/fbcode/7b4b7a52e15abb17/scripts/shangdiy/__aot__/aot#link-tree/scripts/shangdiy/aot.py", line 29, in forward, x = self.fc1(x), File "/data/users/shangdiy/fbsource/buck-out/v2/gen/fbcode/7b4b7a52e15abb17/scripts/shangdiy/__aot__/aot#link-tree/torch/nn/modules/linear.py", line 125, in forward, return F.linear(input, self.weight, self.bias), } )], constant_args=(), kwargs={}, output_view=None, python_kernel_name=extern_kernels.mm, cpp_kernel_name=at::mm_out, ordered_kwargs_for_cpp_kernel=(), op_overload=None, arg_properties=[{}, {}], allarg_properties={}, kwarg_properties=None, unbacked_bindings={}, mutation_outputs=[], origin_node=mm_default_1, origins=OrderedSet([mm_default_1]), stack_traces = {, File "/data/users/shangdiy/fbsource/buck-out/v2/gen/fbcode/7b4b7a52e15abb17/scripts/shangdiy/__aot__/aot#link-tree/scripts/shangdiy/aot.py", line 29, in forward, x = self.fc1(x), File "/data/users/shangdiy/fbsource/buck-out/v2/gen/fbcode/7b4b7a52e15abb17/scripts/shangdiy/__aot__/aot#link-tree/torch/nn/modules/linear.py", line 125, in forward, return F.linear(input, self.weight, self.bias), } ) ``` Test Plan: ``` buck2 run mode/dev-nosan fbcode//caffe2/test/inductor:provenance_tracing ``` Rollback Plan: Differential Revision: D78365534 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158576 Approved by: https://github.com/angelayi	2025-07-18 04:05:17 +00:00
Sherlock Huang	86dbc0ef67	[NativeRT] Remove makeProxyExecutor from ModelRunner interface (#158587 ) Summary: makeProxyExecutor shouldn't be exposed to ModelRunner Interface. Test Plan: CI Rollback Plan: Differential Revision: D78501011 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158587 Approved by: https://github.com/yiming0416, https://github.com/henryoier	2025-07-18 03:20:40 +00:00
Will Constable	89d842fec5	Make torch.distributed.breakpoint() set a long timeout (#158481 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158481 Approved by: https://github.com/d4l3k ghstack dependencies: #158469	2025-07-18 02:18:43 +00:00
Will Constable	ce4554352b	Shunt fx_interpreter graphmodule print on error into tlparse (#158469 ) Include both the error stacktrace and the graphmodule in a new structured trace artifact. Log the shortened version to the console, and also log a hint to look at the tlparse for more. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158469 Approved by: https://github.com/ezyang	2025-07-18 02:18:43 +00:00
Lucas Kabela	583138d170	[Dynamo][Better Engineering] Add typing for comptime, cache, and convert_frame (#158379 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to a critical tracing point for dynamo, primarily for`comptime.py` but also `cache_size.py` and `convert_frame.py`. Running ``` mypy torch/_dynamo/comptime.py torch/_dynamo/cache_size.py torch/_dynamo/convert_frame.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 1837 \| 2215 \| 82.93% \| 45 \| 82 \| 54.88% \| \| This PR \| 2230 \| 2230 \| 100.00% \| 82 \| 82 \| 100.00% \| \| Delta \| +393 \| +15 \| +17.07% \| +37 \| 0 \| +45.12% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158379 Approved by: https://github.com/mlazos	2025-07-18 02:11:57 +00:00
eqy	6fd6fc418d	[B200] Fix flex-attention heuristic for `test_tma_with_customer_kernel_options_cuda` (#158494 ) Otherwise fails with ``` torch._inductor.exc.InductorError: RuntimeError: No valid triton configs. OutOfMemoryError: out of resource: triton_tem_fused__to_copy_ones_sort_sum_zeros_2 Required: 264224 Hardware limit: 232448 Reducing block sizes or `num_stages` may help. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158494 Approved by: https://github.com/drisspg	2025-07-18 02:03:49 +00:00
Will Constable	ddbecdfb66	[DTensor] Document redistribute_costs (#158495 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158495 Approved by: https://github.com/zpcore, https://github.com/XilunWu	2025-07-18 01:43:38 +00:00
CaoE	ef38edb284	Add stride check for attn_mask on non-cpu device (#158424 ) Fixes #158374 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158424 Approved by: https://github.com/Valentine233, https://github.com/drisspg, https://github.com/atalman	2025-07-18 01:10:58 +00:00
CaoE	6673ac746c	Fix test linalg for MKL upgrading (#158312 ) Fixes #158054 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158312 Approved by: https://github.com/albanD	2025-07-18 01:08:33 +00:00
Gabriel Ferns	7b72e5b3ad	Fix Pandas version mismatch upon reinstalling numpy (#158584 ) If you reinstall numpy after having installed pandas, it will error out sometimes if the versions are different enough (see below snippet). This change forces pandas to be reinstalled when installing numpy. It doesn't work in a separate pip call, because then pip takes the version of numpy requested by pandas as the one to install, undoing the command in the first place. ``` (numpy_pandas) [gabeferns@devvm2497.eag0 ~/pt-envs/at (exclamaforte/just-gemm-model)]$ pip list Package Version ------------------ ----------- attrs 25.3.0 build 1.2.2.post1 certifi 2025.7.14 charset-normalizer 3.4.2 cmake 4.0.3 exceptiongroup 1.3.0 expecttest 0.3.0 filelock 3.18.0 fsspec 2025.5.1 hypothesis 6.135.32 idna 3.10 importlib_metadata 8.7.0 Jinja2 3.1.6 lintrunner 0.12.7 MarkupSafe 2.1.5 mpmath 1.3.0 networkx 3.2.1 ninja [1.11.1.4](https://www.internalfb.com/phabricator/paste/view/1.11.1.4) opt-einsum 3.3.0 optree 0.16.0 packaging 25.0 pip 25.1 psutil 7.0.0 pyproject_hooks 1.2.0 python-dateutil 2.9.0.post0 pytz 2025.2 PyYAML 6.0.2 requests 2.32.4 setuptools 78.1.1 six 1.17.0 sortedcontainers 2.4.0 sympy 1.14.0 tomli 2.2.1 typing_extensions 4.14.0 tzdata 2025.2 urllib3 2.5.0 uv 0.7.21 wheel 0.45.1 zipp 3.23.0 (numpy_pandas) [gabeferns@devvm2497.eag0 ~/pt-envs/at (exclamaforte/just-gemm-model)]$ pip install numpy==1.22.4 Collecting numpy==1.22.4 Using cached numpy-1.22.4-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (2.0 kB) Using cached numpy-1.22.4-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (16.8 MB) Installing collected packages: numpy Successfully installed numpy-1.22.4 (numpy_pandas) [gabeferns@devvm2497.eag0 ~/pt-envs/at (exclamaforte/just-gemm-model)]$ pip install pandas==2.0.3 Collecting pandas==2.0.3 Using cached pandas-2.0.3-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (18 kB) Requirement already satisfied: python-dateutil>=2.8.2 in /home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages (from pandas==2.0.3) (2.9.0.post0) Requirement already satisfied: pytz>=2020.1 in /home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages (from pandas==2.0.3) (2025.2) Requirement already satisfied: tzdata>=2022.1 in /home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages (from pandas==2.0.3) (2025.2) Requirement already satisfied: numpy>=1.20.3 in /home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages (from pandas==2.0.3) (1.22.4) Requirement already satisfied: six>=1.5 in /home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages (from python-dateutil>=2.8.2->pandas==2.0.3) (1.17.0) Using cached pandas-2.0.3-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (12.4 MB) Installing collected packages: pandas Successfully installed pandas-2.0.3 (numpy_pandas) [gabeferns@devvm2497.eag0 ~/pt-envs/at (exclamaforte/just-gemm-model)]$ pip install --pre numpy==2.0.2 Collecting numpy==2.0.2 Using cached numpy-2.0.2-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (60 kB) Using cached numpy-2.0.2-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (19.5 MB) Installing collected packages: numpy Attempting uninstall: numpy Found existing installation: numpy 1.22.4 Uninstalling numpy-1.22.4: Successfully uninstalled numpy-1.22.4 Successfully installed numpy-2.0.2 (numpy_pandas) [gabeferns@devvm2497.eag0 ~/pt-envs/at (exclamaforte/just-gemm-model)]$ python Python 3.9.23 (main, Jun 5 2025, 13:40:20) [GCC 11.2.0] :: Anaconda, Inc. on linux Type "help", "copyright", "credits" or "license" for more information. >>> import pandas Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages/pandas/__init__.py", line 22, in <module> from pandas.compat import is_numpy_dev as _is_numpy_dev # pyright: ignore # noqa:F401 File "/home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages/pandas/compat/__init__.py", line 25, in <module> from pandas.compat.numpy import ( File "/home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages/pandas/compat/numpy/__init__.py", line 4, in <module> from pandas.util.version import Version File "/home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages/pandas/util/__init__.py", line 2, in <module> from pandas.util._decorators import ( # noqa:F401 File "/home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages/pandas/util/_decorators.py", line 14, in <module> from pandas._libs.properties import cache_readonly File "/home/gabeferns/.conda/envs/numpy_pandas/lib/python3.9/site-packages/pandas/_libs/__init__.py", line 13, in <module> from pandas._libs.interval import Interval File "pandas/_libs/interval.pyx", line 1, in init pandas._libs.interval ValueError: numpy.dtype size changed, may indicate binary incompatibility. Expected 96 from C header, got 88 from PyObject ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158584 Approved by: https://github.com/huydhn	2025-07-18 00:14:16 +00:00
Nikita Shulga	33c9b414aa	[CI][MPS] Enable test_indexing on MPS (#158582 ) - Skip `test_index_put_accumulate_large_tensor_mps` as it crashes with ``` /com.apple.xbs/Sources/MetalPerformanceShaders/MPSCore/Types/MPSNDArray.mm:829: failed assertion `[MPSNDArray initWithDevice:descriptor:isTextureBacked:] Error: NDArray dimension length > INT_MAX' ``` while running `torch.ones([2**31+5], dtype=torch.int8, device='mps')` - Adjust types for `test_index_put_src_datatype` as index_put on MPS is not implemented for complex (yet) - Adjust `test_index` to avoid using DoubleTensors for MPS Pull Request resolved: https://github.com/pytorch/pytorch/pull/158582 Approved by: https://github.com/dcci, https://github.com/Skylion007, https://github.com/manuelcandales	2025-07-17 23:33:52 +00:00
Lucas Kabela	b0e325c2c8	[Dynamo][Better Engineering] Add type coverage to decorators (#158509 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to an important file in dynamo, `decorators.py` NOTE: Untyped fns are because there is a conflict with `__init__.py` in compiler so we can't type these at this time Running ``` mypy torch/_dynamo/decorators.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 209 \| 908 \| 23.02% \| 9 \| 39 \| 23.08% \| \| This PR \| 870 \| 943 \| 100.00% \| 36 \| 39 \| 100.00% \| \| Delta \| +661 \| +35 \| +76.98% \| +27 \| 0 \| +76.92% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158509 Approved by: https://github.com/williamwen42	2025-07-17 23:31:26 +00:00
Yiming Zhou	f63988ae00	[BE]Clean up old APIs in AOTI c shim (#158400 ) Summary: The shims for aten ops are now generated by torchgen. But there are some still old APIs in `aoti_torch/c/shim.h` This diff moves the old to-be-deprecated APIs for aten ops to a separate header file `shim_deprecated.h` The to-be-deprecated APIs are determined by comparing APIs in `shim.h` and ops in `fallback_ops.py` Test Plan: CI Rollback Plan: Differential Revision: D78378373 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158400 Approved by: https://github.com/jingsh, https://github.com/desertfire	2025-07-17 23:24:50 +00:00
Jeff Daily	2df2e3bb51	[ROCm][CI] Last known good HIP patch (#158596 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/158596 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-17 22:52:16 +00:00
Mwiza Kunda	0ecfb93a0b	Avoid globally modifying torch.testing._internal.common_methods_invocations.wrapper_set_seed (#158548 ) Test modules that depend on the original definition of `wrapper_set_seed` will inadvertently be affected if they import from test_torchinductor_opinfo.py. Additionally, using pytest `test_torchinductor_opinfo.py test_other_module.py` when run in the same process may affect the test behaviour of `test_other_module.py` if the tests depend on `wrapper_set_seed`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158548 Approved by: https://github.com/janeyx99	2025-07-17 22:31:59 +00:00
Paul Ganssle	74f4cf4bd5	Add missing <vector> in c10/util/WaitCounter.h (#158354 ) It seems that `#include <vector>` is being pulled in indirectly, but it is being used directly, so it is best to explicitly include it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158354 Approved by: https://github.com/janeyx99	2025-07-17 22:23:05 +00:00
cyy	1b91954b9f	Suppress volatile type error (#158435 ) Fixes ``` /var/lib/jenkins/workspace/torch/csrc/dynamo/guards.cpp:5320:10: error: compound assignment to object of volatile-qualified type 'volatile char' is deprecated [-Werror,-Wdeprecated-volatile] ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158435 Approved by: https://github.com/janeyx99	2025-07-17 22:21:04 +00:00
Mikayla Gawarecki	41b2c4d119	Reduce random reads for offset metadata when calling torch.load under FakeTensorMode (#157931 ) We already test the `_get_offset` functionality with that TORCH_SERIALIZATION_DEBUG flag that is set in CI, so I didn't add more testing specifically for FakeTensor Pull Request resolved: https://github.com/pytorch/pytorch/pull/157931 Approved by: https://github.com/albanD	2025-07-17 22:17:52 +00:00
Animesh Jain	af6624023e	[dynamo] Skip training flag check id already guarding on nn modules (#158492 ) This might help some legacy models that still have inline_inbuilt_nn_modules False for some reason. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158492 Approved by: https://github.com/StrongerXi	2025-07-17 21:42:19 +00:00
Catherine Lee	a00442421a	[CI][TD] Enable TD on all test configs (#158163 ) I think the main one that was missing is dynamo_wrapped There's also slow and inductor, but the filter later for workflows stops TD from running on those anyways dynamo_wrapped is the second longest jobs for pull right now <img width="1265" height="311" alt="image" src="https://github.com/user-attachments/assets/d4ca034c-a8f0-4b31-a80f-0f4f21fce32a" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158163 Approved by: https://github.com/huydhn, https://github.com/ZainRizvi	2025-07-17 21:05:25 +00:00
PyTorch MergeBot	ced5cf042d	Revert "Cleanup old caffe2 scripts (#158475 )" This reverts commit 94d7f0c1ef9a4cb4db0eb5d6b1ffc55941cbeab1. Reverted https://github.com/pytorch/pytorch/pull/158475 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/158475#issuecomment-3085447409))	2025-07-17 20:58:34 +00:00
Kurt Mohler	1b88da1cac	[MPS] Improve performance of max_pool3d (#157875 ) To check how the changes from this PR affect performance, I wrote a script here: `55ef32a127/max_pool_mps/perf.py`. Before this PR, I get this: ``` =================== max_pool3d =================== 0: 0.013105 ms, max_pool3d, (3, 2, 2, 2), {'kernel_size': 2} 1: 0.038003 ms, max_pool3d, (3, 10, 10, 10), {'kernel_size': 5} 2: 0.212963 ms, max_pool3d, (3, 100, 100, 100), {'kernel_size': 5} 3: 1.224645 ms, max_pool3d, (3, 200, 200, 200), {'kernel_size': 5} 4: 7.317867 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 4, 'padding': 1} 5: 34.679233 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 50, 'padding': 20} 6: 34.626383 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 50, 'padding': 20, 'dilation': 1} 7: 44.835892 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 50, 'padding': 20, 'dilation': 1, 'stride': 40} 8: 0.083579 ms, max_pool3d, (10, 10, 10, 10, 10), {'kernel_size': 2} 9: 0.936575 ms, max_pool3d, (10, 10, 30, 30, 30), {'kernel_size': 2} 10: 5.329883 ms, max_pool3d, (10, 10, 50, 50, 50), {'kernel_size': 2} 11: 11.713617 ms, max_pool3d, (10, 10, 70, 70, 70), {'kernel_size': 2} 12: 25.450454 ms, max_pool3d, (10, 10, 90, 90, 90), {'kernel_size': 2} 13: 0.058375 ms, max_pool3d, (10, 10, 10, 10, 10), {'kernel_size': 2, 'dilation': 2} 14: 3.757558 ms, max_pool3d, (10, 10, 50, 50, 50), {'kernel_size': 2, 'dilation': 2} 15: 33.451588 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 2, 'dilation': 2} ``` After this PR, I get this: ``` =================== max_pool3d =================== 0: 0.007202 ms, max_pool3d, (3, 2, 2, 2), {'kernel_size': 2} 1: 0.018596 ms, max_pool3d, (3, 10, 10, 10), {'kernel_size': 5} 2: 0.130717 ms, max_pool3d, (3, 100, 100, 100), {'kernel_size': 5} 3: 0.966795 ms, max_pool3d, (3, 200, 200, 200), {'kernel_size': 5} 4: 4.095804 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 4, 'padding': 1} 5: 12.833446 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 50, 'padding': 20} 6: 12.859346 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 50, 'padding': 20, 'dilation': 1} 7: 14.080529 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 50, 'padding': 20, 'dilation': 1, 'stride': 40} 8: 0.029283 ms, max_pool3d, (10, 10, 10, 10, 10), {'kernel_size': 2} 9: 0.175700 ms, max_pool3d, (10, 10, 30, 30, 30), {'kernel_size': 2} 10: 0.742750 ms, max_pool3d, (10, 10, 50, 50, 50), {'kernel_size': 2} 11: 1.939596 ms, max_pool3d, (10, 10, 70, 70, 70), {'kernel_size': 2} 12: 4.074821 ms, max_pool3d, (10, 10, 90, 90, 90), {'kernel_size': 2} 13: 0.028425 ms, max_pool3d, (10, 10, 10, 10, 10), {'kernel_size': 2, 'dilation': 2} 14: 0.384375 ms, max_pool3d, (10, 10, 50, 50, 50), {'kernel_size': 2, 'dilation': 2} 15: 2.623346 ms, max_pool3d, (10, 10, 100, 100, 100), {'kernel_size': 2, 'dilation': 2} ``` Every case is improved. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157875 Approved by: https://github.com/malfet	2025-07-17 20:34:12 +00:00
angelayi	66c9bc5062	[export] Add runnable code to export docs (#158506 ) Preview: https://docs-preview.pytorch.org/pytorch/pytorch/158506/export.html Yay I can add runnable code to export docs now Also moved export API reference to a different file. With these changes, we can start to consolidate the [export tutorial](https://docs.pytorch.org/tutorials/intermediate/torch_export_tutorial.html) with the docs on pytorch docs. We just need to move the section on DDE and 0/1 specialization, and then I think we can delete the export tutorial. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158506 Approved by: https://github.com/pianpwk, https://github.com/svekars	2025-07-17 20:15:22 +00:00
Simon Fan	80ac73c057	[ca] reset between tests (#158418 ) CA reset is much faster than dynamo reset, so it's probably okay to run it every time. I'm not sure if this will fix the flaky autograd tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158418 Approved by: https://github.com/jansel	2025-07-17 20:14:29 +00:00
IvanKobzarev	eeb0783fe6	[simple_fsdp][inductor_collectives] rewrite reorder_collectives, sink_waits_iterative (#158062 ) Differential Revision: [D78159013](https://our.internmc.facebook.com/intern/diff/D78159013) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158062 Approved by: https://github.com/wconstab	2025-07-17 20:04:42 +00:00
Edward Yang	ef256ad17b	Make Inductor imports TYPE_CHECKING only (#158524 ) Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158524 Approved by: https://github.com/cyyever, https://github.com/albanD	2025-07-17 19:55:19 +00:00
Wouter Devriendt	fd51bcdd21	check if USE_ROCM is defined (#158571 ) Summary: check if USE_ROCM is defined D78424375 broke some builds: see T231304402 Test Plan: rerunning failed builds Rollback Plan: Reviewed By: Camyll Differential Revision: D78493019 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158571 Approved by: https://github.com/huydhn, https://github.com/malfet	2025-07-17 19:48:26 +00:00
Jack Taylor	7ebbf2cae7	Revert "[PT2][fusion] ban fusions with large accumulated reads (#157563 ) (#158550 ) This reverts commit 8554c8007ddaa8029e7e01bb1af12f358bf597c2 #157563 due to causing a few breakages on ROCm Reverted expected_results.csv to 26807dcf277feb2d99ab88d7b6da526488baea93 > @xuanzhang816 Sorry, but I have to revert this PR yet again because it clearly reintroduced failures on ROCm after the remerge: `f4d8bc46c7/2` and the failures are still showing up on tip-of-tree on HUD Context https://github.com/pytorch/pytorch/pull/157563#issuecomment-3083350857 Needs to be relanded in non bc-breaking way, or sanity checked for correctness. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158550 Approved by: https://github.com/jithunnair-amd, https://github.com/jeffdaily	2025-07-17 19:47:41 +00:00
Xu Han	8dcebaa7b0	[AOTI] add WIN32 implement for create_temp_dir (#158570 ) add Windows implement for `create_temp_dir`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158570 Approved by: https://github.com/angelayi	2025-07-17 19:22:59 +00:00
Divyansh Khanna	7e34f9c292	Add torch._C._log_api_usage_once to datapipes (mapper) (#155489 ) This is to get a better understanding of how datapipes is used right now. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155489 Approved by: https://github.com/ramanishsingh	2025-07-17 19:01:49 +00:00
albanD	25f4d7e482	Use new type statement to fix public API of types (#158487 ) Since type statement breaks older python version, trying to find equivalent behavior without the type mechanics. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158487 Approved by: https://github.com/andrewor14	2025-07-17 18:46:44 +00:00
Kevin Fu	ad223a6c5f	Add FP8 Types (#158430 ) Summary: Add FP8 Types Test Plan: sandcastle Rollback Plan: Differential Revision: D78395110 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158430 Approved by: https://github.com/henryoier	2025-07-17 18:09:56 +00:00
Eli Uriegas	f92a2035e4	ci: Update lint workflow to only run on changed files for PRs (#158518 ) This modifies the lint workflow to use the new get-changed-files workflow to optimize lint execution by only running on files that have actually changed in pull requests. This more closely mirrors the type of behavior that users expect when running lint locally on their PRs. This also leaves the default behavior as a fallback for when you're not running on a pull request. Since lint runs on the pull_request event I'm not really worried about any type of ciflow shenanigans in this. This also splits mypy into its own job since mypy needs to run on all-files all the time. Signed-off-by: Eli Uriegas <eliuriegas@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158518 Approved by: https://github.com/huydhn ghstack dependencies: #158517	2025-07-17 18:00:44 +00:00
Sam Larsen	bff69f25c2	[BE][testing] fix test/dynamo/test_repros:test_longtensor_list (#158458 ) Summary: This test is failing internally because the number of underlying calls to the rng differ by virtue of various library initializations that get sucked in with an internal build. Test Plan: `buck test '@fbcode//mode/opt' fbcode//caffe2/test/dynamo:test_dynamo -- --exact 'caffe2/test/dynamo:test_dynamo - test_repros.py::ReproTests::test_longtensor_list' --run-disabled` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158458 Approved by: https://github.com/jansel	2025-07-17 17:27:00 +00:00
Songhao Jia	6d31d38965	recovering node source from dict (#158373 ) (#158473 ) Summary: this diff recovers NodeSource object from its dict representation, which is crucial for NodeSource serde. Test Plan: ci Rollback Plan: Differential Revision: D78434648 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158473 Approved by: https://github.com/angelayi	2025-07-17 17:00:19 +00:00
PyTorch MergeBot	bfe5674e22	Revert "[cuDNN][SDPA] cuDNN SDPA refactor/cleanup, nested tensor backward, test priority bump for `sm90`, `sm100` (#149282 )" This reverts commit 0797b2b6a80cf70a7accc3d5413186e7693d4451. Reverted https://github.com/pytorch/pytorch/pull/149282 on behalf of https://github.com/wdvr due to reverting as discussed with @drisspg - @eqy please reach out to @drisspg for more info ([comment](https://github.com/pytorch/pytorch/pull/149282#issuecomment-3084759671))	2025-07-17 16:55:55 +00:00
albanD	94d7f0c1ef	Cleanup old caffe2 scripts (#158475 ) Testing on this one is grep based: if there were no reference to that script I can find, I deleted. We can easily add any of these back if needed! Pull Request resolved: https://github.com/pytorch/pytorch/pull/158475 Approved by: https://github.com/seemethere, https://github.com/huydhn, https://github.com/cyyever	2025-07-17 16:50:06 +00:00
PyTorch MergeBot	23550ab735	Revert "DDE-Free select with unbacked index. (#157605 )" This reverts commit 79d7c754ab8ae0e5c3a614521632d2cfbfa0fdba. Reverted https://github.com/pytorch/pytorch/pull/157605 on behalf of https://github.com/laithsakka due to fail pr time benchmarks ([comment](https://github.com/pytorch/pytorch/pull/157605#issuecomment-3084663020))	2025-07-17 16:20:02 +00:00
Xu Han	16b21fa8b2	[AOTI] skip ld and objcopy on Windows. (#158545 ) Skip `ld` and `objcopy` on Windows. They are not support on Windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158545 Approved by: https://github.com/desertfire	2025-07-17 15:43:24 +00:00
Oguz Ulgen	2ecf083b72	Add torch compile force disable caches alias (#158072 ) Bunch of people keep thinking current alias only disables inductor cache because it has the name inductor in it. lets globalize the name Pull Request resolved: https://github.com/pytorch/pytorch/pull/158072 Approved by: https://github.com/ezyang	2025-07-17 15:40:36 +00:00
PyTorch MergeBot	813c76b98d	Revert "Unify torch.tensor and torch.ops.aten.scalar_tensor behavior (#158537 )" This reverts commit 58c7cf9ede6311da5533dbcaf238a912176a6a85. Reverted https://github.com/pytorch/pytorch/pull/158537 on behalf of https://github.com/albanD due to This broke C++ tests ([comment](https://github.com/pytorch/pytorch/pull/158537#issuecomment-3084425920))	2025-07-17 15:06:43 +00:00
PyTorch MergeBot	288bf54a23	Revert "Move off of deprecated API in 2.9 (#158527 )" This reverts commit 9636e2cfd3e995ef977f670ad47e8e895296d992. Reverted https://github.com/pytorch/pytorch/pull/158527 on behalf of https://github.com/albanD due to breaks trunk ([comment](https://github.com/pytorch/pytorch/pull/158527#issuecomment-3084385585))	2025-07-17 14:55:28 +00:00
Xu Han	da4c7b4ced	[AOTI] align signature to model_base.h (#158554 ) Remove `const` keyword, align its signature to `model_base.h` `eeda1a75ac/torch/csrc/inductor/aoti_runtime/model_base.h (L51-L53)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158554 Approved by: https://github.com/desertfire	2025-07-17 14:44:32 +00:00
Xu Han	a04bd11895	[AOTI] Use format_consts_to_cpp on Windows. (#158543 ) `format_consts_to_asm` is not supported on Windows, force use `format_consts_to_cpp` on Windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158543 Approved by: https://github.com/desertfire	2025-07-17 14:40:34 +00:00
bobrenjc93	58c7cf9ede	Unify torch.tensor and torch.ops.aten.scalar_tensor behavior (#158537 ) Fixes #158376 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158537 Approved by: https://github.com/atalman	2025-07-17 13:39:25 +00:00
Ankita George	38c04415a9	[oss][hf][bug fix] Remove buggy consolidation logic (#158380 ) Summary: I tried to add some logic that could optimize for the non-row wise sharded case and do it more efficiently, but this has some bugs, so removing it for now and will find a better algorithm for the non-row wise sharded case to find the maximum number of bytes that we can write at a time. Test Plan: ensure tests pass Rollback Plan: Differential Revision: D78366701 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158380 Approved by: https://github.com/Saiteja64	2025-07-17 13:05:06 +00:00
David Berard	7892f5a007	[inductor][triton] Update HAS_WARP_SPEC to check triton.Config params. Update Triton Hash to top of release/3.4.x stack (#158459 ) Update triton commit hash to `11ec6354315768a85da41032535e3b7b99c5f706`, which is the new release/3.4.x branch in triton-lang/triton. Also, update HAS_WARP_SPEC handling: In triton 3.4, warp spec will have a different interface: num_consumer_groups will be determined automatically by the compiler. This breaks the current Inductor integration, so for now, update HAS_WARP_SPEC to check whether triton.Config takes num_consumer_groups and num_buffers_warp_spec as parameters. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158459 Approved by: https://github.com/atalman	2025-07-17 12:50:46 +00:00
Xuehai Pan	d5af0eca8d	[BE][3/5] fix typos in aten/ (aten/src/ATen/native/) (#157552 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157552 Approved by: https://github.com/albanD ghstack dependencies: #156605, #157637, #157550, #157551	2025-07-17 12:08:34 +00:00
Xuehai Pan	f57ef62ebc	[BE][2/5] fix typos in aten/ (aten/src/ATen/native/) (#157551 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157551 Approved by: https://github.com/albanD ghstack dependencies: #156605, #157637, #157550	2025-07-17 12:08:33 +00:00
Xuehai Pan	4c8b408d16	[BE][1/5] fix typos in aten/ (#157550 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157550 Approved by: https://github.com/albanD ghstack dependencies: #156605, #157637	2025-07-17 12:08:33 +00:00
Xuehai Pan	c8d43cbc6e	[BE][3/6] fix typos in test/ (#157637 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157637 Approved by: https://github.com/yewentao256, https://github.com/albanD ghstack dependencies: #156605	2025-07-17 12:08:33 +00:00
Xuehai Pan	3f8e2e91ad	[BE][15/16] fix typos in torch/ (torch/distributed/tensor/) (#156605 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156605 Approved by: https://github.com/wanchaol, https://github.com/albanD	2025-07-17 12:08:33 +00:00
Luca Wehrstedt	eeda1a75ac	Forward-fix unused variables warning/error (#158549 ) Introduced in https://github.com/pytorch/pytorch/pull/158037, didn't seem to trigger on PR, but trunk CI is failing in some `linux-jammy-cpu-py3.12-gcc11-inductor-*` jobs where this warning is turned into an error. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158549 Approved by: https://github.com/danthe3rd	2025-07-17 09:44:19 +00:00
Jiang, Yanbing	f4d8bc46c7	Enable TF32 as fp32 internal precision for matmul/linear/conv (#157520 ) ### Description This PR is to enable TF32 as fp32 internal precision for matmul/linear/conv in `mkldnn backend`. Since we have refined fp32 precision API in https://github.com/pytorch/pytorch/pull/125888, we can easily extend the API to support TF32 for `mkldnn backend`. ``` torch.backends.mkldnn.matmul.fp32_precision = 'tf32' torch.backends.mkldnn.conv.fp32_precision = "tf32" ``` Related kernel update and UTs update are done. And the wrapper `bf32_on_and _off` is updated to `reduced_f32_on_and_off`, and it can run tests 3 times, one is reduced_f32 OFF, the other two are reduced_f32 ON (including `bf32 ON` and `tf32 ON`). Pull Request resolved: https://github.com/pytorch/pytorch/pull/157520 Approved by: https://github.com/mingfeima, https://github.com/jansel	2025-07-17 08:57:34 +00:00
Luca Wehrstedt	39ac189808	Support DeepSeek-style blockwise scaling scaled-mm for fp8 on Hopper+ (#158037 ) cuBLAS added support for them in CUDA 12.9. It's rather easy to call into them, the hardest thing is allowing the lhs and rhs operands to have different scaling types, as that changes the whole callstack. The scaling format is still detected from the sizes of the scale tensors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158037 Approved by: https://github.com/eqy, https://github.com/drisspg	2025-07-17 08:26:27 +00:00
Sherlock Huang	d76323d417	[NativeRT] Remove normalizeDevice (#158489 ) Summary: In pytorch, tensor.to("cuda") behaves differently from tensor.to("cuda:0). tensor.to("cuda") will read from thread local DeviceGuard, aka cuda::current_device(), to infer the device index. TBEPermute is relying on this behavior to route output tensor to a device specified by current thread. For this reason, we remove the normalizeDevice(), and disallow index-less cuda device in Placement. Device-to-device mapping must be done between concrete device! Test Plan: CI Rollback Plan: Differential Revision: D78443109 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158489 Approved by: https://github.com/henryoier	2025-07-17 06:48:25 +00:00
Kevin Fu	04349f9ee5	[PT2]: Skip AOTI Weight Loading during Init (#158416 ) Summary: AOTI already has weights embedded in .so file. So for the initial load, no need to load the weights again. This allows lowered modules can have different set of weights on different hardwares. Test Plan: ``` MODEL_TYPE=ads_mtml_offsite_cvr_oba_optout_dedicated_model MODEL_ENTITY_ID=895279202 SNAPSHOT_ID=0 MODULE=merge buck2 run mode/dev-nosan -c fbcode.nvcc_arch=a100,h100 -c fbcode.enable_gpu_sections=true fbcode//caffe2/torch/fb/model_transform/fx2trt/packaging:load_net_predictor -- --loadMode=Benchmark --inputNetFile=/data/users/$USER/models/${MODEL_ENTITY_ID}/${SNAPSHOT_ID}/${MODEL_ENTITY_ID}_${SNAPSHOT_ID}.predictor.disagg.gpu.${MODULE} --moduleName ${MODULE} --predictor-hardware-type 1 --submodToDevice "" --benchmarkDontRebatchSamples=true --benchmarkNumIterations 1000 ``` Rollback Plan: Differential Revision: D78383881 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158416 Approved by: https://github.com/henryoier, https://github.com/SherlockNoMad	2025-07-17 06:47:47 +00:00
Jane Xu	09db3a22e8	[BE] Get rid of final mentions of BUILD_SPLIT_CUDA (#158453 ) BUILD_SPLIT_CUDA logic has been removed for a while Differential Revision: [D78418191](https://our.internmc.facebook.com/intern/diff/D78418191/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158453 Approved by: https://github.com/albanD ghstack dependencies: #158358, #158365	2025-07-17 06:47:10 +00:00
Andrey Talman	a38f433be2	[Docker builds] Move from Miniconda to Miniforge (#158370 ) This is related to: https://www.anaconda.com/legal/terms/terms-of-service Trying to fix outage with docker builds. https://github.com/pytorch/pytorch/actions/runs/16298993712/job/46033590799 Rocm and XPU builds since they use Miniforge are not affected ``` #22 ERROR: process "/bin/sh -c bash ./install_conda.sh && rm install_conda.sh install_magma_conda.sh common_utils.sh /opt/conda/requirements-ci.txt /opt/conda/requirements-docs.txt" did not complete successfully: exit code: 1 ------ > [base 14/42] RUN bash ./install_conda.sh && rm install_conda.sh install_magma_conda.sh common_utils.sh /opt/conda/requirements-ci.txt /opt/conda/requirements-docs.txt: 11.93 CondaToSNonInteractiveError: Terms of Service have not been accepted for the following channels. Please accept or remove them before proceeding: 11.93 • https://repo.anaconda.com/pkgs/main 11.93 • https://repo.anaconda.com/pkgs/r 11.93 11.93 To accept a channel's Terms of Service, run the following and replace `CHANNEL` with the channel name/URL: 11.93 ‣ conda tos accept --override-channels --channel CHANNEL ``` Hence solution is: 1. using `` conda tos accept --override-channels --channel defaults`` 2. use Miniforge instead of Miniconda. Using solution 2. Solution Tried that don't work: 1. Using ``CONDA_ALWAYS_YES = true `` 4. Using older version of miniconda ``` [Miniconda3-py310_25.5.1-0-Linux-x86_64.sh](https://repo.anaconda.com/miniconda/Miniconda3-py310_25.5.1-0-Linux-x86_64.sh) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158370 Approved by: https://github.com/seemethere Co-authored-by: Eli Uriegas <1700823+seemethere@users.noreply.github.com>	2025-07-17 06:33:08 +00:00
PyTorch MergeBot	9f37cce693	Revert "[Docker builds] Move from Miniconda to Miniforge (#158370 )" This reverts commit 0a99b026d6bd0f67dc2c0a20fe3228ddc4144854. Reverted https://github.com/pytorch/pytorch/pull/158370 on behalf of https://github.com/laithsakka due to this fail pr time benchmarks ([comment](https://github.com/pytorch/pytorch/pull/158370#issuecomment-3082744071))	2025-07-17 06:28:49 +00:00
drisspg	9636e2cfd3	Move off of deprecated API in 2.9 (#158527 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158527 Approved by: https://github.com/danielvegamyhre	2025-07-17 06:18:13 +00:00
PaliC	d9426a81d2	[BE] Modify PyObjectSlot the assume only a single interpreter is in use (#158407 ) This PR makes some less risky changes to PyObjectSlot as there is a lot of stuff we do not need since there is only one interpreter. Specifically `check_interpreter` and `has_pyobj_nonhermetic` are removed Pull Request resolved: https://github.com/pytorch/pytorch/pull/158407 Approved by: https://github.com/albanD ghstack dependencies: #158288, #158290, #158291	2025-07-17 05:56:26 +00:00
PaliC	0b9fb91f17	[BE] Remove __reduce_deploy__ (#158291 ) This PR removes the integration point torch.fx had with torch::deploy (and another minor change). Note: This PR has some broken mypy errors, but I believe those should have been in the code base beforehand, and should be fixed in a separate PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/158291 Approved by: https://github.com/albanD ghstack dependencies: #158288, #158290	2025-07-17 05:56:26 +00:00
PaliC	a6de309ca1	[BE] Remove torch deploy \| remove torch deploy specific files (#158290 ) This PR removes specific files found in pytorch which are only used for torch::deploy. This is mostly testing code and a debugger. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158290 Approved by: https://github.com/albanD ghstack dependencies: #158288	2025-07-17 05:56:18 +00:00
PaliC	1a4268b811	[BE] remove torch deploy - conditionals (#158288 ) This PR is part of the work to deprecate torch::deploy in OSS. Effectively it does 3 things to get started. 1. Remove test_deploy_interaction as we no longer need to worry about this 2. Remove all torch._running_with_deploy checks and use the False path always (surfaced 1) 3. Remove `USE_DEPLOY` and switch to the default path always Note: MyPy does fail on a bunch of things here as a bunch of older files are touched. It may be better to fix these things on a separate PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/158288 Approved by: https://github.com/albanD	2025-07-17 05:56:07 +00:00
Laith Sakka	79d7c754ab	DDE-Free select with unbacked index. (#157605 ) When select has data dependent input, we cant tell if the actual index shall be index+size or index. to avoid throwing dde, we allocate a new unbacked symbol to represent the storage offset of the output view and we compute its value dynamically at runtime when inductor is lowered. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157605 Approved by: https://github.com/ColinPeppler	2025-07-17 05:08:11 +00:00
FFFrog	415dfabe9b	[Easy] Fix the format (#158450 ) When I modify the code located in test/cpp_extensions/open_registration_extension/torch_openreg/torch_openreg, some unrelated format error occurred. ```Python Lint for torch/_inductor/fx_passes/fuse_attention.py: Error (CODESPELL) spelling error Failed due to ValueError: /pytorch/pytorch/torch/_inductor/fx_passes/fuse_attention.py:587: differnt ==> different Please either fix the error or add the word(s) to the dictionary file. HINT: all-lowercase words in the dictionary can cover all case variations. Lint for torch/fx/traceback.py: Error (MYPY) [assignment] Incompatible types in assignment (expression has type "str", variable has type "None") 101 \| 102 \| def _get_action_string(self): 103 \| if self._action_string is None: 104 \| self._action_string = "+".join([a.name.lower() for a in self.action]) 105 \| return self._action_string 106 \| 107 \| def print_readable(self, indent=0): Error (MYPY) [assignment] Incompatible types in assignment (expression has type "dict[str, Any]", variable has type "None") 121 \| if self._dict is None: 122 \| # Convert the object to a dictionary 123 \| action_string = self._get_action_string() 124 \| self._dict = { 125 \| "name": self.name, 126 \| "target": self.target, 127 \| "graph_id": self.graph_id, Error (MYPY) [return-value] Incompatible return value type (got "None", expected "dict[Any, Any]") 130 \| "from_node": [node.to_dict() for node in self.from_node], 131 \| } 132 \| 133 \| return self._dict 134 \| 135 \| def __eq__(self, other: object): 136 \| if not isinstance(other, NodeSource): ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158450 Approved by: https://github.com/Skylion007	2025-07-17 04:56:10 +00:00
Natalia Gimelshein	8eaa9f2701	Fix mask construction when dispatching index_put to masked_fill (#158472 ) Fixes #158413 Previously trailing Nones in the index were incorrectly handled as implicit broadcasting dims in the mask, whereas they should just be ignored. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158472 Approved by: https://github.com/ezyang	2025-07-17 04:21:43 +00:00
PyTorch UpdateBot	ebf83b8b77	[audio hash update] update the pinned audio hash (#158402 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158402 Approved by: https://github.com/pytorchbot	2025-07-17 04:19:06 +00:00
Raymond Li	24b49b9881	[Fix] Rework CUDA error explanation framework to be less destructive … (#158484 ) …in fbsource Fix-forward for #158395 Added `std::string c10::cuda::get_cuda_error_help(const char* error_string)` to provide a framework for appending clarifying messages to CUDA errors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158484 Approved by: https://github.com/aorenste	2025-07-17 03:36:47 +00:00
Will Constable	1839e8d04b	[DTensor] Assert DTensorSpec has valid placements (#158133 ) This helped identify buggy sharding rules during debugging, why not check it in. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158133 Approved by: https://github.com/XilunWu, https://github.com/zpcore ghstack dependencies: #158132	2025-07-17 02:32:26 +00:00
Yu, Guangye	2ad5c25cfc	Add unified memory APIs for torch.accelerator (#152932 ) # Motivation The following API will be put under torch.accelerator - empty_cache - max_memory_allocated - max_memory_reserved - memory_allocated - memory_reserved - memory_stats - reset_accumulated_memory_stats - reset_peak_memory_stats Pull Request resolved: https://github.com/pytorch/pytorch/pull/152932 Approved by: https://github.com/albanD ghstack dependencies: #138222	2025-07-17 01:56:01 +00:00
Yu, Guangye	1179e33323	Add DeviceAllocator as the base device allocator (#138222 ) # Motivation In line with [RFC] [A device-agnostic Python device memory related API design for stream-based accelerators](https://github.com/pytorch/pytorch/issues/134978), some memory-related APIs are widely used in popular repositories, such as HuggingFace [so many if-else conditional code](https://github.com/search?q=repo%3Ahuggingface%2Faccelerate%20torch.cuda.empty_cache&type=code). We would like to introduce a generic API set under torch.accelerator namespace to generalize these user cases. <div align="center"> <table> <tr> <td> Device-specific memory APIs torch.xxx.foo</td> <td> Device-agnostic memory APIs torch.accelerator.foo</td> </tr> <tr> <td> ```python torch.xxx.empty_cache ``` </td> <td> ```python torch.accelerator.empty_cache ``` </td> </tr> <tr> <td> ```python torch.xxx.reset_peak_memory_stats ``` </td> <td> ```python torch.accelerator.reset_peak_memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.reset_accumulated_memory_stats ``` </td> <td> ```python torch.accelerator.reset_accumulated_memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_stats ``` </td> <td> ```python torch.accelerator.memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_allocated ``` </td> <td> ```python torch.accelerator.memory_allocated ``` </td> </tr> <tr> <td> ```python torch.xxx.max_memory_allocated ``` </td> <td> ```python torch.accelerator.max_memory_allocated ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_reserved ``` </td> <td> ```python torch.accelerator.memory_reserved ``` </td> </tr> <tr> <td> ```python torch.xxx.max_memory_reserved ``` </td> <td> ```python torch.accelerator.max_memory_reserved ``` </td> </tr> </table> </div> # Solution This design follows a similar pattern to `HostAllocator`. We're introducing a base class `DeviceAllocator`, from which `CUDAAllocator` and `XPUAllocator` will inherit. This allows us to provide a unified call path like: `torch.accelerator.empty_cache()` -> `GetDeviceAllocator(allocator)->empty_cache()`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/138222 Approved by: https://github.com/albanD, https://github.com/Camyll	2025-07-17 01:56:01 +00:00
Arsh Zahed	f6d138807f	Always disable ShardingPropagation cache if compiling (#156868 ) Fixes #151106 Addresses issue (2) in #152963 for the DTensor sharding propagation cache being brittle under compile. The existing `_are_we_tracing` from `distributed._functional_collectives`, which mostly determines if currently tracing based on Fake Tensor dispatch mode, is reused here. Test Plan: There are already tests for DTensor + Compile with dynamic shape ([test_dtensor_dynamic](https://github.com/pytorch/pytorch/blob/main/test/distributed/tensor/test_dtensor_compile.py#L260), [test_dynamo_dtensor_from_local_dynamic_shapes](https://github.com/pytorch/pytorch/blob/main/test/distributed/tensor/test_dtensor_compile.py#L402)) that cover the change. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156868 Approved by: https://github.com/xmfan	2025-07-17 01:33:53 +00:00
Jiapeng Li	c09eba877f	[Device] Add support for PrivateUse1 device type in parse_type function (#157609 ) This pull request refactors the `parse_type` function in `c10/core/Device.cpp` to improve the handling of the `PrivateUse1` device type. The main change involves reordering the logic to check for the `PrivateUse1` device type earlier in the function for better clarity and efficiency. This help to migrate existed backend to PrivateUse1 smoothly. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157609 Approved by: https://github.com/jgong5, https://github.com/albanD	2025-07-17 01:27:44 +00:00
Animesh Jain	2179afd714	[easy][guards] Add developer comment for posterity (#158471 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158471 Approved by: https://github.com/StrongerXi	2025-07-17 01:17:04 +00:00
Animesh Jain	d7e1b8b11d	[dynamo] Constant fold torch.autograd._profiler_enabled (#158482 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158482 Approved by: https://github.com/williamwen42, https://github.com/StrongerXi	2025-07-17 01:07:42 +00:00
Xu Han	b6454a9058	[AOT_inductor] model_base.h add Windows include files. (#158477 ) model_base.h add Windows include files. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158477 Approved by: https://github.com/desertfire, https://github.com/jansel	2025-07-17 00:57:48 +00:00
Eli Uriegas	e9367a7a42	ci: Add reusable workflow to get changed files in PRs (#158517 ) Signed-off-by: Eli Uriegas <eliuriegas@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158517 Approved by: https://github.com/huydhn	2025-07-17 00:57:43 +00:00
clr	e78f2ac92b	inductor: Fix crash in split_cat when tensors is a Node (#157155 ) If there is only one node passed to aten::cat, the argument is a single node, rather than a list of nodes with a valid length. Example stack ``` File "/dev/shm/uid-99/be3468a8-seed-nspid4026546656_cgpid14993614-ns-4026546628/torch/_inductor/pattern_matcher.py", line 1115, in apply self.handler(match, match.args, *match.kwargs) File "/dev/shm/uid-99/be3468a8-seed-nspid4026546656_cgpid14993614-ns-4026546628/torch/_inductor/fx_passes/split_cat.py", line 1786, in merge_split_cat_aten if len(cat_inputs) < threshold_to_cat: torch._inductor.exc.InductorError: TypeError: object of type 'Node' has no len() ``` This has failed about 7 internal jobs in the last week, running pytorch trunk code from 06/15 I've attached a test which reproduces this issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157155 Approved by: https://github.com/jansel	2025-07-17 00:57:38 +00:00
Shangdi Yu	82a1ee1135	Refactor Provenance Tracking (#158399 ) Summary: As inductor provenance tracking is getting more use cases, we want to separate the inductor provenance tracking guarding flag from the general `trace.enabled`, so we can enable provenance tracking without all the overhead of `trace.enabled` - change the guard flag from `trace.enabled` to `trace.provenance_tracking`. It is turned on by either `TORCH_COMPILE_DEBUG=1` or `INDUCTOR_PROVENANCE=1`. - Move the provenance tracking logic and variables out of DebugContext, because DebugContext is only enabled with `trace.enabled`. Since the variables are now global variables, added `reset_provenance_globals()` context manager to reset them for each `compile_fx()` call. - Move `set_kernel_post_grad_provenance_tracing` from `util.py` to `debug.py` so now all provenance related logic is in `debug.py`. In the future, if we want to enable it further, we can change the provenance tracking flag to be enabled when `TORCH_TRACE` is set. I think we should do that in a separate PR, so it's easier to revert if this flag change creates any problem. See more motivation in internal Diff Test Plan: ``` buck2 run mode/dev-nosan fbcode//caffe2/test:fx -- -r test_graph_transform_observer buck run mode/dev-nosan fbcode//caffe2/test:fx -- -r graph_provenance buck2 run mode/dev-nosan fbcode//caffe2/test/inductor:provenance_tracing ``` Differential Revision: D78287976 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158399 Approved by: https://github.com/angelayi	2025-07-17 00:23:00 +00:00
Laith Sakka	306dd19216	update expeced results (#158497 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158497 Approved by: https://github.com/xmfan	2025-07-17 00:02:52 +00:00
Howard Huang	1d58476162	[PP] Add eval() API to schedule (#157795 ) These change add an `eval()` API to PP schedules ## Context Currently, you can run "Forward only" for a schedule in two ways: 1. Use a custom schedule `_ScheduleForwardOnly` 2. Do not pass in `loss_fn` in schedule constructor, and no backward computations will be executed. However, this is still limiting because we may want to run forward through the pipeline / calculate the loss, but without backward, e.g. during validation. These changes allow for this. ```python if self.rank == 0: schedule.eval(x) elif self.rank == self.world_size - 1: losses = [] schedule.eval(target=target, losses=losses) else: schedule.eval() ``` TODO: - in later PRs, we will deprecate the `_ScheduleForwardOnly` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157795 Approved by: https://github.com/wconstab	2025-07-16 23:48:45 +00:00
Lucas Kabela	a4d753295e	[Dynamo][Better Engineering] Add enhanced typing support to `_dynamo/eval_frame.py` (#158276 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to the main entrypoint for dynamo, `eval_frame.py` Running ``` mypy torch/_dynamo/eval_frame.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 623 \| 2232 \| 27.91% \| 19 \| 68 \| 27.94% \| \| This PR \| 2285 \| 2285 \| 100.00% \| 68 \| 68 \| 100.00% \| \| Delta \| +1662 \| +63 \| +72.09% \| +49 \| 0 \| +72.06% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158276 Approved by: https://github.com/williamwen42 Co-authored-by: William Wen <williamwen@meta.com>	2025-07-16 23:31:10 +00:00
Frank Lin	a9f902add0	[CUDA] Use runtime driver API for cuStreamWriteValue32 (#158295 ) Reopen https://github.com/pytorch/pytorch/pull/156097 Fixes https://github.com/pytorch/pytorch/issues/154073 Reference: https://github.com/NVIDIA/Fuser/pull/4197 See PR https://github.com/pytorch/pytorch/pull/156097 and https://github.com/pytorch/pytorch/pull/154097 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158295 Approved by: https://github.com/Skylion007, https://github.com/ngimel, https://github.com/eqy, https://github.com/huydhn Co-authored-by: Wei Wang <weiwan@nvidia.com>	2025-07-16 23:14:36 +00:00
Mikayla Gawarecki	e311886e3d	Add transpose to torch/csrc/stable (#158160 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158160 Approved by: https://github.com/janeyx99	2025-07-16 22:50:57 +00:00
angelayi	3cb11877aa	[aoti][mps] Enable test_aot_inductor.py tests (#155598 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155598 Approved by: https://github.com/yushangdi	2025-07-16 22:26:57 +00:00
Lucas Kabela	5951fcd50a	[Dynamo][Better Engineering] Support typing in codegen.py (#158386 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to a critical tracing point for dynamo, primarily for `codegen.py` but also `config.py` Running ``` mypy torch/_dynamo/codegen.py torch/_dynamo/config.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 347 \| 1330 \| 26.09% \| 24 \| 50 \| 48.00% \| \| This PR \| 1334 \| 1334 \| 100.00% \| 50 \| 50 \| 100.00% \| \| Delta \| +987 \| +4 \| +73.91.% \| +26 \| 0 \| +52.00% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158386 Approved by: https://github.com/StrongerXi	2025-07-16 22:09:01 +00:00
Lucas Kabela	ada44e5ba7	[Dynamo][Better Engineering] Add typing to bytecode analysis and transform (#158293 ) As part of better engineering week, we would like to improve out type support to improve dev experience in dynamo This PR adds strict typing support to a critical tracing point for dynamo, `bytecode_transformation.py` and by extension, `bytecode_analysis.py` Running ``` mypy torch/_dynamo/bytecode_transformation.py torch/_dynamo/bytecode_analysis.py --linecount-report /tmp/coverage_log ``` \| -------- \| Lines Unannotated \| Lines Total \| % lines covered \| Funcs Unannotated \| Funcs Total \| % funcs covered \| \| -------- \| ------- \| -------- \| ------- \| ------- \| ------- \| ------- \| \| Main \| 1422 \| 1920 \| 74.06% \| 73 \| 93 \| 78.49% \| \| This PR \| 1968 \| 1968 \| 100.00% \| 93 \| 93 \| 100.00% \| \| Delta \| +546 \| +48 \| +25.94% \| 20 \| 0 \| +21.51% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/158293 Approved by: https://github.com/StrongerXi, https://github.com/Skylion007	2025-07-16 21:50:55 +00:00
Sam Larsen	9df0176408	[BE][testing] Disable test_static_cuda_launcher:test_floats internally (#158296 ) Summary: it seems the check for 'Offd' vs. 'Offf' doesn't work Pull Request resolved: https://github.com/pytorch/pytorch/pull/158296 Approved by: https://github.com/davidberard98	2025-07-16 21:27:40 +00:00
Xilun Wu	94c746bb43	[DTensor][BE] add document to ShardingPropagator.register_op_strategy (#158362 ) Summary Add document to `ShardingPropagator.register_op_strategy` on how to draft `strategy_func` and when to use `schema_info`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158362 Approved by: https://github.com/zpcore	2025-07-16 21:08:59 +00:00
Catherine Lee	473208cb18	[ez][lint] Add pr_time_benchmarks to merge conflictless csv linter (#158353 ) Discovered this when looking at a PR I was trying to revert and was surprised that the PR got rid of the spaces but didn't trigger the linter. Turns out the file was following the rule but wasn't actually being checked Pull Request resolved: https://github.com/pytorch/pytorch/pull/158353 Approved by: https://github.com/seemethere, https://github.com/Camyll	2025-07-16 20:31:07 +00:00
atalman	fb731fe371	Add warning about removed sm50 and sm60 arches (#158301 ) Related to https://github.com/pytorch/pytorch/issues/157517 Detect when users are executing torch build with cuda 12.8/12.9 and running on Maxwell or Pascal architectures. We would like to include reference to the issue: https://github.com/pytorch/pytorch/issues/157517 as well as ask people to install CUDA 12.6 builds if they are running on sm50 or sm60 architectures. Test: ``` >>> torch.cuda.get_arch_list() ['sm_70', 'sm_75', 'sm_80', 'sm_86', 'sm_90', 'sm_100', 'sm_120', 'compute_120'] >>> torch.cuda.init() /home/atalman/.conda/envs/py312/lib/python3.12/site-packages/torch/cuda/__init__.py:263: UserWarning: Found <GPU Name> which is of cuda capability 5.0. PyTorch no longer supports this GPU because it is too old. The minimum cuda capability supported by this library is 7.0. warnings.warn( /home/atalman/.conda/envs/py312/lib/python3.12/site-packages/torch/cuda/__init__.py:268: UserWarning: Support for Maxwell and Pascal architectures is removed for CUDA 12.8+ builds. Please see https://github.com/pytorch/pytorch/issues/157517 Please install CUDA 12.6 builds if you require Maxwell or Pascal support. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158301 Approved by: https://github.com/nWEIdia, https://github.com/albanD	2025-07-16 20:11:18 +00:00
Yiming Zhou	a9ee4250d5	[4/n] Remove references to TorchScript in PyTorch docs (#158317 ) Summary: jit.rst Test Plan: CI Rollback Plan: Differential Revision: D78309840 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158317 Approved by: https://github.com/svekars, https://github.com/zhxchen17	2025-07-16 20:01:34 +00:00
PyTorch MergeBot	14ecc03361	Revert "recovering node source from dict (#158373 )" This reverts commit 4d055982e38f59fdb2a4c9d8855e58548bc42c12. Reverted https://github.com/pytorch/pytorch/pull/158373 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/158373#issuecomment-3080093479))	2025-07-16 19:55:21 +00:00
angelayi	1cc62c2cb9	[export] Update docs (#157750 ) Preview: https://docs-preview.pytorch.org/pytorch/pytorch/157750/export.html Changes: * Rename draft_export.md -> export.draft_export.md for consistency. * Removed non-strict section in export, instead pointed to programming model doc. * Extended "Expressing Dynamism" section to include Dim hints, ShapeCollection, and AdditionalInputs. * Removed Specialization section in favor of programming model doc * Added pt2 archive doc * Cleaned up sidebar Pull Request resolved: https://github.com/pytorch/pytorch/pull/157750 Approved by: https://github.com/pianpwk	2025-07-16 19:53:12 +00:00
fduwjj	f58a680d09	[c10d]Prototype of remote_group_merge (#158287 ) Tentative implementation of merge_remote_group per the proposal here: [docs.google.com/document/d/13R-1t_yESTvmAjcCN-wQjQQadIEu0JNIdS65uZawZzY/edit?tab=t.0#heading=h.3ctbqqopzc89](https://docs.google.com/document/d/13R-1t_yESTvmAjcCN-wQjQQadIEu0JNIdS65uZawZzY/edit?tab=t.0#heading=h.3ctbqqopzc89) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158287 Approved by: https://github.com/d4l3k ghstack dependencies: #157716	2025-07-16 19:33:57 +00:00
PyTorch MergeBot	944a140e90	Revert "[cuda][cupy] Improve cupy device placement when device is provided (#158320 )" This reverts commit 59f9b25f3cfc635053843372ea29ff4bf754da3f. Reverted https://github.com/pytorch/pytorch/pull/158320 on behalf of https://github.com/wdvr due to reverting because most likely causing test/test_numba_integration.py::TestNumbaIntegration::test_from_cuda_array_interface_inferred_strides to fail ([comment](https://github.com/pytorch/pytorch/pull/158320#issuecomment-3079960616))	2025-07-16 19:15:33 +00:00
cyy	79ab84e9b8	Fix invalid formatting (#158436 ) It causes errors under C++20 ``` /Users/runner/work/pytorch/pytorch/pytorch/aten/src/ATen/native/mps/OperationUtils.mm:330:40: error: call to consteval function 'fmt::fstring<>::fstring<std::string, 0>' is not a constant expression ``` Indeed the printed value is treated as format string and it may contain special chars in some cases. While this is not true in our case, it can't be determined in compile time. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158436 Approved by: https://github.com/Skylion007	2025-07-16 18:47:09 +00:00
Jane Xu	2b0f9b1f61	Move c10/macros/Macros.h to headeronly (#158365 ) ^ Differential Revision: [D78361893](https://our.internmc.facebook.com/intern/diff/D78361893/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158365 Approved by: https://github.com/swolchok ghstack dependencies: #158358	2025-07-16 18:46:52 +00:00
Jane Xu	b40f48d191	Move the rest of c10/macros/Export.h (#158358 ) Differential Revision: [D78356975](https://our.internmc.facebook.com/intern/diff/D78356975/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158358 Approved by: https://github.com/swolchok	2025-07-16 18:46:52 +00:00
Songhao Jia	4d055982e3	recovering node source from dict (#158373 ) Summary: this diff recovers NodeSource object from its dict representation, which is crucial for NodeSource serde. Test Plan: ci Rollback Plan: Differential Revision: D78363882 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158373 Approved by: https://github.com/yushangdi	2025-07-16 18:46:09 +00:00
Manuel Candales	bc9091a524	Fix indexing with multi-dimensional boolean mask (#158369 ) Fixes #71673 This fixes a bug in PyTorch indexing, that shows up when mixing multi-dimensional boolean masks with other forms of indexing. Examples: ```python >>> import torch >>> x = torch.ones([2, 2, 3]) >>> m = torch.tensor(((True, False), (False, False))) # (2x2 boolean mask) >>> x[m].shape # this works fine (the boolean mask acts on the 2x2 subspace selecting one row) torch.Size([1, 3]) >>> x[m, 0] # this should produce a tensor of shape (1,) Traceback (most recent call last): File "<stdin>", line 1, in <module> IndexError: The shape of the mask [2, 2] at index 1 does not match the shape of the indexed tensor [2, 3] at index 1 >>> x[m, ::2] # this should produce a tensor of shape (1, 2) Traceback (most recent call last): File "<stdin>", line 1, in <module> IndexError: The shape of the mask [2, 2] at index 1 does not match the shape of the indexed tensor [2, 1, 3] at index 1 >>> x[m, None] # this should produce a tensor of shape (1, 1, 3) Traceback (most recent call last): File "<stdin>", line 1, in <module> IndexError: The shape of the mask [2, 2] at index 1 does not match the shape of the indexed tensor [2, 1, 2, 3] at index 1 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158369 Approved by: https://github.com/ngimel	2025-07-16 18:30:57 +00:00
Denghui Dong	a26bf38927	Don't need to handle PyTrace_EXCEPTION in pyProfileFn (#154392 ) According to the [document](https://python.readthedocs.io/fr/stable/c-api/init.html#c.PyTrace_EXCEPTION) and [comment](https://github.com/python/cpython/blob/3.9/Modules/_lsprof.c#L407), we don't need to handle PyTrace_EXCEPTION in pyProfileFn. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154392 Approved by: https://github.com/sraikund16, https://github.com/cyyever	2025-07-16 18:00:11 +00:00
Yidi Wu	da05b7fb94	[cond] add _FlopCounterMode support for cond (#158067 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158067 Approved by: https://github.com/zou3519 ghstack dependencies: #158077	2025-07-16 17:26:20 +00:00
Yidi Wu	82b1c48292	[hop] add supports_higher_order_operators flag to TorchDispatchMode (#158077 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158077 Approved by: https://github.com/zou3519	2025-07-16 17:26:20 +00:00
yuchengliu1	a369350065	enable compiled autograd on CPU windows (#158432 ) compiled autograd on windows is disabled in PR #144707 because cuda windows cannot compile this code. However these code can be compiled on CPU. This PR enable these code on CPU windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158432 Approved by: https://github.com/jansel, https://github.com/xmfan Co-authored-by: Xu Han <xu.han@outlook.com>	2025-07-16 17:22:37 +00:00
Nichols A. Romero	ff611d971f	[ROCm] check stream graph capture status in memcpy_and_sync inline function (#158165 ) Check for stream graph capture when using hipMemcpyWithStream. Fixes https://github.com/pytorch/pytorch/issues/155684, https://github.com/pytorch/pytorch/issues/155231 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158165 Approved by: https://github.com/jeffdaily	2025-07-16 17:17:34 +00:00
Han, Xu	4805a6ead6	[aot][XPU] switch xpu to use consts cpp build. (#158425 ) Intel compiler is not support `format_consts_to_asm`, let's use `format_consts_to_cpp`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158425 Approved by: https://github.com/jansel	2025-07-16 16:19:33 +00:00
Sam Larsen	a8b9736737	[BE][testing] disable test_custom_op_square internally (#158367 ) Summary: test is failing with `ld.lld: error: unable to find library -laoti_custom_ops` Test Plan: `buck test '@fbcode//mode/opt' fbcode//caffe2/test/inductor:test_aot_inductor_custom_ops -- --exact 'caffe2/test/inductor:test_aot_inductor_custom_ops - test_custom_op_square_cuda (caffe2.test.inductor.test_aot_inductor_custom_ops.AOTInductorTestABICompatibleCuda)' --run-disabled` Differential Revision: [D78364617](https://our.internmc.facebook.com/intern/diff/D78364617) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158367 Approved by: https://github.com/desertfire	2025-07-16 16:16:14 +00:00
Sam Larsen	4b11428cb5	[BE][testing] Skip test_repeated_masked_load internally (#158355 ) Summary: Test is failing internally because of the import from functorch.einops. _Maybe_ there's a way to get this dependence in the TARGETS file, but the obvious things didn't work. I'm wondering if this test is that important to have running in OSS and internally anyway? Test Plan: `buck test '@fbcode//mode/opt' fbcode//caffe2/test/inductor:cuda_repro -- --exact 'caffe2/test/inductor:cuda_repro - test_repeated_masked_load (caffe2.test.inductor.test_cuda_repro.CudaReproTests)' --run-disabled` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158355 Approved by: https://github.com/eellison	2025-07-16 16:15:44 +00:00
Sam Larsen	a04a13c449	[BE][testing] Skip test_triton_interpret internally (#158260 ) Summary: Subprocesses in fbcode are tricky because of .par files. I'm thinking it's not an important enough test to get it running and skipping is fine. Test Plan: `buck test` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158260 Approved by: https://github.com/eellison	2025-07-16 16:14:44 +00:00
tvukovic-amd	a23f4471b9	[ROCm][Windows] Fix finding ROCm/HIP version (#156486 ) This commit fixes Windows build issue related to trying to use rocm-core (rocm-core doesn't exist on HIP SDK) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156486 Approved by: https://github.com/jeffdaily, https://github.com/stellaraccident	2025-07-16 15:31:43 +00:00
Jithun Nair	06a67a8948	Fix sha256 for aotriton ROCm7.0 tarball (#158420 ) Fixes following issue of building PyTorch with ROCm7.0: ``` -- verifying file... file='/var/lib/jenkins/pytorch/build/aotriton_external-prefix/src/aotriton-0.10b-manylinux_2_28_x86_64-rocm7.0-shared.tar.gz' -- SHA256 hash of /var/lib/jenkins/pytorch/build/aotriton_external-prefix/src/aotriton-0.10b-manylinux_2_28_x86_64-rocm7.0-shared.tar.gz does not match expected value expected: '7e29c325d5bd33ba896ddb106f5d4fc7d715274dca7fe937f724fffa82017838' actual: '1e9b3dddf0c7fc07131c6f0f5266129e83ce2331f459fa2be8c63f4ae91b0f5b' -- Hash mismatch, removing... CMake Error at aotriton_external-prefix/src/aotriton_external-stamp/download-aotriton_external.cmake:163 (message): Each download failed! ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158420 Approved by: https://github.com/jeffdaily	2025-07-16 15:24:20 +00:00
PyTorch MergeBot	9513b9d03f	Revert "Support DeepSeek-style blockwise scaling scaled-mm for fp8 on Hopper+ (#158037 )" This reverts commit bc65253369933160a2da3fc786d027a572faf6b7. Reverted https://github.com/pytorch/pytorch/pull/158037 on behalf of https://github.com/lw due to OSX failures are real ([comment](https://github.com/pytorch/pytorch/pull/158037#issuecomment-3079042171))	2025-07-16 15:04:10 +00:00
David Berard	0b19d463d9	forward fix lint (#158448 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158448 Approved by: https://github.com/adamomainz	2025-07-16 14:55:33 +00:00
Ke Wen	5763ec5f8d	[BE] Replace lib with TORCH_INSTALL_LIB_DIR (#158235 ) Their values are actually the same. Just staying in line with other `INSTALL` commands. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158235 Approved by: https://github.com/Skylion007 ghstack dependencies: #158234	2025-07-16 14:20:19 +00:00
Ke Wen	2043f6911e	[BE] Rename libnvshmem_extension to libtorch_nvshmem (#158234 ) `libnvshmem_extension.so` creates an illusion that it is a shared library from NVSHMEM. But indeed it is built from torch source code, for symmetric tensor infrastructure and operations, though leveraging NVSHMEM APIs. Thus this PR renames `libnvshmem_extension.so` to `libtorch_nvshmem.so`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158234 Approved by: https://github.com/albanD	2025-07-16 14:20:19 +00:00
Luca Wehrstedt	bc65253369	Support DeepSeek-style blockwise scaling scaled-mm for fp8 on Hopper+ (#158037 ) cuBLAS added support for them in CUDA 12.9. It's rather easy to call into them, the hardest thing is allowing the lhs and rhs operands to have different scaling types, as that changes the whole callstack. The scaling format is still detected from the sizes of the scale tensors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158037 Approved by: https://github.com/eqy, https://github.com/drisspg	2025-07-16 13:54:09 +00:00
dolpm	51a708ffc6	[nativert] libtorch kernel registry (#157150 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D77451703 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157150 Approved by: https://github.com/georgiaphillips, https://github.com/henryoier	2025-07-16 12:36:55 +00:00
Raymond Li	55d888a616	Add framework for explanations for common CUDA errors (#158395 ) As popularly requested in user groups. Test plan: ``` import torch a = torch.randn(10000) device = torch.device('cuda:1') a = a.to(device) ``` Before: ``` Traceback (most recent call last): File "/data/users/raymo/pytorch/test/cuda.py", line 6, in <module> a = a.to(device) ^^^^^^^^^^^^ torch.AcceleratorError: CUDA error: invalid device ordinal CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1 Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions. ``` After: ``` Traceback (most recent call last): File "/data/users/raymo/pytorch/test/cuda.py", line 6, in <module> a = a.to(device) ^^^^^^^^^^^^ torch.AcceleratorError: CUDA error: invalid device ordinal GPU device may be out of range, do you have enough GPUs? CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1 Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158395 Approved by: https://github.com/aorenste Co-authored-by: Aaron Orenstein <aorenste@fb.com>	2025-07-16 12:31:18 +00:00
Andrey Talman	0a99b026d6	[Docker builds] Move from Miniconda to Miniforge (#158370 ) This is related to: https://www.anaconda.com/legal/terms/terms-of-service Trying to fix outage with docker builds. https://github.com/pytorch/pytorch/actions/runs/16298993712/job/46033590799 Rocm and XPU builds since they use Miniforge are not affected ``` #22 ERROR: process "/bin/sh -c bash ./install_conda.sh && rm install_conda.sh install_magma_conda.sh common_utils.sh /opt/conda/requirements-ci.txt /opt/conda/requirements-docs.txt" did not complete successfully: exit code: 1 ------ > [base 14/42] RUN bash ./install_conda.sh && rm install_conda.sh install_magma_conda.sh common_utils.sh /opt/conda/requirements-ci.txt /opt/conda/requirements-docs.txt: 11.93 CondaToSNonInteractiveError: Terms of Service have not been accepted for the following channels. Please accept or remove them before proceeding: 11.93 • https://repo.anaconda.com/pkgs/main 11.93 • https://repo.anaconda.com/pkgs/r 11.93 11.93 To accept a channel's Terms of Service, run the following and replace `CHANNEL` with the channel name/URL: 11.93 ‣ conda tos accept --override-channels --channel CHANNEL ``` Hence solution is: 1. using `` conda tos accept --override-channels --channel defaults`` 2. use Miniforge instead of Miniconda. Using solution 2. Solution Tried that don't work: 1. Using ``CONDA_ALWAYS_YES = true `` 4. Using older version of miniconda ``` [Miniconda3-py310_25.5.1-0-Linux-x86_64.sh](https://repo.anaconda.com/miniconda/Miniconda3-py310_25.5.1-0-Linux-x86_64.sh) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158370 Approved by: https://github.com/seemethere Co-authored-by: Eli Uriegas <1700823+seemethere@users.noreply.github.com>	2025-07-16 10:52:47 +00:00
bobrenjc93	ac706bfc7f	disable multi kernel rocm (#158299 ) Fixes https://github.com/pytorch/pytorch/issues/158274 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158299 Approved by: https://github.com/huydhn	2025-07-16 10:20:09 +00:00
Hari Krishna Sai Kodali	9d184bda2f	add device generalization support for distributed tests (#156796 ) MOTIVATION To generalize Distributed test cases for non-CUDA devices CHANGES - test/distributed/checkpoint/test_fsspec.py - test/distributed/checkpoint/test_state_dict.py - test/distributed/test_multi_threaded_pg.py Replaced hard coded device names with torch.accelerator.current_accelerator - torch/testing/_internal/distributed/_shard/sharded_tensor/__init__.py support for hccl backend Pull Request resolved: https://github.com/pytorch/pytorch/pull/156796 Approved by: https://github.com/guangyey, https://github.com/ezyang	2025-07-16 09:37:03 +00:00
NikhilAPatel	ea74fdd24a	[Inductor][Triton] Update TMA Compatibility Requirements (#157881 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157881 Approved by: https://github.com/Skylion007, https://github.com/drisspg	2025-07-16 09:31:44 +00:00
Huy Do	e71bb021b9	Add a periodic test for older NVIDIA driver (#158300 ) This is needed because of the botched landing of https://github.com/pytorch/pytorch/pull/156097 which crashed on older NVIDIA drivers `525.*`. I add a periodic job to install the `525.105.17` on CI, then run: 1. A smoke to make sure that CUDA can be initialized 2. And the whole the test suite on the older driver Pull Request resolved: https://github.com/pytorch/pytorch/pull/158300 Approved by: https://github.com/ngimel	2025-07-16 08:18:18 +00:00
Manuel Candales	fb9a5d248f	Fix torch._numpy to match NumPy when empty ellipsis causes advanced indexing separation (#158297 ) Fixes #141563 In NumPy, an ellipsis always acts as a separator between advanced indices, even when the ellipsis doesn't actually match any dimensions. In PyTorch an empty ellipsis doesn't cause a separation. This leads to differing behavior between Numpy and PyTorch in this edge case. This difference in behavior leads to a bug when using torch.compile: ```python >>> import numpy as np >>> f = lambda x: x[:,(0,1),...,(0,1)].shape >>> a = np.ones((3, 4, 5)) >>> f(a) (2, 3) >>> torch.compile(f)(a) (3, 2) ``` Similarly to #157676, this PR doesn't change PyTorch's behavior, but it fixes the translation layer, ensuring torch._numpy compatibility with NumPy. I am marking this PR as fixing #141563, even though PyTorch behavior isn't modified. Notice that there are still some other bugs in PyTorch's advanced indexing, that need to be fixed (mainly regarding proper accounting of dimensions when multidimensional boolean masks are present). But those need to be fixed at the ATen operator level. Examples: - #71673 - #107699 - #158125 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158297 Approved by: https://github.com/soumith	2025-07-16 08:11:53 +00:00
Huamin Li	ddf502c988	[AOTI] add -lstdc++ into aoti link cmd for Meta internal (#158325 ) Differential Revision: D78123716 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158325 Approved by: https://github.com/desertfire	2025-07-16 07:55:08 +00:00
FFFrog	555f356254	[Easy] Show some clear error when torch.ops.load_library fails. (#157524 ) Background: ```Shell torch 2.5.1+cpu torchvision 0.20.1 ``` ```Python import torch import torchvision Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torchvision/__init__.py", line 10, in <module> from torchvision import _meta_registrations, datasets, io, models, ops, transforms, utils # usort:skip File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torchvision/_meta_registrations.py", line 164, in <module> def meta_nms(dets, scores, iou_threshold): File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torch/library.py", line 795, in register use_lib._register_fake(op_name, func, _stacklevel=stacklevel + 1) File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torch/library.py", line 184, in _register_fake handle = entry.fake_impl.register(func_to_register, source) File "/usr/local/anaconda3/envs/test/lib/python3.10/site-packages/torch/_library/fake_impl.py", line 31, in register if torch._C._dispatch_has_kernel_for_dispatch_key(self.qualname, "Meta"): RuntimeError: operator torchvision::nms does not exist ``` Cause: ``` torchvision's .so file lacks some symbol definitions, because these symbols come from CUDA, but the current environment does not have CUDA and GPU. The above error message is very confusing. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157524 Approved by: https://github.com/ezyang	2025-07-16 07:33:22 +00:00
Kaichao You	59f9b25f3c	[cuda][cupy] Improve cupy device placement when device is provided (#158320 ) This is an improvement over https://github.com/pytorch/pytorch/pull/132595 . That PR improves the case where `device` is not given. This PR tries to improve the case where `device` is given but the first step of auto-infer device from `cudaPointerGetAttributes` can be wrong (undesired). See https://github.com/pytorch/pytorch/issues/158316 for more details on when this can happen. I think this is a reasonable improvement, as people expect `torch.as_tensor` + cupy should be zero-copy as much as possible. However, it does change some behaviors, because previously it might incur a device-to-device copy. I will leave it to pytorch developers to see if the improvement is worthwhile. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158320 Approved by: https://github.com/ezyang	2025-07-16 07:12:36 +00:00
saienduri	fedbd1a48e	Enable ROCm 7.0 Alpha docker builds for PyTorch CI (#158390 ) This PR adds ROCm 7.0 alpha docker builds to start testing latest ROCm in PyTorch CI and enable new MI350x hardware. Highlights: * Stop building `pytorch-linux-jammy-rocm-n-1-py3` docker images, as they're not currently used in any CI workflows * Add `pytorch-linux-noble-rocm-alpha-py3` docker images that will use ROCm alpha (newer than latest official release) builds Pull Request resolved: https://github.com/pytorch/pytorch/pull/158390 Approved by: https://github.com/jithunnair-amd, https://github.com/jeffdaily	2025-07-16 06:09:37 +00:00
drisspg	5484890539	Add better typing to avaialbe kernel options for flex attention (#158383 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158383 Approved by: https://github.com/joydddd, https://github.com/BoyuanFeng	2025-07-16 06:06:29 +00:00
Xuehai Pan	61a7b09ef3	[BE][Easy] split build system `requirements.txt` to a separate file (#158111 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158111 Approved by: https://github.com/ezyang	2025-07-16 05:03:30 +00:00
Denghui Dong	e92e3eaf4e	[Profiler] the doc of _ExperimentalConfig is incorrectly truncated by commas (#156586 ) Hi team, Please help review this trivial fix. Without this change: ``` python >>> import torch >>> print(torch._C._profiler._ExperimentalConfig.__init__.__doc__) __init__(self: torch._C._profiler._ExperimentalConfig, profiler_metrics: list[str] = [], profiler_measure_per_kernel: bool = False, verbose: bool = False, performance_events: list[str] = [], enable_cuda_sync_events: bool = False, adjust_profiler_step: bool = False, disable_external_correlation: bool = False, profile_all_threads: bool = False, capture_overload_names: bool = False) -> None capture_overload_names (bool) : whether to include ATen overload names in the profile ``` With this change: ```python >>> import torch >>> print(torch._C._profiler._ExperimentalConfig.__init__.__doc__) __init__(self: torch._C._profiler._ExperimentalConfig, profiler_metrics: list[str] = [], profiler_measure_per_kernel: bool = False, verbose: bool = False, performance_events: list[str] = [], enable_cuda_sync_events: bool = False, adjust_profiler_step: bool = False, disable_external_correlation: bool = False, profile_all_threads: bool = False, capture_overload_names: bool = False) -> None An experimental config for Kineto features. Please note thatbackward compatibility is not guaranteed. profiler_metrics : a list of CUPTI profiler metrics used to measure GPU performance events. If this list contains values Kineto runs in CUPTI profiler mode profiler_measure_per_kernel (bool) : whether to profile metrics per kernel or for the entire measurement duration. verbose (bool) : whether the trace file has `Call stack` field or not. performance_events : a list of profiler events to be used for measurement. enable_cuda_sync_events : for CUDA profiling mode, enable adding CUDA synchronization events that expose CUDA device, stream and event synchronization activities. This feature is new and currently disabled by default. adjust_profiler_step (bool) : whether to adjust the profiler step to match the parent python event duration. This feature is new and currently disabled by default. disable_external_correlation (bool) : whether to disable external correlation profile_all_threads (bool) : whether to profile all threads capture_overload_names (bool) : whether to include ATen overload names in the profile ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156586 Approved by: https://github.com/sraikund16, https://github.com/cyyever	2025-07-16 04:10:49 +00:00
Will Constable	0a9d450168	[DTensor] implement histc (#158298 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158298 Approved by: https://github.com/zpcore, https://github.com/XilunWu	2025-07-16 04:10:32 +00:00
Edward Z. Yang	e265b719bd	Extract out prepare_aot_module_simplified for use in next PR (#158319 ) Also a small amount of extra code cleanup. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158319 Approved by: https://github.com/jingsh ghstack dependencies: #158149, #158150, #158173, #158176, #158213, #158251	2025-07-16 03:59:41 +00:00
Edward Z. Yang	7637c9718a	Move functions from torch._functorch.aot_autograd that are not frontend functions to frontend_utils (#158251 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158251 Approved by: https://github.com/jamesjwu ghstack dependencies: #158149, #158150, #158173, #158176, #158213	2025-07-16 03:59:41 +00:00
Edward Z. Yang	49d0332cef	Introduce stages to aot_dispatch (#158213 ) The starting point for this refactor is that I need access to the fully general joint graph representation in an export-like interface, but I then subsequently need a way to feed this joint graph into the rest of the compilation pipeline so I can get an actual callable that I can run once I've finished modifying it. Previously, people had added export capabilities to AOTAutograd by having an export flag that toggled what exactly the functions return and triggering aot_dispatch to go to a different "export" implementation, but I've found this difficult to understand and has lead to a bit of duplicate code for the export path. So the idea here is to reorganize the structure of the function calls in AOTAutograd. Here, it is helpful to first describe how things used to work: * Start with aot_autograd.py top level functions like aot_function, _aot_export_function and aot_module_simplified. These call: * create_aot_dispatcher_function. This does a bunch of stuff (forward metadata collection) and adds many context managers. This calls: * One of aot_dispatch_base, aot_dispatch_export or aot_dispatch_autograd, which: * Call aot_dispatch_autograd_graph or aot_dispatch_base_graph to actually do the graph capture * Do some base/export/autograd specific post-processing on the graph Notice the pattern of nested function invocations means that there is no way to easily get the graph capture result from the autograd case; furthermore, the export path is "bolted" on to force the entire chain of functions to have a different return result than normal, and no way to resume the rest of the post-processing to actually get a callable. Here is the new structure: * Start with aot_autograd.py top level functions like aot_function, _aot_export_function and aot_module_simplified. These now orchestrate this top level flow: * Start a context manager (stack); this stateful context block takes care of all of the nested context managers which originally necessitated the nested call structure * Call create_aot_state to do initial setup and setup all the context managers on stack. These context managers do NOT exit upon return of this. * Call aot_stage1_graph_capture to do the graph capture * Call aot_stage2_compile or aot_stage2_export depending on what postprocessing you want With this new structure, it's now possible (although not done in this PR) to return the graph after aot_stage1_graph_capture and do something with it, before running aot_stage2_compile to finish the job. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158213 Approved by: https://github.com/jamesjwu ghstack dependencies: #158149, #158150, #158173, #158176	2025-07-16 03:59:32 +00:00
Edward Z. Yang	84dec060b7	Hoist choose_dispatcher to top level, remove unnecessary returns (#158176 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158176 Approved by: https://github.com/jamesjwu ghstack dependencies: #158149, #158150, #158173	2025-07-16 03:56:25 +00:00
Edward Z. Yang	5b0df2565e	Pipeline _create_aot_dispatcher_function (#158173 ) Two main things of note: - Review this diff without whitespace changes - To ensure that context managers correctly propagate to later pipeline stages, I am using the ExitStack trick: there is an ExitStack which is in scope for the entire pipeline, and inside of the individual pipeline stages we push context managers onto this stack when we want them to survive into the next pipeline stage. This is not obviously what the best final form of the code is, but create_aot_dispatcher_function is called from multiple locations so I can't just inline the context managers into the call site. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158173 Approved by: https://github.com/jamesjwu, https://github.com/wconstab ghstack dependencies: #158149, #158150	2025-07-16 03:56:25 +00:00
Songhao Jia	0cb36e2d62	cache dict and string rep for better perf (#158372 ) Summary: NodeSouce should not be updated after created, so that it would be better if we cache its dict and string representation for better perf. Test Plan: ci Rollback Plan: Reviewed By: yushangdi Differential Revision: D78298501 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158372 Approved by: https://github.com/yushangdi	2025-07-16 02:15:32 +00:00
Xu Han	584a0510b3	[inductor] fix windows path for fresh cache. (#158324 ) `normalize_path_separator` for windows path. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158324 Approved by: https://github.com/jansel	2025-07-16 01:54:35 +00:00
yuchengliu1	9768d393fa	add sfdp pattern (#155792 ) add sfdp pattern for MBartForCausalLM/PLBartForCausalLM in transformers==4.44.2. Improve the inference performance of these model. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155792 Approved by: https://github.com/Valentine233, https://github.com/jansel	2025-07-16 01:52:05 +00:00
Jiang, Yanbing	900fba4c07	Update warning of TF32 (#158209 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/158209 Approved by: https://github.com/jansel	2025-07-16 01:28:50 +00:00
PyTorch MergeBot	03852ddc22	Revert "[ROCm] logsumexp on ROCm needs scaling back to natural base. (#156903 )" This reverts commit 1ea9cde598ead20194dbb6c5cb26e74e36e6ad55. Reverted https://github.com/pytorch/pytorch/pull/156903 on behalf of https://github.com/atalman due to Breaks torchao and torchtitan nightly builds ([comment](https://github.com/pytorch/pytorch/pull/156903#issuecomment-3076423488))	2025-07-16 01:28:46 +00:00
Xuan Zhang	8554c8007d	[PT2][fusion] ban fusions with large accumulated reads (#157563 ) Problem: Fusion can accumulate large amount of reads, which leads to significant increase in peak memory utilization. Imagine we have the following code snippet ``` total = torch.rand(N, N) for _ in range(r): x = torch.rand(N, N) total = total + x ``` The default execution is memory efficient as only two tensors of size N-by-N is in memory at any given time. However, with fusion, the additions are fused into a single operation and the execution becomes something like: ``` x_1 = torch.rand(N, N) x_2 = torch.rand(N, N) ... x_r = torch.rand(N, N) total = x_1 + x_2 + ... + x_r ``` Though this is run-time efficient, in the case of large `N` and/or large `r`, this is not memory efficient. [internal only] see [post](https://fb.workplace.com/groups/1075192433118967/permalink/1703374333634104/) for additional details Solution: Our proposed solution is to ban fusions in case where a large amount of reads are accumulated. This is in addition to some existing logics during torch compile. * During lowering (i.e., `ir.py`), the config `realize_acc_reads_threshold`, which is default to be 8, controls _the number of_ buffers can be accumulated for a single operator. However, this is oblivious to the size of the buffers. Hence, we additionally introduce a config `realize_acc_reads_size_threshold` to control _the amount of buffers_ in size that can be accumulated. * During scheduling (i.e., `scheduler.py`), additional fusion will be performed and thus we also need to capture such pattern there. The decisions are implemented under `choices.py`. Results: For a small example similar to be one in the test case (but with larger `N` and higher number of loop repeats), the memory snapshot before and after are shown below. Note the snapshot on the right is zoomed out so that the y-axis of the two snapshots match. <img width="1328" alt="image" src="https://github.com/user-attachments/assets/670b5961-8454-4379-ae0f-62d4e7946c64" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157563 Approved by: https://github.com/jansel, https://github.com/mlazos	2025-07-16 01:05:25 +00:00
Yidi Wu	651b4a68f2	[hop][dynamo] track run-ahead sym variables in side effects (#158273 ) Before the PR, for code like this: ``` class Example2(torch.nn.Module): def forward(self, x, trigger, target): return torch.cond( trigger == 1, lambda: x + target, lambda: x * target, (), ) m = Example2() x = torch.randn(2) trigger = 0 target = 2 args = (x, trigger, target) ep = torch.export.export( m, args, dynamic_shapes=(None, Dim.DYNAMIC, Dim.DYNAMIC) ) ``` dynamo will wrap "target" (i.e. a symInt) twice, once when we speculate the first lambda and find target is a symint and decides to wrap it up, creating a new SymNodeVariable and a placeholder input to the top-level graph. The second time happens when we speculate the second lambda. Tensors are de-duplicated by checking tracked side effects to make sure object with the same id (though different sources) is mapped to the same TensorVaraible. For symints, two things are missing: 1. it's not in the _can_lift_attrs_to_input list (the change in builder.py) 2. it's not in the tracked by runahead_side_effects, so when speculate_subgraph finishes, they're discarded (the change in side_effects.py) Note: the auto lifting mechanism for HOPs happens at proxy level when we trace the subgraph, which is after SymNodeVariable are created (they're created when realizing the args and bind them to subgraph). At that time, builder has created two unique SymNodeVariable for the same symint so the auto lifting in hops cannot de-dup them. Differential Revision: [D78298163](https://our.internmc.facebook.com/intern/diff/D78298163) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158273 Approved by: https://github.com/avikchaudhuri, https://github.com/zou3519	2025-07-15 23:48:20 +00:00
dolpm	144965ca9a	[BE][S538760] get rid of TORCH_CHECK_.* and CHECK macros (#158269 ) Summary: check will be crit, causing program to exit, which is quite dangerous Test Plan: CI Rollback Plan: Differential Revision: D78050595 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158269 Approved by: https://github.com/SherlockNoMad, https://github.com/henryoier	2025-07-15 22:04:12 +00:00
dsashidh	ee0992871c	Add test for user-managed weights with load_state_dict (#157496 ) Summary: Adds a unit test to verify that when 'user_managed=True' is passed to 'update_constant_buffer', the compiled AOTI model properly shares parameter storage with the eager model. The test specifically covers the following: 1. Passes model weights to the AOTI model with 'user_managed=True''. 2. Updates the eager model weights using 'load_state_dict()', which performs in-place 3. Asserts that the compiled AOTI model reflects the updated weights, confirming shared memory behavior. Fixes: #157474 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157496 Approved by: https://github.com/desertfire	2025-07-15 21:17:24 +00:00
Yiming Zhou	05dfd312cf	[3/n] Remove references to TorchScript in PyTorch docs (#158315 ) Summary: - cpp_index.rst - fx.md - jit_builtin_functions.rst - jit_python_reference.md - jit_unsupported.md cpu_threading large_scale_deployment Test Plan: CI Rollback Plan: Differential Revision: D78309320 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158315 Approved by: https://github.com/svekars, https://github.com/zhxchen17	2025-07-15 21:14:18 +00:00
Andrey Talman	abeae997a3	Use brew suggested miniconda install command (#158347 ) Use ```brew install --cask miniconda``` as specified by https://formulae.brew.sh/cask/miniconda Forward fix After: https://github.com/pytorch/pytorch/pull/156898#issuecomment-3074207175 Seeing in CI: ``` Run if [[ -n "$REINSTALL_BREW_MINICONDA" ]]; then ==> Caveats Please run the following to setup your shell: conda init "$(basename "${SHELL}")" Alternatively, manually add the following to your shell init: eval "$(conda "shell.$(basename "${SHELL}")" hook)" ==> Downloading https://repo.anaconda.com/miniconda/Miniconda3-py313_25.5.1-0-MacOSX-arm64.sh Already downloaded: /Users/ec2-user/Library/Caches/Homebrew/downloads/2e356e8b147647692e4da77ce4c0c14eefee65ec86f29cc7e8c21a26ac9397ca--Miniconda3-py313_25.5.1-0-MacOSX-arm64.sh ==> Installing Cask miniconda ==> Running installer script 'Miniconda3-py313_25.5.1-0-MacOSX-arm64.sh' PREFIX=/opt/homebrew/Caskroom/miniconda/base Unpacking payload ... entry_point.py:256: DeprecationWarning: Python 3.14 will, by default, filter extracted tar archives and reject files or modify their metadata. Use the filter argument to control this behavior. entry_point.py:256: DeprecationWarning: Python 3.14 will, by default, filter extracted tar archives and reject files or modify their metadata. Use the filter argument to control this behavior. Installing base environment... Preparing transaction: ...working... done Executing transaction: ...working... done entry_point.py:256: DeprecationWarning: Python 3.14 will, by default, filter extracted tar archives and reject files or modify their metadata. Use the filter argument to control this behavior. installation finished. ==> Linking Binary 'conda' to '/opt/homebrew/bin/conda' 🍺 miniconda was successfully installed! ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158347 Approved by: https://github.com/seemethere	2025-07-15 21:08:25 +00:00
Ti-Tai Wang	3f83e3eeca	[ONNX] Remove legacy registration and dispatcher (#158283 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158283 Approved by: https://github.com/Skylion007, https://github.com/justinchuby ghstack dependencies: #158258, #158262, #158282	2025-07-15 21:00:49 +00:00
Yiming Zhou	0640cfa38c	[2/n] Remove references to TorchScript in PyTorch docs (#158306 ) Summary: Removed jit_language_reference.md Test Plan: CI Rollback Plan: Differential Revision: D78308133 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158306 Approved by: https://github.com/svekars, https://github.com/zhxchen17	2025-07-15 20:57:23 +00:00
Ti-Tai Wang	e4c17d5e1c	[ONNX] Remove fx_onnx_interpreter.py (#158282 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158282 Approved by: https://github.com/Skylion007, https://github.com/justinchuby ghstack dependencies: #158258, #158262	2025-07-15 20:46:06 +00:00
Animesh Jain	cc0faeb80f	[dynamo][guards] Instruction count for guard eval for development work (#158214 ) Its turned off by default. Even the code is hidden before of the define preprocessing flag. It will be used only for development work. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158214 Approved by: https://github.com/StrongerXi ghstack dependencies: #158215	2025-07-15 20:29:23 +00:00
Ti-Tai Wang	205241a0d5	[ONNX] Remove legacy dynamo graph extractor (#158262 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158262 Approved by: https://github.com/justinchuby ghstack dependencies: #158258	2025-07-15 20:21:49 +00:00
Yiming Zhou	19625daf88	[1/n] Remove references to TorchScript in PyTorch docs (#158305 ) Summary: Removed jit_language_reference_v2.md Test Plan: CI Rollback Plan: Differential Revision: D78308009 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158305 Approved by: https://github.com/jingsh, https://github.com/svekars	2025-07-15 20:16:53 +00:00
Sam Larsen	dbf7d421da	[BE][testing] fix aot_inductor_package internally (#158270 ) Summary: We have internal test failure for several aot_inductor_package tests. It looks like we're translating args like: ``` -Wl,--script=/home/slarsen/local/fbsource2/buck-out/v2/gen/fbcode/7ce8f48f92bc4ee6/caffe2/test/inductor/__aot_inductor_package__/aot_inductor_package#link-tree/torch/_inductor/script.ld ``` To: ``` -Wl,--script=/home/slarsen/local/fbsource2/buck-out/v2/gen/fbcode/7ce8f48f92bc4ee6/caffe2/test/inductor/__aot_inductor_package__/aot_inductor_package#link-tree/torch/_inductor//tmp/jZMktZ/tmpsqoxb_cq/data/aotinductor/model/script.ld ``` This PR changes to strings like: ``` -Wl,--script=/tmp/jZMktZ/tmpsqoxb_cq/data/aotinductor/model/script.ld ``` Test Plan: `buck test '@fbcode//mode/opt' fbcode//caffe2/test/inductor:aot_inductor_package --run-disabled` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158270 Approved by: https://github.com/desertfire	2025-07-15 20:15:18 +00:00
Animesh Jain	b86d5cef68	[dynamo][tensor] Skip HASATTR attribute on tensor guards (#158215 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158215 Approved by: https://github.com/StrongerXi	2025-07-15 20:10:47 +00:00
Jane Xu	30587195d3	Migrate c10/macros/cmake_macros.h.in to torch/headeronly (#158035 ) Summary: As above, also changes a bunch of the build files to be better Test Plan: internal and external CI did run buck2 build fbcode//caffe2:torch and it succeeded Rollback Plan: Reviewed By: swolchok Differential Revision: D78016591 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158035 Approved by: https://github.com/swolchok	2025-07-15 19:52:59 +00:00
Aaron Orenstein	250ae2531c	Fix types in graphs.py (#158192 ) Added type annotations for torch/cuda/graphs.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/158192 Approved by: https://github.com/oulgen	2025-07-15 19:49:38 +00:00
Songhao Jia	011026205a	make node source hashable (#158322 ) Summary: as title Test Plan: ci Rollback Plan: Reviewed By: yushangdi Differential Revision: D78296410 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158322 Approved by: https://github.com/yushangdi	2025-07-15 19:31:00 +00:00
Menglu Yu	4657a84bc5	[Optimus][fp8_activation_quantization] Only log when there's some node to be quantized (#158129 ) Summary: We add some extra check on whether there's some node has been marked as should quantize, otherwise we skip the quantizaton and tlparse log. Rollback Plan: Differential Revision: D78173788 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158129 Approved by: https://github.com/Skylion007, https://github.com/avicizhu	2025-07-15 19:22:26 +00:00
Ti-Tai Wang	5606c516fd	[ONNX] Remove legacy Dort (#158258 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158258 Approved by: https://github.com/justinchuby, https://github.com/malfet	2025-07-15 19:14:06 +00:00
Edward Z. Yang	7afb834f93	Inline dispatch_and_compile into its call site. (#158150 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158150 Approved by: https://github.com/jamesjwu, https://github.com/wconstab ghstack dependencies: #158149	2025-07-15 19:08:55 +00:00
Edward Z. Yang	148789ddd8	Avoid AOTAutogradCache.load in stack trace on cache miss path (#158149 ) The general context for the upcoming stack of commits is I am attempting to "pipeline" AOTAutograd. Instead of having function f call function g which is the next "stage" of compilation, instead f should return with its outputs, which are then piped to g for the next stage. This will make it easier to implement early exit / resume pipeline without forcing callback structure, which is good for export-style use cases. It also reduces the size of our stack traces, which makes tools like Perfetto happy. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158149 Approved by: https://github.com/jamesjwu	2025-07-15 19:08:55 +00:00
albanD	3beb915004	Update CODEOWNERS for dataloading (#158348 ) Adding Scott Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/158348 Approved by: https://github.com/scotts, https://github.com/janeyx99	2025-07-15 19:06:18 +00:00
Shangdi Yu	cf3247b74a	Standalone compile API in _Exporter (#158139 ) Given an `package: _ExportPackage`, users can get a ready-to-use workspace in `tmp_dir` by calling: ```python package._compiled_and_package( tmp_dir + "/pt2_pacakge_name.pt2", True, package_example_inputs = True ) ``` `tmp_dir` will contains: - `main.cpp` (an example cpp file that create the models, if package_example_inputs is True, it'll also load the example inputs and run the models) - `CMakeLists.txt` - `pt2_pacakge_name/` (this is where the models are) - `pt2_pacakge_name.pt2` - `inputs.pt` files if package_example_inputs is True Remaining TODOs - support loading contants/weights - the `package_example_inputs = True` option only supports a list of Tensors for now - eventually we should remove the `torch` dependency, and use `SlimTensor`/`StableIValue` instead. Test Plan: ``` python test/inductor/test_aot_inductor_package.py -k test_compile_with_exporter ``` Example generated `main.cpp`: ```cpp #include <dlfcn.h> #include <fstream> #include <iostream> #include <memory> #include <torch/torch.h> #include <vector> #include <torch/csrc/inductor/aoti_torch/tensor_converter.h> #include "package/data/aotinductor/Plus__default/Plus__default.h" #include "package/data/aotinductor/Minus__default/Minus__default.h" using torch::aot_inductor::AOTInductorModelPlus__default; using torch::aot_inductor::AOTInductorModelMinus__default; using torch::aot_inductor::ConstantHandle; using torch::aot_inductor::ConstantMap; int main(int argc, char* argv[]) { std::string device_str = "cpu"; try { c10::Device device(device_str); // Load input tensors for model Plus__default std::vector<at::Tensor> input_tensors1; for (int j = 0; j < 2; ++j) { std::string filename = "Plus__default_input_" + std::to_string(j) + ".pt"; std::ifstream in(filename, std::ios::binary); if (!in.is_open()) { std::cerr << "Failed to open file: " << filename << std::endl; return 1; } std::vector<char> buffer((std::istreambuf_iterator<char>(in)), std::istreambuf_iterator<char>()); torch::IValue ivalue = torch::pickle_load(buffer); input_tensors1.push_back(ivalue.toTensor().to(device)); } // Load input tensors for model Minus__default std::vector<at::Tensor> input_tensors2; for (int j = 0; j < 2; ++j) { std::string filename = "Minus__default_input_" + std::to_string(j) + ".pt"; std::ifstream in(filename, std::ios::binary); if (!in.is_open()) { std::cerr << "Failed to open file: " << filename << std::endl; return 1; } std::vector<char> buffer((std::istreambuf_iterator<char>(in)), std::istreambuf_iterator<char>()); torch::IValue ivalue = torch::pickle_load(buffer); input_tensors2.push_back(ivalue.toTensor().to(device)); } // Create array of input handles auto input_handles1 = torch::aot_inductor::unsafe_alloc_new_handles_from_tensors(input_tensors1); auto input_handles2 = torch::aot_inductor::unsafe_alloc_new_handles_from_tensors(input_tensors2); // Create array for output handles AtenTensorHandle output_handle1; AtenTensorHandle output_handle2; // Create and load models auto constants_map1 = std::make_shared<ConstantMap>(); auto constants_array1 = std::make_shared<std::vector<ConstantHandle>>(); auto model1 = AOTInductorModelPlus__default::Create( constants_map1, constants_array1, device_str, "package/data/aotinductor/Plus__default/"); model1->load_constants(); auto constants_map2 = std::make_shared<ConstantMap>(); auto constants_array2 = std::make_shared<std::vector<ConstantHandle>>(); auto model2 = AOTInductorModelMinus__default::Create( constants_map2, constants_array2, device_str, "package/data/aotinductor/Minus__default/"); model2->load_constants(); // Run the models torch::aot_inductor::DeviceStreamType stream1 = nullptr; model1->run(&input_handles1[0], &output_handle1, stream1, nullptr); torch::aot_inductor::DeviceStreamType stream2 = nullptr; model2->run(&input_handles2[0], &output_handle2, stream2, nullptr); // Convert output handles to tensors auto output_tensor1 = torch::aot_inductor::alloc_tensors_by_stealing_from_handles(&output_handle1, 1); auto output_tensor2 = torch::aot_inductor::alloc_tensors_by_stealing_from_handles(&output_handle2, 1); // Validate outputs std::cout << "output_tensor1" << output_tensor1 << std::endl; std::cout << "output_tensor2" << output_tensor2 << std::endl; return 0; } catch (const std::exception &e) { std::cerr << "Error: " << e.what() << std::endl; return 1; } } ``` Rollback Plan: Differential Revision: D78124705 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158139 Approved by: https://github.com/desertfire	2025-07-15 18:47:56 +00:00
PyTorch MergeBot	46915b1361	Revert "Introduce AcceleratorAllocatorConfig as the common class (#149601 )" This reverts commit 1e8e9f745e43fa38bbfc7b67b30bc66c0e7ebbd6. Reverted https://github.com/pytorch/pytorch/pull/149601 on behalf of https://github.com/huydhn due to See https://github.com/pytorch/pytorch/pull/149601#discussion_r2208325379 ([comment](https://github.com/pytorch/pytorch/pull/149601#issuecomment-3074965720))	2025-07-15 18:40:59 +00:00
Robert Hardwick	8c3f206457	Fix AArch64 segfaults by disabling strict-aliasing in GridSamplerKernel for GCC 12 and above (#158117 ) This PR disables `strict-aliasing` GCC C++ optimization flag on all AArch64 cpus for GCC versions 12 and above. Pull Request #152825 upgraded gcc version from 11 to 13 in manywheel which caused several segmentation faults in unit tests ( not visible in CI workflows because the jammy gcc version has not been updated yet ). We Identified the problem also exists in GCC12 hence the ` __GNUC__ >= 12` Fixes #157626 fixes these tests failures when pytorch is built in GCC12 and above ``` test_ops.py::TestCommonCPU::test_noncontiguous_samples_grid_sampler_2d_cpu_float32 Fatal Python error: Segmentation fault test_ops.py::TestCommonCPU::test_dtypes_grid_sampler_2d_cpu Fatal Python error: Segmentation fault test_ops.py::TestMathBitsCPU::test_neg_view_nn_functional_grid_sample_cpu_float64 free(): invalid next size (fast) test_ops.py::TestCompositeComplianceCPU::test_backward_grid_sampler_2d_cpu_float32 Fatal Python error: Segmentation fault test_ops.py::TestCommonCPU::test_dtypes_nn_functional_grid_sample_cpu Fatal Python error: Segmentation fault ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158117 Approved by: https://github.com/malfet	2025-07-15 18:26:38 +00:00
PyTorch MergeBot	41971335c9	Revert "Refactor CUDAAllocatorConfig to reuse AcceleratorAllocatorConfig (#150312 )" This reverts commit e241a07e6b88aa49d604803bc5a6562f0d9f94d2. Reverted https://github.com/pytorch/pytorch/pull/150312 on behalf of https://github.com/huydhn due to Sorry for reverting your change but because https://github.com/pytorch/pytorch/pull/157908 has been reverted + this PR caused issue earlier, I think it is better to revert the whole stack and reland it from scratch to be sure ([comment](https://github.com/pytorch/pytorch/pull/150312#issuecomment-3074897532))	2025-07-15 18:24:36 +00:00
PyTorch MergeBot	ea5f88dca6	Revert "Deprecate overleap functions in CUDAAllocatorConfig, use AcceleratorAllocatorConfig instead (#156165 )" This reverts commit e40ade5182233f548b25f2732effe3719d16e9ad. Reverted https://github.com/pytorch/pytorch/pull/156165 on behalf of https://github.com/huydhn due to Sorry for reverting your change but because https://github.com/pytorch/pytorch/pull/157908 has been reverted + this PR caused issue earlier, I think it is better to revert the whole stack and reland it from scratch to be sure ([comment](https://github.com/pytorch/pytorch/pull/150312#issuecomment-3074897532))	2025-07-15 18:24:36 +00:00
PyTorch MergeBot	f2ecf6145f	Revert "Enable AcceleratorAllocatorConfig key check (#157908 )" This reverts commit 65fcca4f8c97de82d35d51ad9b790d10433e9b91. Reverted https://github.com/pytorch/pytorch/pull/157908 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it is failing internally per https://github.com/pytorch/pytorch/pull/157908#discussion_r2208204782 ([comment](https://github.com/pytorch/pytorch/pull/157908#issuecomment-3074833696))	2025-07-15 18:17:43 +00:00
PyTorch MergeBot	b26da7741b	Revert "[CI] Fixes CI for CUDA Version > 12.9 (#157385 )" This reverts commit 6c5227ba00a2904365af566c24b4681cd01a041c. Reverted https://github.com/pytorch/pytorch/pull/157385 on behalf of https://github.com/clee2000 due to broke some slow tests test_cpp_extensions_jit.py::TestCppExtensionJIT::test_jit_cuda_archflags [GH job link](https://github.com/pytorch/pytorch/actions/runs/16286465717/job/45986677885) [HUD commit link](`6c5227ba00`) ([comment](https://github.com/pytorch/pytorch/pull/157385#issuecomment-3074737541))	2025-07-15 18:06:52 +00:00
Menglu Yu	243b12e565	[Optimus] add einsum_to_pointwise_pass pattern (#155666 ) Summary: More context: https://docs.google.com/document/d/1ipiskqG13ZKNX1SGygB3QnHcSyXNQ8pACazPIcS4bnI/edit?tab=t.0 Test Plan: ### how to enable ``` torch._inductor.config.pre_grad_fusion_options={ "einsum_to_pointwise_pass": {}, }, ``` ### unit test ``` CUDA_VISIBLE_DEVICES=3 OC_CAUSE=1 buck2 test 'fbcode//mode/dev-nosan' //caffe2/test/inductor:kernel_optimization ``` Buck UI: https://www.internalfb.com/buck2/267263ff-6f5b-4fff-bfc0-d8f013440ba0 Test UI: https://www.internalfb.com/intern/testinfra/testrun/5629499820839168 Network: Up: 61KiB Down: 675KiB (reSessionID-fda8edfc-6eef-4bf0-b268-0f8d2e666571) Loading targets. Remaining 0/1 1 dirs read, 2310 targets declared Analyzing targets. Remaining 0/345 284 actions, 329 artifacts declared Executing actions. Remaining 0/18334 8.0s exec time total Command: test. Finished 6 local Time elapsed: 1:15.5s Tests finished: Pass 2. Fail 0. Fatal 0. Skip 0. Build failure 0 ### local reproduce baseline: \| Metric \| Value \| \|:----------------------\|:------------\| \| Batch size \| 4096 \| \| GPU type \| H100 \| \| Latency \| 196.06 ms \| \| Model size \| 1205.21 MB \| \| Flops \| 7671.30 G \| \| Flops/example \| 1.87 G \| \| TFLOPS/sec \| 39.13 \| \| MFU \| 4.89% \| \| Activation/example \| 1.51 MB \| \| CPU time total \| 602.28 ms \| \| GPU time total \| 798.60 ms \| \| Estimated avg BW \| 234.62 GB/s \| \| Estimated avg BW util \| 9.78% \| Trace link: https://our.intern.facebook.com/intern/perfdoctor/trace_view?filepath=tree/traces/efficient_module_suite/fused_attention_mlp.Jun_09_22_12_38_trace.json.gz&bucket=pyper_traces with the pattern: \| Metric \| Value \| \|:----------------------\|:------------\| \| Batch size \| 4096 \| \| GPU type \| H100 \| \| Latency \| 184.94 ms \| \| Model size \| 1205.21 MB \| \| Flops \| 7671.30 G \| \| Flops/example \| 1.87 G \| \| TFLOPS/sec \| 41.48 \| \| MFU \| 5.18% \| \| Activation/example \| 1.15 MB \| \| CPU time total \| 562.44 ms \| \| GPU time total \| 754.36 ms \| \| Estimated avg BW \| 201.40 GB/s \| \| Estimated avg BW util \| 8.39% \| Trace link: https://our.intern.facebook.com/intern/perfdoctor/trace_view?filepath=tree/traces/efficient_module_suite/fused_attention_mlp.Jun_10_22_03_34_trace.json.gz&bucket=pyper_traces ### E2E baseline: f713998364 with patter: Rollback Plan: Differential Revision: D76400889 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155666 Approved by: https://github.com/Yuzhen11	2025-07-15 17:50:23 +00:00
vfdev	b7b1109f49	Expose opt_einsum in torch.backends (#157740 ) Fixes the following issue: ``` :/tmp# python -c "import torch; print(torch.__version__)" 2.7.1+cu126 :/tmp# python -c "import torch; print(torch.backends.opt_einsum.is_available())" Traceback (most recent call last): File "<string>", line 1, in <module> AttributeError: module 'torch.backends' has no attribute 'opt_einsum' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157740 Approved by: https://github.com/Skylion007, https://github.com/benjaminglass1	2025-07-15 17:46:43 +00:00
PyTorch MergeBot	26807dcf27	Revert "[PT2][fusion] ban fusions with large accumulated reads (#157563 )" This reverts commit c062550a3598d27c2d6572db7c0f4ff90a84cc84. Reverted https://github.com/pytorch/pytorch/pull/157563 on behalf of https://github.com/clee2000 due to broke test_linear_and_cel on main `c062550a35`, caused OOM? Also broken on PR, Dr. CI classification is wrong (claims the test is disabled by an issue but the issue is for a different test). Also I'm pretty sure the expected results json is supposed to have a ton of empty lines, its to prevent merge conflicts, I will add it to the linter ([comment](https://github.com/pytorch/pytorch/pull/157563#issuecomment-3074355331))	2025-07-15 16:35:55 +00:00
PyTorch MergeBot	4f36743f5e	Revert "[simple_fsdp][inductor_collectives] rewrite reorder_collectives, sink_waits_iterative (#158062 )" This reverts commit 5a54db14e3843cfa87fd8d27487dbf2f2dfb6c47. Reverted https://github.com/pytorch/pytorch/pull/158062 on behalf of https://github.com/clee2000 due to sorry I want to revert something else and this is causing a merge conflict, all you should need to do is rebase and remerged ([comment](https://github.com/pytorch/pytorch/pull/158062#issuecomment-3074342140))	2025-07-15 16:31:13 +00:00
dsashidh	05d7288e31	Fix incorrect bin edge description in histogramdd docs (#158275 ) Fixes #124435 This updates the torch.histogramdd documentation to correctly state that bins are inclusive of their left edges, not exclusive as currently written. There was a previous PR addressing this but it was closed due to inactivity. This picks that up and applies the fix. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158275 Approved by: https://github.com/albanD	2025-07-15 16:25:01 +00:00
IvanKobzarev	5a54db14e3	[simple_fsdp][inductor_collectives] rewrite reorder_collectives, sink_waits_iterative (#158062 ) Differential Revision: [D78159013](https://our.internmc.facebook.com/intern/diff/D78159013) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158062 Approved by: https://github.com/wconstab	2025-07-15 14:27:57 +00:00
Aleksandar Samardžić	90618581e9	Fix grouped MM output strides when compiled but not max-autotuned (#158143 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158143 Approved by: https://github.com/ngimel	2025-07-15 11:53:13 +00:00
Andrey Talman	4e13eca713	[BE] Remove CUDA 11.8 artifacts (#158303 ) We are including cufile by default in all CUDA 12+ builds. Since CUDA 11.8 is removed we can safely remove this code Pull Request resolved: https://github.com/pytorch/pytorch/pull/158303 Approved by: https://github.com/Camyll, https://github.com/cyyever	2025-07-15 11:52:08 +00:00
Xiangyang (Mark) Guo	156a377f4c	[AOTI][CPP] add flag TORCHINDUCTOR_CPP_FORCE_INLINE_KERNEL (#157949 ) Summary: Add flag TORCHINDUCTOR_CPP_FORCE_INLINE_KERNEL to force inline the kernel function when TORCHINDUCTOR_CPP_FORCE_INLINE_KERNEL=1. It's disabled by default because force inlining may increase the build time. Differential Revision: D77915987 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157949 Approved by: https://github.com/desertfire	2025-07-15 10:51:43 +00:00
henrylhtsang	6200584193	[cutlass backend][BE] remove force disable cache in tests (#158053 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158053 Approved by: https://github.com/coconutruben	2025-07-15 10:35:34 +00:00
Yu, Guangye	e40ade5182	Deprecate overleap functions in CUDAAllocatorConfig, use AcceleratorAllocatorConfig instead (#156165 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156165 Approved by: https://github.com/albanD ghstack dependencies: #150312	2025-07-15 10:14:35 +00:00
Yu, Guangye	e241a07e6b	Refactor CUDAAllocatorConfig to reuse AcceleratorAllocatorConfig (#150312 ) # Motivation Refactor `CUDAAllocatorConfig` to reuse `AcceleratorAllocatorConfig` and `ConfigTokenizer`. We would deprecate those option that overleap with `AcceleratorAllocatorConfig` in the following PR and keep them only for BC. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150312 Approved by: https://github.com/albanD	2025-07-15 10:14:35 +00:00
Huamin Li	7f9fc7e67c	[Inductor] Add CPU_MAX_FIRST_DIMENSION_DECOMPOSITION and CPU_MAX_OTHER_DIMENSION_DECOMPOSITION for decompose_mm_pass (#158183 ) Differential Revision: D78209993 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158183 Approved by: https://github.com/houseroad	2025-07-15 10:07:25 +00:00
FFFrog	1b389025ba	Refactor and Improve the OpenReg Module (#158090 ) ---- # Refactor and Improve the OpenReg Module ## Background Since PrivateUse1 has become the main path for integrating new devices with PyTorch, there have been some feature requests related to PrivateUse1 regarding interfaces, documentation, reference examples, etc., such as the following: - https://github.com/pytorch/pytorch/issues/155864 - https://github.com/pytorch/pytorch/issues/144955 - https://github.com/pytorch/pytorch/issues/144845 Taking these requests into consideration and combining them with the position of OpenReg, which is currently used as the test backend for PrivateUse1, I'm planning to make the following optimizations: - Optimize the implementation of OpenReg to make it align with the standard specifications for real backend (C++) access, serving as a reference for new device integration code. - Add comprehensive documentation to the [developer notes](https://docs.pytorch.org/docs/main/notes.html) to guide new accelerator integration, functioning as a reference manual. ## Design Principles: - Minimization Principle: Keep the code small and clear; only implement the minimum set of code required for verification and as an integration reference. - Authenticity Principle: Integrate OpenReg in the same way that real accelerators access PyTorch. ## More Infos: Pleaes refer to [this](`6b8020f1ab/test/cpp_extensions/open_registration_extension/torch_openreg/README.md`) for more information about `OpenReg`. ## Current Progress: - Refer to the implementation of [torch_xla](https://github.com/pytorch/xla) to refactor all of OpenReg's code, making it easier to understand. - Ensure all tests in [test/test_openreg.py](https://github.com/FFFrog/pytorch/blob/openreg/test/test_openreg.py) pass after refactoring. ## Next Steps: - Add more features to cover all integration points. - Gradually add user guides and documentation to the [developer notes](https://docs.pytorch.org/docs/main/notes.html). Pull Request resolved: https://github.com/pytorch/pytorch/pull/158090 Approved by: https://github.com/seemethere, https://github.com/albanD	2025-07-15 08:10:05 +00:00
AaronWang04	6c5227ba00	[CI] Fixes CI for CUDA Version > 12.9 (#157385 ) Compute capabilities older than volta (inclusive) is no longer supported in CUDA Version > 12.9 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157385 Approved by: https://github.com/huydhn	2025-07-15 07:04:54 +00:00
wengshiy	c8c221c0b3	[Inductor][Float8] Add float8_e4m3fn into assertion dtype list. (#157684 ) Fix assert issue. Add float8_e4m3fn into dtype list. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157684 Approved by: https://github.com/Xia-Weiwen, https://github.com/leslie-fang-intel, https://github.com/jansel	2025-07-15 06:02:01 +00:00
codingwithsurya	3341c131b7	[SymmMem] Fix NCCL Hang in NVSHMEM Triton Wait Until Test (#158167 ) The `test_triton_wait_until` test was hanging due to an NCCL synchronization issue stemming from mismatched NVSHMEM operations. Specifically, the flag variable was updated using `nvshmemx_signal_op` (a signaling operation), but waited on with `nvshmem_wait_until` (intended for put/get updates). Per NVSHMEM documentation (see documentation reference section below), signal-updated variables require `nvshmem_signal_wait_until` for proper completion guarantees, so the mismatch caused a deadlock and NCCL hang. Fix: - A simple fix was to replace the flag update with a regular `nvshmem_putmem_block` (via `put_kernel`) to match `nvshmem_wait_until`. I also added a fence (`nvshmem_fence`) between data and flag puts on the sender (Rank 1) for ordered delivery. - In a follow-up PR I will add a kernel/test to demonstrate usage of `nvshmemx_signal_op` Testing: - I ran `python test/distributed/test_nvshmem_triton.py` and `python test/distributed/test_nvshmem_triton.py -k test_triton_wait_until` - I also verified with debug prints (Sender completes puts/fence before receiver's wait returns, and assertions confirm correct state). Multiple runs show no hangs or failures. Documentation Referenced: - [NVSHMEM Point-To-Point Synchronization](https://docs.nvidia.com/nvshmem/api/gen/api/sync.html) explicitly states: "the sig_addr object at the calling PE is expected only to be updated as a signal, through the signaling operations available in Section NVSHMEM_PUT_SIGNAL and Section NVSHMEM_PUT_SIGNAL_NBI" - [NVIDIA's Official Ring Broadcast Example](https://docs.nvidia.com/nvshmem/api/examples.html) demonstrates the correct pairing: `nvshmemx_signal_op` with `nvshmem_signal_wait_until` (not `nvshmem_wait_until`) - [NVSHMEM Signaling Operations](https://docs.nvidia.com/nvshmem/api/gen/api/signal.html) documents that signal operations work on special "signal data objects" with specific atomicity guarantees distinct from regular RMA operations Pull Request resolved: https://github.com/pytorch/pytorch/pull/158167 Approved by: https://github.com/Skylion007, https://github.com/fduwjj	2025-07-15 05:57:27 +00:00
Gabriel Ferns	9cd521de4d	Fix torchrec multiprocess tests (#158159 ) Summary: The new version of `get_device_tflops` imported something from testing, which imported common_utils.py, which disabled global flags. Test Plan: Fixing existing tests Rollback Plan: Differential Revision: D78192700 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158159 Approved by: https://github.com/nipung90, https://github.com/huydhn	2025-07-15 05:44:37 +00:00
albanD	058fb1790f	Fix compilation and "import torch" issues for cpython 3.14 (#158184 ) Beginning of process for 3.14 bringup. State of things from this PR: - Nothing too scary looking from the Dynamo CPython side, nothing we heavily rely on seems to be missing @williamwen42 - The existing check that makes torch.compile() nicely fail is working as expected. So all these empty functions shouldn't cause any weirdness. - The `__module__` update changes look suspicious, we should investigate what is the reason and impact of that, in particular for our public API checking @jbschlosser - Leaving the weakref.py thread safety change as a follow up to keep this a bit simpler. I vendored the whole struct in the meantime FYI @ezyang EDIT: The `__module__` change is even more cursed than I though due to changes to Union and Optional type where the `__module__` field cannot be changed anymore. See https://github.com/python/cpython/issues/132139 for details. For now, I'm just skipping the `__module__` setting for 3.14 which will trip the public API checks. Will revisit once I have a final answer on the cpython issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158184 Approved by: https://github.com/msaroufim	2025-07-15 05:06:55 +00:00
Xilun Wu	add0b450bd	[DTensor][BE] improve DTensor ops correctness check utils (#158112 ) Summary Implemented the test pattern described in https://github.com/pytorch/pytorch/pull/157991#discussion_r2196363170 as a util method in `DTensorTestBase`. The difference to `DTensorTestBase._test_op` is: 1. allowing users to specify the `Partial` placement. 2. supporting tree-like output structure. Test so far only adopt `DTensorTestBase._test_op_on_dtensor` in `DistTensorOpsTest.test_split_on_partial`. `pytest test/distributed/tensor/test_tensor_ops.py -s -k test_split_on_partial` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158112 Approved by: https://github.com/Skylion007, https://github.com/zpcore ghstack dependencies: #158051	2025-07-15 04:50:34 +00:00
Xilun Wu	4c1fabf2c9	[DTensor] have split_strategy return OpStrategy instead of TupleStrategy (#158051 ) Summary `split_strategy` used `TupleStrategy` as return type because DTensor sharding propagation's `OpStrategy` support on multi-returns only applies to `Tuple`. However, `TupleStrategy`'s not a good fit for `split` op. `TupleStrategy` was initially introduced to handle the sharding strategy of `foreach_` ops where the input args can be split into independent subsets regarding sharding decisions, so are the outputs. To address the misuse, this PR adds `OpStrategy` propagation for `List[Tensor]` (note that this support is INCOMPLETE because it only checks the return type to be `torch.ListType`). Nevertheless, the logic for `Tuple` returns also made similar assumption so I think it's fine to unblock in such a way. Besides adding `OpStrategy` support to ops having `List[Tensor]` return type, this PR also changes `split_strategy`'s return from `TupleStrategy` to `OpStrategy`. Test* `pytest test/distributed/tensor/test_tensor_ops.py -s -k test_split_on_partial` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158051 Approved by: https://github.com/wconstab, https://github.com/zpcore	2025-07-15 04:50:34 +00:00
Ti-Tai Wang	a2ad16be72	[ONNX] Remove legacy Dort tests (#158294 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158294 Approved by: https://github.com/justinchuby ghstack dependencies: #158255, #158256, #158257	2025-07-15 04:44:14 +00:00
Ti-Tai Wang	5fb07acbc3	[ONNX] Remove legacy modularization (#158257 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158257 Approved by: https://github.com/justinchuby ghstack dependencies: #158255, #158256	2025-07-15 04:36:01 +00:00
Ti-Tai Wang	336bff6d58	[ONNX] Remove legacy graph passes (#158256 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158256 Approved by: https://github.com/justinchuby ghstack dependencies: #158255	2025-07-15 04:27:30 +00:00
Ti-Tai Wang	12151c96d9	[ONNX] Remove legacy io_adapter (#158255 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158255 Approved by: https://github.com/justinchuby	2025-07-15 03:39:18 +00:00
Will Constable	4486a6dbfd	[DTensor] Fix grouped_mm strategy for invalid stride cases (#158245 ) local_tensor input to grouped_mm has a stride requirement. (see `_meta_grouped_mm_common` in meta_registrations.py or `check_valid_strides_and_return_transposed` in native/cuda/Blas.cpp) Don't allow sharding a tensor if its shape would result in an incompatible local_tensor stride. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158245 Approved by: https://github.com/zpcore, https://github.com/XilunWu	2025-07-15 03:29:49 +00:00
Arsh Zahed	a5e68814d5	Allow dynamic shapes for DTensor slice (#157953 ) This PR allows for symints in `gen_slice_strategy` which is the strategy for `aten.slice.Tensor`. Previously, using dynamic shapes with slicing would result in ``` File ".../pytorch/torch/distributed/tensor/_ops/_tensor_ops.py", line 348, in gen_slice_strategy assert isinstance(end, int) torch._dynamo.exc.TorchRuntimeError: Dynamo failed to run FX node with fake tensors: call_function <built-in function getitem>((DTensor(local_tensor=FakeTensor(..., device='cuda:0', size=(s3, 2)), device_mesh=DeviceMesh('cuda', [0, 1]), placements=(Shard(dim=0),)), slice(None, (s77//2), None)), *{}): got AssertionError() ``` Questions before merge: 1. `dim` is still asserted to be int. Is this fine, or is this potentially dynamic as well? 2. I'm using argtype ignore for `normalize_dim`. Should I instead change types for `normalize_dim` and further dependency to be `IntLike` as well? Pull Request resolved: https://github.com/pytorch/pytorch/pull/157953 Approved by: https://github.com/wconstab	2025-07-15 00:54:01 +00:00
James Wu	ef4cca2d79	[precompile] Increment frame and add compile ids when loading packages (#158028 ) When loading a package and calling package.install(backends), we create a new frame and compile id for each package load, so that tlparse and chromium events still show compile times on warm start. There is an argument for not doing this in AOT precompile, as no "compile" occurs. So for now, we put it in `package.install`, which hopefully won't be a thing for AOT precompile. ## Recompiles Recompiles get saved to the same frame and code entry, so on warm start, each recompile will get collapsed into the same entry. Therefore, dynamo compiles that have recompiles on cold start (0/0, 0/1, 0/2, etc) will all get collapsed into a single compile id (0/0), as warm start will load all of the entries properly. ## Graph breaks Graph breaks get their own compile id, and therefore their own code entry. These are replicated on warm start, so if cold start you had 4 different graphs (and therefore 4 compile ids), you'll have 4 compile ids on warm start as well. ## Test plan Added a frame counter check to existing unit tests for automatic dynamic, showing that old and new frame counter between old and new load is the same. This is the chromium event for test_automatic_dynamo_graph_breaks_device_cuda: ``` python test/dynamo/test_package.py -k test_automatic_dynamo_graph_breaks_device_cuda ``` <img width="2216" height="508" alt="image" src="https://github.com/user-attachments/assets/f604ed33-5c31-464b-9320-d67b2e6f57a1" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158028 Approved by: https://github.com/oulgen	2025-07-15 00:53:52 +00:00
Songhao Jia	1c6057fd17	add eq function to NodeSource (#158170 ) Summary: add eq function to NodeSouce by comparing their dict representation. Test Plan: ci Rollback Plan: Differential Revision: D78200762 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158170 Approved by: https://github.com/ezyang, https://github.com/yushangdi	2025-07-15 00:50:06 +00:00
henrylhtsang	7e433d5f42	[cutlass backend] cache a few things for codegen and properties (#158158 ) Differential Revision: [D78193404](https://our.internmc.facebook.com/intern/diff/D78193404/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158158 Approved by: https://github.com/ColinPeppler	2025-07-15 00:18:31 +00:00
Tristan Rice	b7def5ff1c	dist2: add support for passing custom configs directly to PG (#158147 ) This is intended to make it easier to have backend specific "hints" that can be provided by the user to hint about certain options. ```py import torch.distributed._dist2 as dist2 pg = dist2.new_group(backend="my_custom_backend", device=..., timeout=..., foo=1234, bar="1234") pg.allreduce(...) ``` Test plan: ``` pytest test/distributed/test_dist2.py ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158147 Approved by: https://github.com/fduwjj	2025-07-15 00:02:54 +00:00
Simon Fan	7cf31b4a42	[dynamo] fix NamedTupleVariable cloning (#158190 ) FIXES https://github.com/pytorch/pytorch/issues/157945 ## Explanation 1. Some VTs add additional attrs e.g. NamedTupleVariable has "dynamic_attributes" `a0308edb6c/torch/_dynamo/variables/lists.py (L1048-L1051)` 2. VT.clone passes everything by dict, includes "dynamic_attributes" `a0308edb6c/torch/_dynamo/variables/base.py (L255-L259)` 3. Non-handled args become kwargs in VT's `__init__`, `super().__init__()` passes kwargs to Base VT `a0308edb6c/torch/_dynamo/variables/lists.py (L1048-L1051)` 4. Base VT's `__init__` gets unexpected "dynamic_attributes" kwarg `a0308edb6c/torch/_dynamo/variables/base.py (L609-L613)` You could also let Base VT's `__init__` ignore additional kwargs, but that seemed a bit too permissive, and I don't think many VT's add these derived class only attrs. ## After fix ```python ===== __compiled_fn_1_7f9541ed_e166_43fe_8322_c5225ce4207f ===== /home/xmfan/core/miniconda3/envs/0712/lib/python3.12/site-packages/torch/fx/_lazy_graph_module.py class GraphModule(torch.nn.Module): def forward(self, L_x_: "f32[4, 8, 6][48, 6, 1]cpu"): l_x_ = L_x_ # File: /home/xmfan/core/a/torchtitan/wtf.py:10 in forward, code: U, S = torch.linalg.svd(x)[:2] linalg_svd = torch._C._linalg.linalg_svd(l_x_); l_x_ = None U: "f32[4, 8, 8][64, 1, 8]cpu" = linalg_svd[0] S: "f32[4, 6][6, 1]cpu" = linalg_svd[1]; linalg_svd = None # File: /home/xmfan/core/a/torchtitan/wtf.py:11 in forward, code: reduced = U[:, :, :self.k] @ torch.diag_embed(S[:, :self.k]) getitem_3: "f32[4, 8, 5][64, 1, 8]cpu" = U[(slice(None, None, None), slice(None, None, None), slice(None, 5, None))]; U = None getitem_4: "f32[4, 5][6, 1]cpu" = S[(slice(None, None, None), slice(None, 5, None))]; S = None diag_embed: "f32[4, 5, 5][25, 5, 1]cpu" = torch.diag_embed(getitem_4); getitem_4 = None reduced: "f32[4, 8, 5][40, 5, 1]cpu" = getitem_3 @ diag_embed; getitem_3 = diag_embed = None return (reduced,) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158190 Approved by: https://github.com/StrongerXi	2025-07-14 23:39:25 +00:00
Catherine Lee	08799217ae	[CI] Move main branch rocm binary builds to its own workflow (#158161 ) Petition to move out of ciflow/trunk and into ciflow/rocm because it's a long pole for TTS <img width="1192" height="312" alt="image" src="https://github.com/user-attachments/assets/b12a097a-3763-4c62-b09f-094ee9ae1c37" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158161 Approved by: https://github.com/seemethere	2025-07-14 23:07:49 +00:00
Catherine Lee	48315181c7	[CI] Do not run inductor rocm on ciflow/inductor (#158162 ) Petition to only run inductor-rocm on ciflow/inductor-rocm and not ciflow/inductor because it's a long pole for TTS <img width="1266" height="315" alt="image" src="https://github.com/user-attachments/assets/b3587bf7-b1a6-45f3-9b6a-c0e6d473d13b" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158162 Approved by: https://github.com/seemethere	2025-07-14 23:07:45 +00:00
Eli Uriegas	38371f693b	ci: Switch lintrunner-noclang to use linter image (#158261 ) This changes the image the lintrunner jobs utilizes to be the base linter image instead of the CUDA image. This is done to reduce the image size and speed up the build time. This was switched in https://github.com/pytorch/pytorch/pull/110502 when clang used to run in the lintrunner jobs but it is now split out so we can use the default image for non-clang jobs. Difference in pull time (from running job): ~5min --> ~1min (80% reduction), this should result in an overall runtime decrease of ~25min --> ~20min (20% reduction) Signed-off-by: Eli Uriegas <eliuriegas@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158261 Approved by: https://github.com/Camyll, https://github.com/ZainRizvi, https://github.com/atalman, https://github.com/Skylion007	2025-07-14 22:54:51 +00:00
Xuan Zhang	c062550a35	[PT2][fusion] ban fusions with large accumulated reads (#157563 ) Problem: Fusion can accumulate large amount of reads, which leads to significant increase in peak memory utilization. Imagine we have the following code snippet ``` total = torch.rand(N, N) for _ in range(r): x = torch.rand(N, N) total = total + x ``` The default execution is memory efficient as only two tensors of size N-by-N is in memory at any given time. However, with fusion, the additions are fused into a single operation and the execution becomes something like: ``` x_1 = torch.rand(N, N) x_2 = torch.rand(N, N) ... x_r = torch.rand(N, N) total = x_1 + x_2 + ... + x_r ``` Though this is run-time efficient, in the case of large `N` and/or large `r`, this is not memory efficient. [internal only] see [post](https://fb.workplace.com/groups/1075192433118967/permalink/1703374333634104/) for additional details Solution: Our proposed solution is to ban fusions in case where a large amount of reads are accumulated. This is in addition to some existing logics during torch compile. * During lowering (i.e., `ir.py`), the config `realize_acc_reads_threshold`, which is default to be 8, controls _the number of_ buffers can be accumulated for a single operator. However, this is oblivious to the size of the buffers. Hence, we additionally introduce a config `realize_acc_reads_size_threshold` to control _the amount of buffers_ in size that can be accumulated. * During scheduling (i.e., `scheduler.py`), additional fusion will be performed and thus we also need to capture such pattern there. The decisions are implemented under `choices.py`. Results: For a small example similar to be one in the test case (but with larger `N` and higher number of loop repeats), the memory snapshot before and after are shown below. Note the snapshot on the right is zoomed out so that the y-axis of the two snapshots match. <img width="1328" alt="image" src="https://github.com/user-attachments/assets/670b5961-8454-4379-ae0f-62d4e7946c64" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157563 Approved by: https://github.com/jansel, https://github.com/mlazos	2025-07-14 22:27:21 +00:00
yuchengliu1	9345279c6e	skip inductor/test_torchinductor_opinfo in windows (#158225 ) During enabling inductor CI in Windows, `test_torchinductor_opinfo.py` cost too many time (about 12 hours). This UT was seriously exceeding the time limit of CI. The compiler building was slower 4x in Windows than Linux after analyzing. Thus, we decide to skip the UT temporary and @xuhancn will keep searching the solution of compiler building in Windows. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158225 Approved by: https://github.com/jansel Co-authored-by: Xu Han <xu.han@outlook.com>	2025-07-14 22:14:52 +00:00
Joona Havukainen	194539e9c3	Address NaNs if SDPA is called with all values masked from query (#157727 ) Fixes #156707 Detect if all values along the softmax axis are infs and overwrite the outputs for those computations with zeros before the final matmul. The behavior should be aligned with the CPU implementation. These types of cases where all values along the dimension in the attention mask are false leading to the undefined outputs in softmax occur with left padded batches for generation in HF transformers according to the original issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157727 Approved by: https://github.com/malfet	2025-07-14 22:09:35 +00:00
Ethan Wee	bcf50636ba	[CI] Removing --user flag from all pip install commands (#154900 ) Related to https://github.com/pytorch/pytorch/issues/148335 python virtualenv doesn't support using `--user` flag: ``` ERROR: Can not perform a '--user' install. User site-packages are not visible in this virtualenv. + python3 -m pip install --progress-bar off --user ninja==1.10.2 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154900 Approved by: https://github.com/jeffdaily Co-authored-by: Jithun Nair <jithun.nair@amd.com>	2025-07-14 21:09:42 +00:00
fduwjj	6b2bef10af	[c10d] Prototype of `group_split` for dist2 work (#157716 ) This is to implement group_split as proposed in [docs.google.com/document/d/13R-1t_yESTvmAjcCN-wQjQQadIEu0JNIdS65uZawZzY/edit?tab=t.0#heading=h.3ctbqqopzc89](https://docs.google.com/document/d/13R-1t_yESTvmAjcCN-wQjQQadIEu0JNIdS65uZawZzY/edit?tab=t.0#heading=h.3ctbqqopzc89) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157716 Approved by: https://github.com/d4l3k	2025-07-14 21:04:12 +00:00
Huy Do	1e4d8b5a4a	Fix land race typos from #157290 (#158272 ) TSIA, this is a new grammar linter being added recently. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158272 Approved by: https://github.com/clee2000	2025-07-14 20:55:13 +00:00
dolpm	725c327284	[nativert] add memory overlap debug assertion (#157290 ) Summary: better safe than sorry. will throw if memory overlap detected when using planned tensors and debug mode is enabled -- this will make our planning unit tests more robust. Test Plan: ci Rollback Plan: Differential Revision: D77327841 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157290 Approved by: https://github.com/SherlockNoMad, https://github.com/zhxchen17	2025-07-14 19:12:41 +00:00
henrylhtsang	f87d117939	redo of [Inductor][Cutlass] verify cutlass has cache_file attribute before moving...resolves cutlass cute exception (#158206 ) trying to land https://github.com/pytorch/pytorch/pull/156672 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158206 Approved by: https://github.com/lessw2020, https://github.com/Skylion007	2025-07-14 18:50:23 +00:00
Tianyu Liu	5633283574	[reland][DTensor][FSDP2] necessary changes to FSDP and TP to unblock EP (#158204 ) This PR is identical to https://github.com/pytorch/pytorch/pull/157216, which got reverted because of removing an outdated import of `torch._dynamo` https://www.internalfb.com/diff/D78021229?transaction_fbid=1713683499308113 The issue has been fixed by @weifengpy by D78199546, so this PR should be good to re-land. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158204 Approved by: https://github.com/weifengpy	2025-07-14 18:07:21 +00:00
Sudarshan Raghunathan	5b10b0a96f	Slightly improve error message from repeat_interleave kernel (#157996 ) Summary: In many investigations relating to invalid feature values, the three-argument form of `repeat_interleave` currently prints the following message if there is an inconsistency between `sum(repeats)` and `output_size`: ``` Assertion `result_size == cumsum_ptr[size - 1]` failed. ``` This is a bit hard for model authors to understand so I made the error slightly more comprehensible. After the fix the stdout contains the actual values of these parameters: https://fburl.com/mlhub/cfyyhh3q ``` Invalid input! In `repeat_interleave`, the `output_size` argument (949487) must be the same as the sum of the elements in the `repeats` tensor (949687). ``` In many cases, this is potentially useful information since we know for example that the difference between the two values above (949687-949487=200) happens to be the lengths of one of the features. ## What are my concerns with this change? 1. Outputs from `__assert_fail` go to `stderr` whereas `printf` writes to `stdout`. This is not the usual debugging flow where all logs can be found in `stderr`. I could not find a way to redirect `printf` to stderr or `__assert_fail` to stdout 2. Two checks happen instead of one in the error path. I wanted to preserve the semantics of what happens inside `__assert_fail`. 3. I have not seen this pattern in other PyTorch kernels but `repeat_interleave` with three arguments seems special in other ways too. Test Plan: * Built an ephemeral package with my changes: https://www.internalfb.com/intern/servicelab/build/736441058/ * Verified that a job with these changes indeed prints out the expected message to stdout: https://fburl.com/mlhub/jgbqk8eg * I will export to GH and run CI/CD tests. Rollback Plan: steps: - manual.note: content: >- Just reverting this diff should be sufficient. Since this change is in CUDA kernels, I do not believe there is a way to change the error message via a JK. Reviewed By: mradmila Differential Revision: D77904753 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157996 Approved by: https://github.com/ngimel, https://github.com/eqy	2025-07-14 17:55:14 +00:00
James Wu	fb462cec8d	Normalize placeholder names in AOTAutogradCache (#157916 ) This PR adds a pass to sanitize_gm_for_cache which normalizes all placeholder names across input dynamo graphs to AOTAutogradCache. This is safe because nothing underneath AOTAutograd uses the node names on the original dynamo graph: AOTAutograd re-traces with its own nodes, and guards are in terms of original sources rather than placeholder names. Note that the dynamo output graphs traced by tlparse will not show this change because it's done before this sanitization step. The aot autograd outputs also will not change because AOTAutograd's own traced graphs don't use the original placeholders of the dynamo graph. Thus, this change is essentially a no-op from everyone's perspective except for cache key checks. Fixes #157792 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157916 Approved by: https://github.com/zou3519	2025-07-14 17:45:11 +00:00
Catherine Lee	9b0013c6bb	[CI] Update mobile build docker image (#158153 ) The docker image got removed and then the job started building its own -> takes a long time I don't know why it uses the asan image <img width="1906" height="330" alt="image" src="https://github.com/user-attachments/assets/72fbf40c-3cd6-44ea-b61b-6335d2a4b589" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158153 Approved by: https://github.com/Skylion007	2025-07-14 17:35:58 +00:00
PyTorch MergeBot	6ea91f0672	Revert "[Inductor] Set the default value of min_chunk_size to 512 (#150762 )" This reverts commit 3321acc92e24859dbe2ac6499067d1afde5622c3. Reverted https://github.com/pytorch/pytorch/pull/150762 on behalf of https://github.com/huydhn due to Sorry for reverting your change, but an inductor compilation error shows up in trunk ([comment](https://github.com/pytorch/pytorch/pull/150762#issuecomment-3070286787))	2025-07-14 16:58:13 +00:00
PyTorch MergeBot	6fe7456aa1	Revert "Refactor CUDAAllocatorConfig to reuse AcceleratorAllocatorConfig (#150312 )" This reverts commit 03b307575a98dc1d953c9d3521a9489e0e61e70c. Reverted https://github.com/pytorch/pytorch/pull/150312 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it is failing to build PyTorch internally ([comment](https://github.com/pytorch/pytorch/pull/150312#issuecomment-3070218901))	2025-07-14 16:33:48 +00:00
PyTorch MergeBot	e8cca7bac7	Revert "Deprecate overleap functions in CUDAAllocatorConfig, use AcceleratorAllocatorConfig instead (#156165 )" This reverts commit 85857181ebca86e9c709e9922a9d9ef41a9c4ef9. Reverted https://github.com/pytorch/pytorch/pull/156165 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it is failing to build PyTorch internally ([comment](https://github.com/pytorch/pytorch/pull/150312#issuecomment-3070218901))	2025-07-14 16:33:48 +00:00
Guilherme Leobas	59c3cac454	Tag CPython test files with the commit or tag they were copied from. (#158038 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158038 Approved by: https://github.com/XuehaiPan, https://github.com/zou3519 ghstack dependencies: #157799, #157800, #157801, #157802, #156981	2025-07-14 15:42:19 +00:00
Ke Wen	826f12b829	[SymmMem] Avoid library mismatch in CMake search (#157836 ) Before, if NVSHMEM is installed at BOTH system location (e.g. `/usr/local`) and conda location (e.g. `/path/to/conda/lib/python3.10/site-packages/nvidia/nvshmem`, there can be a mismatch in where host lib and device lib are found: ``` -- NVSHMEM_HOME set to: '' -- NVSHMEM wheel installed at: '.conda/envs/pytorch-3.10/lib/python3.10/site-packages/nvidia/nvshmem' -- NVSHMEM_HOST_LIB: '/usr/local/lib/libnvshmem_host.so' -- NVSHMEM_DEVICE_LIB: '.conda/envs/pytorch-3.10/lib/python3.10/site-packages/nvidia/nvshmem/lib/libnvshmem_device.a' -- NVSHMEM_INCLUDE_DIR: '.conda/envs/pytorch-3.10/lib/python3.10/site-packages/nvidia/nvshmem/include' ``` The reason is that CMake prioritize name search over dir search. In the script below, CMake will search all locations for `libnvshmem_host.so` first, before it searches for `.so.3`. ``` find_library(NVSHMEM_HOST_LIB # In pip install case, the lib suffix is `.so.3` instead of `.so` NAMES nvshmem_host nvshmem_host.so.3 HINTS $ENV{NVSHMEM_HOME} ${NVSHMEM_PY_DIR} PATH_SUFFIXES lib lib64 cuda/lib cuda/lib64 lib/x64) ``` This PR adds the `NAMES_PER_DIR` flag, according to CMake's doc: > The NAMES_PER_DIR option tells this command to consider one directory at a time and search for all names in it. After this PR: ``` -- NVSHMEM_HOME set to: '' -- NVSHMEM wheel installed at: '.conda/envs/pytorch-3.10/lib/python3.10/site-packages/nvidia/nvshmem' -- NVSHMEM_HOST_LIB: '.conda/envs/pytorch-3.10/lib/python3.10/site-packages/nvidia/nvshmem/lib/libnvshmem_host.so.3' -- NVSHMEM_DEVICE_LIB: '.conda/envs/pytorch-3.10/lib/python3.10/site-packages/nvidia/nvshmem/lib/libnvshmem_device.a' -- NVSHMEM_INCLUDE_DIR: '.conda/envs/pytorch-3.10/lib/python3.10/site-packages/nvidia/nvshmem/include' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157836 Approved by: https://github.com/fegin, https://github.com/fduwjj ghstack dependencies: #157513, #157695	2025-07-14 14:13:02 +00:00
Ting Lu	86d8af6a6c	Add sm_70 to windows 12.9 build (#158126 ) Please see: https://github.com/pytorch/pytorch/issues/157517 Volta architectures will be kept for 12.8/12.9 builds for release 2.8 (12.8 win build does not need change since already including sm70) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158126 Approved by: https://github.com/Skylion007, https://github.com/atalman	2025-07-14 13:11:10 +00:00
Andrey Talman	0bb733ba23	Add cuda 12.4 build in CI (#157958 ) Fixes to https://github.com/pytorch/pytorch/issues/156747 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157958 Approved by: https://github.com/malfet, https://github.com/Skylion007	2025-07-14 13:01:16 +00:00
morrison-turnansky	0f21fa84fb	Documentation Fix: torch.empty_like memory preservation (#158050 ) updated docs for torch.empty_like to reflect view and dense memory behavior Fixes #158022 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158050 Approved by: https://github.com/ngimel, https://github.com/cyyever	2025-07-14 06:02:54 +00:00
Tom McTiernan	aa11628576	Issue warning with reference to user code rather than torch (#155112 ) Re-raising of #129959 as that was closed. Warning message before: ``` /home/admin/.local/share/hatch/env/virtual/toms-project-1/Qv9k_r_5/dev/lib/python3.10/site-packages/torch/cuda/amp/grad_scaler.py:120: UserWarning: torch.cuda.amp.GradScaler is enabled, but CUDA is not available. Disabling. ``` Warning message after: ``` /path/to/my/code:91: UserWarning: torch.cuda.amp.GradScaler is enabled, but CUDA is not available. Disabling. ``` Helps the user find where the issue stems from in their code. What do you think? (Looks like "skip_file_prefixes" is not available until Python 3.12 minimum...) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155112 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-07-14 05:24:23 +00:00
Nikita Shulga	9ca080db87	[MPS] Extend atomic operations to all int types (#158179 ) That fixes `index_put(..., accumulate=True)` for all dtypes int64 operation is not really atomic, but eventually consistent from the `index_put_accumulate` kernel point of view: i.e. by the end of the operation results in the global memory are indeed accumulation of the operands at given indices Pull Request resolved: https://github.com/pytorch/pytorch/pull/158179 Approved by: https://github.com/dcci, https://github.com/Skylion007 ghstack dependencies: #158064, #158178	2025-07-14 04:25:05 +00:00
Xinya Zhang	1ea9cde598	[ROCm] logsumexp on ROCm needs scaling back to natural base. (#156903 ) Fixes #156012 This is a temporary solution that makes context parallelism working before logsumexp behavior changes landed in AOTriton. After discussion we are not going to release AOTriton 0.10.1 to fix this due to * Even if the interface is not changed, changing the behavior of returned logsumexp tensor should still be considered as an ABI break. Such changes do not fall into the "ABI compatible" category and should be postponed to next release. * AOTriton 0.11 is scheduled to be released before end of July, which is less than five weeks Pull Request resolved: https://github.com/pytorch/pytorch/pull/156903 Approved by: https://github.com/jeffdaily, https://github.com/XilunWu	2025-07-14 02:50:36 +00:00
Michaelgathara	edb92e16ba	feat(dynamo): raise UnsupportedError for ndarray.astype(object) (#157810 ) Fixes #157720 ### What's in this PR? This PR improves the error handling in `torch.compile` for `ndarray.astype('O')` (or `object`). It now explicitly raises a `torch._dynamo.exc.Unsupported` exception with a clear explanation, instead of failing with a less intuitive error during fake tensor propagation. This is achieved by adding a check within `NumpyNdarrayVariable.call_method` for this specific `astype` pattern. A new test, `test_ndarray_astype_object_graph_break`, is also added to `test/test_numpy_interop.py` to verify this new behavior. ### Background Previously, attempting to `torch.compile` a function containing `ndarray.astype('O')` would result in a `TorchRuntimeError` wrapping a `TypeError: data type 'O' not understood`. This error message, originating deep within the tensor mechanism, was not very user-friendly and didn't clearly state why it was unsupported. This change makes the failure more explicit and provides a better user experience by giving a direct, actionable error message. Old Behavior (Error Traceback): ``` torch.dynamo.exc.TorchRuntimeError: Dynamo failed to run FX node with fake tensors: ... got TypeError("data type 'O' not understood") ``` New Behavior (Error Message): ``` torch.dynamo.exc.Unsupported: ndarray.astype(object) Explanation: ndarray.astype('O') or ndarray.astype(object) is not supported by torch.compile, as there is no equivalent to object type in torch. ``` ### Testing A new test has been added to `test_numpy_interop.py` which decorates a function containing `ndarray.astype("O")` with `torch.compile`. The test asserts that a `torch._dynamo.exc.Unsupported` exception is raised, confirming the new error handling works as expected. The test can be run with: `pytest test/test_numpy_interop.py -k test_ndarray_astype_object_graph_break` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157810 Approved by: https://github.com/jansel	2025-07-14 01:22:49 +00:00
Sun, Jiayi	3321acc92e	[Inductor] Set the default value of min_chunk_size to 512 (#150762 ) Change the default value of min_chunk_size from 4096 to 512 to allow more for loops to be parallelized. I tested the Inductor benchmark with this PR on CPU, and saw ~10% improvement in torchbench geomean speedup, and no change in huggingface/timm_models. There are about 15 torchbench models with different degrees of performance improvement, among which functorch_dp_cifar10, opacus_cifar10, hf_Reformer, and pyhpc_turbulent_kinetic_energy have more than 50% performance improvement. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150762 Approved by: https://github.com/leslie-fang-intel, https://github.com/jansel	2025-07-14 01:14:30 +00:00
Valentine233	1f57e0e04d	[CPU] Support GQA for flash attention (#157893 ) As many models require GQA, we support it in flash attention for CPU path. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157893 Approved by: https://github.com/mingfeima, https://github.com/jansel	2025-07-13 09:49:02 +00:00
Yu, Guangye	c68af9af1b	Fix XPU CI UT test_circular_dependencies (#158189 ) # Motivation fix https://github.com/pytorch/pytorch/issues/110040 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158189 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-07-13 09:30:57 +00:00
Nikita Shulga	5aee022d8b	[BE] Move repeated code into helper functions (#158178 ) Namely `index_get_offsets`, giving thread index computes offsets into input, output and indices tensors And `index_apply_indices` applies offests to either input or output tensor index Pull Request resolved: https://github.com/pytorch/pytorch/pull/158178 Approved by: https://github.com/dcci, https://github.com/Skylion007 ghstack dependencies: #158064	2025-07-12 18:24:12 +00:00
Jason Ansel	31326a9ad7	Fix typo in torch.set_float32_matmul_precision docs (#158191 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158191 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-07-12 18:23:11 +00:00
Xuehai Pan	a0308edb6c	[build] remove `wheel` from build requirements (#158027 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158027 Approved by: https://github.com/Skylion007	2025-07-12 16:45:51 +00:00
Bob Ren	9508d73307	remove allow-untyped-defs from torch/ao/nn/intrinsic/quantized/dynamic/modules/linear_relu.py (#157848 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157848 Approved by: https://github.com/Skylion007 ghstack dependencies: #157847	2025-07-12 15:42:12 +00:00
Bob Ren	066bf29334	remove allow-untyped-defs from torch/_higher_order_ops/run_const_graph.py (#157847 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157847 Approved by: https://github.com/Skylion007, https://github.com/zou3519	2025-07-12 15:42:12 +00:00
bobrenjc93	5221448574	multi-kernel matmuls based on varying hint sizes (#156628 ) The core idea is to generate multiple matmul kernels using different hints for symbolic variables, then select the most appropriate one at runtime for each unique shape we encounter. You can find some early experimentation details in these posts: https://fb.workplace.com/groups/8940092306109185/posts/9803850776399996/ https://fb.workplace.com/groups/8940092306109185/posts/9695805170537891/ https://fb.workplace.com/groups/257735836456307/posts/906589324904285/ Here’s a graph illustrating the empirically observed worst-case performance if an oracle always selected the least optimal hint for a given runtime size: ![image](https://github.com/user-attachments/assets/6d90ee06-a572-453e-9cba-03006f343301) This graph illustrates the performance of a hint size of 64 relative to the worst case. Notice that as the runtime sizes increase, the performance gradually approaches the worst case: ![image](https://github.com/user-attachments/assets/85ad49fe-165a-474c-8d03-db2e57654213) This graph shows the performance of a hint size of 4096 — very poor for small sizes, and also suboptimal for some mid-sized shapes: ![image](https://github.com/user-attachments/assets/adea1106-3bc8-40f3-97b0-20d940fb74f1) Finally, here’s the graph that motivated this PR. It illustrates the performance when selecting the best of three kernels generated with three different hints — 64, 256, and 4096: ![image](https://github.com/user-attachments/assets/a7cb0ce5-8139-48b1-b5c9-7670e75cbfce) ## How to review this PR At a high level, this extends @shunting314's multi-kernel abstraction to support varying GEMM choices driven by different hints. A few key points: 1. Unlike reduction kernels, triton template matmuls pass their grid as arguments to the kernel. This PR updates `MultiKernelCall` to support kernels with varying arguments. 2. The `V.graph.sizevars.size_hints` API is extended to accept a `hint_override`, allowing us to substitute the example input’s size hint with a custom value when generating multiple kernels. 3. The choice generation and benchmarking logic is updated to support multiple hint values. One kernel is generated per value in `torch._inductor.config.multi_kernel_hints`, and at runtime, we select the most suitable kernel for the current shape. 4. This PR does not add support for cpp wrapper codegen to keep it scoped. That will be added in the next PR. ## Results The following is a basic test that shows our basic multi kernel working where we no longer show significant variance based on the original hint size: https://gist.github.com/bobrenjc93/ba711d529e65fd65839b34799f6323ec Before ``` Hint\Runtime \| 64 \| 256 \| 4096 --------------------------------------------------- 64 \| 0.0948 \| 0.3124 \| 4.9477 256 \| 0.2243 \| 0.2256 \| 3.3880 4096 \| 0.3384 \| 0.3404 \| 3.3010 ``` After ``` Hint\Runtime \| 64 \| 256 \| 4096 --------------------------------------------------- 64 \| 0.0951 \| 0.2289 \| 3.3013 256 \| 0.0952 \| 0.2258 \| 3.4045 4096 \| 0.0957 \| 0.2231 \| 3.3146 ``` We also see an average speedup of 5.04% for the matrix of all hint/runtime pairs in [64, 4096] for every increment of 64: https://docs.google.com/spreadsheets/d/12TmYUDrAAFASGuP3POXTKPeAvQWIRzKzdrVSIb3vQkA/edit?gid=480268938#gid=480268938 ![Worst Case, multi-kernel](https://github.com/user-attachments/assets/712df23b-87e2-4d9d-95c2-cc25305ba2ed) NB: This is just the beginning and I plan on doing more investigation to see further improve on this initial result. For posterity the script used to generate that matrix is here: https://gist.github.com/bobrenjc93/c211fd0bd97fad8f46b91ad9dee76ad0 HUD benchmark runs: base: https://github.com/pytorch/pytorch/actions/runs/15889871988 head: https://github.com/pytorch/pytorch/actions/runs/15889876842 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156628 Approved by: https://github.com/jansel	2025-07-12 15:08:21 +00:00
Aastha Shah	191693ac85	adding arg values and arg types to Strobelight USDT (#155185 ) Summary: This diff makes changes to the USDT added by RihamSelim in D44636587. The "operator_start" USDT passes in the memory addresses of operator arguments and the argument types. This is so we can record argument values and types in the Strobelight GPUEvent Profiler. The previous diff records the ATEN operator, and this diff lays the groundwork to record ATEN op arguments. Test Plan: I ensured this code builds by running the example in this diff, and testing profiler changes in this diff. Reviewed By: RihamSelim Differential Revision: D75606556 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155185 Approved by: https://github.com/malfet	2025-07-12 12:00:08 +00:00
Xu Han	aacb944079	[aot inductor] fix clang-asan for consts_cpp. (#158175 ) From the perivous PR: https://github.com/pytorch/pytorch/pull/157608 , I added `format_consts_to_cpp` to build consts bytes. But it still raise clang ASAN `stack alloction`, when build large size consts. This PR: 1. add `test_aot_inductor_consts_cpp_build` to stack allocation skip list. 2. add ATTRIBUTE_NO_SANITIZE_ADDRESS to skip ASAN check, because consts array is locate in global area. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158175 Approved by: https://github.com/jansel	2025-07-12 07:14:05 +00:00
William Wen	6b84cb29f9	[dynamo] trace through torch.get_device_module (#157980 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157980 Approved by: https://github.com/anijain2305	2025-07-12 06:25:46 +00:00
Xuehai Pan	7f14b42adf	[BE][2/16] fix typos in torch/ (torch/_*/) (#156312 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156312 Approved by: https://github.com/albanD	2025-07-12 05:47:06 +00:00
PyTorch MergeBot	e90148c91d	Revert "[PT2][fusion] ban fusions with large accumulated reads (#157563 )" This reverts commit 4b9a6f7211123511e856ac8c8524bc332a741241. Reverted https://github.com/pytorch/pytorch/pull/157563 on behalf of https://github.com/huydhn due to Sorry for reverting your change, but I suspect that it might contribute to a string of OOM error in trunk ([comment](https://github.com/pytorch/pytorch/pull/157563#issuecomment-3064678929))	2025-07-12 04:52:11 +00:00
Kaichao You	a529a5daf5	[test][distributed][vllm] stabilize the p2p sharing through ipc (#158089 ) vLLM's RLHF integration `cf75cd2098/examples/offline_inference/rlhf_utils.py (L93)` depends on this hidden feature, adding the test so that PyTorch will not break it in a backward-incompatible way. The goal is to create p2p shared tensors across devices, say sharing process 0's memory on GPU 0, to process 1's memory space on GPU 1, when GPU 0 and GPU 1 can use GPU direct p2p access. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158089 Approved by: https://github.com/houseroad, https://github.com/ngimel	2025-07-12 04:41:13 +00:00
PyTorch MergeBot	e15f4248ad	Revert "[BE][2/16] fix typos in torch/ (torch/_*/) (#156312 )" This reverts commit 7a92b5119654c07d15f5c0818e6ae804b01e836c. Reverted https://github.com/pytorch/pytorch/pull/156312 on behalf of https://github.com/XuehaiPan due to landrace ([comment](https://github.com/pytorch/pytorch/pull/156312#issuecomment-3064672250))	2025-07-12 04:40:52 +00:00
Natalia Gimelshein	9056279f81	don't error out in empty_cache under mempool context (#158152 ) Now instead of erroring out on `empty_cache` call during graph capture or under mempool context, we will just silently do nothing. This used to be the behavior for mempools, cudagraphs used to error out, but it's fine to just ignore the call. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158152 Approved by: https://github.com/zou3519, https://github.com/eqy	2025-07-12 04:37:05 +00:00
Soumith Chintala	f45f6e86b9	Fix torch._numpy advanced indexing to match NumPy when indices are separated (#157676 ) Written with Claude Code. Fixes https://github.com/pytorch/pytorch/issues/157569 Fixes https://github.com/pytorch/pytorch/issues/158134 NumPy and PyTorch handle advanced indexing differently when advanced indices are separated by slices (e.g., arr[:, [0], :, 0]). PyTorch uses "outer" indexing placing result dimensions in original positions, while NumPy uses "vectorized" indexing moving advanced index dimensions to the front. This adds _numpy_style_advanced_indexing() to detect separated advanced indices and transpose results to match NumPy's dimension ordering, ensuring torch._numpy maintains compatibility with NumPy's indexing behavior. Fixes cases like: - arr[:, [0], :, 0] now returns shape (1, 5, 7) instead of (5, 1, 7) - arr[:, [0, 1], :, 0] now returns shape (2, 5, 7) instead of (5, 2, 7) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157676 Approved by: https://github.com/manuelcandales Co-authored-by: Claude <noreply@anthropic.com>	2025-07-12 04:35:04 +00:00
PyTorch MergeBot	9c189ed29a	Revert "multi-kernel matmuls based on varying hint sizes (#156628 )" This reverts commit 6c795306378c47341d58109da03371bba2bec46e. Reverted https://github.com/pytorch/pytorch/pull/156628 on behalf of https://github.com/huydhn due to Sorry for reverting your change but some ROCM jobs went crazy after this lands, so I try to see if reverting helps ([comment](https://github.com/pytorch/pytorch/pull/156628#issuecomment-3064617123))	2025-07-12 03:48:39 +00:00
Ti-Tai Wang	2eff14c445	[ONNX] Delete torch.onnx.dynamo_export (#158130 ) It's deprecated since torch==2.7. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158130 Approved by: https://github.com/justinchuby	2025-07-12 02:30:47 +00:00
Xuehai Pan	7a92b51196	[BE][2/16] fix typos in torch/ (torch/_*/) (#156312 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156312 Approved by: https://github.com/albanD	2025-07-12 01:47:22 +00:00
Michael Gathara	8b97e4dd8c	#IS157973/numpy version issue (#158036 ) Fixes #157973 `THPUtils_unpackNumberAsBool` now recognises `numpy.bool_ scalars` explicitly (using `torch::utils::is_numpy_bool`). If the object is a NumPy boolean, we retrieve its truth value via `PyObject_IsTrue` and return it, avoiding the previous failing path that attempted to treat it as an integer. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158036 Approved by: https://github.com/jansel	2025-07-12 01:36:28 +00:00
Ankita George	627ba41136	[DCP][HF] [ez]Change where sharded tensors are saved (#158069 ) Summary: Previously was saving sharded tensors to same directory as full tensors. But am realizing this doesn't make sense because on load(), you would be loading for a directory which contains both, with no way to distinguish them, so they should be in separate folders. Test Plan: ensure existing tests pass Rollback Plan: Differential Revision: D78108144 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158069 Approved by: https://github.com/teja-rao	2025-07-12 01:02:17 +00:00
Howard Huang	f4406689b8	fix MPCT destroy_pg call (#157952 ) I was seeing hangs / exceptions not raising in some cases. Only call `c10d.destroy_process_group()` for `MultiProcessContinuousTest` in the clean exit case. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157952 Approved by: https://github.com/fduwjj ghstack dependencies: #157589	2025-07-12 00:46:19 +00:00
PyTorch MergeBot	7444debaca	Revert "Fix logdet returning finite values for singular matrices on CUDA (#157910 )" This reverts commit 7d4228dbfd13d1ac8fac2c78c042dbb8314f042d. Reverted https://github.com/pytorch/pytorch/pull/157910 on behalf of https://github.com/huydhn due to Sorry for reverting your change but this seems to fail some internal tests accuracy ([comment](https://github.com/pytorch/pytorch/pull/157910#issuecomment-3064368647))	2025-07-12 00:22:51 +00:00
drisspg	8c928372b3	Make Q Indices optional (#157997 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157997 Approved by: https://github.com/BoyuanFeng, https://github.com/Chillee	2025-07-12 00:16:20 +00:00
anwang	22f3347fd9	[MTIA Aten Backend] Change relu / relu_ back to use relu kernel (#158101 ) # Context In D75803582, we migrated relu/relu_ from out-of-tree to pytorch in-tree. With that, we also changed it to use the ATen op-layer logic: https://www.internalfb.com/code/fbsource/[04ec3fcd0b09b601ae26a785e595ab960a6ba684]/fbcode/caffe2/aten/src/ATen/native/Activation.cpp?lines=512-520 To summarize: The behavior before D75803582: The Relu operator calls this code(https://fburl.com/code/pezspv40) and launches Relu kernel. The behavior after D75803582: The Relu operator uses the ATen logic, which delegates to the clamp_min operator, and no longer launch Relu kernel. ----------------- But according to my discussion with @vvk, we should keep using the Relu kernel, instead of adopting ATen logic that delegates to clamp_min, because MTIA's Relu kernel has special optimization for MTIA device. # This diff Change relu / relu_ to launch relu kernel, which is same as the original behavior before D75803582. Note: this doesn't mean to revert D75803582, because we still want to move relu/relu_ to in-tree. Differential Revision: [D78109262](https://our.internmc.facebook.com/intern/diff/D78109262/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158101 Approved by: https://github.com/albanD	2025-07-12 00:12:29 +00:00
Tristan Rice	0d77364ee3	dist2: cleanup non-option methods on PG (missing, timeouts) (#158123 ) This updates the ProcessGroup.* API to include timeouts on all non-option based overloaded methods. This also adds 2 missing ones `alltoall_base` and `barrier`. Following design in: https://docs.google.com/document/d/13R-1t_yESTvmAjcCN-wQjQQadIEu0JNIdS65uZawZzY/edit?tab=t.0#heading=h.3ctbqqopzc89 Test plan: ``` pytest test/distributed/test_dist2.py ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158123 Approved by: https://github.com/Skylion007, https://github.com/fduwjj	2025-07-12 00:06:37 +00:00
Benjamin Glass	f44a9eee47	[AOTI] Add missing ops to set of C-shim ops which can have nullptr returns (#158073 ) Most added ops are backwards ops, which have not been well-tested previously (thus why they were missed). Necessary ops were identified by manual examination of torch/_meta_registrations.py return values. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158073 Approved by: https://github.com/desertfire	2025-07-11 23:35:26 +00:00
henrylhtsang	ff7dd1776f	[cutlass backend] Global filter ops before situation based filter ops (#157866 ) The idea of this PR is that, sometimes we are filtering ops based not based on the node specific information. For example, we always filter out simt ops. So I want to group them together into a global filtering function. This can help shrink the config space as well. 20s -> 6s for instantiation 3332. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157866 Approved by: https://github.com/ColinPeppler	2025-07-11 23:13:20 +00:00
Tristan Rice	2a8795a981	[c10d] ProcessGroupGloo: support per operation timeouts (#158128 ) This updates ProcessGroupGloo to support per operation timeouts. Previously the timeouts were ignored even if they were set. * This checks if the timeout is `kUnsetTimeout` and conditionally uses the provided timeout or the default timeout from the context. * This exposes `set_timeout` as a standard method on ProcessGroup/Backend so we can test the global timeout. Test plan: ``` pytest test/distributed/test_c10d_gloo.py -v -k allreduce_timeout ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158128 Approved by: https://github.com/H-Huang, https://github.com/fduwjj	2025-07-11 23:09:50 +00:00
Sidharth	a8ec7babcf	[dynamo] expand_hints does exc() to expand graph_break_hints (#158078 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158078 Approved by: https://github.com/williamwen42	2025-07-11 22:51:28 +00:00
Nikita Shulga	beed033b6e	[MPS] Fix `index_kernel` for large tensors (#158064 ) Move `MetalShaderLibrary::bind_tensors` private method to OperatorUtils.h and extract `iter_tensor_offset` method, that returns an offset from the start of the storage associated with given tensor inside the iterator Migrated `index`, `index_put[_accumulate][_serial]` to the new paradigm that does not require additional tensor for indices nor special handling for 32 vs 64-bit offset, which resulted in almost 2x perf gain for 2000x2000 tensor, see results below before ``` [------------------------------------------------------------ -----------------------------------------------------------] \| 11x50x50 \| 11x100x100 \| 11x500x500 \| 11x1000x1000 \| 11x2000x2000 1 threads: ---------------------------------------------------------------------------------------------------------------- __getitem__ (torch.int8, torch.int64) \| 383.5 \| 379.8 \| 470.9 \| 1232.9 \| 4410.3 __getitem__ (torch.float16, torch.int64) \| 379.6 \| 354.5 \| 533.2 \| 1290.3 \| 4442.2 __getitem__ (torch.float32, torch.int64) \| 360.8 \| 338.6 \| 478.6 \| 1348.9 \| 4870.4 Times are in microseconds (us). ``` and after ``` [------------------------------------------------------------ -----------------------------------------------------------] \| 11x50x50 \| 11x100x100 \| 11x500x500 \| 11x1000x1000 \| 11x2000x2000 1 threads: ---------------------------------------------------------------------------------------------------------------- __getitem__ (torch.int8, torch.int64) \| 349.8 \| 330.5 \| 432.6 \| 764.5 \| 1961.2 __getitem__ (torch.float16, torch.int64) \| 342.5 \| 330.7 \| 434.7 \| 741.0 \| 1969.4 __getitem__ (torch.float32, torch.int64) \| 332.2 \| 326.1 \| 445.4 \| 751.3 \| 1972.6 Times are in microseconds (us). ``` While migrating also fixed index_put_accumulate for boolean types, by using compare_and_exchange trick over uint Fixes https://github.com/pytorch/pytorch/issues/153560 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158064 Approved by: https://github.com/dcci	2025-07-11 22:35:44 +00:00
Will Constable	93854e83b7	[DTensor] Rewrite doc of TupleStrategy (#158132 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158132 Approved by: https://github.com/XilunWu	2025-07-11 22:08:57 +00:00
Xuan Zhang	4b9a6f7211	[PT2][fusion] ban fusions with large accumulated reads (#157563 ) Problem: Fusion can accumulate large amount of reads, which leads to significant increase in peak memory utilization. Imagine we have the following code snippet ``` total = torch.rand(N, N) for _ in range(r): x = torch.rand(N, N) total = total + x ``` The default execution is memory efficient as only two tensors of size N-by-N is in memory at any given time. However, with fusion, the additions are fused into a single operation and the execution becomes something like: ``` x_1 = torch.rand(N, N) x_2 = torch.rand(N, N) ... x_r = torch.rand(N, N) total = x_1 + x_2 + ... + x_r ``` Though this is run-time efficient, in the case of large `N` and/or large `r`, this is not memory efficient. [internal only] see [post](https://fb.workplace.com/groups/1075192433118967/permalink/1703374333634104/) for additional details Solution: Our proposed solution is to ban fusions in case where a large amount of reads are accumulated. This is in addition to some existing logics during torch compile. * During lowering (i.e., `ir.py`), the config `realize_acc_reads_threshold`, which is default to be 8, controls _the number of_ buffers can be accumulated for a single operator. However, this is oblivious to the size of the buffers. Hence, we additionally introduce a config `realize_acc_reads_size_threshold` to control _the amount of buffers_ in size that can be accumulated. * During scheduling (i.e., `scheduler.py`), additional fusion will be performed and thus we also need to capture such pattern there. The decisions are implemented under `choices.py`. Results: For a small example similar to be one in the test case (but with larger `N` and higher number of loop repeats), the memory snapshot before and after are shown below. Note the snapshot on the right is zoomed out so that the y-axis of the two snapshots match. <img width="1328" alt="image" src="https://github.com/user-attachments/assets/670b5961-8454-4379-ae0f-62d4e7946c64" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157563 Approved by: https://github.com/jansel, https://github.com/mlazos	2025-07-11 21:07:57 +00:00
dsashidh	4ff9b7fa31	Fix diagnostic message for CUDA version mismatch in cuda.cmake (#157370 ) This PR fixes #157354 It fixes the issue in 'cmake/public/cuda.cmake' where a diagnostic message incorrectly showed an empty CUDA version when 'FindCUDA' and header-reported versions differed. The problem was caused by this line: set(${cuda_version_from_findcuda} ${CUDA_VERSION_STRING}) This incorrectly used the value of cuda_version_from_findcuda as a variable name. As a result the version string wasn't assigned and the error message omitted the version. This has been corrected to: set(cuda_version_from_findcuda ${CUDA_VERSION_STRING}) Now the diagnostic message properly displays the CUDA version reported by FindCUDA. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157370 Approved by: https://github.com/soulitzer	2025-07-11 20:58:35 +00:00
eqy	00ae620b9f	[CUDA] Allow cuDNN or flash attn in `test_activation_checkpointing` pattern match check (#153272 ) Seems more robust than maintaining a mirror of dispatch condition based on compute capability etc Pull Request resolved: https://github.com/pytorch/pytorch/pull/153272 Approved by: https://github.com/soulitzer	2025-07-11 20:58:12 +00:00
PyTorch MergeBot	702a304b07	Revert "[CUDA] Use runtime driver API for cuStreamWriteValue32 (#156097 )" This reverts commit 9a5278225fc5e7b46d54a65ae1a3f049ee49824f. Reverted https://github.com/pytorch/pytorch/pull/156097 on behalf of https://github.com/ngimel due to breaks 525 driver installs ([comment](https://github.com/pytorch/pytorch/pull/156097#issuecomment-3063742807))	2025-07-11 20:36:36 +00:00
eqy	9963845a4e	[CUDA] Support family-conditional compute capabilies in `TORCH_CUDA_ARCH_LIST` (#157999 ) Similar to arch-conditionals, such as 9.0a and 10.0a, family conditionals such as 10.0f enable features specific to a family of architectures, such as between sm100 and sm103 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157999 Approved by: https://github.com/Skylion007 Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-07-11 20:34:59 +00:00
bobrenjc93	6c79530637	multi-kernel matmuls based on varying hint sizes (#156628 ) The core idea is to generate multiple matmul kernels using different hints for symbolic variables, then select the most appropriate one at runtime for each unique shape we encounter. You can find some early experimentation details in these posts: https://fb.workplace.com/groups/8940092306109185/posts/9803850776399996/ https://fb.workplace.com/groups/8940092306109185/posts/9695805170537891/ https://fb.workplace.com/groups/257735836456307/posts/906589324904285/ Here’s a graph illustrating the empirically observed worst-case performance if an oracle always selected the least optimal hint for a given runtime size: ![image](https://github.com/user-attachments/assets/6d90ee06-a572-453e-9cba-03006f343301) This graph illustrates the performance of a hint size of 64 relative to the worst case. Notice that as the runtime sizes increase, the performance gradually approaches the worst case: ![image](https://github.com/user-attachments/assets/85ad49fe-165a-474c-8d03-db2e57654213) This graph shows the performance of a hint size of 4096 — very poor for small sizes, and also suboptimal for some mid-sized shapes: ![image](https://github.com/user-attachments/assets/adea1106-3bc8-40f3-97b0-20d940fb74f1) Finally, here’s the graph that motivated this PR. It illustrates the performance when selecting the best of three kernels generated with three different hints — 64, 256, and 4096: ![image](https://github.com/user-attachments/assets/a7cb0ce5-8139-48b1-b5c9-7670e75cbfce) ## How to review this PR At a high level, this extends @shunting314's multi-kernel abstraction to support varying GEMM choices driven by different hints. A few key points: 1. Unlike reduction kernels, triton template matmuls pass their grid as arguments to the kernel. This PR updates `MultiKernelCall` to support kernels with varying arguments. 2. The `V.graph.sizevars.size_hints` API is extended to accept a `hint_override`, allowing us to substitute the example input’s size hint with a custom value when generating multiple kernels. 3. The choice generation and benchmarking logic is updated to support multiple hint values. One kernel is generated per value in `torch._inductor.config.multi_kernel_hints`, and at runtime, we select the most suitable kernel for the current shape. 4. This PR does not add support for cpp wrapper codegen to keep it scoped. That will be added in the next PR. ## Results The following is a basic test that shows our basic multi kernel working where we no longer show significant variance based on the original hint size: https://gist.github.com/bobrenjc93/ba711d529e65fd65839b34799f6323ec Before ``` Hint\Runtime \| 64 \| 256 \| 4096 --------------------------------------------------- 64 \| 0.0948 \| 0.3124 \| 4.9477 256 \| 0.2243 \| 0.2256 \| 3.3880 4096 \| 0.3384 \| 0.3404 \| 3.3010 ``` After ``` Hint\Runtime \| 64 \| 256 \| 4096 --------------------------------------------------- 64 \| 0.0951 \| 0.2289 \| 3.3013 256 \| 0.0952 \| 0.2258 \| 3.4045 4096 \| 0.0957 \| 0.2231 \| 3.3146 ``` We also see an average speedup of 5.04% for the matrix of all hint/runtime pairs in [64, 4096] for every increment of 64: https://docs.google.com/spreadsheets/d/12TmYUDrAAFASGuP3POXTKPeAvQWIRzKzdrVSIb3vQkA/edit?gid=480268938#gid=480268938 ![Worst Case, multi-kernel](https://github.com/user-attachments/assets/712df23b-87e2-4d9d-95c2-cc25305ba2ed) NB: This is just the beginning and I plan on doing more investigation to see further improve on this initial result. For posterity the script used to generate that matrix is here: https://gist.github.com/bobrenjc93/c211fd0bd97fad8f46b91ad9dee76ad0 HUD benchmark runs: base: https://github.com/pytorch/pytorch/actions/runs/15889871988 head: https://github.com/pytorch/pytorch/actions/runs/15889876842 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156628 Approved by: https://github.com/jansel	2025-07-11 19:38:10 +00:00
Erik Ahlgren	bd364c901d	Fix serialization of nans in torch.export (#155359 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155359 Approved by: https://github.com/angelayi	2025-07-11 19:33:15 +00:00
Simon Mahns	b487003182	[PyTorch Core] MTIA supports arbitrary strides (#157883 ) Summary: Currently, on MTIA the following case will return false ``` options.device().supports_as_strided() ``` As a result, whenever moving a tensor from CPU to MTIA, strides will not be preserved ([see here](`e5edd013ab/aten/src/ATen/native/TensorConversions.cpp (L351)`)). This is a primary reason why deserializing tensors from .pt files will be contiguous. Reviewed By: egienvalue, andyanwang Differential Revision: D77843224 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157883 Approved by: https://github.com/albanD, https://github.com/andyanwang	2025-07-11 18:54:21 +00:00
cyy	b0556110e5	Remove unsafe PyTorchError constructor (#154961 ) Use libfmt in call sites of PyTorchError. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154961 Approved by: https://github.com/albanD	2025-07-11 18:22:53 +00:00
Simon Mahns	1cb0597a89	[PyTorch] Deprecate numpy serialization for MTIA (#157884 ) Summary: NumPy based tensor rebuilding from serialization has been deprecated by other backends (eg. [XLA](https://github.com/pytorch/pytorch/pull/137444)). The new flow has CPU storage being constructed with data from the file and then moved to the target backend device. Furthermore, relying on numpy for serialization will fail loudly when torch.load flips weights_only. Reviewed By: andyanwang Differential Revision: D77843238 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157884 Approved by: https://github.com/albanD	2025-07-11 17:57:33 +00:00
Simon Mahns	157683d862	[Reducer] Remove custom handling of view tensors for MTIA (#157882 ) Summary: Following implementation of the updated ATen Backend for mtia, and diffs enabling in tree view ops (D75266206, D75385411), we can remove custom logic from reducer to handle MTIA view operations. Test Plan: CI Rollback Plan: Reviewed By: egienvalue Differential Revision: D77843212 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157882 Approved by: https://github.com/albanD, https://github.com/andyanwang	2025-07-11 17:56:45 +00:00
PyTorch MergeBot	92ee5bd9f6	Revert "[DTensor][FSDP2] necessary changes to FSDP and TP to unblock EP (#157216 )" This reverts commit d75d30eeb610b164e69d0678a2e2b2dea81eec0f. Reverted https://github.com/pytorch/pytorch/pull/157216 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it turns out that the internal failure was legit ([comment](https://github.com/pytorch/pytorch/pull/157216#issuecomment-3063075001))	2025-07-11 17:07:26 +00:00
Xu Han	c4cdcda754	[aot] add format_consts_to_cpp function for further development. (#157608 ) Changes: 1. Split `format_consts_to_asm` function, which is current way to convert consts to object. 2. Add `format_consts_to_cpp` function, which would support for more compiler support, such as `msvc` and `icx`. 3. Add `config.aot_inductor.use_consts_asm_build` for `format_consts_to_asm` and `format_consts_to_cpp` control. 4. Add UT for `format_consts_to_cpp`. For `format_consts_to_cpp`, I have local tested it: Case: https://docs.pytorch.org/docs/main/torch.compiler_aot_inductor.html Run it and `cat` cpp code: <img width="674" alt="image" src="https://github.com/user-attachments/assets/d47ccf84-06d2-47f5-8a0d-9a43a9020aa3" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157608 Approved by: https://github.com/desertfire, https://github.com/jansel	2025-07-11 17:02:41 +00:00
Xilun Wu	bb3c911c2d	[DTensor] support split op on Partial placement (#157991 ) Summary To enable use case where the input DTensor to `split` op has `Partial()` placement, this PR treats `Partial()` in the same way with `Replicate()`. That means, `split` op only unshards the `Shard(dim=x)` if `x == split_dim` and keep other placement untouched. Test Added a new test because `test_dtensor_ops` doesn't test `Partial()` placement. `pytest test/distributed/tensor/test_tensor_ops.py -s -k test_split_on_partial` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157991 Approved by: https://github.com/zpcore	2025-07-11 16:19:31 +00:00
angelayi	1f1f22991d	Restore fake device (#157972 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/157972 Approved by: https://github.com/ezyang	2025-07-11 16:12:01 +00:00
Luca Wehrstedt	27c50799c1	Use new cuBLAS row-wise fp8 matmul for scaled-mm (#157905 ) Most of the work had already been done by @jeffdaily in #154680, but there was one remaining check that needed to be modified in order for `torch._scaled_mm` to use cuBLAS over CUTLASS when available. I tested this change by rebuilding PyTorch locally with CUDA 12.9 and ran `torch._scaled_mm` under the profiler, and observed that the kernel being launched is called `nvjet_qqtst_128x128_128x6_1x1_h_bz_coopA_algo2_ovscale_TNT` (where `ovscale` stands for "outer vector scaling", I believe, which is how cuBLAS calls this scaling mode). I then benchmarked the new kernels against the old CUTLASS ones on a standard 700W H100 GPU. I used the same approach as in #134781, and obtained these speed-ups: ![image](https://github.com/user-attachments/assets/43dfb816-9ccf-40c5-8b2a-571ce9cb511d) ![image](https://github.com/user-attachments/assets/be7ac6f2-e16c-479b-ad5c-f8039caba4b1) We see that the two kernels perform very closely (I'm surprised, I would have expected cuBLAS to outperform CUTLASS across the board), with some thin/skewed shapes becoming worse but some very large shapes becoming better. I guess the questions are whether we consider this a net-zero change (given that there's improvements _and_ degradations), and how large we consider the burden of maintaining our own CUTLASS kernels. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157905 Approved by: https://github.com/eqy, https://github.com/Skylion007, https://github.com/drisspg	2025-07-11 16:11:55 +00:00
Eddie Yan	0797b2b6a8	[cuDNN][SDPA] cuDNN SDPA refactor/cleanup, nested tensor backward, test priority bump for `sm90`, `sm100` (#149282 ) cleanup tuple/tensor boilerplate in cuDNN SDPA, preparation for nested/ragged tensor backward Pull Request resolved: https://github.com/pytorch/pytorch/pull/149282 Approved by: https://github.com/drisspg Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-07-11 16:07:54 +00:00
Aaron Gokaslan	7a08755c5f	[BE][Ez]: Update ruff to 0.12.2 (#157937 ) Updates to the latest version of ruff and apply some fixes that it flagged and silence a few new lints Pull Request resolved: https://github.com/pytorch/pytorch/pull/157937 Approved by: https://github.com/ezyang	2025-07-11 15:16:20 +00:00
Xuehai Pan	0d17029fea	[BE][6/6] fix typos in test/ (test/distributed/) (#157640 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157640 Approved by: https://github.com/yewentao256, https://github.com/malfet	2025-07-11 14:09:37 +00:00
Xuehai Pan	4283d96bcd	[build] pin `setuptools>=70.1.0` for integrated `bdist_wheel` command (#157783 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157783 Approved by: https://github.com/Skylion007	2025-07-11 12:10:42 +00:00
Victor Komarov	b4476ca378	Add cudaMallocAsync/cudaFreeAsync to cuda_to_hip_mappings (#158056 ) Summary: Adding both functions as they're required for Hipification of https://fburl.com/code/165r7qhr Test Plan: Tested in D78090513 Rollback Plan: Reviewed By: malfet, jiangyurong609 Differential Revision: D78090693 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158056 Approved by: https://github.com/Skylion007	2025-07-11 11:48:19 +00:00
Yu, Guangye	85857181eb	Deprecate overleap functions in CUDAAllocatorConfig, use AcceleratorAllocatorConfig instead (#156165 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156165 Approved by: https://github.com/albanD ghstack dependencies: #149601, #157908, #150312	2025-07-11 11:41:34 +00:00
Yu, Guangye	03b307575a	Refactor CUDAAllocatorConfig to reuse AcceleratorAllocatorConfig (#150312 ) # Motivation Refactor `CUDAAllocatorConfig` to reuse `AcceleratorAllocatorConfig` and `ConfigTokenizer`. We would deprecate those option that overleap with `AcceleratorAllocatorConfig` in the following PR and keep them only for BC. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150312 Approved by: https://github.com/albanD ghstack dependencies: #149601, #157908	2025-07-11 11:25:43 +00:00
Daisy Deng	8088958793	port 4 dynamo test files to Intel GPU (#157779 ) For https://github.com/pytorch/pytorch/issues/114850, we will port test cases to Intel GPU. Six dynamo test files were ported in PR [#156056](https://github.com/pytorch/pytorch/pull/156056) and [#156575](https://github.com/pytorch/pytorch/pull/156575.) In this PR we will port 4 more dynamo test files. We could enable Intel GPU with following methods and try the best to keep the original code styles: - instantiate_device_type_tests() - use "torch.accelerator.current_accelerator()" to determine the accelerator backend - added XPU support in decorators like @requires_gpu - enabled XPU for some test path - added xfailIfXPU to skip xpu test when there is a bug. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157779 Approved by: https://github.com/guangyey, https://github.com/jansel	2025-07-11 10:11:49 +00:00
Xia, Weiwen	e1a20988f3	[Quant][CPU] Enable fp8 qconv (#157076 ) Summary Enable fp8 qconv on CPU. It's part of the plan to enable fp8 static quantization on CPU. This PR only adds FP8 support of the existing int8 qconv op. It does not add a new op nor does it affect frontend or quantization flow. The schema of the qconv op is not changed either. So, the FP8 qconv shares the same op as INT8 qconv and the difference is that src/wei dtype is fp8 instead of int8. The output dtype can be fp8/float32/bfloat16. The implementation uses the oneDNN library. Note: OneDNN does not support quantized fp8 convolution until v3.9 but the version used in PyTorch is v3.7.2. So, the op goes to the reference kernel for now. And we have also update the oneDNN path so that it's compatible with the fp8 dtype. Once oneDNN is upgraded to v3.9 or newer, minimum changes are needed to enable the oneDNN path. And we have ensured that the behavior of the reference kernel is the same as the new oneDNN's implementation. - oneDNN version < 3.9 (now) - Always go to the reference kernel - oneDNN version >= 3.9 (future) - Go to reference kernel on old platforms (without AMX) - Use oneDNN on new platforms (with AMX) Test plan ``` pytest test/quantization/core/test_quantized_op.py -k "qconv and fp8" ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157076 Approved by: https://github.com/leslie-fang-intel, https://github.com/jerryzh168	2025-07-11 10:00:57 +00:00
Mwiza Kunda	ed508cc018	[inductor][triton] Add experimental use_tensor_descriptor config option (#157906 ) Refactor to allow TMA descriptors to be used in general codegen. TMA descriptors can only be generated if the conditions listed in the triton documentation for [make_tensor_descriptor](https://triton-lang.org/main/python-api/generated/triton.language.make_tensor_descriptor.html) are met. Some implementation details: - The `TMACompatibilityChecker` class holds and checks the conditions required for a load / store operation to be represented by a tma descriptor load / store - The current TMA API requires that the innermost block size loads atleast 16 bytes of data. e.g. if the block shape is [YBLOCK, XBLOCK] and the tensor dtype is float32, this requires that XBLOCK >= 4. It is therefore required that the triton heuristics are aware of the minimum block sizes for the IO operations in the kernel. The minimum block sizes are determined in the `TMACompatibilityChecker` class and are passed to the triton heuristics when the block sizes are not static. The heuristic config options are then filtered to ensure that the minimum block size restriction is met. Testing: - Refactored test_torchinductor_strided_blocks.py to also test the `use_tensor_descriptor` option. This requires an upgrade to Triton version 3.4.0: https://github.com/pytorch/pytorch/issues/154206 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157906 Approved by: https://github.com/jansel	2025-07-11 09:32:40 +00:00
charlifu	02724b5f64	[Bugfix][Inductor] Fix dependency list merged incorrectly for a custom op with multiple mutated inputs and None return type. (#157133 ) This is an attempt to fix a memory allocation issue when using `torch.compile` with a custom layernorm kernel in vllm: ```C++ // In-place fused Add and RMS Normalization. ops.def( "fused_add_rms_norm(Tensor! input, Tensor! residual, Tensor weight, " "float epsilon) -> ()"); ops.impl("fused_add_rms_norm", torch::kCUDA, &fused_add_rms_norm); ``` We observed abnormal extra memory allocations with this op enabled using `torch.compile`: <img width="738" alt="{374E9FCF-FB46-4750-8B60-D31E3ADCE00A}" src="https://github.com/user-attachments/assets/6c45e1aa-ccde-4c56-99dc-bf4776d699d5" /> and without this op: <img width="738" alt="{9BB08EFE-FFE3-4D06-82C0-C70BBE6ADD56}" src="https://github.com/user-attachments/assets/56e2ee43-ab87-492d-834c-69e9cafbb0df" /> After investigation, we found that this is because the compiler considers the two buffers for the two mutated inputs `Tensor input` and `Tensor residual` should share a same dependency list, which makes it can not reuse the buffer of `Tensor input`. ``` buf1.users = [ NodeUser(node=ExternKernelSchedulerNode(name='op2'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op9'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op13'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op20'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op24'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op31'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op35'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op42'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op46'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op53'), can_inplace=False, is_weak=False), ] buf16.users = [ NodeUser(node=ExternKernelSchedulerNode(name='op2'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op9'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op13'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op20'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op24'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op31'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op35'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op42'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op46'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op53'), can_inplace=False, is_weak=False), ] ``` ``` op13: ExternKernelSchedulerNode(FallbackKernel) op13.writes = [ StarDep(name='buf17', mode=None), StarDep(name='buf18', mode=None), StarDep(name='buf19', mode=None)] op13.unmet_dependencies = [ StarDep(name='buf13', mode=None), StarDep(name='buf16', mode=None), WeakDep(name='buf11', mutating_buf='buf18'), WeakDep(name='buf12', mutating_buf='buf18'), WeakDep(name='buf13', mutating_buf='buf18'), WeakDep(name='buf2', mutating_buf='buf18'), WeakDep(name='buf3', mutating_buf='buf18')] op13.met_dependencies = [StarDep(name='arg11_1', mode=None)] op13.outputs = [ buf17: FallbackKernel buf17.layout = NoneLayout(device=device(type='cuda', index=0), size=[0], stride=[0]) buf17.aliases = ['buf16', 'buf1'] buf17.users = [ NodeUser(node=ExternKernelSchedulerNode(name='op2'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op9'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op13'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op20'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op24'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op31'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op35'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op42'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op46'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op53'), can_inplace=False, is_weak=False), ] buf18: MutationOutput buf18.layout = NoneLayout(device=device(type='cuda', index=0), size=[0], stride=[0]) buf18.mutations = ['buf16'] buf18.users = [ NodeUser(node=ExternKernelSchedulerNode(name='op14'), can_inplace=False, is_weak=False), NodeUser(node=ExternKernelSchedulerNode(name='op20'), can_inplace=False, is_weak=True), NodeUser(node=ExternKernelSchedulerNode(name='op24'), can_inplace=False, is_weak=True), NodeUser(node=ExternKernelSchedulerNode(name='op31'), can_inplace=False, is_weak=True), NodeUser(node=ExternKernelSchedulerNode(name='op35'), can_inplace=False, is_weak=True), NodeUser(node=ExternKernelSchedulerNode(name='op42'), can_inplace=False, is_weak=True), NodeUser(node=ExternKernelSchedulerNode(name='op46'), can_inplace=False, is_weak=True), NodeUser(node=ExternKernelSchedulerNode(name='op53'), can_inplace=False, is_weak=True), ] buf19: MutationOutput buf19.layout = NoneLayout(device=device(type='cuda', index=0), size=[0], stride=[0]) buf19.mutations = ['buf1'] buf19.users = [NodeUser(node=ExternKernelSchedulerNode(name='op20'), can_inplace=False, is_weak=False)] ] op13.node.kernel = torch.ops._C.fused_add_rms_norm.default ``` Here we can see `buf16` shares the same dependency list with `buf1` because `buf16` and `buf1` are in the aliases list of `buf17`. This is incorrect since those two are two separate tensors. And this makes the compiler could not reuse `buf16` for subsequent ops. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157133 Approved by: https://github.com/jansel	2025-07-11 09:06:31 +00:00
Menglu Yu	44303caabf	[APS] Expose max_autotune lookup table config to frontend (#158070 ) Summary: As titled. We reuse optimus config to receive the yaml config file from users Test Plan: ### how to enable max_autotune lookup table hardcode config ``` inductor.config.post_grad_fusion_options = { "inductor_autotune_lookup_table": <your yaml manifold path> } ``` for example, "manifold://ads_training_p9e/tree/max_autotune/mast_omnifm_v3_1kgpu/mast_omnifm_v3_lookup_table.yaml", see D78052050 Rollback Plan: Reviewed By: PaulZhang12, jackiexu1992 Differential Revision: D77202285 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158070 Approved by: https://github.com/Mingming-Ding	2025-07-11 09:02:52 +00:00
Shivam Raikundalia	11d6ad8b2e	[Docs] Update PT2 Profiler Torch-Compiled Region Image (#158066 ) Summary: In Pytorch 2.5 we added source code attribution to PT2 traces. Each Torch-Compiled Region will now have its frame id and frame compile id associated with it. Update the image in the doc and add a description of this in the doc itself Test Plan: {F1980179183} Rollback Plan: Differential Revision: D78118228 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158066 Approved by: https://github.com/aaronenyeshi	2025-07-11 07:56:45 +00:00
Dmitry Rogozhkin	cd80f9a4c3	xpu: support custom ops with torch.library on xpu backend (#152879 ) Fixes: https://github.com/intel/torch-xpu-ops/issues/1626 This PR started enabling of tests for `torch.library`, but more work is needed. Tests are using `torch._custom_ops` deprecated API planned for removal at pytorch 2.6 (not done). I think cleanup of pytorch would be nice before enabling more tests for xpu. `a2ccda3c60/torch/_custom_op/impl.py (L47)` CC: @EikanWang Pull Request resolved: https://github.com/pytorch/pytorch/pull/152879 Approved by: https://github.com/EikanWang, https://github.com/malfet, https://github.com/guangyey, https://github.com/albanD	2025-07-11 07:36:04 +00:00
Yu, Guangye	442aca44d6	Fix XPU broken CI (#158092 ) # Motivation https://github.com/pytorch/pytorch/pull/157739 introduces the new UT `test_sdpfa` that block XPU CI since `_scaled_dot_product_flash_attention is not supported on XPU yet`. # Additional Context See https://github.com/pytorch/pytorch/actions/runs/16201010860/job/45741815895?pr=138222#step:15:6399 fix https://github.com/pytorch/pytorch/issues/158095 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158092 Approved by: https://github.com/jansel, https://github.com/malfet	2025-07-11 07:23:27 +00:00
Kurt Mohler	d89f30ad45	[MPS] Avoid calling tensor ops in max_pool3d impl (#157874 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157874 Approved by: https://github.com/malfet	2025-07-11 06:47:29 +00:00
zeshengzong	b4fc42ca80	Add `torch.segment_reduce` docs (#154352 ) Fixes #153138 ## Test Result ![image](https://github.com/user-attachments/assets/62346d62-d048-4259-906b-f8261e10b4cc) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154352 Approved by: https://github.com/albanD	2025-07-11 06:16:38 +00:00
zpcore	cec59b76ca	[2/N] cost coverage improvment (#157738 ) Part of plan https://github.com/pytorch/pytorch/issues/157495. Details: 1. Fill in missing redistribute_cost in `cat` and `slice_scatter`; 2. Expand the `cat` strategy based on placement of each input tensor. Previously `cat` only outputs one strategy. Now it output at the level of number_of_input_tensor*number_OpSpec_each_tensor_input_strategy. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157738 Approved by: https://github.com/wconstab	2025-07-11 05:54:16 +00:00
PyTorch MergeBot	ecd73c58ee	Revert "[BE] Replace `std::runtime_error` with `TORCH_CHECK` [2/N] (#152080 )" This reverts commit b85f10ea5006e8ae8fc769f48659ab7ad5eafb69. Reverted https://github.com/pytorch/pytorch/pull/152080 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it is failing some internal tests ([comment](https://github.com/pytorch/pytorch/pull/152080#issuecomment-3060337857))	2025-07-11 03:58:31 +00:00
Boyuan Feng	94995eba07	[Log] add a hook for recompile user context (#157961 ) Users may want compile-related but customized logging info to dynamo_compile. One example is to logging the current training iteration index when recompilation happens. In general, current training iteration index is not available to compiler, since the same compiled function may be called multiple times in the same training iteration. The user could provide the training iteration index in a user hook where torch.compile logs it when recompilation happens. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157961 Approved by: https://github.com/masnesral	2025-07-11 03:41:33 +00:00
Jerry Zhang	11a86ad2fa	Remove pytorch quant docs since we are moving to torchao (#157766 ) Summary: att Test Plan: doc page generated from CI Reviewers: Subscribers: Tasks: Tags: Pull Request resolved: https://github.com/pytorch/pytorch/pull/157766 Approved by: https://github.com/Skylion007	2025-07-11 03:21:47 +00:00
Jordan Fix	dd93883231	[exported_program] Remove _postprocess_graph_module_outputs (#158059 ) Summary: Appears to be dead as of https://github.com/pytorch/pytorch/pull/120019. Test Plan: CI Rollback Plan: Differential Revision: D78112302 Pull Request resolved: https://github.com/pytorch/pytorch/pull/158059 Approved by: https://github.com/angelayi	2025-07-11 02:40:15 +00:00
Bin Bao	326e751d07	[AOTI] Add device guard when launching autotune kernels (#158034 ) Summary: Fix https://github.com/pytorch/pytorch/issues/157737. When launching Triton kernels in the autotune block, we need to consider the fact that the model may not always be on device 0. The reason this was not caught on CI is because test_on_gpu_device1 requires multi_gpu and was not run on a multi_gpu instance. Added test_on_gpu_device1 and other similar multi_gpu tests back. Pull Request resolved: https://github.com/pytorch/pytorch/pull/158034 Approved by: https://github.com/eqy, https://github.com/yushangdi	2025-07-11 02:34:31 +00:00
Soumith Chintala	7d4228dbfd	Fix logdet returning finite values for singular matrices on CUDA (#157910 ) Fixes https://github.com/pytorch/pytorch/issues/154312 Fix logdet returning finite values for singular matrices on CUDA (https://github.com/pytorch/pytorch/issues/154312 https://github.com/pytorch/pytorch/issues/154312) PyTorch's logdet function returns mathematically incorrect finite values for singular matrices on CUDA devices instead of the expected -inf. This occurs because cuSOLVER and LAPACK produce tiny non-zero diagonal elements (~1e-16) instead of exact zeros for singular matrices. Problem: Issue https://github.com/pytorch/pytorch/issues/154312 matrix returns finite values instead of -inf for singular matrices. Solution: Implemented NumPy-style two-tier singularity detection with GPU sync point removal: 1. Primary detection: Use LAPACK's built-in singularity detection via info parameter 2. Backup detection: Apply threshold-based detection for numerical edge cases 3. Zero GPU sync points: Eliminated all .item(), std::get<0>(), and scalar extractions 4. Pure tensor operations: All computations use tensor operations throughout Performance Impact: Based on comprehensive benchmarking across matrix sizes and data types: - Overall Impact: 0.85× average speedup (+18.0% overhead) - CPU Performance: 0.84× average speedup (+18.8% overhead) - CUDA Performance: 0.85× average speedup (+17.3% overhead) Performance Trade-offs: - Small matrices (16×16, 64×64): Higher overhead due to tensor operation setup costs - Large matrices (512×512, 2048×2048): Near-zero overhead, with some cases showing slight improvements - GPU sync elimination: Removes expensive GPU→CPU synchronization bottlenecks Results: - ✅ All singular matrices now correctly return -inf on both CPU and CUDA - ✅ Original issue https://github.com/pytorch/pytorch/issues/154312 matrix now works correctly - ✅ Results match NumPy's slogdet behavior exactly - ✅ Zero GPU synchronization points for improved performance - ✅ Comprehensive edge case testing added Verification: Before: torch.linalg.slogdet(singular_matrix) → finite values (incorrect) After: torch.linalg.slogdet(singular_matrix) → (sign=0, logabsdet=-inf) ✅ The implementation uses pure tensor operations to eliminate GPU sync points while maintaining robust singularity detection through a two-tier approach. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157910 Approved by: https://github.com/lezcano, https://github.com/IvanYashchuk, https://github.com/albanD Co-authored-by: Claude <noreply@anthropic.com>	2025-07-11 02:23:46 +00:00
Yu, Guangye	65fcca4f8c	Enable AcceleratorAllocatorConfig key check (#157908 ) # Motivation Add a mechanism to ensure raise the key if the key is unrecognized in allocator config. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157908 Approved by: https://github.com/albanD ghstack dependencies: #149601	2025-07-11 02:11:08 +00:00
Paul Zhang	905b084690	Add size_hints to cache key (#158026 ) Differential Revision: D78089705 Previously to support overriding autotune configs for post fusion kernels in Inductor with a lookup table, we only keyed on the source code. However, the same source code could have multiple optimal configs, due to the input sizes. With this, we have many collisions in our lookup table, leading to subpar configs. A way around this is to add the size_hints to the lookup key as well Pull Request resolved: https://github.com/pytorch/pytorch/pull/158026 Approved by: https://github.com/jansel	2025-07-11 01:47:50 +00:00
Siddartha Pothapragada	37ccc532f7	Update'unit_batch_dynamic_prepacked' tests to use ASSERT_NEAR instead of ASSERT_EQ (#157860 ) (#157861 ) Summary: Replaced ASSERT_FLOAT_EQ which defaults to fixed kMaxUlps ( = 4-ULP , See gtest-internal.h) with ASSERT_NEAR which lets us set epsilon to 1e-3, (approximately 3 ULPs). This allows for slightly stricter and tunable comparison. Test Plan: Before Fix ✗ Fail: qnnpack:pytorch_qnnpack_testApple - FULLY_CONNECTED_SPARSE_OP_8x1/unit_batch_dynamic_prepacked (0.0s) 'Expected equality of these values: output_dynamic[i * outputChannels() + c] Which is: 9.9160004 accumulators_float[i * outputChannels() + c] Which is: 9.9159956 at 0, 17: reference = 9.9159955978393555, optimized = 9.9160003662109375 ------------------------------ After Fix Everything passes Rollback Plan: Differential Revision: D77911682 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157861 Approved by: https://github.com/kimishpatel, https://github.com/lucylq, https://github.com/malfet	2025-07-11 01:05:50 +00:00
Guilherme Leobas	7599bebead	Add CPython test `test_itertools` (#156981 ) Test the itertools module Pull Request resolved: https://github.com/pytorch/pytorch/pull/156981 Approved by: https://github.com/zou3519 ghstack dependencies: #157799, #157800, #157801, #157802	2025-07-11 00:12:50 +00:00
Guilherme Leobas	397ca98510	Add CPython test `test_with` (#157802 ) Test with statement behavior and dunder methods __enter__ and __exit__ Pull Request resolved: https://github.com/pytorch/pytorch/pull/157802 Approved by: https://github.com/zou3519 ghstack dependencies: #157799, #157800, #157801	2025-07-11 00:12:50 +00:00
Guilherme Leobas	4809f43867	Add CPython test `test_numeric_tower` (#157801 ) Test abstract numeric types and dunder methods like __int__, __float__, __index__, etc. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157801 Approved by: https://github.com/zou3519 ghstack dependencies: #157799, #157800	2025-07-11 00:12:50 +00:00
Guilherme Leobas	0ebf2447da	Add CPython test `test_operator` (#157800 ) Test operators via operator module like add, sub, eq, lt, etc. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157800 Approved by: https://github.com/zou3519 ghstack dependencies: #157799	2025-07-11 00:12:50 +00:00
Guilherme Leobas	91041f559d	Add CPython test `test_bool` (#157799 ) Test dunder methods `__bool__` and `__len__` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157799 Approved by: https://github.com/zou3519, https://github.com/XuehaiPan	2025-07-11 00:12:50 +00:00
zpcore	ae86e8f6c8	[1/N] cost coverage improvment (#157504 ) Part of plan https://github.com/pytorch/pytorch/issues/157495. Details: 1. Fill missing redistribute_cost for ops like `aten::detach`, `aten::bernoulli `, `aten::_to_copy`, `aten::bucketize.Tensor`, `aten::stack`, `aten::clone`, `aten::copy_`, `aten::zero_ `. 2. Fix redistribute_cost error in new_factory_strategy. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157504 Approved by: https://github.com/wconstab	2025-07-10 23:55:45 +00:00
Max Podkorytov	8b68e5b1bb	[ROCm][Inductor][CK] update API for gemm-multiD change (#156122 ) Fixes for the compilation errors in the generated code Pull Request resolved: https://github.com/pytorch/pytorch/pull/156122 Approved by: https://github.com/chenyang78	2025-07-10 23:12:20 +00:00
PyTorch MergeBot	e517066f41	Revert "[dynamo][fsdp] Consistent behavior of int attributes (#157262 )" This reverts commit 178fe7aa98987111a73534375099f4ad255e8b59. Reverted https://github.com/pytorch/pytorch/pull/157262 on behalf of https://github.com/huydhn due to This fails some internal tests and needs to be relanded ([comment](https://github.com/pytorch/pytorch/pull/157262#issuecomment-3059463896))	2025-07-10 23:11:18 +00:00
Edward Z. Yang	1a195bf7d6	Tests for #158030 (#158033 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158033 Approved by: https://github.com/bdhirsh, https://github.com/albanD ghstack dependencies: #158030	2025-07-10 22:51:28 +00:00
Guilherme Leobas	bfcababbcb	[OrderedDict] Implement explicit OrderedDict dunder method call (#154943 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154943 Approved by: https://github.com/zou3519 ghstack dependencies: #154003, #154793, #154794, #154942	2025-07-10 22:50:39 +00:00
Guilherme Leobas	ba71eb496b	[dict] Implement dict.__eq__ and dict.__ne__ (#154942 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154942 Approved by: https://github.com/zou3519 ghstack dependencies: #154003, #154793, #154794	2025-07-10 22:50:39 +00:00
Guilherme Leobas	ba8d19ec02	[dict] Allow Dynamo to trace through explicit dict dunder method call (#154794 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154794 Approved by: https://github.com/mlazos ghstack dependencies: #154003, #154793	2025-07-10 22:50:39 +00:00
Guilherme Leobas	57d64298a0	[dict] Add dict.popitem (#154793 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154793 Approved by: https://github.com/mlazos, https://github.com/zou3519 ghstack dependencies: #154003	2025-07-10 22:50:39 +00:00
Guilherme Leobas	e84710d1e7	[dict] Raise TypeError in dict methods (#154003 ) Raise TypeError in the following scenarios: * #args mismatch * arg is unhashable Pull Request resolved: https://github.com/pytorch/pytorch/pull/154003 Approved by: https://github.com/mlazos, https://github.com/zou3519	2025-07-10 22:50:39 +00:00
George Wigley	9bf41633d7	Allow Custom Time Unit When Printing Profiler Table (#157913 ) ## Overview This PR adds a kwarg to the `table()` method of the profiler allowing users to specify a time unit to be used for all results in the profiling table. The available options are: `s`, `ms` and `us`. If an invalid unit or no unit is provided, then a time unit is selected based on the size of the value (current default behaviour). ## Testing A unit test has been added to verify this works correctly. ## Documentation I couldn't find any documentation specific to the `table()` function beyond doc strings which have been updated. ## Example Output ``` import torch from torch.profiler import profile with profile() as prof: res = torch.mm(torch.rand(1024, 1024), torch.rand(1024, 1024)) print(prof.key_averages().table(time_unit="s")) print(prof.key_averages().table(time_unit="ms")) print(prof.key_averages().table(time_unit="us")) print(prof.key_averages().table()) ``` ``` ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg # of Calls ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ aten::rand 0.04% 0.000s 10.36% 0.014s 0.007s 2 aten::empty 0.04% 0.000s 0.04% 0.000s 0.000s 2 aten::uniform_ 10.27% 0.014s 10.27% 0.014s 0.007s 2 aten::mm 89.64% 0.119s 89.64% 0.119s 0.119s 1 aten::resolve_conj 0.00% 0.000s 0.00% 0.000s 0.000s 3 ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ Self CPU time total: 0.133s ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg # of Calls ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ aten::rand 0.04% 0.055ms 10.36% 13.735ms 6.868ms 2 aten::empty 0.04% 0.054ms 0.04% 0.054ms 0.027ms 2 aten::uniform_ 10.27% 13.626ms 10.27% 13.626ms 6.813ms 2 aten::mm 89.64% 118.892ms 89.64% 118.896ms 118.896ms 1 aten::resolve_conj 0.00% 0.004ms 0.00% 0.004ms 0.001ms 3 ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ Self CPU time total: 132.631ms ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg # of Calls ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ aten::rand 0.04% 55.495us 10.36% 13735.202us 6867.601us 2 aten::empty 0.04% 54.121us 0.04% 54.121us 27.061us 2 aten::uniform_ 10.27% 13625.586us 10.27% 13625.586us 6812.793us 2 aten::mm 89.64% 118892.284us 89.64% 118895.981us 118895.981us 1 aten::resolve_conj 0.00% 3.697us 0.00% 3.697us 1.232us 3 ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ Self CPU time total: 132631.183us ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg # of Calls ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ aten::rand 0.04% 55.495us 10.36% 13.735ms 6.868ms 2 aten::empty 0.04% 54.121us 0.04% 54.121us 27.061us 2 aten::uniform_ 10.27% 13.626ms 10.27% 13.626ms 6.813ms 2 aten::mm 89.64% 118.892ms 89.64% 118.896ms 118.896ms 1 aten::resolve_conj 0.00% 3.697us 0.00% 3.697us 1.232us 3 ---------------------- ------------ ------------ ------------ ------------ ------------ ------------ Self CPU time total: 132.631ms ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157913 Approved by: https://github.com/sraikund16	2025-07-10 22:44:34 +00:00
Tristan Rice	83700b4488	dist2: add group context manager (#157988 ) This adds new context manager based PG management to dist2. This allows for managing the active process group much in the same way as a stream ```py with dist2.process_group(pg): dist2.current_process_group().allreduce(...).wait() ``` matches ```py with torch.cuda.stream(stream): torch.cuda.current_stream().synchronize() ``` Test plan: ``` pytest test/distributed/test_dist2.py -k context ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157988 Approved by: https://github.com/fduwjj	2025-07-10 22:30:19 +00:00
Soumith Chintala	fca7013f85	Fix DCE eliminating random operations by improving is_impure() (#151524 ) (#157981 ) DCE was incorrectly eliminating unused random operations like torch.rand() that have global RNG side effects, causing inconsistent results between eager and compiled execution modes. Root cause: Python random functions (torch.rand, torch.randn, etc.) don't have the _nondeterministic_seeded attribute, so node.is_impure() returns False, allowing DCE to eliminate them despite advancing global RNG state. Solution: Enhanced is_impure() in torch/fx/node.py to recognize Python random functions and mark them as impure when they use global RNG, regardless of the impure_random parameter setting. This ensures consistency between eager and compiled execution even when config.fallback_random=False. Key features: - Handles comprehensive list of random functions: rand, randn, randint, randperm, rand_like, randn_like, randint_like, normal, poisson, bernoulli, multinomial - Generator optimization: Only marks as impure when using global RNG (no generator or generator=None). Operations with explicit generators don't affect global state and can be optimized. - Works with both impure_random=True and impure_random=False cases - Cleaner architecture: addresses root cause rather than working around it Tests: Enhanced test_impure_random to verify both FX tracing and AOT compilation codepaths, ensuring random operations are preserved and eager/compiled execution consistency is maintained. 🤖 Generated with [Claude Code](https://claude.ai/code) Fixes https://github.com/pytorch/pytorch/issues/151524 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157981 Approved by: https://github.com/mlazos Co-authored-by: Claude <noreply@anthropic.com>	2025-07-10 22:24:29 +00:00
Eddie Yan	590607c599	[cuDNN][SDPA] Bump cuDNN frontend submodule version to 1.12.1 (#158044 ) Really we are just interested in this change which fixes an apparent regression for d=256 support on Hopper `bc5f4fd88d` Pull Request resolved: https://github.com/pytorch/pytorch/pull/158044 Approved by: https://github.com/Skylion007	2025-07-10 22:01:18 +00:00
Nikita Shulga	5f1225ef48	[EZ][BE] Delete redundant header (#157966 ) Not sure why it was there in the first place. And why `Indexing.m`` needed to include QScheme.h is also unclear Pull Request resolved: https://github.com/pytorch/pytorch/pull/157966 Approved by: https://github.com/Skylion007	2025-07-10 21:59:36 +00:00
Laith Sakka	96897e721b	Return false in statically_known_multiple_of if numerator has more than 20 unique symbols (#157855 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157855 Approved by: https://github.com/bobrenjc93 ghstack dependencies: #155590, #157845	2025-07-10 21:00:57 +00:00
Laith Sakka	d7e0098bf3	Fix is_unaligned usage of statically_known_true (#157845 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157845 Approved by: https://github.com/ColinPeppler ghstack dependencies: #155590	2025-07-10 21:00:57 +00:00
Sandeep Narendranath Karjala	76ca23c41c	[dynamo] Add FakeProcessGroup support for fx_graph_runnable with distributed collectives (#157162 ) Stack from [ghstack](https://github.com/ezyang/ghstack) (oldest at bottom): Summary: - Modified generate_compiler_repro_string() to automatically detect distributed operations and inject FakeProcessGroup setup code - Added distributed collective tests in test/dynamo/test_fx_graph_runnable.py using FakeProcessGroup API to test distributed collective operations - Generated fx_graph_runnable code now runs successfully standalone when containing distributed operations ```import os os.environ['TORCHINDUCTOR_CACHE_DIR'] = '/var/folders/fd/kcv8m1kn0lqgxz42wvgr46sc0000gn/T/torchinductor_skarjala' import torch from torch import tensor, device import torch.fx as fx from torch._dynamo.testing import rand_strided from math import inf import torch._inductor.inductor_prims import torch.distributed as dist from torch.testing._internal.distributed.fake_pg import FakeStore import torch._dynamo.config import torch._inductor.config import torch._functorch.config import torch.fx.experimental._config torch._functorch.config.functionalize_rng_ops = False torch._functorch.config.fake_tensor_allow_unsafe_data_ptr_access = True torch._functorch.config.unlift_effect_tokens = True isolate_fails_code_str = None # torch version: 2.9.0a0+gitf23d314 # torch cuda version: None # torch git version: f23d31463ca452918e23063409a2bdc55efc0d46 # torch.cuda.is_available()==False, no GPU info collected from torch.nn import * class Repro(torch.nn.Module): def __init__(self) -> None: super().__init__() def forward(self, arg0_1): all_reduce = torch.ops._c10d_functional.all_reduce.default(arg0_1, 'sum', '0') wait_tensor = torch.ops._c10d_functional.wait_tensor.default(all_reduce); all_reduce = None mul = torch.ops.aten.mul.Tensor(wait_tensor, 2) copy_ = torch.ops.aten.copy_.default(arg0_1, wait_tensor); arg0_1 = wait_tensor = copy_ = None return (mul,) def load_args(reader): buf0 = reader.storage(None, 64) reader.tensor(buf0, (4, 4), is_leaf=True) # arg0_1 load_args._version = 0 mod = Repro() if __name__ == '__main__': from torch._dynamo.repro.after_aot import run_repro # Initialize FakeProcessGroup for distributed operations store = FakeStore() dist.init_process_group( backend="fake", rank=0, world_size=2, store=store ) with torch.no_grad(): run_repro(mod, load_args, accuracy=False, command='run', save_dir=None, tracing_mode='real', check_str=None) # To run it separately, do # mod, args = run_repro(mod, load_args, accuracy=False, command='get_args', save_dir=None, tracing_mode='real', check_str=None) # mod(*args) dist.destroy_process_group() ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157162 Approved by: https://github.com/xmfan	2025-07-10 20:30:27 +00:00
Aleksandar Samardžić	a3ec6d64b2	Update test after CUTLASS upgrade (#157903 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157903 Approved by: https://github.com/ngimel	2025-07-10 20:10:20 +00:00
morrison-turnansky	8c5b070d1f	Documentation Fix: torch.tensor.scatter_ docs (#157929 ) updated torch.tensor.scatter_ docs to reflect proper broadcasting behavior Fixes #157419 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157929 Approved by: https://github.com/albanD	2025-07-10 19:22:52 +00:00
Nicolas De Carli	da4e7c77a1	[caffe2] Enable auto vectorization (#157984 ) Summary: We are testing enabling back autovectorization in some codepaths. These resulted in crashes when compiling using clang17, we are now relying on clang19. Test Plan: buck2 build //caffe2/caffe2/fb/transforms:sigrid_interface We are going to deploy it on ads workloads Rollback Plan: Differential Revision: D77448445 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157984 Approved by: https://github.com/Skylion007	2025-07-10 19:19:45 +00:00
Sam Larsen	5bd7804be2	Support caching if joint_custom_pre_pass/joint_custom_post_pass implement the proper interface (#157990 ) Summary: Essentially, treat joint_custom_pre_pass/joint_custom_post_pass the same as post_grad_custom_post_pass/post_grad_custom_pre_pass. Test Plan: More unit tests Pull Request resolved: https://github.com/pytorch/pytorch/pull/157990 Approved by: https://github.com/oulgen	2025-07-10 19:17:11 +00:00
morrison-turnansky	e172309880	Documentation Fix: Torch gather broadcasting (#157920 ) updated torch gather docs to reflect proper broadcasting behavior for specific backends Fixes #157425 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157920 Approved by: https://github.com/albanD	2025-07-10 19:08:51 +00:00
Edward Z. Yang	e2f64eedaf	Fix DTensor handling of conjugate bit. (#158030 ) Fixes https://github.com/pytorch/pytorch/issues/130646 specifically for DTensor Fixes https://github.com/pytorch/torchtitan/issues/267 Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/158030 Approved by: https://github.com/bdhirsh, https://github.com/albanD	2025-07-10 18:28:12 +00:00
zeshengzong	2db1a54465	Add deprecation hint for accelerator APIs (#158013 ) [torch.accelerator.set_device_idx](https://docs.pytorch.org/docs/stable/generated/torch.accelerator.set_device_idx.html#torch.accelerator.set_device_idx) and [torch.accelerator.current_device_idx](https://docs.pytorch.org/docs/stable/generated/torch.accelerator.current_device_idx.html#torch.accelerator.current_device_idx) are deprecated, but not reflect in their docs. ## Test Result ### Before ![image](https://github.com/user-attachments/assets/6e0d8c4a-d5e5-420c-8f3a-b2742f0fe263) ![image](https://github.com/user-attachments/assets/4bd99b15-31dc-4043-82e8-3d2c1dfcb57b) ![image](https://github.com/user-attachments/assets/a3d342da-79f2-4950-b17a-d01257603c97) ### After ![image](https://github.com/user-attachments/assets/faf138a8-bd92-4f31-bd7c-4414aee6da5b) ![image](https://github.com/user-attachments/assets/212456bc-1c6b-48c6-9d8c-075d5096b900) ![image](https://github.com/user-attachments/assets/49bb9c8c-203e-424e-bdc0-0f197239146e) Pull Request resolved: https://github.com/pytorch/pytorch/pull/158013 Approved by: https://github.com/guangyey, https://github.com/albanD	2025-07-10 18:09:22 +00:00
Scott Wolchok	e3f8141c25	Fix UB in BFloat16 round_to_nearest_even (#157942 ) Type punning using unions is undefined behavior in C++ (you may not access a member of a union that is not the active member). bit_cast is the right way. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157942 Approved by: https://github.com/Skylion007	2025-07-10 18:03:39 +00:00
henrylhtsang	a9ac9f2635	[cutlass backend] Change serialization protocol to use more json and cache (#157840 ) Differential Revision: [D77949177](https://our.internmc.facebook.com/intern/diff/D77949177/) What this diff does: * use lru_cache for serialization and deserialization * json dumps more. This seems to help perf. For instantiation level 3332, the loading time decreases from 33s to 20s (roughly 40%) decrease. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157840 Approved by: https://github.com/ColinPeppler ghstack dependencies: #157839	2025-07-10 17:44:33 +00:00
fduwjj	1d0f45d5d1	[c10d][PGNCCL] Cleanup unused params for nccl comm split (#157978 ) Previously we add global ranks as a input params for nccl comm. Now this is not needed, let's clean that up. Differential Revision: [D78051047](https://our.internmc.facebook.com/intern/diff/D78051047) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157978 Approved by: https://github.com/Skylion007	2025-07-10 17:36:23 +00:00
Edward Z. Yang	b40c0b61eb	Make guard collective logging less chatty (#157995 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157995 Approved by: https://github.com/Microve, https://github.com/albanD, https://github.com/Skylion007	2025-07-10 17:18:37 +00:00
henrylhtsang	fb45649df7	[cutlass backend] Make config request key depend on serialization.py and cutlass_utils.py (#157839 ) Differential Revision: [D77893241](https://our.internmc.facebook.com/intern/diff/D77893241/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157839 Approved by: https://github.com/ColinPeppler	2025-07-10 17:09:32 +00:00
Catherine Lee	7caf6c801d	[ez][CI] Add docker instructions for linux build (#157974 ) Copied from linux-test.yml I'm not sure how necessary this is because the wiki also has this info, and has more details about it Pull Request resolved: https://github.com/pytorch/pytorch/pull/157974 Approved by: https://github.com/huydhn	2025-07-10 16:15:28 +00:00
PyTorch MergeBot	493bd625e2	Revert "[BE]: Reduce binary size 40% using aggressive fatbin compression. (#157791 )" This reverts commit 9bdf87e8918b9a3f78d7bcb8a770c19f7c82ac15. Reverted https://github.com/pytorch/pytorch/pull/157791 on behalf of https://github.com/albanD due to Reverting to avoid regressing on the driver supported ([comment](https://github.com/pytorch/pytorch/pull/157791#issuecomment-3058091176))	2025-07-10 16:14:06 +00:00
Shangdi Yu	4781d72faa	[AOTI] codegen for static linkage (#157129 ) Design doc: https://docs.google.com/document/d/1ncV7RpJ8xDwy8-_aCBfvZmpTTL824C-aoNPBLLVkOHM/edit?tab=t.0 (internal) - Add codegen for static linkage - refactor test code for test_compile_after_package tests For now, the following options must be used together with `"aot_inductor.compile_standalone": True`. "aot_inductor.package_cpp_only": True, Will change `"aot_inductor.package_cpp_only"` to be automatically set to True in followup PR. ``` python test/inductor/test_aot_inductor_package.py -k test_compile_after_package python test/inductor/test_aot_inductor_package.py -k test_run_static_linkage_model ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157129 Approved by: https://github.com/desertfire	2025-07-10 16:03:50 +00:00
Aaron Gokaslan	9bdf87e891	[BE]: Reduce binary size 40% using aggressive fatbin compression. (#157791 ) NVCC apparently has a [compression-mode flag](https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/#compress-mode-default-size-speed-balance-none-compress-mode) to tell it how you want to compress the fatbinary since 12.4. This mode defaults to speed (pick a low compression mode that loads the file quickly). Since we are running into PyPi size issues, this will allow us to upload smaller wheel files. From: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html#compress-mode-default-size-speed-balance-none-compress-mode ``` size Uses a compression mode more focused on reduced binary size, at the cost of compression and decompression time. ``` Up to 37.2% reduction in binary size with virtually no drawback (except potentially a little slower loading of the .so at PyTorch startup). 694 MB for CUDA 12.9 builds with 6.0;7.0;7.5;8.0;8.6;9.0;10.0;12.0+PTX vs 1.08GB for CUDA 12.9 builds with 7.5;8.0;8.6;9.0;10.0;12.0+PTX CUDA 12.9 *694MB* vs *1.08GB* CUDA 12.8 *604MB* vs *845MB* This ends up saving PyPi.org approximately 19.6 PiB of bandwidth per month for the CUDA 12.9 case. This will also allow us to add back CUDA 12.8 12.0+PTX which will make the package forward compatible on newer GPUs. Undoing the need for PR https://github.com/pytorch/pytorch/pull/157516 and https://github.com/pytorch/pytorch/pull/157634 <img alt="Screenshot 2025-07-08 at 5 36 44 PM" width="1061" src="https://private-user-images.githubusercontent.com/7563158/463890713-a53ec774-b036-4c0b-a5d5-301756e3644f.png?jwt=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NTIwNzY3OTIsIm5iZiI6MTc1MjA3NjQ5MiwicGF0aCI6Ii83NTYzMTU4LzQ2Mzg5MDcxMy1hNTNlYzc3NC1iMDM2LTRjMGItYTVkNS0zMDE3NTZlMzY0NGYucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI1MDcwOSUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNTA3MDlUMTU1NDUyWiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9Yzg1OGExN2VjYmI3ZDFhNjIwZDk0NTBjOWFlZDIzYzY3MmExYTFiOGZhZjc0NTI1ZTk2YzM3YzdhYzkyYzZlMiZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3QifQ.2-YmmfXrBFuXCrjDCQ_iTgbtbwv9xNFqM6Goc_liDKE"> More details can be found in Nvidia's technical blog for CUDA 12.4: https://developer.nvidia.com/blog/runtime-fatbin-creation-using-the-nvidia-cuda-toolkit-12-4-compiler/ Pull Request resolved: https://github.com/pytorch/pytorch/pull/157791 Approved by: https://github.com/malfet, https://github.com/atalman	2025-07-10 15:51:04 +00:00
Aditya Tewari	f85954e043	Update OpenBLAS commit (#151547 ) Motivation: Update OpenBLAS and change build script to enable SBGEMM kernels . Update pytorch `jammy` builds for aarch64 to use `install_openblas.sh` instead of `conda_install` Link to full [TorchInductor Performance Dashboard AArch64](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Fri%2C%2006%20Jun%202025%2009%3A46%3A35%20GMT&stopTime=Fri%2C%2013%20Jun%202025%2009%3A46%3A35%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cpu%20(aarch64)&lBranch=adi/update_openblas&lCommit=0218b65bcf61971c1861cfe8bc586168b73aeb5f&rBranch=main&rCommit=9d59b516e9b3026948918e3ff8c2ef55a33d13ad) 1. This shows a promising speedup across most of the HF models in benchmark, specifically giving a significant boost to SDPA layers. 2. Overall torch-bench pass-rate (cpp_wrapper mode) increased `[87%, 65/75 → 96%, 72/75]` <img width="676" alt="Screenshot 2025-06-20 at 17 05 15" src="https://github.com/user-attachments/assets/2ca9c1bc-80c6-464a-8db6-b758f2476582" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/151547 Approved by: https://github.com/malfet, https://github.com/snadampal, https://github.com/fadara01 Co-authored-by: Christopher Sidebottom <chris.sidebottom@arm.com> Co-authored-by: Ryo Suzuki <ryo.suzuki@arm.com> Co-authored-by: Ye Tao <ye.tao@arm.com> Co-authored-by: Nikita Shulga <2453524+malfet@users.noreply.github.com>	2025-07-10 14:58:12 +00:00
Sam Larsen	7702855228	[logging] dynamo_timed the synchronize in CachingAutotuner make_launchers (#157747 ) Summary: There's some evidence that some very long compile times are actually attributable to the sync. This should make it easier to say for sure. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157747 Approved by: https://github.com/aorenste, https://github.com/mlazos	2025-07-10 14:48:51 +00:00
Frank Lin	9a5278225f	[CUDA] Use runtime driver API for cuStreamWriteValue32 (#156097 ) Fixes #154073 Reference: https://github.com/NVIDIA/Fuser/pull/4197 See PR #154097 @nWEIdia is currently out of the office, so I’ve temporarily taken over his work. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156097 Approved by: https://github.com/syed-ahmed, https://github.com/wujingyue, https://github.com/atalman Co-authored-by: Wei Wang <weiwan@nvidia.com>	2025-07-10 14:38:18 +00:00
Howard Huang	8532033679	RPC tutorial audit (#157938 ) Fix [T228333894](https://www.internalfb.com/intern/tasks/?t=228333894) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157938 Approved by: https://github.com/AlannaBurke	2025-07-10 14:15:37 +00:00
IvanKobzarev	8dff457f42	[simple_fsdp] Port fx pass to bucket reduce_scatters (#157780 ) Porting fx passes for reduce_scatters bucketing (similar to all_gather bucketing) for simple_fsdp and autoparallel testing. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157780 Approved by: https://github.com/wconstab	2025-07-10 14:04:43 +00:00
rzou	a9537b626c	[standalone_compile] Fix single Tensor outputs from split_module (#157803 ) We assumed that the output in an FX graph would always just be a list[Tensor], even in the single tensor return case. It is possible for the output to be a single Tensor. This can happen by calling torch.fx.split_module on the module. Test Plan: - new test Pull Request resolved: https://github.com/pytorch/pytorch/pull/157803 Approved by: https://github.com/oulgen	2025-07-10 12:49:03 +00:00
Raymond Li	82765dad16	Fix logging of config_suppress_errors and config_inline_inbuilt_nn_modules (#157947 ) Currently ~50% of the time we fail or crash before logging metrics, so moving where this is logged will let us have more comprehensive (less-null) data. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157947 Approved by: https://github.com/masnesral, https://github.com/jovianjaison	2025-07-10 12:05:43 +00:00
David Berard	cd995bfb2a	[inductor] re-enable TMA templates w/ AOTI (#157819 ) Follow-up from #155896: now that AOTI can codegen non-null TMA workspace args, we can re-enable TMA templates w/ AOTI. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157819 Approved by: https://github.com/drisspg	2025-07-10 08:35:29 +00:00
Yu, Guangye	1e8e9f745e	Introduce AcceleratorAllocatorConfig as the common class (#149601 ) # Motivation This PR aims to generalize `AllocatorConfig` to be device-agnostic. Introduce the class `AcceleratorAllocatorConfig` to clarify its scope as a configuration manager for accelerator backends (e.g., CUDA, XPU). The another name `AllocatorConfig` is now reserved for a potential future base class that can unify configuration handling for both CPU and accelerator allocators, should similar requirements arise for the CPU path. # Design Rule ## Overall This class configures memory allocation for both device and host memory. A single `AcceleratorAllocatorConfig` instance is shared across all accelerator backends, such as CUDA and XPU, under the assumption that relevant environment variables apply uniformly to all accelerators. Device-specific configuration extensions are supported via hooks (see `registerDeviceConfigParserHook`). Introduce a new class `ConfigTokenizer` to help process the env variable config key-value pair ## Naming Convention: - Public API names in `AcceleratorAllocatorConfig` should be device-generic. - Members prefixed with `pinned_` are specific to the host/pinned allocator. - Environment variable names should be generic across backends. - Comma-separated key-value pairs in the format: `key:value`. Use square brackets `[]` for list values Example: `key1:123, key2:[val1,val2]` ## Environment Variables: - The default environment variable for configuration is `PYTORCH_ALLOC_CONF`. - For backward compatibility, `PYTORCH_CUDA_ALLOC_CONF` and `PYTORCH_HIP_ALLOC_CONF` are also supported with lower priority. Pull Request resolved: https://github.com/pytorch/pytorch/pull/149601 Approved by: https://github.com/albanD	2025-07-10 07:05:39 +00:00
Xuehai Pan	af3d069094	[BE][Easy] remove unused build-time dependency `astunparse` and change `astunparse.unparse` -> `ast.unparse` (#157907 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157907 Approved by: https://github.com/Skylion007	2025-07-10 07:04:42 +00:00
Meng, Chunhuan	ba0d0de5e6	Enable set SDPA backend by torch.nn.attention.sdpa_kernel on XPU (#156669 ) Introduces support for a new `OVERRIDEABLE` backend in the SDPA module, improves backend selection logic, and adds corresponding tests. In addition, a fallback mechanism was added when a specific backend is unavailable, enhancing user configurability. ### Backend Support and Selection Enhancements: * Added `at::SDPBackend::overrideable` to the list of available SDPA backends in the `Context` class (`aten/src/ATen/Context.h`). * Updated the backend selection logic in `select_sdp_backend_xpu` to include the `OVERRIDEABLE` backend and added a fallback mechanism for unsupported `FLASH_ATTENTION` on XPU. * Adjusted error messaging in `_fused_sdp_choice_xpu` to reflect the inclusion of the `OVERRIDEABLE` backend. (`aten/src/ATen/native/mkldnn/xpu/Attention.cpp`) ### Test Additions for Backend Fallback and Selection: * Added new unit tests to validate fallback behavior for `FLASH_ATTENTION` to `OVERRIDEABLE` and to verify correct backend selection when `MATH` is enabled. (`test/test_transformers.py`,) ### Codebase Updates for Backend Integration: * Introduced `OVERRIDEABLE` as a new member of the `_SDPBackend` enum. (`torch/_C/__init__.pyi.in`) * Extended `_backend_names` and updated related methods to handle the `OVERRIDEABLE` backend. (`torch/nn/attention/__init__.py`) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156669 Approved by: https://github.com/guangyey, https://github.com/drisspg	2025-07-10 06:52:22 +00:00
Pian Pawakapan	4cc13c4af6	[dynamic shapes] avoid unnecessary slices (#157528 ) Fixes #157289, by extending optimization to slices where the end index exceeds the size. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157528 Approved by: https://github.com/angelayi	2025-07-10 06:34:46 +00:00
FFFrog	565fd07909	[Easy] Make the error message shown by THPUtils_unpackLong to be clearer (#157886 ) As the title stated. The error message of `THPUtils_unpackLong` is the same as `THPUtils_unpackInt` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157886 Approved by: https://github.com/Skylion007	2025-07-10 06:26:13 +00:00
Yuanhao Ji	b85f10ea50	[BE] Replace `std::runtime_error` with `TORCH_CHECK` [2/N] (#152080 ) Part of: #148114 Related commits: - #151880 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152080 Approved by: https://github.com/cyyever, https://github.com/albanD	2025-07-10 06:02:47 +00:00
Jithun Nair	fadc936fad	Updates to build and test on Noble (Ubuntu24.04) and py3.12 (#152240 ) This PR enables Ubuntu24.04 testing on CI: * Builds a base docker image using Noble (Ubuntu24.04) and py3.12 for ROCm N version * Builds and tests PyTorch on Ubuntu24.04 as part of the `rocm-mi300` workflow Pull Request resolved: https://github.com/pytorch/pytorch/pull/152240 Approved by: https://github.com/jeffdaily, https://github.com/malfet	2025-07-10 05:55:42 +00:00
Ewart, Timothee	b7860c7863	Implement fast exp for AVX2 and AVX512 for the flash attention (#151441 ) Implement fexp for avx2 and avx512 Cristiano and all propose a clever exp using the IEEE representation with a fine control of the precision, especially useful for mix computation of the flash attention. - Implement Fast Exponential Computation on SIMD Architectures A. Cristiano I. Malossi, Yves Ineichen, Costas Bekas, and Alessandro Curioni - AVX2 and AVX512 float only, up to 20% faster for mix precision flash attention than the current implementation. - For the other types legacy implementation. Precision 1 ULP only valid in hybrid mode fp32 -> f16 due to the cast during the store operation in the flash attention: Benchmark Machine Xeon 6972P, results in TOPs, Python forward pass flash attention numhead 16, Head dimension 64 \|Seq. L.\| PT \| fexp \| \|-------\|------\|------\| \| 512 \| 0.8 \| 1.3 \| \| 1024 \| 1.7 \| 1.7 \| \| 2048 \| 6 \| 6.1 \| \| 4096 \| 16 \| 16.8 \| \| 8192 \| 30.6 \| 32.3 \| \| 16384 \| 40 \| 40.8 \| \| 32768 \| 44.9 \| 51.4 \| \| 65536 \| 45.8 \| 54.4 \| numhead 16, Head dimension 128 \|Seq. L.\| PT \| fexp \| \|-------\|------\|------\| \| 512 \| 2.5 \| 4.1 \| \| 1024 \| 3.3 \| 4 \| \| 2048 \| 11.4 \| 10.5 \| \| 4096 \| 27.4 \| 28.4 \| \| 8192 \| 44.4 \| 46 \| \| 16384 \| 64.2 \| 68.1 \| \| 32768 \| 77.8 \| 83 \| \| 65536 \| 82.1 \| 88.1 \| numhead 16, Head dimension 256 \|Seq. L.\| PT \| fexp \| \|-------\|------\|------\| \| 512 \| 1.7 \| 3.4 \| \| 1024 \| 4.2 \| 6.5 \| \| 2048 \| 14.6 \| 16.1 \| \| 4096 \| 30.1 \| 31.1 \| \| 8192 \| 60 \| 62 \| \| 16384 \| 83.3 \| 87.3 \| \| 32768 \| 98.7 \| 106 \| \| 65536 \| 102.2\| 107.1\| Pull Request resolved: https://github.com/pytorch/pytorch/pull/151441 Approved by: https://github.com/mingfeima	2025-07-10 05:51:31 +00:00
Avik Chaudhuri	9222552572	[non-strict export] uncovered cases of select and slice (#157821 ) Summary: `None` and `Ellipsis` in multi-dimensional indexing was previously not covered. Moreover, we introduce a small optimization for `slice(None)` and a passthrough when symints do not appear in the indexing. The remaining case is where indexing is by tensor, which is fairly complicated; we passthrough in that case. Test Plan: added tests Rollback Plan: Differential Revision: D77943929 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157821 Approved by: https://github.com/pianpwk	2025-07-10 05:48:12 +00:00
Sheng Fu	3584e84c24	Fixed the function to get the origin nodes of fused triton kernel. (#157578 ) Summary: This DIFF is to fix the following issue: In python source code for CompiledFxGraph,the FX graph segment for the Triton kernel is broken. For example, the following function def fn(a, b, c): x = torch.nn.functional.linear(a, b) x = x.sin() x = x.t() + c return x Inductor compiled this FX graph into two nodes: the first one is mm, the second one is a triton kernel for sin + transpose + add. The FX graph segment for the triton kernel is like the following: Graph fragment: %add : [num_users=1] = call_function[target=torch.ops.aten.add.Tensor](args = (%permute_1, %arg2_1), kwargs = {}) Basically only "add" node in the FX graph. The root cause is function caffe2/torch/_inductor/utils.py:gather_origins does not detect the realized node correctly. To fix this issue, the IRNode is checked if it is one of the following IRNode: ir.ComputedBuffer, ir.InputsKernel, ir.InputBuffer, ir.ReinterpretView, ir.TemplateBuffer, If it is one of them, it is realized, otherwise, it is not. Test Plan: buck2 run mode/opt caffe2/test/inductor:provenance_tracing -- caffe2.test.inductor.test_provenance_tracing.TestProvenanceTracingArtifact.test_triton_kernel_to_post_grad_tracing_cuda Rollback Plan: Differential Revision: D77748371 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157578 Approved by: https://github.com/mlazos	2025-07-10 05:34:50 +00:00
Dmitry Rogozhkin	b146ca74f0	docs: add get_default_backend_for_device to distributed documentation (#156783 ) `torch.distributed.get_default_backend_for_device()` API was added to torch 2.6, but is still missing in distributed documentation. This commit addresses the gap. CC: @guangyey, @EikanWang Pull Request resolved: https://github.com/pytorch/pytorch/pull/156783 Approved by: https://github.com/guangyey, https://github.com/malfet	2025-07-10 05:11:30 +00:00
CaoE	eddddea908	Upgrade MKL in CI (#154198 ) This PR is to upgrade MKL in CI as PyTorch release uses MKL 2024.2 while MKL in CI is 2021.4. MKL 2021.4 can't trigger issues like https://github.com/pytorch/pytorch/issues/154477 caused by MKL upgrading in Torch release. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154198 Approved by: https://github.com/leslie-fang-intel, https://github.com/malfet ghstack dependencies: #154585	2025-07-10 05:09:51 +00:00
bobrenjc93	80bcaa4195	have dynamic sources only apply to sizes and not strides (#157960 ) @animesh pointed out using whitelist for strides can result in confusing graphs as follows ``` s60: "Sym(s60)", L_hidden_states_: "bf16[1, 4096, 3072][s60, 3072, 1]cuda:0" ``` We probably want to capture the relationship between sizes and strides anyways so let's make it so the whitelist only makes the sizes dynamic. That same graph now looks lik ethis ``` L_hidden_states_: "bf16[1, 4096, 64][262144, 64, 1]cuda:0" ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157960 Approved by: https://github.com/pianpwk	2025-07-10 05:03:51 +00:00
PyTorch UpdateBot	88cd9f34b0	[audio hash update] update the pinned audio hash (#157873 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157873 Approved by: https://github.com/pytorchbot	2025-07-10 04:59:50 +00:00
zeshengzong	2b19d85d70	FractionalMaxPool3d add kernel_size check (#155549 ) Fixes #96316 ## Test Result ```python >>> import torch >>> from torch.func import jacrev, grad, vmap >>> >>> torch.manual_seed(420) <torch._C.Generator object at 0x7fe4767810d0> >>> >>> input = torch.randn(1, 1, 5, 5, 5, requires_grad=True) >>> >>> def func(input): ... model = torch.nn.FractionalMaxPool3d(kernel_size=0, output_size=(1, 1, 1)) ... output = model(input) ... return output ... >>> >>> func(input).sum().backward() Traceback (most recent call last): File "<stdin>", line 1, in <module> File "<stdin>", line 2, in func File "/home/zong/code/pytorch/torch/nn/modules/pooling.py", line 1054, in __init__ raise ValueError(f"kernel_size must greater than 0, but got {kernel_size}") ValueError: kernel_size must greater than 0, but got 0 ``` ![image](https://github.com/user-attachments/assets/52780ce7-3951-4d1c-95a4-5ce2bf65c727) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155549 Approved by: https://github.com/albanD	2025-07-10 04:55:06 +00:00
CaoE	06a40b6850	Fix MKL error: Inconsistent configuration parameters (#154585 ) Fixes #154477. PyTorch release uses 2024.2 MKL, which has some changes to the usage of DFTI: if `DFTI_NUMBER_OF_TRANSFORMS > 1`, `DFTI_INPUT_DISTANCE` and `DFTI_OUTPUT_DISTANCE` also needs to be explicitly set to a positive integer. In addition, the requirement "the datasets to be transformed cannot contain common elements" should also be satisfied. This means that we need to avoid the case where the input strides have 0. See https://www.intel.com/content/www/us/en/docs/onemkl/developer-reference-dpcpp/2024-2/configuring-data-layouts.html and https://www.intel.com/content/www/us/en/docs/onemkl/developer-reference-c/2024-2/dfti-number-of-transforms.html Pull Request resolved: https://github.com/pytorch/pytorch/pull/154585 Approved by: https://github.com/leslie-fang-intel, https://github.com/soumith, https://github.com/malfet	2025-07-10 03:42:38 +00:00
Shangdi Yu	0a624c2dc5	Fix from_node's graph_id in unlift() (#157943 ) Summary: We should use the node before deepcopy in NodeSource Test Plan: ``` buck run fbcode//caffe2/test:test_export -- -r test_from_node_metadata_export ``` Rollback Plan: Differential Revision: D78022070 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157943 Approved by: https://github.com/angelayi, https://github.com/Gasoonjia	2025-07-10 03:23:55 +00:00
Paul Zhang	4cfc0a3208	[Inductor] Introduce Lookup Table for Overriding Triton Kernel autotune configs post fusion (#157924 ) Summary: Introduce lookup table for kernels post fusion, hashing on inductor generated source code Rollback Plan: Differential Revision: D77866885 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157924 Approved by: https://github.com/jansel	2025-07-10 03:23:50 +00:00
Ankita George	3232b57cd8	Updates to safetensors checkpoint consolidation script to be faster (#157936 ) Summary: - adding mmap-ing - more efficient writing in larger chunks latency from ~150s to ~6s for simple row-wise consolidation of a 7gb model sharded across 4 ranks Test Plan: ran consolidation with the following code: ``` from torch.distributed.checkpoint._consolidate_hf_safetensors import consolidate_safetensors_files import time start_time = time.time() consolidate_safetensors_files(base_path, consolidated_path) end_time = time.time() print(f"Time taken: {end_time - start_time} seconds") ``` With the old code this was taking a couple minutes and this is now down to ~6s. Internal users can find the tensor shards in the manifold path: manifold://ankita_test_bucket/tree/safetensors Rollback Plan: Differential Revision: D77960054 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157936 Approved by: https://github.com/teja-rao, https://github.com/pradeepfn	2025-07-10 02:50:20 +00:00
Ankita George	3404c1f0cf	[HF][DCP] Upload local consolidated files to remote storage if needed (#157371 ) If the final output file is in remote storage, then create a local temp directory to write the files and upload the files to the remotes storage after they are written. Add a new config to the storage writer, `enable_consolidation`, so we don't need to rely on the presence of the `consolidation_output_path` to decide if consolidation is enabled. If `enable_consolidation` is True and `consolidation_output_path` isn't provided, the consolidated safetensors will be added to the same path as the sharded ones. Differential Revision: [D77554585](https://our.internmc.facebook.com/intern/diff/D77554585/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157371 Approved by: https://github.com/pradeepfn	2025-07-10 02:40:25 +00:00
FFFrog	aab949aa96	Deprecated pkg_resources and use distributions instead (#151915 ) As the title stated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151915 Approved by: https://github.com/malfet, https://github.com/atalman, https://github.com/albanD	2025-07-10 01:51:26 +00:00
Edward Yang	6442ae9256	Make the name assert actually do something, and reserve some more names (#157342 ) Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157342 Approved by: https://github.com/albanD	2025-07-10 01:39:40 +00:00
Edward Yang	db188503cb	[BE] Remove stale pyre-fixme (#157816 ) Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157816 Approved by: https://github.com/Skylion007, https://github.com/jingsh, https://github.com/albanD	2025-07-10 01:33:32 +00:00
Edward Yang	693116f765	[doc] DeviceMesh invariant on DTensorSpec (#157806 ) Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157806 Approved by: https://github.com/Skylion007, https://github.com/wanchaol ghstack dependencies: #157805	2025-07-10 01:27:40 +00:00
Edward Yang	9a4ac71b58	[doc] Document an invariant in OpSpec (#157805 ) I am not sure if this is actually true though, please reject this PR if it is not. Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157805 Approved by: https://github.com/wanchaol, https://github.com/zpcore	2025-07-10 01:27:40 +00:00
Michelle Madubuike	8387984257	Improve error message for torch.binomial enforcing float inputs (#157658 ) Fixes #157195 ### Summary: Fixed Issue 157195 by adding a new error message for torch.binomial in aten/src/ATen/native/Distributions.cpp ### Explanation According to the issue, ``` import torch torch.binomial(torch.tensor([10]).long(), torch.tensor([0.5])) ``` `RuntimeError: Found dtype Float but expected Long` It looks like we are getting a Tensor error rather than a binomial function error. Since the error is coming from pytorch/aten/src/ATen/TensorIterator.cpp, it seems like it is trying to align the tensor data to the same datatype for smooth tensor computations instead of giving a binomial function error. I tried using both arguments as longs and both as ints and got the right binomial function error ``` torch.binomial(torch.tensor([10]).long(), torch.tensor([0.5]).long()) NotImplementedError: "binomial_cpu" not implemented for 'Long' ``` ``` torch.binomial(torch.tensor([10.0]).int(), torch.tensor([0.5]).int()) NotImplementedError: "binomial_cpu" not implemented for 'Int' ``` But when I have both as different datatypes, the TensorIterator.cpp error comes back trying to align the datatypes. `RuntimeError: Found dtype Float but expected Long` I then tried finding where the NotImplementation Error was documented and found it in pytorch/aten/src/ATen/Dispatch.h in lines 193 - 211 ``` #define AT_DISPATCH_SWITCH(TYPE, NAME, ...) \ [&] { \ const auto& the_type = TYPE; \ constexpr const char* at_dispatch_name = NAME; \ /* don't use TYPE again in case it is an expensive or side-effect op / \ at::ScalarType _st = ::detail::scalar_type(the_type); \ RECORD_KERNEL_FUNCTION_DTYPE(at_dispatch_name, _st); \ switch (_st) { \ __VA_ARGS__ \ default: \ TORCH_CHECK_NOT_IMPLEMENTED( \ false, \ '"', \ at_dispatch_name, \ "\" not implemented for '", \ toString(_st), \ "'"); \ } \ }() ``` In the AT_DISPATCH_SWITCH* function, it picks a tensor and its datatype and checks if the Tensor datatype matches the supported datatypes. If not we get the Not Implemented error. Unfortunately, I think the AT_DISPATCH_SWITCH function, uses the `common_dtype` from TensorIterator in order to run. So TensorIterator.cpp needs to happen before the AT_DISPATCH_SWITCH function. ### Summary: We are getting the wrong error message because TensorIterator.cpp gets called and errors out due to Tensor datatype mismatch before we can get the right error message in Dispatch.h for torch.binomial not supporting that datatype. ### Options for the Fix Option 1: Make the error message in TensorIterator.cpp more general so it applies to torch.binomial. An error message along the lines `RunTime Error : "Tensor Datatypes", op.target_dtype," and ", common_dtype_, "are different "` Option 2: Add an error message for the binomial function datatype mismatch before the the TensorIterator.cpp error message gets called. Although Option 1 seemed easier I think Option 2 might be better as it is more specific to the binomial function while Option1 would affect all Tensors with datatype mismatch. This PR applies the fix for Option 2 After Fix : ``` torch.binomial(torch.tensor([10]).long(), torch.tensor([0.5])) RuntimeError: Binomial function arguments count and prob must have same datatype of type Float, got: count = Long, prob = Float ``` ``` torch.binomial(torch.tensor([10]).long(), torch.tensor([0.5]).long()) NotImplementedError: "binomial_cpu" not implemented for 'Long' ``` @malfet Pull Request resolved: https://github.com/pytorch/pytorch/pull/157658 Approved by: https://github.com/soulitzer	2025-07-10 00:58:56 +00:00
Brian Hirsh	54a7e5b598	_aot_export_function: allow keeping input mutations in the graph (#157730 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157730 Approved by: https://github.com/ezyang	2025-07-10 00:47:51 +00:00
zeshengzong	ed03492238	Add check `nested_tensor_from_jagged` param `jagged_dim >= 1` (#157770 ) Fixes #157404 ## Test Result ```bash pytest test/test_nestedtensor.py ...............................................s..........ssssss.................................................................................................s.s..sssss..s...ss............................................................. [ 44%] ...........................................................sssss....sss...s.........ss....s....sss.........s.sss...s..s......s............s.sss.ss...............s.....................s....s......................s.s.....s....s..s..ssssssssss [ 59%] sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss..ssssss.ssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ssssssss...............................s........................................... [ 74%] .......sss...................................................................................................................................................................................................................................... [ 89%] ....sss.......................................................................................................................................................... [100%] ==================================================================================================== 1317 passed, 258 skipped in 2504.27s (0:41:44) ==================================================================================================== ``` ![image](https://github.com/user-attachments/assets/dcc8e46d-b88f-4580-b4ad-0999bad33ec9) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157770 Approved by: https://github.com/soulitzer Co-authored-by: Jeffrey Wan <soulitzer@gmail.com>	2025-07-10 00:34:39 +00:00
Pian Pawakapan	752f202ef3	[PGO] include module int attributes in PGO state (#157518 ) Dynamo specializes on int module attributes by default. This includes them in PGO state despite specialization, if they're involved in guards. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157518 Approved by: https://github.com/bobrenjc93	2025-07-09 23:57:54 +00:00
Tristan Rice	ed051c3084	torch.distributed: add initial _dist2 prototype API (#157841 ) This adds the initial dist2 API as proposed in https://docs.google.com/document/d/13R-1t_yESTvmAjcCN-wQjQQadIEu0JNIdS65uZawZzY/edit?tab=t.0#heading=h.3ctbqqopzc89 This is a WIP experimental API and is a sandbox for a number of new features and quality of life improvements/changes to c10d. Test plan: ``` pytest test/distributed/test_dist2.py ``` Docs ``` cd docs make html ``` ![Screenshot 2025-07-08 at 13-39-23 Object Oriented Distributed API - torch distributed _dist2 — PyTorch main documentation](https://github.com/user-attachments/assets/9c03a7ec-09e5-42b9-8478-1ec28bc2b6bd) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157841 Approved by: https://github.com/fduwjj	2025-07-09 23:40:43 +00:00
Xuan Zhang	39456edbba	[PT2][memory] mutation size correctness (#157562 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157562 Approved by: https://github.com/yf225	2025-07-09 22:14:20 +00:00
Aaron Gokaslan	a1dad2f2d2	[BE][Ez]: Autotype torch/profiler with ruff ANN (#157923 ) Apply ruff autotyping fixes to add annotations to torch profiler Pull Request resolved: https://github.com/pytorch/pytorch/pull/157923 Approved by: https://github.com/albanD, https://github.com/sraikund16	2025-07-09 22:07:50 +00:00
Colin Peppler	53ab73090e	[inductor] support unbacked symint in sdpfa (#157739 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157739 Approved by: https://github.com/laithsakka	2025-07-09 22:01:29 +00:00
Ti-Tai Wang	08e9dd280f	[ONNX] Support symbolic arguments in onnx exporter (#157734 ) Previous to this PR, torch.onnx.export(..., dynamo=True, veriy=True, report=True) does not support symbolic arguments. Such examples are like follwing: ```python class M(torch.nn.Module): def forward(self, a, x): return a + torch.tensor(1) + x op = torch.onnx.export(M(), (1, torch.ones(2)), dynamic_shapes=(torch.export.Dim.DYNAMIC, {0: torch.export.Dim.DYNAMIC}), dynamo=True, report=True) ``` symbolic arguments are like constant arguments that they don't have tensor_meta wither. Besides, torch.export.export supports model inputs having constants, which is different from the legacy issue: https://github.com/pytorch/pytorch/issues/99534 where we tried to get the FX directly from dynamo export. Thus, `_remove_non_tensor` is deleted from args processing. NOTE: If the ConstantArugment shows up in exported_program, it was kept to align the length of inputs to nn.Module, but it's irrelevant to the model graph, hwich is why in ONNX model the input is omitted. The test `test_constant_argument_user_input_is_omitted_in_onnx_graph` needs #157719 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157734 Approved by: https://github.com/justinchuby	2025-07-09 21:15:45 +00:00
Aaron Gokaslan	163f0d8f2a	[BE][Ez]: Auto add return type annotations for methods in torch/nn/module (#157925 ) Automatically type a bunch of methods in nn.Module using ruff's type inference rules Pull Request resolved: https://github.com/pytorch/pytorch/pull/157925 Approved by: https://github.com/albanD	2025-07-09 21:12:25 +00:00
Ryan Guo	f742b32a2f	[dynamo] Avoid recompiling over unused objects (#156891 ) Dynamo was aggressively specializing on lazy VTs over `set_name_hint` in `STORE_FAST`, etc., and `isinstance` in `LOAD_FAST_CHECK`. This causes regional `torch.compile` from optimizing ComfyUI GGUF + LoRA to either (1). exceed the recompialtion limit of 8, which results in suboptimal performance, and (2). even if recompilation limit is increased, the compilation time gets unnecessarily high (180s v.s. 20s for Flux). This patch fixes the recompilation issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156891 Approved by: https://github.com/williamwen42, https://github.com/mlazos	2025-07-09 20:14:34 +00:00
Jane Xu	317520bf6e	Add an ovrsource target for torch/headeronly (#157912 ) Summary: no idea how this works Test Plan: will things just pass? Rollback Plan: Differential Revision: D77965219 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157912 Approved by: https://github.com/albanD	2025-07-09 19:32:03 +00:00
PyTorch MergeBot	dfa2649434	Revert "[Inductor] Fix epilogue fusion decision with 1 Triton caller as choice (#156500 )" This reverts commit c48d0f4643b7a69ebe24069e932ce1465a31cdbe. Reverted https://github.com/pytorch/pytorch/pull/156500 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/156500#issuecomment-3053680762))	2025-07-09 18:56:10 +00:00
Shangdi Yu	52772765e0	Change AOTI_RUNTIME_DEVICE_CHECK to be device device specific (#157818 ) Summary: Change AOTI_RUNTIME_DEVICE_CHECK to the following depending on device: AOTI_RUNTIME_CUDA_CHECK AOTI_RUNTIME_XPU_CHECK AOTI_RUNTIME_CPU_CHECK Currently in the codebase, only `AOTI_RUNTIME_CUDA_CHECK` is used. This shouldn't change anything as of now, but we do this to prepare for simultaneouly loading multiple backends (e..g CPU and CUDA) in AOTI standalone. We don't want people writing `AOTI_RUNTIME_DEVICE_CHECK` for both CPU and CUDA checks. This could cause compilation problems when we statically link both CPU and CUDA models. Test Plan: CI Rollback Plan: Reviewed By: muchulee8 Differential Revision: D77742977 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157818 Approved by: https://github.com/jingsh	2025-07-09 18:34:56 +00:00
Jack Francis Dalton	c54778625e	Update `is_sparse` doc to mention that it is sparse_coo specific (#157378 ) ## Issue being addressed `is_sparse` presents itself as determining if a tensor is sparse. HOWEVER, it only does checks against the tensor for `sparse_coo`. This has lead to confusion from developers as when non-coo sparse tensors are provided it return false, despite those tensors being sparse. ## Considered Remedy Fixing this is do-able however would result in complexity as existing systems may depend on this behavior remaining consistent, and even inside of pytorch is_sparse is used by `bform` which states that it supports only `sparse_csr and sparse_coo` meaning additional work/thought would have to go into solving for `sparse_csc` and `sparse_bsr` ## Remedy provided in this PR In lieu of these complications the lowest risk highest gain action was to add clear warning messaging to the function for now to avoid confusion to developers utilizing the function. The rest of the function behavior remains identical ## Issue content Addresses issue number: #101385 Original issue: https://github.com/pytorch/pytorch/issues/101385 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157378 Approved by: https://github.com/soulitzer	2025-07-09 18:22:14 +00:00
mori360	81c7445eb9	[FSDP2] Use reduceOpSum for world size 1 (#157529 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157529 Approved by: https://github.com/Skylion007, https://github.com/lw, https://github.com/weifengpy	2025-07-09 18:08:48 +00:00
Shivam Raikundalia	28aae93f24	[Memory Snapshot] Fix Linter for Global Annotations flag in Snapshot (#157858 ) Summary: We added the ability to make Annotating Global or Local based on an input flag in PyTorch but didn't add the args to the linter Reviewed By: mzzchy Differential Revision: D77959409 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157858 Approved by: https://github.com/mzzchy	2025-07-09 17:28:22 +00:00
Xiangyang (Mark) Guo	b354328ecd	[AOTI] add flag AOT_INDUCTOR_ENABLE_LTO (#157773 ) Add env var AOT_INDUCTOR_ENABLE_LTO to enable clang's ThinLTO by setting AOT_INDUCTOR_ENABLE_LTO=1. The LTO is disabled by default because it may increase the build time. Rollback Plan: Differential Revision: D77899195 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157773 Approved by: https://github.com/desertfire	2025-07-09 16:54:19 +00:00
Tianyu Liu	d75d30eeb6	[DTensor][FSDP2] necessary changes to FSDP and TP to unblock EP (#157216 ) This is to unblock "dp2ep" Expert Parallel + TP integration in torchtitan https://github.com/pytorch/torchtitan/pull/1324. It does two things: 1. Slightly modifies the glue code for FSDP/HSDP + TP to work with FSDP/HSDP + EP and FSDP/HSDP + EP + TP. I kept the name `FSDPParam._tp_spec` to make the change minimal. We can consider renaming it in the future if it confuses people, but I heard @wanchaol has a plan to rewrite DTensor strided sharding entirely. 2. Lifts the check of `_validate_tp_mesh_dim` for `torch.distributed.tensor.parallel.parallelize_module`, as in EP or EP+TP this check is too strict. In particular it assumes a DeviceMesh must have `mesh_dim_names` which is not always true. I'm also removing the file `torch/distributed/tensor/parallel/_utils.py` it belongs entirely, as the other check `_deprecate_warnings`, added two years ago, is not used any more. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157216 Approved by: https://github.com/wanchaol, https://github.com/weifengpy	2025-07-09 16:49:34 +00:00
PyTorch MergeBot	cb711c8fa0	Revert "[BE] always use `uv pip` if possible in `pip_init.py` for `lintrunner init` (#157199 )" This reverts commit 754699610b0abec2fe3f5a73269b1dd09a330445. Reverted https://github.com/pytorch/pytorch/pull/157199 on behalf of https://github.com/malfet due to It breaks lintrunner init` for default environments, see https://github.com/pytorch/pytorch/issues/152999 ([comment](https://github.com/pytorch/pytorch/pull/157199#issuecomment-3053279711))	2025-07-09 16:26:47 +00:00
Nikita Shulga	981c99fdff	Uninstall brew miniconda while running MacOS testing (#156898 ) That results in torch.compile being unable to produce working artifacts But reinstall it later, when done Should fix https://github.com/pytorch/pytorch/issues/156833 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156898 Approved by: https://github.com/seemethere, https://github.com/atalman	2025-07-09 16:02:55 +00:00
FFFrog	054cd4ca28	[CPU Generator] Remove the unused CPUGeneratorImplStateLegacy in set_state (#153934 ) As the title stated. The old state named CPUGeneratorImplStateLegacy in set_state will not been used, so just remove it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153934 Approved by: https://github.com/Skylion007, https://github.com/albanD, https://github.com/malfet, https://github.com/atalman	2025-07-09 15:45:19 +00:00
Svetlana Karslioglu	f4d60a68dd	Adding a change to kick off the theme pull (#157732 ) Adding a small change so that Docker container is rebuild and reflects the latest changes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157732 Approved by: https://github.com/malfet	2025-07-09 15:43:00 +00:00
PyTorch MergeBot	6defd5084e	Revert "[PT2][memory] mutation size correctness (#157562 )" This reverts commit 86670b39fa3df63a652a9a06b59b73f92d70c392. Reverted https://github.com/pytorch/pytorch/pull/157562 on behalf of https://github.com/xuanzhang816 due to internal_test_failure ([comment](https://github.com/pytorch/pytorch/pull/157562#issuecomment-3053115025))	2025-07-09 15:38:29 +00:00
Catherine Lee	b4e3c9ea34	[ez][CI][testing] Set upload artifacts while running to default true if in CI (#157868 ) I was confused about why the distributed tests weren't showing up quickly on HUD, its because the call of run_tests.py for distributed didn't include upload artifacts while running flag, so set it to default to IS_CI so I don't need to put the flag everywhere Pull Request resolved: https://github.com/pytorch/pytorch/pull/157868 Approved by: https://github.com/huydhn	2025-07-09 15:21:25 +00:00
Aaron Gokaslan	fcc682be4b	[BE][Ez]: Fully type nn.utils.clip_grad (#154801 ) Full types clip_grad and exposed typing annotations that were hidden by a bad decorator Pull Request resolved: https://github.com/pytorch/pytorch/pull/154801 Approved by: https://github.com/jansel	2025-07-09 14:27:51 +00:00
Aaron Gokaslan	ed6ae20cf0	[BE][Ez]: Update mimalloc submodule to 2.2.4 (#157794 ) Fixes a few minor bugfixes with the previous release and better compiler support. Should be a NOOP. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157794 Approved by: https://github.com/atalman	2025-07-09 14:03:07 +00:00
Jane Xu	02a9d9095f	[BE] remove commented out code in c10/ovrsource_defs.bzl (#157856 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157856 Approved by: https://github.com/swolchok, https://github.com/albanD	2025-07-09 13:28:56 +00:00
Aidyn-A	86eaf452c3	[Easy][Profiler] Fix pattern matcher of profiler (#157711 ) Per title, as it fails with the following error if "+PTX" was used in `TORCH_CUDA_ARCH_LIST`: ``` File "/usr/local/lib/python3.12/dist-packages/torch/profiler/_pattern_matcher.py", line 313, in skip has_tf32 = all(int(arch[3:]) >= 80 for arch in torch.cuda.get_arch_list()) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/torch/profiler/_pattern_matcher.py", line 313, in <genexpr> has_tf32 = all(int(arch[3:]) >= 80 for arch in torch.cuda.get_arch_list()) ^^^^^^^^^^^^^ ValueError: invalid literal for int() with base 10: 'pute_120' ``` Because slicing `arch[3:]` will not end up on having only digits for `compute_120` element of `torch.cuda.get_arch_list()`: ```python >>> torch.cuda.get_arch_list() ['sm_75', 'sm_80', 'sm_86', 'sm_90', 'sm_100', 'sm_120', 'compute_120'] ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157711 Approved by: https://github.com/Skylion007, https://github.com/sraikund16	2025-07-09 12:09:46 +00:00
Ting Lu	297daa1d30	[aarch64] Add sm_80 to CUDA SBSA build (#157843 ) related to https://github.com/pytorch/pytorch/issues/152690 This adds sm_80 to CUDA SBSA builds (12.9), so that we will be able to support Ampere family (e.g: sm_86) and Ada family (e.g: sm_89) on CUDA SBSA builds. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157843 Approved by: https://github.com/Skylion007, https://github.com/atalman	2025-07-09 11:46:34 +00:00
FFFrog	a355158fcb	[Easy] Fix the compilation warning (#157889 ) Background: ```Shell [1376/2332] Building CUDA object caffe2/CMakeFiles/torch_...h/csrc/distributed/c10d/symm_mem/NCCLSymmetricMemory.cu.o /root/Git.d/pytorch/pytorch/torch/csrc/distributed/c10d/ProcessGroupNCCL.hpp(450): warning #68-D: integer conversion resulted in a change of sign size_t numelIn_ = -1; ^ Remark: The warnings can be suppressed with "-diag-suppress <warning-number>" /root/Git.d/pytorch/pytorch/torch/csrc/distributed/c10d/ProcessGroupNCCL.hpp(451): warning #68-D: integer conversion resulted in a change of sign size_t numelOut_ = -1; ^ /root/Git.d/pytorch/pytorch/torch/csrc/distributed/c10d/ProcessGroupNCCL.hpp(450): warning #68-D: integer conversion resulted in a change of sign size_t numelIn_ = -1; ^ Remark: The warnings can be suppressed with "-diag-suppress <warning-number>" /root/Git.d/pytorch/pytorch/torch/csrc/distributed/c10d/ProcessGroupNCCL.hpp(451): warning #68-D: integer conversion resulted in a change of sign size_t numelOut_ = -1; ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157889 Approved by: https://github.com/mlazos	2025-07-09 11:41:02 +00:00
Xuehai Pan	4dce5b71a0	[build] modernize build-frontend: `python setup.py develop/install` -> `[uv ]pip install --no-build-isolation [-e ].` (#156027 ) Modernize the development installation: ```bash # python setup.py develop python -m pip install --no-build-isolation -e . # python setup.py install python -m pip install --no-build-isolation . ``` Now, the `python setup.py develop` is a wrapper around `python -m pip install -e .` since `setuptools>=80.0`: - pypa/setuptools#4955 `python setup.py install` is deprecated and will emit a warning during run. The warning will become an error on October 31, 2025. - `9c4d383631/setuptools/command/install.py (L58-L67)` > ```python > SetuptoolsDeprecationWarning.emit( > "setup.py install is deprecated.", > """ > Please avoid running ``setup.py`` directly. > Instead, use pypa/build, pypa/installer or other > standards-based tools. > """, > see_url="https://blog.ganssle.io/articles/2021/10/setup-py-deprecated.html", > due_date=(2025, 10, 31), > ) > ``` - pypa/setuptools#3849 Additional Resource: - [Why you shouldn't invoke setup.py directly](https://blog.ganssle.io/articles/2021/10/setup-py-deprecated.html) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156027 Approved by: https://github.com/ezyang	2025-07-09 11:24:27 +00:00
Xuehai Pan	fc0376e8b1	[BE][2/6] fix typos in test/ (test/test_*.py) (#157636 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157636 Approved by: https://github.com/yewentao256, https://github.com/mlazos ghstack dependencies: #156311, #156609	2025-07-09 11:02:23 +00:00
Xuehai Pan	ffe11b2bf2	[BE] fix typo in torch/distributed/tensor/: childs -> children (#156609 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156609 Approved by: https://github.com/wanchaol, https://github.com/cyyever ghstack dependencies: #156311	2025-07-09 11:02:23 +00:00
Xuehai Pan	4cc8b60d1b	[BE][1/16] fix typos in torch/ (#156311 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156311 Approved by: https://github.com/albanD	2025-07-09 11:02:22 +00:00
Syed Tousif Ahmed	f5bbaa2253	Fixes typo in nccl_window_registration test (#157293 ) As mentioned here: https://github.com/pytorch/pytorch/pull/155134#discussion_r2175605192 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157293 Approved by: https://github.com/Skylion007	2025-07-09 11:01:18 +00:00
Xuehai Pan	924fc52e18	[BE] add a linter to check consistency for cmake minimum version in requirements (#156961 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156961 Approved by: https://github.com/ezyang, https://github.com/malfet	2025-07-09 10:44:17 +00:00
PyTorch MergeBot	b83d8827bc	Revert "Deprecate DataLoader pin_memory_device param (#146821 )" This reverts commit ab655816b8f76f511fb2262d45276d8d1b13d59c. Reverted https://github.com/pytorch/pytorch/pull/146821 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/146821#issuecomment-3052093902))	2025-07-09 10:29:31 +00:00
thenumberouscode	6f23f53599	[inductor] fix tensor.to(uint8) error when tensor src type is float (#157267 ) The cpu inductor processes .to(torch.uint8) incorrectly, leading to numerical inconsistencies. The convert_float_to_int8 function may return incorrect results for negative inputs, such as -2.xx, when the data type is uint8_t, producing 0 instead of 255. This issue stems from the clamping logic; we should avoid converting min_val to uint8_t too early Fixes https://github.com/pytorch/pytorch/issues/156788 @leslie-fang-intel Pull Request resolved: https://github.com/pytorch/pytorch/pull/157267 Approved by: https://github.com/leslie-fang-intel	2025-07-09 07:03:38 +00:00
Menglu Yu	e3f2597b45	[Optimus] Fix normalization pass in the aten IR (#157857 ) Summary: We found there's a special case in recent APS model where the input tensor has smaller size compared to the split size. It will be automatically truncated in split.Tensor thus we add extra condition check for split_with_sizes when do the normalization. Test Plan: ### unit ``` buck2 test 'fbcode//mode/dev-nosan' fbcode//caffe2/test/inductor:split_cat_fx_aten_passes -- test_split_aten_normalization ``` Buck UI: https://www.internalfb.com/buck2/2ecd1ef8-8efe-4245-b4c8-282c23645b3c Test UI: https://www.internalfb.com/intern/testinfra/testrun/7599824648585787 Network: Up: 3.9GiB Down: 9.2GiB (reSessionID-1396c91e-0dd2-457b-a49b-a6ab1f2a7d8f) Loading targets. Remaining 0/5344 99617 dirs read, 1074949 targets declared Analyzing targets. Remaining 0/123279 4988547 actions, 5966764 artifacts declared Executing actions. Remaining 0/728058 209:52:59.9s exec time total Command: test. Finished 12466 local, 209448 remote, 1226 cache (1% hit) 42:10.5s exec time cached (0%) Time elapsed: 26:07.6s Tests finished: Pass 2. Fail 0. Fatal 0. Skip 0. Build failure 0 ### E2E before fix: aps-afoc_apop_pt2_v0-db2fe0449a after fix: aps-afoc_apop_pt2_v0-755ad0cdc6 Rollback Plan: Differential Revision: D77961394 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157857 Approved by: https://github.com/anijain2305	2025-07-09 05:38:15 +00:00
Shangdi Yu	effe376db0	Adding aoti_standalone config (#157731 ) Summary: When `compile_standalone` is True, we set `package_cpp_only` to True as well. We raise an error if `package_cpp_only` is explicitly set to False in config. Test Plan: ``` buck2 run mode/dev-nosan fbcode//caffe2/test/inductor:test_aot_inductor -- -r TestAOTInductorConfig ``` Rollback Plan: Differential Revision: D77889754 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157731 Approved by: https://github.com/desertfire	2025-07-09 04:30:04 +00:00
Xu Han	fcbf7c749a	[Windows][Inductor] normalize_path_separator compiler path (#157835 ) Fixes #157673 For the call trace: ``` ...... File "D:\Programs\Python\virtualenvs\torch_code-afvE469o\lib\site-packages\torch\_inductor\codegen\common.py", line 2569, in reduction return self.kernel.reduction(dtype, src_dtype, reduction_type, value) File "D:\Programs\Python\virtualenvs\torch_code-afvE469o\lib\site-packages\torch\_inductor\codegen\cpp.py", line 2155, in reduction self._gen_parallel_reduction_buffers(acc, acc_type, reduction_type, init_dtype) File "D:\Programs\Python\virtualenvs\torch_code-afvE469o\lib\site-packages\torch\_inductor\codegen\cpp.py", line 1942, in _gen_parallel_reduction_buffers reduction_prefix_array( File "D:\Programs\Python\virtualenvs\torch_code-afvE469o\lib\site-packages\torch\_inductor\codegen\cpp.py", line 335, in reduction_prefix_array if cpp_builder.is_msvc_cl() File "D:\Programs\Python\virtualenvs\torch_code-afvE469o\lib\site-packages\torch\_inductor\cpp_builder.py", line 317, in is_msvc_cl return _is_msvc_cl(get_cpp_compiler()) File "D:\Programs\Python\virtualenvs\torch_code-afvE469o\lib\site-packages\torch\_inductor\cpp_builder.py", line 240, in _is_msvc_cl subprocess.check_output([cpp_compiler, "/help"], stderr=subprocess.STDOUT) torch._inductor.exc.InductorError: UnicodeDecodeError: 'utf-8' codec can't decode byte 0xd3 in position 0: invalid continuation byte ``` On non-English language pack msvc environment, compiler path has raised `utf-8` issue. I add the `normalize_path_separator` to normalize the compiler path and avoid the issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157835 Approved by: https://github.com/jansel	2025-07-09 04:02:20 +00:00
soulitzer	8bda95228f	[autograd] Avoid creating and recording event when unnecessary (#157503 ) Today, we always create and record an events in two places: 1) Upon seeing the first producer, we record an event on the producer, and we wait for this event in two places: (1) when the engine goes to run the consumer, the consumer stream waits for this event. (2) prior to doing accumulation, the accumulation stream waits for this event. 2) After doing accumulation, we record an event on the accumulation stream and wait for this event in a single place: when the engine goes to run the consumer. We do not actually need to record the event in the cases where the 1st producer stream is the same as the consumer and as the accumulation stream, and where the accumulation stream is the same as the consumer stream. Removing this unnecessary create + record event should save a few us for each instance avoided. Fixes https://github.com/pytorch/pytorch/issues/157407 ---- Manual test plan: - [x] @eqy to confirm perf is restored - [x] Running the repro originally reported before/after the patch Pull Request resolved: https://github.com/pytorch/pytorch/pull/157503 Approved by: https://github.com/eqy ghstack dependencies: #155715	2025-07-09 03:36:14 +00:00
florian	8d070187e3	fix type hints for interpolation functions (#157202 ) Fixes #129053 Previously interpolate had a bad signature and not correct type hints. This fixes this issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157202 Approved by: https://github.com/ezyang, https://github.com/albanD	2025-07-09 03:11:37 +00:00
Jing Xu	c515385b0a	Add Intel GPU info collection to the collect env script (#157351 ) https://github.com/pytorch/pytorch/pull/137846 was mistakenly closed. Reopen a PR to land the PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157351 Approved by: https://github.com/guangyey, https://github.com/malfet	2025-07-09 03:01:41 +00:00
Nikita Shulga	d6237721c0	[Build] Make PyTorch compilable with gcc-14 on ARM (#157867 ) Fixes numerous ICEs in vreg allocations for SVE+BF16 ``` /pytorch/aten/src/ATen/ParallelOpenMP.h:25:9: error: unrecognizable insn: 25 \| #pragma omp parallel \| ^~~ (insn 257 256 258 30 (set (reg:VNx8BF 449 [ bf16_vec1_217 ]) (unspec:VNx8BF [ (reg:VNx8BF 455) (reg:VNx8BF 456) ] UNSPEC_IORF)) "/pytorch/aten/src/ATen/cpu/vec/sve/vec_bfloat16.h":228:31 discrim 1 -1 (nil)) during RTL pass: vregs /pytorch/aten/src/ATen/ParallelOpenMP.h:25:9: internal compiler error: in extract_insn, at recog.cc:2812 0xd73c33 internal_error(char const, ...) ???:0 0xd73d1f fancy_abort(char const, int, char const) ???:0 0x890053 _fatal_insn(char const, rtx_def const, char const, int, char const) ???:0 0x890087 _fatal_insn_not_found(rtx_def const, char const, int, char const) ???:0 0x1379093 extract_insn(rtx_insn) ???:0 ``` And one in RTL-expand pass while compiling Activation.cpp ``` during RTL pass: expand In file included from /pytorch/aten/src/ATen/native/cpu/Activation.cpp:12, from /pytorch/build/aten/src/ATen/native/cpu/Activation.cpp.DEFAULT.cpp:1: /pytorch/aten/src/ATen/native/cpu/Activation.cpp: In lambda function: /pytorch/aten/src/ATen/native/cpu/Activation.cpp:94:7: internal compiler error: Segmentation fault 94 \| }); \| ^ /pytorch/aten/src/ATen/Dispatch.h:201:7: note: in definition of macro 'AT_DISPATCH_SWITCH' 201 \| __VA_ARGS__ \ \| ^~~~~~~~~~~ /pytorch/aten/src/ATen/Dispatch.h:72:3: note: in expansion of macro 'AT_PRIVATE_CASE_TYPE_USING_HINT' 72 \| AT_PRIVATE_CASE_TYPE_USING_HINT(enum_type, scalar_t, __VA_ARGS__) \| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ /pytorch/aten/src/ATen/Dispatch.h:214:3: note: in expansion of macro 'AT_DISPATCH_CASE' 214 \| AT_DISPATCH_CASE(at::ScalarType::Double, __VA_ARGS__) \ \| ^~~~~~~~~~~~~~~~ /pytorch/aten/src/ATen/Dispatch.h:218:34: note: in expansion of macro 'AT_DISPATCH_CASE_FLOATING_TYPES' 218 \| AT_DISPATCH_SWITCH(TYPE, NAME, AT_DISPATCH_CASE_FLOATING_TYPES(__VA_ARGS__)) \| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ /pytorch/aten/src/ATen/native/cpu/Activation.cpp:70:5: note: in expansion of macro 'AT_DISPATCH_FLOATING_TYPES' 70 \| AT_DISPATCH_FLOATING_TYPES(input.scalar_type(), "log_sigmoid_cpu", [&] { \| ^~~~~~~~~~~~~~~~~~~~~~~~~~ 0xd73c33 internal_error(char const, ...) ???:0 0x134f987 rebuild_jump_labels(rtx_insn) ???:0 ``` Interestingly enough, attempt to compile `Unfold2d.cpp` for `-march=armv8-a+sve` (i.e. without sve+bf16) support also causes ICE ``` /pytorch/aten/src/ATen/native/cpu/Unfold2d.cpp:221:1: error: unrecognizable insn: 221 \| } \| ^ (insn 2918 2917 2919 296 (set (reg:VNx8BI 5917) (unspec:VNx16BI [ (reg:VNx8BI 5920) (reg:VNx8BI 5922) (const_vector:VNx4BI [ (const_int 0 [0]) repeated x8 ]) ] UNSPEC_TRN1_CONV)) "/usr/include/aarch64-linux-gnu/bits/string_fortified.h":29:33 discrim 1 -1 (expr_list:REG_EQUAL (const_vector:VNx8BI [ (const_int 1 [0x1]) repeated x9 (const_int 0 [0]) (const_int 1 [0x1]) repeated x2 (const_int 0 [0]) repeated x4 ]) (nil))) during RTL pass: vregs ``` Which could be worked around by adding ```patch diff --git a/aten/src/ATen/native/cpu/Unfold2d.cpp b/aten/src/ATen/native/cpu/Unfold2d.cpp index 8ef0741e77af0a..59c76505dd6246 100644 --- a/aten/src/ATen/native/cpu/Unfold2d.cpp +++ b/aten/src/ATen/native/cpu/Unfold2d.cpp @@ -169,6 +169,10 @@ static void unfolded2d_acc_channels_last( / note: due to write issues, this one cannot be parallelized as well as * unfolded2d_copy / +#if defined(__GNUC__) && __GNUC__ == 14 && defined(__ARM_FEATURE_SVE) +// Workaround for gcc-14.2.0 ICE during RTL pass: vregs when compiling for SVE +__attribute__((optimize("no-tree-vectorize"))) +#endif void unfolded2d_acc_kernel( ScalarType dtype, void finput_data, ``` Fixes https://github.com/pytorch/pytorch/issues/157842 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157867 Approved by: https://github.com/atalman, https://github.com/Skylion007	2025-07-09 02:59:08 +00:00
Hao Zhang(张浩)	ab8874bd26	Suppress warning when using native arch for jit loading cuda extensions. (#156923 ) Previeusly, if users want to let pytorch determine the cuda arch when jit loading cuda extensions, they should left environment variable `TORCH_CUDA_ARCH_LIST` empty, but which will raise an warning. This commit add an option to set `TORCH_CUDA_ARCH_LIST=native`, to tell pytorch users want to use native cuda arch intentionally. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156923 Approved by: https://github.com/ezyang	2025-07-09 02:51:20 +00:00
fduwjj	bc6e0661a6	Fix more H100 CI (#157829 ) Follow @d4l3k 's fix in https://github.com/pytorch/pytorch/pull/157826/files. Two more fixes might be needed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157829 Approved by: https://github.com/davidberard98, https://github.com/d4l3k	2025-07-09 01:28:05 +00:00
Bin Bao	e5edd013ab	[AOTI] Skip test_simple_multi_arch_embed_kernel_binary_True_cuda (#157301 ) Summary: For https://github.com/pytorch/pytorch/issues/156930, still no clue on what went wrong as it is not reproducible locally, but somehow the problem seems only exists when embed_kernel_binary is True. Let's skip it for now. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157301 Approved by: https://github.com/yushangdi	2025-07-09 01:18:36 +00:00
xinan.lin	75f489d37f	[Break XPU][Inductor UT] Align tolerance of newly added case with cuda. (#157702 ) Align tolerance with cuda for the newly added case `test_comprehensive_logcumsumexp_xpu_float16` in #157512. Fixes #157697 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157702 Approved by: https://github.com/jansel	2025-07-09 00:55:01 +00:00
Tristan Rice	3eb7084f7a	[ci] fix h100-distributed (#157826 ) This was broken by https://github.com/pytorch/pytorch/pull/157341 This should resolve the permission issue Pull Request resolved: https://github.com/pytorch/pytorch/pull/157826 Approved by: https://github.com/fduwjj, https://github.com/Skylion007, https://github.com/huydhn	2025-07-09 00:27:55 +00:00
PyTorch MergeBot	86251eff40	Revert "Introduce AcceleratorAllocatorConfig as the common class (#149601 )" This reverts commit 55108074c0795be3b617d3b13b06794f63e1f8ca. Reverted https://github.com/pytorch/pytorch/pull/149601 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/149601#issuecomment-3050628047))	2025-07-09 00:07:31 +00:00
Tristan Rice	1b3d69b59f	Work: block_current_stream API (#156883 ) This implements a new `wait_stream` API in Work that matches how `wait` works for ProcessGroupNCCL for CPU based backends such as Gloo. The idea is to support Gloo communication overlap in FSDPv2/HSDP with minimal changes to FSDP. There was a previous attempt to make FSDPv2 use Work.wait but given the extensive stream semantics used it doesn't play nicely. https://github.com/pytorch/pytorch/pull/148780 This uses a "Baton" CUDA kernel which spinlocks on a pinned CPU tensor waiting for it to be set. Test plan: ``` pytest test/distributed/test_c10d_gloo.py -v -k wait_stream pytest test/distributed/test_c10d_nccl.py -v -k wait_stream ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156883 Approved by: https://github.com/kwen2501, https://github.com/fduwjj	2025-07-08 23:55:46 +00:00
Blaine Burton Rister	92f41ccc26	[Inductor] Support precomputed size args in the FX backend. (#157758 ) # Feature If a Triton kernel has a complicated indexing expression, Inductor may decide to precompute it on the host and pass it to the kernel as an argument. This happens in situations like broadcasts with dynamic shapes. This PR adds support for this feature to Inductor's FX IR backend. We generate FX IR for precomputed size args in 3 steps: 1. In `PythonWrapperCodegen`, this PR refactors the relevant code to use a `SymbolicCallArgLine` instead of raw Python strings. This stores a (symbol, expr) pair. (Prior to this PR, it was (str, expr), but changing this to a symbol makes it easier to do substitutions later on.) 2. In `WrapperFxCodegen`, keep a dict of {symbol: expr} arg defs which gets updated whenever we see a `SymbolicCallArgLine`. 3. When the FX backend sees a `KernelCallLine`, it uses this dict to replace symbolic call args with their definitions. In the longer run, it might be desirable to emit FX nodes defining these symbolic call args. That way, we could reuse the size computation when the same kernel is called multiple times. However, I wasn't sure if there was an existing way to generate FX nodes from a sympy expression, and implementing that seemed like overkill for the present purposes. # Test plan Added a new CI test exercising this feature. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157758 Approved by: https://github.com/jansel	2025-07-08 23:22:17 +00:00
Simon Fan	95bc3da9f8	[c10d] support dynamic shapes for all_to_all_single_autograd (#157521 ) `all_to_all_single_autograd` is not an op, all the code executed until the `all_to_all_single` dispatch is visible to the compiler. This means the `all_to_all_single_autograd` wrapper code must support symints in order to be traceable with dynamic shapes. FIXES https://github.com/pytorch/pytorch/issues/157479 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157521 Approved by: https://github.com/wconstab	2025-07-08 23:19:59 +00:00
Sidharth	9f18482d41	[dynamo] removing string literals for weblink generation (#157820 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157820 Approved by: https://github.com/williamwen42	2025-07-08 23:08:06 +00:00
Nikita Shulga	c5b46b5408	[BE] Standardize CPU capabilities name (#157809 ) It's weird to call default x86 CPU capability `NO AVX`, when in reality it's something different. Also it's a bit strange to have it assigned different names on different platforms Fixes https://github.com/pytorch/pytorch/issues/157538 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157809 Approved by: https://github.com/Skylion007	2025-07-08 23:06:09 +00:00
Andrey Talman	179dcc10e4	Add sm_70 arch for linux cuda 12.8 and 12.9 builds (#157558 ) Please see: https://github.com/pytorch/pytorch/issues/157517 We would like to keep Volta architectures by default for release 2.8 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157558 Approved by: https://github.com/Skylion007, https://github.com/Camyll, https://github.com/seemethere, https://github.com/malfet	2025-07-08 23:02:10 +00:00
Sam Larsen	7a41f20794	[inductor] Quiesce Triton compile worker pool after each dynamo compile (#156187 ) For internal usages, keeping the Triton compile worker pool active for the lifetime of the process has caused some challenges, e.g., it slows down and muddies profiling due to the huge number of threads on a box: N threads = 8 ranks * 32 subprocs * M threads started by torch. Also, each subproc can use more than 1GB each. This PR adds the functionality to shutdown worker subprocs after each dynamo compile when using the SubprocPool implementation. The idea is to leave the main sidecar process running, but signal it to tear down its internal ProcessPoolExecutor when compile is finished. Restarting the ProcessPoolExecutor is relatively fast, e.g., 500ms because the ProcessPoolExecutor forks from the sidecar. Changes: * Do not start the ProcessPoolExecutor automatically when compile_fx is imported. Instead, start the sidecar process only. The sidecar process imports torch, so is still slow to start. * Introduce wakeup() and quiesce() calls to the implementation to start and stop the ProcessPoolExecutor. * Add a context manager to automatically quiesce() at the end of dynamo compilation. * Signal a wakeup() in compile_fx only when we have cuda devices. * Add a killswitch so we can turn of quiescing. Testing: For correctness, the stacked change at https://github.com/pytorch/pytorch/pull/156534 enables the feature for OSS so it's exercised in CI. For performance, because of recent compile-time variance (see https://github.com/pytorch/pytorch/issues/152566), it's pretty hard to glean whether there's a regression.... * Training: https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Tue%2C%2017%20Jun%202025%2021%3A32%3A04%20GMT&stopTime=Tue%2C%2024%20Jun%202025%2021%3A32%3A04%20GMT&granularity=hour&mode=training&dtype=amp&deviceName=cuda%20(h100)&lBranch=gh/masnesral/210/head&lCommit=1b7315031c3bfad66a1a01700167a9ca1a2ae5f1&rBranch=main&rCommit=eab45643f22e58ee12d95d8b0162d51ca0a50801 * Inference: https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Tue%2C%2017%20Jun%202025%2021%3A32%3A04%20GMT&stopTime=Tue%2C%2024%20Jun%202025%2021%3A32%3A04%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cuda%20(h100)&lBranch=gh/masnesral/210/head&lCommit=1b7315031c3bfad66a1a01700167a9ca1a2ae5f1&rBranch=main&rCommit=eab45643f22e58ee12d95d8b0162d51ca0a50801 The wins (mostly for inference) don't make sense, but I'm also skeptical of the losses (mostly for training). I can't repro any of the slowdowns locally. Furthermore, check out the benchmarking results for the stacked diff, which actually enables the quiescing functionality for OSS. That should only slow down compile since there can only be overhead to stop and start the workers. But the results are somehow better: * Training: https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Tue%2C%2017%20Jun%202025%2021%3A32%3A04%20GMT&stopTime=Tue%2C%2024%20Jun%202025%2021%3A32%3A04%20GMT&granularity=hour&mode=training&dtype=amp&deviceName=cuda%20(h100)&lBranch=gh/masnesral/214/head&lCommit=41943253882a019b8ceafcd2bf4cd6acbe0cbca9&rBranch=main&rCommit=eab45643f22e58ee12d95d8b0162d51ca0a50801 * Inference: https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Tue%2C%2017%20Jun%202025%2021%3A32%3A04%20GMT&stopTime=Tue%2C%2024%20Jun%202025%2021%3A32%3A04%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cuda%20(h100)&lBranch=gh/masnesral/214/head&lCommit=41943253882a019b8ceafcd2bf4cd6acbe0cbca9&rBranch=main&rCommit=eab45643f22e58ee12d95d8b0162d51ca0a50801 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156187 Approved by: https://github.com/aorenste, https://github.com/jansel	2025-07-08 22:53:13 +00:00
Animesh Jain	178fe7aa98	[dynamo][fsdp] Consistent behavior of int attributes (#157262 ) Reimpl of https://github.com/pytorch/pytorch/pull/150954 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157262 Approved by: https://github.com/bdhirsh	2025-07-08 22:11:33 +00:00
PyTorch MergeBot	2e14069081	Revert "[DTensor][FSDP2] necessary changes to FSDP and TP to unblock EP (#157216 )" This reverts commit 777eca9f16aeecd7c362a235cf25e6b8e6eda57f. Reverted https://github.com/pytorch/pytorch/pull/157216 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it seems to fail a distributed test in trunk ([comment](https://github.com/pytorch/pytorch/pull/157216#issuecomment-3050258896))	2025-07-08 20:48:51 +00:00
angelayi	391473cca0	[export] Fix lift constants bug (#157719 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/157719 Approved by: https://github.com/yushangdi	2025-07-08 20:33:53 +00:00
lucasb-eyer	b9dc2fa4f7	Add legacy note to autograd.profiler doc. (#157459 ) Via google search I got to `torch.autograd.profiler` and implemented my code with it. Only to be taken by surprise finding `torch.profile.profiler`, which has a note saying the autograd one is legacy. This just adds such note to `autograd.profiler` to avoid this confusion and waste of time to future people in my situation. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157459 Approved by: https://github.com/sraikund16	2025-07-08 20:33:23 +00:00
zpcore	a73d9e0aec	Fix einsum strategy shard dim > ndim (#157593 ) Previously we didn't constrain Shard dim to be <= the tensor's ndim. This cause an invalid strategy like `(RR, RS(2)) -> RS(2),` for einsum `bmk,kn->bmn` on the 2d mesh. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157593 Approved by: https://github.com/wconstab, https://github.com/wanchaol	2025-07-08 20:27:17 +00:00
Huy Do	06b3265cb1	Increase nightly C++ docs build timeout to 6h (#157759 ) This job has been timing out since May `261897734a/1`, maybe it's time to figure out if this makes sense. Issues https://github.com/pytorch/pytorch/issues/157763 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157759 Approved by: https://github.com/malfet	2025-07-08 19:28:48 +00:00
Ankita George	dea4864ce0	HF loads dcp - don't do a full deserialize on every file (#157715 ) Summary: These changes in D76442012 got reverted after the PR landed due to aps_models/ads/launchers/pearl/tests/ne/e2e_deterministic_tests:pearl_e2e_ne_tests failing with `Config not loaded due to no timely response from configerator. Likely configerator_proxy or falcon_proxy are not healthy`, but that test failing is definitely transient and unrelated to my changes, so re-creating the diff Test Plan: ensure tests pass Rollback Plan: Differential Revision: D77871099 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157715 Approved by: https://github.com/meetv18	2025-07-08 18:13:27 +00:00
Zeina Migeed	4f5be56612	[Pyrefly][Refactor] Replace dict() calls with literal dict syntax for improved readability (#157735 ) There are 31 places that I spotted which construct literal dictionaries. This PR refactors dictionary construction by replacing` dict(...) `calls with `literal {...}` syntax where applicable. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157735 Approved by: https://github.com/ezyang, https://github.com/Skylion007	2025-07-08 18:10:33 +00:00
Howard Huang	0f31445139	Add stack trace of exception to MultiProcContinousTest (#157589 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157589 Approved by: https://github.com/Skylion007	2025-07-08 17:54:35 +00:00
Shangdi Yu	5b4e0255d7	Check FakeScriptObject in _resolve_name_collision (#157736 ) Summary: Fix https://github.com/pytorch/pytorch/issues/157401 torch.equal cannot handle FakeScriptObject inputs. Test Plan: ``` buck run fbcode//mode/dev-nosan //caffe2/test/inductor:torchbind -- -r test_aoti_torchbind_name_collision ``` Rollback Plan: Differential Revision: D77894081 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157736 Approved by: https://github.com/angelayi	2025-07-08 17:51:46 +00:00
Weishi.Deng	44d0800d60	[Intel GPU] Set higher tolerance for squeezenet1_1 with bf16 (#156920 ) We need to increase the tolerance slightly to ensure that certain models pass the accuracy check on the XPU device. This pull request preserves the original tolerance threshold for CUDA/CPU devices and introduces a new key, higher_bf16_xpu, which only affects the XPU device. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156920 Approved by: https://github.com/soulitzer	2025-07-08 17:49:54 +00:00
Nikita Shulga	a5c61eb78d	[MPS][BE] Delete `as_strided_tensorimpl_mps` (#157772 ) Because it's just copy-n-paste of `as_strided_tensorimpl` with call to `updateTensorBaseShape`, which is not called/used anywhere else. Fixes https://github.com/pytorch/pytorch/issues/152701 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157772 Approved by: https://github.com/Skylion007	2025-07-08 17:02:36 +00:00
henrylhtsang	bbe681ed51	[cutlass backend][BE][ez] Make matmul layouts be row x column (#156656 ) Differential Revision: [D77184232](https://our.internmc.facebook.com/intern/diff/D77184232/) Motivation: * This is the case we care the most. * We are caching the kernels for this row x column layout. So testing on them can potentially make ci run faster. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156656 Approved by: https://github.com/ColinPeppler	2025-07-08 16:57:33 +00:00
Tianyu Liu	ed911747c2	[dtensor] add support for fused optimizer with parameters across multiple meshes (#157682 ) We are seeing more and more use cases where parameters in a model (under the same optimizer group) are put on different meshes. E.g. - when FSDP and TP are both applied, some parameters are sharded only on the FSDP mesh but not TP mesh (see https://github.com/pytorch/pytorch/pull/153268). - in [dp2ep Expert Parallel](https://github.com/pytorch/torchtitan/pull/1324), the routed experts are sharded on the (global FSDP \ EP) mesh for smaller FSDP and on the EP mesh for EP, whereas other params are sharded on the global FSDP mesh for FSDP. This PR is, in some sense, a continuation of https://github.com/pytorch/pytorch/pull/147869 to tackle the problem when fused optimizers are used. In such cases, the [`fused_adam`](https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/native/native_functions.yaml#L15786) / `fused_adamw` has a scalar tensor arg `state_steps` which gets automatically cast to DTensor on the default [`compute_mesh`](https://github.com/pytorch/pytorch/blob/main/torch/distributed/tensor/_dispatch.py#L350) (one of the multiple meshes), even though the it could correspond to different meshes. To avoid hitting the cross-mesh propagation exception in `common_pointwise_strategy` and followup redistribute problems, we manually set the target mesh and placements to be the same as input mesh and placements, so that no redistribute will be triggered. This also helps bypass the situation where [`generate_redistribute_costs`](https://github.com/pytorch/pytorch/pull/157682/files#diff-eea32a36dd2d4e58307bc5229402e48048b2ecaef64a7c085495fba1ee10ac89R597) returns infinite cost due to cross mesh redistribute. Moreover, this PR has minimal scope (restricted to the `fused_ops`) and doesn't need to modify other files such as `_sharding_prop.py`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157682 Approved by: https://github.com/wanchaol	2025-07-08 15:58:30 +00:00
Tianyu Liu	777eca9f16	[DTensor][FSDP2] necessary changes to FSDP and TP to unblock EP (#157216 ) This is to unblock "dp2ep" Expert Parallel + TP integration in torchtitan https://github.com/pytorch/torchtitan/pull/1324. It does two things: 1. Slightly modifies the glue code for FSDP/HSDP + TP to work with FSDP/HSDP + EP and FSDP/HSDP + EP + TP. I kept the name `FSDPParam._tp_spec` to make the change minimal. We can consider renaming it in the future if it confuses people, but I heard @wanchaol has a plan to rewrite DTensor strided sharding entirely. 2. Lifts the check of `_validate_tp_mesh_dim` for `torch.distributed.tensor.parallel.parallelize_module`, as in EP or EP+TP this check is too strict. In particular it assumes a DeviceMesh must have `mesh_dim_names` which is not always true. I'm also removing the file `torch/distributed/tensor/parallel/_utils.py` it belongs entirely, as the other check `_deprecate_warnings`, added two years ago, is not used any more. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157216 Approved by: https://github.com/wanchaol, https://github.com/weifengpy	2025-07-08 15:57:37 +00:00
Aaron Gokaslan	476874b37f	[BE]: Update NCCL to 2.27.5 (#157108 ) Update NCCL to 2.27.5. Minor version, improves Blackwell, Symmem FP8 support, and fixes a bug with MNVVL. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157108 Approved by: https://github.com/atalman	2025-07-08 15:40:54 +00:00
Steven Troxler	5dc75f72d4	Simplify the base classes of `_PyFutureMeta` (#157757 ) Summary: I'm fairly sure the use of a custom metaclass is a holdover from pre-3.7 where Generic used a custom metaclass so we had to use multiple inheritance to avoid import-time failures. At this point, `type(Generic)` is just `type` so it isn't needed, and we will get the least metaclass from our base classes, which means the `type(torch._C.Future)` isn't needed either, it will happen automatically just by inheritance. Test Plan: I'm fairly confident from local testing that this should be a no-op. But also, Pytorch CI should give us pretty strong signal that this change doesn't break anything in case there's some edge case I missed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157757 Approved by: https://github.com/ezyang, https://github.com/Skylion007	2025-07-08 15:39:56 +00:00
Nikita Shulga	f88d7a7a34	[BE] Do not add `.` after troubleshooting_url (#157753 ) As it gets included into auto-hrefed URLs in say github logs to point to non existing location For example from https://github.com/pytorch/pytorch/actions/runs/16130448756/job/45517004735?pr=157749#step:18:27 > W0708 00:23:20.150000 67082 torch/_dynamo/convert_frame.py:1047] [0/8] To diagnose recompilation issues, see [https://pytorch.org/docs/main/torch.compiler_troubleshooting.html.](https://pytorch.org/docs/main/torch.compiler_troubleshooting.html.) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157753 Approved by: https://github.com/zou3519, https://github.com/jansel	2025-07-08 15:38:24 +00:00
Nikita Shulga	98bb0c0e78	[CI][MacOS] Add `VENV_PATH` to search path (#157749 ) When building/testing PyTorch on MacOS Shoudl prevent some flakiness when conda environment overtakes CI/CD Pull Request resolved: https://github.com/pytorch/pytorch/pull/157749 Approved by: https://github.com/atalman, https://github.com/huydhn	2025-07-08 15:37:45 +00:00
PyTorch MergeBot	76fe88fa56	Revert "Cleanup leftover miniconda brew installation (#156898 )" This reverts commit 214e2959dcdbf91a999d5c0a5d40c91e4442e8c5. Reverted https://github.com/pytorch/pytorch/pull/156898 on behalf of https://github.com/malfet due to Breaks TorchVision builds ([comment](https://github.com/pytorch/pytorch/pull/156898#issuecomment-3049281232))	2025-07-08 14:54:42 +00:00
Xuan Zhang	86670b39fa	[PT2][memory] mutation size correctness (#157562 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157562 Approved by: https://github.com/yf225	2025-07-08 14:02:20 +00:00
Wang, Chuanqi	c78bbdf410	[BE] Update xpu driver repo for CD used almalinux 8.10 (#157356 ) XPU CD docker image built on `quay.io/pypa/manylinux_2_28_x86_64`, which based on almalinux 8.10 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157356 Approved by: https://github.com/EikanWang, https://github.com/malfet	2025-07-08 13:59:46 +00:00
rzou	b9afdd9bcc	Add flag to fx.passes.split_module to normalize input names (#157733 ) This is useful for vLLM, which runs AOTAutograd directly on graphs after they have been split. I created a new flag for this instead of reusing `keep_original_node_name` (please let me know if you think I should reuse this). The reasoning is: - The names of the placeholder nodes is different from the targets of the placehoder nodes. The targets are the actual input names. - Backwards compatibility: this API has been out for ~4 years, it looks public, and it has extensive public use. For example, this change would actually be BC-breaking to vLLM (they rely on the subgraph input names being different at the moment). Test Plan: - new tests Pull Request resolved: https://github.com/pytorch/pytorch/pull/157733 Approved by: https://github.com/ezyang	2025-07-08 13:47:24 +00:00
cyy	7381c77724	Use CMake wholearchive group (#156393 ) Use CMake wholearchive group to simplify code. It may also support more OSes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156393 Approved by: https://github.com/ezyang	2025-07-08 12:20:29 +00:00
zeshengzong	ab655816b8	Deprecate DataLoader pin_memory_device param (#146821 ) Following [ #131858 suggestion](https://github.com/pytorch/pytorch/pull/131858#pullrequestreview-2517760602) to optimize DataLoader code Pull Request resolved: https://github.com/pytorch/pytorch/pull/146821 Approved by: https://github.com/divyanshk Co-authored-by: Divyansh Khanna <divyanshkhanna09@gmail.com>	2025-07-08 09:24:53 +00:00
Aleksei Nikiforov	41e8b826d0	S390x update test marks (#157541 ) Update s390x test marks test_logs_out from test/dynamo/test_logging.py is updated and no longer fails on s390x. test_qengine from test/test_torch.py doesn't work on s390x: no QEngine is available. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157541 Approved by: https://github.com/huydhn	2025-07-08 09:08:33 +00:00
pralay	5430990bd7	Added philox based RNG context for HPU device in Dtensor scenarios (#156581 ) In this PR, we are enabling `HPU` device-specific function calls for random operations. These calls will manage the setting and unsetting of the `context of Random Number Generator`. While HPU devices typically utilize a `Mersenne-based RNG`, Dtensor-specific random operations employ an `offset-based (Philox) RNG tracker` which is specifically integrated with `CUDA` in scope. To integrate a similar offset-based RNG tracker within the `HPU backend`, a backend-specific device handle function is necessary to identify the execution context of these random operations. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156581 Approved by: https://github.com/jeromean, https://github.com/wanchaol	2025-07-08 08:50:24 +00:00
Yu, Guangye	55108074c0	Introduce AcceleratorAllocatorConfig as the common class (#149601 ) # Motivation This PR aims to generalize `AllocatorConfig` to be device-agnostic. Introduce the class `AcceleratorAllocatorConfig` to clarify its scope as a configuration manager for accelerator backends (e.g., CUDA, XPU). The another name `AllocatorConfig` is now reserved for a potential future base class that can unify configuration handling for both CPU and accelerator allocators, should similar requirements arise for the CPU path. # Design Rule ## Overall This class configures memory allocation for both device and host memory. A single `AcceleratorAllocatorConfig` instance is shared across all accelerator backends, such as CUDA and XPU, under the assumption that relevant environment variables apply uniformly to all accelerators. Device-specific configuration extensions are supported via hooks (see `registerDeviceConfigParserHook`). Introduce a new class `ConfigTokenizer` to help process the env variable config key-value pair ## Naming Convention: - Public API names in `AcceleratorAllocatorConfig` should be device-generic. - Members prefixed with `pinned_` are specific to the host/pinned allocator. - Environment variable names should be generic across backends. - Comma-separated key-value pairs in the format: `key:value`. Use square brackets `[]` for list values Example: `key1:123, key2:[val1,val2]` ## Environment Variables: - The default environment variable for configuration is `PYTORCH_ALLOC_CONF`. - For backward compatibility, `PYTORCH_CUDA_ALLOC_CONF` and `PYTORCH_HIP_ALLOC_CONF` are also supported with lower priority. Pull Request resolved: https://github.com/pytorch/pytorch/pull/149601 Approved by: https://github.com/albanD	2025-07-08 08:40:47 +00:00
Xuehai Pan	84b77ec128	[BE] add a minimal linter to check `pyproject.toml` consistency (#156017 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156017 Approved by: https://github.com/ezyang	2025-07-08 08:17:36 +00:00
IvanKobzarev	8134684d44	[inductor collectives] sink waits iterative (#157708 ) Differential Revision: [D77861763](https://our.internmc.facebook.com/intern/diff/D77861763) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157708 Approved by: https://github.com/wconstab ghstack dependencies: #157706	2025-07-08 07:17:10 +00:00
Huy Do	2af7c67e48	Mitigate some flaky tests in trunk (#157756 ) (not really fix these issues, but we should be able to close them. This also allows CI from the PR to test them) Fixes https://github.com/pytorch/pytorch/issues/156579 Fixes https://github.com/pytorch/pytorch/issues/156580 Fixes https://github.com/pytorch/pytorch/issues/126867 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157756 Approved by: https://github.com/clee2000	2025-07-08 07:07:11 +00:00
Jithun Nair	38757d94f1	Enable target-determination (TD) for ROCm CI (#156545 ) Target determination sorts the tests in a PR CI run based on heuristics about which tests are more relevant to the PR's changes. This can help provide faster CI signal as well as help alleviate capacity concerns as job durations should decrease due to catching failures earlier. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156545 Approved by: https://github.com/jeffdaily, https://github.com/clee2000	2025-07-08 06:27:40 +00:00
Yu, Guangye	1b58e7adab	fix storage use_count (#157694 ) # Motivation https://github.com/pytorch/pytorch/pull/155451 decoupled `torch._C._storage_Use_Count` from CUDA and introduced a corresponding unit test: `815545f2dd/test/test_torch.py (L257-L262)` However, this test fails when PyTorch is built with debug assertions enabled. @clee2000 disabled this UT in https://github.com/pytorch/pytorch/pull/156731. The root cause is that `_cdata` is obtained from an `intrusive_ptr`, not a `weak_intrusive_ptr`. As a result, calling `c10::weak_intrusive_ptr::use_count` on it triggers the internal assertion: `815545f2dd/c10/util/intrusive_ptr.h (L912-L917)` For example: ```python a = torch.randn(10, device=device) # refcount=1, weakcount=1 prev_cf = torch._C._storage_Use_Count(a.untyped_storage()._cdata) # violate the assertation ``` This violates the expected invariant inside `weak_intrusive_ptr::use_count`, which assumes the pointer was originally constructed from a valid `weak_intrusive_ptr`. Actually, `storage_impl` is obtained from an `intrusive_ptr`. `815545f2dd/torch/csrc/Module.cpp (L2105-L2109)` # Solution Use `c10::intrusive_ptr::use_count` instead. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157694 Approved by: https://github.com/albanD	2025-07-08 05:53:12 +00:00
Xuehai Pan	8186af5a26	[BE][Easy] set end-of-line for `.bat` file to CRLF in `.editorconfig` (#156032 ) See also: `54976bca10/.gitattributes (L1)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156032 Approved by: https://github.com/seemethere, https://github.com/ezyang	2025-07-08 05:40:57 +00:00
Xuehai Pan	bdacf08b86	[BE][Easy] add `.editorconfig` setting for C/C++/CUDA/ObjC (#157692 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157692 Approved by: https://github.com/ezyang	2025-07-08 05:37:15 +00:00
drisspg	987314aa96	Split batch-num-heads grid dim between y and z (#157745 ) for #157018 doesn't totally fix the problem but should help alot Pull Request resolved: https://github.com/pytorch/pytorch/pull/157745 Approved by: https://github.com/Chillee	2025-07-08 05:17:43 +00:00
Nikita Shulga	39a8f66d59	[BE] Use `simdgroup_size` constexpr (#157751 ) Instead of every shader defining it separately, move it to `c10/metal/common.h` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157751 Approved by: https://github.com/Skylion007, https://github.com/dcci ghstack dependencies: #157746	2025-07-08 03:46:20 +00:00
Nikita Shulga	0b73f7c871	[EZ][BE] Move array def to `c10/metal/common.h` (#157746 ) And use proper type aliasing instead of weird _ARRAY_NS Also use `uint64_t` instead of `ulong` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157746 Approved by: https://github.com/Skylion007, https://github.com/dcci	2025-07-08 03:46:20 +00:00
Avanish Tiwari	a4c7e7f983	[PowerPC]: Fixed build issue that occur because of datatype f8 enablement for onednn in qlinear and prepack (#157469 ) Getting the build issue because of enablement of data type fp8 for onednn in qlinear and qlinear_prepack file after this commit c2185dc4a5626848df37cad214b73d5ae7dd4f17 Currrently cpuinfo is disable for power system because of that it is giving below error. Error: ‘cpuinfo_has_x86_amx_int8’ was not declared in this scope Made a required changes and now build issue got fixed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157469 Approved by: https://github.com/malfet	2025-07-08 03:45:06 +00:00
cyy	3ee8828c87	[1/N] Don't use CUDA.cmake module (#157188 ) Small changes before removing CUDA.cmake. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157188 Approved by: https://github.com/ezyang	2025-07-08 03:05:35 +00:00
Valentine233	f56bfb3030	[CPU] Fix memory access for sbgemm bf16 (#156585 ) Fixes #156022. 1. The original dtype conversion overwrites the whole `n_ldc_` instead of `n_m_` with stride `ldc_`, causing the potential memory issue. 2. Fix the None value issue in attention backward UT, as the sbgemm bf16 could be used. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156585 Approved by: https://github.com/mingfeima, https://github.com/aditew01, https://github.com/ezyang	2025-07-08 02:36:28 +00:00
zpcore	12f9942b10	Fix slice op redistribute_cost compute (#157178 ) For slice op backward, my understanding is that the `redistribute_cost` attribute is incorrectly assigned to previous placement strategy: `0decd966af/torch/distributed/tensor/_ops/_tensor_ops.py (L399-L400)` The mistake is hard to be tested since we didn't enforce the `redistribute_cost` for `strategy.strategies` with size one: `2815ade9a8/torch/distributed/tensor/_sharding_prop.py (L491-L499)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157178 Approved by: https://github.com/XilunWu	2025-07-08 02:28:59 +00:00
Ke Wen	c5589074e6	[SymmMem] find_path does not search /usr/local/lib (#157695 ) This PR uses `find_library` to replace `find_path`. It also searches for NVSHMEM host lib and device lib separately. Tested against system install location: /usr/local/lib and /usr/local/include. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157695 Approved by: https://github.com/Skylion007 ghstack dependencies: #157513	2025-07-08 01:21:59 +00:00
PyTorch MergeBot	30a1cc11a4	Revert "[CI][MacOS] Add `VENV_PATH` to search path (#157749 )" This reverts commit 85111cd165f108ffabb4a90083d59d7a867ebd9f. Reverted https://github.com/pytorch/pytorch/pull/157749 on behalf of https://github.com/huydhn due to It looks like lint was not green, so revert and reland I guess ([comment](https://github.com/pytorch/pytorch/pull/157749#issuecomment-3047032909))	2025-07-08 01:18:16 +00:00
PyTorch MergeBot	19a01382bc	Revert "[SymmMem] find_path does not search /usr/local/lib (#157695 )" This reverts commit 3effe0c293219b00a0eae7e139fe2d9aed84bc03. Reverted https://github.com/pytorch/pytorch/pull/157695 on behalf of https://github.com/kwen2501 due to Changing it to be landable on 2.8 branch ([comment](https://github.com/pytorch/pytorch/pull/157695#issuecomment-3047020152))	2025-07-08 01:12:01 +00:00
zeshengzong	df72078fe1	[dynamo] Replace unimplemented with unimplemented_v2 in `torch/_dynamo/variables/torch.py` (#157344 ) Fixes part of #147913 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157344 Approved by: https://github.com/williamwen42 Co-authored-by: William Wen <william.wen42@gmail.com>	2025-07-08 00:46:56 +00:00
Nikita Shulga	85111cd165	[CI][MacOS] Add `VENV_PATH` to search path (#157749 ) When building/testing PyTorch on MacOS Shoudl prevent some flakiness when conda environment overtakes CI/CD Pull Request resolved: https://github.com/pytorch/pytorch/pull/157749 Approved by: https://github.com/atalman, https://github.com/huydhn	2025-07-08 00:38:37 +00:00
Aaron Orenstein	edf7bb4f51	Fix unbound local when an error occurs before pool is initialized (#156750 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156750 Approved by: https://github.com/jamesjwu	2025-07-08 00:28:21 +00:00
dependabot[bot]	bbb930aba2	Bump urllib3 from 2.2.2 to 2.5.0 in /tools/build/bazel (#156390 ) Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.2.2 to 2.5.0. - [Release notes](https://github.com/urllib3/urllib3/releases) - [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst) - [Commits](https://github.com/urllib3/urllib3/compare/2.2.2...2.5.0) --- updated-dependencies: - dependency-name: urllib3 dependency-version: 2.5.0 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2025-07-07 17:13:21 -07:00
Bob Ren	60b41de0ca	remove allow-untyped-defs from torch/ao/nn/quantized/modules/rnn.py (#157234 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157234 Approved by: https://github.com/jingsh ghstack dependencies: #157231, #157232	2025-07-08 00:11:52 +00:00
Bob Ren	e38a335d7f	remove allow-untyped-defs from torch/backends/cusparselt/__init__.py (#157232 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157232 Approved by: https://github.com/jingsh ghstack dependencies: #157231	2025-07-08 00:11:52 +00:00
Bob Ren	9d8cf24b3b	remove allow-untyped-defs from torch/_classes.py (#157231 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157231 Approved by: https://github.com/jingsh	2025-07-08 00:11:52 +00:00
James Wu	be56a8d7ac	Automatically load and save dynamo entries via caching_precompile (#155913 ) This PR adds a new config option, `caching_precompile`, and a `DynamoCache`, which loads and saves Dynamo Cache entries automatically. It also hooks up DynamoCache to PrecompileContext, so that we can save multiple cache entries. When this configuration is turned on, we: - Automatically create and initialize a CompilePackage on every torch.compile - Automatically use BundledAutogradcache - Automatically save the CompilePackage entry to DynamoCache after every compile You can also use PrecompileContext.serialize() to manually serialize a full object. I've added unit tests to exhibit this behavior. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155913 Approved by: https://github.com/zhxchen17	2025-07-07 23:57:17 +00:00
Ke Wen	3effe0c293	[SymmMem] find_path does not search /usr/local/lib (#157695 ) This PR uses `find_library` to replace `find_path`. It also searches for NVSHMEM host lib and device lib separately. Tested against system install location: /usr/local/lib and /usr/local/include. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157695 Approved by: https://github.com/Skylion007 ghstack dependencies: #157513	2025-07-07 23:16:45 +00:00
IvanKobzarev	2fde2090d0	[inductor_collectives] Make reorder_collectives_preserve_peak pass grouping nodes (#157706 ) Differential Revision: [D77861765](https://our.internmc.facebook.com/intern/diff/D77861765) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157706 Approved by: https://github.com/wconstab	2025-07-07 23:13:58 +00:00
rzou	5d8d126249	Fix einops x torch.compile interaction (#157600 ) Fixes https://github.com/pytorch/pytorch/issues/157451 If/when einops releases a version greater than 0.8.1, it will just break (without this patch). The history is: - Between 2.6 and 2.7, we tried to delete the einops import (#142847) - That didn't work so well, so we applied a hotfix in 2.7.1. (#153925) - The hotfix wasn't completely correct (0.8.1 is the latest version of einops, so the condition in the hotfix just always evaluates to True!) - It turns out we didn't need to delete the einops import. We already do not eagerly import einops. - I reverted the code back to the state it was in in 2.6. https://github.com/pytorch/pytorch/blob/release/2.6/torch/_dynamo/decorators.py Test Plan: - We have testing in CI for einops 0.6.1, 0.7.0, and 0.8.1. Wait for CI. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157600 Approved by: https://github.com/guilhermeleobas, https://github.com/anijain2305 ghstack dependencies: #157416	2025-07-07 23:04:02 +00:00
thenumberouscode	378c121d5e	Remove unnecessary warnings during the ATen compilation process. (#157703 ) Comparing uint32_t(num_threads()) with int(kCUDABlockReduceMaxThreads) always results in a compilation warning. Just change the return type of kCUDABlockReduceMaxThreads to uint32_t to avoid it. Fixes https://github.com/pytorch/pytorch/issues/157701 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157703 Approved by: https://github.com/malfet, https://github.com/Skylion007	2025-07-07 22:49:38 +00:00
Gabriel Ferns	7e83d50845	Inductor logging + analysis of torch.profile (#149697 ) Prereqs: - https://github.com/pytorch/pytorch/pull/152708 Features: 1. Adds inductor's estimate of flops and bandwidth to the json trace events that perfetto uses. 1. Only use the tflops estimation from triton if we don't have the info from the datasheet because Triton's estimates are inaccurate. I have a backlog item to fix triton flops estimation upstream. New `DeviceInfo` class, and new function `get_device_tflops`. 1. New helpers `countable_fx` and `count_flops_fx` helps get the flops of an `fx.Node`. 1. Extends Triton `torch.profiler` logging to `DebugAutotuner`. 1. New script `profile_analysis.py`: `--augment_trace` adds perf estimates to any perfetto json trace, `--analyze` creates a summary table of these perf estimates, and `--diff` will compare two traces side by side: ```python Device(NVIDIA H100, 0): Kernel Name \| resnet Kernel Count \| resnet FLOPS \| resnet bw gbps \| resnet Dur (ms) \| resnet Achieved FLOPS % \| resnet Achieved Bandwidth % \| newresnet Kernel Count \| newresnet FLOPS \| newresnet bw gbps \| newresnet Dur (ms) \| newresnet Achieved FLOPS % \| newresnet Achieved Bandwidth % --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- triton_poi_fused__native_batch_norm_legi \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 triton_red_fused__native_batch_norm_legi \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 triton_poi_fused__native_batch_norm_legi \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 void cutlass::Kernel2<cutlass_80_tensoro \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 triton_red_fused__native_batch_norm_legi \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 void cutlass::Kernel2<cutlass_80_tensoro \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 triton_poi_fused__native_batch_norm_legi \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 triton_red_fused__native_batch_norm_legi \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 triton_per_fused__native_batch_norm_legi \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 triton_poi_fused__native_batch_norm_legi \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 triton_per_fused__native_batch_norm_legi \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 triton_poi_fused__native_batch_norm_legi \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 void cublasLt::splitKreduce_kernel<32, 1 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 triton_per_fused__native_batch_norm_legi \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 triton_poi_fused__native_batch_norm_legi \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 void cask_plugin_cudnn::xmma_cudnn::init \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 triton_per_fused__native_batch_norm_legi \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 triton_per_fused__native_batch_norm_legi \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 void cutlass::Kernel2<cutlass_80_tensoro \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 triton_poi_fused__native_batch_norm_legi \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 triton_poi_fused__native_batch_norm_legi \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 triton_poi_fused__native_batch_norm_legi \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 triton_per_fused__native_batch_norm_legi \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 triton_poi_fused__native_batch_norm_legi \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 triton_per_fused__native_batch_norm_legi \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 triton_poi_fused__native_batch_norm_legi \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 triton_per_fused__native_batch_norm_legi \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 std::enable_if<!(false), void>::type int \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 triton_poi_fused_add_copy__38 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 triton_poi_fused_convolution_0 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 triton_poi_fused_convolution_1 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 void convolve_common_engine_float_NHWC<f \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 triton_per_fused__native_batch_norm_legi \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 triton_per_fused__native_batch_norm_legi \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 triton_poi_fused__native_batch_norm_legi \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 void at::native::(anonymous namespace):: \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 void at::native::vectorized_elementwise_ \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/149697 Approved by: https://github.com/eellison, https://github.com/shunting314	2025-07-07 22:13:34 +00:00
Shangdi Yu	6f05d58f2b	[AOTI] Split aoti_runtime/model.h to prepare for model static linking (#157592 ) Summary: Prepare for https://github.com/pytorch/pytorch/pull/157129. We split the file so we can re-use `model.h` part for codegen a separate header for each model in static linkage. Test Plan: CI Rollback Plan: Differential Revision: D77761249 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157592 Approved by: https://github.com/desertfire	2025-07-07 22:13:22 +00:00
lucasb-eyer	a7eb153bba	[MemoryViz] Add file selector button (#157647 ) In some linux desktop environments like mine, there is no drag and dropping of files. Which made the memoryviz impossible for me to use. So this adds a file selector button as an alternative. Tested that it works locally, and also works with multiple files. ![image](https://github.com/user-attachments/assets/dcb61d68-6c6f-42f6-a075-1783d747d1b0) And the button remains when something is loaded, to allow loading something else, but it moves out of the way to save vertical space: ![image](https://github.com/user-attachments/assets/4239d13c-3d80-4790-9696-0906c75e14e6) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157647 Approved by: https://github.com/sraikund16	2025-07-07 22:03:51 +00:00
Zeina Migeed	ed6df0e324	correctly import torch.version (#157584 ) The structure is ``` torch/ __init__.py version.py ``` When we import torch, only `torch/__init__.py` is executed by default. The submodules like `version.py` are not automatically imported or attached to the torch module. So without anything in `__init__.py`, `torch.version` may not be found. So in this PR, we make the import explicit. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157584 Approved by: https://github.com/ezyang	2025-07-07 21:43:35 +00:00
Ankita George	5c79a55e7e	[oss] Add version to metadata (#155343 ) Summary: We want to add versioning to DCP to the metadata so that whenever planner logic changes, we can use the version on save to determine how to load the data Test Plan: added a test Rollback Plan: Differential Revision: D76135887 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155343 Approved by: https://github.com/teja-rao	2025-07-07 20:57:30 +00:00
Andrey Talman	3d06ff82a8	[release] Triton pin update to 3.4 (#156664 ) Triton pin update issue: https://github.com/pytorch/pytorch/issues/154206 Please see post: https://dev-discuss.pytorch.org/t/2-8-final-rc-release-postponed-by-a-week/3101 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156664 Approved by: https://github.com/davidberard98	2025-07-07 20:52:25 +00:00
Olga Gerasimova	2efa5eaa65	swa avoid stream sync (#157705 ) Summary: When AveragedModel updates_parameters it calls self.n_averaged == 0 for each parameter, where n_averated is a buffer on GPU. Moving check before the cycle to call sync once It improves update_parameter from 74ms to 57ms ~22% improvement {F1980011097} {F1980011111} Test Plan: CI Rollback Plan: Differential Revision: D77723025 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157705 Approved by: https://github.com/albanD, https://github.com/Skylion007, https://github.com/janeyx99	2025-07-07 20:47:35 +00:00
zpcore	c2510fcd86	Fix index_put propagate strategy arg unpack error (#157671 ) Fix `index_put` propagate strategy didn't consider optional arg `accumulate`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157671 Approved by: https://github.com/fmassa, https://github.com/wconstab	2025-07-07 20:18:18 +00:00
Kurt Mohler	510c398a4f	Add `max_pool3d` backward pass for MPS (#157498 ) Note on backward precision over fp16: A float16 number has 10 bits of mantissa, 5 bits of exponent, and 1 bit for the sign. If the sign bit is positive, then with a mantissa $m$ and exponent $e$ represented in base 10, the number that the float16 format represents is $(1 + m / 1024) \exp2(e)$. ([source](https://en.wikipedia.org/wiki/Half-precision_floating-point_format)) Consider adding two numbers $a$ and $b$ which have arbitrary mantissas, and say their exponents are $e_a = 1$ (so $2 \le a \lt 4$) and $e_b=-3$ (so $0.175 \le b \lt 0.25$). Assume that the result has the same exponent as $a$. Since the exponents differ by 4, we'll effectively need to truncate the 4 rightmost bits of $b$'s mantissa, which would introduce a maximum error on the order of $(2^4 / 1024) \exp2(-3) \approx 0.002$. The error is nearly the same if $e_b = -2$ (so $0.25 \le b \lt 0.5$), where the 3 rightmost bits are truncated, giving a maximum error on the order of $(2^3 / 1024) \exp2(-2) \approx 0.002$. Same for $e_b=-1$. So if we're adding up nine different numbers that all have exponents -3, -2, or -1, and they sum to a number with exponent 1, then we would expect a maximum error of several times greater than 0.002. In my comments above, summing those particular nine numbers in different ways gave results that ranged between 3.1816 and 3.1758, a difference of $0.0058 \approx 2.9 * 0.002$. That's within the acceptable bounds, and we can safely just increase the error tolerance used in test_output_grad_match for the case of max_pool3d_backward with float16. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157498 Approved by: https://github.com/malfet	2025-07-07 19:46:44 +00:00
fduwjj	63a96eaeb8	[DeviceMesh] Add error when users try to slice non contiguous flattened dim submesh (#157523 ) With https://github.com/pytorch/pytorch/issues/157393, we want to first throw a clearer error for users and then fix it in the long-term Pull Request resolved: https://github.com/pytorch/pytorch/pull/157523 Approved by: https://github.com/fegin ghstack dependencies: #157501	2025-07-07 19:43:51 +00:00
fduwjj	2b8d3b1b2b	[DeviceMesh] Use user set backend and pg option even for the global mesh (#157501 ) Short term solution to https://github.com/pytorch/pytorch/issues/156593. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157501 Approved by: https://github.com/fegin, https://github.com/lw	2025-07-07 19:43:51 +00:00
Abhishek Nandy	bf1ebe0531	Fix typo: 'paramter' → 'parameter' in dynamo variable comment (#157651 ) This PR fixes a minor typo in a comment in `torch/_dynamo/variables/torch.py`, changing 'paramter' to the correct spelling 'parameter'. These small but meaningful changes help improve code readability and maintain the overall quality of the codebase. Thanks for your time and review! Pull Request resolved: https://github.com/pytorch/pytorch/pull/157651 Approved by: https://github.com/Skylion007	2025-07-07 19:42:44 +00:00
Sam Larsen	433a247102	[logging] [redo] dynamo_timed for CachingAutotuner.coordinate_descent_tuning (#156840 ) Summary: This is a redo of https://github.com/pytorch/pytorch/pull/156517, but with pt2_compile_events logging disabled. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156840 Approved by: https://github.com/jamesjwu	2025-07-07 19:09:48 +00:00
Wang, Chuanqi	8a47f9d03b	[CI] Fix xpu ci test sccache issue (#157693 ) With PR #157341 land, it broken the PXU CI test on sccache which has been disabled by #143851. Re-disable it Pull Request resolved: https://github.com/pytorch/pytorch/pull/157693 Approved by: https://github.com/atalman, https://github.com/huydhn	2025-07-07 18:29:38 +00:00
yifanmao	9e5f4a844c	[FSDP2] Fix issue with set_reduce_scatter_divide_factor errors and MixedPrecisionPolicy (#155964 ) fix https://github.com/pytorch/pytorch/issues/155223 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155964 Approved by: https://github.com/weifengpy	2025-07-07 17:09:29 +00:00
cyy	7c1f627828	Fix 'dllimport attribute ignored on inline function' (#157670 ) There are lots of warnings in builds: ``` 2025-07-05T16:59:46.9208806Z C:\actions-runner\_work\pytorch\pytorch\build\aten\src\ATen\core\TensorBody.h(5043,29): warning: 'at::Tensor::less_' redeclared inline; 'dllimport' attribute ignored [-Wignored-attributes] 2025-07-05T16:59:46.9209030Z 5043 \| inline at::Tensor & Tensor::less_(const at::Scalar & other) const { 2025-07-05T16:59:46.9209104Z \| ^ 2025-07-05T16:59:46.9209671Z C:\actions-runner\_work\pytorch\pytorch\build\aten\src\ATen\core\TensorBody.h(5048,29): warning: 'at::Tensor::less_' redeclared inline; 'dllimport' attribute ignored [-Wignored-attributes] 2025-07-05T16:59:46.9209860Z 5048 \| inline at::Tensor & Tensor::less_(const at::Tensor & other) const ``` This PR has fixed them and turned the warning into an error. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157670 Approved by: https://github.com/albanD	2025-07-07 16:57:48 +00:00
henrylhtsang	b3b4d28f4c	[submodule][cutlass] Update pin to b995f93 v4.0.0 (#157376 ) @Skylion007 seems afk. https://github.com/pytorch/pytorch/pull/153541 https://github.com/NVIDIA/cutlass/releases/tag/v4.0.0 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157376 Approved by: https://github.com/drisspg, https://github.com/Skylion007	2025-07-07 16:55:47 +00:00
PyTorch MergeBot	ae1094b72b	Revert "[WIP] Automatically load and save dynamo entries via caching_precompile (#155913 )" This reverts commit e466dab164d9236bfe5817ec8e4d24c7b9d3e392. Reverted https://github.com/pytorch/pytorch/pull/155913 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it seems to fail a test in trunk ([comment](https://github.com/pytorch/pytorch/pull/155913#issuecomment-3045914878))	2025-07-07 16:53:35 +00:00
Guilherme Leobas	eda0a9cc90	[list] Add list.__delitem__ (#156339 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156339 Approved by: https://github.com/zou3519 ghstack dependencies: #153969, #156148, #156242, #156270, #156271	2025-07-07 14:51:32 +00:00
Guilherme Leobas	d74ccf4ffe	[list] Add list.__mul__ and list.__imul__ (#156271 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156271 Approved by: https://github.com/zou3519 ghstack dependencies: #153969, #156148, #156242, #156270	2025-07-07 14:51:32 +00:00
Guilherme Leobas	689fba032d	Implement list.__add__ and list.__iadd__ (#156270 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156270 Approved by: https://github.com/Skylion007, https://github.com/zou3519 ghstack dependencies: #153969, #156148, #156242	2025-07-07 14:51:25 +00:00
Guilherme Leobas	c1d69d5dd5	[list] Implement `list.remove` (#156242 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156242 Approved by: https://github.com/Skylion007, https://github.com/zou3519 ghstack dependencies: #153969, #156148	2025-07-07 14:51:17 +00:00
Guilherme Leobas	e49acfc5c5	[list] Raise exception in invalid list method call (#156148 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156148 Approved by: https://github.com/zou3519 ghstack dependencies: #153969	2025-07-07 14:51:10 +00:00
Guilherme Leobas	034e996d37	[list] Implement list.count (#153969 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153969 Approved by: https://github.com/zou3519, https://github.com/XuehaiPan	2025-07-07 14:51:03 +00:00
Yahaya Suleiman	16c3b4143b	[gtest][listing] Enable gtest json listing for the fbcode/caffe2 project (#156816 ) *SUMMARY* The main function in this tests overrides that of the Gtest framework which contains it's `RUN_ALL_TESTS()` function. The main function in this test is called conditionally when conditions apply, in this case, when the C10_MOBILE directive is provided. This is wrong as we always want to call the `RUN_ALL_TEST()` function. In this PR, we only make the test suite available for cases that apply, i.e if the C10_MOBILE directive exist which represents the caching allocator and is only exposed on mobile *TEST PLAN* This tests should run in modes where it applies which should be covered in the CI run. Below shows a sample run in the dev-nosan mode which do not have the cache allocator BEFORE ``` buck test fbcode//caffe2:cpu_caching_allocator_test Discovered 0. Pass 0. Fail 0. Fatal 0. Skip 0. Timeout 0 ⚠ Listing failed: caffe2:cpu_caching_allocator_test Listing tests failed with error: Failed to read from /data/users/ysuleiman/fbsource/buck-out/v2/test/buck-out/v2/test_discovery/fbcode/6dcc55a61c1b90b3/default/tpx_execution_dir/gtest_output_file.json. Listing process stdout: , stderr: ``` AFTER ``` buck test '@fbcode//mode/dev-nosan' fbcode//caffe2:cpu_caching_allocator_test Analyzing targets. Remaining 0/46242 1871690 actions, 2251668 artifacts declared Executing actions. Remaining 0/257870 83:28:24.4s exec time total Command: test. Finished 10 remote, 112314 cache (99% hit) 83:22:43.5s exec time cached (99%) Time elapsed: 2:57.7s Tests finished: Pass 0. Fail 0. Fatal 0. Skip 0. Build failure 0 NO TESTS RAN ``` Rollback Plan: steps: - manual.note: content: Revert this diff Reviewed By: patskovn Differential Revision: D77229077 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156816 Approved by: https://github.com/kimishpatel	2025-07-07 14:16:43 +00:00
Henry Tsang	54a4d34d10	[fbcode] switch to cutlass-4 (#157579 ) Summary: Update cutlass version to 4. For most use cases. Test Plan: testing in progress Rollback Plan: Differential Revision: D77605011 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157579 Approved by: https://github.com/drisspg, https://github.com/Skylion007	2025-07-07 14:12:33 +00:00
PyTorch UpdateBot	78684e27ac	[xla hash update] update the pinned xla hash (#156584 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned xla hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156584 Approved by: https://github.com/pytorchbot	2025-07-07 12:09:20 +00:00
PyTorch UpdateBot	40e39ae21f	Update slow tests (#157696 ) This PR is auto-generated weekly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/weekly.yml). Update the list of slow tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157696 Approved by: https://github.com/pytorchbot	2025-07-07 12:09:06 +00:00
James Wu	e466dab164	[WIP] Automatically load and save dynamo entries via caching_precompile (#155913 ) This PR adds a new config option, `caching_precompile`, and a `DynamoCache`, which loads and saves Dynamo Cache entries automatically. It also hooks up DynamoCache to PrecompileContext, so that we can save multiple cache entries. When this configuration is turned on, we: - Automatically create and initialize a CompilePackage on every torch.compile - Automatically use BundledAutogradcache - Automatically save the CompilePackage entry to DynamoCache after every compile You can also use PrecompileContext.serialize() to manually serialize a full object. I've added unit tests to exhibit this behavior. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155913 Approved by: https://github.com/zhxchen17	2025-07-07 11:56:30 +00:00
Aleksei Nikiforov	d27d36136c	Don't try installing missing cuda dependencies on s390x (#157540 ) Don't try installing missing cuda dependencies on s390x Fixes #157409 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157540 Approved by: https://github.com/seemethere, https://github.com/huydhn	2025-07-07 09:16:38 +00:00
haozhe.zhu	815545f2dd	[inductor] enable bf32 for mkldnn linear pointwise/binary in inductor (#127294 ) When `torch.backends.mkldnn.matmul.fp32_precision=='bf16'`, we also enabled mkldnn linear in inductor path and allow to run with bf16 computation data type. Testplan: ``` python test/inductor/test_mkldnn_pattern_matcher.py -k test_linear_unary python test/inductor/test_mkldnn_pattern_matcher.py -k test_linear_fp32 python test/inductor/test_mkldnn_pattern_matcher.py -k test_multi_linear_share_same_input ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/127294 Approved by: https://github.com/jgong5, https://github.com/jansel Co-authored-by: Jiang, Yanbing <yanbing.jiang@intel.com>	2025-07-07 06:03:41 +00:00
Valentine233	d26ca5de05	Support transpose and pack for bit8 (#156065 ) To be used by CPU INT8 SDPA in torchao. https://github.com/pytorch/ao/pull/2380 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156065 Approved by: https://github.com/mingfeima, https://github.com/ezyang	2025-07-07 01:40:47 +00:00
Lei	2022588295	Fix: Ensure writeback handles NO_SHARD correctly by flattening tensors before copying (#154369 ) Fixes #151223 Because FSDP stores original parameters as views into a flattened tensor, changing the flattened parameter’s tensor directly can desynchronize the views. With the NO_SHARD strategy this caused a shape mismatch error when writing back modified parameters. Ensured writeback handles NO_SHARD correctly by flattening tensors before copying. The logic now flattens the source parameter or gradient when the strategy is unsharded to maintain the expected 1‑D shape for writeback operations Pull Request resolved: https://github.com/pytorch/pytorch/pull/154369 Approved by: https://github.com/weifengpy	2025-07-06 09:20:31 +00:00
Xuehai Pan	02715d0876	[BE][5/6] fix typos in test/ (test/dynamo/) (#157639 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157639 Approved by: https://github.com/yewentao256, https://github.com/jansel ghstack dependencies: #157638	2025-07-06 06:34:25 +00:00
Xuehai Pan	17687eb792	[BE][4/6] fix typos in test/ (test/inductor/) (#157638 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157638 Approved by: https://github.com/yewentao256, https://github.com/jansel	2025-07-06 06:34:25 +00:00
Sv. Lockal	7cda4017dd	Fix torch.utils.cpp_extension parser for clang version 20.1.7+libcxx (#157666 ) When CC and CXX compiler is set to clang, and clang was compiled with libc++, compilation of torchvision fails with: ``` File "/usr/lib/python3.12/site-packages/torch/utils/cpp_extension.py", line 585, in build_extensions compiler_name, compiler_version = self._check_abi() ^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/site-packages/torch/utils/cpp_extension.py", line 1034, in _check_abi _, version = get_compiler_abi_compatibility_and_version(compiler) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.12/site-packages/torch/utils/cpp_extension.py", line 449, in get_compiler_abi_compatibility_and_version if tuple(map(int, version)) >= minimum_required_version: ^^^^^^^^^^^^^^^^^^^^^^^^ ValueError: invalid literal for int() with base 10: '7+libcxx' ``` Compiler identification is a valid semantic version: ``` $ clang -dumpfullversion -dumpversion 20.1.7+libcxx ``` After adjusting parser of version, clang is able to compile extensions successfully. Fixes #157665 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157666 Approved by: https://github.com/msaroufim	2025-07-06 01:35:00 +00:00
Tom Ritchford	3e56a9cdfb	More testing of Python arithmetic operators between tensors and scalars (see 157266) (#157632 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157632 Approved by: https://github.com/ezyang, https://github.com/Skylion007	2025-07-05 17:48:27 +00:00
dsashidh	ee9ac36c23	Fixing misspelling in documentation (#157565 ) Fixes #157564 Fixes misspelling of the word parameter in documentation Pull Request resolved: https://github.com/pytorch/pytorch/pull/157565 Approved by: https://github.com/awgu, https://github.com/cyyever	2025-07-05 17:04:13 +00:00
Sandeep Narendranath Karjala	9be5860bc3	[dynamo] Fix dynamic shapes handling in after_aot repro generation (#157136 ) Summary: - Extract symbolic variables directly from graph placeholders and arguments - Add symbolic variable definitions to generated repro code - Add unit tests with ToyModel for testing Pull Request resolved: https://github.com/pytorch/pytorch/pull/157136 Approved by: https://github.com/xmfan ghstack dependencies: #157021	2025-07-05 15:38:41 +00:00
Abhishek Nandy	548c9d8281	Fix typo: 'paramter' → 'parameter' in quantization model report test (#157646 ) This PR addresses a minor typo in the file `test/quantization/fx/test_model_report_fx.py`: - Corrected the word "paramter" to "parameter" for better readability and accuracy. While it's a small change, correcting such typographical errors contributes to maintaining the overall quality and professionalism of the codebase. Thank you for your time and consideration in reviewing this PR. I'm happy to make any further adjustments if needed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157646 Approved by: https://github.com/yewentao256, https://github.com/ezyang	2025-07-05 12:28:36 +00:00
Abhishek Nandy	71a650ad56	Fix typo: 'Intializing' → 'Initializing' in test_parametrization.py (#157362 ) This pull request fixes a minor typo in the doc comments of `test/nn/test_parametrization.py`. - Replaced `'Intializing'` with `'Initializing'` in two docstring comments to improve clarity and maintain consistency across the codebase. This is a non-functional change and does not impact behavior or test outcomes. Thank you for maintaining such a high-quality codebase. Please let me know if any adjustments are needed. I'd be happy to help! Pull Request resolved: https://github.com/pytorch/pytorch/pull/157362 Approved by: https://github.com/ezyang	2025-07-05 12:21:15 +00:00
bobrenjc93	2471cc3355	[pc] verify max autotune is in generated source code (#157650 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157650 Approved by: https://github.com/aorenste ghstack dependencies: #157305, #157614, #157619	2025-07-05 07:55:11 +00:00
bobrenjc93	db00e1699a	[pc] introduce ProgressiveCompilationState and clear callback (#157619 ) followup from https://github.com/pytorch/pytorch/pull/157305 where @aorenste correctly suggested clearing callback. this refactor introduces a new dataclass so we don't need to check nullability for each field Pull Request resolved: https://github.com/pytorch/pytorch/pull/157619 Approved by: https://github.com/aorenste ghstack dependencies: #157305, #157614	2025-07-05 07:55:11 +00:00
bobrenjc93	5ea832e5f6	[pc] migrate progression futures from list to deque (#157614 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157614 Approved by: https://github.com/aorenste ghstack dependencies: #157305	2025-07-05 07:55:03 +00:00
Nikita Shulga	a952956d05	Add isnan exit condition to special ops (#157464 ) They might have been slow on CUDA-11.3, but this version of CUDA is long gone. More fundamental underlying issue were linear complexity of the recursive polynomial definitions for higher order polynomials, for example see this loop from implementation of Chebyshev polynomial of the first kind `7081b8233a/aten/src/ATen/native/Math.h (L2969-L2973)` which were tested by `test_compare_cpu` using following values (as sample index 16) `7081b8233a/torch/testing/_internal/opinfo/core.py (L2079)` Luckily chebyshev polynomials for absolute values higher than 1 pretty quickly reach infinity, see below ``` python3 -c "import torch;print(torch.special.chebyshev_polynomial_v(torch.nextafter(torch.tensor(1.0), torch.tensor(2.0)), torch.tensor(1e6)))" tensor(nan) ``` Which is not the case for Laguerre polynomials, but it's probably fine to just limit it to 1e7 Before ``` $ PYTORCH_TEST_WITH_SLOW=1 python test_ops.py -k chebyshev_polynomial_ ssssssss..ssssss..ssssss..ssssssssssssssssssssss..ssssss/home/ubuntu/py3.10-nightly/lib/python3.10/site-packages/torch/backends/cuda/__init__.py:131: UserWarning: This API is going to be deprecated, please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:78.) return torch._C._get_cublas_allow_tf32() ....ssssssssssss..ssssss..ssssss............ssssssssssssssssssssssssssssssssssss..ssssssssssssss..ssssss..ssssssssssssssssssssssssssssss..ssssss....ssssssssssss..ssssss..ssssss............ssssssssssssssssssssssssssssssssssss..ssssss..ssssssssssssss..ssssss..ssssss..ssssssssssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssssssssssss ---------------------------------------------------------------------- Ran 432 tests in 8.575s OK (skipped=344) ``` After ``` $ PYTORCH_TEST_WITH_SLOW=1 python test_ops.py -k chebyshev_polynomial_ ssssssss........................ssssssssssssssss......../home/ubuntu/pytorch/torch/backends/cuda/__init__.py:131: UserWarning: This API is going to be deprecated, please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /home/ubuntu/pytorch/aten/src/ATen/Context.cpp:78.) return torch._C._get_cublas_allow_tf32() ........................................................................................xxxxxxxx................ssssssssssssssssssssssss........................................................................................................ssssssss........................ssssssss........................................................................................ssssssss ---------------------------------------------------------------------- Ran 432 tests in 45.580s OK (skipped=72, expected failures=8) ``` Fixes https://github.com/pytorch/pytorch/issues/79528 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157464 Approved by: https://github.com/Skylion007, https://github.com/dcci ghstack dependencies: #157488	2025-07-05 04:19:50 +00:00
yewentao256	63e87d6d05	[Refactor] Add maybe unused flag to remove warning (#157655 ) Fixes #157653 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157655 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-07-05 03:23:39 +00:00
yewentao256	f7127b9b94	[Refactor] Remove unused variables (#157654 ) Fixes #157653 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157654 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-07-05 02:12:15 +00:00
Adeniji Adekunle James	44f5b93122	fix: correct sentence punctuation in cuDNN note (#157623 ) Fixes #ISSUE_NUMBER This PR fixes a small punctuation issue in the PyTorch README. Specifically: Added a missing full stop at the end of the sentence: "Note: You could refer to the cuDNN Support Matrix for cuDNN versions with the various supported CUDA, CUDA driver and NVIDIA hardware." Added comma for clarity between "CUDA driver" and "NVIDIA hardware". These edits improve the readability and grammatical correctness of the documentation. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157623 Approved by: https://github.com/Skylion007	2025-07-05 01:37:33 +00:00
Abhishek Nandy	e0fd48be7d	Fix typo: 'occurances' → 'occurrences' in mobile model test (#157629 ) This PR addresses a typo in the file `test/mobile/model_test/gen_test_model.py`. ### Changes: - Corrected "occurances" to the correct spelling "occurrences" - Renamed associated variables to reflect this change for consistency and clarity This is a non-functional, cleanup-only PR to improve code readability. Thanks to the PyTorch team for maintaining such a high-quality codebase Pull Request resolved: https://github.com/pytorch/pytorch/pull/157629 Approved by: https://github.com/Skylion007	2025-07-05 01:36:42 +00:00
Abhishek Nandy	43f7216327	Fix typo: 'paramters' → 'parameters' in ATen tunable README (#157575 ) This PR addresses a minor typo in the documentation file aten/src/ATen/cuda/tunable/README.md, where paramters has been corrected to parameters for improved clarity and consistency. Context Accurate and clear documentation is crucial for helping developers and contributors understand PyTorch internals. This small fix contributes to the overall quality and readability of the project. Thank you to the PyTorch team and maintainers for your continued efforts in building such an incredible framework. I'm happy to contribute in any way I can — even if just with a small doc improvement like this one. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157575 Approved by: https://github.com/eqy	2025-07-05 01:14:45 +00:00
Ke Wen	8a8fac1131	[SymmMem] Move code to where it is used (#157611 ) `maybe_initialize_env_vars` and `initialize_nvshmem_with_store` are only used in `NVSHMEMSymmetricMemory.cu`. Moving them there. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157611 Approved by: https://github.com/Skylion007 ghstack dependencies: #157513	2025-07-04 23:37:49 +00:00
Huy Do	bcc98bb2a4	Update _linux-test to support B200 runner (#157341 ) This unblocks https://github.com/pytorch/test-infra/issues/6869. The key changes to call out: * B200 needs OIDC to access ECR and upload stats to S3, so we need to set `id-token: write` in `_linux-test`. All workflows calling `_linux-test` also need to be updated accordingly * Connecting sccache to S3 on B200 doesn't seem to work, so I disable it. It still works locally though. ### Testing https://github.com/pytorch/pytorch/actions/runs/16055549292/job/45312298376 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157341 Approved by: https://github.com/nWEIdia, https://github.com/atalman, https://github.com/malfet	2025-07-04 23:19:24 +00:00
Xuehai Pan	524e827095	[build] modernize build-backend: `setuptools.build_meta:__legacy__` -> `setuptools.build_meta` (#155998 ) Change `build-system.build-backend`: `setuptools.build_meta:__legacy__` -> `setuptools.build_meta`. Also, move static package info from `setup.py` to `pyproject.toml`. Now the repo can be installed from source via `pip` command instead of `python setup.py develop`: ```bash python -m pip install --verbose --editable . python -m pip install --verbose --no-build-isolation --editable . ``` In addition, the SDist is also buildable: ```bash python -m build --sdist python -m install dist/torch-*.tar.gz # build from source using SDist ``` Note that we should build the SDist with a fresh git clone if we will upload the output to PyPI. Because all files under `third_party` will be included in the SDist. The SDist file will be huge if the git submodules are initialized. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155998 Approved by: https://github.com/ezyang, https://github.com/cyyever, https://github.com/atalman ghstack dependencies: #157557	2025-07-04 19:25:14 +00:00
Tom Ritchford	9968edd002	Fix #153942 (#153943 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153943 Approved by: https://github.com/malfet	2025-07-04 18:25:18 +00:00
Andrey Talman	7275f28045	Fix cuda 12.9 aarch64 GPU builds. Update CUDA_STABLE variable. (#157630 ) This contains 2 fixes that required in main and will need to be cherry-picked to Release 2.8 branch: 1. The PR https://github.com/pytorch/pytorch/pull/155819 missed to include triton change. 2. CUDA STABLE variable needs to be set to 12.8. Updating CUDA stable updates full static build Pull Request resolved: https://github.com/pytorch/pytorch/pull/157630 Approved by: https://github.com/Skylion007, https://github.com/jeanschmidt	2025-07-04 18:08:31 +00:00
zhxchen17	7be862ab8f	[dynamo] Relax DUPLICATED_INPUT to be serializable. (#157492 ) Since we don't actually rely on any real data while building DUPLICATE_INPUT guard, we can safely serialize it with sources and it should be able to reconstruct the guard correctly in the new process. Therefore we don't really need to prevent serializing it. Differential Revision: [D77683302](https://our.internmc.facebook.com/intern/diff/D77683302/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157492 Approved by: https://github.com/jamesjwu, https://github.com/jansel	2025-07-04 15:19:34 +00:00
Xuehai Pan	336f1e2d35	[AOTI] Fix AOT inductor CMake build dependency order (#157557 ) compile_model.py -> aoti_custom_class -> torch The custom command requires `torch` to be installed. `8408522976/test/cpp/aoti_inference/compile_model.py (L1-L7)` Fixes CI failure on trunk: - https://github.com/pytorch/pytorch/actions/runs/16041370426/job/45275085572#step:22:18348 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157557 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-07-04 14:33:36 +00:00
Abhishek Nandy	a46ea8a364	Fix typo: 'initalized' → 'initialized' in alias analysis test (#157628 ) This PR corrects a small spelling error in `test/jit/test_alias_analysis.py`. - "initalized" → "initialized" This is a minor comment correction and does not affect functionality or logic. Thank you for maintaining this amazing codebase. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157628 Approved by: https://github.com/Skylion007	2025-07-04 13:41:53 +00:00
zeshengzong	f41d017aa6	Add device check in `mse_loss` (#155089 ) Fixes #154978 ## Test Result ```python >>> import torch >>> import numpy as np >>> import torch.nn as nn >>> import torch.distributions.normal as norm >>> device = torch.device(('cuda' if torch.cuda.is_available() else 'cpu')) >>> print('Using {}'.format(device)) Using cuda >>> m = nn.Sequential(nn.Linear(1, 128).cuda(), nn.Tanh(), nn.Linear(128, 128).cuda(), nn.Tanh(), nn.Linear(128, 128).cuda(), nn.Tanh()) >>> m.to(device, dtype=None, non_blocking=False) Sequential( (0): Linear(in_features=1, out_features=128, bias=True) (1): Tanh() (2): Linear(in_features=128, out_features=128, bias=True) (3): Tanh() (4): Linear(in_features=128, out_features=128, bias=True) (5): Tanh() ) >>> opt = torch.optim.Adam(m.parameters(), lr=0.001) >>> print('Number of trainable parameters: ', sum((p.numel() for p in m.parameters() if p.requires_grad))) Number of trainable parameters: 33280 >>> input_tensor = torch.tensor(77.0, device=device) >>> target = torch.tensor(66.0) >>> loss_function = nn.MSELoss() >>> print('Loss Function: ', loss_function) Loss Function: MSELoss() >>> loss = loss_function(input_tensor, target) Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/home/zong/code/pytorch/torch/nn/modules/module.py", line 1767, in _wrapped_call_impl return self._call_impl(args, kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/zong/code/pytorch/torch/nn/modules/module.py", line 1778, in _call_impl return forward_call(args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/zong/code/pytorch/torch/nn/modules/loss.py", line 610, in forward return F.mse_loss(input, target, reduction=self.reduction) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/zong/code/pytorch/torch/nn/functional.py", line 3903, in mse_loss return torch._C._nn.mse_loss( ^^^^^^^^^^^^^^^^^^^^^^ RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155089 Approved by: https://github.com/cyyever, https://github.com/albanD	2025-07-04 12:37:48 +00:00
William Wen	52e4e41cbc	[dynamo] do not issue lru_cache warning for functions in the top-level torch namespace (#157598 ) `lru_cache` usage warning was being raised for `torch.get_device_module()`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157598 Approved by: https://github.com/Sidharth123-cpu	2025-07-04 08:17:50 +00:00
Jason Ansel	64f2ec77f8	[inductor] Fix fractional_max_pool2d 3D input causing assertion error (#156912 ) Fixes #156682 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156912 Approved by: https://github.com/angelayi	2025-07-04 06:09:28 +00:00
Laith Sakka	fdc5b42a8f	_broadcast_shapes gso generalizations (#157008 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157008 Approved by: https://github.com/ColinPeppler ghstack dependencies: #155590	2025-07-04 05:56:42 +00:00
bobrenjc93	d58ed04d89	[async-compile] add progressive compile mode (#157305 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157305 Approved by: https://github.com/aorenste	2025-07-04 04:18:50 +00:00
PyTorch UpdateBot	386bc9e2e9	[audio hash update] update the pinned audio hash (#156905 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156905 Approved by: https://github.com/pytorchbot	2025-07-04 04:06:59 +00:00
PyTorch MergeBot	f2e712ca14	Revert "Fix is_unaligned usage of statically_known_true (#157400 )" This reverts commit b359571c6043b40c4ae4fbb07135fd0f04902e21. Reverted https://github.com/pytorch/pytorch/pull/157400 on behalf of https://github.com/malfet due to It break tests, see `99c1a6bdd9/1` ([comment](https://github.com/pytorch/pytorch/pull/157400#issuecomment-3034353539))	2025-07-04 03:57:08 +00:00
Ke Wen	99c1a6bdd9	[SymmMem] Find NVSHMEM from system installation (#157513 ) Previously we only search for NVSHMEM from pip install location. This PR adds search in system locations deemed default by CMake. Related: #157453 untars NVSHMEM into `/usr/local` on our CI machines. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157513 Approved by: https://github.com/atalman, https://github.com/Skylion007	2025-07-04 03:34:44 +00:00
Geon-Woo Kim	4ed1b03f72	Add missing graph and memory related symbols to cuda_to_hip_mappings (#157435 ) (#157573 ) Summary: This PR adds missing CUDA symbols in `cuda_to_hip_mappings`. Test Plan: Tested in D77642700. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157573 Approved by: https://github.com/Skylion007 Co-authored-by: Geon-Woo Kim <gwkim@meta.com>	2025-07-04 03:03:04 +00:00
Ke Wen	8f9a191db6	[SymmMem] Fix CI name mismatch; remove TORCH_SYMMMEM requirement (#157597 ) Thanks @huydhn for spotting two name mismatches in the CI configs. We were matching against "test_h100_symm_mem" instead of "h100-symm-mem". Also, replaced `TORCH_SYMMMEM` env setting with programmatic method: `symm_mem.set_backend(...)` Further, skips a hanged test in `test_nvshmem_trion.py`. (#TODO @codingwithsurya ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157597 Approved by: https://github.com/fduwjj, https://github.com/huydhn	2025-07-04 01:43:08 +00:00
Mustafa Quraish	ef97bd4713	[torch] Add MTIA to the list of devices supporting foreach/fused kernels (#157583 ) Summary: We currently have foreach kernel implementations for MTIA, and for when we don't we internally decompose the ops. Anyone using this list for compatibility checks should be sending through the foreach kernels. Reviewed By: egienvalue, scottxu0730 Differential Revision: D77751248 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157583 Approved by: https://github.com/egienvalue	2025-07-04 01:15:24 +00:00
rzou	f0b388665e	Add dynamo_timed to bytecode hook (#157587 ) Test Plan: - ran tlparse on vLLM and saw this Pull Request resolved: https://github.com/pytorch/pytorch/pull/157587 Approved by: https://github.com/jingsh, https://github.com/BoyuanFeng	2025-07-04 01:11:03 +00:00
Chong Gu	c9a5bf09ba	[FP8] FP8 for SwishLayerNorm (#157574 ) Summary: Add a pass use_triton_fp8_swish_replace_normal_swish to replace _triton_swish_rms_norm with its counterpart that supports fp8 triton_swish_rms_norm, and turn on fp8 during inference. Test Plan: ``` buck2 run mode/opt mode/inplace -c fbcode.platform010_cuda_version=12.4 -c fbcode.nvcc_arch=h100 caffe2/torch/fb/model_transform/experimental/benchmark:mts_gpu_benchmark -- --lower-backend=AOT_INDUCTOR --model-snapshot-id=899072727_0 --node-replacement-dict="{}" --gpu-trace --add-passes=use_triton_fp8_swish_replace_normal_swish ``` The perf improvement on the 100x model with this pass is roughly ~7%, details are recorded [here](https://docs.google.com/document/d/1eIV_OTQyQcf_DlEDxwycTwhyGxT5OJkLzs8cPL6EMYc/edit?tab=t.0) Rollback Plan: Reviewed By: frank-wei Differential Revision: D76531303 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157574 Approved by: https://github.com/frank-wei	2025-07-04 01:06:21 +00:00
Guilherme Leobas	dfcda613b6	Ensure Dynamo can trace through explicit dunder method call (#154366 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154366 Approved by: https://github.com/zou3519 ghstack dependencies: #153150, #152991, #154539, #153553, #154063, #154064, #154065, #154066, #154263	2025-07-04 00:46:05 +00:00
Guilherme Leobas	0e7f02fe2e	[Dynamo] [FrozensetSubclass] Add support for user defined frozensets (#154263 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154263 Approved by: https://github.com/williamwen42 ghstack dependencies: #153150, #152991, #154539, #153553, #154063, #154064, #154065, #154066	2025-07-04 00:46:05 +00:00
Guilherme Leobas	308b88bde9	[Dynamo] [Set] Add comparison for set subclass (#154066 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154066 Approved by: https://github.com/Skylion007 ghstack dependencies: #153150, #152991, #154539, #153553, #154063, #154064, #154065	2025-07-04 00:45:58 +00:00
Guilherme Leobas	c51da57b55	[Dynamo] [Set] Raise TypeError in set.union(...) and "__or__" (#154065 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154065 Approved by: https://github.com/williamwen42 ghstack dependencies: #153150, #152991, #154539, #153553, #154063, #154064	2025-07-04 00:45:50 +00:00
Guilherme Leobas	f9544f1f0c	[Dynamo] [Set] Raise TypeError if object is unhashable (#154064 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154064 Approved by: https://github.com/Skylion007 ghstack dependencies: #153150, #152991, #154539, #153553, #154063	2025-07-04 00:45:42 +00:00
Guilherme Leobas	11c71053e0	[Dynamo] [Set] Implement some binop operators for dict/set/frozenset/dict_keys (#154063 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154063 Approved by: https://github.com/williamwen42, https://github.com/zou3519 ghstack dependencies: #153150, #152991, #154539, #153553	2025-07-04 00:45:34 +00:00
Guilherme Leobas	22abe6ded4	[Dynamo] [SetSubclass] Add support for user defined sets (#153553 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153553 Approved by: https://github.com/williamwen42, https://github.com/zou3519 ghstack dependencies: #153150, #152991, #154539	2025-07-04 00:45:25 +00:00
Guilherme Leobas	2b82c61f04	[Generator] Implement generator.__contains__ (#154539 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154539 Approved by: https://github.com/williamwen42, https://github.com/zou3519 ghstack dependencies: #153150, #152991	2025-07-04 00:45:18 +00:00
Guilherme Leobas	f651e28f80	[FrozenSet] Fixes for FrozenSet (#152991 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152991 Approved by: https://github.com/zou3519 ghstack dependencies: #153150	2025-07-04 00:45:11 +00:00
Guilherme Leobas	e7167dbacf	[Set] Support sets in VariableBuilder (#153150 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153150 Approved by: https://github.com/zou3519	2025-07-04 00:45:03 +00:00
Shuai Yang	6c42afe196	Introduce sync_cross_rank_decision (#156287 ) Summary: This is an improvement over `_broadcast_rank0_decision` where we uses the rank0's decision to broadcast to every rank. The issue of `_broadcast_rank0_decision` is that we observed large variance on the peak memory usage. One cause is that different ranks receive different dynamic shaped tensors and the hints of those tensors are different in different ranks. If we only rely on rank0's decision and it's unlucky to get unrepresentative hints, then the decision it makes may not be suitable for other ranks. Here, we introduce `sync_cross_rank_decision` which comes up with the decision after comparing all ranks' local decision, it will: 1. all gather decisions from all ranks; 2. test each decision on the current rank and get its estimated memory usage; 3. all reduce estimated memory usage with ReduceOp.MAX, so that we know the maximum memory usage of each decision on all ranks; 4. pick the decision which gives us minimum maximum memory memory usage; A graph to show more details https://internalfb.com/excalidraw/EX484509 After applying sync_cross_rank_decision, we observed that the variance are much smaller Rollback Plan: Differential Revision: D76714005 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156287 Approved by: https://github.com/fmassa, https://github.com/bdhirsh	2025-07-03 23:43:53 +00:00
Sheng Qin	f7130c097e	[nativert] Move Executor to PyTorch core (#157514 ) Test Plan: CI Rollback Plan: Differential Revision: D77693984 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157514 Approved by: https://github.com/zhxchen17	2025-07-03 23:31:54 +00:00
Scott Wolchok	ad86c05b78	efficient zero_mask implementation for vec128_*_neon (#155766 ) Differential Revision: [D76481039](https://our.internmc.facebook.com/intern/diff/D76481039/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155766 Approved by: https://github.com/malfet	2025-07-03 23:27:03 +00:00
Laith Sakka	b359571c60	Fix is_unaligned usage of statically_known_true (#157400 ) Summary: - symbolic shapes statically_known_true usage is wrong, this API is meant to be used for SymNodes. what is needed is V.graph.sizevars.statically_known_true. or V.graph.sizevars.statically_known_Equals or ideally V.graph.sizevars.statically_known_multiple_of. - The construction using == 0 is not symbolic, this used to always return false for symbolic inputs. Differential Revision: D77619293 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157400 Approved by: https://github.com/ColinPeppler	2025-07-03 23:26:36 +00:00
Aaron Gokaslan	a6fab82b16	[BE]: Fix NVSHMEM builds, add missing 12.9 dependency and update to latest for 2.8RC (#157453 ) Fixed our bad builds of nvshmem, (we were not building or testing before) and also updates to the latest version. Newest versions has critical support for things that would actually make it useful, like bfloat16 and float16 support. This is a proper fix for: https://github.com/pytorch/pytorch/pull/157411 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157453 Approved by: https://github.com/kwen2501, https://github.com/atalman	2025-07-03 22:55:18 +00:00
Teja	dd3e7170c2	Add async checkpointing impl to experimental checkpointer and add a builder API (#156927 ) 1. Adds an AsyncCheckpointer with out-of-process checkpointing and state_dict_stager with shared memory, pinned memory and Zero Overhead Support. 2. Adds two conveinient functions to create sync/async checkpointers Differential Revision: [D77336833](https://our.internmc.facebook.com/intern/diff/D77336833/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156927 Approved by: https://github.com/pradeepfn	2025-07-03 22:49:20 +00:00
Mark Saroufim	7081b8233a	[BE] Accelerator agnostic timer.py (#157131 ) Farewell to a lot of if statements - benefit is this now also supports mps synchronization Still need to think of a good test strategy for the privateUse1 removal, granted I'm not sure what the semantics of something like https://docs.pytorch.org/docs/stable/generated/torch.cpu.synchronize.html actually since CPU is probably synchronous? Pull Request resolved: https://github.com/pytorch/pytorch/pull/157131 Approved by: https://github.com/albanD	2025-07-03 22:23:04 +00:00
IvanKobzarev	7b392bac13	all_gather_bucketing fx pass (#157396 ) Porting passes to bucket all_gathers The main logic of the pass is done via 1. Searching for all all_gathers from the buckets Copying tests from @wconstab PR to test compatibility with reordering. Test checks only compatibility, as because of (3) the joint all_gather will be scheduled already as early as possible and no space for reordering. Pass changes: Using mutation ops to match performance of fsdp, in future the perfect scenario will be to have only functional graph, that inductor does all memory optimizations on its own without mutable ops. Inductor changes: Adding foreach_copy_ lowering Pull Request resolved: https://github.com/pytorch/pytorch/pull/157396 Approved by: https://github.com/wconstab	2025-07-03 22:07:42 +00:00
Abhishek Nandy	19ae5afdaa	Fix typo: 'recieve' → 'receive' in comments (#157544 ) This PR corrects minor typos in developer-facing comments: - Replaces 'recieve' with 'receive' in: - `FunctionalTensorWrapper.cpp` - `make_boxed_from_unboxed_functor.h` These changes improve code readability and maintain comment correctness. Thank you for reviewing! Pull Request resolved: https://github.com/pytorch/pytorch/pull/157544 Approved by: https://github.com/soulitzer	2025-07-03 19:11:15 +00:00
Xuehai Pan	3fd84a8592	[BE][PYFMT] migrate PYFMT for `torch/[a-c]*/` to `ruff format` (#144554 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/144554 Approved by: https://github.com/soulitzer	2025-07-03 18:56:07 +00:00
Manuel Candales	d56f11a1f2	[MPS] Implement logcumsumexp metal kernel (#156858 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156858 Approved by: https://github.com/malfet ghstack dependencies: #157512	2025-07-03 18:16:25 +00:00
Manuel Candales	794b95d54b	Enable Half dtype for logcumsumexp_backward (#157512 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157512 Approved by: https://github.com/malfet	2025-07-03 18:13:38 +00:00
rzou	e3fe001d9e	Add einops x torch.compile testing in PyTorch CI (#157416 ) Fixes #146782. This PR adds testing for multiple einops versions in PyTorch CI. This occurs in a new "einops" CI job that runs for both Python 3.9 and 3.13 (aka, what we test Dynamo over). Test Plan: - wait for CI Pull Request resolved: https://github.com/pytorch/pytorch/pull/157416 Approved by: https://github.com/guilhermeleobas, https://github.com/arogozhnikov, https://github.com/anijain2305	2025-07-03 17:36:39 +00:00
henrylhtsang	660dbea909	[cutlass backend] modify presets ahead of cutlass 4 upgrade (#157522 ) Differential Revision: [D77707409](https://our.internmc.facebook.com/intern/diff/D77707409/) Also asking in https://github.com/NVIDIA/cutlass/issues/2435 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157522 Approved by: https://github.com/coconutruben	2025-07-03 17:13:24 +00:00
Wanchao Liang	5cfe4377d6	[dtensor] Rework partial propagation in pointwise op and support mul (#157340 ) I am trying to see if I can easily add the linearity support for aten.mul to allow Partial placement to propagate through. But it turns out that I have to completely rework the current linearity propagation. In short, before this PR, linearity mainly support aten.add and some trival ops. It is done by allowing input Partial to propagate, and in the meanwhile, redistribute Replicate inputs to Partial to preserve the single device semantic, i.e suppose we want to execute `aten.add(lhs, rhs)` on 2 ranks: * `lhs` is partial, value on rank 0: `r0`, lhs value on rank 1: `r1` * `rhs` is replicate, value: `a` Then in order to preserve single device semantic (which should produce the value of `a + r0 + r1`), we do `rhs/world_size` first, then add `rhs` to `lhs`. This means every operand would first need be partial, then we can add them together. But this become non-true for multiplicative operations, like `aten.mul`, for `aten.mul`, assuming the same `aten.mul(lhs, rhs)` and value, we don't need to divide lhs by world_size to preserve single device semantic, b.c. `a* (r0+r1) = a* r0 + a* r1` So to accomodate the difference of add/mul, in this PR I: * change linearity to be a int to support different linearity types, add linearity and multiplicative are separate * add checks to ensure only a subset of partial types can support linearity (namely partial-sum/avg) * handle the linearity type plumbing through the pointwise ops. * add `mul.Tensor/Scalar` to be the multiplicative linearity * added the tests to show that the partial placements can be propagated with `aten.mul` Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/157340 Approved by: https://github.com/zpcore	2025-07-03 17:04:08 +00:00
henrylhtsang	898179331e	[cutlass backend] fix CutlassTensor post-renaming (#157408 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157408 Approved by: https://github.com/mlazos ghstack dependencies: #157402	2025-07-03 17:02:21 +00:00
PyTorch MergeBot	2e64e45b0b	Revert "[build] modernize build-backend: `setuptools.build_meta:__legacy__` -> `setuptools.build_meta` (#155998 )" This reverts commit 404008e3efdabeaf5b140a3aff77131461c33a0a. Reverted https://github.com/pytorch/pytorch/pull/155998 on behalf of https://github.com/malfet due to Broke inductor_cpp, wrapper see `e472daa809/1` ([comment](https://github.com/pytorch/pytorch/pull/155998#issuecomment-3032915058))	2025-07-03 16:47:07 +00:00
Sandeep Narendranath Karjala	e472daa809	[dynamo] Add fx_graph_runnable test coverage (#157021 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157021 Approved by: https://github.com/StrongerXi, https://github.com/xmfan Co-authored-by: Simon Fan <xmfan@meta.com>	2025-07-03 16:42:06 +00:00
Nikita Shulga	ec816d73b4	[MPS] Add `shifted_chebyshev_polynomial_[tuvw]` (#157488 ) For eager and inductor As for all other chebyshev ops, logic is simply compiled from `94716db222/aten/src/ATen/native/cuda/Math.cuh (L2821)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157488 Approved by: https://github.com/dcci	2025-07-03 15:48:37 +00:00
namgyu-youn	f17f658125	[profiler] add more CUDA API for kernel launcher (#156016 ) Add more kernel detection options, resolving TODO - References : [NVIDIA - docs](https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__EXECUTION.html) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156016 Approved by: https://github.com/albanD Co-authored-by: albanD <desmaison.alban@gmail.com>	2025-07-03 15:26:42 +00:00
PyTorch MergeBot	c9174a20f7	Revert "[BE] Unskip special ops (#157464 )" This reverts commit e124a0d88ca2aa04bfaca2dcabf5de6244048e45. Reverted https://github.com/pytorch/pytorch/pull/157464 on behalf of https://github.com/clee2000 due to caused slow test config to time out [GH job link](https://github.com/pytorch/pytorch/actions/runs/16037776972/job/45254574100) [HUD commit link](`e124a0d88c`) ([comment](https://github.com/pytorch/pytorch/pull/157464#issuecomment-3032676989))	2025-07-03 15:24:15 +00:00
PyTorch MergeBot	b6276a425f	Revert "[MPS] Add `shifted_chebyshev_polynomial_[tuvw]` (#157488 )" This reverts commit 9620994067b18e846a097d1e99af85ec2426ef0a. Reverted https://github.com/pytorch/pytorch/pull/157488 on behalf of https://github.com/clee2000 due to caused slow test config to time out [GH job link](https://github.com/pytorch/pytorch/actions/runs/16037776972/job/45254574100) [HUD commit link](`e124a0d88c`) ([comment](https://github.com/pytorch/pytorch/pull/157464#issuecomment-3032676989))	2025-07-03 15:24:15 +00:00
Abhishek Nandy	a0e0abd037	Fix typo: 'intialized' → 'initialized' in test_modules.py (#157226 ) This PR fixes a minor typo in `test/jit/test_modules.py`: - Before: `intialized` - After: `initialized` There are no functional code changes — this is a comment-only fix to improve clarity and consistency. Thank you to the PyTorch team for maintaining this outstanding project. Please let me know if anything else is needed. With appreciation, Abhishek Nandy [@abhitorch81](https://github.com/abhitorch81) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157226 Approved by: https://github.com/Skylion007	2025-07-03 14:56:02 +00:00
Abhishek Nandy	b221be9140	Fix typo: 'intial_query_grad' → 'initial_query_grad' in test_transformers.py (#157306 ) This is a minor typo fix in `test/test_transformers.py`: - Renamed `intial_query_grad` to `initial_query_grad` for improved clarity and correctness in test variable naming. There are no functional or logic changes — this PR is aimed purely at improving readability and maintaining code quality. Thanks to the PyTorch team for their work and review time Please feel free to suggest if this needs any adjustment. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157306 Approved by: https://github.com/Skylion007	2025-07-03 14:08:12 +00:00
atalman	8408522976	Remove +PTX from CUDA 12.8 builds (#157516 ) Remove +PTX from CUDA 12.8 builds and small refactor in build_cuda.sh. Removing +PTX reduces binary size required to be able to upload binaries to pypi Pull Request resolved: https://github.com/pytorch/pytorch/pull/157516 Approved by: https://github.com/malfet, https://github.com/ptrblck, https://github.com/tinglvv	2025-07-03 13:19:19 +00:00
Alexander Grund	c329a8f19c	Fix CPU bitwise shifts for out-of-limit values in VSX-vec (#157463 ) Similar to #96659 this implements the conditionals handling the out-of-limit values in the shift amounts (rhs) for the vectorized VSX code using the same logic as the scalar code. Fixes #109777 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157463 Approved by: https://github.com/jgong5	2025-07-03 10:41:33 +00:00
Tugsbayasgalan (Tugsuu) Manlaibaatar	5dfd8a9c7a	Remove is_jit_trace option (#157387 ) Summary: Title Test Plan: CI Rollback Plan: Differential Revision: D77319249 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157387 Approved by: https://github.com/pianpwk	2025-07-03 09:20:27 +00:00
Shawn Xu	8c2e450082	[PT][FSDP] fail `set_allocate_memory_from_process_group` if used together with custom comm hooks (#157487 ) Summary: This is a follow up after the PR to add comm override support: https://github.com/pytorch/pytorch/pull/155189 The previous PR loosely checks the allocation mixin classes, which isn't really safe as the actual hook may still override the behavior. This may lead to unnecessary confusion for no good use case. So for now we just make the 2 sets of APIs largely incompatible: 1. setting custom comms after `set_allocate_memory_from_process_group_for_comm()` is ok. 2. setting `set_allocate_memory_from_process_group_for_comm()` after custom comms is ko. Basically `set_allocate_memory_from_process_group_for_comm` is like a drop in hammer while the `set_custom_all_gather/reduce_scatter()` are like finer-grained scalpels that require more code crafted. We can revisit this if there's use case in between but for now they can be largely viewed independent from each other (even tho we do share some of the underlying pieces for now, that could be subject to change and should not be exposed to end users). Test Plan: added UT Differential Revision: D77681620 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157487 Approved by: https://github.com/weifengpy	2025-07-03 07:00:35 +00:00
Sheng Fu	2bb33e7a08	Fixed triton kernel in ET due to Triton version change. (#157484 ) Summary: Fixed triton kernel in ET due to Triton version change. Test Plan: buck2 run mode/opt param_bench/fb/integration_tests:test_et_replay Rollback Plan: Differential Revision: D77398841 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157484 Approved by: https://github.com/davidberard98	2025-07-03 06:16:23 +00:00
Tanima Dey	4ce6e6ec88	XCCL changes for DDP (#155497 ) Add XCCL documentation for DDP Pull Request resolved: https://github.com/pytorch/pytorch/pull/155497 Approved by: https://github.com/guangyey, https://github.com/AlannaBurke Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-07-03 05:18:08 +00:00
Will Constable	382598ef87	Fix unsafe collective reorder past wait (#157489 ) Covers the case where the output of one collective feeds the input of another collective. e.g. TP + FSDP - all_gather(tp+dp sharded param on TP dim) -> allgather dp_sharded buffer on DP dim Fixes a bug where the reordering pass specifically exempted wait nodes from dependencies. Note: this exemption was incorrect, so it should be removed. But it was also put there for a reason, to help move collectives past wait nodes that are not related to that collective. After this fix, reordering performance may be worse and we need to find a smarter way to decide if a particular wait node is a blocker for a given collective. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157489 Approved by: https://github.com/IvanKobzarev ghstack dependencies: #156879	2025-07-03 05:04:19 +00:00
Will Constable	dc524efb4d	Move logging into inner method for reorder pass (#156879 ) The reason for inner/outer method is to keep the outer method conforming to the typedef for a comms graph pass which returns one obj, while allowing unit tests to call the inner method that returns more metadata useful for testing the pass. The logs should be in the inner part, so they are functional also during unit testing. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156879 Approved by: https://github.com/IvanKobzarev	2025-07-03 05:04:19 +00:00
Klaus Zimmermann	5d5a5b3501	Fix GITHUB_OUTPUT syntax in create_release.yml workflow (#157466 ) #149919 fixed a number of linting issues, however, the conversion of the deprecated `::set-output` command to the new `>> $GITHUB_OUTPUT` redirect syntax went wrong, resulting in [failing uploads of the 2.8.0 rc1-rc3 pre-release tarballs](https://github.com/pytorch/pytorch/actions/runs/15892205745/job/44816789782). This PR fixes that. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157466 Approved by: https://github.com/clee2000, https://github.com/atalman	2025-07-03 04:57:52 +00:00
Xuehai Pan	404008e3ef	[build] modernize build-backend: `setuptools.build_meta:__legacy__` -> `setuptools.build_meta` (#155998 ) Change `build-system.build-backend`: `setuptools.build_meta:__legacy__` -> `setuptools.build_meta`. Also, move static package info from `setup.py` to `pyproject.toml`. Now the repo can be installed from source via `pip` command instead of `python setup.py develop`: ```bash python -m pip install --verbose --editable . python -m pip install --verbose --no-build-isolation --editable . ``` In addition, the SDist is also buildable: ```bash python -m build --sdist python -m install dist/torch-*.tar.gz # build from source using SDist ``` Note that we should build the SDist with a fresh git clone if we will upload the output to PyPI. Because all files under `third_party` will be included in the SDist. The SDist file will be huge if the git submodules are initialized. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155998 Approved by: https://github.com/ezyang, https://github.com/cyyever, https://github.com/atalman	2025-07-03 04:10:44 +00:00
henrylhtsang	b642a5c118	[cutlass backend] Add dynamo timed (#157410 ) Differential Revision: [D77631592](https://our.internmc.facebook.com/intern/diff/D77631592/) Before: ![Screenshot 2025-07-01 at 4 08 06 PM](https://github.com/user-attachments/assets/8f6445aa-50c7-456f-b5ac-b2749eb9bf40) After (different run): ![Screenshot 2025-07-01 at 5 11 09 PM](https://github.com/user-attachments/assets/7513d312-c4dc-4e39-9718-c63eb641bc30) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157410 Approved by: https://github.com/jingsh	2025-07-03 04:03:20 +00:00
fduwjj	493f42a541	[symm_mem] Create a one side get api for symm mem (#157294 ) Doing similar like what we did in https://github.com/pytorch/pytorch/pull/156443 so that we can also have a one-side get API for symmetric memory. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157294 Approved by: https://github.com/kwen2501	2025-07-03 03:52:05 +00:00
fduwjj	662c1cfed2	[c10d][PGNCCL] Add waitcounter for watchdog and heartbeat monitoring thread (#157480 ) We want to have a wait counter for both side thread so that we can monitor its lifecycle. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157480 Approved by: https://github.com/d4l3k	2025-07-03 02:47:06 +00:00
Yu, Guangye	5cc4e856fd	Add device_id to XPU device properties (#156481 ) # Motivation Some older Intel iGPUs may share the same device name across different hardware products. (See [device name example](`aaa01c06f9/shared/source/dll/devices/devices_base.inl (L190-L199)`)) To help disambiguate which specific iGPU product is being used, we introduce the use of a [device id](https://github.com/intel/llvm/blob/sycl/sycl/doc/extensions/supported/sycl_ext_intel_device_info.md#device-id). This device id corresponds to the Device ID in [official Intel product specification](https://www.intel.com/content/www/us/en/products/sku/232155/intel-core-i71360p-processor-18m-cache-up-to-5-00-ghz/specifications.html) and enables more accurate identification and troubleshooting for user issues. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156481 Approved by: https://github.com/EikanWang, https://github.com/albanD	2025-07-03 01:22:11 +00:00
Valentine233	7597988f1b	[fake tensor] fix issue of no attribute tags (#156689 ) Fixes #156688 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156689 Approved by: https://github.com/leslie-fang-intel, https://github.com/atalman	2025-07-03 01:16:01 +00:00
Nikita Shulga	9620994067	[MPS] Add `shifted_chebyshev_polynomial_[tuvw]` (#157488 ) For eager and inductor As for all other chebyshev ops, logic is simply compiled from `94716db222/aten/src/ATen/native/cuda/Math.cuh (L2821)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157488 Approved by: https://github.com/dcci ghstack dependencies: #157464	2025-07-02 23:29:35 +00:00
Nikita Shulga	e124a0d88c	[BE] Unskip special ops (#157464 ) They were slow on CUDA-11.3, which has long been gone, let's see if they work now Before ``` $ python test_ops.py -k chebyshev_polynomial_ ssssssss..ssssss..ssssss..ssssssssssssssssssssss..ssssss/home/ubuntu/py3.10-nightly/lib/python3.10/site-packages/torch/backends/cuda/__init__.py:131: UserWarning: This API is going to be deprecated, please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:78.) return torch._C._get_cublas_allow_tf32() ....ssssssssssss..ssssss..ssssss............ssssssssssssssssssssssssssssssssssss..ssssssssssssss..ssssss..ssssssssssssssssssssssssssssss..ssssss....ssssssssssss..ssssss..ssssss............ssssssssssssssssssssssssssssssssssss..ssssss..ssssssssssssss..ssssss..ssssss..ssssssssssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssss..ssssssssssssss ---------------------------------------------------------------------- Ran 432 tests in 8.575s OK (skipped=344) ``` After ``` $ python test_ops.py -k chebyshev_polynomial_ ssssssss........................ssssssssssssssss......../home/ubuntu/py3.10-nightly/lib/python3.10/site-packages/torch/backends/cuda/__init__.py:131: UserWarning: This API is going to be deprecated, please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:78.) return torch._C._get_cublas_allow_tf32() ........................................................................................ssssssss................ssssssssssssssssssssssss........................................................................................................ssssssss........................ssssssss........................................................................................ssssssss ---------------------------------------------------------------------- Ran 432 tests in 42.379s OK (skipped=80) ``` Fixes https://github.com/pytorch/pytorch/issues/79528 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157464 Approved by: https://github.com/Skylion007	2025-07-02 23:16:52 +00:00
Laith Sakka	7cfd054075	[attempt 2] Compute contiguity symbolically to avoid dde, and introduce c++ sym_is_contiguous (#157472 ) Summary: When we compute contiguity for a tensor with dynamic shapes we first: 1) Try to compute it without guarding. 2) If all shapes hinted, compute it with potentially adding guards. 3) if any input is not hinted, compute it symbolically. sym_is_contiguous return a SymBool that is then either evaluated or guard_or_false can be called on it to avoid data dependent errors. ex: bool is_contiguous = input.sym_is_contiguous().guard_or_false(__FILE__, __LINE__); is_contiguous_or_false is a helper function that does that. In this PR I only handle default contiguity, will follow up with changes for other formats like channel_last . We use this patter in this PR for several locations to avoid DDEs. Test Plan: contbuild & OSS CI, Rollback Plan: Reviewed By: malfet Differential Revision: D77639021 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157472 Approved by: https://github.com/aorenste	2025-07-02 23:12:29 +00:00
Xuehai Pan	d40aaa42ee	[BE][16/16] fix typos in torch/ (torch/utils/) (#156606 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156606 Approved by: https://github.com/albanD ghstack dependencies: #156318, #156320, #156602, #156604	2025-07-02 22:55:29 +00:00
Xuehai Pan	11c07c848c	[BE][14/16] fix typos in torch/ (torch/fx/) (#156604 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156604 Approved by: https://github.com/jingsh ghstack dependencies: #156318, #156320, #156602	2025-07-02 22:55:29 +00:00
Xuehai Pan	db259bd6b8	[BE][12/16] fix typos in torch/ (#156602 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156602 Approved by: https://github.com/justinchuby, https://github.com/albanD ghstack dependencies: #156318, #156320	2025-07-02 22:55:29 +00:00
Xuehai Pan	d5cdc36943	[BE][10/16] fix typos in torch/ (torch/csrc/jit/) (#156320 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156320 Approved by: https://github.com/albanD ghstack dependencies: #156318	2025-07-02 22:55:29 +00:00
Xuehai Pan	541584d22e	[BE][8/16] fix typos in torch/ (torch/csrc/jit/) (#156318 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156318 Approved by: https://github.com/albanD	2025-07-02 22:55:29 +00:00
henrylhtsang	c0e155a8d2	[cutlass backend] Use alignment of D for EVT / Float8 (#157402 ) I encountered an C++ compile error from running cutlass backend tests when upgrading cutlass version. It seems like Nvidia added "static_assert(detail::is_aligned<ElementC_, AlignmentC, ElementD_, AlignmentD>()," `b995f93317/include/cutlass/epilogue/collective/builders/sm90_builder.inl (L297)` However, it seems codegen have the wrong alignment for D. For C, 1 is okay since it is void. But for D, this is probably wrong. ``` void, cutlass::layout::ColumnMajor, 1, cutlass::bfloat16_t, cutlass::layout::RowMajor, 1, ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157402 Approved by: https://github.com/ColinPeppler, https://github.com/mlazos	2025-07-02 22:55:00 +00:00
Animesh Jain	48560eef80	[dynamo] Fix bug in dict(mapping_proxy) (#157467 ) Fixes https://github.com/pytorch/pytorch/issues/157284 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157467 Approved by: https://github.com/jansel, https://github.com/StrongerXi Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-07-02 22:13:02 +00:00
Catherine Lee	fd4f704905	[ez][CI] Print set output in CI (#157477 ) Print what the output that's getting set is for better debugging It's probably bad there are 4 of these, but I'm also not sure if imports will behave correctly Pull Request resolved: https://github.com/pytorch/pytorch/pull/157477 Approved by: https://github.com/huydhn	2025-07-02 21:47:19 +00:00
Catherine Lee	60e66d11ab	[CI] Keep-going on main (#157470 ) Run an experiment where we turn on keep going on main. Revert this PR to cancel the experiment There have been a couple of changes that make it so that HUD will show the failure early even while the job is in progress, so triaging for reverts should still be able to happen quickly Pull Request resolved: https://github.com/pytorch/pytorch/pull/157470 Approved by: https://github.com/huydhn, https://github.com/ZainRizvi, https://github.com/malfet	2025-07-02 21:42:46 +00:00
Will Constable	4b4c2a7b1d	Support complex numbers in DTensor redistribute (#157329 ) Add complex number unwrapping in functional collectives used by DTensor. Complex tensors are not directly supported by underlying comm kernels (e.g. nccl) but complex tensors can be viewed as real tensors of a higher rank (added size-2 tensor dim represents real vs im component). Collective output is then viewed as complex to restore the original/expected shape and dtype. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157329 Approved by: https://github.com/XilunWu	2025-07-02 21:37:16 +00:00
Benjamin Glass	af9c92b4cb	[CI] Remove redundant accuracy benchmarks for cpp_wrapper (#155966 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155966 Approved by: https://github.com/desertfire	2025-07-02 20:58:08 +00:00
Catherine Lee	c09cf29d7d	[ez][BE] Tag deletion script to delete any old ciflow + autorevert tags (#157468 ) Change the branch/tag deletion script that runs once per day to delete more tags Previous: only delete ciflow tags that didn't correspond to an open PR New: delete ciflow tags attached to commits that are > 7 days old. Also delete `trunk/<sha>` (I think they are for autorevert) tags that are attached to commits that are > 7 days old It's hard to figure out when the actual tag was pushed or created, so instead it looks at the commit date, which might lead to unexpected behavior if the tag was pushed much later than the commit (ex triggering periodic later to bisect). I think it's ok though since you don't really need the tag after the workflow runs Pull Request resolved: https://github.com/pytorch/pytorch/pull/157468 Approved by: https://github.com/izaitsevfb	2025-07-02 20:42:32 +00:00
Catherine Lee	6f60cfe9b1	[ez] Add super().setUp() in test_ops::TestFakeTensor (#157475 ) Noticed some disable issues getting a bunch of comments, so I took a look One day I'll write a better check for this Pull Request resolved: https://github.com/pytorch/pytorch/pull/157475 Approved by: https://github.com/huydhn	2025-07-02 20:34:00 +00:00
zhxchen17	e20784f228	[dynamo] Support BUILTIN_MATCH serialization. (#157016 ) Serialize BUILTIN_MATCH since they are all stored in __builtin__ dict. Also fixed an issue that the wrong global scope is passed to CheckFunctionManager while loading guards. Previously we can always reuse the compile-time global scope for evaluating guards because the compile-time and runtime global scope are always the same. For precompile, we need to serialize the compile-time global scope for loading only. We need to point the CheckFunctionManager to the new global scope after loading is finished for evaluating guards. Differential Revision: [D77159313](https://our.internmc.facebook.com/intern/diff/D77159313/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157016 Approved by: https://github.com/jansel, https://github.com/jamesjwu	2025-07-02 20:24:24 +00:00
Colin Peppler	172853547a	[inductor] more size_hint_or_throw usage (#157394 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157394 Approved by: https://github.com/jingsh	2025-07-02 20:20:59 +00:00
Catherine Lee	e0ab1b538a	[ez][BE] Remove max jobs override for CI build jobs (#157473 ) Basically reverts #147487 since it's not needed anymore Not an exact revert because some things have already been removed in a different PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/157473 Approved by: https://github.com/huydhn	2025-07-02 20:12:28 +00:00
Nikita Shulga	3f569f9af7	[BE] Remove extra semicolon (#157486 ) Fixes ``` /Users/nshulga/git/pytorch/pytorch/torch/nativert/executor/GraphExecutorBase.cpp:16:58: warning: extra ';' outside of a function is incompatible with C++98 [-Wc++98-compat-extra-semi] 16 \| execPlan_(ExecutionPlanner{graph_}.createPlan()) {}; \| ^ 1 warning generated. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157486 Approved by: https://github.com/seemethere, https://github.com/atalman, https://github.com/Skylion007	2025-07-02 19:56:21 +00:00
Nicolas Macchioni	94716db222	[BE][DCE] eliminate remnants of global gemm cache (#157327 ) Summary: The global gemm cache has not been maintained in ~1 year, and the only entry point (`search_autotune_cache`) was recently deprecated. Meaning, this is now dead code that we can remove. Test Plan: CI Rollback Plan: Differential Revision: D77520979 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157327 Approved by: https://github.com/jansel	2025-07-02 19:52:35 +00:00
atalman	06f39a71b6	Add Release 2.8 CUDA matrix. Update Release schedule for 2.7.1 and 2.9 (#157482 ) This PR: - Adds Release 2.8 CUDA matrix - Update Release 2.9 schedule, to make it more similar to 2.5 release schedule. Mid Oct release - Update 2.7.1 release day Pull Request resolved: https://github.com/pytorch/pytorch/pull/157482 Approved by: https://github.com/Camyll	2025-07-02 19:52:24 +00:00
Ahmad Sharif	36dd598bda	layernorm tests: Tweak test thresholds for comparing tensors (#156699 ) After I landed this PR: https://github.com/pytorch/pytorch/pull/156600, this test was failing internally on large tensors because the differences were greater than tolerances on some cuda devices. We now raise the tolerances for larger tensors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156699 Approved by: https://github.com/eqy, https://github.com/ngimel	2025-07-02 19:33:38 +00:00
dolpm	32983ea698	[nativert] continue to move generated static dispatch kernels (#157460 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D77623080 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157460 Approved by: https://github.com/zhxchen17	2025-07-02 19:28:13 +00:00
Nikita Shulga	5e636d664a	[BE] `@serialTest` decorator must be called (#157388 ) Otherwise it turns test into a trivial one(that always succeeds), as following example demonstrates ```python import torch from torch.testing._internal.common_utils import serialTest, run_tests, TestCase class MegaTest(TestCase): @serialTest def test_foo(self): if hasattr(self.test_foo, "pytestmark"): print("foo has attr and it is", self.test_foo.pytestmark) print("foo") @serialTest() def test_bar(self): if hasattr(self.test_bar, "pytestmark"): print("bar has attr and it is", self.test_bar.pytestmark) print("bar") if __name__ == "__main__": run_tests() ``` That will print ``` test_bar (__main__.MegaTest.test_bar) ... bar has attr and it is [Mark(name='serial', args=(), kwargs={})] bar ok test_foo (__main__.MegaTest.test_foo) ... ok ---------------------------------------------------------------------- Ran 2 tests in 0.013s ``` Added assert that arg is boolean in the decorator to prevent such silent skips in the future Pull Request resolved: https://github.com/pytorch/pytorch/pull/157388 Approved by: https://github.com/clee2000	2025-07-02 19:15:19 +00:00
Dhia-naouali	eaf32fffb7	fixed a tiny typo in torch.compiler.md (#157462 ) Fixes #157444 there was a typo in [docs/source/torch.compiler.md](https://github.com/pytorch/pytorch/blob/main/docs/source/torch.compiler.md) : see -> seen Pull Request resolved: https://github.com/pytorch/pytorch/pull/157462 Approved by: https://github.com/Skylion007, https://github.com/svekars	2025-07-02 19:15:15 +00:00
Xuehai Pan	0e9d8032a3	[build] remove cmake cache and reconfigure again if it is invalid (#156958 ) See also: - astral-sh/uv#14269 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156958 Approved by: https://github.com/Skylion007 ghstack dependencies: #156742	2025-07-02 18:46:32 +00:00
xadupre	0105cd89ab	[ONNX] Fix conversion of attention - 4D (#157130 ) Fixes a wrong conversion to onnx while investigation #149662. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157130 Approved by: https://github.com/gramalingam, https://github.com/justinchuby, https://github.com/titaiwangms Co-authored-by: Justin Chu <justinchuby@users.noreply.github.com>	2025-07-02 18:05:10 +00:00
dolpm	d5d14ee823	[nativert] create persistent value helper (#157286 ) Summary: att Test Plan: CI Reviewed By: georgiaphillips Differential Revision: D74300519 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157286 Approved by: https://github.com/SherlockNoMad	2025-07-02 17:15:52 +00:00
Prajesh Praveen Anchalia	156bc243f0	Back out "Include c++ stack traces when we hit constraint violation (#155603 )" (#157406 ) Summary: Original commit changeset: 4b3fdaa8f2c6 Original Phabricator Diff: D76434787 Meta: https://fb.workplace.com/groups/1286739428954016/permalink/1535462614081695/ Test Plan: Meta: Revert D76434787 for S536719 Rollback Plan: Differential Revision: D77626334 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157406 Approved by: https://github.com/bobrenjc93	2025-07-02 16:51:07 +00:00
James Wu	bd6b5fddbf	[Precompile] [easy] Serialize requires_grad for tensors when serializing guards (#157372 ) Need to keep requires_grad on the tensor when serializing/deserializing guards. This matters when there's a TENSOR_MATCH guard on a tensor that requires_grad. Added a unit test. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157372 Approved by: https://github.com/jansel, https://github.com/zhxchen17 ghstack dependencies: #156433	2025-07-02 16:34:37 +00:00
Witold Dziurdz	54701a0c94	Add is_hidden_event method to KinetoEvent Python interface (#155214 ) Fixes #155213 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155214 Approved by: https://github.com/sraikund16	2025-07-02 16:29:21 +00:00
Paul Zhang	0edc1b91f7	[Inductor] Disable decompose_k for AMD (#157283 ) Differential Revision: D77544250 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157283 Approved by: https://github.com/bdhirsh	2025-07-02 15:21:46 +00:00
Abhishek Nandy	9f5276dc07	Fix typo: 'Intializes' → 'Initializes' in _distributed_c10d.pyi docst… (#157455 ) Description: This PR fixes a small documentation typo in torch/_C/_distributed_c10d.pyi, correcting: Intializes → Initializes This helps improve clarity in internal docstrings for maintainers and contributors. Let me know if further changes are needed. Thanks for your time and the amazing work on PyTorch! Pull Request resolved: https://github.com/pytorch/pytorch/pull/157455 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-07-02 15:19:05 +00:00
Guilherme Leobas	9d175bc7e6	Fixes for CPython int/float tests (#155978 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155978 Approved by: https://github.com/zou3519	2025-07-02 15:04:00 +00:00
Xuehai Pan	b096341963	[BE] use `pathlib.Path` instead of `os.path.*` in `setup.py` (#156742 ) Resolves: - https://github.com/pytorch/pytorch/pull/155998#discussion_r2164376634 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156742 Approved by: https://github.com/malfet	2025-07-02 14:57:58 +00:00
David Berard	82eefaedd9	[inductor][user triton] sanitize triple-quoted docstrings in kernel definitions (#157322 ) Fixes #155006 Inductor sometimes codegens triton kernel definitions into a triple-quoted text block. If the text block itself contains triple-quotes, this breaks. Notably, this can happen for user-defined triton kernels, where the user may have added a docstring in their triton kernel. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157322 Approved by: https://github.com/zou3519, https://github.com/drisspg	2025-07-02 14:02:01 +00:00
PyTorch MergeBot	c553c55be7	Revert "Fix full_like decomposition to preserve strides (#144765 )" This reverts commit 01b0f09931d47bd2716398a0c335b2807dc3074d. Reverted https://github.com/pytorch/pytorch/pull/144765 on behalf of https://github.com/jeanschmidt due to Seems to be breaking internal tests see [D77652778](https://www.internalfb.com/diff/D77652778), @jansel may you help get this PR merged? ([comment](https://github.com/pytorch/pytorch/pull/144765#issuecomment-3027975098))	2025-07-02 13:56:03 +00:00
PyTorch MergeBot	d5a89178b0	Revert "[dynamo] Add fx_graph_runnable test coverage (#157021 )" This reverts commit 77676753ecabf6a6645bdd3abfe01939e5751e76. Reverted https://github.com/pytorch/pytorch/pull/157021 on behalf of https://github.com/jeanschmidt due to New tests are red internally, more details on [D77652793](https://www.internalfb.com/diff/D77652793). Maybe codev could be a better strategy to merge this PR faster... ([comment](https://github.com/pytorch/pytorch/pull/157021#issuecomment-3027952946))	2025-07-02 13:48:41 +00:00
William Wen	bdb7819166	[dynamo, nested graph breaks] remove recursive cell/freevar in instruction tx (#154078 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154078 Approved by: https://github.com/StrongerXi, https://github.com/jansel	2025-07-02 13:36:14 +00:00
Bin Bao	34c8033fd3	Fix a div_mod bug in generic_math.h (#157383 ) Summary: There is a bug in integer div_mod that when the remainder is 0 and the divisor is negative, mod operation produces a negative number. Fixed in this PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157383 Approved by: https://github.com/angelayi, https://github.com/jingsh	2025-07-02 12:22:57 +00:00
William Wen	ab2294d828	[dynamo] fix _torchdynamo_orig_callable naming issues (#156901 ) `_torchdynamo_orig_callable` was being used in two distinct places: - to get the original user function from nested eval_frame.py decorators - to get the original backend from nested convert_frame.py callbacks We rename ~the first usage to `_torchdynamo_orig_fn`~ and the second to `_torchdynamo_orig_backend` in order to distinguish these cases. UPDATE: seems like both internal and OSS users depend on `_torchdynamo_orig_callable`, but it only seems in the first context. We should thus keep the original name for the first case then. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156901 Approved by: https://github.com/StrongerXi, https://github.com/jansel	2025-07-02 09:53:55 +00:00
Dylan Maloy	3173616532	[nativert] start to move generated static dispatch kernels (#157403 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D77622952 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157403 Approved by: https://github.com/georgiaphillips	2025-07-02 08:42:01 +00:00
PyTorch MergeBot	8c0df6fe17	Revert "[dynamo][fsdp] Consistent behavior of int attributes (#157262 )" This reverts commit 42b48ee67229286127390000f103a11dfc8901f5. Reverted https://github.com/pytorch/pytorch/pull/157262 on behalf of https://github.com/jeanschmidt due to Newly introduced tests are red in internal runs, check D77593713 ([comment](https://github.com/pytorch/pytorch/pull/157262#issuecomment-3026944993))	2025-07-02 08:30:39 +00:00
Shawn Xu	0364db7cd1	[PT] support custom all_gather and reduce_scatter comms (#155189 ) Summary: This change introduces 2 comm override APIs: `set_custom_all_gather` and `set_custom_reduce_scatter` to allow for custom behavior respectively. This allow users to control how the comm buffers are allocated and the exact comm implementation for flexibility. For details, see docstring in `Comm` in `_fsdp_api.py` Related PR: https://github.com/pytorch/pytorch/pull/150564 Test Plan: CI Differential Revision: D75714362 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155189 Approved by: https://github.com/weifengpy	2025-07-02 06:58:45 +00:00
haozhe.zhu	f8c0a4bd28	[inductor] enable bf32 test for mkldnn conv (#127293 ) Enable more test on inductor conv + bf32 Testplan: ``` python test/inductor/test_mkldnn_pattern_matcher.py -k test_conv2d_unary_cpu python test/inductor/test_mkldnn_pattern_matcher.py -k test_conv3d_unary_cpu python test/inductor/test_mkldnn_pattern_matcher.py -k test_conv_transpose2d_unary python test/inductor/test_mkldnn_pattern_matcher.py -k test_conv2d_binary python test/inductor/test_mkldnn_pattern_matcher.py -k test_conv3d_binary ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/127293 Approved by: https://github.com/jgong5 ghstack dependencies: #126050, #126054 Co-authored-by: Jiang, Yanbing <yanbing.jiang@intel.com>	2025-07-02 01:49:01 +00:00
haozhe.zhu	4c8eb65efb	allow to use bf16 as fp32 internal precision for mkldnn conv backward (#126054 ) Used for CI since depends on ideep update. Allow to use `BF16` as the internal computation data types by `torch.backends.mkldnn.conv.fp32_precision="bf16"` ### TestPlan python test/test_mkldnn.py -k conv ### Benchmarking FP32 conv2d backward vs. BF16 internal computation conv backward on SPR Single core: Input \| fp32 ms \| bf16 internal ms \| Speed up -- \| -- \| -- \| -- IC: 64, OC: 256, kernel: 1, stride: 1, N: 256, H: 56, W: 56, G: 1, pad: 0 \| 461.6734\| 358.3779\| 1.48 IC: 128, OC: 512, kernel: 1, stride: 1, N: 256, H: 28, W: 28, G: 1, pad: 0 \| 358.3779 \| 247.8631\| 1.46 IC: 256, OC: 256, kernel: 3, stride: 1, N: 1, H: 16, W: 16, G: 1, pad: 0 \| 4.3783\| 3.8513\| 1.14 56 cores: Input \| fp32 ms \| bf16 internal ms \| Speed up -- \| -- \| -- \| -- IC: 64, OC: 256, kernel: 1, stride: 1, N: 256, H: 28, W: 28, G: 1, pad: 0 \| 16.6119 \| 12.2047 \| 1.38 IC: 128, OC: 512, kernel: 1, stride: 1, N: 256, H: 28, W: 28, G: 1, pad: 0 \| 12.0016 \| 8.6711 \| 1.38 IC: 256, OC: 1024, kernel: 1, stride: 1, N: 256, H: 14, W: 14, G: 1, pad: 0 \| 20.5947 \| 15.9366 \| 1.29 IC: 1024, OC: 256, kernel: 1, stride: 1, N: 256, H: 14, W: 14, G: 1, pad: 0 \| 40.0952 \| 32.2222 \| 1.24 IC: 256, OC: 256, kernel: 3, stride: 1, N: 1, H: 16, W: 16, G: 1, pad: 0 \| 162.7449 \| 142.3054 \| 1.14 Pull Request resolved: https://github.com/pytorch/pytorch/pull/126054 Approved by: https://github.com/jgong5 ghstack dependencies: #126050 Co-authored-by: Jiang, Yanbing <yanbing.jiang@intel.com>	2025-07-02 01:40:13 +00:00
haozhe.zhu	5a2db5152d	allow to use bf16 as fp32 internal precision for mkldnn conv (#126050 ) Allow to use `BF16` as the internal computation data types by `torch.backends.mkldnn.conv.fp32_precision="bf16"` ### TestPlan python test/test_mkldnn.py -k conv ### Benchmarking FP32 conv2d vs. BF16 internal computation conv2d on SPR Single core: Input \| fp32 ms \| bf16 internal ms \| Speed up -- \| -- \| -- \| -- IC: 64, OC: 256, kernel: 1, stride: 1, N: 256, H: 56, W: 56, G: 1, pad: 0 \| 185.5071 \| 83.4749 \| 2.22 IC: 128, OC: 512, kernel: 1, stride: 1, N: 256, H: 28, W: 28, G: 1, pad: 0 \| 194.7558 \| 79.1683\| 2.46 IC: 256, OC: 256, kernel: 3, stride: 1, N: 1, H: 16, W: 16, G: 1, pad: 0 \| 1.9213 \| 1.3690 \| 1.40 56 cores: Input \| fp32 ms \| bf16 internal ms \| Speed up -- \| -- \| -- \| -- IC: 64, OC: 256, kernel: 1, stride: 1, N: 256, H: 28, W: 28, G: 1, pad: 0 \| 6.5804 \| 7.4349 \| 0.89 IC: 128, OC: 512, kernel: 1, stride: 1, N: 256, H: 28, W: 28, G: 1, pad: 0 \| 4.9940 \| 3.8093 \| 1.31 IC: 256, OC: 1024, kernel: 1, stride: 1, N: 256, H: 14, W: 14, G: 1, pad: 0 \| 8.8359 \| 5.5802 \| 1.58 IC: 1024, OC: 256, kernel: 1, stride: 1, N: 256, H: 14, W: 14, G: 1, pad: 0 \| 16.5800 \| 9.2367 \| 1.80 IC: 256, OC: 256, kernel: 3, stride: 1, N: 1, H: 16, W: 16, G: 1, pad: 0 \| 79.5436 \| 38.3861 \| 2.07 Pull Request resolved: https://github.com/pytorch/pytorch/pull/126050 Approved by: https://github.com/jgong5, https://github.com/jansel Co-authored-by: Jiang, Yanbing <yanbing.jiang@intel.com>	2025-07-02 01:31:23 +00:00
Nikita Shulga	0a63053fe9	Don't store flamegraph to tmp folder (#157374 ) Where it's accessible(and mutable) by multiple users. Instead use `~/.cache` folder instead Pull Request resolved: https://github.com/pytorch/pytorch/pull/157374 Approved by: https://github.com/eqy ghstack dependencies: #157373	2025-07-02 00:46:51 +00:00
Animesh Jain	bb476310a4	[dynamo][guards] Stash root guard manager pointer in the LeafGuard (#157325 ) Preparing to simplify the recompilation reason codebase. This PR was 95% done by using AI tools. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157325 Approved by: https://github.com/jansel	2025-07-02 00:42:43 +00:00
Ankita George	fa1c20ae92	Fix test consolidate hf safetensors (#157386 ) Need to change an argument name that was changed in the test so that it doesn't throw Differential Revision: [D77604210](https://our.internmc.facebook.com/intern/diff/D77604210/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157386 Approved by: https://github.com/meetv18 ghstack dependencies: #154743, #156705	2025-07-02 00:16:21 +00:00
Sandeep Narendranath Karjala	77676753ec	[dynamo] Add fx_graph_runnable test coverage (#157021 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157021 Approved by: https://github.com/StrongerXi, https://github.com/xmfan Co-authored-by: Simon Fan <xmfan@meta.com>	2025-07-02 00:10:01 +00:00
Chong Gu	617e3f69f8	[FP8] Fix Benchmarking for certain Priors (#155722 ) Summary: For priors like layer norm, the order of the weight quantization kernel might be different and therefore have a different suffix, so we use regular expression instead. Test Plan: Trying this on model id 737772166 with ``` buck2 run mode/opt mode/inplace -c fbcode.platform010_cuda_version=12 -c fbcode.nvcc_arch=h100 caffe2/torch/fb/model_transform/experimental/benchmark:mts_gpu_benchmark -- --lower-backend=AOT_INDUCTOR --model-snapshot-id=737772166_0 --trace-aot-inductor-module=True --disable-acc-tracer=False --batch-size=1024 --node_replacement_dict "{'(autotune)':{'(1000+,1000+)':'fp8_float_model_dynamic_quantization_rowwise'}" ``` will allow more linears to be correctly replaced with fp8. An example of the gpu trace can be found in https://www.internalfb.com/intern/perfdoctor/trace_view?filepath=tree/hpc/new/models/feed/benchmark/libkineto_activities_773108_f58b57e208c04787acd3bcb01a3e8771.json.gz&bucket=gpu_traces. Rollback Plan: Differential Revision: D76092551 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155722 Approved by: https://github.com/Skylion007	2025-07-02 00:01:23 +00:00
PyTorch MergeBot	ab6cb34480	Revert "[inductor][user triton] sanitize triple-quoted docstrings in kernel definitions (#157322 )" This reverts commit 563fd95563c5edd732ae260b3bd3d0c38822ab57. Reverted https://github.com/pytorch/pytorch/pull/157322 on behalf of https://github.com/davidberard98 due to fails on rocm ([comment](https://github.com/pytorch/pytorch/pull/157322#issuecomment-3025826951))	2025-07-01 23:21:37 +00:00
PyTorch MergeBot	c6a27bae36	Revert "[do not revert] Compute contiguity symbolically to avoid dde, and introduce c++ sym_is_contiguous (#155590 )" This reverts commit d0a9629435aaceb5acbf31aad70f2109cb8a3ea2. Reverted https://github.com/pytorch/pytorch/pull/155590 on behalf of https://github.com/laithsakka due to was asked by to land this internally ([comment](https://github.com/pytorch/pytorch/pull/155590#issuecomment-3025796794))	2025-07-01 22:58:14 +00:00
David Berard	563fd95563	[inductor][user triton] sanitize triple-quoted docstrings in kernel definitions (#157322 ) Fixes #155006 Inductor sometimes codegens triton kernel definitions into a triple-quoted text block. If the text block itself contains triple-quotes, this breaks. Notably, this can happen for user-defined triton kernels, where the user may have added a docstring in their triton kernel. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157322 Approved by: https://github.com/zou3519, https://github.com/drisspg	2025-07-01 22:51:11 +00:00
PyTorch MergeBot	6ef70edd9a	Revert "Inductor logging + analysis of torch.profile (#149697 )" This reverts commit 47f10d0ad0dda281c886ff08ac2f938207027316. Reverted https://github.com/pytorch/pytorch/pull/149697 on behalf of https://github.com/malfet due to Looks like it's breaking ROCM tests, see https://hud.pytorch.org/hud/pytorch/pytorch/main/1?per_page=50&name_filter=rocm%20%2F%20linux-jammy ([comment](https://github.com/pytorch/pytorch/pull/149697#issuecomment-3025673908))	2025-07-01 22:11:53 +00:00
Xuehai Pan	3df6360e8c	[BE][Easy][setup] use `super().method(...)` in command subclasses in `setup.py` (#156044 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156044 Approved by: https://github.com/albanD ghstack dependencies: #156741	2025-07-01 22:09:10 +00:00
Laith Sakka	d0a9629435	[do not revert] Compute contiguity symbolically to avoid dde, and introduce c++ sym_is_contiguous (#155590 ) When we compute contiguity for a tensor with dynamic shapes we first: 1) Try to compute it without guarding. 2) If all shapes hinted, compute it with potentially adding guards. 3) if any input is not hinted, compute it symbolically. sym_is_contiguous return a SymBool that is then either evaluated or guard_or_false can be called on it to avoid data dependent errors. ex: bool is_contiguous = input.sym_is_contiguous().guard_or_false(__FILE__, __LINE__); is_contiguous_or_false is a helper function that does that. In this PR I only handle default contiguity, will follow up with changes for other formats like channel_last . We use this patter in this PR for several locations to avoid DDEs. Differential Revision: [D77183032](https://our.internmc.facebook.com/intern/diff/D77183032) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155590 Approved by: https://github.com/ezyang	2025-07-01 21:39:38 +00:00
Animesh Jain	22edb457c9	[invoke_subgraph][partitioner] Add meta val on run_and_save_rng ops (#157319 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157319 Approved by: https://github.com/zou3519	2025-07-01 21:02:08 +00:00
Nikita Shulga	e5f6ffd810	[BE] Replace `checkcall("chmod")` with `os.chmod` (#157373 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157373 Approved by: https://github.com/clee2000, https://github.com/eqy, https://github.com/Skylion007	2025-07-01 20:46:25 +00:00
Nikita Shulga	019e30e3b8	[BE] Decorate LargeTensorTest with serialTests (#157382 ) May be it'll help make M2-15 jobs more stable, as that was the last test run before OOM Pull Request resolved: https://github.com/pytorch/pytorch/pull/157382 Approved by: https://github.com/clee2000	2025-07-01 20:35:42 +00:00
Bob Ren	4500a4aa50	remove allow-untyped-defs from torch/backends/mps/__init__.py (#157227 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157227 Approved by: https://github.com/Skylion007	2025-07-01 20:00:19 +00:00
Ke Wen	6bc263809d	[SymmMem] Add NVSHMEM_CHECK macro (#157174 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157174 Approved by: https://github.com/fduwjj, https://github.com/fegin	2025-07-01 19:50:28 +00:00
angelayi	ffac0de07e	[export] Remove stack trace from input/output (#157302 ) Fixes https://github.com/pytorch/pytorch/issues/157183 https://github.com/pytorch/pytorch/pull/156257 consolidated the path for saving stack traces, but missed the part where stacktraces are not added to placeholder/output nodes in proxy_tensor tracing [(code)](https://github.com/pytorch/pytorch/pull/156257/files#diff-6960ce90e7162c0953b1ca07e92e7f0f2f6ba63b427b42df593e20cc6a096bb7L1107). Pull Request resolved: https://github.com/pytorch/pytorch/pull/157302 Approved by: https://github.com/yushangdi	2025-07-01 19:16:28 +00:00
Isuru Fernando	01b0f09931	Fix full_like decomposition to preserve strides (#144765 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/144765 Approved by: https://github.com/amjames, https://github.com/jansel	2025-07-01 19:13:22 +00:00
PyTorch MergeBot	6401d1d53d	Revert "Fused RMSNorm implementation (#153666 )" This reverts commit e1aee86646aa6d1b9cb9d34351e43936401c5efc. Reverted https://github.com/pytorch/pytorch/pull/153666 on behalf of https://github.com/davidberard98 due to causing build failures on main branch [GH job link](https://github.com/pytorch/pytorch/actions/runs/16007148842/job/45156382001) [HUD commit link](`e1aee86646`) ([comment](https://github.com/pytorch/pytorch/pull/153666#issuecomment-3025146176))	2025-07-01 18:46:45 +00:00
PyTorch MergeBot	3a5677a380	Revert "ci: Add ability to test images for build-triton-wheel (#156894 )" This reverts commit 0e47312ae5a687f0aed61db753d03180118cddc4. Reverted https://github.com/pytorch/pytorch/pull/156894 on behalf of https://github.com/seemethere due to causing issues in downstream builds see https://github.com/pytorch/pytorch/pull/156664 for more info ([comment](https://github.com/pytorch/pytorch/pull/156894#issuecomment-3025137790))	2025-07-01 18:43:34 +00:00
Jack Taylor	02608e560a	[ROCm] Add more shards for inductor dashboard, more frequent runs (#157288 ) Also increases regularity of dashboard runs on ROCm. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157288 Approved by: https://github.com/jeffdaily	2025-07-01 18:27:30 +00:00
AaronWang04	e1aee86646	Fused RMSNorm implementation (#153666 ) Relevant #72643 Benchmarked versus unfused torch implementation and torch.compile implementation. Around 9x speedup vs unfused implementation on cuda and slightly faster vs inductor compile on 5090. ```py import torch import torch.nn as nn class RMSNorm(nn.Module): def __init__(self, dim, eps=1e-5): super().__init__() self.eps = eps self.scale = nn.Parameter(torch.ones(dim)) def forward(self, x): norm_x = x.norm(2, dim=-1, keepdim=True) rms_x = norm_x * torch.rsqrt(torch.tensor(x.shape[-1], dtype=x.dtype)) x_normed = x / (rms_x + self.eps) return self.scale * x_normed def benchmark_rmsnorm_cuda(input_shape, normalized_dim, num_iterations=100, warmup_iterations=10, dtype=torch.float16): rms_norm_layer = torch.nn.RMSNorm(normalized_dim, device='cuda', dtype=dtype) input_data = torch.randn(input_shape, device='cuda', dtype=dtype) for _ in range(warmup_iterations): _ = rms_norm_layer(input_data) torch.cuda.synchronize() start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() for _ in range(num_iterations): _ = rms_norm_layer(input_data) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = elapsed_time_ms / num_iterations print(f"--- RMSNorm CUDA Benchmark ---") print(f"Input Shape: {input_shape}") print(f"Normalized Dimension: {normalized_dim}") print(f"Benchmark Iterations: {num_iterations}") print(f"--- Fused Implementation ---") print(f"Average Time per Iteration: {avg_time_ms:.4f} ms") print(f"Total Time for {num_iterations} Iterations: {elapsed_time_ms:.3f} ms") compiled_rms_norm = torch.compile(RMSNorm(dim=normalized_dim)).cuda() for _ in range(warmup_iterations): _ = compiled_rms_norm(input_data) torch.cuda.synchronize() start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() for _ in range(num_iterations): _ = compiled_rms_norm(input_data) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = elapsed_time_ms / num_iterations print(f"--- TorchCompile Implementation ---") print(f"Average Time per Iteration: {avg_time_ms:.4f} ms") print(f"Total Time for {num_iterations} Iterations: {elapsed_time_ms:.3f} ms") print("-" * 50) if __name__ == '__main__': parameter_sets = [ {'batch_size': 16, 'sequence_length': 256, 'hidden_features': 512, 'dtype': torch.float16}, {'batch_size': 32, 'sequence_length': 512, 'hidden_features': 768, 'dtype': torch.float16}, {'batch_size': 64, 'sequence_length': 1024, 'hidden_features': 1024, 'dtype': torch.float16}, {'batch_size': 32, 'sequence_length': 512, 'hidden_features': 768, 'dtype': torch.float32}, {'batch_size': 8, 'sequence_length': 2048, 'hidden_features': 2048, 'dtype': torch.float16}, ] num_benchmark_iterations = 200 num_warmup_iterations = 20 for params in parameter_sets: batch_size = params['batch_size'] sequence_length = params['sequence_length'] hidden_features = params['hidden_features'] data_type = params.get('dtype', torch.float16) shape = (batch_size, sequence_length, hidden_features) norm_dim_to_normalize = hidden_features print(f"Benchmarking with: BS={batch_size}, SeqLen={sequence_length}, Hidden={hidden_features}, DType={data_type}") benchmark_rmsnorm_cuda(input_shape=shape, normalized_dim=norm_dim_to_normalize, num_iterations=num_benchmark_iterations, warmup_iterations=num_warmup_iterations, dtype=data_type) ``` Here are the triton compile tests ran on a 5090 (comparing this branch vs main) ```py import torch import torch.nn as nn from torch._inductor.utils import run_and_get_code, run_fw_bw_and_get_code torch.manual_seed(0) device = torch.device("cuda") for batch in range(0, 9): for i in range(9, 16): normalized_shape_arg = (2batch, 2i) input_tensor = torch.randn(2batch, 2i, device=device, requires_grad=True) weight_tensor = torch.randn(2batch, 2i,device=device, requires_grad=True) model = torch.nn.functional.rms_norm compiled_model = torch.compile(model) loss = torch.randn_like(input_tensor) num_iter = 5 for j in range(num_iter): output = compiled_model(input_tensor, normalized_shape_arg, weight_tensor) output.backward(loss) start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) start_event.record() num_iter = 10 for j in range(num_iter): output = compiled_model(input_tensor, normalized_shape_arg, weight_tensor) output.backward(loss) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = round(elapsed_time_ms / num_iter, 5) print(2batch, 2i, avg_time_ms) ``` main ``` 32 512 0.1812 32 1024 0.19021 32 2048 0.18871 32 4096 0.17019 32 8192 0.21944 32 16384 0.38871 32 32768 0.83282 64 512 0.14705 64 1024 0.13987 64 2048 0.14111 64 4096 0.21699 64 8192 0.43141 64 16384 0.90652 64 32768 2.18573 128 512 0.19361 128 1024 0.1963 128 2048 0.20122 128 4096 0.38888 128 8192 0.93795 128 16384 2.23437 128 32768 5.50079 256 512 0.16722 256 1024 0.22856 256 2048 0.39421 256 4096 0.96621 256 8192 2.48746 256 16384 5.53571 256 32768 11.97932 ``` current branch ``` 32 512 0.16328 32 1024 0.18104 32 2048 0.15508 32 4096 0.14356 32 8192 0.20111 32 16384 0.45974 32 32768 0.94799 64 512 0.16874 64 1024 0.18701 64 2048 0.16107 64 4096 0.20152 64 8192 0.46568 64 16384 0.96599 64 32768 2.21661 128 512 0.14982 128 1024 0.15565 128 2048 0.22241 128 4096 0.46128 128 8192 0.88883 128 16384 2.3097 128 32768 5.84448 256 512 0.14346 256 1024 0.2007 256 2048 0.45927 256 4096 0.87876 256 8192 2.10571 256 16384 5.73948 256 32768 12.98581 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153666 Approved by: https://github.com/ngimel	2025-07-01 18:22:24 +00:00
Nikita Shulga	1c8844d9e7	[MPS] Switch Cholesky decomp to column wise (#157014 ) Everything should go thru a generalized kernels, and Metal kernels should work with the same sizes and strides as CPU or CUDA backends to avoid problems with `torch.compile` that relies on the meta kernels to tell what its ouput going to look like. To avoid returning tensors with different layout depending on whether upper parameter is true or false, templatize `factorDiagonalBlock`, `applyTRSM` and `applySYRK` to take upper/lower (actually row-wise vs column-wise) as template argument and call appropriate templates from host TODOs: - Rename upper parameter to something more sensible and add comments - Use simd_groupsize instead of hardcoded 32 everywhere Fixes https://github.com/pytorch/pytorch/issues/156658 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157014 Approved by: https://github.com/Skylion007, https://github.com/dcci ghstack dependencies: #157179	2025-07-01 18:00:59 +00:00
xinan.lin	720c2c46b1	[Inductor UT][XPU] Reduce the runtime of the test case test_comprehensive_nn_functional_max_pool2d_xpu. (#157357 ) This test case has over a thousand input samples, causing it to run for more than 30 minutes, which triggers the timeout mechanism and breaks the XPU CI. This PR limit the sample number as one for this XPU case . Pull Request resolved: https://github.com/pytorch/pytorch/pull/157357 Approved by: https://github.com/chuanqi129, https://github.com/jansel	2025-07-01 17:47:49 +00:00
Xuehai Pan	3bc6bdc866	[BE] add type annotations and run `mypy` on `setup.py` (#156741 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156741 Approved by: https://github.com/aorenste	2025-07-01 17:09:05 +00:00
Gabriel Ferns	47f10d0ad0	Inductor logging + analysis of torch.profile (#149697 ) Prereqs: - https://github.com/pytorch/pytorch/pull/152708 Features: 1. Adds inductor's estimate of flops and bandwidth to the json trace events that perfetto uses. 1. Only use the tflops estimation from triton if we don't have the info from the datasheet because Triton's estimates are inaccurate. I have a backlog item to fix triton flops estimation upstream. New `DeviceInfo` class, and new function `get_device_tflops`. 1. New helpers `countable_fx` and `count_flops_fx` helps get the flops of an `fx.Node`. 1. Extends Triton `torch.profiler` logging to `DebugAutotuner`. 1. New script `profile_analysis.py`: `--augment_trace` adds perf estimates to any perfetto json trace, `--analyze` creates a summary table of these perf estimates, and `--diff` will compare two traces side by side: ```python Device(NVIDIA H100, 0): Kernel Name \| resnet Kernel Count \| resnet FLOPS \| resnet bw gbps \| resnet Dur (ms) \| resnet Achieved FLOPS % \| resnet Achieved Bandwidth % \| newresnet Kernel Count \| newresnet FLOPS \| newresnet bw gbps \| newresnet Dur (ms) \| newresnet Achieved FLOPS % \| newresnet Achieved Bandwidth % --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- triton_poi_fused__native_batch_norm_legi \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 triton_red_fused__native_batch_norm_legi \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 triton_poi_fused__native_batch_norm_legi \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 void cutlass::Kernel2<cutlass_80_tensoro \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 triton_red_fused__native_batch_norm_legi \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 void cutlass::Kernel2<cutlass_80_tensoro \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 triton_poi_fused__native_batch_norm_legi \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 triton_red_fused__native_batch_norm_legi \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 triton_per_fused__native_batch_norm_legi \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 triton_poi_fused__native_batch_norm_legi \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 triton_per_fused__native_batch_norm_legi \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 triton_poi_fused__native_batch_norm_legi \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 void cublasLt::splitKreduce_kernel<32, 1 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 triton_per_fused__native_batch_norm_legi \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 triton_poi_fused__native_batch_norm_legi \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 void cask_plugin_cudnn::xmma_cudnn::init \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 triton_per_fused__native_batch_norm_legi \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 triton_per_fused__native_batch_norm_legi \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 void cutlass::Kernel2<cutlass_80_tensoro \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 triton_poi_fused__native_batch_norm_legi \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 triton_poi_fused__native_batch_norm_legi \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 triton_poi_fused__native_batch_norm_legi \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 triton_per_fused__native_batch_norm_legi \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 triton_poi_fused__native_batch_norm_legi \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 triton_per_fused__native_batch_norm_legi \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 triton_poi_fused__native_batch_norm_legi \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 triton_per_fused__native_batch_norm_legi \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 std::enable_if<!(false), void>::type int \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 triton_poi_fused_add_copy__38 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 triton_poi_fused_convolution_0 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 triton_poi_fused_convolution_1 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 void convolve_common_engine_float_NHWC<f \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 triton_per_fused__native_batch_norm_legi \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 triton_per_fused__native_batch_norm_legi \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 triton_poi_fused__native_batch_norm_legi \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 void at::native::(anonymous namespace):: \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 void at::native::vectorized_elementwise_ \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/149697 Approved by: https://github.com/eellison, https://github.com/shunting314	2025-07-01 16:51:03 +00:00
zhxchen17	0f9c1b374f	[dynamo] Ensure global state guard is preserved across serialization. (#157285 ) Currently, every time we construct a GLOBAL_STATE guard, we always create a fresh guard based on the current global state. For precompile, we want to create a GLOBAL_STATE guard always based on some external sources, e.g. serialized global states. This can also be applied with the normal case where we just pass in the global state guard from Python. Differential Revision: [D77400988](https://our.internmc.facebook.com/intern/diff/D77400988/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157285 Approved by: https://github.com/jansel	2025-07-01 15:46:34 +00:00
Xuehai Pan	b146e1a264	[BE] remove duplicates in generated `torch._VF.__all__` (#157365 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157365 Approved by: https://github.com/Skylion007	2025-07-01 15:43:20 +00:00
zhxchen17	c78fce9e79	[dynamo] show frame information when recompilation is triggered on fail_on_recompile (#156433 ) adding more information to the error message for debugging. example error message: ``` Detected recompile when torch.compile stance is 'fail_on_recompile'. filename: 'caffe2/test/dynamo/test_misc.py', function name: 'fn', line number: 0 Failed on the following precompiled guards: TREE_GUARD_MANAGER: +- RootGuardManager \| +- LAMBDA_GUARD: isinstance(L['x'], bool) GuardDebugInfo( result=0, verbose_code_parts=["isinstance(L['x'], bool)"], num_guards_executed=1) ``` Differential Revision: [D76987126](https://our.internmc.facebook.com/intern/diff/D76987126/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156433 Approved by: https://github.com/jamesjwu	2025-07-01 15:15:58 +00:00
PyTorch MergeBot	023887fc5a	Revert "Switch to standard pep517 sdist generation (#152098 )" This reverts commit f16053f0c9a09fa337fbf85aaf64f88712b8dcdb. Reverted https://github.com/pytorch/pytorch/pull/152098 on behalf of https://github.com/malfet due to IMO this PR needs to be split into few helper ones, with better test plan ([comment](https://github.com/pytorch/pytorch/pull/152098#issuecomment-3024223880))	2025-07-01 14:14:52 +00:00
PyTorch MergeBot	1586521461	Revert "Compute contiguity symbolically to avoid dde, and introduce c++ sym_is_contiguous (#155590 )" This reverts commit 2c76f31221e117b217b8a6a96a5405f626d2218a. Reverted https://github.com/pytorch/pytorch/pull/155590 on behalf of https://github.com/jeanschmidt due to Breaking 1000s of internal builds, it cant be properly landed internally, there are no options except revert and codev. ([comment](https://github.com/pytorch/pytorch/pull/155590#issuecomment-3023503929))	2025-07-01 11:23:00 +00:00
PyTorch MergeBot	534c454e77	Revert "[xla hash update] update the pinned xla hash (#156584 )" This reverts commit b1a54fab9bcb0cc167773f9a885d4170447e1c68. Reverted https://github.com/pytorch/pytorch/pull/156584 on behalf of https://github.com/jeanschmidt due to Need to revert in order to revert https://github.com/pytorch/pytorch/pull/155590 ([comment](https://github.com/pytorch/pytorch/pull/156584#issuecomment-3023492421))	2025-07-01 11:20:05 +00:00
PyTorch MergeBot	13bf2655c1	Revert "HF loads dcp - don't do a full deserialize on every file (#155942 )" This reverts commit 117db5601d78cbc746b35eef71fc815e042e903f. Reverted https://github.com/pytorch/pytorch/pull/155942 on behalf of https://github.com/jeanschmidt due to Newly introduced tests are red internally, more details on D76442012 ([comment](https://github.com/pytorch/pytorch/pull/155942#issuecomment-3023473036))	2025-07-01 11:15:08 +00:00
PyTorch MergeBot	0bce390269	Revert "[dynamo] Add fx_graph_runnable test coverage (#157021 )" This reverts commit 20e40492b046b9287726d3ec656117e4dc38f0e2. Reverted https://github.com/pytorch/pytorch/pull/157021 on behalf of https://github.com/jeanschmidt due to New tests are red internally, more details on D77471538 ([comment](https://github.com/pytorch/pytorch/pull/157021#issuecomment-3023455082))	2025-07-01 11:10:45 +00:00
Bob Ren	a767e50adc	remove allow-untyped-defs from torch/fx/experimental/migrate_gradual_types/util.py (#157236 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157236 Approved by: https://github.com/ezyang	2025-07-01 10:36:48 +00:00
Jeff Daily	210632fae1	[ROCm] support experimental CU carveout (#149466 ) Fixes #149280. Follow up to #147966, but now available for ROCm. Since hipblaslt does not support HIPBLASLT_MATMUL_DESC_CU_COUNT_TARGET, we instead create a hipStream that has a CU mask applied. We pass this masked stream to hipblaslt instead of pytorch's current stream. We ensure stream ordering between streams using hipEvents and stream synchronization. Pull Request resolved: https://github.com/pytorch/pytorch/pull/149466 Approved by: https://github.com/malfet, https://github.com/atalman	2025-07-01 08:54:52 +00:00
Jason Ansel	0596323c35	Better fix for `__index__` SymInt issue (#157201 ) This improves on #156928 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157201 Approved by: https://github.com/ezyang	2025-07-01 07:06:46 +00:00
PyTorch MergeBot	c202a7329a	Revert "Fixes for CPython int/float tests (#155978 )" This reverts commit 23491519d288dedb2a54cfad5fef7fcb2ad8eade. Reverted https://github.com/pytorch/pytorch/pull/155978 on behalf of https://github.com/XuehaiPan due to sys.get_int_max_str_digits is not always available ([comment](https://github.com/pytorch/pytorch/pull/155978#issuecomment-3021990027))	2025-07-01 06:16:49 +00:00
Xuehai Pan	754699610b	[BE] always use `uv pip` if possible in `pip_init.py` for `lintrunner init` (#157199 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157199 Approved by: https://github.com/ezyang	2025-07-01 06:07:29 +00:00
Isuru Fernando	8f0998aafe	Check F2C BLAS for OpenBLAS and other vendors (#143846 ) This issue came from https://github.com/conda-forge/pytorch-cpu-feedstock/issues/180. MKL follows the F2C convention for returning single precision floats as doubles and uses the G77 convention for returning complex valued scalars. OpenBLAS does the opposite. There is a check for this already, but it's done only when the Generic BLAS vendor code path is used and this PR moves that code to `Dependencies.cmake` to make it work when the BLAS vendor is OpenBLAS and others Pull Request resolved: https://github.com/pytorch/pytorch/pull/143846 Approved by: https://github.com/rgommers, https://github.com/atalman	2025-07-01 05:56:24 +00:00
Ethan Wee	04bd7e6850	[ROCm] Remove use of `warpsize` on host-side compilation (#156979 ) Changes needed for ROCm7.0: * `warpSize` is _not_ a compile-time constant on device-side compilation for ROCm anymore * `warpSize` is _not_ defined on host-side compilation, hence `at::cuda::warp_size()` must be used to query warpsize at runtime * Redefining `C10_WARP_SIZE` to be a compile-time constant, with a reasonable value for device-side compilation, but an unreasonable value of 1 for host-side compilation Pull Request resolved: https://github.com/pytorch/pytorch/pull/156979 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-07-01 04:55:31 +00:00
Nikita Shulga	c811f41cf5	[BE] Remove unused variable from Pooling.metal (#157332 ) Fixes following compilation warning ``` /Users/nshulga/git/pytorch/pytorch/aten/src/ATen/native/mps/kernels/Pooling.metal:101:21: warning: unused variable 'indices_sizes' [-Wunused-variable] constant int64_t* indices_sizes = params.indices_sizes.data(); ^ ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157332 Approved by: https://github.com/clee2000, https://github.com/huydhn, https://github.com/dcci	2025-07-01 04:28:04 +00:00
drisspg	4d5d627e5f	Remove super spammy log (#157157 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157157 Approved by: https://github.com/davidberard98	2025-07-01 03:51:58 +00:00
Jason Ansel	b40981c630	Fix incorrect stride handling in adaptive_avg_pool3d (#157326 ) Fixes #157248 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157326 Approved by: https://github.com/eqy ghstack dependencies: #157242	2025-07-01 03:03:48 +00:00
Andy Lugo	b5ce77c1f5	[ROCm] Initial AITER Integration for mha_bwd asm kernels (#152630 ) Generates AITER plumbing via cmake. Calls into fav3 asm bwd CK kernels. Update submodule composable kernel for this change Pull Request resolved: https://github.com/pytorch/pytorch/pull/152630 Approved by: https://github.com/xw285cornell, https://github.com/yoyoyocmu	2025-07-01 02:53:27 +00:00
Catherine Lee	f40efde2a4	[CI] Add prebuild command option, set prebuild command option for CI to build flash attention (#156236 ) Build flash attention separately in build using 2 jobs since it OOMs on more, then the rest of the job uses 6 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156236 Approved by: https://github.com/malfet	2025-07-01 02:53:22 +00:00
Sidharth	3ed4384f5b	[dynamo] temporarily disabling generation of weblinks for torch v2.8 release (#157299 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157299 Approved by: https://github.com/williamwen42	2025-07-01 02:31:17 +00:00
Ti-Tai Wang	c174f3a6a5	[ONNX] Delete deprecated tutorial page link (#157310 ) Related to https://github.com/pytorch/tutorials/issues/3420 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157310 Approved by: https://github.com/justinchuby	2025-07-01 01:18:26 +00:00
Prachi Gupta	6dc2b22269	[ROCm][SymmetricMemory] Performance improvements for two-shot allreduce (#156746 ) The biggest bottleneck that we found with two-shot allreduce was that the compiler was serializing all the load operations for some reason. To avoid these load delays, we've added de-serialization of loads. Along with this improvement, we also found that on AMD GPUs a different block and thread size gives a nice performance boost. Here are the bandwidth numbers I am getting with this PR: ![image](https://github.com/user-attachments/assets/57005856-4cb5-43cd-8e9c-46869f75ab0b) The rows that are green are the tensor sizes that we are interested in because two-shot is only used for bigger sizes (one-shot is used for smaller sizes). As we can see, our baseline numbers wrt to fbgemm numbers were consistently underperforming. However, with this deserialize change, most of the tensor sizes have a performance boost (positive %) for the green tensors. There's one tensor with negative performance, but that's within error margin. co-authored by: @amd-hhashemi https://github.com/pytorch/FBGEMM/issues/4072 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156746 Approved by: https://github.com/jeffdaily Co-authored-by: Hashem Hashemi <hashem.hashemi@amd.com>	2025-07-01 00:37:30 +00:00
fuwenguang	f860992db5	Add a custom profiler configuration option (#151656 ) We aim to pass some configuration options to our custom Kineto backend via ExperimentalConfig,, so we added a `custom_profiler_config` parameter. Requires https://github.com/pytorch/kineto/pull/1077 , Pull Request resolved: https://github.com/pytorch/pytorch/pull/151656 Approved by: https://github.com/sraikund16	2025-07-01 00:36:09 +00:00
Ankita George	b60569ed94	HF - consolidate shards of safetensors files to full tensors in finish step (#156705 ) Title - we can consolidate the shards to a full tensors, optionally behind a flag, in the finish step of DCP.save also adds the thread count argument which is configurable for users, before we were just using the default of 1. Re-creating https://github.com/pytorch/pytorch/pull/155940 bc it got into a bad detached state Differential Revision: [D77231774](https://our.internmc.facebook.com/intern/diff/D77231774/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156705 Approved by: https://github.com/saumishr ghstack dependencies: #154743	2025-07-01 00:30:48 +00:00
Nikita Shulga	4ebd269065	[Testing] Remove duplicate MPSInductor tests (#157328 ) They were added there before test_torchinductor were running in CI, but now the same are covered by `GPUTests.test_pointwise_*_mps` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157328 Approved by: https://github.com/huydhn	2025-07-01 00:21:22 +00:00
Bob Ren	7709ff5512	[remove untyped defs] batch 1 (#157011 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157011 Approved by: https://github.com/Skylion007	2025-06-30 23:54:40 +00:00
Scott Wolchok	fee2377f9e	Reapply D77381084 / #156964 : Rename torch::standalone to headeronly (#157251 ) Was reverted due to internal failure which should be fixed now. I believe Jane wants this reapplied and picked to release, and she's out this week. Original summary: headeronly is more clear, let's change the name before anyone depends on standalone Differential Revision: [D77520173](https://our.internmc.facebook.com/intern/diff/D77520173/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157251 Approved by: https://github.com/janeyx99, https://github.com/Skylion007, https://github.com/desertfire	2025-06-30 23:25:30 +00:00
Aaron Ang	3dda80e990	Overload `mul_overflows` for `size_t` (#155736 ) Partially fixes https://github.com/pytorch/executorch/pull/11537. We want to extend `mul_overflows` to support `size_t` in ExecuTorch. The current workflow in ET checks that the `c10` mirrors exactly as in PT, so the tests are failing. See comment: https://github.com/pytorch/executorch/pull/11537#issuecomment-2963821312 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155736 Approved by: https://github.com/swolchok	2025-06-30 22:57:28 +00:00
Animesh Jain	42b48ee672	[dynamo][fsdp] Consistent behavior of int attributes (#157262 ) Reimpl of https://github.com/pytorch/pytorch/pull/150954 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157262 Approved by: https://github.com/bdhirsh	2025-06-30 22:32:52 +00:00
Ankita George	a9352bd25e	Script for consolidation of sharded safetensor files (#154743 ) Script to consolidate sharded safetensors files with DCP into full tensors. This relies on file system operations to read and copy bytes directly instead of the traditional approach of loading and re-sharding and then saving again, because users will have models that are larger than allotted memory. Differential Revision: [D75536985](https://our.internmc.facebook.com/intern/diff/D75536985/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154743 Approved by: https://github.com/saumishr	2025-06-30 22:25:58 +00:00
zhxchen17	f096820d0f	[precompile] Detect source code changes for save/load. (#156432 ) Go through all dynamo traced functions and compute checksum for them. While loading a precompilation back to memory, we will always check the checksum and refuse to load when source code changes are detected. Differential Revision: [D76987123](https://our.internmc.facebook.com/intern/diff/D76987123/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156432 Approved by: https://github.com/jansel, https://github.com/jamesjwu	2025-06-30 21:16:15 +00:00
PyTorch MergeBot	d3efd73234	Revert "[cutlass backend][BE][ez] Make matmul layouts be row x column (#156656 )" This reverts commit 84c588e5eada9e7921608065edc444a15c22cb1c. Reverted https://github.com/pytorch/pytorch/pull/156656 on behalf of https://github.com/henrylhtsang due to breaking fbcode A100 tests ([comment](https://github.com/pytorch/pytorch/pull/156656#issuecomment-3020769914))	2025-06-30 21:16:04 +00:00
Animesh Jain	3684be056d	[dynamo] Fix source for lru_cache method (#157292 ) Fixes - https://github.com/pytorch/pytorch/issues/157273 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157292 Approved by: https://github.com/zou3519, https://github.com/malfet, https://github.com/jansel	2025-06-30 20:53:57 +00:00
Guilherme Leobas	23491519d2	Fixes for CPython int/float tests (#155978 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155978 Approved by: https://github.com/zou3519	2025-06-30 19:42:11 +00:00
Klaus Zimmermann	f16053f0c9	Switch to standard pep517 sdist generation (#152098 ) Generate source tarball with PEP 517 conform build tools instead of the custom routine in place right now. Closes #150461. The current procedure for generating the source tarball consists in creation of a source tree by manual copying and pruning of source files. This PR replaces that with a call to the standard [build tool](https://build.pypa.io/en/stable/), which works with the build backend to produce an sdist. For that to work correctly, the build backend also needs to be configured. In the case of Pytorch, the backend currently is (the legacy version of) the setuptools backend, the source dist part of which is mostly configured via the `MANIFEST.in` file. The resulting source distribution can be used to install directly from source with `pip install ./torch-{version}.tar.gz` or to build wheels directly from source with `pip wheel ./torch-{version}.tar.gz`; both should be considered experimental for now. ## Issues ### sdist name According to PEP 517, the name of the source distribution file must coincide with the project name, or [more precisely](https://peps.python.org/pep-0517/#source-distributions), the source distribution of a project that generates `{NAME}-{...}.whl` wheels are required to be named `{NAME}-{...}.tar.gz`. Currently, the source tarball is called `pytorch-{...}.tar.gz`, but the generated wheels and python package are called `torch-{...}`. ### Symbolic Links The source tree at the moment contains a small number of symbolic links. This [has been seen as problematic](https://github.com/pypa/pip/issues/5919) largely because of lack of support on Windows, but also because of [a problem in setuptools](https://github.com/pypa/setuptools/issues/4937). Particularly unfortunate is a circular symlink in the third party `ittapi` module, which can not be resolved by replacing it with a copy. PEP 721 (now integrated in the [Source Distribution Format Specification](https://packaging.python.org/en/latest/specifications/source-distribution-format/#source-distribution-archive-features)) allows for symbolic links, but only if they don't point outside the destination directory and if they don't contain `../` in their target. The list of symbolic links currently is as follows: <details> \|source\|target\|problem\|solution\| \|-\|-\|-\|-\| \| `.dockerignore` \| `.gitignore` \| ✅ ok (individual file) \|\| \| `docs/requirements.txt` \| `../.ci/docker/requirements-docs.txt` \|❗`..` in target\|swap source and target[^1]\| \| `functorch/docs/source/notebooks` \| `../../notebooks/` \|❗`..` in target\|swap source and target[^1]\| \| `.github/ci_commit_pins/triton.txt` \| `../../.ci/docker/ci_commit_pins/triton.txt` \| ✅ ok (omitted from sdist)\|\| \| `third_party/flatbuffers/docs/source/CONTRIBUTING.md` \| `../../CONTRIBUTING.md` \|❗`..` in target\|omit from sdist[^2]\| \| `third_party/flatbuffers/java/src/test/java/DictionaryLookup` \| `../../../../tests/DictionaryLookup` \|❗`..` in target\|omit from sdist[^3]\| \| `third_party/flatbuffers/java/src/test/java/MyGame` \| `../../../../tests/MyGame` \|❗`..` in target\|omit from sdist[^3]\| \| `third_party/flatbuffers/java/src/test/java/NamespaceA` \| `../../../../tests/namespace_test/NamespaceA` \|❗`..` in target\|omit from sdist[^3]\| \| `third_party/flatbuffers/java/src/test/java/NamespaceC` \| `../../../../tests/namespace_test/NamespaceC` \|❗`..` in target\|omit from sdist[^3]\| \| `third_party/flatbuffers/java/src/test/java/optional_scalars` \| `../../../../tests/optional_scalars` \|❗`..` in target\|omit from sdist[^3]\| \| `third_party/flatbuffers/java/src/test/java/union_vector` \| `../../../../tests/union_vector` \|❗`..` in target\|omit from sdist[^3]\| \| `third_party/flatbuffers/kotlin/benchmark/src/jvmMain/java` \| `../../../../java/src/main/java` \|❗`..` in target\|omit from sdist[^3]\| \| `third_party/ittapi/rust/ittapi-sys/c-library` \| `../../` \|❗`..` in target\|omit from sdist[^4]\| \| `third_party/ittapi/rust/ittapi-sys/LICENSES` \| `../../LICENSES` \|❗`..` in target\|omit from sdist[^4]\| \| `third_party/opentelemetry-cpp/buildscripts/pre-merge-commit` \| `./pre-commit` \|✅ ok (individual file)\|\| \| `third_party/opentelemetry-cpp/third_party/prometheus-cpp/cmake/project-import-cmake/sample_client.cc` \| `../../push/tests/integration/sample_client.cc` \|❗`..` in target\|omit from sdist[^5]\| \| `third_party/opentelemetry-cpp/third_party/prometheus-cpp/cmake/project-import-cmake/sample_server.cc` \| `../../pull/tests/integration/sample_server.cc` \|❗`..` in target\|omit from sdist[^5]\| \| `third_party/opentelemetry-cpp/third_party/prometheus-cpp/cmake/project-import-pkgconfig/sample_client.cc` \| `../../push/tests/integration/sample_client.cc` \|❗`..` in target\|omit from sdist[^5]\| \| `third_party/opentelemetry-cpp/third_party/prometheus-cpp/cmake/project-import-pkgconfig/sample_server.cc` \| `../../pull/tests/integration/sample_server.cc` \|❗`..` in target\|omit from sdist[^5]\| \| `third_party/XNNPACK/tools/xngen` \| `xngen.py` \| ✅ ok (individual file)\|\| </details> The introduction of symbolic links inside the `.ci/docker` folder creates a new problem, however, because Docker's `COPY` command does not allow symlinks in this way. We work around that by using `tar ch` to dereference the symlinks before handing them over to `docker build`. [^1]: These resources can be naturally considered to be part of the docs, so moving the actual files into the place of the current symlinks and replacing them with (unproblematic) symlinks can be said to improve semantics as well. [^2]: The flatbuffers docs already actually use the original file, not the symlink and in the most recent releases, starting from flatbuffers-25.1.21 the symlink is replaced by the actual file thanks to a documentation overhaul. [^3]: These resources are flatbuffers tests for java and kotlin and can be omitted from our sdist. [^4]: We don't need to ship the rust bindings for ittapi. [^5]: These are demonstration examples for how to link to prometheus-cpp using cmake and can be omitted. ### Nccl Nccl used to be included as a submodule. However, with #146073 (first released in v2.7.0-rc1), the submodule was removed and replaced with a build time checkout procedure in `tools/build_pytorch_libs.py`, which checks out the required version of nccl from the upstream repository based on a commit pin recorded in `.ci/docker/ci_commit_pins/nccl-cu{11,12}.txt`. This means that a crucial third party dependency is missing from the source distribution and as the `.ci` folder is omitted from the source distribution, it is not possible to use the build time download. However, it is possible to use a system provided Nccl using the `USE_SYSTEM_NCCL` environment variable, which now also is the default for the official Pytorch wheels. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152098 Approved by: https://github.com/atalman	2025-06-30 19:07:34 +00:00
Wanchao Liang	c7b6c98d10	[tp] improve parallelize_module API to support more cases (#157182 ) This PR improves the parallelize_module API to support more corner cases: 1. if the plan entry specified as "", it should apply the style to the current module 2. if the plan entry does not have a corresponding submodule to apply, raise a warning and ignore this plan entry As working on this PR, I also found that the while-loop inside is actually not necessary and could produce some nasty on the fly modifying while iterating behavior.. So I removed the while loop Pull Request resolved: https://github.com/pytorch/pytorch/pull/157182 Approved by: https://github.com/tianyu-l	2025-06-30 18:10:44 +00:00
PyTorch MergeBot	d5e6f42094	Revert "Use std::string_view in torchgen (#157050 )" This reverts commit 064288cbab94c9931ca2296a2b9723e864f9050a. Reverted https://github.com/pytorch/pytorch/pull/157050 on behalf of https://github.com/jeanschmidt due to Seems to have broken internal builds, more details on D77449943. @ezyang may I count on your help to get those changes merged? ([comment](https://github.com/pytorch/pytorch/pull/157050#issuecomment-3020222668))	2025-06-30 18:08:54 +00:00
PyTorch MergeBot	efbf07e7ea	Revert "[dynamo] Fix issue with tensors passed as view() shapes (#156928 )" This reverts commit 75f3e5a88df60caef27fd9c9df3fd51161378fcc. Reverted https://github.com/pytorch/pytorch/pull/156928 on behalf of https://github.com/jeanschmidt due to Breaks a internal test, more details can be found on D77449971 ([comment](https://github.com/pytorch/pytorch/pull/156928#issuecomment-3020186268))	2025-06-30 17:56:01 +00:00
Avanish Tiwari	5e18bc3331	[PowerPC] Fixed build issue for vsx vec256 complexfloat and scaled_mm_out_cpu (#155255 ) Pytorch build is failing on power system from this commit ec24f8f58a74502c5a2488f5d9e85a817616dda0 *Build Failure Logs* Error related to mkldnn ``` pytorch/aten/src/ATen/native/Blas.cpp:302:26: error: ‘cpuinfo_has_x86_amx_int8’ was not declared in this scope 302 \| if ((!mixed_dtype && cpuinfo_has_x86_amx_int8()) \|\| \| ^~~~~~~~~~~~~~~~~~~~~~~~ pytorch/aten/src/ATen/native/Blas.cpp:303:25: error: ‘cpuinfo_has_x86_amx_fp16’ was not declared in this scope 303 \| (mixed_dtype && cpuinfo_has_x86_amx_fp16())) { \| ^~~~~~~~~~~~~~~~~~~~~~~~ ``` Error related to vec256 complex float redefinition ``` aten/src/ATen/cpu/vec/vec256/vsx/vec256_complex_float_vsx.h:19:7: error: specialization of ‘at::vec::DEFAULT::Vectorized<c10::complex<float> >’ after instantiation 19 \| class Vectorized<ComplexFlt> { \| ^~~~~~~~~~~~~~~~~~~~~~ aten/src/ATen/cpu/vec/vec256/vsx/vec256_complex_float_vsx.h:19:7: error: redefinition of ‘class at::vec::DEFAULT::Vectorized<c10::complex<float> >’  aten/src/ATen/cpu/vec/vec256/vsx/vec256_complex_float_vsx.h:633:18: error: ‘const class at::vec::DEFAULT::Vectorized<c10::complex<float> >’ has no member named ‘abs_2_’ 633 \| auto abs_a = a.abs_2_(); \| ^~~~~~ aten/src/ATen/cpu/vec/vec256/vsx/vec256_complex_float_vsx.h:634:18: error: ‘const class at::vec::DEFAULT::Vectorized<c10::complex<float> >’ has no member named ‘abs_2_’ 634 \| auto abs_b = b.abs_2_(); \| ^~~~~~ /aten/src/ATen/cpu/vec/vec256/vsx/vec256_complex_float_vsx.h:666:17: error: ‘const class at::vec::DEFAULT::Vectorized<c10::complex<float> >’ has no member named ‘vec0’ 666 \| vec_add(a.vec0(), b.vec0()), vec_add(a.vec1(), b.vec1())}; aten/src/ATen/cpu/vec/vec256/vsx/vec256_complex_float_vsx.h:673:17: error: ‘const class at::vec::DEFAULT::Vectorized<c10::complex<float> >’ has no member named ‘vec0’ 673 \| vec_sub(a.vec0(), b.vec0()), vec_sub(a.vec1(), b.vec1())}; \| ^~~~ aten/src/ATen/cpu/vec/vec256/vsx/vec256_complex_float_vsx.h:680:27: error: ‘const class at::vec::DEFAULT::Vectorized<c10::complex<float> >’ has no member named ‘vec0’ 680 \| vec_and(a.vec0(), b.vec0()), vec_and(a.vec1(), b.vec1())}; ``` *With this changes build logs* ``` Building wheel torch-2.8.0a0+gita3098a7 -- Building version 2.8.0a0+gita3098a7 -- Checkout nccl release tag: v2.26.5-1 cmake -GNinja -DBLAS=OpenBLAS -DBUILD_PYTHON=True -DBUILD_TEST=True -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/home/avanish/OfficeWork2025/JuneWork/pytorch_5Jun/pack/torch_night_5Jun/pytorch/torch -DCMAKE_PREFIX_PATH=/home/avanish/OfficeWork2025/JuneWork/pyenv/pytorch_5Jun/lib/python3.12/site-packages -DPython_EXECUTABLE=/home/avanish/OfficeWork2025/JuneWork/pyenv/pytorch_5Jun/bin/python -DTORCH_BUILD_VERSION=2.8.0a0+gita3098a7 -DUSE_MKLDNN=ON -DUSE_MKLDNN_CBLAS=ON -DUSE_NUMPY=True -DUSE_OPENMP=ON /home/avanish/OfficeWork2025/JuneWork/pytorch_5Jun/pack/torch_night_5Jun/pytorch cmake --build . --target install --config Release running build_ext -- Building with NumPy bindings -- Not using cuDNN -- Not using CUDA -- Not using XPU -- Using MKLDNN -- Not using Compute Library for the Arm architecture with MKLDNN -- Using CBLAS in MKLDNN -- Not using NCCL -- Building with distributed package: -- USE_TENSORPIPE=True -- USE_GLOO=True -- USE_MPI=False -- Building Executorch -- Not using ITT Copying functorch._C from functorch/functorch.so to /home/avanish/OfficeWork2025/JuneWork/pytorch_5Jun/pack/torch_night_5Jun/pytorch/build/lib.linux-ppc64le-cpython-312/functorch/_C.cpython-312-powerpc64le-linux-gnu.so copying functorch/functorch.so -> /home/avanish/OfficeWork2025/JuneWork/pytorch_5Jun/pack/torch_night_5Jun/pytorch/build/lib.linux-ppc64le-cpython-312/functorch/_C.cpython-312-powerpc64le-linux-gnu.so building 'torch._C' extension creating build/temp.linux-ppc64le-cpython-312/torch/csrc ``` This patch will fix the pytorch build issue on power, and i am able to build successfully. Hi @malfet @albanD Please review this PR for pytorch build issue that we are observing on power. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155255 Approved by: https://github.com/albanD, https://github.com/malfet	2025-06-30 17:54:37 +00:00
Wanchao Liang	2815eea0d0	[dtensor] relax device_mesh argument constraint in local_map (#157049 ) This PR relaxes the device_mesh argument constraint in the local_map API. The current restriction is too strict, i.e. all the input arguments must have the same device mesh if they are DTensors. But many times user might want to pass in DTensors to this function that lives on different device mesh, i.e. weight and activation could live in different device mesh. When using the local_map, we are extracting the local tensors from DTensors, and as long as the placements user specified match with the actual DTensor placements, user knows clearly that the inputs are intended to live in different mesh. So this PR removes the same mesh check and update doc to clearly document the behavior. The `device_mesh` argument now serves for a main purpose, allow user to specify the device_mesh for the output DTensor reconstruction Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/157049 Approved by: https://github.com/Chillee, https://github.com/zpcore	2025-06-30 17:51:48 +00:00
Jason Ansel	f8cc4c0af8	[inductor] Update triton_key import to support latest Triton (#157242 ) With Triton main things were failing with: ```py File "/home/jansel/pytorch/torch/_inductor/codecache.py", line 205, in get_system from triton.compiler.compiler import triton_key torch._dynamo.exc.BackendCompilerFailed: backend='inductor' raised: ImportError: cannot import name 'triton_key' from 'triton.compiler.compiler' (/home/jansel/pytorch/triton/compiler/compiler.py) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157242 Approved by: https://github.com/aorenste	2025-06-30 17:51:43 +00:00
Ankita George	117db5601d	HF loads dcp - don't do a full deserialize on every file (#155942 ) Differential Revision: [D76442012](https://our.internmc.facebook.com/intern/diff/D76442012/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155942 Approved by: https://github.com/saumishr ghstack dependencies: #155707	2025-06-30 17:45:10 +00:00
Laith Sakka	ed5d6d2a20	python definitely_contiguous-> is_contiguous_or_false (#156515 ) We probably can avoid having those in python as well and just depend on c++ impl after we land https://github.com/pytorch/pytorch/pull/155590 but that is for a different PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156515 Approved by: https://github.com/bobrenjc93	2025-06-30 17:31:51 +00:00
PyTorch MergeBot	c038719731	Revert "Inductor logging + analysis of torch.profile (#149697 )" This reverts commit 347ace4c7ac2dbb14799089c30bd01a9ac312791. Reverted https://github.com/pytorch/pytorch/pull/149697 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it seems to fail on ROCm ([comment](https://github.com/pytorch/pytorch/pull/149697#issuecomment-3020006655))	2025-06-30 16:58:54 +00:00
Yukio Siraichi	b54eac2a5e	Upgrade to DLPack 1.0. (#145000 ) This PR makes the necessary changes in order to upgrade PyTorch DLPack support to version 1.0. In summary, we add support for the following: - Support both `DLManagedTensor` and `DLManagedTensorVersioned` when producing and consuming DLPack capsules - New parameter for `__dlpack__` method: `max_version` - Version checks: - Fallback to old implementation if no `max_version` or if version lower than 1.0 - Check that the to-be-consumed capsule is of version up to 1.X In order to accommodate these new specifications, this PR adds the following main changes: - `torch._C._to_dlpack_versioned` Python API (Module.cpp): new Python API for creating a versioned DLPack capsule (called by `__dlpack__` method) - `DLPackTraits<T>` class (DLConvertor.h): select the correct traits (e.g. capsule name, conversion functions) depending on which DLPack tensor class is being used - `toDLPackImpl<T>` function (DLConvertor.cpp): populates the common fields of both classes - `fromDLPackImpl<T>` function (DLConvertor.cpp): constructs a tensor from a DLPAck capsule - `fillVersion<T>` function (DLConvertor.cpp): populates the version field for `DLManagedTensorVersioned` (no-op for `DLManagedTensor`) - `tensor_fromDLPackImpl<T>` function (tensor_new.cpp): outer function for constructing a tensor out of a DLPack capsule that also marks the capsule as used Pull Request resolved: https://github.com/pytorch/pytorch/pull/145000 Approved by: https://github.com/albanD	2025-06-30 16:58:06 +00:00
Han, Xu	39b71d11fc	[Inductor] add pedantic to limit inductor code follow standard. (#156914 ) ### Background: During my development work, I found Windows msvc don't support to compile zero size array, please reference: https://github.com/pytorch/pytorch/issues/153180 As discussed with MSFT engineer, we found zero size array don't align to c++ standard, though gcc/clang can support it. When we add `-pedantic` option to gcc, it should check and raise c++ standard strictly. Reference: https://github.com/pytorch/pytorch/issues/153180#issuecomment-2986676878 So this PR add `-pedantic` to torch inductor build option list to constraint codegen generate c++ standard well code. Additional, It also fixed a halide zero size array code. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156914 Approved by: https://github.com/jansel	2025-06-30 16:29:08 +00:00
Tom Ritchford	e3afbb0362	[inductor] Add typing to _inductor/ir.py (#149958 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/149958 Approved by: https://github.com/Skylion007	2025-06-30 15:56:35 +00:00
eqy	3b4b5f8d47	[SDPA] Fix `alloc_with_matching_layout` stride sorting (#157145 ) Otherwise dims with "zero" stride get moved before contiguous dims (stride 1). Need to move the fix from #149282 to here as #154340 moved the original definition from `MHA.cpp`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157145 Approved by: https://github.com/Skylion007	2025-06-30 15:43:29 +00:00
PyTorch MergeBot	da1f337bc4	Revert "Fixes for CPython int/float tests (#155978 )" This reverts commit fab53dfdf1d89cecd5e82b12cced9b6dd217e87c. Reverted https://github.com/pytorch/pytorch/pull/155978 on behalf of https://github.com/guilhermeleobas due to failing in trunk ([comment](https://github.com/pytorch/pytorch/pull/155978#issuecomment-3019457531))	2025-06-30 14:49:44 +00:00
Guilherme Leobas	fab53dfdf1	Fixes for CPython int/float tests (#155978 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155978 Approved by: https://github.com/zou3519	2025-06-30 14:15:47 +00:00
PyTorch UpdateBot	ffaed8c569	Update slow tests (#155448 ) This PR is auto-generated weekly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/weekly.yml). Update the list of slow tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155448 Approved by: https://github.com/pytorchbot	2025-06-30 12:08:52 +00:00
PyTorch UpdateBot	b1a54fab9b	[xla hash update] update the pinned xla hash (#156584 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned xla hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156584 Approved by: https://github.com/pytorchbot	2025-06-30 11:23:06 +00:00
LifengWang	ccb67f39b4	Enable the AMP precision with freezing for CPU nightly test (#152298 ) Hi, @desertfire. Since we recommend users to use AMP precision and run with `--freezing` for CPU x86 Inductor inference, we suggest adding the AMP freezing test to the CPU nightly tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152298 Approved by: https://github.com/desertfire, https://github.com/huydhn Co-authored-by: zengxian <xiangdong.zeng@intel.com>	2025-06-30 09:17:17 +00:00
morrison-turnansky	f79689bd3d	updated matplotlib version in docs requirements (#155931 ) Fixes #155199 The issue on main is due an outdated version of matplotlib. I have bumped the version so that it is compatible with Numpy 2.0 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155931 Approved by: https://github.com/malfet	2025-06-30 02:05:53 +00:00
Isalia20	a1282b1823	[MPS] Add boilerplate sparse code support (#157238 ) This PR makes minimal changes to support sparse tensors on MPS. In the followup PRs I'll start adding different operations slowly so we can fix the issue of https://github.com/pytorch/pytorch/issues/129842 which is highly requested(I assume because of whisper using sparse tensors) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157238 Approved by: https://github.com/malfet	2025-06-30 01:53:45 +00:00
Bin Bao	771be85704	[AOTI] Print out error msg when nvcc compiler fails (#157203 ) Summary: To debug https://github.com/pytorch/pytorch/issues/156930. Not able to reproduce the problem locally. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157203 Approved by: https://github.com/jansel Co-authored-by: Jason Ansel <jansel@meta.com>	2025-06-30 01:30:55 +00:00
Burak Turk	86ced14453	increment pending_callbacks_counter before initation the pt2 compile callbacks (#157185 ) Summary: Since we increment the counter after performing the callback, it leads to the assertion error when callback raises an error and increment never happens. Let's increment first to avoid it. Test Plan: tba Rollback Plan: Differential Revision: D77475650 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157185 Approved by: https://github.com/xmfan	2025-06-30 01:23:59 +00:00
Jason Ansel	12cb06e574	[inductor] Increase tolerance for test_comprehensive_nn_functional_linear_cuda_float16 (#156962 ) Fixes #156514 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156962 Approved by: https://github.com/jamesjwu	2025-06-30 00:54:20 +00:00
cyy	c27f83dd91	Remove old ASAN Docker images (#157197 ) The old ASAN jobs have been replaced. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157197 Approved by: https://github.com/Skylion007	2025-06-30 00:30:56 +00:00
Jake Stevens	11f7e2f145	[caffe][executorch] rename to avoid shadow in irange (#157107 ) Summary: D76832520 switched Executorch to use the caffe c10 headers. This copy contains a shadow, which is treated as an error for certain embedded compile flows. Simple rename to avoid. Test Plan: CI Rollback Plan: Differential Revision: D77446104 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157107 Approved by: https://github.com/Skylion007	2025-06-30 00:17:09 +00:00
dolpm	018e9826a2	[nativert] hook up memory planning to execution frame (#157053 ) Summary: pretty simple. if planner exists, which implies that planning is enabled, create a manager for each frame. the associated serial executor will use the withMemoryPlannner fn to ensure the deallocation is done after execution completes. Test Plan: CI Differential Revision: D73635809 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157053 Approved by: https://github.com/henryoier, https://github.com/georgiaphillips	2025-06-30 00:06:37 +00:00
Jason Ansel	41f6acef83	Update pr_time_benchmarks expected results (#157214 ) The job has been unstable Pull Request resolved: https://github.com/pytorch/pytorch/pull/157214 Approved by: https://github.com/laithsakka	2025-06-29 19:12:13 +00:00
PyTorch MergeBot	29f76ec0f3	Revert "[BE] use `pathlib.Path` instead of `os.path.*` in `setup.py` (#156742 )" This reverts commit 2380115f9738f97cf706affefd647d2cb6dfbb3f. Reverted https://github.com/pytorch/pytorch/pull/156742 on behalf of https://github.com/malfet due to Looks like it broke all ROCM tests, see `721d2580db/1` ([comment](https://github.com/pytorch/pytorch/pull/156742#issuecomment-3016937704))	2025-06-29 18:10:03 +00:00
Simon Fan	721d2580db	[dynamo][callbacks] temporarily disable TRITON_AUTOTUNING (#157186 ) Differential Revision: D77476551 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157186 Approved by: https://github.com/burak-turk	2025-06-29 17:20:55 +00:00
Nick Riasanovsky	aec569da23	[Triton] [Inductor[ Add tt.descriptor_store to get_tma_stores (#157212 ) Summary: Fixes a gap in the Triton update where the traverse would break because `get_tma_stores` didn't handle both TMA APIs. Test Plan: `buck test -m ovr_config//triton:beta 'fbcode//mode/dev-nosan' fbcode//ads_mkl/ops/tests:gdpa_dcpp_test -- --exact 'ads_mkl/ops/tests:gdpa_dcpp_test - test_gdpa_dcpp (ads_mkl.ops.tests.gdpa_dcpp_test.GdpaDCPPTest)'` Rollback Plan: Differential Revision: D77501582 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157212 Approved by: https://github.com/davidberard98	2025-06-29 16:44:52 +00:00
Jason Ansel	b147b6c0e3	Increase tolerance for test_corrcoef_cuda_int32 (#157206 ) Fixes #156988 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157206 Approved by: https://github.com/Skylion007	2025-06-29 16:30:54 +00:00
Patryk Ozga	e959dd017d	[TSAN][live speech translation] Fix A data race in caffe2 (#156378 ) Summary: noticed that context quantized_engine is accessed and written from multiple threads Test Plan: ➜ fbsource buck test --flagfile fbcode/mode/dev-tsan //xplat/assistant/integration_test/tests/supernova/speechtranslation:live_speech_translation_en_fr_tests -- --exact 'fbsource//xplat/assistant/integration_test/tests/supernova/speechtranslation:live_speech_translation_en_fr_tests - Translate/LiveSpeechTranslationTests.LiveSpeechTranslationEnFr/silence___fr_en' Rollback Plan: Differential Revision: D76921416 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156378 Approved by: https://github.com/jerryzh168, https://github.com/cyyever	2025-06-29 07:23:20 +00:00
bobrenjc93	9d677389cb	[async compile] make it more obvious that we support backwards (#157204 ) current failing with ``` (/home/bobren/local/a/pytorch-env) [13:02] devgpu009:/home/bobren/local/a/pytorch python test/inductor/test_compile_subprocess.py -k GPUTests.test_async /home/bobren/local/a/pytorch/torch/backends/cudnn/__init__.py:115: UserWarning: PyTorch was compiled without cuDNN/MIOpen support. To use cuDNN/MIOpen, rebuild PyTorch making sure the library is visible to the build system. warnings.warn( /home/bobren/local/a/pytorch/torch/_inductor/ops_handler.py:741: UserWarning: undefined OpHandler.__getstate__, please add missing op schema warnings.warn(f"undefined OpHandler.{name}, please add missing op schema") /home/bobren/local/a/pytorch/torch/_inductor/ops_handler.py:741: UserWarning: undefined OpHandler.__getstate__, please add missing op schema warnings.warn(f"undefined OpHandler.{name}, please add missing op schema") W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] Unable to pickle input graph or example inputs W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] Traceback (most recent call last): W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] File "/home/bobren/local/a/pytorch/torch/_inductor/compile_fx_ext.py", line 484, in serialize_compile W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] ).serialize() W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] File "/home/bobren/local/a/pytorch/torch/_inductor/compile_fx_ext.py", line 210, in serialize W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] return _WireProtocolPickledInput(GraphPickler.dumps(self)) W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] File "/home/bobren/local/a/pytorch/torch/fx/_graph_pickler.py", line 124, in dumps W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] pickler.dump(obj) W0628 13:02:30.666000 3610483 torch/_inductor/compile_fx_ext.py:491] [0/0] AttributeError: Can't pickle local object 'make_opaque_bitwise_fn.<locals>.BitwiseFn' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157204 Approved by: https://github.com/aorenste	2025-06-29 05:38:54 +00:00
Gabriel Ferns	347ace4c7a	Inductor logging + analysis of torch.profile (#149697 ) Prereqs: - https://github.com/pytorch/pytorch/pull/152708 Features: 1. Adds inductor's estimate of flops and bandwidth to the json trace events that perfetto uses. 1. Only use the tflops estimation from triton if we don't have the info from the datasheet because Triton's estimates are inaccurate. I have a backlog item to fix triton flops estimation upstream. New `DeviceInfo` class, and new function `get_device_tflops`. 1. New helpers `countable_fx` and `count_flops_fx` helps get the flops of an `fx.Node`. 1. Extends Triton `torch.profiler` logging to `DebugAutotuner`. 1. New script `profile_analysis.py`: `--augment_trace` adds perf estimates to any perfetto json trace, `--analyze` creates a summary table of these perf estimates, and `--diff` will compare two traces side by side: ```python Device(NVIDIA H100, 0): Kernel Name \| resnet Kernel Count \| resnet FLOPS \| resnet bw gbps \| resnet Dur (ms) \| resnet Achieved FLOPS % \| resnet Achieved Bandwidth % \| newresnet Kernel Count \| newresnet FLOPS \| newresnet bw gbps \| newresnet Dur (ms) \| newresnet Achieved FLOPS % \| newresnet Achieved Bandwidth % --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- triton_poi_fused__native_batch_norm_legi \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 triton_red_fused__native_batch_norm_legi \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 triton_poi_fused__native_batch_norm_legi \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 void cutlass::Kernel2<cutlass_80_tensoro \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 triton_red_fused__native_batch_norm_legi \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 void cutlass::Kernel2<cutlass_80_tensoro \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 triton_poi_fused__native_batch_norm_legi \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 triton_red_fused__native_batch_norm_legi \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 triton_per_fused__native_batch_norm_legi \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 triton_poi_fused__native_batch_norm_legi \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 triton_per_fused__native_batch_norm_legi \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 triton_poi_fused__native_batch_norm_legi \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 void cublasLt::splitKreduce_kernel<32, 1 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 triton_per_fused__native_batch_norm_legi \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 triton_poi_fused__native_batch_norm_legi \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 void cask_plugin_cudnn::xmma_cudnn::init \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 triton_per_fused__native_batch_norm_legi \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 triton_per_fused__native_batch_norm_legi \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 void cutlass::Kernel2<cutlass_80_tensoro \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 triton_poi_fused__native_batch_norm_legi \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 triton_poi_fused__native_batch_norm_legi \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 triton_poi_fused__native_batch_norm_legi \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 triton_per_fused__native_batch_norm_legi \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 triton_poi_fused__native_batch_norm_legi \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 triton_per_fused__native_batch_norm_legi \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 triton_poi_fused__native_batch_norm_legi \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 triton_per_fused__native_batch_norm_legi \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 std::enable_if<!(false), void>::type int \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 triton_poi_fused_add_copy__38 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 triton_poi_fused_convolution_0 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 triton_poi_fused_convolution_1 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 void convolve_common_engine_float_NHWC<f \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 triton_per_fused__native_batch_norm_legi \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 triton_per_fused__native_batch_norm_legi \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 triton_poi_fused__native_batch_norm_legi \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 void at::native::(anonymous namespace):: \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 void at::native::vectorized_elementwise_ \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/149697 Approved by: https://github.com/eellison, https://github.com/shunting314	2025-06-29 05:00:47 +00:00
Xuehai Pan	f8293116f5	[BE][13/16] fix typos in torch/ (torch/ao/) (#156603 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156603 Approved by: https://github.com/msaroufim	2025-06-29 04:34:04 +00:00
Abdourrahmane Kabbaj	1913c915e0	Fixes issue #156414 : Fixes bug in implementation of _combine_histograms. (#156457 ) Fixes #156414 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156457 Approved by: https://github.com/jerryzh168	2025-06-29 04:30:28 +00:00
Saiteja Samudrala	2796f31b5e	[DCP] OSS Zero Overhead Checkpointing Implementation (#156207 ) Summary: This diff updates DCP driver code/APIs to support Zero Overhead Checkpointing Test Plan: Test with TorchTitan on this PR: https://github.com/pytorch/torchtitan/pull/1287 Differential Revision: D72391401 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156207 Approved by: https://github.com/teja-rao	2025-06-29 03:19:48 +00:00
Nichols A. Romero	bccb8473fe	[ROCm] Allow use of rocSOLVER for Cholesky inversion. (#157154 ) Fixes https://github.com/pytorch/pytorch/issues/155046 This change allows Cholesky inversion to use rocSOLVER. This is now also the default on ROCm for Cholesky inversion which aligns with the behavior on NVIDIA (which defaults to cuSOLVER for this linear algebra operation). This fix also gets around a memory access fault encountered in MAGMA for large matrices. MAGMA can still be forced on ROCm by doing: ``` torch.backends.cuda.preferred_linalg_library(backend='magma') ``` Ran all Cholesky UT on ROCm and there were no regressions. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157154 Approved by: https://github.com/jeffdaily	2025-06-29 01:53:02 +00:00
Laith Sakka	6cc490d40b	simplify max(1,x) to x when x known >=1 (#157189 ) Creating contiguous strides creates an expression max(1, x). Often we know that x >= 1, in which case we should simplify max(1, x) to x. This appeared in two situations: 1) An internal user complained about statically_known_true(x == max(1, x)) failing (internal link: https://fb.workplace.com/groups/1028545332188949/permalink/1232958568414290). This https://github.com/pytorch/pytorch/pull/155938 won't be needed with this. 3) Not simplifying the above could result in wrong ConstraintViolationErrors. Because we assume non-trival single arg guards shall evaporate see the logic in the function issue_guard in symbolic_shapes.py with this change we longer throw ConstraintViolationErrors with the program bellow this is blocking landing this [PR](https://github.com/pytorch/pytorch/pull/155590) from landing internally. Due to internal export tests throwing ConstraintViolationErrors. like ``` Constraints violated (width)! - Not all values of width = L['x'].size()[3] in the specified range 224 <= width <= 455 satisfy the generated guard max(1, 1 + (((-1) + L['x'].size()[3]) // 2)) == (1 + (((-1) + L['x'].size()[3]) // 2)). ```` ``` x = torch.rand(10) torch._dynamo.mark_dynamic(x, 0, max=20, min=5) @torch.compile(fullgraph=True, dynamic=True) def func(x): if max(1, (-1 + x.size()[0]//2)) == (-1+x.size()[0]//2): return x400 else: return (x10)*100 func(x) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157189 Approved by: https://github.com/pianpwk	2025-06-29 01:16:30 +00:00
Yidi Wu	836bb1941b	[hop] support torch.func.functional_call in hop subgraph (#155886 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155886 Approved by: https://github.com/zou3519	2025-06-28 23:47:46 +00:00
Xuehai Pan	2380115f97	[BE] use `pathlib.Path` instead of `os.path.*` in `setup.py` (#156742 ) Resolves: - https://github.com/pytorch/pytorch/pull/155998#discussion_r2164376634 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156742 Approved by: https://github.com/malfet	2025-06-28 23:31:15 +00:00
Xuehai Pan	90b973a2e2	[BE] parse CMake version from `cmake -E capabilities` instead of `cmake --version` (#157073 ) `cmake -E capabilities` produces a JSON format that is more machine-friendly. ```console $ cmake --version cmake version 4.0.3 CMake suite maintained and supported by Kitware (kitware.com/cmake). $ cmake -E capabilities \| jq '.version.string' "4.0.3" $ cmake -E capabilities \| jq { "debugger": true, "fileApi": { "requests": [ { "kind": "codemodel", "version": [ { "major": 2, "minor": 8 } ] }, { "kind": "configureLog", "version": [ { "major": 1, "minor": 0 } ] }, { "kind": "cache", "version": [ { "major": 2, "minor": 0 } ] }, { "kind": "cmakeFiles", "version": [ { "major": 1, "minor": 1 } ] }, { "kind": "toolchains", "version": [ { "major": 1, "minor": 0 } ] } ] }, "generators": [ { "extraGenerators": [], "name": "Watcom WMake", "platformSupport": false, "toolsetSupport": false }, { "extraGenerators": [ "Kate" ], "name": "Ninja Multi-Config", "platformSupport": false, "toolsetSupport": false }, { "extraGenerators": [ "CodeBlocks", "CodeLite", "Eclipse CDT4", "Kate", "Sublime Text 2" ], "name": "Ninja", "platformSupport": false, "toolsetSupport": false }, { "extraGenerators": [], "name": "Xcode", "platformSupport": false, "toolsetSupport": true }, { "extraGenerators": [ "CodeBlocks", "CodeLite", "Eclipse CDT4", "Kate", "Sublime Text 2" ], "name": "Unix Makefiles", "platformSupport": false, "toolsetSupport": false } ], "serverMode": false, "tls": true, "version": { "isDirty": false, "major": 4, "minor": 0, "patch": 3, "string": "4.0.3", "suffix": "" } } ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157073 Approved by: https://github.com/Skylion007	2025-06-28 23:20:10 +00:00
AaronWang04	772d590415	[CUTLASS] [CUDA] SM100 GroupMM (#156203 ) Closes https://github.com/pytorch/pytorch/issues/156202 PR adds blackwell support for GroupMM Most of the code that is used for SM90 can be reused, kernel schedule has to be changed in accordance with https://docs.nvidia.com/cutlass/media/docs/cpp/blackwell_functionality.html Did some preliminary benchmarking of H200 vs B200 Script ```py import torch print(torch.__file__) device = torch.device("cuda") dtype = torch.bfloat16 shapes = [ (16, 128000, 7168, 7168), (128, 1, 2048, 7168) ] for batch, M, N, K in shapes: a = torch.randn(batch, M, K, device=device, dtype=dtype) b = torch.randn(batch, N, K, device=device, dtype=dtype) start_event = torch.cuda.Event(enable_timing=True) end_event = torch.cuda.Event(enable_timing=True) for i in range(5): c = torch._grouped_mm(a, b) num_iter = 50 start_event.record() for i in range(num_iter): c = torch._grouped_mm(a, b) end_event.record() torch.cuda.synchronize() elapsed_time_ms = start_event.elapsed_time(end_event) avg_time_ms = elapsed_time_ms / num_iter print(f"batch: {batch}\tM: {M}\tN: {N}\tK: {K}") print(f"Time per Iteration:\t {avg_time_ms:.4f} ms") ``` On H200 ``` batch: 16 M: 128000 N: 7168 K: 7168 Time per Iteration: 298.6668 ms batch: 128 M: 1 N: 2048 K: 7168 Time per Iteration: 4.1462 ms ``` B200 ``` batch: 16 M: 128000 N: 7168 K: 7168 Time per Iteration: 190.7458 ms batch: 128 M: 1 N: 2048 K: 7168 Time per Iteration: 3.0680 ms ``` nsys nvprof ``` root@16930b42ffc6:/workspace/pytorch# nsys nvprof python gemm_test.py WARNING: python and any of its children processes will be profiled. Collecting data... batch: 16 M: 128000 N: 7168 K: 7168 Time per Iteration: 192.6420 ms batch: 128 M: 1 N: 2048 K: 7168 Time per Iteration: 1.2255 ms Generating '/tmp/nsys-report-6a53.qdstrm' [1/7] [========================100%] report1.nsys-rep [2/7] [========================100%] report1.sqlite [3/7] Executing 'nvtx_sum' stats report SKIPPED: /workspace/pytorch/report1.sqlite does not contain NV Tools Extension (NVTX) data. [4/7] Executing 'cuda_api_sum' stats report Time (%) Total Time (ns) Num Calls Avg (ns) Med (ns) Min (ns) Max (ns) StdDev (ns) Name -------- --------------- --------- ------------ ------------ -------- ----------- ------------ --------------------------------- 98.9 10586895744 2 5293447872.0 5293447872.0 73786464 10513109280 7381715954.2 cudaDeviceSynchronize 1.0 104084608 5 20816921.6 33552480.0 100800 34786208 18048125.3 cudaMalloc 0.1 5694304 4 1423576.0 1416656.0 1258560 1602432 181668.1 cudaGetDeviceProperties_v2_v12000 0.1 5430496 130 41773.0 4560.0 2496 3854368 345761.8 cudaLaunchKernel 0.0 587584 110 5341.7 4992.0 4224 16992 1482.0 cudaLaunchKernelExC_v11060 0.0 119200 660 180.6 128.0 96 4128 206.7 cudaGetDriverEntryPoint_v11030 0.0 68352 660 103.6 64.0 32 4928 224.6 cuTensorMapEncodeTiled 0.0 34976 49 713.8 224.0 160 6720 1343.4 cudaStreamIsCapturing_v10000 0.0 32992 4 8248.0 7456.0 4128 13952 4804.4 cudaEventRecord 0.0 16928 4 4232.0 3600.0 1728 8000 2764.7 cudaEventQuery 0.0 16288 4 4072.0 3568.0 1952 7200 2396.1 cudaEventCreateWithFlags 0.0 13632 4 3408.0 2672.0 544 7744 3408.7 cudaEventDestroy 0.0 1056 1 1056.0 1056.0 1056 1056 0.0 cuModuleGetLoadingMode [5/7] Executing 'cuda_gpu_kern_sum' stats report Time (%) Total Time (ns) Instances Avg (ns) Med (ns) Min (ns) Max (ns) StdDev (ns) Name -------- --------------- --------- ----------- ----------- --------- --------- ----------- ---------------------------------------------------------------------------------------------------- 99.0 10549232845 55 191804233.5 192944479.0 165746368 203645313 5353204.3 void cutlass::device_kernel<at::cuda::detail::enable_3x_kernel_for_sm10<cutlass::gemm::kernel::Gemm… 0.6 67327135 55 1224129.7 1330656.0 924320 1364928 182180.4 void cutlass::device_kernel<at::cuda::detail::enable_3x_kernel_for_sm10<cutlass::gemm::kernel::Gemm… 0.3 34854783 20 1742739.1 1597856.0 10080 3899616 818421.2 void at::native::<unnamed>::distribution_elementwise_grid_stride_kernel<float, (int)4, void at::nat… 0.0 354880 110 3226.2 3296.0 1920 4160 554.4 void at::cuda::detail::prepare_grouped_gemm_data<cutlass::bfloat16_t, cutlass::bfloat16_t, cutlass:… ``` The kernel names are too long to be shown via nvprof, I pasted this from nsight systems ``` small kernel 1SM 100.0% 1.286 ms 1 1.286 ms 1.286 ms 1.286 ms 1.286 ms 0 ns void cutlass::device_kernel<at::cuda::detail::enable_3x_kernel_for_sm10<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<cute::tuple<int, int, int>>, cutlass::gemm::collective::CollectiveMma<cutlass::gemm::MainloopSm100ArrayTmaUmmaWarpSpecialized<(int)3, (int)8, (int)2, cute::tuple<cute::C<(int)2>, cute::C<(int)1>, cute::C<(int)1>>>, cute::tuple<cute::C<(int)128>, cute::C<(int)256>, cute::C<(int)64>>, cutlass::bfloat16_t, cute::tuple<long, cute::C<(int)1>, cute::C<(int)0>> , cutlass::bfloat16_t, cute::tuple<cute::C<(int)1>, long, cute::C<(int)0>> , cute::TiledMMA<cute::MMA_Atom<cute::SM100_MMA_F16BF16_SS<cutlass::bfloat16_t, cutlass::bfloat16_t, float, (int)128, (int)256, (cute::UMMA::Major)0, (cute::UMMA::Major)1, (cute::UMMA::ScaleIn)0, (cute::UMMA::ScaleIn)0>>, cute::Layout<cute::tuple<cute::C<(int)1>, cute::C<(int)1>, cute::C<(int)1>>, cute::tuple<cute::C<(int)0>, cute::C<(int)0>, cute::C<(int)0>>>, cute::tuple<cute::Underscore, cute::Underscore, cute::Underscore>>, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<(int)3, (int)4, (int)3>, cute::smem_ptr_flag_bits<(int)16>, cute::Layout<cute::tuple<cute::C<(int)8>, cute::C<(int)64>>, cute::tuple<cute::C<(int)64>, cute::C<(int)1>>>>, void, cute::identity, cute::SM90_TMA_LOAD_MULTICAST, cute::ComposedLayout<cute::Swizzle<(int)3, (int)4, (int)3>, cute::smem_ptr_flag_bits<(int)16>, cute::Layout<cute::tuple<cute::C<(int)64>, cute::C<(int)8>>, cute::tuple<cute::C<(int)1>, cute::C<(int)64>>>>, void, cute::identity>, cutlass::epilogue::collective::CollectiveEpilogue<cutlass::epilogue::Sm100PtrArrayTmaWarpSpecialized<(int)4, (int)2, (int)64, (bool)1, (bool)0>, cute::tuple<cute::C<(int)128>, cute::C<(int)256>, cute::C<(int)64>>, cute::tuple<cute::Layout<cute::C<(int)128>, cute::C<(int)1>>, cute::Layout<cute::C<(int)64>, cute::C<(int)1>>>, cutlass::bfloat16_t, cute::tuple<long, cute::C<(int)1>, cute::C<(int)0>> , cutlass::bfloat16_t, cute::tuple<long, cute::C<(int)1>, cute::C<(int)0>> , cutlass::epilogue::fusion::FusionCallbacks<cutlass::epilogue::Sm100PtrArrayTmaWarpSpecialized<(int)4, (int)2, (int)64, (bool)1, (bool)0>, cutlass::epilogue::fusion::LinearCombination<cutlass::bfloat16_t, float, cutlass::bfloat16_t, float, (cutlass::FloatRoundStyle)2>, cute::tuple<cute::C<(int)128>, cute::C<(int)256>, cute::C<(int)64>>, cute::tuple<cute::Layout<cute::C<(int)128>, cute::C<(int)1>>, cute::Layout<cute::C<(int)64>, cute::C<(int)1>>>, >, cute::SM100::TMEM::LOAD::SM100_TMEM_LOAD_32dp32b64x, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<(int)3, (int)4, (int)3>, cute::smem_ptr_flag_bits<(int)16>, cute::Layout<cute::tuple<cute::C<(int)8>, cute::C<(int)64>>, cute::tuple<cute::C<(int)64>, cute::C<(int)1>>>>, cute::AutoVectorizingCopyWithAssumedAlignment<(int)128>, cute::SM90_TMA_STORE, cute::ComposedLayout<cute::Swizzle<(int)3, (int)4, (int)3>, cute::smem_ptr_flag_bits<(int)16>, cute::Layout<cute::tuple<cute::C<(int)8>, cute::C<(int)64>>, cute::tuple<cute::C<(int)64>, cute::C<(int)1>>>>, cute::AutoVectorizingCopyWithAssumedAlignment<(int)128>, cute::AutoVectorizingCopyWithAssumedAlignment<(int)128>>, void, void>>>(T1::Params) large kernel 2SM 100.0% 194.178 ms 1 194.178 ms 194.178 ms 194.178 ms 194.178 ms 0 ns void cutlass::device_kernel<at::cuda::detail::enable_3x_kernel_for_sm10<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<cute::tuple<int, int, int>>, cutlass::gemm::collective::CollectiveMma<cutlass::gemm::MainloopSm100ArrayTmaUmmaWarpSpecialized<(int)5, (int)8, (int)2, cute::tuple<cute::C<(int)2>, cute::C<(int)1>, cute::C<(int)1>>>, cute::tuple<cute::C<(int)256>, cute::C<(int)256>, cute::C<(int)64>>, cutlass::bfloat16_t, cute::tuple<long, cute::C<(int)1>, cute::C<(int)0>> , cutlass::bfloat16_t, cute::tuple<cute::C<(int)1>, long, cute::C<(int)0>> , cute::TiledMMA<cute::MMA_Atom<cute::SM100_MMA_F16BF16_2x1SM_SS<cutlass::bfloat16_t, cutlass::bfloat16_t, float, (int)256, (int)256, (cute::UMMA::Major)0, (cute::UMMA::Major)1, (cute::UMMA::ScaleIn)0, (cute::UMMA::ScaleIn)0>>, cute::Layout<cute::tuple<cute::C<(int)1>, cute::C<(int)1>, cute::C<(int)1>>, cute::tuple<cute::C<(int)0>, cute::C<(int)0>, cute::C<(int)0>>>, cute::tuple<cute::Underscore, cute::Underscore, cute::Underscore>>, cute::SM100_TMA_2SM_LOAD, cute::ComposedLayout<cute::Swizzle<(int)3, (int)4, (int)3>, cute::smem_ptr_flag_bits<(int)16>, cute::Layout<cute::tuple<cute::C<(int)8>, cute::C<(int)64>>, cute::tuple<cute::C<(int)64>, cute::C<(int)1>>>>, void, cute::identity, cute::SM100_TMA_2SM_LOAD, cute::ComposedLayout<cute::Swizzle<(int)3, (int)4, (int)3>, cute::smem_ptr_flag_bits<(int)16>, cute::Layout<cute::tuple<cute::C<(int)64>, cute::C<(int)8>>, cute::tuple<cute::C<(int)1>, cute::C<(int)64>>>>, void, cute::identity>, cutlass::epilogue::collective::CollectiveEpilogue<cutlass::epilogue::Sm100PtrArrayTmaWarpSpecialized<(int)4, (int)2, (int)64, (bool)1, (bool)0>, cute::tuple<cute::C<(int)128>, cute::C<(int)256>, cute::C<(int)64>>, cute::tuple<cute::Layout<cute::C<(int)128>, cute::C<(int)1>>, cute::Layout<cute::C<(int)64>, cute::C<(int)1>>>, cutlass::bfloat16_t, cute::tuple<long, cute::C<(int)1>, cute::C<(int)0>> , cutlass::bfloat16_t, cute::tuple<long, cute::C<(int)1>, cute::C<(int)0>> , cutlass::epilogue::fusion::FusionCallbacks<cutlass::epilogue::Sm100PtrArrayTmaWarpSpecialized<(int)4, (int)2, (int)64, (bool)1, (bool)0>, cutlass::epilogue::fusion::LinearCombination<cutlass::bfloat16_t, float, cutlass::bfloat16_t, float, (cutlass::FloatRoundStyle)2>, cute::tuple<cute::C<(int)128>, cute::C<(int)256>, cute::C<(int)64>>, cute::tuple<cute::Layout<cute::C<(int)128>, cute::C<(int)1>>, cute::Layout<cute::C<(int)64>, cute::C<(int)1>>>, >, cute::SM100::TMEM::LOAD::SM100_TMEM_LOAD_32dp32b64x, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<(int)3, (int)4, (int)3>, cute::smem_ptr_flag_bits<(int)16>, cute::Layout<cute::tuple<cute::C<(int)8>, cute::C<(int)64>>, cute::tuple<cute::C<(int)64>, cute::C<(int)1>>>>, cute::AutoVectorizingCopyWithAssumedAlignment<(int)128>, cute::SM90_TMA_STORE, cute::ComposedLayout<cute::Swizzle<(int)3, (int)4, (int)3>, cute::smem_ptr_flag_bits<(int)16>, cute::Layout<cute::tuple<cute::C<(int)8>, cute::C<(int)64>>, cute::tuple<cute::C<(int)64>, cute::C<(int)1>>>>, cute::AutoVectorizingCopyWithAssumedAlignment<(int)128>, cute::AutoVectorizingCopyWithAssumedAlignment<(int)128>>, void, void>>>(T1::Params) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156203 Approved by: https://github.com/syed-ahmed, https://github.com/drisspg	2025-06-28 23:02:00 +00:00
Jeff Daily	996206e66f	cublaslt/hipblaslt persistent workspace (#156495 ) Similar to cublas/hipblas, LT now allocates one workspace per handle+stream combo. - fixes hipblaslt issue where memory use increased during graph capture - preserves CUDA env var TORCH_CUBLASLT_UNIFIED_WORKSPACE - moves LT workspace and size from CUDABlas.cpp into CublasHandlePool.cpp, new APIs - size_t getCUDABlasLtWorkspaceSize() - void* getCUDABlasLtWorkspace() Fixes https://github.com/ROCm/pytorch/issues/2286. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156495 Approved by: https://github.com/eqy	2025-06-28 22:38:43 +00:00
Edenzzzz	0629dfb860	Fix FSDP offload pin_memory bug (#157147 ) Fixes #157146 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157147 Approved by: https://github.com/weifengpy	2025-06-28 21:09:11 +00:00
blorange-amd	67f8270516	[ROCm] test_hip_device_count safely runs on 1 GPU systems (#156398 ) Fixes test_cuda.py::TestCuda::test_hip_device_count on single gpu scenario Pull Request resolved: https://github.com/pytorch/pytorch/pull/156398 Approved by: https://github.com/jeffdaily	2025-06-28 20:17:26 +00:00
Yidi Wu	aeffb68d34	[schema_upgrader] add C++ upgrader for json based upgrading (#156761 ) Differential Revision: [D77459912](https://our.internmc.facebook.com/intern/diff/D77459912) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156761 Approved by: https://github.com/angelayi	2025-06-28 18:15:06 +00:00
Yidi Wu	064a7db7fc	[invoke_subgraph] turn on supports_input_mutation by default (#157177 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157177 Approved by: https://github.com/anijain2305	2025-06-28 18:14:47 +00:00
PyTorch MergeBot	2eb744c08d	Revert "[BE] parse CMake version from `cmake -E capabilities` instead of `cmake --version` (#157073 )" This reverts commit 0c58bdd8fb5f269aef100af8e2c43cfcf5f1f9dd. Reverted https://github.com/pytorch/pytorch/pull/157073 on behalf of https://github.com/XuehaiPan due to break libtorch build on Windows ([comment](https://github.com/pytorch/pytorch/pull/157073#issuecomment-3015273679))	2025-06-28 13:40:19 +00:00
Xuehai Pan	0c58bdd8fb	[BE] parse CMake version from `cmake -E capabilities` instead of `cmake --version` (#157073 ) `cmake -E capabilities` produces a JSON format that is more machine-friendly. ```console $ cmake --version cmake version 4.0.3 CMake suite maintained and supported by Kitware (kitware.com/cmake). $ cmake -E capabilities \| jq '.version.string' "4.0.3" $ cmake -E capabilities \| jq { "debugger": true, "fileApi": { "requests": [ { "kind": "codemodel", "version": [ { "major": 2, "minor": 8 } ] }, { "kind": "configureLog", "version": [ { "major": 1, "minor": 0 } ] }, { "kind": "cache", "version": [ { "major": 2, "minor": 0 } ] }, { "kind": "cmakeFiles", "version": [ { "major": 1, "minor": 1 } ] }, { "kind": "toolchains", "version": [ { "major": 1, "minor": 0 } ] } ] }, "generators": [ { "extraGenerators": [], "name": "Watcom WMake", "platformSupport": false, "toolsetSupport": false }, { "extraGenerators": [ "Kate" ], "name": "Ninja Multi-Config", "platformSupport": false, "toolsetSupport": false }, { "extraGenerators": [ "CodeBlocks", "CodeLite", "Eclipse CDT4", "Kate", "Sublime Text 2" ], "name": "Ninja", "platformSupport": false, "toolsetSupport": false }, { "extraGenerators": [], "name": "Xcode", "platformSupport": false, "toolsetSupport": true }, { "extraGenerators": [ "CodeBlocks", "CodeLite", "Eclipse CDT4", "Kate", "Sublime Text 2" ], "name": "Unix Makefiles", "platformSupport": false, "toolsetSupport": false } ], "serverMode": false, "tls": true, "version": { "isDirty": false, "major": 4, "minor": 0, "patch": 3, "string": "4.0.3", "suffix": "" } } ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157073 Approved by: https://github.com/Skylion007	2025-06-28 13:35:30 +00:00
Erik Schultheis	cdb144fcf0	Display a warning when overwriting `CMAKE_CUDA_ARCHITECTURES` (#156123 ) Really, pytorch shoudn't be messing with basic _global_ cmake configuration like this, but without a careful analysis what all depends on this behaviour, I'm not confident to propose a change. But at least notifying the user that something wonky is going on seems like a good idea. @drisspg Pull Request resolved: https://github.com/pytorch/pytorch/pull/156123 Approved by: https://github.com/drisspg, https://github.com/msaroufim Co-authored-by: Mark Saroufim <marksaroufim@meta.com>	2025-06-28 11:22:09 +00:00
fduwjj	8147c4a904	[symm_mem] Create a dedicated ci flow for symmetric memory and only use 4 GPUs (#157181 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157181 Approved by: https://github.com/kwen2501, https://github.com/huydhn	2025-06-28 08:33:50 +00:00
Sheng Qin	88c6199db0	[nativert] Move KernelFactory to PyTorch core (#156913 ) Summary: Kernel factory handles the kernel nodes initializations and different type of kernels executions. Test Plan: CI Rollback Plan: Differential Revision: D77346836 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156913 Approved by: https://github.com/zhxchen17	2025-06-28 06:34:24 +00:00
Aidyn-A	51eb8e8f84	[ATen][CUDA][CUB] Implement changes to CCCL (CUB/Thrust/LibCUDACXX) usage in ATen (#153373 ) A major release of CCCL 3.0.0 will introduce some bc-breaking changes. Namely iterators like TransformInputIterator and ConstantInputIterator were moved from CUB to Thrust, some operators like Max and Sum were moved to LibCUDACXX. For the more info on changes please visit: https://nvidia.github.io/cccl/cccl/3.0_migration_guide.html This is a follow up to PR #147493. A description from the original PR: > Several cub iterators have been deprecated and removed in the latest CCCL (cub) development https://github.com/NVIDIA/cccl/pull/3831. This PR replaced the usage of those cub iterators with thrust iterators. > > Some cub thread operators were also deprecated and removed in https://github.com/NVIDIA/cccl/pull/3918. This PR replaced those operators with libcudacxx ops. > > This might also affect ROCM usability a bit. > > This patch is tested to work with CCCL commit at `82befb0894` > > Tracking of CCCL/CUB deprecations in the most recent development https://github.com/NVIDIA/cccl/issues/101 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153373 Approved by: https://github.com/cyyever, https://github.com/atalman	2025-06-28 05:44:52 +00:00
Lukas Geiger	a92b24cd83	Prevent cudaStreamSync when indexing GPU tensors with boolean CPU mask (#156384 ) `index_put` with a boolean mask (`target[mask] = src`) causes a `cudaStreamSynchronize`. When both `mask` and `target` tensors are on GPU this is expected. However, the sync can be prevented if the `mask` is a CPU tensor. Internally a new index tensor is created with `mask.nonzero()` so we can use a non-blocking copy to transfer it to the GPU since it cannot be accidentally mutated by the user between its creation and the device copy. @ngimel Let me know if I'm missing something. I think this is useful since users can't prevent a sync simply by making sure all tensors are on the same device as with other ops. Instead one would need to do something like this which is much less readable ```python indices = mask.nonzero().squeeze(1).to("cuda", non_blocking=True) target[indices] = src ``` Fixes #12461 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156384 Approved by: https://github.com/ngimel	2025-06-28 05:41:16 +00:00
Justin Chu	5692cbb818	[ONNX] Delete symbolic caffe2 (#157102 ) Caffe2 is removed from pytorch. This is a clean up. Pull Request resolved: https://github.com/pytorch/pytorch/pull/157102 Approved by: https://github.com/titaiwangms, https://github.com/cyyever	2025-06-28 05:22:02 +00:00
cyy	30d2648a4a	Install nvperf_host together with cupti (#156668 ) Because cupti depends on nvperf_host, as discussed in https://github.com/pytorch/pytorch/pull/154595 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156668 Approved by: https://github.com/Skylion007	2025-06-28 04:26:36 +00:00
zpcore	adf6dd1e44	Fix `aten::index_put` args Dtensor type mismatch and add a propagation strategy (#156240 ) We notice model code contains indexing syntax like [nanogpt model code](`f144fe9095/torchbenchmark/models/nanogpt/model.py (L240)`), which causes training fail in the backward pass when using DTensor. In the code, `x = x[:, [-1], :]` calls the index op and in the backward pass, it will trigger `aten.index_put.default` with the second argument to be of type `torch::List<std::optional<Tensor>>`, e.g., `[None, tensor([-1], device='cuda:0')]`. We are unable to unwarp the op info into Dtensor based on the current logic [here](`2625c70aec/torch/distributed/tensor/_dispatch.py (L339-L358)`). We need to set runtime_schema_info for the op and enable needs_pytree to support the conversion of tensor list arg. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156240 Approved by: https://github.com/wanchaol	2025-06-28 04:09:41 +00:00
PyTorch MergeBot	f810480dbe	Revert "[schema_upgrader] add C++ upgrader for json based upgrading (#156761 )" This reverts commit 61712e6f2ba58cce354a742d918934ec7293ee43. Reverted https://github.com/pytorch/pytorch/pull/156761 on behalf of https://github.com/ydwu4 due to break linter test, which doesn't show up in the pr ([comment](https://github.com/pytorch/pytorch/pull/156761#issuecomment-3014918800))	2025-06-28 03:58:25 +00:00
Eli Uriegas	0e47312ae5	ci: Add ability to test images for build-triton-wheel (#156894 ) This wasn't available prior making it difficult to test if manywheel image changes would affect triton wheel builds. Signed-off-by: Eli Uriegas <eliuriegas@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156894 Approved by: https://github.com/atalman, https://github.com/clee2000, https://github.com/malfet ghstack dependencies: #156893	2025-06-28 03:41:18 +00:00
Teja	ef6dfa06a9	Create a base Checkpointer and SyncCheckpointer and add dist barrier impl and (#156926 ) In preparation to adding async checkpointing, this diff adds 1. Change Checkpointer to an Abstract base class and adds a sync checkpointer implementation. 2. torch.distributed.barrier() as one of the barrier choices. Differential Revision: [D77341314](https://our.internmc.facebook.com/intern/diff/D77341314/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156926 Approved by: https://github.com/pradeepfn	2025-06-28 02:48:29 +00:00
David Berard	e8217ad8be	[inductor][static launcher] Skip correctness test for test_floats (#157023 ) https://github.com/triton-lang/triton/issues/6176 causes kernels that take fp64 scalar inputs to generate wrong results. Until we get around to fixing this, just skip the accuracy check (it'll fail on Triton's launcher anyway). Differential Revision: [D77407307](https://our.internmc.facebook.com/intern/diff/D77407307) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157023 Approved by: https://github.com/jamesjwu	2025-06-28 02:19:10 +00:00
fduwjj	e3320965b4	[sym_mem] Further Fix NCCL symm mem unit test (#157156 ) We still see CI failures because of error "RuntimeError: CUDA driver error: invalid device ordinal". So upon discussion, we might also need a GPU number skip macro for the test itself: Fixes #156569 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157156 Approved by: https://github.com/kwen2501, https://github.com/fegin	2025-06-28 02:17:13 +00:00
Nikita Shulga	a1e4f1f98a	[MPS] Reimplement `tri[ul]` as Metal shaders (#157179 ) And add in-place flavor, as it is currently broken for non-contig tensors Pull Request resolved: https://github.com/pytorch/pytorch/pull/157179 Approved by: https://github.com/dcci	2025-06-28 01:33:18 +00:00
Orvid King	c14110056f	[caffe2] Allow the elimination of implicit calls to strlen when using the RECORD_FUNCTION macros (#153567 ) Summary: With the way these were written, any string literals that were being passed in, like `__func__`, were only ever passed down as a `const char*`, so this switches it over to take a `std::string_view` at the deepest part. This also has the side effect of allowing `std::string_view` to be passed to the `RECORD_FUNCTION` macros as well. Test Plan: contbuilds Rollback Plan: Differential Revision: D74681042 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153567 Approved by: https://github.com/Skylion007, https://github.com/swolchok	2025-06-28 01:11:00 +00:00
PyTorch MergeBot	1e4c5b666a	Revert "[dynamo] fix _torchdynamo_orig_callable naming issues (#156901 )" This reverts commit eb9efb37c8f315f1d30e86d5797490c6a8666889. Reverted https://github.com/pytorch/pytorch/pull/156901 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it seems to break some internal tests D77411594 ([comment](https://github.com/pytorch/pytorch/pull/156901#issuecomment-3014734151))	2025-06-28 00:37:01 +00:00
Yidi Wu	61712e6f2b	[schema_upgrader] add C++ upgrader for json based upgrading (#156761 ) Differential Revision: [D77459912](https://our.internmc.facebook.com/intern/diff/D77459912) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156761 Approved by: https://github.com/angelayi	2025-06-27 23:50:19 +00:00
morrison-turnansky	2815ade9a8	updated adafactor doc #154862 (#155248 ) updated adafactor doc to reflect difference in implementation vs original paper Fixes #154862 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155248 Approved by: https://github.com/janeyx99 Co-authored-by: Jane (Yuan) Xu <31798555+janeyx99@users.noreply.github.com>	2025-06-27 23:23:19 +00:00
nautsimon	feea575082	[MTIA ATen Backend] Add dispatch keys for add.out (#156952 ) Migrate add.out Differential Revision: [D77352482](https://our.internmc.facebook.com/intern/diff/D77352482/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156952 Approved by: https://github.com/malfet, https://github.com/huydhn ghstack dependencies: #156944, #156945, #156946, #156947, #156948, #156949, #156950, #156951	2025-06-27 22:49:00 +00:00
nautsimon	253cbadade	[MTIA ATen Backend] Add dispatch keys for rsub.Tensor / rsub.Scalar / sub.out (#156951 ) Migrate rsub.Tensor / rsub.Scalar / sub.out Differential Revision: [D77015033](https://our.internmc.facebook.com/intern/diff/D77015033/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156951 Approved by: https://github.com/malfet ghstack dependencies: #156944, #156945, #156946, #156947, #156948, #156949, #156950	2025-06-27 22:49:00 +00:00
nautsimon	b6b2871555	[MTIA ATen Backend] Add dispatch keys for fmod / abs.out / logical_not.out (#156950 ) Migrate fmod / abs.out / logical_not.out Differential Revision: [D77220217](https://our.internmc.facebook.com/intern/diff/D77220217/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156950 Approved by: https://github.com/malfet ghstack dependencies: #156944, #156945, #156946, #156947, #156948, #156949	2025-06-27 22:48:48 +00:00
nautsimon	a95bee9ed6	[MTIA ATen Backend] Add dispatch key for div.out (#156949 ) Migrate div.out Differential Revision: [D77063371](https://our.internmc.facebook.com/intern/diff/D77063371/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156949 Approved by: https://github.com/malfet ghstack dependencies: #156944, #156945, #156946, #156947, #156948	2025-06-27 22:48:39 +00:00
nautsimon	f30e072cb4	[MTIA ATen Backend] Add dispatch keys for mul.Scalar_out / mul.out (#156948 ) Migrate mul.Scalar_out / mul.out Differential Revision: [D77011801](https://our.internmc.facebook.com/intern/diff/D77011801/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156948 Approved by: https://github.com/malfet ghstack dependencies: #156944, #156945, #156946, #156947	2025-06-27 22:48:32 +00:00
nautsimon	66ad843583	[MTIA ATen Backend] Add dispatch keys for gt.Tensor_out / gt.Scalar_out (#156947 ) Migrate gt.Tensor_out / gt.Scalar_out Differential Revision: [D77009468](https://our.internmc.facebook.com/intern/diff/D77009468/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156947 Approved by: https://github.com/malfet ghstack dependencies: #156944, #156945, #156946	2025-06-27 22:48:25 +00:00
nautsimon	f0a5a3b453	[MTIA ATen Backend] Add dispatch keys for ne.Tensor_out / ne.Scalar_out (#156946 ) Migrate ne.Tensor_out / ne.Scalar_out Differential Revision: [D77008139](https://our.internmc.facebook.com/intern/diff/D77008139/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156946 Approved by: https://github.com/malfet ghstack dependencies: #156944, #156945	2025-06-27 22:48:18 +00:00
Dylan Maloy	cd1a924dba	[nativert] get rid of sigmoid naming (#157134 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D77451215 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157134 Approved by: https://github.com/zhxchen17, https://github.com/jingsh	2025-06-27 22:41:52 +00:00
Jane Xu	d283fc79b1	chunk_size should always be int64_t for Foreach functors (#156872 ) See https://github.com/pytorch/pytorch/issues/156261#issuecomment-3002394773 Testing is a valid q--it is pretty expensive to test such large tensors for all these ops. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156872 Approved by: https://github.com/Skylion007, https://github.com/eqy ghstack dependencies: #156876, #156871	2025-06-27 22:35:34 +00:00
Jane Xu	5a0926a26e	Stop skipping entire foreach tests, just skip the profiler portion (#156871 ) Instead of skipping the whole test as the CUPTI team figures out what is wrong, let's temporarily skip the profiler check portion. It is high pri to add it back to ensure foreach ops are actually performant. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156871 Approved by: https://github.com/albanD ghstack dependencies: #156876	2025-06-27 22:35:34 +00:00
Sandeep Narendranath Karjala	20e40492b0	[dynamo] Add fx_graph_runnable test coverage (#157021 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157021 Approved by: https://github.com/StrongerXi, https://github.com/xmfan	2025-06-27 21:35:56 +00:00
Morrison Turnansky	130d4973bd	Documentation update torch.clone #156644 (#157007 ) updated torch clone docs to reflect implemented memory behavior Fixes #156644 Pull Request resolved: https://github.com/pytorch/pytorch/pull/157007 Approved by: https://github.com/malfet, https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-27 21:10:09 +00:00
nautsimon	3ee75b7eac	[MTIA ATen Backend] Add dispatch keys for le.Tensor_out / le.Scalar_out (#156945 ) Migrate le.Tensor_out / le.Scalar_out Differential Revision: [D77002317](https://our.internmc.facebook.com/intern/diff/D77002317/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156945 Approved by: https://github.com/malfet ghstack dependencies: #156944	2025-06-27 21:03:19 +00:00
nautsimon	6b7767fc8d	[MTIA ATen Backend] Add dispatch keys for ge.Tensor_out / ge.Scalar_out (#156944 ) Migrate ge.Tensor_out / ge.Scalar_out Differential Revision: [D77002145](https://our.internmc.facebook.com/intern/diff/D77002145/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156944 Approved by: https://github.com/malfet	2025-06-27 21:02:27 +00:00
PyTorch MergeBot	0decd966af	Revert "Fixes for CPython int/float tests (#155978 )" This reverts commit 216bd6091ec52865052282eced7e6d5d2a4b4fb4. Reverted https://github.com/pytorch/pytorch/pull/155978 on behalf of https://github.com/huydhn due to Some tests are still failing in trunk ([comment](https://github.com/pytorch/pytorch/pull/155978#issuecomment-3014185210))	2025-06-27 19:39:41 +00:00
Philip Maybank	7c51619e7f	Fix Float16 CooperativeReduction Test Failure (#154516 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154516 Approved by: https://github.com/jansel, https://github.com/jeffdaily	2025-06-27 19:31:49 +00:00
Jane Xu	4048a144ab	Address richard's comments on libtorch_stable_abi note (#156324 ) Followups from #155984 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156324 Approved by: https://github.com/zou3519	2025-06-27 19:19:12 +00:00
Tugsbayasgalan (Tugsuu) Manlaibaatar	dcb97cd519	Remove unneccesary code to check autograd state (#156855 ) Summary: Title Test Plan: CI Rollback Plan: Differential Revision: D77317627 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156855 Approved by: https://github.com/zhxchen17 Co-authored-by: Camyll Harajli <camyllh@meta.com>	2025-06-27 19:18:06 +00:00
Lucas Beyer	8a88c6e85a	[nit] fix xavier init doc (#157100 ) Remove part of the documentation that is irrelevant and confusing at best, probably a copy-paste mistake: <img src="https://github.com/user-attachments/assets/77fa5734-5a5a-4f8d-80a5-bc3269668e07" width="500"> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157100 Approved by: https://github.com/mikaylagawarecki	2025-06-27 19:13:40 +00:00
PyTorch MergeBot	75a7d9e868	Revert "python definitely_contiguous-> is_contiguous_or_false (#156515 )" This reverts commit 4c0091fda65b714fa73671a15e379f814af153e0. Reverted https://github.com/pytorch/pytorch/pull/156515 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it seems to cause some torch.export failures internally ([comment](https://github.com/pytorch/pytorch/pull/156515#issuecomment-3014104570))	2025-06-27 19:07:06 +00:00
Svetlana Karslioglu	2860f5c4f5	Remove mentioning of TorchScript in Export doc (#156969 ) Remove mentioning of TorchScript Pull Request resolved: https://github.com/pytorch/pytorch/pull/156969 Approved by: https://github.com/angelayi Co-authored-by: Angela Yi <yiangela7@gmail.com>	2025-06-27 17:59:15 +00:00
vfdev	456b7451c7	Minor error message fix in device_mesh.py (#157096 ) Fixed error message: On main: ``` KeyError: ("Invalid mesh_dim_names ('dp_shard', 'dp_shard') specified. ", 'Found mesh dim indices to slice: [(1,), (1,)]. ', 'Mesh dim indices should be in ascending order.') ``` On PR: ``` KeyError: Invalid mesh_dim_names ('dp_shard', 'dp_shard') specified. Found mesh dim indices to slice: [(1,), (1,)]. Mesh dim indices should be in ascending order.' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/157096 Approved by: https://github.com/Skylion007	2025-06-27 17:42:29 +00:00
Justin Chu	36fd1ac932	[ONNX] Bump onnxscript api for torch 2.8 (#157017 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157017 Approved by: https://github.com/titaiwangms, https://github.com/malfet	2025-06-27 17:39:17 +00:00
henrylhtsang	84c588e5ea	[cutlass backend][BE][ez] Make matmul layouts be row x column (#156656 ) Differential Revision: [D77184232](https://our.internmc.facebook.com/intern/diff/D77184232/) Motivation: * This is the case we care the most. * We are caching the kernels for this row x column layout. So testing on them can potentially make ci run faster. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156656 Approved by: https://github.com/ColinPeppler	2025-06-27 17:15:45 +00:00
Wanchao Liang	b22b93a6ba	[2/n] rewrite load balancing and sharding in context parallel (#155442 ) This PR rewrite how load balancing and sharding works in the current context parallel implementation. Why the changes? We should NOT expose another layer of "sharding" concept as it would confuse the user about its difference with DTensor sharding. The current CP perform sharding weirdly simply because it mixed the concept of load balancing and sharding. I think load balancing and sharding need to be decoupled to separate layers: * The load balancing layer is responsible to reorder the input sequence so that the attention computation are evenly balanced across rows/ranks. * Sharding is a separate layer after it, it simply take the input reordered by the load balancer and shard it exactly as how DTensor shard tensor sequentially In this PR: * I removed the "Sharder" and "LoadBalancer" mixed usage, and simply generate a roundrobin indices when the mask is a casual mask * use `distribute_tensor` to perform the sharding. We still keep the local shard instead of the DTensor objects to allow maximum compatibility with arbitrary model architecture given DTensor op coverage is not high enough. One alternative design is to still keep the LoadBalancer and add the indices generation and restore to be the protocol of the LoadBalancer. I thought through it and think we might want to directly expose the load_balancing indices as an argument instead of a dedicated class interface, so I removed it here. More discussion on this is welcomed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155442 Approved by: https://github.com/XilunWu ghstack dependencies: #155441	2025-06-27 17:06:42 +00:00
Wanchao Liang	f7c730107e	[1/n] refactor the ring attention implementation (#155441 ) as titled, I'm working on a series of changes to make ring attention impl and DTensor works better together, this PR specifically refactor the current implemtnation to: * remove dead/unused code * restructure the functions to make them stay organized * refactor to remove/make error message better Pull Request resolved: https://github.com/pytorch/pytorch/pull/155441 Approved by: https://github.com/fegin	2025-06-27 17:06:42 +00:00
Tuan Trieu	eeaefa1336	Fix UnbackedSymint rebinding - check unbacked before renaming (#156911 ) Differential Revision: D77249427 Due to memoization and graph order update, it can happen that a backed symbol is passed into compute_unbacked_bindings and lead to failure. An example as follow: - There are 2 boolean indexing operators (e.g. op1 and op2) with the same mask. - A unbacked symint is generated from op1, and then op2 reuses the unbacked symint due to a nonzero_memo in nonzero's fake implementation and no rebinding is needed for op2. - Since op1 generated the unbacked symint, its meta has "unbacked_bindings" field filled and op2's meta doesn't have it. - Output from op1 and op2 are later concated with others with backed symint, so that the unbacked symint can be replaced by a backed symint. - In Inductor, during fake tensor prop, there is no memoi because new fake tensor is always generated (for the same node). op1 generates an unbacked symint and the unbacked can be rebound successfully to the backed symint. Since there is no memoi, op2 also generates a new unbacked symint, but no rebinding can happen because op2's meta doesn't have "unbacked_bindings". And "compute_unbacked_bindings/_rename_unbacked_to" fails to assert op2's old symbol to be unbacked. From discussion with [@ezyang](https://www.internalfb.com/intern/profile/?id=503862770), there is no easy way to fix this issue. - We can try to enable memoization for fake tensor prop in Inductor, however, we need to ensure that op1 is visited before op2 during Inductor fake tensor prop for this to work (op2's meta doesn't have "unbacked_bindings" so no rebinding can happen and we need to do rebinding from op1. But there are passes such as reorder_for_locality that can change the graph order so this doesn't work. - A simple hack is to just replace the unbacked symbol in op2 by the backed symbol. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156911 Approved by: https://github.com/ezyang	2025-06-27 16:57:04 +00:00
Guilherme Leobas	216bd6091e	Fixes for CPython int/float tests (#155978 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155978 Approved by: https://github.com/zou3519	2025-06-27 16:41:00 +00:00
fduwjj	d0cfa3e5bf	[c10d] Move the include of header file of TraceUtils.h into NCCLUtil.cpp instead of keeping in hpp (#156909 ) We have seen complaint about compilation failure of `NCCLSymmetricMemory.cu` and the reason is because we include <torch/csrc/distributed/c10d/TraceUtils.h> inside NCCLUtil.hpp this is not necessary so we want to move the include to cpp. Differential Revision: [D77346675](https://our.internmc.facebook.com/intern/diff/D77346675) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156909 Approved by: https://github.com/kwen2501	2025-06-27 16:30:49 +00:00
Nikita Shulga	21b5dc7a6a	[CD] Add python-3.14.0b3 to docker image (#156889 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156889 Approved by: https://github.com/albanD, https://github.com/atalman ghstack dependencies: #157033	2025-06-27 16:24:39 +00:00
Andrey Talman	d158e9ea82	Update nightly PyTorch version to 2.8.0->2.9.0 (#156965 ) Same as https://github.com/pytorch/pytorch/pull/149038 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156965 Approved by: https://github.com/Camyll, https://github.com/malfet	2025-06-27 16:22:08 +00:00
Jason Ansel	60abb0d327	[dynamo] Better error for invalid @contextlib.contextmanager usage (#156924 ) Fixes #156716 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156924 Approved by: https://github.com/williamwen42	2025-06-27 15:50:36 +00:00
Zizeng Meng	ff8b53c056	[Kineto] Add MTIA_INSIGHT to kineto_shim (#156853 ) Summary: Add MTIA_INSIGHT to kMtiaTypes in kineto_shim.cpp For insight, user can use MTIA_INSIGHT_VERBOSE_TRACES=0 to disable the profiler. So, we can enable it by default Test Plan: {F1979756361} When the environment var isn't set, it uses 0. Rollback Plan: Differential Revision: D77315882 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156853 Approved by: https://github.com/sraikund16	2025-06-27 15:30:14 +00:00
Aleksandar Samardžić	5118a8f8a5	Rename mm_scaled_grouped.py to mm_grouped.py (#156849 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156849 Approved by: https://github.com/amjames, https://github.com/Skylion007	2025-06-27 15:02:22 +00:00
rzou	aa2d54148d	Add AOTDispatcher config to set backward autocast behavior (#156356 ) This PR adds a new config `backward_pass_autocast`, to set the backward autocast behavior. It does not change the existing behavior. The reason why we need this is that torch.compile acquires a forward and backward graph at the time of the forward pass. This means that implemented naively, if there are any context managers active outside the call to torch.compile, the backward graph will also get the behaviors from those context managers. This PR gives users a way to tweak the autocast behavior of the backward pass. Please see torch._functorch.config for the options to the `backward_pass_autocast` config. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156356 Approved by: https://github.com/bdhirsh ghstack dependencies: #155354	2025-06-27 14:58:58 +00:00
Howard Huang	adf9644440	Add pg transport and tests (#154653 ) Add PG transport and tests under `torch/distributed/checkpoint/` ### API: ```python def send_checkpoint(self, dst_ranks: list[int], state_dict: object) -> None: def recv_checkpoint(self, src_rank: int) -> object: ``` ### Tests: ``` python test/distributed/checkpoint/test_pg_transport.py ``` ### Example: Under `_pg_transport_example.py` (in https://github.com/pytorch/pytorch/pull/155810) ``` torchrun --nproc_per_node=2 -m torch.distributed.checkpoint._pg_transport_example -- --device cuda ``` Differential Revision: [D76044919](https://our.internmc.facebook.com/intern/diff/D76044919) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154653 Approved by: https://github.com/meetv18	2025-06-27 14:53:34 +00:00
Vasiliy Kuznetsov	414ad47045	revamp dtype documentation for 2025 (#156087 ) The dtype documentation has not been updated in awhile, let's do a revamp. 1. combine the duplicated docs for dtypes from `tensors.rst` and `tensor_attributes.rst` to live in `tensor_attributes.rst`, and link to that page from `tensors.rst` 2. split the dtype table into floating point and integer dtypes 3. add the definition of shell dtype 4. add the float8 and MX dtypes as shell dtypes to the dtype table 5. remove legacy quantized dtypes from the table 6. add the definition of various dtype suffixes ("fn", etc) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156087 Approved by: https://github.com/albanD	2025-06-27 13:10:23 +00:00
rzou	43523bf168	Fix silent incorrectness arising from incorrect alias information (#152011 ) Fixes #136662 There are two problems: 1) canonicalize_view_scatter_ops adds some new nodes into the graph. These new nodes cause the alias info on the graph to be wrong. To fix this, we try to run FakeTensorUpdater on the graph again. 2) FakeTensorUpdater's alias information is wrong. It tries to skip nodes that it thinks have "equivalent" FakeTensor metadata. It should not be allowed to do this if any users of the node can alias the node. The example is if we have `x = foo(...); y = x.view(...)`. If the user replaces `foo` with a new `bar` node and sets bar.meta["val"] correctly, then FakeTensorUpdater still needs to update y's meta["val"] to be a view of the new bar node. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152011 Approved by: https://github.com/yf225	2025-06-27 12:45:03 +00:00
Jason Ansel	75f3e5a88d	[dynamo] Fix issue with tensors passed as view() shapes (#156928 ) Fixes #156720 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156928 Approved by: https://github.com/ezyang	2025-06-27 08:52:31 +00:00
Jason Ansel	588b5fb94b	Optimize TorchHigherOrderOperatorVariable.make() with lookup table (#157022 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157022 Approved by: https://github.com/zou3519	2025-06-27 07:36:12 +00:00
tvukovic-amd	968f90ce73	[ROCm][Windows] Fixing undefined symbol linker error after exposing MIOpen symbols (#156479 ) Fixing undefined symbol linker error after [exposing MIOpen symbols](https://github.com/pytorch/pytorch/pull/154545). This fix: - Hipifies `aten/src/ATen/miopen` and `aten/src/ATen/native/miopen` files - Adds `aten/src/ATen/miopen` and `aten/src/ATen/native/miopen` hipified source files to `all_hip_cpp` list Pull Request resolved: https://github.com/pytorch/pytorch/pull/156479 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-27 07:23:32 +00:00
PyTorch MergeBot	4a80ddfbe7	Revert "Fix reinplace pass handling of view input + mutable custom op (#156729 )" This reverts commit b754b1fa43d20f5b31e17c396487ab56991912da. Reverted https://github.com/pytorch/pytorch/pull/156729 on behalf of https://github.com/davidberard98 due to breaks lint: [GH job link](https://github.com/pytorch/pytorch/actions/runs/15918483073/job/44900430950) [HUD commit link](`b754b1fa43`) ([comment](https://github.com/pytorch/pytorch/pull/156729#issuecomment-3011867746))	2025-06-27 06:38:58 +00:00
cyy	064288cbab	Use std::string_view in torchgen (#157050 ) Let the generated code use std::sv Pull Request resolved: https://github.com/pytorch/pytorch/pull/157050 Approved by: https://github.com/ezyang	2025-06-27 06:36:10 +00:00
Laith Sakka	cc3ea2d840	remove gso from Linear.cpp (#156899 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156899 Approved by: https://github.com/ColinPeppler	2025-06-27 06:30:50 +00:00
zeshengzong	cf0749c92f	Use expecttest in test_compiled_optimizers.py (#155308 ) Fixes #141262 ## Test Result ```bash pytest test/inductor/test_compiled_optimizers.py -vv ``` ![image](https://github.com/user-attachments/assets/1886fb71-ff05-46e7-988c-82d36358a834) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155308 Approved by: https://github.com/mlazos, https://github.com/msaroufim Co-authored-by: Mark Saroufim <marksaroufim@gmail.com>	2025-06-27 06:29:51 +00:00
Laith Sakka	cbcffce48a	address remaining straight forward gso in meta_registrations (#156902 ) Those are all straight forward generalization of existing checks, Pull Request resolved: https://github.com/pytorch/pytorch/pull/156902 Approved by: https://github.com/ColinPeppler	2025-06-27 06:19:54 +00:00
Menglu Yu	640703d95f	add torch.concat to normalization pass (#156574 ) Summary: In the normalization pass, we also add torch.concat to it to normalize it as torch.cat Test Plan: ``` buck2 test 'fbcode//mode/dev-nosan' fbcode//caffe2/test/inductor:split_cat_fx_passes -- test_cat_normalization ``` Buck UI: https://www.internalfb.com/buck2/597fd4f1-0aa7-4372-8a66-5a690d9b63a4 Test UI: https://www.internalfb.com/intern/testinfra/testrun/1688850152284203 Network: Up: 84KiB Down: 34KiB (reSessionID-3916e009-7117-41ce-b6f9-089873aa50dd) Executing actions. Remaining 0/3 1.1s exec time total Command: test. Finished 2 local Time elapsed: 3:47.1s Tests finished: Pass 2. Fail 0. Fatal 0. Skip 0. Build failure 0 Rollback Plan: Differential Revision: D77125331 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156574 Approved by: https://github.com/Mingming-Ding	2025-06-27 06:07:26 +00:00
Daisy Deng	1155c53e7d	Port three dynamo test to Intel GPU (#156575 ) For https://github.com/pytorch/pytorch/issues/114850, we will port test cases to Intel GPU. Two dynamo test files were ported in PR [#156056](https://github.com/pytorch/pytorch/pull/156056). In this PR we will port 3 more dynamo test files. We could enable Intel GPU with following methods and try the best to keep the original code styles: - instantiate_device_type_tests() - use "torch.accelerator.current_accelerator()" to determine the accelerator backend - added XPU support in decorators like @requires_gpu - enabled XPU for some test path. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156575 Approved by: https://github.com/guangyey, https://github.com/jansel Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-06-27 05:56:22 +00:00
Jason Ansel	51853b358e	[dynamo] Improve error message for cond aliasing (#156963 ) See #156724 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156963 Approved by: https://github.com/zou3519, https://github.com/williamwen42	2025-06-27 05:31:46 +00:00
David Berard	6b05842e47	[test][inductor] fix test_conv_cat failure (#155852 ) This test is currently failing because triton_poi_fused_cat_2 has changed to triton_poi_fused_cat_3. I have not investigated why the extra kernel is generated, but this test has been failing on trunk for a while (and I verified locally that it is failing). Pull Request resolved: https://github.com/pytorch/pytorch/pull/155852 Approved by: https://github.com/FindHao, https://github.com/Skylion007	2025-06-27 05:11:11 +00:00
Laith Sakka	2c76f31221	Compute contiguity symbolically to avoid dde, and introduce c++ sym_is_contiguous (#155590 ) When we compute contiguity for a tensor with dynamic shapes we first: 1) Try to compute it without guarding. 2) If all shapes hinted, compute it with potentially adding guards. 3) if any input is not hinted, compute it symbolically. sym_is_contiguous return a SymBool that is then either evaluated or guard_or_false can be called on it to avoid data dependent errors. ex: bool is_contiguous = input.sym_is_contiguous().guard_or_false(__FILE__, __LINE__); is_contiguous_or_false is a helper function that does that. In this PR I only handle default contiguity, will follow up with changes for other formats like channel_last . We use this patter in this PR for several locations to avoid DDEs. Differential Revision: [D77183032](https://our.internmc.facebook.com/intern/diff/D77183032) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155590 Approved by: https://github.com/ezyang	2025-06-27 04:59:52 +00:00
Will Feng	b754b1fa43	Fix reinplace pass handling of view input + mutable custom op (#156729 ) Fixes #153389. Using approach https://github.com/pytorch/pytorch/issues/153389#issuecomment-3006049928 suggested by Richard. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156729 Approved by: https://github.com/zou3519	2025-06-27 04:54:17 +00:00
Divyansh Khanna	e6d8ed02cb	PyTorch Data Sampler benchmark (#156974 ) ## Motivation Many PRs optimizing samplers (for eg https://github.com/pytorch/pytorch/pull/147706, https://github.com/pytorch/pytorch/pull/137423) are leveraging an adhoc script for benchmarking samplers. The script and outputs are often copied over in PRs. We want to begin centralizing benchmarks for torch.utils.data components. ## What ? * This PR adds a new sub-folder in `benchmarks` for `data`. This is aimed to cover benchmarking scripts for torch.utils.data components like dataloader and sampler. * Specifically, this PR includes a simple script to time samplers. This is often "copy-pasted" in PRs optimizing samplers. Having it in a centralized location should prevent that, and allow a common standard. ## Output ``` Benchmark Results: +--------------+-------------+----------------+-----------+-----------+ \| Batch Size \| Drop Last \| Original (s) \| New (s) \| Speedup \| +==============+=============+================+===========+===========+ \| 4 \| True \| 0.004 \| 0.0088 \| -119.62% \| +--------------+-------------+----------------+-----------+-----------+ \| 4 \| False \| 0.0083 \| 0.009 \| -9.23% \| +--------------+-------------+----------------+-----------+-----------+ \| 8 \| True \| 0.003 \| 0.0074 \| -147.64% \| +--------------+-------------+----------------+-----------+-----------+ \| 8 \| False \| 0.0054 \| 0.0075 \| -38.72% \| +--------------+-------------+----------------+-----------+-----------+ \| 64 \| True \| 0.0021 \| 0.0056 \| -161.92% \| +--------------+-------------+----------------+-----------+-----------+ \| 64 \| False \| 0.0029 \| 0.0055 \| -92.50% \| +--------------+-------------+----------------+-----------+-----------+ \| 640 \| True \| 0.002 \| 0.0055 \| -168.75% \| +--------------+-------------+----------------+-----------+-----------+ \| 640 \| False \| 0.0024 \| 0.0062 \| -161.35% \| +--------------+-------------+----------------+-----------+-----------+ \| 6400 \| True \| 0.0021 \| 0.0055 \| -160.13% \| +--------------+-------------+----------------+-----------+-----------+ \| 6400 \| False \| 0.0021 \| 0.0068 \| -215.46% \| +--------------+-------------+----------------+-----------+-----------+ \| 64000 \| True \| 0.0042 \| 0.0065 \| -55.29% \| +--------------+-------------+----------------+-----------+-----------+ \| 64000 \| False \| 0.0029 \| 0.0077 \| -169.56% \| +--------------+-------------+----------------+-----------+-----------+ ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156974 Approved by: https://github.com/ramanishsingh	2025-06-27 04:49:43 +00:00
codingwithsurya	195ef1bce8	[SymmMem] Refactor NVSHMEM tests: separate Triton tests into dedicated file (#156685 ) ## Summary Moved the Triton-specific NVSHMEM tests in `test_nvshmem.py` into a dedicated `test_nvshmem_triton.py` file. Also put the shared Triton JIT kernels at the top-level of new file for reusability. ## Testing ```bash TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem_triton.py ``` All 16 original tests pass with no functionality changes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156685 Approved by: https://github.com/mandroid6, https://github.com/kwen2501 ghstack dependencies: #156684	2025-06-27 04:38:37 +00:00
David Berard	b6c00dfe24	[user triton] AOT inductor support for device-side TMA (#155896 ) Tests: `python test/inductor/test_aot_inductor.py -vvv -k device_tma` Device-side TMA in Triton allows the kernel author to construct the TMA descriptor on the device (which composes with things like autotuning much better). However, it also requires a scratch space to be provided into which the TMA descriptor will be constructed. In the new TMA API (tl.make_tensor_descriptor), this is implemented using a "global scratch space" - a tensor which is allocated beforehand and then passed in as an argument for the kernel. To support this in AOTI, this PR: * records the global scratch space needed (triton_heuristics.py), so that it can be used during AOTI codegen * allocates global scratch, if needed (cuda/device_op_overrides.py) * plumbs `device_idx_` into the triton caller function, so that global scratch can be allocated on the right device) * updates tests to verify this works for dynamically shaped inputs This PR should support both inductor-generated device-side TMA (e.g. persistent TMA mm) and user-defined triton kernels that contain device-side TMA (which is the test I ran to verify this works) Note: this overrides any user-provided allocator function (typically with eager triton code, the user must provide their own custom allocator function that is used to allocate scratch space). For Meta reviewers, here is a tlparse from running `python test/inductor/test_aot_inductor.py -vvv -k test_triton_kernel_on_device_tma_dynamic_True_tma_version_new_cuda` https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmpFg13g1/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 Differential Revision: [D77352139](https://our.internmc.facebook.com/intern/diff/D77352139) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155896 Approved by: https://github.com/desertfire	2025-06-27 04:28:04 +00:00
Nikita Shulga	710b92cf3b	[BE][BugFix] Install Python-3.13 correctly (#157033 ) Fixes temporary workaround introduced by https://github.com/pytorch/builder/pull/1827 I.e. it's been downloading latest 3.13 branch rather than 3.13.0 release Simplify nogil version handling Pull Request resolved: https://github.com/pytorch/pytorch/pull/157033 Approved by: https://github.com/wdvr, https://github.com/huydhn	2025-06-27 04:19:59 +00:00
leslie-fang-intel	1eea2c4fe3	[Inductor][CPP] Fix perf regression of functorch_maml_omniglot (#156526 ) Summary Fix the performance regression of `functorch_maml_omniglot` in TorchBench. The issue reported in [#151523](https://github.com/pytorch/pytorch/issues/151523) occurs only when a parallel reduction is performed under the vectorized loop and a scalar kernel is used for the tail loop. Previously, we addressed this regression in [#151887](https://github.com/pytorch/pytorch/pull/151887) by disabling all cases where a parallel reduction occurs under the vectorized loop. However, for `functorch_maml_omniglot`, we found that a masked vector kernel is used in the tail loop instead of the scalar kernel in the job of `inductor_torchbench_cpu_smoketest_perf`. In this PR, we refine the fix by excluding the cases where a masked vector kernel is used in the tail loop, rather than disabling all such scenarios. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156526 Approved by: https://github.com/CaoE	2025-06-27 03:09:24 +00:00
dolpm	7392470da4	[nativert] alias analyzer + layout planner/manager to pytorch core (#156897 ) Summary: att Test Plan: ci - unit tests still have some unresolved deps but will move them later. Rollback Plan: Differential Revision: D77320950 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156897 Approved by: https://github.com/zhxchen17	2025-06-27 03:01:22 +00:00
Raman Kumar	382c6190c1	complex.pow(2) on GPU by replacing with complex * complex to avoid numerical instability (#152373 ) Fixes #150951 Summary: For complex.pow(2) on GPU: Uses complex * complex directly. Produces results consistent with CPU implementation. Eliminates spurious imaginary components for real inputs. 🧪 Tests Added unit tests to verify correctness of the new kernel path. Verified numerical consistency with CPU results. This change is backward-compatible and only affects the specific case of pow(2) on complex tensors on GPU. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152373 Approved by: https://github.com/ezyang	2025-06-27 02:21:59 +00:00
PyTorch MergeBot	e290a4c645	Revert "Rename torch::standalone to headeronly (#156964 )" This reverts commit 7e54c02a35b905e758497b856a1953eb009ba836. Reverted https://github.com/pytorch/pytorch/pull/156964 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/156964#issuecomment-3011136947))	2025-06-27 02:20:33 +00:00
Ke Wen	4ab4d29cbe	[BE] Remove SymmMem allocator destruct log (#157020 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/157020 Approved by: https://github.com/fduwjj	2025-06-27 02:10:54 +00:00
PyTorch MergeBot	56c69bedcc	Revert "[dynamo] Better error for invalid @contextlib.contextmanager usage (#156924 )" This reverts commit 863327ae496471654344e1e04ccaa713a44a135d. Reverted https://github.com/pytorch/pytorch/pull/156924 on behalf of https://github.com/jansel due to Likely same issue as #156963 ([comment](https://github.com/pytorch/pytorch/pull/156924#issuecomment-3011087802))	2025-06-27 01:57:05 +00:00
Tugsbayasgalan (Tugsuu) Manlaibaatar	8e8bbfc803	Remove ts to export retracer (#156857 ) Summary: This is probably not used anymore Test Plan: CI Rollback Plan: Reviewed By: SherlockNoMad Differential Revision: D77318582 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156857 Approved by: https://github.com/SherlockNoMad	2025-06-27 01:54:24 +00:00
Ryan Guo	a4b59498c5	Fix fake kernel for the `out=...` variant of `unbind_copy` (#156643 ) `unbind_copy(..., out=...)` returns None rather than the `out` argument (see https://github.com/pytorch/pytorch/issues/130829#issuecomment-2283936222), but the old fake kernel didn't account for that and caused an assertion failure in `pushPyOutToStack`. This patch fixes that. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156643 Approved by: https://github.com/zou3519, https://github.com/jansel, https://github.com/bdhirsh ghstack dependencies: #156642	2025-06-27 01:34:07 +00:00
Ryan Guo	89aa708b39	[core] Dispatch to `at::nansum_out` rather than `at::native::nansum_out` (#156642 ) Calling `at::native::nansum_out` causes the fake kernel to dispatch to a `make_reduction` call and then segfaults later due to the `mutable_data_ptr` call in `TensorIteratorBase::build`. It also causes fake tensor propagation issue in Dynamo. The added tests demonstrate the aforementioned 2 issues. This patch fixes it by dispatching to `at::nansum_out` instead. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156642 Approved by: https://github.com/zou3519	2025-06-27 01:34:07 +00:00
Jason Ansel	863327ae49	[dynamo] Better error for invalid @contextlib.contextmanager usage (#156924 ) Fixes #156716 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156924 Approved by: https://github.com/williamwen42	2025-06-27 01:02:01 +00:00
Jane Xu	7e54c02a35	Rename torch::standalone to headeronly (#156964 ) Summary: headeronly is more clear, let's change the name before anyone depends on standalone Test Plan: CI should pass! Rollback Plan: Differential Revision: D77381084 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156964 Approved by: https://github.com/swolchok, https://github.com/albanD, https://github.com/desertfire	2025-06-27 01:00:14 +00:00
Nicolas Macchioni	3bdd5ae334	[PT2] deprecate `force_same_precision`, guarded by JK (#156789 ) Summary: cuBLAS used to have strict alignment requirements for TF32 usage, even if TF32 was enabled by users; this caused a numeric SEV in the past, when Triton would use TF32 even if cuBLAS could not due to failing the alignment checks we believe that cuBLAS no longer has alignment requirements for TF32 usage, based on some testing in D77265581; we'd like to deprecate `force_same_precision` since it no longer functions as expected changing the default to False in fbcode, guarded by a jk so that we can quickly revert to the original behavior if needed Test Plan: CI Rollback Plan: Differential Revision: D77265930 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156789 Approved by: https://github.com/jhadidjojo, https://github.com/masnesral	2025-06-27 00:43:06 +00:00
PyTorch MergeBot	6215e90b7b	Revert "[dynamo] Improve error message for cond aliasing (#156963 )" This reverts commit 9c39bc24807a5843f8affdf56bd71836760dc554. Reverted https://github.com/pytorch/pytorch/pull/156963 on behalf of https://github.com/huydhn due to Sorry for reverting your PR, but the failures are legit ([comment](https://github.com/pytorch/pytorch/pull/156963#issuecomment-3010870664))	2025-06-27 00:31:00 +00:00
PyTorch MergeBot	e3977e843d	Revert "Fix silent incorrectness arising from incorrect alias information (#152011 )" This reverts commit 2d39a48d524021995269411bd49fe792e59d9f94. Reverted https://github.com/pytorch/pytorch/pull/152011 on behalf of https://github.com/Camyll due to cannot land internally. owner will update and reland to fix ([comment](https://github.com/pytorch/pytorch/pull/152011#issuecomment-3010723960))	2025-06-26 23:54:13 +00:00
William Wen	eb9efb37c8	[dynamo] fix _torchdynamo_orig_callable naming issues (#156901 ) `_torchdynamo_orig_callable` was being used in two distinct places: - to get the original user function from nested eval_frame.py decorators - to get the original backend from nested convert_frame.py callbacks We rename the first usage to `_torchdynamo_orig_fn` and the second to `_torchdynamo_orig_backend` in order to distinguish these cases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156901 Approved by: https://github.com/StrongerXi, https://github.com/jansel ghstack dependencies: #156527	2025-06-26 23:51:08 +00:00
William Wen	6089ebcf6d	[dynamo] fix segfault due to dangling CacheEntry backend pointer (#156527 ) Fixes https://github.com/pytorch/pytorch/issues/155057 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156527 Approved by: https://github.com/anijain2305, https://github.com/jansel	2025-06-26 23:51:08 +00:00
Kurt Mohler	e0447bb5f8	Add `max_pool3d` for MPS (#156467 ) Fixes #100674 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156467 Approved by: https://github.com/malfet	2025-06-26 23:33:50 +00:00
Manuel Candales	1fff6356d9	[MPS] Optimize cummin/cummax metal kernels (#156794 ) Performance improvement (M4 Max 64GB, macOS 15.5): ``` \| Current \| Previous cummin-dim0-32x32 (torch.float16) \| 103.4 \| 102.5 cummin-dim0-128x128 (torch.float16) \| 112.2 \| 133.6 cummin-dim0-512x512 (torch.float16) \| 146.9 \| 233.1 cummin-dim0-1024x1024 (torch.float16) \| 193.6 \| 364.2 cummin-dim1-32x32 (torch.float16) \| 102.0 \| 94.4 cummin-dim1-128x128 (torch.float16) \| 103.0 \| 109.9 cummin-dim1-512x512 (torch.float16) \| 109.1 \| 227.0 cummin-dim1-1024x1024 (torch.float16) \| 140.5 \| 985.1 cummin-1d-100 (torch.float16) \| 101.8 \| 100.7 cummin-1d-10000 (torch.float16) \| 112.8 \| 805.0 cummin-1d-1000000 (torch.float16) \| 1343.8 \| 70545.6 cummin-dim0-32x32 (torch.float32) \| 104.6 \| 102.7 cummin-dim0-128x128 (torch.float32) \| 112.3 \| 137.2 cummin-dim0-512x512 (torch.float32) \| 146.6 \| 209.7 cummin-dim0-1024x1024 (torch.float32) \| 194.0 \| 340.1 cummin-dim1-32x32 (torch.float32) \| 100.1 \| 99.2 cummin-dim1-128x128 (torch.float32) \| 101.4 \| 111.9 cummin-dim1-512x512 (torch.float32) \| 110.3 \| 250.7 cummin-dim1-1024x1024 (torch.float32) \| 141.4 \| 987.9 cummin-1d-100 (torch.float32) \| 101.0 \| 100.6 cummin-1d-10000 (torch.float32) \| 112.9 \| 794.7 cummin-1d-1000000 (torch.float32) \| 1311.7 \| 71995.3 cummin-dim0-32x32 (torch.bfloat16) \| 105.8 \| 105.9 cummin-dim0-128x128 (torch.bfloat16) \| 111.9 \| 135.7 cummin-dim0-512x512 (torch.bfloat16) \| 147.1 \| 231.9 cummin-dim0-1024x1024 (torch.bfloat16) \| 191.2 \| 327.7 cummin-dim1-32x32 (torch.bfloat16) \| 101.8 \| 91.3 cummin-dim1-128x128 (torch.bfloat16) \| 100.2 \| 108.5 cummin-dim1-512x512 (torch.bfloat16) \| 108.9 \| 222.0 cummin-dim1-1024x1024 (torch.bfloat16) \| 140.1 \| 936.9 cummin-1d-100 (torch.bfloat16) \| 103.0 \| 106.6 cummin-1d-10000 (torch.bfloat16) \| 113.1 \| 795.8 cummin-1d-1000000 (torch.bfloat16) \| 1296.8 \| 68667.4 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156794 Approved by: https://github.com/malfet ghstack dependencies: #156860	2025-06-26 23:30:20 +00:00
Jason Ansel	9c39bc2480	[dynamo] Improve error message for cond aliasing (#156963 ) See #156724 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156963 Approved by: https://github.com/zou3519, https://github.com/williamwen42	2025-06-26 23:12:00 +00:00
Laith Sakka	e6ed4074e8	update expected results (#157010 ) <img width="1490" alt="Screenshot 2025-06-26 at 12 30 46 PM" src="https://github.com/user-attachments/assets/4df626d4-3010-4362-974c-fb96fa68b29f" /> <img width="904" alt="Screenshot 2025-06-26 at 12 28 29 PM" src="https://github.com/user-attachments/assets/42626892-27e1-4e69-9efc-c9baf80c5384" /> <img width="752" alt="Screenshot 2025-06-26 at 12 29 05 PM" src="https://github.com/user-attachments/assets/0b1afb30-5868-4ba6-9985-2cc7994a4227" /> PR https://github.com/pytorch/pytorch/pull/152011 added slight regression <br class="Apple-interchange-newline"> Pull Request resolved: https://github.com/pytorch/pytorch/pull/157010 Approved by: https://github.com/zou3519	2025-06-26 21:56:57 +00:00
William Wen	80d89974c1	[dynamo] raise hard error if error is encountered while tracing resume function prologue (#154564 ) This should prevent bad resume function prologues from slipping by. In particular, graph breaks in resume function prologues will now hard error. Implementation details: - The resume function prologue is surrounded by `LOAD_CONST arg, STORE_FAST __is_tracing_resume_prologue` instructions. The first sequence has `arg=True` and the second sequence has `arg=False`. - InstructionTranslator will know when it is tracing a resume function prologue when it detects `STORE_FAST __is_tracing_resume_prologue`. The top of stack will be True to mark the start of the prologue, False to mark the end. - When `convert_frame.py` detects that an error occurred while the InstructionTranslator was tracing a resume function prologue, we will wrap the exception and hard error Pull Request resolved: https://github.com/pytorch/pytorch/pull/154564 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289, #154782, #156762, #155166	2025-06-26 21:40:38 +00:00
William Wen	6df6eacce8	[dynamo] handle fullgraph toggle using nested torch.compile (#155166 ) See added test for the case that this PR handles. In particular, the semantics for nested torch.compile with toggled fullgraph settings was strange before - `@torch.compile(fullgraph=True)` overrides the existing fullgraph setting, while `@torch.compile(fullgraph=False)` does not. Note that this change will add an extra frame to any inlined torch.compile'd function (which I don't expect to happen frequently). Pull Request resolved: https://github.com/pytorch/pytorch/pull/155166 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289, #154782, #156762	2025-06-26 21:40:38 +00:00
William Wen	dcb8982969	[dynamo] move error_on_graph_break out of config (#156762 ) error_on_graph_break doesn't need to be in config, so we move it out. It should make the functorch_maml_omniglot regression less severe. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156762 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289, #154782	2025-06-26 21:40:38 +00:00
William Wen	36666033ab	[dynamo] fix set_fullgraph for nested calls (#154782 ) - Make the fullgraph argument of set_fullgraph a positional argument - Fix behavior on nested calls by updating `tracer.error_on_graph_break` in more places. In particular, a tracer's error_on_graph_break is set to the inlined tracer's error_on_graph_break upon the latter's exit. We also track error_on_graph_break in the speculation log now, since if we encounter a nested graph break, we will restart analysis and we need to somehow remember the error_on_graph_break setting after attempting to run the nested function (but we don't actually trace into it in the restart analysis). Pull Request resolved: https://github.com/pytorch/pytorch/pull/154782 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289	2025-06-26 21:40:38 +00:00
William Wen	7b7eafe7ba	[dynamo] add set_fullgraph decorator/context manager (#154289 ) Implements https://github.com/pytorch/pytorch/issues/144908. Implementation notes: - `set_fullgraph` is implemented using `patch_config`, which changes config correctly during runtime and tracing. - Moved setting `config.error_on_graph_break` from convert_frame.py to eval_frame.py. This is because this should only be done at the top-level decorated function. If we kept this in convert_frame.py, we would be changing `config.error_on_graph_break` on every top-level frame, which causes confusing behavior (see added test for example). - InstructionTranslator reads from `config.error_on_graph_break` every `step()`. This is to determine the value of `config.error_on_graph_break` at the time of the graph break, because tracer cleanup will restore the value of `config.error_on_graph_break` . - `convert_frame.py` determines whether we should abort tracing (fullgraph=True) or continue (fullgraph=False) by reading the value of the tracer's `error_on_graph_break`. If there is no tracer (failed to initialize), then default to reading `config.error_on_graph_break`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154289 Approved by: https://github.com/jansel, https://github.com/zou3519 ghstack dependencies: #154283	2025-06-26 21:40:38 +00:00
William Wen	1c3f5e902d	[dynamo] control one_graph behavior additionally through config (#154283 ) `torch.compile` now always goes through `torch._dynamo._optimize`. fullgraph is now implemented in `torch.compile` by looking at `config.error_on_graph_break`. Export still goes through `torch._dynamo._optimize_assert`, which uses `tx.one_graph` instead of `config.error_on_graph_break`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154283 Approved by: https://github.com/jansel, https://github.com/anijain2305	2025-06-26 21:40:38 +00:00
Ke Wen	fc10d4b1d6	[SymmMem] Allow selection of allocation backend (#156661 ) Stack from [ghstack](https://github.com/ezyang/ghstack) (oldest at bottom): Today the only way to choose allocation backend is via env `TORCH_SYMMMEM=...`. This is a bit hard to set in CI on test file basis. (The env has to be set before program is loaded). This PR added a programmatic way -- a `set_backend` API. Implementation: Since this API is slightly more dynamic than static registration, at static time each backend registers its availability rather than filling itself as the allocator directly. Later when `set_backend` is called, the allocator would actually fill in the device-to-allocation `map_`. Though added, `set_backend` is not a necessary API for user to call -- one backend is still registered as the default at static time. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156661 Approved by: https://github.com/ngimel, https://github.com/fduwjj	2025-06-26 21:37:44 +00:00
dolpm	262654ee51	[nativert] move constantfolder to libtorch (#156918 ) Summary: att -- unit tests will be migrated later, since they still have unresolved deps. Test Plan: ci Rollback Plan: Differential Revision: D77159278 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156918 Approved by: https://github.com/henryoier, https://github.com/zhxchen17	2025-06-26 21:26:37 +00:00
Geevarghese George	7f6e7103a3	Convert to markdown: jit_python_reference.rst, jit_unsupported.rst, jit_utils.rst, library.rst (#155404 ) Fixes #155024 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155404 Approved by: https://github.com/svekars	2025-06-26 21:09:46 +00:00
angelayi	aff9c1eec5	[aoti][mps] Add fused_rms and sdpa_mps fallback ops (#156844 ) Needed for llama3.1 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156844 Approved by: https://github.com/desertfire ghstack dependencies: #156843	2025-06-26 21:03:05 +00:00
angelayi	17dab018e3	[aoti][mps] Fix deduplication of kernels (#156843 ) Previously I was not correctly deduplicating kernels generated by mps, so it would generate multiple of the same kernel. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156843 Approved by: https://github.com/desertfire	2025-06-26 21:03:05 +00:00
Joshua Su	977abe786d	fix 'register_foward_pre_hook not supported on ScriptModule' error (#156904 ) Summary: Encountered 'register_foward_pre_hook not supported on ScriptModule' error when trying to publish CFR MTML with placing remote_ro module in remote. Issue may come from the fact that the local net from torchArrow is already scriptModule before gen_app_graph pass. {F1979770267} Test Plan: hg checkout 1ff14dfaade4ac1f3cbbf38fbd72f7fdd5cdcd16 bash hstu_blocker.sh Rollback Plan: Reviewed By: RenfeiChen-FB Differential Revision: D77341370 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156904 Approved by: https://github.com/jingsh	2025-06-26 20:59:24 +00:00
Dylan Maloy	81759afed4	[nativert] clean up some migration side-effects (#156919 ) Summary: explicit torch::nativert namespace usage + // manual declarations Test Plan: ci Rollback Plan: Differential Revision: D77328855 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156919 Approved by: https://github.com/zhxchen17	2025-06-26 20:28:32 +00:00
codingwithsurya	b6e625e34f	[SymmMem] Remove redundant dist.barrier in Triton NVSHMEM tests & add device‐side signal_op support (#156684 ) ## Summary This PR removes unnecessary `dist.barrier` calls up in our Triton NVSHMEM test suite and adds signal_op support, which is a lightweight device-side signaling mechanism. Added test for this in our `wait_until` kernel and corresponding `core.extern` wrapper. Why did we drop the `dist.barrier()` calls? We dropped the host‐side dist.barrier() in all Triton NVSHMEM tests (except the raw put/get cases) because every other test already uses NVSHMEM collectives or device‐side sync primitives (fence/quiet/signal/wait), making the extra barrier redundant. This keeps synchronization entirely on the GPU and leverages NVSHMEM’s native ordering guarantees for clearer, more efficient tests. `test_triton_wait_until` update - Rank 1: after `put_kernel` writes the data, launches `signal_op_kernel` to atomically set Rank 0's flag via `nvshmemx_signal_op` - Rank 0: drops its old `dist.barrier()` and simply calls `wait_until_kernel` to spin-wait on the device flag, then asserts data correctness - Changes made per [this comment](https://github.com/pytorch/pytorch/pull/156472#discussion_r2159734046) ## Testing ```bash TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156684 Approved by: https://github.com/kwen2501, https://github.com/mandroid6	2025-06-26 20:16:06 +00:00
Manuel Candales	43a09189c6	[MPS] Add benchmark for scan with indices (#156860 ) Baseline performance on M4 Max 64GB (macOS 15.5): ``` [-------------------------------- --------------------------------] \| eager \| compile 1 threads: --------------------------------------------------------- cummin-dim0-32x32 (torch.float16) \| 102.5 \| 115.0 cummin-dim0-128x128 (torch.float16) \| 133.6 \| 147.8 cummin-dim0-512x512 (torch.float16) \| 233.1 \| 243.1 cummin-dim0-1024x1024 (torch.float16) \| 364.2 \| 385.2 cummin-dim1-32x32 (torch.float16) \| 94.4 \| 109.8 cummin-dim1-128x128 (torch.float16) \| 109.9 \| 122.5 cummin-dim1-512x512 (torch.float16) \| 227.0 \| 233.8 cummin-dim1-1024x1024 (torch.float16) \| 985.1 \| 1010.5 cummin-1d-100 (torch.float16) \| 100.7 \| 114.3 cummin-1d-10000 (torch.float16) \| 805.0 \| 879.1 cummin-1d-1000000 (torch.float16) \| 70545.6 \| 71310.3 cummin-dim0-32x32 (torch.float32) \| 102.7 \| 115.5 cummin-dim0-128x128 (torch.float32) \| 137.2 \| 143.8 cummin-dim0-512x512 (torch.float32) \| 209.7 \| 222.0 cummin-dim0-1024x1024 (torch.float32) \| 340.1 \| 389.9 cummin-dim1-32x32 (torch.float32) \| 99.2 \| 107.8 cummin-dim1-128x128 (torch.float32) \| 111.9 \| 119.3 cummin-dim1-512x512 (torch.float32) \| 250.7 \| 255.1 cummin-dim1-1024x1024 (torch.float32) \| 987.9 \| 1013.2 cummin-1d-100 (torch.float32) \| 100.6 \| 114.6 cummin-1d-10000 (torch.float32) \| 794.7 \| 862.2 cummin-1d-1000000 (torch.float32) \| 71995.3 \| 71963.5 cummin-dim0-32x32 (torch.bfloat16) \| 105.9 \| 113.9 cummin-dim0-128x128 (torch.bfloat16) \| 135.7 \| 147.9 cummin-dim0-512x512 (torch.bfloat16) \| 231.9 \| 240.7 cummin-dim0-1024x1024 (torch.bfloat16) \| 327.7 \| 366.9 cummin-dim1-32x32 (torch.bfloat16) \| 91.3 \| 103.3 cummin-dim1-128x128 (torch.bfloat16) \| 108.5 \| 117.4 cummin-dim1-512x512 (torch.bfloat16) \| 222.0 \| 233.6 cummin-dim1-1024x1024 (torch.bfloat16) \| 936.9 \| 982.5 cummin-1d-100 (torch.bfloat16) \| 106.6 \| 112.4 cummin-1d-10000 (torch.bfloat16) \| 795.8 \| 819.6 cummin-1d-1000000 (torch.bfloat16) \| 68667.4 \| 68557.9 Times are in microseconds (us). ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156860 Approved by: https://github.com/malfet	2025-06-26 18:44:16 +00:00
PyTorch MergeBot	9fe2d156a9	Revert "[dynamo] fix segfault due to dangling CacheEntry backend pointer (#156527 )" This reverts commit 5ad2bee2c8a7defd2580bb138145a49c37146fcc. Reverted https://github.com/pytorch/pytorch/pull/156527 on behalf of https://github.com/Camyll due to failing test assertions ([comment](https://github.com/pytorch/pytorch/pull/156527#issuecomment-3009231797))	2025-06-26 17:32:34 +00:00
Nicolas Macchioni	13efb2c858	[BE] Deprecate `search_autotune_cache` (#155302 ) We haven't had the offline cache populated in > 1 year, this should be safe; if this passes, we can finally go through and rip out the offline cache logic Pull Request resolved: https://github.com/pytorch/pytorch/pull/155302 Approved by: https://github.com/masnesral	2025-06-26 17:30:08 +00:00
Nikita Shulga	039a1ce0eb	[BE] Remove CXX11_ABI references from cpp_builder.py (#156896 ) As all Linux builds are CXX11_ABI compatible at this point Pull Request resolved: https://github.com/pytorch/pytorch/pull/156896 Approved by: https://github.com/desertfire, https://github.com/jansel	2025-06-26 17:28:01 +00:00
Laith Sakka	e15ea965a1	remove guard_size_oblivious from unbind. (#148815 ) unbind will always specialize on dim, because it determine the number of output tensors. guard_size_oblivious is not useful there and more confusing probably for code readers added a comment and a test that verifies the specialization. Pull Request resolved: https://github.com/pytorch/pytorch/pull/148815 Approved by: https://github.com/pianpwk	2025-06-26 17:16:32 +00:00
Shangdi Yu	61eaaa21a4	Better error message when no .so/cpp files are found (#156863 ) Summary: Sample error message: ``` RuntimeError: Failed to find a generated cpp file or so file for model 'forward' in the zip archive. Available models in the archive: model To load a specific model, please provide its name using the `model_name` parameter when calling AOTIModelPackageLoader() or torch._inductor.package.load_package. The following files were loaded from the archive: c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/data/aotinductor/model/cqdxv6zki2oiiytjeqrg774uxlxgqdemhdxn5dycn4nnc3rmcd7w.cubin c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/data/aotinductor/model/c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper.cpp c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/data/aotinductor/model/ctmp7adn3spwyscdotllyj4yx3vrqcnxk3thkpgdcax7zvqmyyp3.kernel.cpp c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/data/aotinductor/model/c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper_metadata.json c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/data/aotinductor/model/ctmp7adn3spwyscdotllyj4yx3vrqcnxk3thkpgdcax7zvqmyyp3.kernel_metadata.json c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/data/aotinductor/model/c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper.so c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/archive_format c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/archive_version c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/.data/version c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/byteorder c7l7jkswdq7ud6gpvpmunx76hi3c357l7epyc7oofeemzeoy7euo.wrapper/.data/serialization_id ``` Test Plan: ``` buck2 run @//mode/dev-nosan //caffe2/test/inductor:aot_inductor_package -- -r "test_loading_wrong_model" ``` Rollback Plan: Differential Revision: D77320485 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156863 Approved by: https://github.com/tugsbayasgalan	2025-06-26 17:13:29 +00:00
PyTorch MergeBot	21990fbad9	Revert "[cond] support gen_schema for cond (#154193 )" This reverts commit 6de41ce0f899604c3f8b33e1f8d37eb89b3a963e. Reverted https://github.com/pytorch/pytorch/pull/154193 on behalf of https://github.com/Camyll due to issue landing internally, discussed with Yidi offline ([comment](https://github.com/pytorch/pytorch/pull/154193#issuecomment-3009160081))	2025-06-26 17:10:00 +00:00
Isuru Fernando	c808af514d	Support deterministic upsample trilinear backward (#154239 ) Fixes https://github.com/pytorch/pytorch/issues/154183 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154239 Approved by: https://github.com/eellison, https://github.com/albanD	2025-06-26 15:02:27 +00:00
IvanKobzarev	2f94f69b7c	[aotd] Support mutations of the same input in fw and bw (#155354 ) Original issue: https://github.com/pytorch/pytorch/issues/154820 The issue happens when there is a mutation for the same input in forward AND in backward. AOTD emited copy_ after joint_function tracing. This made this fx-node to correspond to the side effects of both mutations (in forward and in backward). After that partitioner can put it either in forward or in backward. The fix: 1/ Introduce joint_function.handle that allows to set "post_forward" callback, to be able to check inputs state after forward We do not want to apply the mutation after joint, if we already applied it in forward. For that we need "mutation_counter" and memorize the version of mutation that we applied for forward mutation. 2/ Exposing mutation_counter to python We want to keep invariant that copy_ exist only in the end of joint graph. 3/ We memorize mutation_counter and state of the inputs after forward, using the handle post_forward. Emit post_forward mutations after joint graph fully traced. add for post_forward mutations "must_be_in_forward" tag (similar to existing "must_be_in_backward") to keep them in forward. 4/ Ban recompute of the source of mutation. Recompute can apply the same op (e.g. add) in forward and backward. For this set MUST_SAVE for the source of mutation in forward. proxy_tensor changes: By default proxy tensor updates tensor_tracker. In this case applied mutations will be chained. But we want that this copy_ will be independent and applied just to primals. For this introducing a contextmanager to be able to disable update of tensor_tracker for adding forward mutations. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155354 Approved by: https://github.com/bdhirsh	2025-06-26 14:05:54 +00:00
Anatoly Myachev	197c1869f5	[Inductor][CLN] Remove unused default configs in `flex_attention.py` (#156700 ) They probably became unusable after `03023f178c` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156700 Approved by: https://github.com/jataylo, https://github.com/drisspg	2025-06-26 13:24:09 +00:00
rzou	2d39a48d52	Fix silent incorrectness arising from incorrect alias information (#152011 ) Fixes #136662 There are two problems: 1) canonicalize_view_scatter_ops adds some new nodes into the graph. These new nodes cause the alias info on the graph to be wrong. To fix this, we try to run FakeTensorUpdater on the graph again. 2) FakeTensorUpdater's alias information is wrong. It tries to skip nodes that it thinks have "equivalent" FakeTensor metadata. It should not be allowed to do this if any users of the node can alias the node. The example is if we have `x = foo(...); y = x.view(...)`. If the user replaces `foo` with a new `bar` node and sets bar.meta["val"] correctly, then FakeTensorUpdater still needs to update y's meta["val"] to be a view of the new bar node. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152011 Approved by: https://github.com/yf225	2025-06-26 13:05:08 +00:00
haozhe.zhu	53e0b9c393	refine fp32 precision api (#125888 ) Based on the [conversation](https://github.com/pytorch/pytorch/issues/121791), we plan to drop the "highest, high, medium" to represent fp32 internal computation data types . Instead, we will directly use the algorithm to represent it. ### Design Choice: Directly use algorithms name like "TF32", "BF16". #### Pros - The names are more informative. 'tf32' is more informative than a simple "high". - Easier to extend new algorithm like `tf32x3` #### Cons - "HIGHEST, HIGH, MEDIUM" indicated the relative precision between different algorithms. However, we can have more documents to discuss them. ### We provide a layered structure for backends/operators. ('f32' is short for 'fp32_precision') ![image](https://github.com/user-attachments/assets/f89143e5-d6a1-4865-9351-9a50439f5067) ### We provide 3 fp32 compute precision can be set: - "ieee": Not allowed to use any other internal computation data types . - "tf32": Allowed to use tf32 as internal computation data types. - "bf16": Allowed to use bf16 as internal computation data types. - "none": Precision's are not set. Can be override by its father node. ### Overriding Precision Settings Child node can be override by its father node if it is set to default. For current default settings: ``` backend = generic, op = all, precision setting = none backend = cuda, op = all, precision setting = none backend = cuda, op = conv, precision setting = tf32 backend = cuda, op = rnn, precision setting = tf32 backend = cuda, op = matmul, precision setting = none backend = matmul, op = all, precision setting = none backend = matmul, op = conv, precision setting = none backend = matmul, op = rnn, precision setting = none backend = matmul, op = matmul, precision setting = none ``` - If the user set `torch.backends.mkldnn.fp32_precision="bf16"`, his child nodes `torch.backends.mkldnn.matmul.fp32_precision` / `torch.backends.mkldnn.conv.fp32_precision` / `torch.backends.mkldnn.rnn.fp32_precision` will also be override to "bf16". - If the user set `torch.backends.fp32_precision="bf16"`, `torch.backends.mkldnn.fp32_precision` and his child nodes will also we override to "bf16". ### Backward Compatible Since new API allow user to have more fine-grained control. There will be some conflict. For example, previous `torch.backends.cudnn.allow_tf32` are not enough to represent the status for `torch.backends.cudnn.rnn.fp32_precision="ieee"` and `torch.backends.cudnn.conv.fp32_precision="tf32"`. Therefore, our goal for backward compatible is - If the user only uses previous APIs, it will work as previous expectations. - If the user use new API to change the status to an un-representable status for old API, and try to access the status by old API. We will raise Runtime Error and point the document for user. ### Test Plan ``` python test/test_cuda.py -k test_fp32_precision_with_tf32 python test/test_cuda.py -k test_fp32_precision_with_float32_matmul_precision python test/test_cuda.py -k test_invalid_status_for_legacy_api python test/test_mkldnn.py -k test_mlkdnn_get_set python test/test_mkldnn.py -k test_generic_precision python test/test_mkldnn.py -k test_invalid python test/test_mkldnn.py -k test_default_use_parent ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/125888 Approved by: https://github.com/jgong5, https://github.com/albanD Co-authored-by: Jiang, Yanbing <yanbing.jiang@intel.com>	2025-06-26 10:32:20 +00:00
Ting Lu	de45c5f673	[aarch64] Add back NCCL lib to cuda arm wheel (#156888 ) We discovered that when importing latest 12.9 arm nightly wheel, it is missing the NCCL lib. With the use of USE_SYSTEM_NCCL=1, we need to copy the libnccl.so lib into our big wheel environment, so that it can be dynamically linked at runtime. https://github.com/pytorch/pytorch/pull/152835 enabled USE_SYSTEM_NCCL=1, which would use the system NCCL by default, and it would no longer use the one built from libtorch_cuda.so. With this PR, we add back the libnccl.so to be used at runtime. In this way, we also provide the flexibility to use different versions of NCCL from what came with the original pytorch build. related - https://github.com/pytorch/pytorch/issues/144768 ``` Python 3.12.3 (main, Jun 18 2025, 17:59:45) [GCC 13.3.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> import torch Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/usr/local/lib/python3.12/dist-packages/torch/__init__.py", line 417, in <module> from torch._C import * # noqa: F403 ^^^^^^^^^^^^^^^^^^^^^^ ImportError: libnccl.so.2: cannot open shared object file: No such file or directory ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156888 Approved by: https://github.com/atalman	2025-06-26 10:24:18 +00:00
Mark Saroufim	18b01afa9e	load inline user overridable gencode (#156850 ) Fixes https://github.com/pytorch/pytorch/issues/156815 As far as testing goes * I tried to use cuobjdump but that was kinda goofy `bccd9393a5` the problem was that the name of the cubin will have a single gencode always * Another idea was to read stderr and check that the right amount of gencodes is there `0beadc01b3` this helped a lot to convince me locally that this test works, the test passed on my dev gpu but was failing in CI and I suspect it's because of a bad interaction with subprocesses * Last approach was to have a simpler unit test to check which flags get added by default, this is not as comprehensive as the previous ideas but it works and is fast so will opt for this since I'm convinced testing is working per my own experiments and customers Pull Request resolved: https://github.com/pytorch/pytorch/pull/156850 Approved by: https://github.com/malfet	2025-06-26 10:15:08 +00:00
Klaus Zimmermann	bbf1a6feac	Add dist_info to non-building setup.py commands (#156709 ) This adds the `dist_info` command to the list of non-building commands of `setup.py`, which avoids the current situation where simple metadata generation with any packaging tool already triggers a build. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156709 Approved by: https://github.com/Skylion007	2025-06-26 08:38:39 +00:00
Xuehai Pan	455dfd2589	Fix macOS build with `USE_MPS=OFF` (#156847 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156847 Approved by: https://github.com/angelayi	2025-06-26 07:15:41 +00:00
Jane Xu	50b2069b61	Move out super large one off foreach_copy test (#156876 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156876 Approved by: https://github.com/albanD, https://github.com/jeffdaily	2025-06-26 06:02:38 +00:00
Nicolas Macchioni	dfc31b3345	[BE] comments + try to get rid of secondary `make_autotune_fn` (#156358 ) Not sure this will work, but let's try it on the unit tests. The only thing I am worried about is the counters drifting off from their true values, so let the unit tests check that Pull Request resolved: https://github.com/pytorch/pytorch/pull/156358 Approved by: https://github.com/masnesral	2025-06-26 05:54:01 +00:00
Laith Sakka	0d01bafc34	remove gso from set_storage_meta__symint (#156525 ) We already check that inputs are hinted? i dont see value here for it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156525 Approved by: https://github.com/pianpwk	2025-06-26 05:42:05 +00:00
Eli Uriegas	127695eb5c	ci: Add ciflow trigger for build-triton-wheel (#156893 ) Signed-off-by: Eli Uriegas <eliuriegas@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156893 Approved by: https://github.com/malfet	2025-06-26 04:38:38 +00:00
FFFrog	0a16818d5b	[OpenReg] Remove the unit.skip for test_serialization (#156804 ) This bugs was fixed by this [PR](https://github.com/pytorch/pytorch/pull/147095) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156804 Approved by: https://github.com/albanD ghstack dependencies: #156588, #156589	2025-06-26 03:59:50 +00:00
FFFrog	98e594b565	[OpenReg][2/N] Migrate cpp_extensions_open_device_registration to OpenReg (#156589 ) ---- - serialization - dlpack Next Steps: - The rest of `test/test_cpp_extensions_open_device_registration.py` is about the fallback mechanism. In order to keep it consistent with other accelerator usage (C++ registration), the implementation of OpenReg needs to be refactored: * Simulate multiple device memory in a single process (a brief RFC will be submitted this week) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156589 Approved by: https://github.com/albanD ghstack dependencies: #156588	2025-06-26 03:59:50 +00:00
FFFrog	a730c65fe3	[OpenReg][1/N] Migrate cpp_extensions_open_device_registration to OpenReg (#156588 ) ---- - fake tensor - named tensor - custom autograd function Pull Request resolved: https://github.com/pytorch/pytorch/pull/156588 Approved by: https://github.com/albanD	2025-06-26 03:59:50 +00:00
fduwjj	4585c33e74	[symm_mem] Fix nccl test for symm mem (#156752 ) Try not to call set_device to Fixes #156569 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156752 Approved by: https://github.com/kwen2501	2025-06-26 02:59:38 +00:00
Edward Yang	7521cd9111	[BE] Typo fix (#156836 ) Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156836 Approved by: https://github.com/albanD, https://github.com/jingsh, https://github.com/Skylion007 ghstack dependencies: #156830, #156831	2025-06-26 02:48:55 +00:00
Edward Yang	68e023cbbb	[BE] Add missing type for storage dict (#156831 ) For some reason, this one always bleats when I run mypy on OSX, so shut it up. Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156831 Approved by: https://github.com/mikaylagawarecki, https://github.com/atalman, https://github.com/malfet ghstack dependencies: #156830	2025-06-26 02:48:55 +00:00
Edward Yang	df9e5a276b	[BE] Add type and docs for _process_export_inputs (#156830 ) Done using claude code and manual review. Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156830 Approved by: https://github.com/tugsbayasgalan, https://github.com/malfet	2025-06-26 02:48:55 +00:00
henrylhtsang	81bf278537	[cutlass] rename cutlass python lib to python-cutlass (#156655 ) Differential Revision: [D77173366](https://our.internmc.facebook.com/intern/diff/D77173366/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156655 Approved by: https://github.com/Skylion007	2025-06-26 02:47:14 +00:00
bobrenjc93	8da774d81f	[ez] Add docblock for SchedulerNode.codegen (#156718 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156718 Approved by: https://github.com/BoyuanFeng ghstack dependencies: #156466, #156445, #156625, #156717	2025-06-26 02:43:50 +00:00
Catherine Lee	78ee2ee90e	Fix environment and push env var for docker image builds for binary builds (#156910 ) Changes WITH_PUSH and the environment check to be ok with giving credentials to push to docker io if its on the main branch, a tag starting with v, or the release branch Credentials for pushing to docker io are in the environment, so without the environment, you can't push to docker io. You also don't do the push unless WITH_PUSH is true binary builds on release branch were failing because they pull from docker io, but the docker build wasn't pushing to docker io because it was either on the release branch (didn't have credentials https://github.com/pytorch/pytorch/actions/runs/15888166271/job/44813180986) or it was on the tag (doesn't have WITH_PUSH) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156910 Approved by: https://github.com/atalman	2025-06-26 02:06:57 +00:00
atalman	5db9a2b54a	[BE] Install Helion without dependencies (#156706 ) After: https://github.com/pytorch/pytorch/pull/155513 Please see comment: https://github.com/pytorch/pytorch/pull/155513#issuecomment-2998085740 Here are the logs: https://github.com/pytorch/pytorch/actions/runs/15838529400/job/44646874281?pr=156664#step:6:16372 Looks like current workflow is : Build triton - triton-3.4.0+git5389ed79-cp310-cp310-linux_x86_64.whl Install Helion - Overwrite triton with production 3.3.1 and install production torch Reinstall triton as final docker build step - triton-3.4.0+git5389ed79-cp310-cp310-linux_x86_64.whl This makes it somewhat messy since we install both torch and triton from prod. This is something we want to avoid when building underlining docker images for CI Log: ``` #55 311.4 + pip_install helion #55 311.4 + as_jenkins conda run -n py_3.10 pip install --progress-bar off helion #55 311.4 + sudo -E -H -u jenkins env -u SUDO_UID -u SUDO_GID -u SUDO_COMMAND -u SUDO_USER env PATH=/usr/local/nvidia/bin:/usr/local/cuda/bin:/opt/conda/envs/py_3.10/bin:/opt/conda/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin LD_LIBRARY_PATH= conda run -n py_3.10 pip install --progress-bar off helion #55 393.6 Collecting helion #55 393.6 Downloading helion-0.0.7-py3-none-any.whl.metadata (14 kB) #55 393.6 Collecting filecheck (from helion) #55 393.6 Downloading filecheck-1.0.2-py3-none-any.whl.metadata (5.8 kB) #55 393.6 Collecting torch>=2.7.0 (from helion) #55 393.6 Downloading torch-2.7.1-cp310-cp310-manylinux_2_28_x86_64.whl.metadata (29 kB) #55 393.6 Requirement already satisfied: typing-extensions>=4.0.0 in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from helion) (4.14.0) #55 393.6 Requirement already satisfied: filelock in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from torch>=2.7.0->helion) (3.18.0) #55 393.6 Requirement already satisfied: sympy>=1.13.3 in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from torch>=2.7.0->helion) (1.13.3) #55 393.6 Requirement already satisfied: networkx in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from torch>=2.7.0->helion) (2.8.8) #55 393.6 Requirement already satisfied: jinja2 in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from torch>=2.7.0->helion) (3.1.6) #55 393.6 Requirement already satisfied: fsspec in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from torch>=2.7.0->helion) (2025.5.1) #55 393.6 Collecting nvidia-cuda-nvrtc-cu12==12.6.77 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cuda_nvrtc_cu12-12.6.77-py3-none-manylinux2014_x86_64.whl.metadata (1.5 kB) #55 393.6 Collecting nvidia-cuda-runtime-cu12==12.6.77 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cuda_runtime_cu12-12.6.77-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.5 kB) #55 393.6 Collecting nvidia-cuda-cupti-cu12==12.6.80 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cuda_cupti_cu12-12.6.80-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.6 kB) #55 393.6 Collecting nvidia-cudnn-cu12==9.5.1.17 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cudnn_cu12-9.5.1.17-py3-none-manylinux_2_28_x86_64.whl.metadata (1.6 kB) #55 393.6 Collecting nvidia-cublas-cu12==12.6.4.1 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cublas_cu12-12.6.4.1-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.5 kB) #55 393.6 Collecting nvidia-cufft-cu12==11.3.0.4 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cufft_cu12-11.3.0.4-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.5 kB) #55 393.6 Collecting nvidia-curand-cu12==10.3.7.77 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_curand_cu12-10.3.7.77-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.5 kB) #55 393.6 Collecting nvidia-cusolver-cu12==11.7.1.2 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cusolver_cu12-11.7.1.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.6 kB) #55 393.6 Collecting nvidia-cusparse-cu12==12.5.4.2 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cusparse_cu12-12.5.4.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.6 kB) #55 393.6 Collecting nvidia-cusparselt-cu12==0.6.3 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cusparselt_cu12-0.6.3-py3-none-manylinux2014_x86_64.whl.metadata (6.8 kB) #55 393.6 Collecting nvidia-nccl-cu12==2.26.2 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_nccl_cu12-2.26.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (2.0 kB) #55 393.6 Collecting nvidia-nvtx-cu12==12.6.77 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_nvtx_cu12-12.6.77-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.6 kB) #55 393.6 Collecting nvidia-nvjitlink-cu12==12.6.85 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_nvjitlink_cu12-12.6.85-py3-none-manylinux2010_x86_64.manylinux_2_12_x86_64.whl.metadata (1.5 kB) #55 393.6 Collecting nvidia-cufile-cu12==1.11.1.6 (from torch>=2.7.0->helion) #55 393.6 Downloading nvidia_cufile_cu12-1.11.1.6-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl.metadata (1.5 kB) #55 393.6 Collecting triton==3.3.1 (from torch>=2.7.0->helion) #55 393.6 Downloading triton-3.3.1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.metadata (1.5 kB) #55 393.6 Requirement already satisfied: setuptools>=40.8.0 in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from triton==3.3.1->torch>=2.7.0->helion) (80.9.0) #55 393.6 Requirement already satisfied: mpmath<1.4,>=1.1.0 in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from sympy>=1.13.3->torch>=2.7.0->helion) (1.3.0) #55 393.6 Requirement already satisfied: MarkupSafe>=2.0 in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from jinja2->torch>=2.7.0->helion) (3.0.2) #55 393.6 Downloading helion-0.0.7-py3-none-any.whl (149 kB) #55 393.6 Downloading torch-2.7.1-cp310-cp310-manylinux_2_28_x86_64.whl (821.2 MB) #55 393.6 Downloading nvidia_cublas_cu12-12.6.4.1-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (393.1 MB) #55 393.6 Downloading nvidia_cuda_cupti_cu12-12.6.80-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (8.9 MB) #55 393.6 Downloading nvidia_cuda_nvrtc_cu12-12.6.77-py3-none-manylinux2014_x86_64.whl (23.7 MB) #55 393.6 Downloading nvidia_cuda_runtime_cu12-12.6.77-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (897 kB) #55 393.6 Downloading nvidia_cudnn_cu12-9.5.1.17-py3-none-manylinux_2_28_x86_64.whl (571.0 MB) #55 393.6 Downloading nvidia_cufft_cu12-11.3.0.4-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (200.2 MB) #55 393.6 Downloading nvidia_cufile_cu12-1.11.1.6-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (1.1 MB) #55 393.6 Downloading nvidia_curand_cu12-10.3.7.77-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (56.3 MB) #55 393.6 Downloading nvidia_cusolver_cu12-11.7.1.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (158.2 MB) #55 393.6 Downloading nvidia_cusparse_cu12-12.5.4.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (216.6 MB) #55 393.6 Downloading nvidia_cusparselt_cu12-0.6.3-py3-none-manylinux2014_x86_64.whl (156.8 MB) #55 393.6 Downloading nvidia_nccl_cu12-2.26.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (201.3 MB) #55 393.6 Downloading nvidia_nvjitlink_cu12-12.6.85-py3-none-manylinux2010_x86_64.manylinux_2_12_x86_64.whl (19.7 MB) #55 393.6 Downloading nvidia_nvtx_cu12-12.6.77-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl (89 kB) #55 393.6 Downloading triton-3.3.1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (155.6 MB) #55 393.6 Downloading filecheck-1.0.2-py3-none-any.whl (23 kB) #55 393.6 Installing collected packages: nvidia-cusparselt-cu12, triton, nvidia-nvtx-cu12, nvidia-nvjitlink-cu12, nvidia-nccl-cu12, nvidia-curand-cu12, nvidia-cufile-cu12, nvidia-cuda-runtime-cu12, nvidia-cuda-nvrtc-cu12, nvidia-cuda-cupti-cu12, nvidia-cublas-cu12, filecheck, nvidia-cusparse-cu12, nvidia-cufft-cu12, nvidia-cudnn-cu12, nvidia-cusolver-cu12, torch, helion #55 393.6 Attempting uninstall: triton #55 393.6 Found existing installation: triton 3.4.0+git5389ed79 #55 393.6 Uninstalling triton-3.4.0+git5389ed79: #55 393.6 Successfully uninstalled triton-3.4.0+git5389ed79 #55 393.6 Successfully installed filecheck-1.0.2 helion-0.0.7 nvidia-cublas-cu12-12.6.4.1 nvidia-cuda-cupti-cu12-12.6.80 nvidia-cuda-nvrtc-cu12-12.6.77 nvidia-cuda-runtime-cu12-12.6.77 nvidia-cudnn-cu12-9.5.1.17 nvidia-cufft-cu12-11.3.0.4 nvidia-cufile-cu12-1.11.1.6 nvidia-curand-cu12-10.3.7.77 nvidia-cusolver-cu12-11.7.1.2 nvidia-cusparse-cu12-12.5.4.2 nvidia-cusparselt-cu12-0.6.3 nvidia-nccl-cu12-2.26.2 nvidia-nvjitlink-cu12-12.6.85 nvidia-nvtx-cu12-12.6.77 torch-2.7.1 triton-3.3.1 #55 393.6 #55 DONE 428.8s #56 [final 1/30] COPY --from=triton-builder /opt/triton /opt/triton #56 DONE 0.0s #57 [final 2/30] RUN if [ -n "yes" ] \|\| [ -n "" ]; then pip install /opt/triton/*.whl; chown -R jenkins:jenkins /opt/conda; fi #57 0.823 Processing /opt/triton/triton-3.4.0+git5389ed79-cp310-cp310-linux_x86_64.whl #57 2.263 Requirement already satisfied: setuptools>=40.8.0 in /opt/conda/envs/py_3.10/lib/python3.10/site-packages (from triton==3.4.0+git5389ed79) (80.9.0) #57 2.589 Installing collected packages: triton #57 6.405 Successfully installed triton-3.4.0+git5389ed79 #57 6.405 WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning. #57 DONE 86.5s ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156706 Approved by: https://github.com/oulgen, https://github.com/malfet	2025-06-26 02:05:47 +00:00
fduwjj	b50075343a	[distributed] Enable H100 test for all distributed related changes (#156721 ) We want to run H100 CI for distributed related changes. We already have a labeling of oncall:distributed when touching distributed related code: `4491326fb0/.github/labeler.yml (L94)`. So we want to leverage that. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156721 Approved by: https://github.com/huydhn	2025-06-26 01:51:41 +00:00
James Wu	e581f015ee	Bump STATIC_CUDA_LAUNCHER_VERSION to 2 (#156726 ) Differential Revision: [D77241813](https://our.internmc.facebook.com/intern/diff/D77241813) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156726 Approved by: https://github.com/oulgen	2025-06-26 01:50:51 +00:00
Xia, Weiwen	b5bfbba184	[Quant][CPU] fix fake_quantize_per_tensor_affine of inf values (#155109 ) Fixes #154328 Summary Fail reason: The input value is infinity in float and it has undefined behavior to convert it to int64_t. On X86, it will be converted to the min value of int64_t, which is not expected. Fix: Clamping `(input * inv_scale + zero_point)` to `[quant_min, quant_max]` before converting it to int64_t. Test plan ``` pytest test/quantization/core/test_workflow_ops.py -k test_fake_quantize_per_tensor_affine_inf ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155109 Approved by: https://github.com/leslie-fang-intel, https://github.com/jerryzh168	2025-06-26 01:24:36 +00:00
Nikita Shulga	214e2959dc	Cleanup leftover miniconda brew installation (#156898 ) That results in torch.compile being unable to produce working artifacts Should fix https://github.com/pytorch/pytorch/issues/156833 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156898 Approved by: https://github.com/seemethere, https://github.com/atalman	2025-06-26 01:02:04 +00:00
Laith Sakka	4c0091fda6	python definitely_contiguous-> is_contiguous_or_false (#156515 ) We probably can avoid having those in python as well and just depend on c++ impl after we land https://github.com/pytorch/pytorch/pull/155590 but that is for a different PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156515 Approved by: https://github.com/bobrenjc93	2025-06-26 00:47:14 +00:00
Laith Sakka	85df746892	refresh expected numbers (#156877 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156877 Approved by: https://github.com/huydhn	2025-06-26 00:03:09 +00:00
Mikayla Gawarecki	2c6324a1eb	Delete sections referencing torchscript in serialization docs (#156648 ) Address [T228333890](https://www.internalfb.com/intern/tasks/?t=228333890) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156648 Approved by: https://github.com/svekars	2025-06-25 23:41:24 +00:00
albanD	a25d1443fa	Mark TorchServe as all emeritus (#156865 ) As per title and to follow the broader tutorial cleanup work. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156865 Approved by: https://github.com/svekars, https://github.com/malfet, https://github.com/seemethere	2025-06-25 23:34:57 +00:00
bobrenjc93	451b525bf0	[ez] add docblock and comments to simd.split_and_set_ranges (#156717 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156717 Approved by: https://github.com/BoyuanFeng ghstack dependencies: #156445	2025-06-25 23:07:28 +00:00
Shangdi Yu	204db27a0c	Consolidate stack trace in Tracer (#156257 ) Summary: - Consolidate the stack trace recording code in TracerBase and PythonKeyTracer - Change `make_fx`'s arg name to be consistent with TracerBase member name `record_stack_traces` We move the stack trace logic from `create_proxy` to `create_node` so all inherited classes of TracerBase and re-use the same stack trace logic. Test Plan: ``` buck run caffe2/test:test_export -- -r test_stack_trace ``` Rollback Plan: Pull Request resolved: https://github.com/pytorch/pytorch/pull/156257 Approved by: https://github.com/angelayi, https://github.com/zou3519	2025-06-25 23:07:10 +00:00
Isalia20	653c52fe52	[MPS] Fix batch norm incorrect gradient (#156867 ) Fixes #156555 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156867 Approved by: https://github.com/malfet	2025-06-25 23:05:49 +00:00
Jane Xu	acaf6ba3c6	Organize BUCK for torch/standalone (#156503 ) Summary: Undo highlevel BUCKification in favor of something more organized by moving it to the dir itself Test Plan: CI Rollback Plan: Reviewed By: swolchok Differential Revision: D76920013 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156503 Approved by: https://github.com/swolchok	2025-06-25 22:56:15 +00:00
Dylan Maloy	d98fa4a103	implement SR's storage group planning algorithm (#156715 ) Summary: att Test Plan: tested on a localnet. it's ~15% worse performance than greedy-by-size, but more performant. local: gbs: 110656b dsg: 131584b local_ro: gbs: 38208 dsg: 44544 Differential Revision: D75653840 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156715 Approved by: https://github.com/zhxchen17	2025-06-25 22:43:40 +00:00
Laith Sakka	1e7e21ec5d	unify dynamic shapes API namings 3 (guard_int, guard_int_seq) (#155973 ) evaluate_static_shape -> guard_int evaluate_static_shapes -> guard_int_seq Pull Request resolved: https://github.com/pytorch/pytorch/pull/155973 Approved by: https://github.com/bobrenjc93	2025-06-25 22:40:28 +00:00
Yidi Wu	61f6aa36b9	[resubmit][export] add _union_dataclass to support comparing dataclasses that inherits from union. (#156765 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156765 Approved by: https://github.com/zhxchen17	2025-06-25 22:32:12 +00:00
William Wen	53057fc16a	[dynamo] update base variable call_method hint with note on comprehensions (#156769 ) Internal xref: https://fb.workplace.com/groups/1075192433118967/permalink/1696822194289318/ List/dict comprehensions in Python <= 3.11 result in potentially weird graph breaking behavior because comprehensions result in implicit function calls, which Dynamo may end up tracing as top-level frames, resulting in iterators being passed as arguments to the compiled region. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156769 Approved by: https://github.com/StrongerXi	2025-06-25 21:55:55 +00:00
dolpm	95a7d1912a	[sigmoid] add layout planner to executor (#156852 ) Summary: if memory planning is enabled in the runtime config, we will create a copy in the executor here. Test Plan: ci Differential Revision: D73635622 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156852 Approved by: https://github.com/zhxchen17	2025-06-25 21:41:09 +00:00
Yidi Wu	6de41ce0f8	[cond] support gen_schema for cond (#154193 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154193 Approved by: https://github.com/zou3519 ghstack dependencies: #155644	2025-06-25 21:19:58 +00:00
Yidi Wu	3257c8f74c	[cond] preserve merged phs meta for subgraph (#155644 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155644 Approved by: https://github.com/zou3519	2025-06-25 21:19:58 +00:00
James Wu	e7a66166ce	[precompile] When using BundledAOTAutogradCache, disable FXGraphCache (#156611 ) The goal of this PR is to fix a specific bug when turning precompile on/off between caching runs. If you try to turn on BundledAOTAutogradCacheEntry today in between local runs, the FXGraphCache may randomly hit between the two runs, because FXGraphCache knows nothing about AOTAutogradCache's config. When FXGraphCache hits, it immediately will call make_launchers() immediately on the triton code it launches, which then causes an assertion failure because pickle should not be called after make_launchers. One way to resolve the bug is just to add whether precompile is enabled to teh FxGraph cache key. But the better fix for this, however, is higher level/philosophical: When using BundledAOTAutogradCacheEntry, the entire CompiledFxGraph is saved directly to the cache entry, and we expect the two caches to work in sync, i.e. as one cache. So to simplify the programming model, we disable FxGraphCache when BundledAOTAUtogradCache is turned on. BundledAOTAutogradCacheEntry is only used for precompile use cases now; if we wanted to use BundledAOTAutogradCache for traditional caching use cases, there's a bunch of further work, one of which would be to re-enable FxGraphCache in the event that BundledAOTAutogradCache has to bypass. However, for precompile, this is not a scenario that should happen: we should always expect the entire callable to be saveable, and we should expect to never bypass. So we don't do that change for now. Added a unit test demonstrating this behavior. Also updated existing unit tests to show that all fx graph cache operations are now 0 (but all tests still pass). Pull Request resolved: https://github.com/pytorch/pytorch/pull/156611 Approved by: https://github.com/zhxchen17	2025-06-25 21:01:42 +00:00
Dmitry Nikolaev	fe1f1a38df	add test_batchnorn_2D and 3D tests (#156498 ) New set of batchnorm tests to verify NCHW 2D/3D BatchNorm This test also allows to add and configure different BatchNorm tests (dtypes, NCHW/NHWC, Mixed) in the future based on: - Train [test_batchnorm_cudnn_nhwc](`1051b93192/test/test_nn.py (L4985)`) - Inference [test_batchnorm_nhwc_cuda](`1051b93192/test/test_nn.py (L5130)`) ``` test_batchnorm_3D_inference_NCHW_vs_cpu_float32 (__main__.TestNN.test_batchnorm_3D_inference_NCHW_vs_cpu_float32) ... ok (0.113s) test_batchnorm_3D_inference_NCHW_vs_cpu_mixed_bfloat16 (__main__.TestNN.test_batchnorm_3D_inference_NCHW_vs_cpu_mixed_bfloat16) ... ok (0.057s) test_batchnorm_3D_inference_NCHW_vs_cpu_mixed_float16 (__main__.TestNN.test_batchnorm_3D_inference_NCHW_vs_cpu_mixed_float16) ... ok (0.063s) test_batchnorm_3D_inference_NCHW_vs_native_float32 (__main__.TestNN.test_batchnorm_3D_inference_NCHW_vs_native_float32) ... ok (0.059s) test_batchnorm_3D_inference_NCHW_vs_native_mixed_bfloat16 (__main__.TestNN.test_batchnorm_3D_inference_NCHW_vs_native_mixed_bfloat16) ... ok (0.006s) test_batchnorm_3D_inference_NCHW_vs_native_mixed_float16 (__main__.TestNN.test_batchnorm_3D_inference_NCHW_vs_native_mixed_float16) ... ok (0.006s) test_batchnorm_3D_train_NCHW_vs_cpu_float32 (__main__.TestNN.test_batchnorm_3D_train_NCHW_vs_cpu_float32) ... ok (0.007s) test_batchnorm_3D_train_NCHW_vs_cpu_mixed_bfloat16 (__main__.TestNN.test_batchnorm_3D_train_NCHW_vs_cpu_mixed_bfloat16) ... ok (0.005s) test_batchnorm_3D_train_NCHW_vs_cpu_mixed_float16 (__main__.TestNN.test_batchnorm_3D_train_NCHW_vs_cpu_mixed_float16) ... ok (0.005s) test_batchnorm_3D_train_NCHW_vs_native_float32 (__main__.TestNN.test_batchnorm_3D_train_NCHW_vs_native_float32) ... ok (0.003s) test_batchnorm_3D_train_NCHW_vs_native_mixed_bfloat16 (__main__.TestNN.test_batchnorm_3D_train_NCHW_vs_native_mixed_bfloat16) ... skip: bfloat16 NCHW train failed due to native tolerance issue (0.001s) test_batchnorm_3D_train_NCHW_vs_native_mixed_float16 (__main__.TestNN.test_batchnorm_3D_train_NCHW_vs_native_mixed_float16) ... skip: 3D float16 NCHW train failed on ROCm<7.0 (0.001s) test_batchnorm_2D_inference_NCHW_vs_cpu_float32 (__main__.TestNN.test_batchnorm_2D_inference_NCHW_vs_cpu_float32) ... ok (0.016s) test_batchnorm_2D_inference_NCHW_vs_cpu_mixed_bfloat16 (__main__.TestNN.test_batchnorm_2D_inference_NCHW_vs_cpu_mixed_bfloat16) ... ok (0.003s) test_batchnorm_2D_inference_NCHW_vs_cpu_mixed_float16 (__main__.TestNN.test_batchnorm_2D_inference_NCHW_vs_cpu_mixed_float16) ... ok (0.003s) test_batchnorm_2D_inference_NCHW_vs_native_float32 (__main__.TestNN.test_batchnorm_2D_inference_NCHW_vs_native_float32) ... ok (0.054s) test_batchnorm_2D_inference_NCHW_vs_native_mixed_bfloat16 (__main__.TestNN.test_batchnorm_2D_inference_NCHW_vs_native_mixed_bfloat16) ... ok (0.002s) test_batchnorm_2D_inference_NCHW_vs_native_mixed_float16 (__main__.TestNN.test_batchnorm_2D_inference_NCHW_vs_native_mixed_float16) ... ok (0.001s) test_batchnorm_2D_train_NCHW_vs_cpu_float32 (__main__.TestNN.test_batchnorm_2D_train_NCHW_vs_cpu_float32) ... ok (0.007s) test_batchnorm_2D_train_NCHW_vs_cpu_mixed_bfloat16 (__main__.TestNN.test_batchnorm_2D_train_NCHW_vs_cpu_mixed_bfloat16) ... ok (0.004s) test_batchnorm_2D_train_NCHW_vs_cpu_mixed_float16 (__main__.TestNN.test_batchnorm_2D_train_NCHW_vs_cpu_mixed_float16) ... ok (0.004s) test_batchnorm_2D_train_NCHW_vs_native_float32 (__main__.TestNN.test_batchnorm_2D_train_NCHW_vs_native_float32) ... ok (0.003s) test_batchnorm_2D_train_NCHW_vs_native_mixed_bfloat16 (__main__.TestNN.test_batchnorm_2D_train_NCHW_vs_native_mixed_bfloat16) ... skip: bfloat16 NCHW train failed due to native tolerance issue (0.001s) test_batchnorm_2D_train_NCHW_vs_native_mixed_float16 (__main__.TestNN.test_batchnorm_2D_train_NCHW_vs_native_mixed_float16) ... ok (0.002s) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156498 Approved by: https://github.com/jeffdaily	2025-06-25 20:38:02 +00:00
angelayi	48e7b62d3a	[dynamo] Add immutable pytree to trace_rules (#156772 ) Fixes https://github.com/pytorch/pytorch/issues/155426 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156772 Approved by: https://github.com/williamwen42	2025-06-25 20:08:47 +00:00
Pavan Balaji	e99a2a2dba	[PG/nccl] Simplify uniqueHash management (#156790 ) Summary: ncclUniqueID is only relevant when a comm is created using ncclCommCreate or ncclCommCreateConfig. If a comm is created with ncclCommSplit, this field is unset, causing its usage to create unexpected behavior. This patch creates a unique hash key for each comm, irrespective of how the comm is created. Test Plan: CI Reviewers: Subscribers: Tasks: Tags: Pull Request resolved: https://github.com/pytorch/pytorch/pull/156790 Approved by: https://github.com/fduwjj, https://github.com/kwen2501	2025-06-25 20:06:08 +00:00
James Wu	070aa59e49	Refactor DynamoStore into disk and in memory implementations (#155818 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155818 Approved by: https://github.com/zhxchen17	2025-06-25 18:24:28 +00:00
Yulu Jia	6c24c6633a	[torch][test] skip `test_transformer_backend_inductor_fullgraph_True` (#156763 ) Summary: "Traceable FSDP2" is not being maintained anymore. Test Plan: ``` buck test @//mode/opt caffe2/test/distributed/_composable:fully_shard_compile -- test_transformer_backend_inductor_fullgraph_True ``` https://www.internalfb.com/intern/testinfra/testconsole/testrun/16044073764394232/ Rollback Plan: Differential Revision: D77264408 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156763 Approved by: https://github.com/xunnanxu, https://github.com/yf225	2025-06-25 18:15:23 +00:00
PaliC	09ffba3cf7	[docs] Decorator to create a deprecation warning (#155127 ) This PR adds the `@deprecate` decorator for internal functions which we are prepping for deprecation. Add it on top of an internal function to emit a deprecation warning + allow bc with the non internal version of the function. Tested with `python test/test_utils.py TestDeprecate.test_deprecated ` Furthermore, testing with a modified version of the tes in the pr gives something like this which is what we want ``` /home/sahanp/repos/pytorch/test/test_utils.py:1239: UserWarning: deprecated_api is DEPRECATED, please consider using an alternative API(s). deprecated_api(1, 2) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155127 Approved by: https://github.com/albanD Co-authored-by: albanD <desmaison.alban@gmail.com>	2025-06-25 18:09:04 +00:00
henrylhtsang	4bc3e4b497	[cutlass backend] Move cutlass key to cutlass_library (#156654 ) Differential Revision: [D77188311](https://our.internmc.facebook.com/intern/diff/D77188311/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156654 Approved by: https://github.com/ColinPeppler, https://github.com/jingsh ghstack dependencies: #156651	2025-06-25 17:55:57 +00:00
Arijit Mukhopadhyay	c1a629f76d	Update device for perf dashboard on AMD runners (#156809 ) Uses arch_device naming convention for storing perf dashboard logs on AMD runners based on the following PR https://github.com/pytorch/test-infra/pull/6793 Updated from zen_cpu_x86 to cpu_x86_zen Fixes https://github.com/pytorch/test-infra/issues/6823 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156809 Approved by: https://github.com/desertfire, https://github.com/malfet	2025-06-25 17:34:49 +00:00
henrylhtsang	e071837594	[cutlass backend] compile and link for .so files (#155876 ) Differential Revision: [D76482736](https://our.internmc.facebook.com/intern/diff/D76482736/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155876 Approved by: https://github.com/coconutruben, https://github.com/ColinPeppler	2025-06-25 17:01:56 +00:00
zhxchen17	1051b93192	[export] Implement _compile_and_package for ExportPackage. (#156638 ) add a method to implement weight sharing. Differential Revision: [D76132005](https://our.internmc.facebook.com/intern/diff/D76132005/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156638 Approved by: https://github.com/tugsbayasgalan	2025-06-25 16:00:40 +00:00
Andrey Talman	8eb3c5b7a1	[release] delete tag-docker-images.sh as not required anymore (#156737 ) Thanks to @clee2000 This is no longer required since the docker images use hash as tag: https://github.com/pytorch/pytorch/actions/runs/15844298044/job/44662813176#step:15:92 ``` Login Succeeded ++ docker manifest inspect docker.io/pytorch/manylinux2_28-builder:cuda12.9-5011468da53e13424002bd211cc919a0ec0e8b09 ++ jq '[.layers[].size, .config.size] \| add / 1024 / 1024' + IMAGE_SIZE=9322.26076889038 + echo 'Compressed size of image in MB: 9322.26076889038' + set -e + docker inspect --type=image docker.io/pytorch/manylinux2_28-builder:cuda12.9-5011468da53e13424002bd211cc919a0ec0e8b09 Compressed size of image in MB: 9322.26076[88](https://github.com/pytorch/pytorch/actions/runs/15844298044/job/44662813176#step:15:90)9038 + retry docker pull docker.io/pytorch/manylinux2_28-builder:cuda12.9-5011468da53e13424002bd211cc919a0ec0e8b09 + docker pull docker.io/pytorch/manylinux2_28-builder:cuda12.9-5011468da53e13424002bd211cc919a0ec0e8b09 cuda12.9-5011468da53e13424002bd211cc919a0ec0e8b09: Pulling from pytorch/manylinux2_28-builder ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156737 Approved by: https://github.com/clee2000	2025-06-25 15:17:06 +00:00
PyTorch MergeBot	029e2b05c2	Revert "[Quant][CPU] fix fake_quantize_per_tensor_affine of inf values (#155109 )" This reverts commit 19ffb5e6f7606436249742b0f3efc0bab244dc55. Reverted https://github.com/pytorch/pytorch/pull/155109 on behalf of https://github.com/albanD due to The corresponding test still breaks on rocm ([comment](https://github.com/pytorch/pytorch/pull/155109#issuecomment-3004698438))	2025-06-25 13:05:40 +00:00
Xia, Weiwen	c2185dc4a5	[Quant][CPU] Enable fp8 qlinear (#155678 ) Summary Enable fp8 qlinear on CPU. It's part of the plan to enable fp8 static quantization on CPU. This PR only adds FP8 support of the existing int8 qlinear op. It does not add a new op nor does it affect frontend or quantization flow. The schema of the qlinear op is not changed either. So, the FP8 qlinear shares the same op as INT8 qlinear and the difference is that src/wei dtype is fp8 instead of int8. The output dtype can be fp8/float32/bfloat16. The implementation uses the oneDNN library. The differences of qlinear from `_scaled_mm` are that - Qlinear supports post op fusion while `_scaled_mm` does not - Weights are prepacked for qlinear Test plan ``` pytest test/quantization/core/test_quantized_op.py -k "qlinear and fp8" ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155678 Approved by: https://github.com/leslie-fang-intel, https://github.com/jerryzh168	2025-06-25 10:01:08 +00:00
Xia, Weiwen	19ffb5e6f7	[Quant][CPU] fix fake_quantize_per_tensor_affine of inf values (#155109 ) Fixes #154328 Summary Fail reason: The input value is infinity in float and it has undefined behavior to convert it to int64_t. On X86, it will be converted to the min value of int64_t, which is not expected. Fix: Clamping `(input * inv_scale + zero_point)` to `[quant_min, quant_max]` before converting it to int64_t. Test plan ``` pytest test/quantization/core/test_workflow_ops.py -k test_fake_quantize_per_tensor_affine_inf ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155109 Approved by: https://github.com/leslie-fang-intel, https://github.com/jerryzh168	2025-06-25 09:28:54 +00:00
Aleksei Nikiforov	0ab075a69e	Fix docker image build for s390x (#156687 ) Add upstream patch for onnxruntime updating eigen dependency URL and hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156687 Approved by: https://github.com/seemethere	2025-06-25 09:09:22 +00:00
Teja Rao	4918502d2e	bug fix for losing shape on wrapper tensor for DTensor (#156774 ) Summary: Wrapper tensor for DTensor is losing shape in offload_tensor. This PR fixes this bug. Test Plan: updated the test. Test fails with old code and passes with the fix. Rollback Plan: Differential Revision: D77269733 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156774 Approved by: https://github.com/mikaylagawarecki	2025-06-25 08:14:16 +00:00
Xinya Zhang	d9577df312	[ROCm] Bump AOTriton to 0.10b (#156499 ) Notable new features/optimizations for SDPA operators on AMD systems from AOTriton 0.10b: * Official support of gfx950/gfx1201 * Experimental support of gfx1101/gfx1151/gfx1150/gfx1200 * Reduce libaotriton.so binary size by over 80%. + Without this optimization the binary size of `libaotriton.so` could be over 100MiB due to 2x more supported architectures compared with 0.9b. Now it is only about 11MiB. * Support sliding window attention (SWA) in `_flash_attention_forward/backward`. Should fix #154582 See https://github.com/ROCm/aotriton/releases/tag/0.10b for full details, including Known Problems. Notable changes to SDPA backend: * `std::optional<int64_t>` `window_size_left/right` are directly passed to ROCM's SDPA backend, because the default value `-1` is meaningful to AOTriton's backend and bottom-right aligned causal mask is implemented with negative `window_size_left/right` * Some code clean up around `USE_CK_FLASH_ATTENTION` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156499 Approved by: https://github.com/jeffdaily, https://github.com/jithunnair-amd	2025-06-25 07:09:03 +00:00
Xuehai Pan	62272d5b24	[BE][Easy][setup] wrap over long error messages and redirect them to `stderr` in `setup.py` (#156043 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156043 Approved by: https://github.com/jingsh	2025-06-25 06:57:59 +00:00
Sheng Qin	6c008e2fb5	[nativert] Move ParallelGraphExecutor to PyTorch core (#156751 ) Summary: `ParallelGraphExecutor` inherits from `GraphExecutorBase` and executes all nodes in the graph in a parallel manner Test Plan: CI Rollback Plan: Differential Revision: D77088996 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156751 Approved by: https://github.com/zhxchen17, https://github.com/dolpm	2025-06-25 06:54:45 +00:00
Isuru Fernando	44a5f93462	[dynamo] allow symints in list.__setitem__ (#156197 ) Fixes https://github.com/pytorch/pytorch/issues/155174 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156197 Approved by: https://github.com/StrongerXi	2025-06-25 06:20:35 +00:00
Xuehai Pan	162ca185ff	[BE][PYFMT] migrate PYFMT for `torch/_[a-h]*/` to `ruff format` (#144551 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/144551 Approved by: https://github.com/ezyang ghstack dependencies: #148186	2025-06-25 06:16:06 +00:00
morrison-turnansky	9642c75689	added stubs for jit tree views (#156504 ) Fixes #156488 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156504 Approved by: https://github.com/ezyang	2025-06-25 06:15:17 +00:00
yuchengliu1	c60327ba74	avoid to declare an unknown bound array without any element (#156543 ) Fixes #153180 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156543 Approved by: https://github.com/jansel Co-authored-by: Xu Han <xu.han@outlook.com>	2025-06-25 06:14:57 +00:00
Wang, Chuanqi	4237ee3c33	[XPU] Add periodic run for xpu worklfow (#156698 ) Enable XPU periodic testing in xpu.yml workflow directly. It works for https://github.com/pytorch/pytorch/issues/114850. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156698 Approved by: https://github.com/atalman, https://github.com/huydhn	2025-06-25 05:57:52 +00:00
leslie-fang-intel	194c221e0a	Update the UT of test_decompose_mm_cpu (#154100 ) Summary Fixes #153616 Based on the latest decomposed heuristic in `daca611465/torch/_inductor/fx_passes/decompose_mem_bound_mm.py (L79-L82)`, for the shape in this test case `[m=1, k=64, n=32]`, the result should be decomposed. The previous CI didn't capture this failure due to the UT skip described in https://github.com/pytorch/pytorch/pull/153245. So this PR should be verified in CI after https://github.com/pytorch/pytorch/pull/153245 landed. Test Plan ``` python -u -m pytest -s -v test/inductor/test_decompose_mem_bound_mm.py -k test_decompose_mm_cpu ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154100 Approved by: https://github.com/jansel	2025-06-25 05:45:58 +00:00
Yidi Wu	f5f4beaf56	[invoke_subgraph] make collect_meta_analysis fake prop cachable (#156347 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156347 Approved by: https://github.com/anijain2305, https://github.com/zou3519 ghstack dependencies: #156260	2025-06-25 04:29:22 +00:00
Yidi Wu	558d7f7db0	[invoke_subgraph] make same subgraph share get_attr target (#156260 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156260 Approved by: https://github.com/anijain2305, https://github.com/zou3519	2025-06-25 04:29:22 +00:00
Aaron Orenstein	568ca89bac	Add a crash handler to async compile subprocesses (#155068 ) When the async compile subprocesses crash in C++ they tend to just silently die instead of leaving any kind of trace. This installs a crash handler so that if they SEGV, ILL, or ABRT they'll attempt to output a backtrace instead. While in there I also cleaned up the CLANGTIDY warnings coming from Module.cpp. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155068 Approved by: https://github.com/masnesral	2025-06-25 03:27:28 +00:00
Natalia Gimelshein	beb52f5c0a	use more efficient implementation for broadcasted indexing in determi… (#156744 ) …nistic scatter_add per title Pull Request resolved: https://github.com/pytorch/pytorch/pull/156744 Approved by: https://github.com/suo	2025-06-25 02:59:50 +00:00
Yu, Guangye	9b498d3bb2	Update docs for torch.device (#156686 ) # Motivation Update the doc, to make `torch.device`'s constructor officially support the following methods: - A device string, which is a string representation of the device type and optionally the device ordinal. - A device type and a device ordinal. - A device ordinal, which is treated as the current accelerator type. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156686 Approved by: https://github.com/albanD	2025-06-25 02:12:36 +00:00
bobrenjc93	3608737347	[ez] fix typo in comment (#156402 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156402 Approved by: https://github.com/BoyuanFeng ghstack dependencies: #156397	2025-06-25 02:07:36 +00:00
Ryan Guo	d06a406656	[dynamo] Graph break on `torch.Tensor.data` assignment with mismatched dtype (#156623 ) Fixes #152162. Discussed with @bdhirsh and decided this is the easiest workaround for now. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156623 Approved by: https://github.com/bdhirsh	2025-06-25 02:03:04 +00:00
FFFrog	e8cf5ff564	Fix the Problems About Defining Static Variable in Inline Function (#147095 ) Refer to https://github.com/pytorch/pytorch/issues/125465 for more informations - Remove unused header files - Move common functionality to separate files to reduce dependencies between picklers and unpicklers - Move the inline function that defines the static variable to .cc Differential Revision: [D76266755](https://our.internmc.facebook.com/intern/diff/D76266755) Pull Request resolved: https://github.com/pytorch/pytorch/pull/147095 Approved by: https://github.com/cyyever, https://github.com/albanD Co-authored-by: Edward Yang <ezyang@meta.com>	2025-06-25 01:59:10 +00:00
cyy	41910d7a94	Move use of c10::string_view to std::string_view (#152509 ) Eliminate use of c10::string_view in OSS. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152509 Approved by: https://github.com/ezyang	2025-06-25 01:57:49 +00:00
Valentine233	02c7ab2f9b	[cpp wrapper] add AOTI shim for collective ops (#154492 ) Implementations: 1. Move collective ops to c10d namespace, so that we can call them externally. 2. Add AOTI shims for collective ops. Testing 1. Add c10d functional UT for cpu. 2. Include the above one in cpp wrapper UT. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154492 Approved by: https://github.com/desertfire	2025-06-25 01:20:05 +00:00
Teja Rao	d797038ea9	[dcp_poc] Introduce a new simple rank local checkpointer (#156142 ) Summary: Adds an experimental implementation for a rank local checkpointer with save and load with partial load, blind load and in-place load. This uses an new API and simpler format. Plan to add async checkpointing, IO layer, pluggable storage backend, layout customization, Resharding, deduplication etc are not implemented. Test Plan: unit tests Reviewed By: saumishr Differential Revision: D75426560 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156142 Approved by: https://github.com/saumishr	2025-06-25 01:19:40 +00:00
Pavan Balaji	0d8e4e2327	[PG/nccl] improvements to eager init (#156748 ) Summary: Cleanup eager init management, to detect and throw a warning when multiple p2p are issued on the same PG in eager init mode. Test Plan: CI Pull Request resolved: https://github.com/pytorch/pytorch/pull/156748 Approved by: https://github.com/wconstab, https://github.com/kwen2501, https://github.com/Skylion007	2025-06-25 01:04:37 +00:00
Vijay Kestur	06930706a1	Improve documentation for torch.lobpcg (#156139 ) The changes are documentation changes to the function lobpcg. There are three changes to the doc. 1. Match doc arg description to be in the same order as the parameters to the function. 2. Update documentation for arg `n` to indicate that when arg `x` is specified value of `n` is ignored if set. 3. Add warning that `m` must be bigger than 3 x the number of requested eigenpairs. Fixes #152107 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156139 Approved by: https://github.com/soulitzer	2025-06-25 00:39:34 +00:00
PyTorch MergeBot	3dd872e6d5	Revert "Add DeviceAllocator as the base device allocator (#138222 )" This reverts commit 92409b6c89fbfbd3caa79c81b1e3d9e7917d3bc7. Reverted https://github.com/pytorch/pytorch/pull/138222 on behalf of https://github.com/Camyll due to internal build failures ([comment](https://github.com/pytorch/pytorch/pull/138222#issuecomment-3002206756))	2025-06-25 00:11:35 +00:00
PyTorch MergeBot	6459a5c7a9	Revert "Add unified memory APIs for torch.accelerator (#152932 )" This reverts commit 35e44067c4d9cc9be2652c0b9098885c5a321029. Reverted https://github.com/pytorch/pytorch/pull/152932 on behalf of https://github.com/Camyll due to internal build failures ([comment](https://github.com/pytorch/pytorch/pull/138222#issuecomment-3002206756))	2025-06-25 00:11:35 +00:00
PyTorch MergeBot	fd4bb29410	Revert "[logging] dynamo_timed for CachingAutotuner.coordinate_descent_tuning (#156517 )" This reverts commit fb75dea2c1b93c78dccf08d5fd5e20b362ecd405. Reverted https://github.com/pytorch/pytorch/pull/156517 on behalf of https://github.com/Camyll due to internal reverted ([comment](https://github.com/pytorch/pytorch/pull/156517#issuecomment-3002172049))	2025-06-24 23:45:13 +00:00
IvanKobzarev	313a6a8ef9	[pt2][pr_time_benchmarks] Refresh instructions count after disabled test (#156738 ) https://github.com/pytorch/pytorch/issues/153987 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156738 Approved by: https://github.com/laithsakka	2025-06-24 23:45:02 +00:00
PyTorch MergeBot	4bd18e31e5	Revert "Add fx_graph_runnable tests boilerplate (#156552 )" This reverts commit 0a2ec7681d2af973d8daaf7905431a088739dc90. Reverted https://github.com/pytorch/pytorch/pull/156552 on behalf of https://github.com/Camyll due to breaking internal ([comment](https://github.com/pytorch/pytorch/pull/156552#issuecomment-3002159473))	2025-06-24 23:34:21 +00:00
Catherine Lee	2ff3280c77	[ez] Disable some failing periodic tests (#156731 ) test_torch.py::TestTorchDeviceTypeCUDA::test_storage_use_count_cuda: Added in https://github.com/pytorch/pytorch/pull/150059 Fails in debug mode [GH job link](https://github.com/pytorch/pytorch/actions/runs/15856606665/job/44706020831) [HUD commit link](`4491326fb0`) inductor/test_inductor_freezing.py::FreezingGpuTests::test_cpp_wrapper_cuda: [GH job link](https://github.com/pytorch/pytorch/actions/runs/15856606665/job/44707119967) [HUD commit link](`4491326fb0`) started failing after moving to new cuda version https://github.com/pytorch/pytorch/pull/155234 I'll ping people if this gets merged Pull Request resolved: https://github.com/pytorch/pytorch/pull/156731 Approved by: https://github.com/huydhn	2025-06-24 23:02:21 +00:00
bobrenjc93	d8bb5ac260	[ez] fix typo in select_algorithm.py (#156625 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156625 Approved by: https://github.com/Skylion007, https://github.com/BoyuanFeng ghstack dependencies: #156445	2025-06-24 23:01:58 +00:00
Mwiza Kunda	ce97a5dcfa	[Inductor] Restrict block analysis to only match integer dims and strides (#149615 ) Restrict block analysis to only match dimension sizes and strides that are integers. E.g. `sympy` can match index expressions like `ModularIndexing(xindex, 4, 4)) + 4(ModularIndexing(xindex, 32, 2))` with the candidate below that is invalid. ```python match_expr = stride_mod0_((xindex//(dim_mod1_dim_mod2_dim_mod3_dim_mod4_))) + stride_mod1_(ModularIndexing(xindex, dim_mod2_dim_mod3_dim_mod4_, dim_mod1_)) + stride_mod2_(ModularIndexing(xindex, dim_mod3_dim_mod4_, dim_mod2_)) + stride_mod3_(ModularIndexing(xindex, dim_mod4_, dim_mod3_)) + stride_mod4_(ModularIndexing(xindex, 1, dim_mod4_)) match={ dim_mod4_: 32, dim_mod3_: 2, stride_mod3_: 4, dim_mod2_: 1/16, dim_mod1_: 4, stride_mod1_: 1, stride_mod4_: 0, stride_mod2_: 0, stride_mod0_: 0 } ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/149615 Approved by: https://github.com/blaine-rister	2025-06-24 22:43:12 +00:00
Paul Zhang	c48d0f4643	[Inductor] Fix epilogue fusion decision with 1 Triton caller as choice (#156500 ) Differential Revision: D76904773 In the current scheduler logic, if a template buffer is only a Triton template, which can result from only 1 Triton choice in the autotuning, the fusion won't be benchmarked. This can lead to an edge case in which a Triton GEMM template from the autotune lookup table can have a problematic fusion, leading to shared memory requirements above the hardware limit. `(256, 128, 64, 4, 8, 8)` is such a config, where we have seen fusion with a `.to(torch.float32)` can lead to this issue, `out of resource: shared memory, Required: 264224, Hardware limit: 232448`. We benchmark the fusion for this case to ensure it's safe. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156500 Approved by: https://github.com/jansel	2025-06-24 22:33:47 +00:00
Scott Wolchok	e96f530af5	Remove unnecessary use of c10::SmallVector from moments_utils (#156714 ) It's just making arrays of a particular size. (If it was resizing the vectors, we'd see compile errors.) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156714 Approved by: https://github.com/Skylion007	2025-06-24 22:30:10 +00:00
Jane Xu	4ee4863232	Fix #156261 _foreach_copy indexing (#156719 ) Fixes #156261 Thanks to @ngimel's fast eyes For testing, I had experimented with a broader test case change but found that creating a tensor of 2**31+1 size was too expensive to do more than just a few times. Note that while the test case does not run in CI, I did run it locally to ensure it passes with new changes and fails without. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156719 Approved by: https://github.com/albanD	2025-06-24 21:58:44 +00:00
Yiming Zhou	310e8361c5	[nativert] Move PrimKernelRegistry to PyTorch core (#156506 ) Summary: Torch Native Runtime RFC: pytorch/rfcs#72 PrimKernelRegistry manages a small subset of kernel registry in NativeRT. Including ListPack, ListUnpack, Input, Output, VarConcat, VarStack Test Plan: Internal unittests Differential Revision: D77034945 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156506 Approved by: https://github.com/zhxchen17	2025-06-24 21:42:41 +00:00
Jeff Daily	fa0ea57f5e	[ROCm][CD] upgrade to 6.4.1 patch release (#156636 ) During https://github.com/pytorch/pytorch/pull/156112, we missed upgrading the manylinux and libtorch docker images. Fixes #155292 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156636 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-24 21:41:42 +00:00
Isuru Fernando	3efb22e091	Enable C++ dynamic shape guards by default (#140756 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/140756 Approved by: https://github.com/anijain2305, https://github.com/laithsakka	2025-06-24 21:10:17 +00:00
Laith Sakka	26f7ca3972	Unify dynamic shapes APIs naming 2 (expect_true and check) attempt2 (#156518 ) Summary: The functions guard_lt, guard_equals, and guard_leq work similarly to torch.check and expect_true, but they operate on SymPy expressions. Notably, guard_equals applies local replacements before comparison, which might be better extracted into a separate function. This pull request standardizes naming conventions to match symbolic_shapes.py. Specifically, - it introduces size_vars.expect_true and size_vars.check. - guard_lt becomes check_lt - guard_leq becomes check_leq - guard_equals becomes check_equals I am also seeing a couple of wrong usages !! that i will fix in the next PR Test Plan: OSS and cont Rollback Plan: Differential Revision: D77054177 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156518 Approved by: https://github.com/bobrenjc93	2025-06-24 21:01:38 +00:00
zeshengzong	dfef1e4408	Optimize dim description in torch.max (#156153 ) Fixes #156071 ## Test Result ### Before ![image](https://github.com/user-attachments/assets/8dd0d952-277a-4197-b323-d68ae1438171) ### After ![image](https://github.com/user-attachments/assets/4af5388e-ca9e-4268-a7c4-cf16b09b899f) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156153 Approved by: https://github.com/albanD	2025-06-24 20:50:40 +00:00
PyTorch MergeBot	1dc1eedd43	Revert "[dynamo] Graph break on `torch.Tensor.data` assignment with mismatched dtype (#156623 )" This reverts commit c1ad4b8e7a16f54c35a3908b56ed7d9f95eef586. Reverted https://github.com/pytorch/pytorch/pull/156623 on behalf of https://github.com/albanD due to Breaks Dynamo tests in trunk ([comment](https://github.com/pytorch/pytorch/pull/156623#issuecomment-3001806841))	2025-06-24 20:44:42 +00:00
PyTorch MergeBot	aa280ea19f	Revert "Remove remaining CUDA 12.4 CI code (#155412 )" This reverts commit 9fed2addedb42da86b657165fe14eadc911232cf. Reverted https://github.com/pytorch/pytorch/pull/155412 on behalf of https://github.com/Camyll due to cuda 12.4 still needed ([comment](https://github.com/pytorch/pytorch/pull/155412#issuecomment-3001711830))	2025-06-24 20:05:39 +00:00
PyTorch MergeBot	19f851ce10	Revert "Simplify nvtx3 CMake handling, always use nvtx3 (#153784 )" This reverts commit 099d0d6121125062ebc05771c8330cb7cd8d053a. Reverted https://github.com/pytorch/pytorch/pull/153784 on behalf of https://github.com/Camyll due to breaking internal tests and cuda 12.4 builds still used in CI ([comment](https://github.com/pytorch/pytorch/pull/153784#issuecomment-3001702310))	2025-06-24 20:02:07 +00:00
Edward Yang	376c16703c	Document each of the private member variables on ExportedProgram (#156704 ) Authored with claude code and then reviewed by hand. If you don't like it, tell me. Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156704 Approved by: https://github.com/albanD, https://github.com/zhxchen17, https://github.com/jingsh	2025-06-24 19:56:40 +00:00
Ryan Guo	c1ad4b8e7a	[dynamo] Graph break on `torch.Tensor.data` assignment with mismatched dtype (#156623 ) Fixes #152162. Discussed with @bdhirsh and decided this is the easiest workaround for now. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156623 Approved by: https://github.com/bdhirsh	2025-06-24 19:33:11 +00:00
henrylhtsang	f97f03c7ef	[cutlass backend] delete pip cutlass path since nvidia stops supporting nvidia-cutlass (#156651 ) Differential Revision: [D77186982](https://our.internmc.facebook.com/intern/diff/D77186982/) source: https://pypi.org/project/nvidia-cutlass/ If users want to use it, they can install pytorch through wheel, git clone cutlass, and specify cutlass path via TORCHINDUCTOR_CUTLASS_DIR Pull Request resolved: https://github.com/pytorch/pytorch/pull/156651 Approved by: https://github.com/mlazos	2025-06-24 18:32:15 +00:00
Sidharth	a00a697c17	[dynamo] updated version of detecting any differences between PRs unimplemented_v2() callsites and graph_break_registry json file (#156237 ) This PR runs an automatic check as part of dynamo_wrapped to make sure that all unimplemented_v2() callsites are mapped to the JSON file. It also fixes the issue of the CI not able to expand the hints, which was the root cause of the previous workflow failure. If not, the dev gets a message giving them instructions on how to update the JSON file. I also updated a dynamic gb_type to static and updated its test_error_message to include the GBID link for the graph break (before the link would not be produced). Testing: I ran the file with the argument to ensure all cases were covered, and also tested the test in CI. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156237 Approved by: https://github.com/williamwen42	2025-06-24 18:12:23 +00:00
Manuel Candales	2d7e6c6241	[MPS] Revert cumsum/cumprod to MPSGraph implementation (#156708 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156708 Approved by: https://github.com/malfet	2025-06-24 18:12:18 +00:00
Dylan Maloy	af284b45d5	[sigmoid] layout planner alias analyzer (#156676 ) Summary: we need a mechanism that provided the functionschemas for each kernel will be able to trace aliasing behaviour s.t., we have correct value lifetimes when we plan. Test Plan: ci + unit tests Reviewed By: SherlockNoMad Differential Revision: D73635213 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156676 Approved by: https://github.com/zhxchen17	2025-06-24 18:11:03 +00:00
Guilherme Leobas	644cc58dff	Add CPython exception tests (#150789 ) ---- * test_baseexception.py * test_exceptions.py * test_exception_variations.py * test_raise.py * test_sys.py Minor changes were made to each test to run them inside Dynamo One can reproduce the changes by downloading the tests from CPython and applying the diff: ```bash for f in "test_raise" "test_sys" "test_exceptions" "test_baseexception" "test_exception_variations"; do wget -O "test/dynamo/cpython/3_13/${f}.py" "https://raw.githubusercontent.com/python/cpython/refs/heads/3.13/Lib/test/${f}.py" git apply "test/dynamo/cpython/3_13/${f}.diff" done ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150789 Approved by: https://github.com/zou3519	2025-06-24 18:06:42 +00:00
William Wen	5ad2bee2c8	[dynamo] fix segfault due to dangling CacheEntry backend pointer (#156527 ) Fixes https://github.com/pytorch/pytorch/issues/155057 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156527 Approved by: https://github.com/anijain2305, https://github.com/jansel	2025-06-24 17:57:14 +00:00
Ruben Rodriguez Buchillon	4491326fb0	[inductor] select_algorithm: add preprocessing fns (#156464 ) Summary: # Why - keep code cleaner - modular way to hook up preprocessing steps - expand testability of flows that change which choices are provided e.g. to test performance models and lookup tables by running torch.compile # What - similar to feedback_saver_fns, now there are preprocessing_fns - the existing regex logic is exported into those as a proof of concept Test Plan: ``` buck2 run mode/opt scripts/coconutruben/torchmm:experiment 2>&1 \| tee /tmp/epx038 ``` This does not exercise the logic, it just shows that it's safe right now Rollback Plan: Differential Revision: D76946993 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156464 Approved by: https://github.com/masnesral	2025-06-24 16:44:40 +00:00
Dmitry Nikolaev	6e17315cd3	Skip FSDP tests if device count is less then requested world_size value (#155836 ) Usually `world_size=torch.cuda.device_count()` for FSDPTest-based tests But distributed test class `TestFullyShardAllGatherExtensionsMultiProcess` [forces to use `world_size=2`](`0a6e1d6b9b/test/distributed/_composable/fsdp/test_fully_shard_extensions.py (L170)`) even for 1 GPU. Then NCCL fails with errors: ``` HIP_VISIBLE_DEVICES=0 python distributed/_composable/fsdp/test_fully_shard_extensions.py -v -k test_all_gather_extensions_train_parity ... ncclInvalidUsage: This usually reflects invalid usage of NCCL library. Duplicate GPU detected : rank 1 and rank 0 both on CUDA device c000 Duplicate GPU detected : rank 0 and rank 1 both on CUDA device c000 ``` The test method [has `@skip_if_lt_x_gpu(2)` decorator](`0a6e1d6b9b/test/distributed/_composable/fsdp/test_fully_shard_extensions.py (L209)`), but test fails during test class initialization before decorator activation This PR will skip FSDPtest-based tests if `world_size > torch.cuda.device_count()` ``` HIP_VISIBLE_DEVICES=0 python distributed/_composable/fsdp/test_fully_shard_extensions.py -v -k test_all_gather_extensions_train_parity ... dist init r=0, world=2 dist init r=1, world=2 SKIPPED [15.5507s] (Need at least 2 CUDA devices) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155836 Approved by: https://github.com/jeffdaily	2025-06-24 16:38:23 +00:00
Tom Ritchford	e2c9d8d641	Fix non-bitwise type annotations for Tensor operators (see #145838 ) (#146845 ) Fix https://github.com/pytorch/pytorch/issues/145838 Pull Request resolved: https://github.com/pytorch/pytorch/pull/146845 Approved by: https://github.com/Skylion007	2025-06-24 15:41:34 +00:00
Catherine Lee	cb853945a7	[ez][CI] Update viable strict: change concurrency group to cancel in progress (#156619 ) Should help with https://github.com/pytorch/pytorch/issues/156425 The one I saw today was because the job was waiting for an environment deployment approval for mergebot environment, which I think comes from something like a temporary github outage or a dropped webhook since it should have permissions as it was on the main branch, and other runs are fine The run is https://github.com/pytorch/pytorch/actions/runs/15820977440 but you can't see anything about waiting for deployment anymore My solution is to change the concurrency group so that it will cancel in progress jobs if there is one. My hope is that if one gets stuck, the next one will cancel and re do the environment check. I don't know how to replicate this because apparently you're just supposed to fail if you don't match the protection rules https://github.com/pytorch/pytorch/actions/runs/15830920815 The job runs every 30 minutes so there might be an issue if this job needs to run for >30 minutes to find a green sha, but takes <5 minutes to run usually so I think its ok Pull Request resolved: https://github.com/pytorch/pytorch/pull/156619 Approved by: https://github.com/atalman	2025-06-24 15:37:43 +00:00
Jeddie Ji	4c59edf0c5	[nativert] Move call_torchbind_kernel (#156571 ) Summary: Move call_torchbind_kernel target from internal sigmoid to pytorch Test Plan: Test Internally: buck2 test mode/dev-nosan caffe2/test/cpp/nativert:op_kernel_test buck build //sigmoid/core/kernels:kernel_factory and all sandcastle tests Rollback Plan: Differential Revision: D77118592 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156571 Approved by: https://github.com/zhxchen17	2025-06-24 15:24:06 +00:00
leslie-fang-intel	795a6a0aff	Update github first merge rule (#156583 ) Summary Update the merge rules for `CPU Frontend` and `Autocast`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156583 Approved by: https://github.com/atalman	2025-06-24 14:04:22 +00:00
Guilherme Leobas	dd78d6e7ea	Add CPython generator/contextlib tests (#150796 ) Tests: * test_generator.py * test_generator_stop.py * test_contextlib.py Minor changes were made to each test to run them inside Dynamo. We intentionally didn't copy the binary files stored in `python/Lib/test/archivetestdata` for security reasons. There's a single test that requires a binary file and it is skipped because of that. The tests were downloaded from CPython 3.13 and the diff was generated using `git diff` to apply the changes: ```bash for f in "test_contextlib" "test_generators" "test_generator_stop"; do wget -O "test/dynamo/cpython/3_13/${f}.py" "https://raw.githubusercontent.com/python/cpython/refs/heads/3.13/Lib/test/${f}.py" git apply "test/dynamo/cpython/3_13/${f}.diff" done ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150796 Approved by: https://github.com/williamwen42	2025-06-24 13:15:04 +00:00
Nikita Shulga	3a7ff829c5	Fix MacOS MP hang in Python-3.12+ (#155698 ) By leaking resource_tracker destructor (introduced by https://github.com/python/cpython/issues/88887 ) at exit, as at this point handle to child process might no longer be valid Also, switch CI from using `setup-miniconda` to `setup-python` as an integration test for the fix as all data loader tests will hang otherwise - Remove `CONDA_RUN` macro... - Hack the search path in `macos-test.sh` to put both python and python3 aliases first in the path (not sure what other action are messing with path environment variable) Fixes https://github.com/pytorch/pytorch/issues/153050 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155698 Approved by: https://github.com/atalman	2025-06-24 12:13:35 +00:00
Xuehai Pan	f5e6e52f25	[BE][PYFMT] migrate PYFMT for `test/inductor/` to `ruff format` (#148186 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/148186 Approved by: https://github.com/jansel	2025-06-24 11:12:11 +00:00
Mark Saroufim	4e8dd11be1	simplify nvrtc discovery login in compile_kernel (#156674 ) Followup from https://github.com/pytorch/pytorch/pull/156332 Tested a bunch while I was working on https://github.com/pytorch/pytorch/pull/156380 Works just fine on dev gpus Pull Request resolved: https://github.com/pytorch/pytorch/pull/156674 Approved by: https://github.com/malfet	2025-06-24 08:55:40 +00:00
Mark Saroufim	ce73b0c53f	Validate custom op support for compile_kernel (#156332 ) Follow-up work from #151484 - just makes sure that compile_kernel composes nicely with custom ops by writing some new tests, no new code functionality is added benchmark failure in CI is unrelated to this change, CI is green Pull Request resolved: https://github.com/pytorch/pytorch/pull/156332 Approved by: https://github.com/zou3519, https://github.com/malfet	2025-06-24 08:21:21 +00:00
Yu, Guangye	35e44067c4	Add unified memory APIs for torch.accelerator (#152932 ) # Motivation The following API will be put under torch.accelerator - empty_cache - max_memory_allocated - max_memory_reserved - memory_allocated - memory_reserved - memory_stats - reset_accumulated_memory_stats - reset_peak_memory_stats Pull Request resolved: https://github.com/pytorch/pytorch/pull/152932 Approved by: https://github.com/albanD ghstack dependencies: #138222	2025-06-24 07:57:48 +00:00
cyy	ce1a07570d	Fix TORCH_CUDA_ARCH_LIST (#156667 ) Before the fix, `TORCH_CUDA_ARCH_LIST` variable contains string `TORCH_CUDA_ARCH_LIST` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156667 Approved by: https://github.com/ngimel	2025-06-24 07:27:53 +00:00
fengqing.lu	04178d347c	[Reland] [Intel GPU] Make SDPA output has the same stride as Query. (#154340 ) Fixes [#153903](https://github.com/pytorch/pytorch/issues/153903). Currently the output tensor of SDPA XPU is always defined as contiguous stride, while CPU/CUDA flash_attention and cudnn_attention allocate output tensor with stride the same as Query. This PR aligns XPU's behavior with CUDA/CPU to make XPU compatible to CPU/CUDA's modeling code. The function `alloc_with_matching_layout` is copied from cudnn `8c16d0e404/aten/src/ATen/native/cudnn/MHA.cpp (L874)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154340 Approved by: https://github.com/guangyey, https://github.com/drisspg	2025-06-24 06:09:59 +00:00
Ti-Tai Wang	a7b29c88b1	[ONNX] Preserve all legacy exporter params in fallback (#156659 ) Fixes #151693 Previous to this PR, the fallback does not take care of all user parameters. This pr preserves them to ensure a smooth transition for users. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156659 Approved by: https://github.com/justinchuby	2025-06-24 05:28:55 +00:00
Yu, Guangye	a6a8641c8a	Fix UT failure on non-cuda backend (#156577 ) # Motivation `HAS_TRITON` is a generic API that could return `True` on xpu backend. It will result in these cases failing on xpu. So we should use `HAS_CUDA` (equivalently `torch.cuda.is_available() && HAS_TRITON`) to avoid these failures. Please refer to https://github.com/pytorch/pytorch/actions/runs/15813693789/job/44569593370#step:15:2129 # Additional Context This PR aims to fix the CI failure soon. We will have a dedicated PR to generalize these UT to be generic. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @chenyang78 @kadeng @chauhang @amjames @daisyden Fix https://github.com/pytorch/pytorch/issues/156576 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156577 Approved by: https://github.com/jansel	2025-06-24 05:24:24 +00:00
zeshengzong	495c317005	Replace deprecated `is_compiling` method (#154476 ) Replace depreacted `is_compiling` in `torch._dynamo` with `torch.compiler` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154476 Approved by: https://github.com/eellison	2025-06-24 05:16:40 +00:00
Boyuan Feng	1044934878	[CUDAGraph] add config `cudagraph_capture_sizes` (#156551 ) Users may want CUDAGraph for certain sizes and fallback for other sizes. As discussed in Issue #121968, we would like to use cudagraph for [batch size [1,2,3,...,16]](https://github.com/pytorch/pytorch/issues/121968#issuecomment-2259942345) and fallback for others. Another use case is [vllm](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/cuda_piecewise_backend.py#L114-L119), where 67 batch sizes (i.e., [1,2,4,8,16,24,32,...,512]) are captured and all other sizes fallback. This PR implements the feature with `torch._inductor.config.triton.cudagraph_capture_sizes`. When it is specified, we only capture cudagraph for these shapes. When it is None (by default), we capture cudagraph for all shapes. Example: ```python import torch torch._inductor.config.triton.cudagraph_capture_sizes = [(2,3), (4,5), (6, 2), (7,3)] def f(x): return x + 1 f = torch.compile(f, mode="reduce-overhead", dynamic=False) def run(batch_size, seq_len, d): x = torch.randn((batch_size, seq_len, d), device="cuda") # Need to mark the dimension as dynamic. Automated-dynamic # may have some ux issues on matching `cudagraph_capture_sizes` # with the actual dynamic shapes, since there are specialization and # multiple dynamo graphs. torch._dynamo.mark_dynamic(x, 0) torch._dynamo.mark_dynamic(x, 1) for _ in range(3): f(x) for i in range(2, 10): for j in range(2, 10): run(i, j, 8) num_cudagraph = torch._inductor.cudagraph_trees.get_container(0).tree_manager.new_graph_id() assert num_cudagraph.id == 4 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156551 Approved by: https://github.com/bobrenjc93	2025-06-24 05:14:49 +00:00
Ahmad Sharif	899d3d3e9e	Don't call `sum()` on a tensor that is not summable in layer_norm (#156600 ) Don't call `sum()` on a tensor that is default constructed. Previously we could call `sum()` on a tensor that was default-contructed. That would lead to an error like this: ``` Traceback (most recent call last): File "/home/ahmads/.conda/envs/pt3/lib/python3.12/unittest/case.py", line 58, in testPartExecutor yield File "/home/ahmads/.conda/envs/pt3/lib/python3.12/unittest/case.py", line 634, in run self._callTestMethod(testMethod) File "/home/ahmads/.conda/envs/pt3/lib/python3.12/unittest/case.py", line 589, in _callTestMethod if method() is not None: ^^^^^^^^ File "/home/ahmads/personal/pytorch/torch/testing/_internal/common_utils.py", line 3191, in wrapper method(args, kwargs) File "/home/ahmads/personal/pytorch/test/test_nn.py", line 7235, in test_layer_norm_backwards_eps ln_out_cuda.backward(grad_output_cuda) File "/home/ahmads/personal/pytorch/torch/_tensor.py", line 647, in backward torch.autograd.backward( File "/home/ahmads/personal/pytorch/torch/autograd/__init__.py", line 354, in backward _engine_run_backward( File "/home/ahmads/personal/pytorch/torch/autograd/graph.py", line 829, in _engine_run_backward return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ RuntimeError: tensor does not have a device Exception raised from device_default at /home/ahmads/personal/pytorch/c10/core/TensorImpl.h:1265 (most recent call first): C++ CapturedTraceback: #4 std::_Function_handler<std::shared_ptr<c10::LazyValue<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > const> (), c10::SetStackTraceFetcher(std::function<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > ()>)::{lambda()#1}>::_M_invoke(std::_Any_data const&) from Logging.cpp:0 #5 c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) from ??:0 #6 c10::detail::torchCheckFail(char const, char const, unsigned int, char const) from ??:0 #7 at::TensorBase::options() const from :0 #8 at::meta::resize_reduction(at::impl::MetaBase&, at::Tensor const&, c10::OptionalArrayRef<long>, bool, c10::ScalarType, bool) from :0 #9 at::meta::structured_sum_dim_IntList::meta(at::Tensor const&, c10::OptionalArrayRef<long>, bool, std::optional<c10::ScalarType>) from ??:0 #10 at::(anonymous namespace)::wrapper_CompositeExplicitAutogradNonFunctional_sum_dim_IntList(at::Tensor const&, c10::OptionalArrayRef<long>, bool, std::optional<c10::ScalarType>) from RegisterCompositeExplicitAutogradNonFunctional_0.cpp:0 #11 c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (at::Tensor const&, c10::OptionalArrayRef<long>, bool, std::optional<c10::ScalarType>), &at::(anonymous namespace)::wrapper_CompositeExplicitAutogradNonFunctional_sum_dim_IntList>, at::Tensor, c10::guts::typelist::typelist<at::Tensor const&, c10::OptionalArrayRef<long>, bool, std::optional<c10::ScalarType> > >, at::Tensor (at::Tensor const&, c10::OptionalArrayRef<long>, bool, std::optional<c10::ScalarType>)>::call(c10::OperatorKernel, c10::DispatchKeySet, at::Tensor const&, c10::OptionalArrayRef<long>, bool, std::optional<c10::ScalarType>) from RegisterCompositeExplicitAutogradNonFunctional_0.cpp:0 #12 at::_ops::sum_dim_IntList::call(at::Tensor const&, c10::OptionalArrayRef<long>, bool, std::optional<c10::ScalarType>) from ??:0 #13 void at::native::(anonymous namespace)::LaunchGammaBetaBackwardCUDAKernel<float, float>(float const, float const, float const, float const, long, long, at::Tensor, at::Tensor, CUstream_st) from ??:0 #14 void at::native::(anonymous namespace)::LayerNormBackwardKernelImplInternal<float>(at::Tensor const&, at::Tensor const&, at::Tensor const&, at::Tensor const&, at::Tensor const&, long, long, at::Tensor, at::Tensor, at::Tensor) from ??:0 #15 at::native::(anonymous namespace)::LayerNormBackwardKernelImpl(at::Tensor const&, at::Tensor const&, at::Tensor const&, at::Tensor const&, at::Tensor const&, long, long, at::Tensor, at::Tensor, at::Tensor) from ??:0 #16 at::native::layer_norm_backward_cuda(at::Tensor const&, at::Tensor const&, c10::ArrayRef<long>, at::Tensor const&, at::Tensor const&, std::optional<at::Tensor> const&, std::optional<at::Tensor> const&, std::array<bool, 3ul>) from ??:0 #17 at::(anonymous namespace)::(anonymous namespace)::wrapper_CUDA__native_layer_norm_backward(at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, at::Tensor const&, at::Tensor const&, std::optional<at::Tensor> const&, std::optional<at::Tensor> const&, std::array<bool, 3ul>) from RegisterCUDA_0.cpp:0 ``` Now we only call `sum(0)` on tensors that are defined and properly guard the `sum(0)` and assignment. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156600 Approved by: https://github.com/eqy, https://github.com/ngimel	2025-06-24 05:00:42 +00:00
Edward Z. Yang	17eb649d55	Implement guard collectives (optimized version) (#156562 ) This is a remix of https://github.com/pytorch/pytorch/pull/155558 Instead of mediating guard collective via a config option, in this one it's done via a `set_stance` like API. The motivation is that checking for the config value on entry on torch.compile is apparently quite expensive, according to functorch_maml_omniglot. So this makes it a bit cheaper. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156562 Approved by: https://github.com/Microve	2025-06-24 04:59:49 +00:00
morrison-turnansky	73772919d2	remove deprecated numpy.typing.mypy_plugin in mypy.ini (#156601 ) Fixes #156489 removed deprecated numpy plugin in mypy.ini @ezyang Pull Request resolved: https://github.com/pytorch/pytorch/pull/156601 Approved by: https://github.com/ezyang	2025-06-24 04:56:08 +00:00
Xuehai Pan	6d5c789ad5	[BE][PYFMT] migrate PYFMT for `test/[a-h]*/` to `ruff format` (#144555 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/144555 Approved by: https://github.com/ezyang ghstack dependencies: #144551, #144554	2025-06-24 04:53:54 +00:00
PyTorch MergeBot	e600e044a7	Revert "[aotd] Support mutations of the same input in fw and bw (#155354 )" This reverts commit 3f920f3d8f5bd15d2222758f21f9a5d36e4dad1f. Reverted https://github.com/pytorch/pytorch/pull/155354 on behalf of https://github.com/malfet due to Not sure why CI was green, but it breaks tons of tests, see `930b575389/1` ([comment](https://github.com/pytorch/pytorch/pull/155354#issuecomment-2998780884))	2025-06-24 04:42:14 +00:00
fduwjj	930b575389	[symm_mem] Add sym mem test into ptd h100 ci (#156634 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156634 Approved by: https://github.com/ngimel, https://github.com/mori360	2025-06-24 03:43:22 +00:00
tvukovic-amd	b2d473c8f8	[ROCm][Windows] Fix rocsolver undefined symbol error (#156591 ) Fix undefined symbol error while using `rocsolver_ssyevd_strided_batched` call in `aten/src/ATen/native/cuda/linalg/BatchLinearAlgebraLib.cpp`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156591 Approved by: https://github.com/jeffdaily	2025-06-24 03:28:45 +00:00
fduwjj	87d615efab	[fr] Use a vector to temporarily keep the reference to future object to avoid block (#156653 ) At the end of the scope when std::async is launched, a wait will be called which could makes the code blocking, this is not expected for monitoring thread. Instead, let's use a vector to contain the reference to it. So no blocking will happen. And at the end of loop, wait will still be called but it is ok since all the checks or dump has already been finished. Differential Revision: [D77190380](https://our.internmc.facebook.com/intern/diff/D77190380) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156653 Approved by: https://github.com/kwen2501	2025-06-24 03:25:04 +00:00
cyy	b09bd414a6	Deprecate c10::string (#155084 ) Now there is no mention of c10::string in OSS. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155084 Approved by: https://github.com/ezyang	2025-06-24 03:03:06 +00:00
Simon Fan	0a2ec7681d	Add fx_graph_runnable tests boilerplate (#156552 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156552 Approved by: https://github.com/StrongerXi	2025-06-24 02:41:38 +00:00
dolpm	9665702c64	[nativert] reland D76832891 remove designated initializer cpp20 (#156565 ) Summary: fix windows build broke in https://github.com/pytorch/pytorch/pull/156508 Test Plan: ci Rollback Plan: Differential Revision: D77080420 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156565 Approved by: https://github.com/zhxchen17	2025-06-24 02:38:08 +00:00
Andrey Talman	6a3d00aa3b	Add Windows cuda 12.9.1 build (#156630 ) Without Support for SegmentReduce.cu Test PR confirmed by Removing SegmentReduce.cu windows build for CUDA 12.9 can succeed Related to: https://github.com/pytorch/pytorch/issues/156181 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156630 Approved by: https://github.com/malfet Co-authored-by: Ting Lu <tingl@nvidia.com> Co-authored-by: Nikita Shulga <2453524+malfet@users.noreply.github.com>	2025-06-24 02:15:49 +00:00
Sidharth	a9ef7c4d04	[dynamo] update to lru_cache message and updated user stack trace in debug mode (#156639 ) I had to create a new PR for this because of @atalman request of temporary reverting the previous PR to restore diff train sync. Nothing has changed from this PR and the original one. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156639 Approved by: https://github.com/atalman	2025-06-24 01:52:13 +00:00
Paul Zhang	86996c15dc	[Inductor] Allow exhaustive autotuning across all GEMM options (#156610 ) Differential Revision: D76843916 Exhaustive autotuning is meant to autotune GEMM configs across the entire search space of possible configs. Some of these configs can cause extremely long compilation times and OOMs, especially with configs of the following nature: Excessive register spillage Using much larger amounts of shared memory than available on the hardware This diff prunes out those configs to make exhaustive autotuning more viable, along with supporting exhaustive autotuning for persistent+tma template and decompose_k. Previously, exhaustive autotuning would hang, now we are able to tune shapes in ~5 minutes. Below is a sample log for autotuning with exhaustive: ``` AUTOTUNE mm(1152x21504, 21504x1024) strides: [21504, 1], [1, 21504] dtypes: torch.bfloat16, torch.bfloat16 mm 0.1167 ms 100.0% triton_mm_6270 0.1172 ms 99.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=256, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4, num_consumer_groups=0, num_buffers_warp_spec=0 triton_mm_6522 0.1183 ms 98.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4, num_consumer_groups=0, num_buffers_warp_spec=0 triton_mm_persistent_tma_7482 0.1190 ms 98.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, A_ROW_MAJOR=True, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, B_ROW_MAJOR=False, EVEN_K=True, GROUP_M=8, NUM_SMS=132, TMA_SIZE=128, USE_FAST_ACCUM=False, num_stages=5, num_warps=4, num_consumer_groups=0, num_buffers_warp_spec=0 triton_mm_persistent_tma_7483 0.1195 ms 97.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, A_ROW_MAJOR=True, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, B_ROW_MAJOR=False, EVEN_K=True, GROUP_M=8, NUM_SMS=132, TMA_SIZE=128, USE_FAST_ACCUM=False, num_stages=5, num_warps=8, num_consumer_groups=0, num_buffers_warp_spec=0 triton_mm_6523 0.1274 ms 91.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8, num_consumer_groups=0, num_buffers_warp_spec=0 triton_mm_6267 0.1285 ms 90.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=256, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4, num_consumer_groups=0, num_buffers_warp_spec=0 triton_mm_6519 0.1287 ms 90.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4, num_consumer_groups=0, num_buffers_warp_spec=0 triton_mm_persistent_tma_7480 0.1298 ms 89.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, A_ROW_MAJOR=True, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, B_ROW_MAJOR=False, EVEN_K=True, GROUP_M=8, NUM_SMS=132, TMA_SIZE=128, USE_FAST_ACCUM=False, num_stages=4, num_warps=4, num_consumer_groups=0, num_buffers_warp_spec=0 triton_mm_persistent_tma_7312 0.1302 ms 89.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, A_ROW_MAJOR=True, BLOCK_K=64, BLOCK_M=64, BLOCK_N=256, B_ROW_MAJOR=False, EVEN_K=True, GROUP_M=8, NUM_SMS=132, TMA_SIZE=128, USE_FAST_ACCUM=False, num_stages=4, num_warps=4, num_consumer_groups=0, num_buffers_warp_spec=0 SingleProcess AUTOTUNE benchmarking takes 298.7185 seconds and 21.2569 seconds precompiling for 2210 choices INFO:tritonbench.utils.triton_op:Took 333894.46ms to get benchmark function for pt2_matmul_maxautotune ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156610 Approved by: https://github.com/jansel	2025-06-24 01:42:05 +00:00
Isuru Fernando	40a785103c	[dynamo] fix debugging code_parts for relational guards (#154753 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154753 Approved by: https://github.com/anijain2305 ghstack dependencies: #154772	2025-06-24 01:38:29 +00:00
Isuru Fernando	849468034d	[dynamo] fix selecting shape guards (#154772 ) Not all LAMBDA_GUARDs are shape guards. Only the epilogue guards are lambda guards Pull Request resolved: https://github.com/pytorch/pytorch/pull/154772 Approved by: https://github.com/anijain2305	2025-06-24 01:38:29 +00:00
Ankita George	5dd9652389	Clean up HF components (#155707 ) Differential Revision: [D76427358](https://our.internmc.facebook.com/intern/diff/D76427358/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155707 Approved by: https://github.com/saumishr	2025-06-24 00:07:37 +00:00
Francisco Massa	ca5a40395d	[partitioner] Fix _broadcast_on_rank0 to use deterministic hash function (#153734 ) Summary: I was using python's hash, which is not deterministic across different interpreter runs. Use hashlib instead. Test Plan: Run using it https://www.internalfb.com/mlhub/pipelines/runs/mast/aps-rebase_sanity_128bs_8t_cc-8e17be61ce?job_attempt=1&version=0&tab=summary&env=prod Differential Revision: D74882405 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153734 Approved by: https://github.com/Microve	2025-06-24 00:06:23 +00:00
Georgia Phillips	24063ad109	Fix native static dispatch kernels (#156331 ) Summary: Fix for native static dispatch kernels not taking effect Test Plan: ``` buck2 test //sigmoid/backend/test:static_kernels_ops_test buck2 run mode/opt caffe2/torch/fb/model_transform/fx2trt/packaging:load_net_predictor -- --loadMode=BenchmarkByOp --inputNetFile=/data/users/$USER/models/${MODEL_ENTITY_ID}/${SNAPSHOT_ID}/${MODEL_ENTITY_ID}_${SNAPSHOT_ID}${SUFFIX} --moduleName=${MODULE} --submodToDevice "" --pytorch_predictor_sigmoid_static_dispatch_enable=true --pytorch_predictor_sigmoid_graph_passes_enable=true --benchmarkEnableProfiling=true --load_lowered_merge=3 --using_aoti_lowering_allowlist=false --requestFilePath=/data/users/georgiaphillips/replayer/inputs/742055223/0/mix/742055223_0_mix.inputs.recordio --benchmarkNumIterations=2 ``` Rollback Plan: Reviewed By: dolpm Differential Revision: D76559764 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156331 Approved by: https://github.com/Skylion007, https://github.com/jingsh	2025-06-24 00:05:49 +00:00
Shivam Raikundalia	380e30a723	[EZ/Profiler] Change 'b' to 'B' in FunctionEvent Frontend (#156250 ) Summary: Fixes https://github.com/pytorch/pytorch/issues/149311 Test Plan: Just changes string output ``` ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg Self CUDA Self CUDA % CUDA total CUDA time avg CPU Mem Self CPU Mem # of Calls ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ void at::native::vectorized_elementwise_kernel<4, at... 0.00% 0.000us 0.00% 0.000us 0.000us 60.993us 0.97% 60.993us 1.848us 0 B 0 B 33 ... ``` Rollback Plan: Differential Revision: D76857251 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156250 Approved by: https://github.com/sanrise	2025-06-23 23:25:04 +00:00
Yuanyuan Chen	07bb097698	Fix clang-tidy bugprone* warnings (#148529 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/148529 Approved by: https://github.com/ezyang	2025-06-23 23:09:56 +00:00
IvanKobzarev	3f920f3d8f	[aotd] Support mutations of the same input in fw and bw (#155354 ) Original issue: https://github.com/pytorch/pytorch/issues/154820 The issue happens when there is a mutation for the same input in forward AND in backward. AOTD emited copy_ after joint_function tracing. This made this fx-node to correspond to the side effects of both mutations (in forward and in backward). After that partitioner can put it either in forward or in backward. The fix: 1/ Introduce joint_function.handle that allows to set "post_forward" callback, to be able to check inputs state after forward We do not want to apply the mutation after joint, if we already applied it in forward. For that we need "mutation_counter" and memorize the version of mutation that we applied for forward mutation. 2/ Exposing mutation_counter to python We want to keep invariant that copy_ exist only in the end of joint graph. 3/ We memorize mutation_counter and state of the inputs after forward, using the handle post_forward. Emit post_forward mutations after joint graph fully traced. add for post_forward mutations "must_be_in_forward" tag (similar to existing "must_be_in_backward") to keep them in forward. 4/ Ban recompute of the source of mutation. Recompute can apply the same op (e.g. add) in forward and backward. For this set MUST_SAVE for the source of mutation in forward. proxy_tensor changes: By default proxy tensor updates tensor_tracker. In this case applied mutations will be chained. But we want that this copy_ will be independent and applied just to primals. For this introducing a contextmanager to be able to disable update of tensor_tracker for adding forward mutations. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155354 Approved by: https://github.com/bdhirsh	2025-06-23 22:25:45 +00:00
Scott Wolchok	c82a174cea	Extract CPU log_softmax kernels to header (#156243 ) This allows sharing them with ExecuTorch. Differential Revision: [D76830114](https://our.internmc.facebook.com/intern/diff/D76830114/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156243 Approved by: https://github.com/janeyx99	2025-06-23 21:31:16 +00:00
Paul Zhang	96e4c95cd8	[Inductor] Subgraph as a choice symbolic expression as input (#156185 ) Differential Revision: D76514984 Fix subgraph as a choice for when a symbolic shape is inputted as an expression, i.e. 256 * s0, which typically happens in the backwards pass. The current logic assumes that all symbolic shapes are single inputs, i.e. standalone s0 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156185 Approved by: https://github.com/masnesral	2025-06-23 21:29:17 +00:00
PyTorch MergeBot	b1d62febd0	Revert "Use official CUDAToolkit module in CMake (#154595 )" This reverts commit 08dae945ae380d80efbaf140a95abfc5d96e5100. Reverted https://github.com/pytorch/pytorch/pull/154595 on behalf of https://github.com/malfet due to It breaks on some local setup with no clear diagnostic, but looks like it fails to find cuFile ([comment](https://github.com/pytorch/pytorch/pull/154595#issuecomment-2997959344))	2025-06-23 21:15:31 +00:00
anwang	31e1274597	[MTIA Aten Backend] Migrate max.dim_max / min.dim_min (#156568 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate max.dim_max / min.dim_min to in-tree. Differential Revision: [D77095185](https://our.internmc.facebook.com/intern/diff/D77095185/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156568 Approved by: https://github.com/malfet ghstack dependencies: #156502, #156539, #156554	2025-06-23 20:43:39 +00:00
Colin Peppler	dfdd636cfa	[aoti] Check longlong upperbound for codegening input size check (#156522 ) Summary: Fixes ``` error: integer literal is too large to be represented in any integer type 38979 \| if (arg410_1_size[0] > 1171368248680556527362) { ``` Test Plan: ci Differential Revision: D77057898 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156522 Approved by: https://github.com/jingsh, https://github.com/desertfire	2025-06-23 20:38:34 +00:00
anwang	edd9c09e73	[MTIA Aten Backend] Migrate isnan (#156554 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate isnan to in-tree. Differential Revision: [D77094811](https://our.internmc.facebook.com/intern/diff/D77094811/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156554 Approved by: https://github.com/malfet ghstack dependencies: #156502, #156539	2025-06-23 20:22:32 +00:00
anwang	070e580d30	[MTIA Aten Backend] Migrate _log_softmax.out / _log_softmax_backward_data.out (#156539 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate _log_softmax.out / _log_softmax_backward_data.out to in-tree. Differential Revision: [D77044380](https://our.internmc.facebook.com/intern/diff/D77044380/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156539 Approved by: https://github.com/malfet ghstack dependencies: #156502	2025-06-23 19:56:01 +00:00
anwang	93cd16512f	[MTIA Aten Backend] Migrate maximum.out / minimum.out / cos.out / erf.out / exp.out (#156502 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate maximum.out / minimum.out / cos.out / erf.out / exp.out to in-tree. Differential Revision: [D76917384](https://our.internmc.facebook.com/intern/diff/D76917384/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156502 Approved by: https://github.com/malfet	2025-06-23 19:56:01 +00:00
atalman	ee4d343499	Revert "[dynamo] handle fullgraph toggle using nested torch.compile (#155166 )" (#156624 ) This reverts changes to [test/dynamo/test_repros.py](https://github.com/pytorch/pytorch/compare/main...atalman:revert_only_portion_of_file?expand=1#diff-4c82a5798a61d4cceb176b2700ba6fdd7c3e72d575b8e7e22458589139459caa) Missed by: `ee3d9969cc (diff-036cb21341ff8e390cc250e74fe9e3f0f15f259ea4bec4abcce49d95febf1553)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156624 Approved by: https://github.com/Camyll	2025-06-23 19:30:08 +00:00
Shangdi Yu	56b3bf0c74	[nativert] Move HigherOrderKernel (#156507 ) Summary: Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 As part of the effort to open source TorchNativeRuntime (or what we call Sigmoid), we are moving the implementation to torch/: fbcode/sigmoid/kernels -> fbcode/caffe2/torch/nativert/kernels Test Plan: CI Differential Revision: D77032074 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156507 Approved by: https://github.com/zhxchen17	2025-06-23 19:29:27 +00:00
PyTorch MergeBot	d061a02e6e	Revert "[invoke_subgraph] make same subgraph share get_attr target (#156260 )" This reverts commit 39dd2f4d7defc63164a7969bfac0d0c62ffac900. Reverted https://github.com/pytorch/pytorch/pull/156260 on behalf of https://github.com/ydwu4 due to no signal, it breaks linter tests. ([comment](https://github.com/pytorch/pytorch/pull/156260#issuecomment-2997478798))	2025-06-23 18:24:10 +00:00
PyTorch MergeBot	35d03398e5	Revert "[invoke_subgraph] make collect_meta_analysis fake prop cachable (#156347 )" This reverts commit f179b7198522e6d93bd103efba1a1ebd5a2cf891. Reverted https://github.com/pytorch/pytorch/pull/156347 on behalf of https://github.com/ydwu4 due to no signal, it breaks linter tests. ([comment](https://github.com/pytorch/pytorch/pull/156347#issuecomment-2997453729))	2025-06-23 18:19:29 +00:00
Tom Ritchford	98a34e8d4b	Move code out of individual token linters (#152256 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152256 Approved by: https://github.com/Skylion007	2025-06-23 18:16:33 +00:00
Prachi Gupta	da910e603a	[ROCm] update state check for test_trace_while_active* (#153545 ) When timing is enabled, ROCR runtime used to sleep for a small amount which ensured that the application saw the correct state. However, for perf reasons this sleep was removed and now the state is not guaranteed to be "started". That's why I updated the test state check to be either "started" or "scheduled" Pull Request resolved: https://github.com/pytorch/pytorch/pull/153545 Approved by: https://github.com/jeffdaily, https://github.com/pruthvistony Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-23 17:58:14 +00:00
PyTorch MergeBot	55ef7b15e0	Revert "[dynamo] fixes to lru_cache message and adding user stack trace in debug mode (#156463 )" This reverts commit afbf5420b8745099bf7d871f5a4fb6dec338f825. Reverted https://github.com/pytorch/pytorch/pull/156463 on behalf of https://github.com/atalman due to This is temoprary revert, to restore diff train sync. We should be good to reland this change ([comment](https://github.com/pytorch/pytorch/pull/156463#issuecomment-2997335541))	2025-06-23 17:44:36 +00:00
Boyuan Feng	a95504b10f	[torchbench] update environment setup script (#156465 ) Existing torchbench `Makefile` installs all models from torchbench, which could easily take 30 minutes, even if a developer only want to run 1 model. This PR adds a config to only install torchbench models we want to run. Example usage: ``` # Install 1 torchbench model make build-deps TORCHBENCH_MODELS="alexnet" # Install 3 torchbench models make build-deps TORCHBENCH_MODELS="alexnet basic_gnn_gcn BERT_pytorch" # Install all models make build-deps # Install all models make build-deps TORCHBENCH_MODELS="" ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156465 Approved by: https://github.com/ezyang	2025-06-23 17:41:29 +00:00
PyTorch MergeBot	e583b88819	Revert "[Draft][CUDA] Use runtime driver API for cuStreamWriteValue32 (#156097 )" This reverts commit ac86ec0e60370c037e018137f2048cafd47c5c28. Reverted https://github.com/pytorch/pytorch/pull/156097 on behalf of https://github.com/atalman due to internal breakage ([comment](https://github.com/pytorch/pytorch/pull/156097#issuecomment-2997314638))	2025-06-23 17:36:44 +00:00
Yidi Wu	f179b71985	[invoke_subgraph] make collect_meta_analysis fake prop cachable (#156347 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156347 Approved by: https://github.com/anijain2305, https://github.com/zou3519 ghstack dependencies: #156260	2025-06-23 17:10:07 +00:00
Yidi Wu	39dd2f4d7d	[invoke_subgraph] make same subgraph share get_attr target (#156260 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156260 Approved by: https://github.com/anijain2305, https://github.com/zou3519	2025-06-23 17:10:07 +00:00
Prachi Gupta	276c790010	[ROCm][SymmetricMemory] Avoid bf16 to float conversion during reduce (#155587 ) This PR helps improve the performance of one-shot and two-shot allreduce as reported here: https://github.com/pytorch/FBGEMM/issues/4072 One-Shot: ![image](https://github.com/user-attachments/assets/69fe0d53-6636-42e1-90e0-e5efb989f59f) As shown in the numbers presented above, symmetric memory performance prior to the PR (baseline) was on average about 26% less than fbgemm's number reported in the issue above. After this PR, we are seeing 16% improvement on average as compared to fbgemm and 59% as compared to our baseline numbers. Two-Shot: ![image](https://github.com/user-attachments/assets/e5c8a288-303e-4d50-814b-4348e589e1fc) Similarly, in two-shot, we were originally underperforming by 12%. We have improved by 22% after this PR as compared to symmetric memory performance prior to this PR. However, two-shot performance is still about 23% lower than fbgemm. This work is still in progress and will be pushing those changes through a separate PR. Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/155587 Approved by: https://github.com/jeffdaily	2025-06-23 16:14:01 +00:00
Nikita Shulga	5a533f74a1	Checkout optional submodules when publishing a release tarball (#156615 ) This includes Eigen and nccl for now Pull Request resolved: https://github.com/pytorch/pytorch/pull/156615 Approved by: https://github.com/huydhn	2025-06-23 16:08:22 +00:00
Aby Mathew C	6835ba1b34	Register hpu device to fake backend (#156076 ) ## MOTIVATION This PR intends to add hpu ( Intel Gaudi) also to the list of devices that will be supported by the "fake" distributed backend and the process group that will be created. ## CHANGES - Add "hpu" to the list of devices @ankurneog, @EikanWang Pull Request resolved: https://github.com/pytorch/pytorch/pull/156076 Approved by: https://github.com/d4l3k, https://github.com/EikanWang, https://github.com/albanD	2025-06-23 16:08:08 +00:00
Ke Wen	cc410d3761	[SymmMem] Rename all_to_all_vdev ops (#156582 ) `all_to_all_vdev` are not binding of NVSHMEM APIs. Removing the `nvshmem_` prefix. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156582 Approved by: https://github.com/fduwjj ghstack dependencies: #155134	2025-06-23 15:57:36 +00:00
Ryan Guo	640f5a7090	[dynamo] Support builtin bool on non-constant VTs (#155863 ) In practice `bool(...)` is either constant folded by Dynamo or used for branching (so most of its emulation logic lived in `InstructionTranslator.generic_jump`. This patch adds a dedicated `bool` hanlder (only for symbolic bool/int/float for now), and fixes #136075. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155863 Approved by: https://github.com/williamwen42	2025-06-23 15:53:15 +00:00
Simon Fan	6b45af38a5	[easy] better copy_misaligned_inputs assertion failure message (#154472 ) internal xref: https://fb.workplace.com/groups/1075192433118967/permalink/688540560729579/ Pull Request resolved: https://github.com/pytorch/pytorch/pull/154472 Approved by: https://github.com/williamwen42	2025-06-23 15:39:15 +00:00
Randolf Scholz	2e9bd03f60	Implemented `Size.__radd__` (#152554 ) Fixes #144334 Builds on top of #146834 by @khushi-411 The needed trick was to add `PyNumberMethods` because these Number Protocol appears to be responsible for `__radd__` (see https://stackoverflow.com/q/18794169) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152554 Approved by: https://github.com/albanD Co-authored-by: Khushi Agrawal <khushiagrawal411@gmail.com> Co-authored-by: albanD <desmaison.alban@gmail.com>	2025-06-23 15:38:37 +00:00
Nikita Shulga	3cbae6dde8	[MPSInductor][BE] Fix multistage reduction check (#156567 ) From less than max threadgroup size to less or equal to that, which eliminates redundant trivial loops. I.e. it changes shader code generated for ```python import torch def f(x): var, mean = torch.var_mean(x, dim=2, keepdim = True) return x / var, var torch.compile(f)(torch.rand(1, 16, 1024, dtype=torch.float32, device='mps')) ``` from ```metal [[max_total_threads_per_threadgroup(1024)]] kernel void generated_kernel( device float* out_ptr1, device float* out_ptr2, constant float* in_ptr0, uint2 thread_pos [[thread_position_in_grid]], uint2 group_pos [[thread_position_in_threadgroup]] ) { auto xindex = thread_pos.x; auto r0_index = thread_pos.y; int x0 = xindex; threadgroup float3 tmp_acc_0[1024]; tmp_acc_0[r0_index * 1] = 0.0; for(auto r0_1_cnt = 0; r0_1_cnt < 1; ++r0_1_cnt) { int r0_1 = 1 * r0_index + r0_1_cnt; auto tmp0 = in_ptr0[r0_1 + 1024x0]; tmp_acc_0[r0_index 1] = ::c10:🤘:welford_combine(tmp_acc_0[r0_index * 1], float3(tmp0, 0.0, 1.0)); } auto tmp1 = c10:🤘:threadgroup_welford_combine(tmp_acc_0, 1024); auto tmp2 = 1023.0; auto tmp3 = tmp1.y / tmp2; out_ptr1[x0] = static_cast<float>(tmp3); for(auto r0_1_cnt = 0; r0_1_cnt < 1; ++r0_1_cnt) { int r0_1 = 1 * r0_index + r0_1_cnt; auto tmp4 = in_ptr0[r0_1 + 1024x0]; auto tmp5 = tmp4 / tmp3; out_ptr2[r0_1 + 1024x0] = static_cast<float>(tmp5); } } ``` to ```metal [[max_total_threads_per_threadgroup(1024)]] kernel void generated_kernel( device float* out_ptr1, device float* out_ptr2, constant float* in_ptr0, uint2 thread_pos [[thread_position_in_grid]], uint2 group_pos [[thread_position_in_threadgroup]] ) { auto xindex = thread_pos.x; auto r0_index = thread_pos.y; int r0_1 = r0_index; int x0 = xindex; threadgroup float tmp_acc_0[1024]; auto tmp0 = in_ptr0[r0_1 + 1024x0]; tmp_acc_0[r0_index 1] = tmp0; auto tmp1 = c10:🤘:threadgroup_welford_reduce(tmp_acc_0, 1024); auto tmp2 = 1023.0; auto tmp3 = tmp1.y / tmp2; out_ptr1[x0] = static_cast<float>(tmp3); auto tmp4 = tmp0 / tmp3; out_ptr2[r0_1 + 1024*x0] = static_cast<float>(tmp4); } `` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156567 Approved by: https://github.com/dcci ghstack dependencies: #156566	2025-06-23 14:49:26 +00:00
Manuel Candales	e28925aa75	[MPS] Activation kernels: do compute at float precision (#155735 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155735 Approved by: https://github.com/malfet ghstack dependencies: #155304, #155316, #155462, #155479, #155571, #155586	2025-06-23 14:48:57 +00:00
PyTorch MergeBot	f5e1b24945	Revert "Enable Leak Sanitizer (#154584 )" This reverts commit c79c7bbe615265b6b3d7df39d6d5a68afd7d6b2a. Reverted https://github.com/pytorch/pytorch/pull/154584 on behalf of https://github.com/cyyever due to Need to suppress more output ([comment](https://github.com/pytorch/pytorch/pull/154584#issuecomment-2995792265))	2025-06-23 10:08:40 +00:00
PyTorch MergeBot	4f70fbbd16	Revert "Use CMake wholearchive group (#156393 )" This reverts commit d1b4e0fa9a5feb22fc6de1d36dc4c9dac685caed. Reverted https://github.com/pytorch/pytorch/pull/156393 on behalf of https://github.com/etaf due to This PR is breaking XPU windows build. ([comment](https://github.com/pytorch/pytorch/pull/156393#issuecomment-2995576362))	2025-06-23 09:03:19 +00:00
Yu, Guangye	92409b6c89	Add DeviceAllocator as the base device allocator (#138222 ) # Motivation In line with [RFC] [A device-agnostic Python device memory related API design for stream-based accelerators](https://github.com/pytorch/pytorch/issues/134978), some memory-related APIs are widely used in popular repositories, such as HuggingFace [so many if-else conditional code](https://github.com/search?q=repo%3Ahuggingface%2Faccelerate%20torch.cuda.empty_cache&type=code). We would like to introduce a generic API set under torch.accelerator namespace to generalize these user cases. <div align="center"> <table> <tr> <td> Device-specific memory APIs torch.xxx.foo</td> <td> Device-agnostic memory APIs torch.accelerator.foo</td> </tr> <tr> <td> ```python torch.xxx.empty_cache ``` </td> <td> ```python torch.accelerator.empty_cache ``` </td> </tr> <tr> <td> ```python torch.xxx.reset_peak_memory_stats ``` </td> <td> ```python torch.accelerator.reset_peak_memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.reset_accumulated_memory_stats ``` </td> <td> ```python torch.accelerator.reset_accumulated_memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_stats ``` </td> <td> ```python torch.accelerator.memory_stats ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_allocated ``` </td> <td> ```python torch.accelerator.memory_allocated ``` </td> </tr> <tr> <td> ```python torch.xxx.max_memory_allocated ``` </td> <td> ```python torch.accelerator.max_memory_allocated ``` </td> </tr> <tr> <td> ```python torch.xxx.memory_reserved ``` </td> <td> ```python torch.accelerator.memory_reserved ``` </td> </tr> <tr> <td> ```python torch.xxx.max_memory_reserved ``` </td> <td> ```python torch.accelerator.max_memory_reserved ``` </td> </tr> </table> </div> # Solution This design follows a similar pattern to `HostAllocator`. We're introducing a base class `DeviceAllocator`, from which `CUDAAllocator` and `XPUAllocator` will inherit. This allows us to provide a unified call path like: `torch.accelerator.empty_cache()` -> `GetDeviceAllocator(allocator)->empty_cache()`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/138222 Approved by: https://github.com/albanD	2025-06-23 08:49:30 +00:00
Bob Ren	d5781c8d21	remove allow-untyped-defs from torch/fx/passes/utils/fuser_utils.py (#156538 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156538 Approved by: https://github.com/ezyang	2025-06-23 08:18:16 +00:00
Pandya, Vivek Vasudevbhai	e0ae4ecca8	Refactor cpp codegen to support overridable class attributes. (#155553 ) - Refactored CppKernelProxy and CppScheduling to use class-level attributes (kernel_cls, kernel_proxy_cls) for backend-specific kernel customization. - Avoids method duplication (e.g., codegen_functions, codegen_node) for backend-specific overrides thus reduces downstream maintenance when upgrading Torch. - Ensures type safety with annotations while keeping core logic centralized and extensible. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155553 Approved by: https://github.com/leslie-fang-intel, https://github.com/jgong5	2025-06-23 07:36:30 +00:00
cyy	67ee0c6725	Remove outdated Android workarounds of nearbyintf (#151292 ) This PR uses std::nearbyint on all supported platforms. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151292 Approved by: https://github.com/ezyang	2025-06-23 06:28:15 +00:00
cyy	d1b4e0fa9a	Use CMake wholearchive group (#156393 ) Use CMake wholearchive group to simplify code. It may also support more OSes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156393 Approved by: https://github.com/ezyang	2025-06-23 06:22:34 +00:00
cyy	099d0d6121	Simplify nvtx3 CMake handling, always use nvtx3 (#153784 ) Fall back to third-party NVTX3 if system NVTX3 doesn't exist. We also reuse the `CUDA::nvtx3` target for better interoperability. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153784 Approved by: https://github.com/ezyang	2025-06-23 06:12:46 +00:00
Michael Lazos	31659964a5	[Cutlass] Fix buffer missing issues (#155897 ) Handles constants and constant folding with aoti. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155897 Approved by: https://github.com/henrylhtsang	2025-06-23 05:58:39 +00:00
cyy	c79c7bbe61	Enable Leak Sanitizer (#154584 ) It enables Leak Sanitizer and also provides a suppression file. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154584 Approved by: https://github.com/ezyang	2025-06-23 05:20:27 +00:00
Yuanyuan Chen	9fed2added	Remove remaining CUDA 12.4 CI code (#155412 ) Because no 12.4 job. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155412 Approved by: https://github.com/ezyang	2025-06-23 05:16:38 +00:00
Nikita Shulga	4cd6e96bf0	[MPSInductor] Fix nested loop var elimination (#156566 ) As reduction resuts must be kept around Add regression test that is specific for this issue Fixes https://github.com/pytorch/pytorch/issues/156426 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156566 Approved by: https://github.com/dcci	2025-06-23 04:35:16 +00:00
Xuehai Pan	d55dc00f84	[BE][11/16] fix typos in torch/ (torch/csrc/distributed/) (#156321 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156321 Approved by: https://github.com/jingsh ghstack dependencies: #156313, #156314, #156315, #156316, #156317, #156319	2025-06-23 02:57:50 +00:00
Xuehai Pan	5b210bb3a6	[BE][9/16] fix typos in torch/ (torch/csrc/) (#156319 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156319 Approved by: https://github.com/albanD ghstack dependencies: #156313, #156314, #156315, #156316, #156317	2025-06-23 02:57:50 +00:00
Xuehai Pan	ced90016c1	[BE][7/16] fix typos in torch/ (torch/csrc/) (#156317 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156317 Approved by: https://github.com/albanD ghstack dependencies: #156313, #156314, #156315, #156316	2025-06-23 02:57:41 +00:00
Xuehai Pan	cec2977ed2	[BE][6/16] fix typos in torch/ (#156316 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156316 Approved by: https://github.com/albanD ghstack dependencies: #156313, #156314, #156315	2025-06-23 02:57:34 +00:00
Xuehai Pan	4ccc0381de	[BE][5/16] fix typos in torch/ (torch/distributed/) (#156315 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156315 Approved by: https://github.com/Skylion007, https://github.com/albanD ghstack dependencies: #156313, #156314	2025-06-23 02:57:28 +00:00
Xuehai Pan	1b2146fc6d	[BE][4/16] fix typos in torch/ (torch/_dynamo/) (#156314 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156314 Approved by: https://github.com/jingsh ghstack dependencies: #156313	2025-06-23 02:57:19 +00:00
Xuehai Pan	6ff6630375	[BE][3/16] fix typos in torch/ (torch/_inductor/) (#156313 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156313 Approved by: https://github.com/jingsh	2025-06-23 02:57:12 +00:00
leslie-fang-intel	c55eef79f8	[Inductor][CPP] Enable a config to use a small dequant buffer for woq int4 (#156395 ) Summary Add a configuration option to enable a smaller dequantization buffer for WOQ INT4 CPP GEMM template. This can improve the performance of the WOQ INT4 GEMM template in cases where M is small. In such scenarios, matrix B cannot be effectively reused across matrix A, and we found that reducing the Kc block size can lead to better performance. Test Plan ``` python test/inductor/test_cpu_select_algorithm.py -k test_int4_woq_mm_with_small_buffer_config ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156395 Approved by: https://github.com/jansel ghstack dependencies: #156407, #156387	2025-06-23 02:00:42 +00:00
leslie-fang-intel	3c7079959c	[Inductor][CPP] Enable WOQ int4 concat linear (#156387 ) Summary Enable the concat linear optimization pass in Inductor for woq int4 linear. Test Plan ``` python test/inductor/test_cpu_select_algorithm.py -k test_int4_concat_woq_mm ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156387 Approved by: https://github.com/CaoE, https://github.com/jansel ghstack dependencies: #156407	2025-06-23 01:52:00 +00:00
Jack Taylor	03023f178c	FlexAttn config refactor + ROCm optimisations (#156307 ) This PR primarily unifies the flex attention config logic with the GEMM/Conv config approach https://github.com/pytorch/pytorch/pull/147452 this will make it much easier to handle optimisation pathways for particular triton backends. This PR also introduces: 1. Introduces an exhaustive tuning mode for flex attention via TORCHINDUCTOR_MAX_AUTOTUNE_FLEX_SEARCH_SPACE="EXHAUSTIVE" to allow for wide scale benchmarking for perf investigation use cases. 3. Updates configs for ROCm flex autotune path providing perf optimisations AMD perf numbers on score mod benchmark (default inputs) flex_attn \| mode \| Speedup (Avg) \| Speedup (Max) -- \| -- \| -- \| -- fwd \| autotune before PR \| 2.608 \| 20.56 fwd \| autotune after PR \| 2.862 \| 22 fwd \| exhaustive_autotune \| 2.943 \| 22.471 bwd \| autotune before PR \| 2.196 \| 9.831 bwd \| autotune after PR \| 2.423 \| 11.331 bwd \| exhaustive_autotune \| 2.566 \| 13.87 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156307 Approved by: https://github.com/drisspg, https://github.com/jansel	2025-06-22 22:27:38 +00:00
Varun Thumbe	a5cbb2bcb3	Improve All to All Perf for inter-node use-case (#156376 ) (#156389 ) Summary: For 16 GPU use-case. NVSHMEM can drive only upto 49GB/s with 8 thread blocks per peer for all to all V use-case. Increasing that to 16 threads per block is able to max out the perf. Test Plan: Verify on two hosts Host1: TORCH_SYMMMEM=NVSHMEM torchrun --nnodes=2 --nproc_per_node=8 --master_addr ${master_ip} --node_rank=0 comms.py -- master-ip ${master_ip} --b 4 --e 256M --n 500 --f 2 --z 1 --collective all_to_allv --backend nccl --device cuda Host2: TORCH_SYMMMEM=NVSHMEM torchrun --nnodes=2 --nproc_per_node=8 --master_addr ${master_ip} --node_rank=1 comms.py -- master-ip ${master_ip} --b 4 --e 256M --n 100 --f 2 --z 1 --collective all_to_allv --backend nccl --device cuda Rollback Plan: Differential Revision: D76937048 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156389 Approved by: https://github.com/kwen2501	2025-06-22 20:45:46 +00:00
FFFrog	a28e6ae38f	[OpenReg][2/N] Migrate cpp_extensions_open_device_registration to OpenReg (#156401 ) As the title stated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156401 Approved by: https://github.com/albanD ghstack dependencies: #156400	2025-06-22 18:40:38 +00:00
FFFrog	1d522325b4	[OpenReg][1/N] Migrate cpp_extensions_open_device_registration to OpenReg (#156400 ) As the title stated. Changes: - add resize_ for OpenReg - migrate related tests into test_openreg.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/156400 Approved by: https://github.com/albanD	2025-06-22 18:40:38 +00:00
Aaron Orenstein	54b8087f63	Improve torch.ops typing (#154555 ) Summary: Cloned https://github.com/pytorch/pytorch/pull/153558 from benjaminglass1 and fixed internal typing errors. Fixes longstanding issue where direct references to aten operations are seen as untyped by type checkers. This is accomplished by setting attributes on several classes more consistently, so that `__getattr__` can return a single type in all other cases. Decisions made along the way: 1. `torch.ops.higher_order` is now implemented by a single-purpose class. This was effectively true before, but the class implementing it attempted to be generalized unnecessarily. Fixing this simplified typing for the `_Ops` class. 2. `__getattr__` is only called when all other lookup methods have failed, so several constant special-cases in the function could be implemented as class variables. The remainder of this PR is fixing up all the bugs exposed by the updated typing, as well as all the nitpicky typing issues. Test Plan: CI Differential Revision: D75497142 Co-authored-by: Benjamin Glass <bglass@quansight.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/154555 Approved by: https://github.com/Skylion007, https://github.com/malfet, https://github.com/zou3519, https://github.com/benjaminglass1	2025-06-22 15:52:27 +00:00
James Wu	10fb98a004	[Precompile] Hook up backend="inductor" (#155387 ) This PR adds the necessary things to register and record backend ids from BundledAOTAutogradCacheEntry. One TODO to point out; in this diff, if there are multiple backends that would have the same AOTAutogradCache key (traditional cache key, not backend_id), we just end up serializing the same BundledAOTAutogradCache entry multiple times. This is not ideal obviously, so we'll want to deduplicate these and just track the different keys that one BundledAOTAutogradCacheEntry is associated with instead. This shouldn't be super hard to do, though, as we just need to run a deduplication step on call to `serialize()`, I think. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155387 Approved by: https://github.com/oulgen	2025-06-22 15:05:08 +00:00
PyTorch MergeBot	b5c8b8d09f	Revert "[dynamo] control one_graph behavior additionally through config (#154283 )" This reverts commit b46eb1ccaff944cdcd43e9ce3958819226d2952f. Reverted https://github.com/pytorch/pytorch/pull/154283 on behalf of https://github.com/ezyang due to All of this is responsible for regression, see https://github.com/pytorch/pytorch/pull/156561 ([comment](https://github.com/pytorch/pytorch/pull/154283#issuecomment-2994242583))	2025-06-22 14:22:07 +00:00
PyTorch MergeBot	5e56db59d4	Revert "[dynamo] add set_fullgraph decorator/context manager (#154289 )" This reverts commit 2c372a0502578e0136a84423c3f49c19c26d6bb7. Reverted https://github.com/pytorch/pytorch/pull/154289 on behalf of https://github.com/ezyang due to All of this is responsible for regression, see https://github.com/pytorch/pytorch/pull/156561 ([comment](https://github.com/pytorch/pytorch/pull/154283#issuecomment-2994242583))	2025-06-22 14:22:07 +00:00
PyTorch MergeBot	c10eeb5bad	Revert "[dynamo] fix set_fullgraph for nested calls (#154782 )" This reverts commit 537b0877a87948bc221301a518fdbc1cf772bc7e. Reverted https://github.com/pytorch/pytorch/pull/154782 on behalf of https://github.com/ezyang due to All of this is responsible for regression, see https://github.com/pytorch/pytorch/pull/156561 ([comment](https://github.com/pytorch/pytorch/pull/154283#issuecomment-2994242583))	2025-06-22 14:22:07 +00:00
PyTorch MergeBot	ee3d9969cc	Revert "[dynamo] handle fullgraph toggle using nested torch.compile (#155166 )" This reverts commit 24dc33b37b50ec92da08fc693dd83e7c87b74f8b. Reverted https://github.com/pytorch/pytorch/pull/155166 on behalf of https://github.com/ezyang due to All of this is responsible for regression, see https://github.com/pytorch/pytorch/pull/156561 ([comment](https://github.com/pytorch/pytorch/pull/154283#issuecomment-2994242583))	2025-06-22 14:22:07 +00:00
PyTorch MergeBot	f1331f3f1b	Revert "[BE][3/16] fix typos in torch/ (torch/_inductor/) (#156313 )" This reverts commit 3627270bdf17b0fb6f528ca1cb87d6f2ec32680a. Reverted https://github.com/pytorch/pytorch/pull/156313 on behalf of https://github.com/atalman due to export/test_torchbind.py::TestCompileTorchbind::test_compile_error_on_input_aliasing_contents_backend_aot_eager [GH job link](https://github.com/pytorch/pytorch/actions/runs/15804799771/job/44548489912) [HUD commit link](`c95f7fa874`) ([comment](https://github.com/pytorch/pytorch/pull/156313#issuecomment-2994171213))	2025-06-22 12:31:57 +00:00
PyTorch MergeBot	5b427c92a8	Revert "[BE][4/16] fix typos in torch/ (torch/_dynamo/) (#156314 )" This reverts commit ead741c5fb0036e0fc95b79d4fe1af3a426e1306. Reverted https://github.com/pytorch/pytorch/pull/156314 on behalf of https://github.com/atalman due to export/test_torchbind.py::TestCompileTorchbind::test_compile_error_on_input_aliasing_contents_backend_aot_eager [GH job link](https://github.com/pytorch/pytorch/actions/runs/15804799771/job/44548489912) [HUD commit link](`c95f7fa874`) ([comment](https://github.com/pytorch/pytorch/pull/156313#issuecomment-2994171213))	2025-06-22 12:31:57 +00:00
PyTorch MergeBot	145d4cdc11	Revert "[BE][5/16] fix typos in torch/ (torch/distributed/) (#156315 )" This reverts commit c2f0292bd5b4b3206f5b295e96f81cd6c178eb18. Reverted https://github.com/pytorch/pytorch/pull/156315 on behalf of https://github.com/atalman due to export/test_torchbind.py::TestCompileTorchbind::test_compile_error_on_input_aliasing_contents_backend_aot_eager [GH job link](https://github.com/pytorch/pytorch/actions/runs/15804799771/job/44548489912) [HUD commit link](`c95f7fa874`) ([comment](https://github.com/pytorch/pytorch/pull/156313#issuecomment-2994171213))	2025-06-22 12:31:57 +00:00
PyTorch MergeBot	3f44fdc03d	Revert "[BE][6/16] fix typos in torch/ (#156316 )" This reverts commit b210cf1ea56bcd9f937a2805d9e70d8684d25ee4. Reverted https://github.com/pytorch/pytorch/pull/156316 on behalf of https://github.com/atalman due to export/test_torchbind.py::TestCompileTorchbind::test_compile_error_on_input_aliasing_contents_backend_aot_eager [GH job link](https://github.com/pytorch/pytorch/actions/runs/15804799771/job/44548489912) [HUD commit link](`c95f7fa874`) ([comment](https://github.com/pytorch/pytorch/pull/156313#issuecomment-2994171213))	2025-06-22 12:31:57 +00:00
PyTorch MergeBot	035a68d25a	Revert "[BE][7/16] fix typos in torch/ (torch/csrc/) (#156317 )" This reverts commit ee72815f1180fe2d8bcdb23493999256169ac2fa. Reverted https://github.com/pytorch/pytorch/pull/156317 on behalf of https://github.com/atalman due to export/test_torchbind.py::TestCompileTorchbind::test_compile_error_on_input_aliasing_contents_backend_aot_eager [GH job link](https://github.com/pytorch/pytorch/actions/runs/15804799771/job/44548489912) [HUD commit link](`c95f7fa874`) ([comment](https://github.com/pytorch/pytorch/pull/156313#issuecomment-2994171213))	2025-06-22 12:31:56 +00:00
PyTorch MergeBot	1d3bca40ed	Revert "[BE][9/16] fix typos in torch/ (torch/csrc/) (#156319 )" This reverts commit a23ccaa8479e038e79532759a64e9947c0fac43d. Reverted https://github.com/pytorch/pytorch/pull/156319 on behalf of https://github.com/atalman due to export/test_torchbind.py::TestCompileTorchbind::test_compile_error_on_input_aliasing_contents_backend_aot_eager [GH job link](https://github.com/pytorch/pytorch/actions/runs/15804799771/job/44548489912) [HUD commit link](`c95f7fa874`) ([comment](https://github.com/pytorch/pytorch/pull/156313#issuecomment-2994171213))	2025-06-22 12:31:56 +00:00
PyTorch MergeBot	4b55871e06	Revert "[BE][11/16] fix typos in torch/ (torch/csrc/distributed/) (#156321 )" This reverts commit c95f7fa874a3116f1067f9092456ee7281003614. Reverted https://github.com/pytorch/pytorch/pull/156321 on behalf of https://github.com/atalman due to export/test_torchbind.py::TestCompileTorchbind::test_compile_error_on_input_aliasing_contents_backend_aot_eager [GH job link](https://github.com/pytorch/pytorch/actions/runs/15804799771/job/44548489912) [HUD commit link](`c95f7fa874`) ([comment](https://github.com/pytorch/pytorch/pull/156321#issuecomment-2994163667))	2025-06-22 12:27:36 +00:00
Sidharth	afbf5420b8	[dynamo] fixes to lru_cache message and adding user stack trace in debug mode (#156463 ) This PR refers to the issue: https://github.com/pytorch/pytorch/issues/155352 This PR uses torch._dynamo.utils.warn_once so that this warning only emits once, clarifies in the warning that silent incorrectness is potential, not observed, Doesn't warn for functions that come from torch.* As of right now with this code change the terminal outputs: if the code came from torch.* : Nothing, as we shouldn't warn for functions that come from torch.* else: /data/users/ssubbarao8/pytorch/torch/_dynamo/variables/functions.py:1565: UserWarning: Dynamo detected a call to a `functools.lru_cache`-wrapped function. Dynamo ignores the cache wrapper and directly traces the wrapped function. Silent incorrectness is only a potential risk, not something we have observed. Enable TORCH_LOGS="+dynamo" for a DEBUG stack trace. torch._dynamo.utils.warn_once(msg) If the user runs the command 'TORCH_LOGS="+dynamo" python foo4.py', in the debug logs it shows(this log below is based on chillee's repro: /data/users/ssubbarao8/pytorch/torch/_dynamo/variables/functions.py:1565: UserWarning: Dynamo detected a call to a `functools.lru_cache`-wrapped function. Dynamo ignores the cache wrapper and directly traces the wrapped function. Silent incorrectness is only a potential risk, not something we have observed. Enable TORCH_LOGS="+dynamo" for a DEBUG stack trace. torch._dynamo.utils.warn_once(msg) V0619 21:00:16.504000 956424 torch/_dynamo/variables/functions.py:1575] [0/0] call to a lru_cache` wrapped function from user code at: /data/users/ssubbarao8/pytorch/foo4.py:9 V0619 21:00:16.504000 956424 torch/_dynamo/variables/functions.py:1575] [0/0] File "/data/users/ssubbarao8/pytorch/foo4.py", line 9, in <module> V0619 21:00:16.504000 956424 torch/_dynamo/variables/functions.py:1575] [0/0] torch.compile(foo, backend="eager")(torch.randn(4)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156463 Approved by: https://github.com/williamwen42	2025-06-22 11:40:28 +00:00
Sidharth	aeaf6b59e2	[dynamo] Weblink generation when unimplemented_v2() is called (#156033 ) This PR includes the GBID weblink whenever a user encounters a graph break. I also had to include the JSON file in setup.py, so it can be part of the files that are packaged in during CI. It also fixes the issue of the hardcoded error messages stripping away one of the '/' in 'https'. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156033 Approved by: https://github.com/williamwen42	2025-06-22 11:39:31 +00:00
Xuehai Pan	c95f7fa874	[BE][11/16] fix typos in torch/ (torch/csrc/distributed/) (#156321 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156321 Approved by: https://github.com/jingsh ghstack dependencies: #156313, #156314, #156315, #156316, #156317, #156319	2025-06-22 08:43:49 +00:00
Xuehai Pan	a23ccaa847	[BE][9/16] fix typos in torch/ (torch/csrc/) (#156319 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156319 Approved by: https://github.com/albanD ghstack dependencies: #156313, #156314, #156315, #156316, #156317	2025-06-22 08:43:49 +00:00
Xuehai Pan	ee72815f11	[BE][7/16] fix typos in torch/ (torch/csrc/) (#156317 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156317 Approved by: https://github.com/albanD ghstack dependencies: #156313, #156314, #156315, #156316	2025-06-22 08:43:41 +00:00
Xuehai Pan	b210cf1ea5	[BE][6/16] fix typos in torch/ (#156316 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156316 Approved by: https://github.com/albanD ghstack dependencies: #156313, #156314, #156315	2025-06-22 08:43:33 +00:00
Xuehai Pan	c2f0292bd5	[BE][5/16] fix typos in torch/ (torch/distributed/) (#156315 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156315 Approved by: https://github.com/Skylion007, https://github.com/albanD ghstack dependencies: #156313, #156314	2025-06-22 08:43:26 +00:00
Xuehai Pan	ead741c5fb	[BE][4/16] fix typos in torch/ (torch/_dynamo/) (#156314 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156314 Approved by: https://github.com/jingsh ghstack dependencies: #156313	2025-06-22 08:43:18 +00:00
Xuehai Pan	3627270bdf	[BE][3/16] fix typos in torch/ (torch/_inductor/) (#156313 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156313 Approved by: https://github.com/jingsh	2025-06-22 08:43:09 +00:00
cyy	08dae945ae	Use official CUDAToolkit module in CMake (#154595 ) Use CUDA language in CMake and remove forked FindCUDAToolkit.cmake. Some CUDA targets are also renamed with `torch::` prefix. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154595 Approved by: https://github.com/albanD	2025-06-22 05:44:29 +00:00
Edward Yang	1d993fa309	Don't change set_skip_guard_eval_unsafe for DisableContext, since compiler won't run (#156490 ) Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156490 Approved by: https://github.com/anijain2305	2025-06-22 00:51:32 +00:00
Edward Yang	333e0e6147	Make build-deps drop builds into current venv again (#156200 ) Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156200 Approved by: https://github.com/malfet	2025-06-22 00:45:02 +00:00
Laith Sakka	74ebd8d14e	use guard_or_false for expand utils reduction (#155868 ) This is classic broadcast like pattern. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155868 Approved by: https://github.com/bobrenjc93	2025-06-21 23:42:19 +00:00
Syed Tousif Ahmed	f70c80105e	Enables NCCL symmetric memory kernels through mempool registration (#155134 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155134 Approved by: https://github.com/kwen2501 Co-authored-by: Ke Wen <kw2501@meta.com>	2025-06-21 23:24:04 +00:00
Isalia20	9e132b770e	[CUDA] Skip test on low vram machines (#156548 ) I noticed some jobs error out after merging #155397 due to the test requiring >15GB GPU memory to execute and some of the machines it's running on has 8GB GPUs. This PR adds the skip option on those machines. CC: @eqy @ngimel Pull Request resolved: https://github.com/pytorch/pytorch/pull/156548 Approved by: https://github.com/eqy, https://github.com/malfet	2025-06-21 22:32:57 +00:00
codingwithsurya	e4ae60a413	[SymmMem] Add NVSHMEM Quiet support to Triton (#156475 ) This PR introduces device-side NVSHMEM completion guarantees via the quiet API in Triton, enabling GPU kernels to ensure all pending remote memory operations are fully complete before proceeding with subsequent operations. Changes: - Added a new `core.extern` wrapper for `nvshmem_quiet` in `nvshmem_triton.py` - Implemented `test_triton_quiet` in `test/distributed/test_nvshmem.py`, including: - A Triton kernel that performs `putmem_block` followed by `quiet()` to ensure completion - Flag-based signaling only after `quiet()` completes, guaranteeing data delivery - Consumer validation that when the completion flag arrives, all data transfers are guaranteed complete Tests: `$ TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_quiet` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156475 Approved by: https://github.com/kwen2501 ghstack dependencies: #156472, #156473, #156474	2025-06-21 22:19:58 +00:00
Xuan Zhang	c2d1b225e6	[PT2][partitioners] raise getitems in partitioners to allow earlier release of buffers (#155809 ) Problem & Solution: Assume we have something like: ``` x = some_op(...) x0 = x[0] do_something_with_and_is_last_use_of(x0) do_a_bunch_of_other_things() x1 = x[1] ``` In this case, the memory associated with `x0` cannot be released until `x1 = x[1]`. Since `x1 = x[1]` does not use additional memory, it would be beneficial to move and `x1 = x[1]` and all such `getitem` operations to be immediately after `x = some_op(...)` such as ``` x = some_op(...) x0 = x[0] x1 = x[1] do_something_with_and_is_last_use_of(x0) do_a_bunch_of_other_things() ``` Results: For instance, for the `res2net101_26w_4s` model in pytorch benchmark, when running with `aot_eager` backend and with `activation_memory_budget=0.4`, the peak memory are * baseline: 7.73GiB * with the chage: 6.45GiB As a sanity check, for the same setting with `inductor` backend, the peak memory is not regressed. cc and credit to @ShatianWang for noticing this issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155809 Approved by: https://github.com/fmassa, https://github.com/bdhirsh	2025-06-21 19:57:21 +00:00
codingwithsurya	04b91a9e43	[SymmMem] Add NVSHMEM Fence support to Triton (#156474 ) This PR introduces device-side NVSHMEM memory ordering via the fence API in Triton, enabling GPU kernels to enforce completion and ordering of remote memory operations before subsequent operations proceed. Changes: - Added a new `core.extern` wrapper for `nvshmem_fence` in `nvshmem_triton.py` - Implemented `test_triton_fence` in `test/distributed/test_nvshmem.py`, including: - A Triton kernel that performs two ordered `putmem_block` operations separated by `fence()` calls - Final fence before flag update to ensure all data transfers complete before signaling - Consumer validation that both buffers contain expected values when flag arrives, proving ordering guarantees Tests: `$ TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_fence` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156474 Approved by: https://github.com/mandroid6, https://github.com/kwen2501 ghstack dependencies: #156472, #156473	2025-06-21 18:57:05 +00:00
Simon Fan	c06c2569ee	[ca] Support TorchDispatchMode via pass through (#156516 ) The CA initial trace just proxies nodes without dispatching any ops, we should hide it from ambient TorchDispatchModes In terms of differences with eager autograd engine: - For function mode, CA additionally disables/re-enables `_set_multithreading_enabled` - For dispatch mode: - accumulate grad doesn't go down the stealing path (inaccurate compile-time refcount) so the grad `detach` ops are `copy_` instead - Since we always initial trace with dynamic shapes, and we filter out sizes, there's 1 aten.empty.memory_format for each mark_dynamic'd scalar Pull Request resolved: https://github.com/pytorch/pytorch/pull/156516 Approved by: https://github.com/jansel ghstack dependencies: #156374, #156509	2025-06-21 18:33:47 +00:00
Simon Fan	5f2f343e1e	[ca] suggest to disable compiled autograd for trace-time NotImplementedErrors (#156509 ) Example: ```python File "/home/xmfan/core/a/pytorch/torch/autograd/graph.py", line 829, in _engine_run_backward return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ NotImplementedError: TorchDispatchMode not yet implemented for compiled autograd. You can disable compiled autograd for this operation by: 1. Relocating the unsupported autograd call outside the compiled region. 2. Wrapping the unsupported autograd call within a scope that disables compiled autograd. 3. Configuring the specific compilation unit to disable compiled autograd. 4. Globally disabling compiled autograd at the application's initialization. ``` No duplicate error messages for python side trace-time errors ```python ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/xmfan/core/a/pytorch/torch/_dynamo/compiled_autograd.py", line 344, in begin_capture raise NotImplementedError( NotImplementedError: Found tensor of type <class 'torch.nn.utils._expanded_weights.expanded_weights_impl.ExpandedWeight'>, which is not supported by FakeTensorMode. You can turn off compiled autograd by either: 1. Moving the unsupported autograd call outside of the torch.compile'd region. 2. Wrapping the unsupported autograd call in the torch._dynamo.compiled_autograd._disable() context manager. 3. Setting torch._dynamo.config.compiled_autograd=False for the torch.compile call containing the unsupported autograd call. 4. Setting torch._dynamo.config.compiled_autograd=False at the start of the program. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156509 Approved by: https://github.com/jansel ghstack dependencies: #156374	2025-06-21 18:33:46 +00:00
Simon Fan	f1968a5e76	[ca] skip on some PYTORCH_TEST_WITH_DYNAMO=1 autograd tests (#156374 ) These aren't supported. Not sure how they passed CI Pull Request resolved: https://github.com/pytorch/pytorch/pull/156374 Approved by: https://github.com/jansel	2025-06-21 18:33:38 +00:00
Animesh Jain	fab85fc5f9	[compile][hierarchical compilation] Release nested_compile_region API (#156449 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156449 Approved by: https://github.com/zou3519, https://github.com/jansel	2025-06-21 15:14:59 +00:00
Sam Larsen	fb75dea2c1	[logging] dynamo_timed for CachingAutotuner.coordinate_descent_tuning (#156517 ) Summary: Discussed internally at https://fburl.com/workplace/v3hllrs9. With coordinate descent tuning enabled, we're missing the dynamo_timed logging. Test Plan: `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 TORCHINDUCTOR_COORDINATE_DESCENT_TUNING=1 buck run mode/opt caffe2/benchmarks/dynamo:torchbench -- --training --backend=inductor --only nanogpt --repeat 1 --performance --cold-start-latency` * tlparse: https://fburl.com/bh2hxw4z * dynamo_compile: https://fburl.com/scuba/dynamo_compile/sandbox/u88ogw39 * pt2_compile_events: https://fburl.com/scuba/pt2_compile_events/yqljow6c Rollback Plan: Differential Revision: D77053918 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156517 Approved by: https://github.com/mengluy0125	2025-06-21 14:17:19 +00:00
atalman	a47ca4fc74	Revert "[dynamo] Weblink generation when unimplemented_v2() is called (#156033 )" (#156546 ) Broke multiple CI jobs: dynamo/test_reorder_logs.py::ReorderLogsTests::test_constant_mutation [GH job link](https://github.com/pytorch/pytorch/actions/runs/15792695433/job/44521220864) [HUD commit link](`9de23d0c29`) This reverts commit 9de23d0c29dfac8dc0f6f234bdbcd85a6375fa81. PyTorch bot revert failed: https://github.com/pytorch/pytorch/pull/156033 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156546 Approved by: https://github.com/jansel	2025-06-21 14:10:12 +00:00
PyTorch MergeBot	d846e21355	Revert "[nativert] move layout planner algorithms to libtorch (#156508 )" This reverts commit eab45643f22e58ee12d95d8b0162d51ca0a50801. Reverted https://github.com/pytorch/pytorch/pull/156508 on behalf of https://github.com/atalman due to [GH job link](https://github.com/pytorch/pytorch/actions/runs/15793524714/job/44524067679) [HUD commit link](`eab45643f2`) ([comment](https://github.com/pytorch/pytorch/pull/156508#issuecomment-2993589983))	2025-06-21 13:42:40 +00:00
Isalia20	1cfdcb975a	[CUDA] fix illegal memory access in attention (#155397 ) Fixes https://github.com/pytorch/pytorch/issues/150054 CI seemed to be messed up in the old one, old PR: https://github.com/pytorch/pytorch/pull/155145 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155397 Approved by: https://github.com/ngimel	2025-06-21 12:32:00 +00:00
fduwjj	cd75cf3cab	[symm_mem] Add one side put API for nvshvem (#156443 ) `nvshmem_put(Tensor tensor, int peer)`, where `tensor` must be a symmetric tensor, i.e. rendezvoused before this call. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156443 Approved by: https://github.com/kwen2501 Co-authored-by: Ke Wen <kw2501@meta.com>	2025-06-21 12:16:36 +00:00
codingwithsurya	4ff0e033c1	[SymmMem] Add NVSHMEM signal_wait_until support to Triton (#156473 ) This PR introduces device-side NVSHMEM signal synchronization via the signal_wait_until API in Triton, enabling GPU kernels to block until a signal variable meets a specified condition. This replaces previous barrier-based synchronization patterns with more efficient signal-based coordination between PEs. Changes: - Added a new `core.extern` wrapper for `nvshmem_signal_wait_until` in `nvshmem_triton.py` - Updated existing `test_triton_put_signal` and `test_triton_put_signal_add` tests to use `signal_wait_until` instead of `dist.barrier()` for proper device-side synchronization ([per feedback](https://github.com/pytorch/pytorch/pull/156211#discussion_r2153035675)) - Implemented `test_triton_signal_wait_until` with: - Producer-consumer pattern where Rank 0 puts data and signals completion via `putmem_signal_block` - Consumer (Rank 1) uses `signal_wait_until` to block until the signal variable reaches the expected value - End-to-end validation of both data transfer and signal synchronization Tests: `$ TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_signal_wait_until` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156473 Approved by: https://github.com/kwen2501, https://github.com/mandroid6 ghstack dependencies: #156472	2025-06-21 10:55:40 +00:00
Laith Sakka	8485f19507	remove gso from vector_norm (#156530 ) guard_or_false here does same thing that guard_size_oblivuous do, note that size is >=0 and this is size like by definition since its a tensor size Pull Request resolved: https://github.com/pytorch/pytorch/pull/156530 Approved by: https://github.com/bobrenjc93	2025-06-21 08:42:36 +00:00
sanchitintel	6ffa03ef9e	[Inductor-CPU] int8 WoQ concat linear (#153004 ) ### Summary int8 WoQ GEMM concat linear optimization pertaining to the same activation applied to 3 sets of weights of the same shape. ### Perf data GPT-J 128 input tokens, 128 output tokens. 32 physical cores of one socket of Intel(R) Xeon(R) 6972P (Xeon Gen 5). tcmalloc & Intel OpenMP were preloaded. \| May 8 nightly first token latency \| First token latency with this implementation \| Rest token latency with May 8 nightly \| Rest token latency with this implementation combined with #149373 \| \|---\|---\|---\|---\| \|202 ms \| 190 ms \| 33 ms \| 30 ms\| Pull Request resolved: https://github.com/pytorch/pytorch/pull/153004 Approved by: https://github.com/leslie-fang-intel, https://github.com/chunyuan-w, https://github.com/jansel Co-authored-by: Anthony Shoumikhin <anthony@shoumikh.in>	2025-06-21 08:40:09 +00:00
Laith Sakka	35321b2ad6	remove make_fast_binary_impl from make_fast_binary_impl (#156528 ) This was added in https://github.com/pytorch/pytorch/pull/133584. Take slow path when we cant determine fast path is valid. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156528 Approved by: https://github.com/bobrenjc93	2025-06-21 08:27:54 +00:00
dolpm	eab45643f2	[nativert] move layout planner algorithms to libtorch (#156508 ) Summary: tt Test Plan: ci Rollback Plan: Differential Revision: D76832891 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156508 Approved by: https://github.com/zhxchen17	2025-06-21 07:35:40 +00:00
Scott Wolchok	bf50d71553	Add missing `inline namespace CPU_CAPABILITY` to Gelu/Elu.h (#156512 ) As I recently learned the hard way (#156243), it is necessary to put kernel code that uses Vectorized in headers in this namespace. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156512 Approved by: https://github.com/malfet	2025-06-21 06:26:23 +00:00
codingwithsurya	e3b44edfd8	[SymmMem] Add NVSHMEM wait_until support to Triton (#156472 ) This PR introduces device-side NVSHMEM synchronization via the wait_until API in Triton, enabling GPU kernels to block until a remote flag reaches a specified value. It also adds a corresponding end-to-end test to validate correct behavior across PEs. Changes: - Added a new `core.extern` wrapper for `nvshmem_longlong_wait_until` in `nvshmem_triton.py`. - Implemented `test_triton_wait_until` in `test/distributed/test_nvshmem.py`, including: - A simple Triton kernel that calls `nvshmem.wait_until` on a symmetric memory flag. - Coordination logic where Rank 0 blocks until Rank 1 atomically sets the flag and transfers data. Tests: `$ TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_wait_until` ```python @triton.jit def put_kernel(dst_ptr, src_ptr, numel: tl.constexpr, peer: tl.constexpr): nvshmem.putmem_block(dst_ptr, src_ptr, numel, peer) @triton.jit def wait_until_kernel(ivar_ptr, cmp_op: tl.constexpr, cmp_val: tl.constexpr): nvshmem.wait_until(ivar_ptr, cmp_op, cmp_val) ... if rank == 0: print(f"[RANK 0] About to call wait_until_kernel - this will BLOCK until rank 1 sets flag to 21") wait_until_kernel[(1, 1, 1)](ivar_ptr, cmp_op=NVSHMEM_CMP_EQ, cmp_val=flag_val, extern_libs=nvshmem_lib) print(f"[RANK 0] WAIT IS OVER! Flag was set, checking data now...") print(f"[RANK 0] Current out buffer contents: {out.tolist()}") torch.testing.assert_close(out, val * torch.ones(numel, dtype=dtype, device=self.device)) print(f"[RANK 0] ✓ DATA VERIFICATION PASSED! Got expected values.") if rank == 1: print(f"[RANK 1] About to PUT 8 elements of value 13 to rank 0") put_kernel[(1, 1, 1)](dst_ptr, src_ptr, numel=numel, peer=peer, extern_libs=nvshmem_lib) print(f"[RANK 1] About to PUT flag value 21 to wake up rank 0") put_kernel[(1, 1, 1)](dst_ptr, src_ptr, numel=1, peer=peer, extern_libs=nvshmem_lib) print(f"[RANK 1] FLAG PUT complete! Rank 0 should wake up now.") ... ``` Output: ``` [RANK 0] About to call wait_until_kernel - this will BLOCK until rank 1 sets flag to 21 [RANK 1] About to PUT 8 elements of value 13 to rank 0 [RANK 1] About to PUT flag value 21 to wake up rank 0 [RANK 1] FLAG PUT complete! Rank 0 should wake up now. [RANK 0] WAIT IS OVER! Flag was set, checking data now... [RANK 0] Current out buffer contents: [13, 13, 13, 13, 13, 13, 13, 13] [RANK 0] ✓ DATA VERIFICATION PASSED! Got expected values. [RANK 0] Test completed successfully! 🎉 [RANK 1] Test completed successfully! 🎉 ... ---------------------------------------------------------------------- Ran 1 test in 18.773s OK ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156472 Approved by: https://github.com/kwen2501	2025-06-21 06:18:31 +00:00
Pian Pawakapan	92c79f36db	[PGO] frame-specific whitelist logging (#155959 ) Summary: In D75617963, we started logging dynamic whitelist suggestions to PT2 Compile Events. The whitelists were aggregated across all frames, intending to avoid manual work for the user (e.g. if frame 0/1 saw L['x'] turn dynamic, and later 1/1 saw L['y'], we'd log "L['x'],L['y']" on frame 1/1). This switches to frame-specific whitelists, as attributing dynamism changes to certain frames was difficult, and suggestions are sometimes polluted by problematic frames (e.g. optimizer states). The globally aggregated whitelist is still available in tlparse, by looking at the final `put_local_code_state_*` entry. Test Plan: loggercli codegen GeneratedPt2CompileEventsLoggerConfig Rollback Plan: Differential Revision: D76628834 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155959 Approved by: https://github.com/bobrenjc93	2025-06-21 06:15:51 +00:00
Sidharth	9de23d0c29	[dynamo] Weblink generation when unimplemented_v2() is called (#156033 ) This PR includes the GBID weblink whenever a user encounters a graph break. I also had to include the JSON file in setup.py, so it can be part of the files that are packaged in during CI. It also fixes the issue of the hardcoded error messages stripping away one of the '/' in 'https'. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156033 Approved by: https://github.com/williamwen42	2025-06-21 05:47:54 +00:00
Aby Mathew C	b8ace6f951	Make dtensor tests device agnostic (#155687 ) ## MOTIVATION This PR is a continuation of https://github.com/pytorch/pytorch/pull/154840 and we are trying to make the tests more device agnostic by removing hard coded references to any particular device. Please refer to this RFC as well: https://github.com/pytorch/rfcs/pull/66 ## CHANGES 1. test_convolution_ops.py: - Replace "cuda" with self.device_type 2. test_random_ops.py: - Remove setting and using TYPE_DEVICE variable since device_type is set as per the environment (device) in DTensorTestBase class. - Replace "cuda" with self.device_type Pull Request resolved: https://github.com/pytorch/pytorch/pull/155687 Approved by: https://github.com/EikanWang, https://github.com/d4l3k	2025-06-21 04:51:59 +00:00
anwang	f3ec16c26a	[MTIA Aten Backend][3/n] Migrate mm.out from out-of-tree to in-tree (#154393 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate mm.out from out-of-tree to in-tree. We dispatch mm.out to MTIA separately from CPU/CUDA. So this diff adds the file `MTIAOps.cpp` under `ATen/native/mtia` to hold the dispatched functions. In future we can split `MTIAOps.cpp` to categorized ops files. Differential Revision: [D74743849](https://our.internmc.facebook.com/intern/diff/D74743849/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154393 Approved by: https://github.com/albanD, https://github.com/egienvalue, https://github.com/nautsimon	2025-06-21 04:31:04 +00:00
drisspg	88b9c285e0	Workaround for e4m2 dtype (#156461 ) Found in: https://github.com/pytorch/ao/pull/2408 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156461 Approved by: https://github.com/vkuzo	2025-06-21 04:01:44 +00:00
soulitzer	554b568040	Add internal use only utility to allow externally visible side effects within HOPs (#155715 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155715 Approved by: https://github.com/zou3519	2025-06-21 03:55:28 +00:00
Arsh Zahed	c09b054878	Add runtime profiler info for AOTDispatcher prologue (#155785 ) Fixes #155721 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155785 Approved by: https://github.com/bdhirsh	2025-06-21 03:34:07 +00:00
fduwjj	fd8ea3c8a3	[symm_mem] Add nccl as a backend for symmetric memory (#155740 ) Running unit test: TORCH_SYMMMEM=NCCL TORCH_DISTRIBUTED_DEBUG=INFO TORCH_CPP_LOG_LEVEL=INFO pytest test/distributed/test_nccl.py -k test_nccl_symmem_alloc Pull Request resolved: https://github.com/pytorch/pytorch/pull/155740 Approved by: https://github.com/kwen2501	2025-06-21 03:22:23 +00:00
Nikita Shulga	ee56e9f8a8	[BE] Make Eigen an optional dependency (#155955 ) Whose version is controlled by `eigen_pin.txt`, but which will be installed only if BLAS providers could not be found. Why this is good for CI: we don't really build with Eigen ever and gitlab can be down when github is up, which causes spurious CI failures in the past, for example. Remove eigen submodule and replace it with eigen_pin.txt Fixes https://github.com/pytorch/pytorch/issues/108773 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155955 Approved by: https://github.com/atalman	2025-06-21 03:02:02 +00:00
Xuehai Pan	b4228a94d1	Split the exclude pattern for `CODESPELL` linter (#156229 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156229 Approved by: https://github.com/albanD ghstack dependencies: #156080, #156081	2025-06-21 02:47:40 +00:00
Xuehai Pan	e3507c3777	[BE] fix typos in functorch/ and scripts/ (#156081 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156081 Approved by: https://github.com/albanD ghstack dependencies: #156080	2025-06-21 02:47:40 +00:00
Xuehai Pan	2ccfd14e23	[BE] fix typos in docs/ (#156080 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156080 Approved by: https://github.com/cyyever, https://github.com/albanD	2025-06-21 02:47:32 +00:00
clr	9aaa184105	dynamo: Don't crash when someone tries to access a non existent list member (#156335 ) dynamo: Don't crash when someone tries to access a non existent list member Test added which reproduces the failure. Note that I'm using the new unimplemented_v2 API. Let me know if people have a strong preference that I use something else. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156335 Approved by: https://github.com/jansel	2025-06-21 02:26:31 +00:00
Frank Lin	ac86ec0e60	[Draft][CUDA] Use runtime driver API for cuStreamWriteValue32 (#156097 ) Fixes #154073 Reference: https://github.com/NVIDIA/Fuser/pull/4197 See PR #154097 @nWEIdia is currently out of the office, so I’ve temporarily taken over his work. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156097 Approved by: https://github.com/ngimel Co-authored-by: Wei Wang <weiwan@nvidia.com>	2025-06-21 01:34:41 +00:00
Yiming Zhou	e98dd95446	[nativert] Move SerialGraphExecutor to PyTorch core (#156459 ) Summary: `SerialGraphExecutor` inherits from `GraphExecutorBase` and executes all nodes in the graph in a serial manner Test Plan: CI Rollback Plan: Differential Revision: D76917966 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156459 Approved by: https://github.com/zhxchen17, https://github.com/jingsh	2025-06-21 01:32:06 +00:00
bobrenjc93	a67eb1a0d6	[ez] remove unused functions (#156466 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156466 Approved by: https://github.com/jingsh	2025-06-21 00:38:34 +00:00
Animesh Jain	2ee23175d9	[dynamo][guards] Catch exception and return false in the backend match (#156341 ) Its difficult to write a test. I found this while debugging a sefgault. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156341 Approved by: https://github.com/williamwen42	2025-06-21 00:13:26 +00:00
Ke Wen	0f0c010714	[c10d] init_process_group supports index-only device id (#156214 ) Before: ``` acc = torch.accelerator.current_accelerator() if acc: local_idx = ... dist.init_process_group( device_id=torch.device(acc.type, local_idx) ) ``` After: ``` dist.init_process_group(device_id=local_idx) ``` That is, `init_process_group` checks `torch.accelerator.current_accelerator()` internally. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156214 Approved by: https://github.com/guangyey, https://github.com/albanD	2025-06-21 00:02:37 +00:00
Justin Chu	fbbab794ef	[ONNX] Implement Attention-23 (#156431 ) Implement Attention-23 using sdpa and flexattention. - I used copilot for this. - Also updated the conversion logic to remove trailing None inputs. @gramalingam @kunal-vaishnavi @titaiwangms Pull Request resolved: https://github.com/pytorch/pytorch/pull/156431 Approved by: https://github.com/titaiwangms Co-authored-by: kunal-vaishnavi <115581922+kunal-vaishnavi@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>	2025-06-20 23:54:57 +00:00
Menglu Yu	0ad88a2224	Support environement var for autotune log (#156254 ) Summary: Titled Test Plan: See the scadcastle signal Rollback Plan: Differential Revision: D76860928 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156254 Approved by: https://github.com/Mingming-Ding	2025-06-20 23:06:33 +00:00
Nicolas Macchioni	6098209bff	[BE][5/X] Phase out usage of use_max_autotune() (#156269 ) These look to be the last call sites using `use_max_autotune(...)`, so remove those and `use_max_autotune(...)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156269 Approved by: https://github.com/masnesral	2025-06-20 22:37:45 +00:00
Animesh Jain	5ab257c74c	[invoke_subgraph] Make invoke_subgraph cacheable (#156448 ) Its unclear to me what happens if the subgraph itself is not cacheable. Imo, there is nothing special about invoke_subgraph to prevent any caching. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156448 Approved by: https://github.com/oulgen, https://github.com/zou3519	2025-06-20 21:20:23 +00:00
Scott Wolchok	e2351f2dcf	fix apparent copy-paste bug in log_softmax reduced-precision fp kernel (#156379 ) This looks like a bug. Check if trying to fix it breaks existing tests; if not, will look into why no test coverage caught it Pull Request resolved: https://github.com/pytorch/pytorch/pull/156379 Approved by: https://github.com/janeyx99	2025-06-20 20:54:53 +00:00
Guilherme Leobas	b8fc5e0c0d	skip flaky test in CPython 3.13 tests (#155561 ) Changed files: * test_math.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/155561 Approved by: https://github.com/zou3519	2025-06-20 20:25:35 +00:00
PyTorch MergeBot	754c04aa06	Revert "[dynamo] raise hard error if error is encountered while tracing resume function prologue (#154564 )" This reverts commit 0aed855b2bde6d9bd045bb20cc24544a9f2fb72b. Reverted https://github.com/pytorch/pytorch/pull/154564 on behalf of https://github.com/ezyang due to regresses functorch_maml_omniglot ([comment](https://github.com/pytorch/pytorch/pull/154564#issuecomment-2992685744))	2025-06-20 20:18:24 +00:00
Bhagirath Mehta	de1930a429	Add ONNX dynamo metadata documentation (#155816 ) Describe auto-generated metadata when calling torch.onnx.export Pull Request resolved: https://github.com/pytorch/pytorch/pull/155816 Approved by: https://github.com/justinchuby, https://github.com/titaiwangms Co-authored-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>	2025-06-20 20:12:22 +00:00
bobrenjc93	a69e27ca5a	Remove unused MultiKernelCall import from inductor codegen (#156158 ) Since it's now actually used within async_compile.multi_kernel ``` def multi_kernel(self, args, kwargs) -> Any: from torch._inductor.codegen.multi_kernel import MultiKernelCall # no need to call this in parallel since the sub-kernels are already parallel tasks return MultiKernelCall(args, **kwargs) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156158 Approved by: https://github.com/jansel, https://github.com/shunting314	2025-06-20 19:55:24 +00:00
Shangdi Yu	e5ea24fb27	[nativert] Move auto_functionalize_kernel (#156454 ) Summary: Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 As part of the effort to open source TorchNativeRuntime (or what we call Sigmoid), we are moving the Pytree implementation to torch/: fbcode/sigmoid/kernels -> fbcode/caffe2/torch/nativert/kernels Copied from original auto_functionalize Diff Summary D53776805: This is a non-functional kernel implementation for auto_functionalize In AutoFunctionalizeKernel, I directly call the underlying target without making a clone of mutating inputs. This would mutates the input tensors inplace, which is unsafe in general. However, Sigmoid is not doing any graph optimization, or node reordering at the moment, so it's ok do take this short cut. In the proper functional implementation, it will make a clone of the mutating input tensor return these new instance of tensors as AutoFunctionalizeKernel output. If the original exported program has some "bufferMutation" or "userInputMutation" fields, it will also need to honor such mutations in Sigmoid. Test Plan: See internal for test plan Differential Revision: D76926383 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156454 Approved by: https://github.com/zhxchen17	2025-06-20 19:53:16 +00:00
Jane Xu	eb331b59fe	Add shim fallback for narrow (#156496 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156496 Approved by: https://github.com/albanD	2025-06-20 19:47:00 +00:00
Aleksandar Samardžić	6ed85bfe6a	Refine alignment check along dynamic dimension for grouped MMs (#155466 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155466 Approved by: https://github.com/ngimel	2025-06-20 19:42:57 +00:00
Nikita Shulga	ef6d2cee7a	[BE][MPS] Refactor core matmul logic into matmul_core (#155969 ) In preparation of adding integer addmm, move matmul computation part into matmul_inner function Change callstack from group_id, thread_id_in_group to thread_id, threadid_in_group, which eliminates the need of calculating the index Pull Request resolved: https://github.com/pytorch/pytorch/pull/155969 Approved by: https://github.com/Skylion007	2025-06-20 18:54:38 +00:00
sekyondaMeta	18e4c461fb	Update index.md (#155143 ) Related to: https://github.com/pytorch/pytorch/issues/152134 Update to index.md to add language for Stable and Unstable Pull Request resolved: https://github.com/pytorch/pytorch/pull/155143 Approved by: https://github.com/AlannaBurke, https://github.com/atalman Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-20 18:53:32 +00:00
Kevin Fu	502486d946	[PT2]Add weight and constant config path template (#156359 ) Summary: At title. Test Plan: N/A Rollback Plan: Differential Revision: D76925510 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156359 Approved by: https://github.com/SherlockNoMad	2025-06-20 18:46:01 +00:00
Jane Xu	4b6cbf528b	Add C shim fallback for fill_ (#156245 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156245 Approved by: https://github.com/desertfire	2025-06-20 18:45:48 +00:00
PyTorch MergeBot	208ec60e72	Revert "[BE] Make Eigen an optional dependency (#155955 )" This reverts commit 1b50c12584909bda00009f4f0fd0d38ec792d019. Reverted https://github.com/pytorch/pytorch/pull/155955 on behalf of https://github.com/atalman due to need to revert eigen test ([comment](https://github.com/pytorch/pytorch/pull/155955#issuecomment-2992512124))	2025-06-20 18:43:52 +00:00
PyTorch MergeBot	d309cd1d50	Revert "[BE][MPS] Refactor core matmul logic into matmul_core (#155969 )" This reverts commit 769d754ab2469813a3b790ec58c25c466099dd3d. Reverted https://github.com/pytorch/pytorch/pull/155969 on behalf of https://github.com/atalman due to need to revert eigen test ([comment](https://github.com/pytorch/pytorch/pull/155969#issuecomment-2992502683))	2025-06-20 18:40:38 +00:00
PyTorch MergeBot	96d082d06b	Revert "[InductorBench] Fix accuracy validation logic for MPS (#156385 )" This reverts commit 242eb19c8383b4b197963a8a564475d52c85ac66. Reverted https://github.com/pytorch/pytorch/pull/156385 on behalf of https://github.com/malfet due to Has some bug in error handling ([comment](https://github.com/pytorch/pytorch/pull/156385#issuecomment-2992441769))	2025-06-20 18:17:18 +00:00
Shunting Zhang	39270430c9	[inductor] force min num-split (off by default) (#155941 ) This is a fix for the 10% QPS regression of some internal model (internal doc: [here](https://docs.google.com/document/d/19EiSZSS_SNUNfRg3jmevyrDs9nVpyvyGX_LHfiz-SbU/edit?tab=t.0#heading=h.dim0r28ztzu5) and [here](https://docs.google.com/document/d/1DjRWJPl1cgpceaj8YXTyw6FubGb43Vw-lTAETF9XXnI/edit?tab=t.0#heading=h.ld0vvn8o77sp) ). The regression is caused by un-representable example inputs for compilation with dynamic shapes. While the general problem is hard to solve and requires more work, for this specific one, there is a quick fix. When we compile LayerNormBackward with small xnumel and large rnumel, we do split reduction. With un-representative inputs, rnumel may be something in the range like 4K and we pick a small num-split (9 in this specific case). Later on when we get an inputs with larger rnumel (100K range. no recompile due to dynamic shape enabled), the small num-split does not introduce enough parallelism and cause sub-optimal performance. The quick fix is to force a minimum value for num_split. Let's say we split a reduction [xnueml, rnueml] to two in this order: - [xnumel * num_split, rnumel / num_split] - [xnumel, num_split] A larger num_split always introduce more parallelism for kernel 1. It may results in more work in kernel 2. But if we set the minimum num_split to something not too large (like 256), for kernel2 each row may still be able to get done by reduction with a few or even a single warp. There may not be slow down for kernel 2. Here are some benchmarking results. ``` import torch from triton.testing import do_bench import functools from torch._inductor import config from torch._dynamo.decorators import mark_dynamic import os @torch.compile(dynamic=True) def f(x): return x.sum(dim=0) N = 512 C = functools.partial(torch.randn, device="cuda") x_small = C(4096, N) x_large = C(4096 * 1000, N) if os.getenv("HINT_WITH_SMALL_INPUT") == "1": x = x_small else: x = x_large mark_dynamic(x, 0) f(x) ms = do_bench(lambda: f(x_large)) # 4.03ms if hint with large input. Output code: https://gist.github.com/shunting314/0be562a0c14f8ec0852b12bbf53d7a15 # 8.32ms if hint with small input. Output code: https://gist.github.com/shunting314/79b924c266d5c562703c3bdfb48d8272 # 3.92ms if hint with small input, and force min num split: Output code: https://gist.github.com/shunting314/c82917a1849b698bf4d2be2fde2fd2ba print(ms) ``` This test mimic what we see in the original problem. - If we compile with large inputs and benchmark for large inputs, latency is 4.03ms - if we compile with small input but benchmark for large inputs, we get more than 2x slowdown. latency is 8.32ms - with the fix, even if we compile with small input and benchmark for large inputs, latency is 3.92ms. The perf is slightly better than the first case. So it's possible that the heuristic to decide num-split has room to improve The minimum num-split restriction could be applied for dynamic shape case solely, but I found it can also help for static shape cases a little bit. So I plan to apply it without checking dynamic shape for now unless I see red signals in thorough perf test. - Outer reduction with static shape: https://gist.github.com/shunting314/6a670a818e63533479399c4dbea5b29a . The fix improve perf from 0.01 ms to 0.009 ms - Inner reduction with static shape: https://gist.github.com/shunting314/f12f20099126130b953e55ad325c0f62 Perf is neutral (0.011 ms v.s. 0.011ms) A thorough perf test is running here: https://github.com/pytorch/pytorch/actions/runs/15642912325 # Update for not applying the change to static shape: from the perf test result [here](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Mon%2C%2009%20Jun%202025%2020%3A57%3A15%20GMT&stopTime=Mon%2C%2016%20Jun%202025%2020%3A57%3A15%20GMT&granularity=hour&mode=training&dtype=amp&deviceName=cuda%20(h100)&lBranch=gh/shunting314/210/head&lCommit=62b8e191e027842d402fb046a429732616f87570&rBranch=main&rCommit=5b9db4335e61c1c903cb0769282cbea588e49036), it looks like the change hurts perf for static shape case. I think one reason is the change may increase the number of kernels and lose some fusion opportunities. Check the following code for example: ``` import torch from torch._inductor import config aten = torch.ops.aten def f(x): return aten.bernoulli(x).sum() x = torch.randn(8000 * 3, dtype=torch.bfloat16, device="cuda") torch.compile(f)(x) ``` With the change the bernoulli kernel would NOT be able to fuse with the first layer reduction due to 8000 * 3 is not divisible by 256. Potentially we could improve the change to always pick num-split greater than 256 and divisible by rnumel . But I'll simply apply the change for dynamic shape for now since that's the original issue. Another perf test only applying min-num-split to dynamic shape [here](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Wed%2C%2011%20Jun%202025%2018%3A14%3A04%20GMT&stopTime=Wed%2C%2018%20Jun%202025%2018%3A14%3A04%20GMT&granularity=hour&mode=training&dtype=amp&deviceName=cuda%20(h100)&lBranch=gh/shunting314/210/head&lCommit=e7b2cf55f30a585acd4d907fc9127fcb30a256cc&rBranch=main&rCommit=d3d655ad14ee4cd1c135ac57bbf75d5623fc9fa6) Differential Revision: [D76625617](https://our.internmc.facebook.com/intern/diff/D76625617) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155941 Approved by: https://github.com/jansel, https://github.com/bobrenjc93	2025-06-20 18:01:28 +00:00
Jane Xu	55dae0bf7a	Add a basic shim and stable::Tensor is_contiguous API (#156228 ) Add a limited is_contiguous in shim, stable::Tensor API with a test case Pull Request resolved: https://github.com/pytorch/pytorch/pull/156228 Approved by: https://github.com/desertfire	2025-06-20 17:59:52 +00:00
Catherine Lee	49ee1e7106	[CI] Reuse old whl: loosen check for deleted files, do not handle renames (#156138 ) Make the check for deleted files only be for files in the torch folder since docs only changes could not get through this Use `--no-renames` to make both the old name and the old name show up in the diff. Without it I think only the new name shows up in git diff Pull Request resolved: https://github.com/pytorch/pytorch/pull/156138 Approved by: https://github.com/huydhn, https://github.com/malfet, https://github.com/cyyever	2025-06-20 17:58:04 +00:00
Mwiza Kunda	e31f205292	[Inductor] Adjust boundary checking of dimensions using YBLOCK (#149504 ) Apply the same logic introduced in https://github.com/pytorch/pytorch/pull/139751 to triton kernels using block ptrs. Here, if ynumel / YBLOCK > max_y_grids, dimensions dependent on YBLOCK need to be boundary checked, even if the block shape in such dimensions is a multiple of an expression in YBLOCK. This is because ynumel / YBLOCK % get_max_y_grids() may not be zero, so redundant programs will be launched that will attempt to read / write OOB. Pull Request resolved: https://github.com/pytorch/pytorch/pull/149504 Approved by: https://github.com/blaine-rister Co-authored-by: blaine-rister <145300525+blaine-rister@users.noreply.github.com>	2025-06-20 17:43:38 +00:00
Frost Mitchell	d83ff89d3b	Add toggle functionality for XPU profiler (#155135 ) Fixes #154898 by adding ability to toggle XPU profiler on and off (which has already been added in pytorch/kineto#1088 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155135 Approved by: https://github.com/guangyey, https://github.com/sraikund16	2025-06-20 17:27:48 +00:00
Nikita Shulga	1b50c12584	[BE] Make Eigen an optional dependency (#155955 ) Whose version is controlled by `eigen_pin.txt`, but which will be installed only if BLAS providers could not be found. Why this is good for CI: we don't really build with Eigen ever and gitlab can be down when github is up, which causes spurious CI failures in the past, for example. Remove eigen submodule and replace it with eigen_pin.txt Fixes https://github.com/pytorch/pytorch/issues/108773 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155955 Approved by: https://github.com/atalman ghstack dependencies: #155947, #155954	2025-06-20 17:21:27 +00:00
Xuehai Pan	63360e64da	[BE][Easy] do not install yanked `types-pkg-resources` in lint environment (#156462 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156462 Approved by: https://github.com/ezyang	2025-06-20 16:00:43 +00:00
PyTorch MergeBot	1036f6d114	Revert "[ROCm] Bump AOTriton to 0.10b (#156290 )" This reverts commit 34d8e64ef64d88324092a2028884c54c13e086b3. Reverted https://github.com/pytorch/pytorch/pull/156290 on behalf of https://github.com/atalman due to failing multiple internal tests ([comment](https://github.com/pytorch/pytorch/pull/156290#issuecomment-2992072727))	2025-06-20 15:35:25 +00:00
PyTorch MergeBot	b4442f42a9	Revert "Upgrade to DLPack 1.0. (#145000 )" This reverts commit 6e185c53124e1b5a0fe391959060c1249178bcb6. Reverted https://github.com/pytorch/pytorch/pull/145000 on behalf of https://github.com/atalman due to failing internal tests ([comment](https://github.com/pytorch/pytorch/pull/145000#issuecomment-2992055400))	2025-06-20 15:32:47 +00:00
PyTorch MergeBot	edd45f3a02	Revert "[Precompile] Hook up backend="inductor" (#155387 )" This reverts commit 2c68c3e8d5e9a235f5861be6486de4959f80c840. Reverted https://github.com/pytorch/pytorch/pull/155387 on behalf of https://github.com/atalman due to dynamo/test_precompile_context.py::PrecompileContextTests::test_basic [GH job link](https://github.com/pytorch/pytorch/actions/runs/15772892021/job/44464141039) [HUD commit link](`2c68c3e8d5`) ([comment](https://github.com/pytorch/pytorch/pull/155387#issuecomment-2992044073))	2025-06-20 15:30:04 +00:00
Hari Krishna Sai Kodali	e1f28fe17b	add device generalisation support for distributed tests (#152471 ) ### MOTIVATION To generalize Distributed test cases for non-CUDA devices ### CHANGES - test/distributed/optim/test_zero_redundancy_optimizer.py - test/distributed/test_c10d_logger.py - test/distributed/test_compute_comm_reordering.py Replaced hard coded device names with get_devtype from torch.testing._internal.common_fsdp. DistributedTestBase is used instead of MultiProcessTestCase, to make use of helper functions. - torch/testing/_internal/common_distributed.py extended common utility functions Pull Request resolved: https://github.com/pytorch/pytorch/pull/152471 Approved by: https://github.com/d4l3k	2025-06-20 07:35:42 +00:00
William Wen	0aed855b2b	[dynamo] raise hard error if error is encountered while tracing resume function prologue (#154564 ) This should prevent bad resume function prologues from slipping by. In particular, graph breaks in resume function prologues will now hard error. Implementation details: - The resume function prologue is surrounded by `LOAD_CONST arg, STORE_FAST __is_tracing_resume_prologue` instructions. The first sequence has `arg=True` and the second sequence has `arg=False`. - InstructionTranslator will know when it is tracing a resume function prologue when it detects `STORE_FAST __is_tracing_resume_prologue`. The top of stack will be True to mark the start of the prologue, False to mark the end. - When `convert_frame.py` detects that an error occurred while the InstructionTranslator was tracing a resume function prologue, we will wrap the exception and hard error Pull Request resolved: https://github.com/pytorch/pytorch/pull/154564 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289, #154782, #155166	2025-06-20 07:03:29 +00:00
William Wen	24dc33b37b	[dynamo] handle fullgraph toggle using nested torch.compile (#155166 ) See added test for the case that this PR handles. In particular, the semantics for nested torch.compile with toggled fullgraph settings was strange before - `@torch.compile(fullgraph=True)` overrides the existing fullgraph setting, while `@torch.compile(fullgraph=False)` does not. Note that this change will add an extra frame to any inlined torch.compile'd function (which I don't expect to happen frequently). Pull Request resolved: https://github.com/pytorch/pytorch/pull/155166 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289, #154782	2025-06-20 07:03:29 +00:00
William Wen	537b0877a8	[dynamo] fix set_fullgraph for nested calls (#154782 ) - Make the fullgraph argument of set_fullgraph a positional argument - Fix behavior on nested calls by updating `tracer.error_on_graph_break` in more places. In particular, a tracer's error_on_graph_break is set to the inlined tracer's error_on_graph_break upon the latter's exit. We also track error_on_graph_break in the speculation log now, since if we encounter a nested graph break, we will restart analysis and we need to somehow remember the error_on_graph_break setting after attempting to run the nested function (but we don't actually trace into it in the restart analysis). Pull Request resolved: https://github.com/pytorch/pytorch/pull/154782 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289	2025-06-20 07:03:16 +00:00
William Wen	2c372a0502	[dynamo] add set_fullgraph decorator/context manager (#154289 ) Implements https://github.com/pytorch/pytorch/issues/144908. Implementation notes: - `set_fullgraph` is implemented using `patch_config`, which changes config correctly during runtime and tracing. - Moved setting `config.error_on_graph_break` from convert_frame.py to eval_frame.py. This is because this should only be done at the top-level decorated function. If we kept this in convert_frame.py, we would be changing `config.error_on_graph_break` on every top-level frame, which causes confusing behavior (see added test for example). - InstructionTranslator reads from `config.error_on_graph_break` every `step()`. This is to determine the value of `config.error_on_graph_break` at the time of the graph break, because tracer cleanup will restore the value of `config.error_on_graph_break` . - `convert_frame.py` determines whether we should abort tracing (fullgraph=True) or continue (fullgraph=False) by reading the value of the tracer's `error_on_graph_break`. If there is no tracer (failed to initialize), then default to reading `config.error_on_graph_break`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154289 Approved by: https://github.com/jansel, https://github.com/zou3519 ghstack dependencies: #154283	2025-06-20 07:03:07 +00:00
William Wen	b46eb1ccaf	[dynamo] control one_graph behavior additionally through config (#154283 ) `torch.compile` now always goes through `torch._dynamo._optimize`. fullgraph is now implemented in `torch.compile` by looking at `config.error_on_graph_break`. Export still goes through `torch._dynamo._optimize_assert`, which uses `tx.one_graph` instead of `config.error_on_graph_break`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154283 Approved by: https://github.com/jansel, https://github.com/anijain2305	2025-06-20 07:02:57 +00:00
James Wu	2c68c3e8d5	[Precompile] Hook up backend="inductor" (#155387 ) This PR adds the necessary things to register and record backend ids from BundledAOTAutogradCacheEntry. One TODO to point out; in this diff, if there are multiple backends that would have the same AOTAutogradCache key (traditional cache key, not backend_id), we just end up serializing the same BundledAOTAutogradCache entry multiple times. This is not ideal obviously, so we'll want to deduplicate these and just track the different keys that one BundledAOTAutogradCacheEntry is associated with instead. This shouldn't be super hard to do, though, as we just need to run a deduplication step on call to `serialize()`, I think. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155387 Approved by: https://github.com/oulgen	2025-06-20 06:38:29 +00:00
Xuehai Pan	d5b4a32960	[BE] fix `PYPROJECT` linting errors in `test/` and `tools/` (#156021 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156021 Approved by: https://github.com/Skylion007	2025-06-20 06:19:05 +00:00
Nikita Shulga	4cbbc8b458	[MPS] Implement backward pass for interpolate_trilinear (#156373 ) Backwards pass simply iterates over all 8 points current point contributed to, and back propagates them with the respective weights TODO: Benchmark the performance of similar loop for the forward pas (i.e. compiler should be able to do loop unrolling, so no point of unrolling it by hand) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156373 Approved by: https://github.com/dcci ghstack dependencies: #156375	2025-06-20 05:41:24 +00:00
angelayi	c37ddcaefb	Fix torchgen update-aoti-shim (#156323 ) will remove the fill changes before landing and let Jane merge her changes! Pull Request resolved: https://github.com/pytorch/pytorch/pull/156323 Approved by: https://github.com/janeyx99	2025-06-20 05:23:06 +00:00
leslie-fang-intel	f7a5ad6c29	[Inductor][CPP] Fix WOQ int4 accuracy issue when NC large than one (#156407 ) Summary There is an accuracy issue when `Nc_block` is greater than 1 in WOQ int4 GEMM. Previously, we used the slice `{%- set tile_W = kernel.slice_nd(W, [("n_start", "n_start + n_size"), ("k_start * Nr / 2", "k_end * Nr / 2")]) %}`, which means that each `ni` in `Nc_block` takes the exact same N slice from `n_start` to `n_start + n_size`, leading to the accuracy problem. This accuracy issue is exposed by [PR #156174](https://github.com/pytorch/pytorch/pull/156174), which changes `block_N` from 64 to 32. This change increases the likelihood of `Nc_block` being greater than 1, making it more likely to trigger the issue. This PR will fix this accuracy issue. Test Plan ``` python test/inductor/test_cpu_select_algorithm.py -k test_int4_woq_mm_amx_Nc_larger_than_one ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156407 Approved by: https://github.com/CaoE	2025-06-20 03:08:02 +00:00
Cui, Yifeng	72c8751b61	Align meta deducing for fft_r2c with fft_r2c_mkl on XPU (#156048 ) There is a memory layout mismatching between `fft_r2c` XPU and Inductor meta deducing. Original `fft_r2c` Inductor meta deducing for XPU backend is aligned with CPU (fallback). This PR is to correct the Inductor meta deducing and update the torch-xpu-ops commit to [intel/torch-xpu-ops@`3a9419c`](`3a9419c8bb`). The XPU implementation first performs the R2C transform on the last dimension, followed by iterative C2C transforms on the remaining dimensions. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156048 Approved by: https://github.com/guangyey, https://github.com/etaf, https://github.com/jansel	2025-06-20 01:41:03 +00:00
CaoE	159a39ad34	Add an option for cpp_wrapper to compile entry and kernel separately (#156050 ) Fixes #156037. Compiling entry and kernel separately has a non-negligible impact on the performance. This PR is to add an option for cpp_wrapper to control whether to compile entry and kernel separately, and turn it off by default. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156050 Approved by: https://github.com/leslie-fang-intel, https://github.com/benjaminglass1, https://github.com/jansel	2025-06-20 01:11:16 +00:00
atalman	ebab279942	Forward fix inductor benchmark after #150287 (#156455 ) Looks like https://github.com/pytorch/pytorch/pull/150287 stack fixed some inductor tests HUD: https://hud.pytorch.org/hud/pytorch/pytorch/main/1?per_page=50&name_filter=inductor-periodic%20%2F%20linux-jammy-cpu-py3.9-gcc11-inductor Pull Request resolved: https://github.com/pytorch/pytorch/pull/156455 Approved by: https://github.com/huydhn	2025-06-20 00:04:15 +00:00
cyy	3c2324c64a	[2/N] Fix cppcoreguidelines-init-variables suppression (#146237 ) This PR removes all `cppcoreguidelines-init-variables` suppressions. Pull Request resolved: https://github.com/pytorch/pytorch/pull/146237 Approved by: https://github.com/ezyang	2025-06-19 23:26:42 +00:00
Aaron Orenstein	52f873adc2	Add logging for async compile worker statistics (#155820 ) Add some on-exit logging to the async compile workers. When you use `TORCH_LOGS=async_compile` (or `all`) it will now report how many workers were enqueued & dequeued (should be the same) as well as queuing time (how long workers sat on the queue before starting to run) and maximum depth (how many workers were waiting to start. Tested manually by running a larger internal model and then lowering the number of available workers to see the time and depth get longer. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155820 Approved by: https://github.com/masnesral	2025-06-19 23:10:15 +00:00
Yiming Zhou	c60d8188d2	[nativert] Move GraphExecutorBase to PyTorch core (#156196 ) Summary: Moves GraphExecutorBase class to PyTorch core. GraphExecutorBase is a lightweight abstraction to execute a graph with execution frames without actually owning the graph nor the weights. This is introduced to decouple the state management of the top level runtime from the kernel executions so that sub graphs from higher order ops can be supported. Torch Native Runtime RFC: pytorch/rfcs#72 Test Plan: CI Rollback Plan: Differential Revision: D76830436 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156196 Approved by: https://github.com/zhxchen17	2025-06-19 22:42:35 +00:00
Xinya Zhang	34d8e64ef6	[ROCm] Bump AOTriton to 0.10b (#156290 ) Notable new features/optimizations for SDPA operators on AMD systems from AOTriton 0.10b: * Official support of gfx950/gfx1201 * Experimental support of gfx1101/gfx1151/gfx1150/gfx1200 * Reduce libaotriton.so binary size by over 80%. + Without this optimization the binary size of `libaotriton.so` could be over 100MiB due to 2x more supported architectures compared with 0.9b. Now it is only about 11MiB. * Support sliding window attention (SWA) in `_flash_attention_forward/backward`. Should fix #154582 See https://github.com/ROCm/aotriton/releases/tag/0.10b for full details, including Known Problems. Notable changes to SDPA backend: * `std::optional<int64_t>` `window_size_left/right` are directly passed to ROCM's SDPA backend, because the default value `-1` is meaningful to AOTriton's backend and bottom-right aligned causal mask is implemented with negative `window_size_left/right` * Some code clean up around `USE_CK_FLASH_ATTENTION` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156290 Approved by: https://github.com/jithunnair-amd, https://github.com/jeffdaily	2025-06-19 21:13:58 +00:00
Ti-Tai Wang	3644b41a7c	[ONNX] Note on attention op symbolic function (#156441 ) Follow up https://github.com/pytorch/pytorch/pull/156367 Explain why num_heads is provided when ONNX Attention op does not need it in torch case: The thread: https://github.com/pytorch/pytorch/pull/156367#discussion_r2155727038 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156441 Approved by: https://github.com/justinchuby	2025-06-19 21:00:05 +00:00
Dmitry Rogozhkin	443b5b43c3	xpu: fix AOT compilation in sycl cpp extension (#156364 ) Commit fixes AOT compilation in sycl cpp extension which got accidentally dropped on aca2c99a652 (fallback to JIT compilation had happened). Commit also fixes override logic for default sycl targets allowing flexibility to specify targets externally. Further, commit extends test coverage to cover such a case and fixes issue in the test where consequent tests executed same (first) compiled extension due to name conflicts. Fixes: #156249 Fixes: aca2c99a652 ("xpu: get xpu arch flags at runtime in cpp_extensions (#152192)") CC: @pengxin99, @guangyey Pull Request resolved: https://github.com/pytorch/pytorch/pull/156364 Approved by: https://github.com/ezyang	2025-06-19 20:11:38 +00:00
Ke Wen	d32deb664a	[c10d] Disable NCCL NVLS when using deterministic mode (#156381 ) via setting env `NCCL_ALGO=^NVLS`. Note that this setting must be made before the first NCCL init. Otherwise, it won't take effect. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156381 Approved by: https://github.com/ngimel	2025-06-19 20:09:24 +00:00
Huy Do	69f2e09cc2	Add more shards to H100 benchmark, and also run it more frequently (#156429 ) There are 32 H100 `linux.aws.h100` and they are still not fully utilized with more than half staying idle, so we could add more shards to finish the whole suite within 4 hours. I add 1 more for `TIMM` and 3 more for `TorchBench` using the duration from a sample run https://github.com/pytorch/pytorch/actions/runs/15753185459/job/44411825090 With this computing power, we could also run the whole suite every 4 hours now. I could run this less frequently later if I see queueing Pull Request resolved: https://github.com/pytorch/pytorch/pull/156429 Approved by: https://github.com/atalman	2025-06-19 20:02:56 +00:00
Catherine Lee	aac0e8f0e9	[build] Create target for flash attention (#156235 ) Create a target for flash attention? so it can be built using ninja flash_attention Pull Request resolved: https://github.com/pytorch/pytorch/pull/156235 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-06-19 20:02:38 +00:00
Nikita Shulga	c2f4cc59a7	[MPS] Fix bug in 3d coords calculation (#156375 ) Which was not caught by CI beforehand, as all 3D examples right now are symmetric, so add an uneven shape to `sample_inputs_interpolate` Though it's indirectly tested by `test_upsample_nearest3d` inductor test Pull Request resolved: https://github.com/pytorch/pytorch/pull/156375 Approved by: https://github.com/atalman	2025-06-19 19:56:15 +00:00
Xuehai Pan	c0ee01c2fb	tools/nightly.py: only download `torch` via pip and install dependenices via `uv` (#156409 ) Setup time (cpu-only): 70s -> 27.6s -> 17.4s The tool can setup the pinned NVIDIA dependencies correctly: ```console $ make setup-env-cuda PYTHON="${HOMEBREW_PREFIX}/bin/python3.13" && source venv/bin/activate make setup-env PYTHON="/home/linuxbrew/.linuxbrew/bin/python3.13" NIGHTLY_TOOL_OPTS="pull --cuda" make[1]: Entering directory '/home/PanXuehai/Projects/pytorch' /home/linuxbrew/.linuxbrew/bin/python3.13 tools/nightly.py pull --cuda log file: /home/PanXuehai/Projects/pytorch/nightly/log/2025-06-19_21h16m16s_94cd1471-4d0f-11f0-b120-b88584c06696/nightly.log Creating virtual environment Removing existing venv: /home/PanXuehai/Projects/pytorch/venv Creating venv (Python 3.13.4): /home/PanXuehai/Projects/pytorch/venv Installing packages Upgrading package(s) (https://download.pytorch.org/whl/nightly/cu128): - uv - pip - setuptools - packaging - wheel - build[uv] Looking in indexes: https://pypi.tuna.tsinghua.edu.cn/simple, https://download.pytorch.org/whl/nightly/cu128 Collecting uv Using cached `f2e96cec5e/uv-0.7.13-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl` (17.8 MB) Requirement already satisfied: pip in ./venv/lib/python3.13/site-packages (25.1.1) Collecting setuptools Using cached `17031897da/setuptools-80.9.0-py3-none-any.whl` (1.2 MB) Collecting packaging Using cached `38679034af/packaging-25.0-py3-none-any.whl` (66 kB) Collecting wheel Using cached `87f3254fd8/wheel-0.45.1-py3-none-any.whl` (72 kB) Collecting build[uv] Using cached `80633736cd/build-1.2.2.post1-py3-none-any.whl` (22 kB) Collecting pyproject_hooks (from build[uv]) Using cached `12818598c3/pyproject_hooks-1.2.0-py3-none-any.whl` (10 kB) Installing collected packages: wheel, uv, setuptools, pyproject_hooks, packaging, build Successfully installed build-1.2.2.post1 packaging-25.0 pyproject_hooks-1.2.0 setuptools-80.9.0 uv-0.7.13 wheel-0.45.1 Installing packages took 6.251 [s] Creating virtual environment took 9.050 [s] Downloading packages Downloading package(s) (https://download.pytorch.org/whl/nightly/cu128): torch Looking in indexes: https://pypi.tuna.tsinghua.edu.cn/simple, https://download.pytorch.org/whl/nightly/cu128 Collecting torch Using cached https://download.pytorch.org/whl/nightly/cu128/torch-2.8.0.dev20250619%2Bcu128-cp313-cp313-manylinux_2_28_x86_64.whl.metadata (30 kB) Using cached https://download.pytorch.org/whl/nightly/cu128/torch-2.8.0.dev20250619%2Bcu128-cp313-cp313-manylinux_2_28_x86_64.whl (1040.3 MB) Saved /tmp/pip-download-xeqmhrww/torch-2.8.0.dev20250619+cu128-cp313-cp313-manylinux_2_28_x86_64.whl Successfully downloaded torch Downloaded 1 file(s) to /tmp/pip-download-xeqmhrww: - torch-2.8.0.dev20250619+cu128-cp313-cp313-manylinux_2_28_x86_64.whl Downloading packages took 6.284 [s] Unpacking wheel file Unpacking to: /tmp/wheel-kugk2os0/torch-2.8.0.dev20250619+cu128...OK Unpacking wheel file took 15.107 [s] Installing dependencies Installing packages Installing package(s) (https://download.pytorch.org/whl/nightly/cu128): - filelock - typing-extensions>=4.10.0 - setuptools; python_version >= "3.12" - sympy>=1.13.3 - networkx - jinja2 - fsspec - nvidia-cuda-nvrtc-cu12==12.8.93; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cuda-runtime-cu12==12.8.90; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cuda-cupti-cu12==12.8.90; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cudnn-cu12==9.10.2.21; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cublas-cu12==12.8.4.1; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cufft-cu12==11.3.3.83; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-curand-cu12==10.3.9.90; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cusolver-cu12==11.7.3.90; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cusparse-cu12==12.5.8.93; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cusparselt-cu12==0.7.1; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-nccl-cu12==2.27.3; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-nvshmem-cu12==3.2.5; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-nvtx-cu12==12.8.90; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-nvjitlink-cu12==12.8.93; platform_system == "Linux" and platform_machine == "x86_64" - nvidia-cufile-cu12==1.13.1.3; platform_system == "Linux" and platform_machine == "x86_64" - pytorch-triton==3.3.1+gitc8757738; platform_system == "Linux" - numpy - cmake - ninja - packaging - ruff - mypy - pytest - hypothesis - ipython - rich - clang-format - clang-tidy - sphinx Using Python 3.13.4 environment at: venv Resolved 78 packages in 2.95s Installed 76 packages in 93ms + alabaster==1.0.0 + asttokens==3.0.0 + attrs==24.2.0 + babel==2.17.0 + certifi==2024.8.30 + charset-normalizer==3.3.2 + clang-format==20.1.6 + clang-tidy==20.1.0 + cmake==3.25.0 + decorator==5.2.1 + docutils==0.21.2 + executing==2.2.0 + filelock==3.18.0 + fsspec==2025.5.1 + hypothesis==6.135.11 + idna==3.10 + imagesize==1.4.1 + iniconfig==2.1.0 + ipython==9.3.0 + ipython-pygments-lexers==1.1.1 + jedi==0.19.2 + jinja2==3.1.6 + markdown-it-py==3.0.0 + markupsafe==2.1.5 + matplotlib-inline==0.1.7 + mdurl==0.1.2 + mpmath==1.3.0 + mypy==1.16.1 + mypy-extensions==1.0.0 + networkx==3.5 + ninja==1.11.1.4 + numpy==2.3.0 + nvidia-cublas-cu12==12.8.4.1 + nvidia-cuda-cupti-cu12==12.8.90 + nvidia-cuda-nvrtc-cu12==12.8.93 + nvidia-cuda-runtime-cu12==12.8.90 + nvidia-cudnn-cu12==9.10.2.21 + nvidia-cufft-cu12==11.3.3.83 + nvidia-cufile-cu12==1.13.1.3 + nvidia-curand-cu12==10.3.9.90 + nvidia-cusolver-cu12==11.7.3.90 + nvidia-cusparse-cu12==12.5.8.93 + nvidia-cusparselt-cu12==0.7.1 + nvidia-nccl-cu12==2.27.3 + nvidia-nvjitlink-cu12==12.8.93 + nvidia-nvshmem-cu12==3.2.5 + nvidia-nvtx-cu12==12.8.90 + parso==0.8.4 + pathspec==0.12.1 + pexpect==4.9.0 + pluggy==1.6.0 + prompt-toolkit==3.0.51 + ptyprocess==0.7.0 + pure-eval==0.2.3 + pygments==2.19.1 + pytest==8.4.1 + pytorch-triton==3.3.1+gitc8757738 + requests==2.32.3 + rich==14.0.0 + roman-numerals-py==3.1.0 + ruff==0.12.0 + snowballstemmer==3.0.1 + sortedcontainers==2.4.0 + sphinx==8.2.3 + sphinxcontrib-applehelp==2.0.0 + sphinxcontrib-devhelp==2.0.0 + sphinxcontrib-htmlhelp==2.1.0 + sphinxcontrib-jsmath==1.0.1 + sphinxcontrib-qthelp==2.0.0 + sphinxcontrib-serializinghtml==2.0.0 + stack-data==0.6.3 + sympy==1.14.0 + traitlets==5.14.3 + typing-extensions==4.14.0 + urllib3==2.2.3 + wcwidth==0.2.13 Installing packages took 3.080 [s] Installing dependencies took 3.080 [s] Pulling nightly PyTorch Found released git version 5622038e20ddb12b9a011c9a9128190d71a21cba Found nightly release version 2625c70aecc6eced1dbe108279feab7509733bef Already up to date. Pulling nightly PyTorch took 0.017 [s] Moving nightly files into repo Moving nightly files into repo took 4.898 [s] Writing pytorch-nightly.pth Writing pytorch-nightly.pth took 0.021 [s] ------- PyTorch Development Environment set up! Please activate to enable this environment: $ source /home/PanXuehai/Projects/pytorch/venv/bin/activate make[1]: Leaving directory '/home/PanXuehai/Projects/pytorch' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156409 Approved by: https://github.com/ezyang ghstack dependencies: #156408	2025-06-19 19:42:15 +00:00
Xuehai Pan	71faa7e5b9	tools/nightly.py: use `uv pip install` instead of `pip install` (#156408 ) Setup time: 70s -> 27.6s Pull Request resolved: https://github.com/pytorch/pytorch/pull/156408 Approved by: https://github.com/ezyang	2025-06-19 19:42:15 +00:00
Chen Haifeng	134dfb3fe6	[dynamo] Fix cycle reference problem caused by recursive collect_temp_source in codegen (#155791 ) Recursive function collect_temp_source with closure in PyCodegen caused cycle reference issue when torch.compile is used. This issue may cause major tensors will not freed timely even there are no user references to these tensors. We saw OOM issues because of this problem in many cases including training and inference using torch.compile. The fix is to use iterative function implementation to replace the recursive function implementation. Fixes #155778 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155791 Approved by: https://github.com/ezyang	2025-06-19 19:37:44 +00:00
Shangdi Yu	e4c9f6d9a2	[nativert] Move c10_kernel (#156208 ) Summary: Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 As part of the effort to open source TorchNativeRuntime (or what we call Sigmoid), we are moving the Pytree implementation to torch/: fbcode/sigmoid/kernels -> fbcode/caffe2/torch/nativert/kernels Test Plan: ``` buck run fbcode//mode/dev-nosan //caffe2/test/cpp/nativert:c10_kernel_test ``` Differential Revision: D76825830 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156208 Approved by: https://github.com/zhxchen17	2025-06-19 17:36:23 +00:00
Dmitry Nikolaev	f402eed4d9	[ROCm] Enable BF16 NCHW Mixed batchnorm on MIOpen if ROCm>=6.4 (#154611 ) This PR enables MIOpen for BF16 NCHW Mixed batchnorm if MIOpen version >=3.4 (ROCm >= 6.4) CUDAHooks::versionMIOpen() was added to detect MIOpen version Pull Request resolved: https://github.com/pytorch/pytorch/pull/154611 Approved by: https://github.com/jeffdaily, https://github.com/jithunnair-amd	2025-06-19 17:22:37 +00:00
Doru Bercea	085f270a00	[ROCm] Enable more parallelism for multi-dimensional reductions (#155806 ) Enable more parallelism for multi-dimensional reductions. In the case of multi-dimensional reductions the grid often start with a single active block. In such cases, we need to allow the parallelism to be extended along the y-direction of the grid to avoid having a single block running. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155806 Approved by: https://github.com/Skylion007, https://github.com/jeffdaily	2025-06-19 17:19:40 +00:00
Shangdi Yu	eaf704914e	[aoti] package weights to disk and dedup (#155241 ) We package the weights and save them in `data/weights/` (`WEIGHTS_DIR`). In addition, we store a `weights_config.json` in the model folder for each model to specify which weight file corresponding to which weight name. Models can share weights. We dedup the weights based on their underlying storage (`tensor.untyped_storate()`). - Use `"aot_inductor.package_constants_on_disk": True` config to produce the `Weights` in aot_compile - If we see `Weights` in aoti_files, we'll automatically package them to disk - `"aot_inductor.package_constants_on_disk"` config and `"aot_inductor.package_constants_in_so"` config work independently. - Use `load_pt2(package_path, load_weights_from_disk=True)` to load the weights from disk. `load_weights_from_disk` defaults to False. Test Plan: ``` buck2 run @//mode/dev-nosan //caffe2/test/inductor:aot_inductor_package -- -r "test_package_shared_weights" ``` Tested with whisper at https://github.com/pytorch-labs/torchnative/pull/7 Rollback Plan: Differential Revision: D74747190 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155241 Approved by: https://github.com/desertfire	2025-06-19 17:17:17 +00:00
Yukio Siraichi	6e185c5312	Upgrade to DLPack 1.0. (#145000 ) This PR makes the necessary changes in order to upgrade PyTorch DLPack support to version 1.0. In summary, we add support for the following: - Support both `DLManagedTensor` and `DLManagedTensorVersioned` when producing and consuming DLPack capsules - New parameter for `__dlpack__` method: `max_version` - Version checks: - Fallback to old implementation if no `max_version` or if version lower than 1.0 - Check that the to-be-consumed capsule is of version up to 1.X In order to accommodate these new specifications, this PR adds the following main changes: - `torch._C._to_dlpack_versioned` Python API (Module.cpp): new Python API for creating a versioned DLPack capsule (called by `__dlpack__` method) - `DLPackTraits<T>` class (DLConvertor.h): select the correct traits (e.g. capsule name, conversion functions) depending on which DLPack tensor class is being used - `toDLPackImpl<T>` function (DLConvertor.cpp): populates the common fields of both classes - `fromDLPackImpl<T>` function (DLConvertor.cpp): constructs a tensor from a DLPAck capsule - `fillVersion<T>` function (DLConvertor.cpp): populates the version field for `DLManagedTensorVersioned` (no-op for `DLManagedTensor`) - `tensor_fromDLPackImpl<T>` function (tensor_new.cpp): outer function for constructing a tensor out of a DLPack capsule that also marks the capsule as used Pull Request resolved: https://github.com/pytorch/pytorch/pull/145000 Approved by: https://github.com/albanD	2025-06-19 16:27:42 +00:00
Raman-RH	6eb6f198e1	update codebase structure documentation to include mps (#156297 ) 📚 The doc update adding description about mps folder in code structure guide @albanD @malfet @svekars @sekyondaMeta Pull Request resolved: https://github.com/pytorch/pytorch/pull/156297 Approved by: https://github.com/ezyang	2025-06-19 16:16:29 +00:00
Zhengxu Chen	7f0cddfb55	[dynamo] Add documentation for guard_filter_fn (#156114 ) Summary: Adding a section of doc for guard_filter_fn. Test Plan: CI Rollback Plan: Differential Revision: D76756743 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156114 Approved by: https://github.com/jansel	2025-06-19 16:13:12 +00:00
Benjamin Glass	c9afcffed0	[AOTInductor] Call most runtime fallback ops without calling into Python (#154142 ) Uses the new aoti_torch_call_dispatcher interface to call runtime fallback ops without calling back into Python. This supports a limited subset of input and output datatypes, but a significant majority of remaining fallback ATen ops are covered. Fixes #150988 Fixes #153478 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154142 Approved by: https://github.com/desertfire	2025-06-19 15:27:15 +00:00
PyTorch MergeBot	317af4c87b	Revert "[cuDNN][64-bit indexing] update conv depthwise 64bit indexing dispatch condition to match native kernel (#156140 )" This reverts commit a5f59cc2eab3a5201712c52fe48c268357ba4f3c. Reverted https://github.com/pytorch/pytorch/pull/156140 on behalf of https://github.com/atalman due to breaks internal builds ([comment](https://github.com/pytorch/pytorch/pull/156140#issuecomment-2988441548))	2025-06-19 15:09:29 +00:00
Jeff Daily	ab3393e923	[ROCm][CI] fix mi300 test failure after 6.4.1 update (#156368 ) Fixes failures such as https://github.com/pytorch/pytorch/actions/runs/15739699156/job/44365395854: `test/test_linalg.py::TestLinalgCUDA::test_broadcast_batched_matmul_cuda` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156368 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-19 15:02:40 +00:00
PyTorch MergeBot	0b62465b99	Revert "Refine alignment check along dynamic dimension for grouped MMs (#155466 )" This reverts commit 830a335a7da5fec00395d440ba568749cb4e2e9e. Reverted https://github.com/pytorch/pytorch/pull/155466 on behalf of https://github.com/atalman due to breaks internal builds ([comment](https://github.com/pytorch/pytorch/pull/155466#issuecomment-2988285117))	2025-06-19 14:25:38 +00:00
Kaichao You	fec8af8b98	[bugfix] [build] guard cuda version for ipc with fabric handle (#156394 ) https://github.com/pytorch/pytorch/pull/156074 adds the support of ipc with fabric handle, but the code cannot compile for cuda < 12.3 (in particular, e.g. cuda 11.8). this pr improves the support by adding some compilation-time check against cuda versions. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156394 Approved by: https://github.com/ngimel	2025-06-19 13:54:01 +00:00
Nikita Shulga	769d754ab2	[BE][MPS] Refactor core matmul logic into matmul_core (#155969 ) In preparation of adding integer addmm, move matmul computation part into matmul_inner function Change callstack from group_id, thread_id_in_group to thread_id, threadid_in_group, which eliminates the need of calculating the index Pull Request resolved: https://github.com/pytorch/pytorch/pull/155969 Approved by: https://github.com/Skylion007	2025-06-19 13:22:41 +00:00
xinan.lin	8cb0c4a4da	[Intel GPU][AOTI] Add xpu mkldnn ops support for AOTInductor. (#154586 ) This PR is closely related to the previous one in the stack(https://github.com/pytorch/pytorch/pull/150287). The previous PR enabled MKLDNN ops for XPU, which caused several test cases to fail in test_aot_inductor.py. This PR addresses those failing cases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154586 Approved by: https://github.com/EikanWang, https://github.com/desertfire ghstack dependencies: #150287	2025-06-19 13:17:22 +00:00
xinan.lin	83259cf7a7	[Inductor][Intel GPU] Support mkldnn Conv post op fusion for XPU. (#150287 ) This PR adds support for MKLDNN Conv post-op fusion in the Inductor Intel GPU backend under freezing mode. The implementation reuses the CPU's MKLDNN pattern fusion mechanism, as well as the corresponding Inductor unit tests for CPU MKLDNN pattern fusion. The performance improvement: \| Suite \| Inductor Speedup (Baseline) \| Inductor Speedup (Compared) \| Acc Failed \| Perf Failed \| Inductor Perf Ratio \| Speedup \| \|-------------\|-----------------------------\|------------------------------\|------------\|--------------\|----------------------\|----------\| \| Huggingface \| 2.134838 \| 2.125740314 \| 0 \| 0 \| 1.001462504 \| 100.43% \| \| Torchbench \| 1.808558 \| 1.675100479 \| 0 \| 0 \| 1.075722187 \| 107.97% \| \| Timm \| 2.343893 \| 2.070476653 \| 0 \| 0 \| 1.131023832 \| 113.21% \| Pull Request resolved: https://github.com/pytorch/pytorch/pull/150287 Approved by: https://github.com/ZhiweiYan-96, https://github.com/EikanWang, https://github.com/jansel	2025-06-19 13:17:22 +00:00
Ting Lu	0504480f37	Add CUDA 12.9 libtorch nightly (#155895 ) https://github.com/pytorch/pytorch/issues/155196 with libtorch docker added, we can add the build script Pull Request resolved: https://github.com/pytorch/pytorch/pull/155895 Approved by: https://github.com/atalman	2025-06-19 13:15:42 +00:00
Daisy Deng	ccb1f687d6	Port two dynamo test cases for Intel GPU (#156056 ) For https://github.com/pytorch/pytorch/issues/114850, we will port more cases to Intel GPU. This PR is for 2 dynamo cases. We adopted "torch.accelerator.current_accelerator()" to determine the backend, and added XPU support in decorators like @requires_gpu, also enabled XPU for some test path. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156056 Approved by: https://github.com/guangyey, https://github.com/jansel	2025-06-19 12:49:04 +00:00
PyTorch MergeBot	a8fe982993	Revert "[build] Create target for flash attention (#156235 )" This reverts commit 6d02321472ee0761092166dd273eb3ec386cf0c0. Reverted https://github.com/pytorch/pytorch/pull/156235 on behalf of https://github.com/ZainRizvi due to Weird, but seems to have broken trunk: test_jit_fuser_te.py::TestTEFuserDynamic::test_skip_grad_in_check [GH job link](https://github.com/pytorch/pytorch/actions/runs/15748768079/job/44390494621) [HUD commit link](`6d02321472`) ([comment](https://github.com/pytorch/pytorch/pull/156235#issuecomment-2987784207))	2025-06-19 11:47:27 +00:00
codingwithsurya	4da98351b9	[SymmMem] Add NVSHMEM PUT with Signal support to Triton (#156211 ) Adds NVSHMEM PUT with Signal operation support for Triton kernels: - Added`putmem_signal_block` core.extern wrapper for nvshmemx_putmem_signal_block - Added kernel for 2-rank PUT operation with atomic SET signaling (`test_triton_put_signal_set`) - Added kernel for 2-rank PUT operation with atomic ADD signaling (`test_triton_put_signal_add`) Tests: `$ TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py` `TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_put_signal_set` `TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_put_signal_add` ```python @skipIfRocm @requires_triton() def test_triton_put_signal_set(self) -> None: @triton.jit def put_signal_kernel(dst_ptr, src_ptr, numel: tl.constexpr, sig_ptr, signal_val: tl.constexpr, sig_op: tl.constexpr, peer: tl.constexpr): nvshmem.putmem_signal_block(dst_ptr, src_ptr, numel, sig_ptr, signal_val, sig_op, peer) # ... setup code ... val = 11 inp = symm_mem.empty(numel, dtype=dtype, device=self.device).fill_(val) out = symm_mem.empty(numel, dtype=dtype, device=self.device).fill_(-1) # destination buffer # Signal flag buffer - starts at 0, will be set to 1 upon completion flag = symm_mem.empty(1, dtype=torch.int64, device=self.device).fill_(0) peer = 1 - rank NVSHMEM_SIGNAL_SET = 0 # atomic set operation SIGNAL_VAL = 1 # completion signal value if rank == 0: # Rank 0 atomically: (1) puts data to rank 1, (2) sets rank 1's flag to 1 put_signal_kernel[(1, 1, 1)](dst_ptr, src_ptr, numel=numel, sig_ptr=sig_ptr, signal_val=SIGNAL_VAL, sig_op=NVSHMEM_SIGNAL_SET, peer=peer, extern_libs=nvshmem_lib) dist.barrier() # Rank 1 can check flag to know data transfer completed! print(f"[Rank {rank}] inp buffer: {inp}") print(f"[Rank {rank}] out buffer: {out}") print(f"[Rank {rank}] flag buffer: {flag}") ``` ``` [Rank 0] inp buffer: tensor([11, 11, 11, 11, 11, 11, 11, 11], device='cuda:0', dtype=torch.int8) [Rank 0] out buffer: tensor([-1, -1, -1, -1, -1, -1, -1, -1], device='cuda:0', dtype=torch.int8) [Rank 0] got data from peer 1 [Rank 0] flag buffer: tensor([0], device='cuda:0') [Rank 1] inp buffer: tensor([11, 11, 11, 11, 11, 11, 11, 11], device='cuda:1', dtype=torch.int8) [Rank 1] out buffer: tensor([11, 11, 11, 11, 11, 11, 11, 11], device='cuda:1', dtype=torch.int8) [Rank 1] got data from peer 0 [Rank 1] flag buffer: tensor([1], device='cuda:1') ---------------------------------------------------------------------- Ran 2 tests in 17.046s OK ``` Working as expected! Data is received, and flag set to 1 for completion signal! ```python @skipIfRocm @requires_triton() def test_triton_put_signal_add(self) -> None: @triton.jit def put_signal_kernel(dst_ptr, src_ptr, numel: tl.constexpr, sig_ptr, signal_val: tl.constexpr, sig_op: tl.constexpr, peer: tl.constexpr): nvshmem.putmem_signal_block(dst_ptr, src_ptr, numel, sig_ptr, signal_val, sig_op, peer) # ... setup code ... # Signal buffer (uint64 flag) flag = symm_mem.empty(1, dtype=torch.int64, device=self.device).fill_(0) peer = 1 - rank NVSHMEM_SIGNAL_ADD = 5 # atomic add operation SIGNAL_VAL = 16 # Signal value to add if rank == 0: # Rank 0 puts into Rank 1 and adds to signal put_signal_kernel[(1, 1, 1)](dst_ptr, src_ptr, numel=numel, sig_ptr=sig_ptr, signal_val=SIGNAL_VAL, sig_op=NVSHMEM_SIGNAL_ADD, peer=peer, extern_libs=nvshmem_lib) dist.barrier() print(f"[Rank {rank}] inp buffer: {inp}") print(f"[Rank {rank}] out buffer: {out}") print(f"[Rank {rank}] flag buffer: {flag}") ``` ``` [Rank 0] inp buffer: tensor([11, 11, 11, 11, 11, 11, 11, 11], device='cuda:0', dtype=torch.int8) [Rank 0] out buffer: tensor([-1, -1, -1, -1, -1, -1, -1, -1], device='cuda:0', dtype=torch.int8) [Rank 0] got data from peer 1 [Rank 0] flag buffer: tensor([0], device='cuda:0') [Rank 1] inp buffer: tensor([11, 11, 11, 11, 11, 11, 11, 11], device='cuda:1', dtype=torch.int8) [Rank 1] out buffer: tensor([11, 11, 11, 11, 11, 11, 11, 11], device='cuda:1', dtype=torch.int8) [Rank 1] got data from peer 0 [Rank 1] flag buffer: tensor([16], device='cuda:1') ---------------------------------------------------------------------- Ran 1 test in 17.145s OK ``` The flag transition from [0] → [16] confirms both data delivery and atomic signal completion in a single operation! Pull Request resolved: https://github.com/pytorch/pytorch/pull/156211 Approved by: https://github.com/kwen2501, https://github.com/mandroid6	2025-06-19 10:24:30 +00:00
bobrenjc93	348e2a76df	s/defer_runtime_assert/guard_or_defer_runtime_assert (#156397 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156397 Approved by: https://github.com/laithsakka	2025-06-19 10:18:28 +00:00
Yuan Yao	02080c2cd9	Fix num_heads inference in ONNX Attention-23 exporter (#156367 ) Fixes issue in torch-onnx exporter for Attention: https://github.com/pytorch/pytorch/issues/156105 Previously the number of heads attributes inferred by the exporter is incorrect. It should be read from input dimension -3 not dimension 3: ![image](https://github.com/user-attachments/assets/26f10e15-bc98-42ac-807a-2e089a7d996a) But in fact, [torch sdpa](https://docs.pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html) doesn't support combined num_heads and head_size dimensions like [ONNX](https://onnx.ai/onnx/operators/onnx__Attention.html) does, so this num_heads attribute is not needed. Extending support to rank>4 can be left as future work if there is use case for that. The translation logic will look like: Reshape(Q,K,V to 4d) -> Attention -> Reshape(Y to original rank). Pull Request resolved: https://github.com/pytorch/pytorch/pull/156367 Approved by: https://github.com/justinchuby, https://github.com/titaiwangms	2025-06-19 09:40:01 +00:00
Ke Wen	8fcda2c60d	[SymmMem] Add runtime detection of NVSHMEM (#156291 ) so that we can pick the default backend for SymmetricMemory without fully relying on env var `TORCH_SYMMMEM=CUDA \| NVSHMEM` On Python side, the following API is added: `torch.distributed._symmetric_memory.is_nvshmem_available()` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156291 Approved by: https://github.com/Skylion007 ghstack dependencies: #155506, #155835, #155968, #155971, #155975, #156116, #156117	2025-06-19 08:26:11 +00:00
Pian Pawakapan	eabf7cd3c5	[export] update docs for Dims (#156262 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/156262 Approved by: https://github.com/angelayi	2025-06-19 06:25:21 +00:00
Pian Pawakapan	ec0276103f	[PGO] fix whitelist scalar bug (#156194 ) Test Plan: test_pgo Rollback Plan: Differential Revision: D76830552 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156194 Approved by: https://github.com/bobrenjc93	2025-06-19 05:51:21 +00:00
Xuehai Pan	1c960c5638	[Makefile] lazily setup `lintrunner` on first `make lint` run (#156058 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156058 Approved by: https://github.com/ezyang	2025-06-19 05:43:35 +00:00
Nikita Shulga	242eb19c83	[InductorBench] Fix accuracy validation logic for MPS (#156385 ) As it does not support full fp64, validate against float32 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156385 Approved by: https://github.com/Skylion007	2025-06-19 05:37:51 +00:00
Junjie Wang (PyTorch)	ce8180a61d	[c10d] Disable stack trace call in logging (#156362 ) Summary: We noticed std::future_error: Broken promise errors in logging, so let's disable for now and will investigate more. Test Plan: CI Rollback Plan: Differential Revision: D76929722 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156362 Approved by: https://github.com/fegin	2025-06-19 05:11:57 +00:00
Yiming Zhou	a21806f038	[ez][export] Better error message for schema check in torch.export.load (#156361 ) Summary: Fixes https://github.com/pytorch/pytorch/issues/156354 torch.export.load() only supports files generated by torch.export.save() Test Plan: CI Rollback Plan: Differential Revision: D76928725 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156361 Approved by: https://github.com/zhxchen17	2025-06-19 04:50:56 +00:00
Laith Sakka	3f69e3b3a0	Add view_simple as meta function for view, and avoid calling reshape_view_helper for unbacked (#154757 ) address https://github.com/pytorch/pytorch/issues/153303 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154757 Approved by: https://github.com/bobrenjc93, https://github.com/leslie-fang-intel	2025-06-19 04:50:18 +00:00
Simon Fan	3bec588bf5	[aot][ca] save bw_module in AOTAutogradCache (#151860 ) Compiled Autograd retraces AOT's bw_module at backward runtime into a larger graph, and today this runs into an issue on warm cache runs because the bw_module is not restored. This PR adds it to the cache, by first stripping it bare from unserializable metadata. I also intentionally differentiate the cached and non-cached versions to avoid accidental attempts of AOT compilation with a restored bw_module (would probably crash). The bw_module's generated code is then serialized, and at compiled autograd runtime, it is restored via symbolic_trace. This also means that presence of tensor constructors will be lifted as constants. Something we will address separately. Note that since the cache entry may be used by runs that use compiled autograd and runs that do not, we need to cache both the lowered backward and the bw_module. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151860 Approved by: https://github.com/jamesjwu ghstack dependencies: #156120	2025-06-19 03:47:41 +00:00
Catherine Lee	6d02321472	[build] Create target for flash attention (#156235 ) Create a target for flash attention? so it can be built using ninja flash_attention Pull Request resolved: https://github.com/pytorch/pytorch/pull/156235 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-06-19 03:35:04 +00:00
Wang, Chuanqi	77518d1a13	[CI] fix xpu-smi hang in XPU test container (#156171 ) Apply same fix #155443 for XPU test container, refer https://github.com/pytorch/pytorch/actions/runs/15589866881/job/43907973867#step:15:911 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156171 Approved by: https://github.com/huydhn	2025-06-19 02:48:11 +00:00
Teja Rao	19ffdf4ea0	[dcp] add new checkpoint staging to preserve storage sharing and support mutable state_dicts (#155192 ) Summary: This implements staging in way that doesnt mess up checkpointing semantics. We want to be close to torch.save/load semantics and when async checkpointing is used it messes up shared storages, doesnt handle custom objects or tensors well. EG: users passes a state_dict with a cuda tensor in datatype. this is deepcloned causing the staging tensor to be created on GPU. This can cause ooms is hard to debug. This diffs hooks into deepcopy of storages to move them to cpu using the cached storages created for async checkpoint staging. This allows reusing storages created for staging to avoid recreating them on each checkpoint while also being flexible enough to handle any changes - clean up old storages or create new ones as needed. Lifetime of staging storages is tied to the original storage object. when the original storage object is gc-ed, we delete the corresponding staging storage from cache possibly causing it to gc-ed is there are no other references. I am using data_ptr of the storage to keep track of this. Please share thoughts on this. The alternative is to use fqn's instead of storage_id and verify the underlying storage object has same shape/size,etc to make the caching logic work. Current implementation is much simpler and cleaner. The API: ``` # construct a stager once per job in checkpointing. stager = StateDictStager(pin_memory=pin_memory, share_memory=share_memory) # do this on every checkpoint: with staging_context(stager): cpu_state_dict = copy.deepcopy(state_dict) ``` Also, adds support for pinned-memory. One problem this implementation does not address is that we lose the original device. The only alternatives here are - pickle synchronously like torch.save but with special handling for storages. It is valuable to keep state_dict throughout the checkpointing process. so users can manipulate and debug as needed. so we need to unpickle in the background process. I think this is flexible, not performant and not very different to current solution but needs more code. One idea if we really want to address is this to stick the original device in a some variable on storage and then use it recover on load side. I think we do not need this for now and can be explicit about losing device type for async checkpointing. Update: Note: Due to reservations on hooking into deepcopy to customize it, the PR is now updated to use deepcopy like logic to clone the state_dict. There are some caveats to this solution: 1. Duplicated deepcopy code to hook into for tensors. There is a risk of this code getting outdated with python version changes. This is needed to handle several different types like NamedTuples, frozen dataclasses, nested dataclasses. deepcopy logic is relying on reduce_ex to get a function with which these can be constructed. 2. Since we are bypassing deepcopy and adding custom logic to clone a tensor, we are missing some of the functionality that exists in deepcopy for torch.Tensor like _clear_non_serializable_cached_data(), or other logic. Would like thoughts on which logic or if everything should be copied? 3. If any object implemented deepcopy , we will not be able to handle any tensors in the attrs with this logic because they likely just call copy.deepcopy on the attrs instead of this deepcopy logic. We are taking care of subclasses of torch.Tensor to workaround this. The new API: ``` # construct a stager once per job in checkpointing. stager = StateDictStager(pin_memory=pin_memory, share_memory=share_memory) # do this on every checkpoint: cpu_state_dict = copy.stage(state_dict) ``` Test Plan: unit tests Differential Revision: D75993324 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155192 Approved by: https://github.com/mikaylagawarecki, https://github.com/pradeepfn	2025-06-19 02:04:21 +00:00
Luca Wehrstedt	d4ad280429	Enable querying the build and runtime NCCL versions (#156305 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156305 Approved by: https://github.com/wconstab, https://github.com/Skylion007, https://github.com/fegin	2025-06-19 02:00:08 +00:00
Thanh Ha	bc9bd2a766	Use linux.2xlarge runner (#156351 ) The cuda version of this job uses a linux.2xlarge here so matching that to see if this job really needs a 12xlarge system or not. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156351 Approved by: https://github.com/jeffdaily, https://github.com/cyyever	2025-06-19 01:50:56 +00:00
bobrenjc93	e5a1197191	Fix fx tracing for mark dynamic (#156346 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156346 Approved by: https://github.com/tony-ivchenko	2025-06-19 01:03:09 +00:00
Ben Koopman	6959b5febe	Context on torch.cuda.memory._record_memory_history max_entries (#155889 ) Context on torch.cuda.memory._record_memory_history buffer behavior ## Description Answer questions: - Can I keep _record_memory_history() always enabled with the default max_entries=sys.maxsize (9223372036854775807)? Will it consume a significant amount of CPU RAM? - If I set max_entries to a lower value, e.g. 2000, will it keep the first 2000 entries and then stop recording or will it keep the most recent 2000 entries before each snapshot (fifo-style)? - What is the expected size on disk of the snapshots? Some KBs, MBs? Fixes #129674 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155889 Approved by: https://github.com/ngimel	2025-06-19 00:44:43 +00:00
Jeff Daily	6303cc41b7	[ROCm] support CUDA_KERNEL_ASSERT using abort() (#155262 ) We won't have the full message that __assert_fail would provide, but at least we won't silently do nothing. Fixes #155045. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155262 Approved by: https://github.com/hongxiayang, https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-18 23:52:35 +00:00
Feng Shi	b8c2d4c259	add a corner test case of dynamic sizes for combo kernel (#156035 ) Summary: Added a unit test case for a corner case of combo kernel where all below are true: 1. more than 1 dimensions are dynamic size 2. no_x_dim presistent reduce op Test Plan: ``` buck2 test mode/opt caffe2/test/inductor:combo_kernels -- test_dynamic_shapes_persistent_reduction_no_x_dim_2 ``` Rollback Plan: Differential Revision: D76699002 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156035 Approved by: https://github.com/mlazos	2025-06-18 22:57:09 +00:00
Scott Wolchok	76d07e919f	Unbreak //c10/util:base (#156216 ) Missing dep. Bifferential Revision: [D76840057](https://our.internmc.facebook.com/intern/diff/D76840057/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156216 Approved by: https://github.com/janeyx99, https://github.com/desertfire	2025-06-18 22:44:20 +00:00
Meet Vadakkanchery	9bfefda296	[DCP][PyTorch Staging APIs][2/x] Handle 0-elem case + ShardedTensor copy for staging (#156092 ) Summary: ### Diff Context 1. Sometimes, a tensor might have non-zero size and 0 numel. In this case, pinning memory will fail so we take a best guess at how to replicate the tensor below to maintain symmetry in the returned state dict. 2. ShardedTensor copying was not handled originally in PyTorch state_dict copy APIs, handled in this diff. Test Plan: CI Differential Revision: D75553096 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156092 Approved by: https://github.com/pradeepfn	2025-06-18 22:41:25 +00:00
dolpm	a5b4463d60	[nativert] session state (#156190 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D76827309 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156190 Approved by: https://github.com/zhxchen17	2025-06-18 22:40:44 +00:00
Yiming Zhou	6918758f55	[export] Update documents for ExportGraphSiganture (#156244 ) Summary: Fixes https://github.com/pytorch/pytorch/issues/156184 The current document for ExportGraphSignature doesn't reflect `torch.export.export()` returns non-functional graph by default. And users may get confused. Test Plan: Document change only. CI Rollback Plan: Differential Revision: D76849097 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156244 Approved by: https://github.com/yushangdi	2025-06-18 22:37:34 +00:00
Justin Chu	1e474cc9c8	[ONNX] Fix how shapes are computed for float4 (#156353 ) Changed the way we compute shapes for unpacked float4. Previously we always added a last dimension [2] to existing shape, but this doesn't really make sense because it prevents use from being able to represent any shape other than those with a list dim [2]. I updated the logic to be `[shape[:-1], shape[-1]2]` which doubles the last dimension. This is more in line with what we see in practice when people are using 4bit types, and it allows us to represent any shape with an even dimension at the end, which is much more reasonable in my opinion. Also clarified in https://github.com/pytorch/pytorch/pull/148791#discussion_r2155395647 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156353 Approved by: https://github.com/titaiwangms	2025-06-18 22:28:02 +00:00
henrylhtsang	9afee0fa96	[inductor] Set num_workers to number of available cpu divided by number of available gpu (#156201 ) internal: https://fb.workplace.com/groups/1075192433118967/posts/1689562705015267/?comment_id=1690284241609780&notif_id=1749770611538976&notif_t=work_group_comment&ref=notif Right now it doesn't have the divided by 2 logic yet. Not sure how to tell if we are on a dev machine. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156201 Approved by: https://github.com/masnesral	2025-06-18 22:15:32 +00:00
anwang	e5a0b73ce9	[MTIA Aten Backend] Migrate logical_and.out (#156286 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate logical_and.out to in-tree Differential Revision: [D76874551](https://our.internmc.facebook.com/intern/diff/D76874551/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156286 Approved by: https://github.com/nautsimon, https://github.com/jingsh ghstack dependencies: #155634, #156046, #156047, #156283, #156284, #156285	2025-06-18 21:57:05 +00:00
PyTorch MergeBot	bfccfa0b31	Revert "[Draft][CUDA] Use runtime driver API for cuStreamWriteValue32 (#156097 )" This reverts commit cf90c9f8d1632777ec5f4b6ccaa14bc5bf259e9c. Reverted https://github.com/pytorch/pytorch/pull/156097 on behalf of https://github.com/atalman due to break internal tests ([comment](https://github.com/pytorch/pytorch/pull/156097#issuecomment-2985785811))	2025-06-18 21:48:50 +00:00
dolpm	f5eb42e4c0	[nativert] move layoutplanneralgorithm to libtorch (#156205 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D76831634 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156205 Approved by: https://github.com/zhxchen17	2025-06-18 21:46:38 +00:00
anwang	d1c924c68a	[MTIA Aten Backend] Migrate lt.Tensor_out / lt.Scalar_out (#156285 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate t.Tensor_out / lt.Scalar_out to in-tree. Differential Revision: [D76873997](https://our.internmc.facebook.com/intern/diff/D76873997/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156285 Approved by: https://github.com/nautsimon ghstack dependencies: #155634, #156046, #156047, #156283, #156284	2025-06-18 21:40:26 +00:00
anwang	5c7e1d39ab	[MTIA Aten Backend] Migrate logit (#156284 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate logit to in-tree. Differential Revision: [D76871451](https://our.internmc.facebook.com/intern/diff/D76871451/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156284 Approved by: https://github.com/nautsimon ghstack dependencies: #155634, #156046, #156047, #156283	2025-06-18 21:36:27 +00:00
anwang	706e236b08	[MTIA Aten Backend] Migrate logical_or.out / log.out / log2.out (#156283 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate logical_or.out / log.out / log2.out to in-tree. Differential Revision: [D76857072](https://our.internmc.facebook.com/intern/diff/D76857072/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156283 Approved by: https://github.com/nautsimon ghstack dependencies: #155634, #156046, #156047	2025-06-18 21:27:58 +00:00
anwang	ab81fb846c	[MTIA Aten Backend] Migrate remainder.Tensor_out / reciprocal.out / neg.out (#156047 ) Migrate remainder.Tensor_out / reciprocal.out / neg.out Differential Revision: [D76696710](https://our.internmc.facebook.com/intern/diff/D76696710/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156047 Approved by: https://github.com/nautsimon ghstack dependencies: #155634, #156046	2025-06-18 21:17:34 +00:00
anwang	c26ce593d8	[MTIA Aten Backend] Migrate nan_to_num.out (#156046 ) Migrate nan_to_num.out Differential Revision: [D76696155](https://our.internmc.facebook.com/intern/diff/D76696155/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156046 Approved by: https://github.com/nautsimon ghstack dependencies: #155634	2025-06-18 21:14:13 +00:00
anwang	2f1c5c4131	[MTIA Aten Backend] Achieve CPU fallback by overriding registration (#155634 ) # Context MTIA supports CPU fallback, and people can set it using env vars. By migrating aten backend to in-tree, we also need to provide this support. # This diff Suggested by Alban(pytorch core), instead of skipping registration, this diff achieves CPU fallback by doing additional registration and override. The benefits of this approach: 1. The previous solution has problem handling ops that have default dispatch key(e.g. CompositeImplicitAutograd), and can't really achieve CPU fallback. 2. The CPU fallback related logic can be aggregated in aten_mtia_cpu_fallback.cpp. ---------------- p.s. D76314740 also tried reusing the yaml parsing logic in mtia's python script, but realized that the env vars are only available in runtime but not compile/codegen time Differential Revision: [D76376644](https://our.internmc.facebook.com/intern/diff/D76376644/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155634 Approved by: https://github.com/nautsimon, https://github.com/albanD	2025-06-18 21:10:18 +00:00
Mu-Chu Lee	e99cc126a4	[AOTInductor] Reuse input information instead of directly applying unbacked_symint_fallback (#156133 ) Summary: When we encounter unbacked symint during autotuning, we try to reuse existing symbols from user provided inputs, then fallback. Test Plan: python test/inductor/test_aot_inductor.py -k test_triton_dynamic_launcher_grid Rollback Plan: Differential Revision: D76769711 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156133 Approved by: https://github.com/jingsh	2025-06-18 20:53:21 +00:00
PyTorch MergeBot	728cf6721e	Revert "[PT2]load dense delta by trimming prefixes (#155872 )" This reverts commit c74fd35050a7241f0c439501ef735aa6cdde751f. Reverted https://github.com/pytorch/pytorch/pull/155872 on behalf of https://github.com/malfet due to Broke lint, internal has been backed out ([comment](https://github.com/pytorch/pytorch/pull/155872#issuecomment-2985542895))	2025-06-18 20:05:56 +00:00
Kevin Fu	c74fd35050	[PT2]load dense delta by trimming prefixes (#155872 ) Summary: In PT2 with GPU with AOTI, weight names are like ```merge.submod_0._run_on_acc_0.main_module.user_embedding_arch.relevance_pmas.ig_feed.pos_emb``` but when publishing delta snapshots, lowering is skipped so weights are like ```merge.main_module.user_embedding_arch.relevance_pmas.ig_feed.pos_emb``` so when loading delta weights in original model runner, we need to: 1. Redo tensorName -> weight idx look up, because the weight ordering may be different. 2. use trimmed tensorName to find the correct weight path. Note that with this diff, delta snapshot loading still does NOT use xl weights. This should be fine for now as we are still publishing full model with non-xl weights. Test Plan: Merge only: ``` MODEL_TYPE=mtml_ctr_instagram_model MODULE=merge MODEL_ENTITY_ID=900234243 SNAPSHOT_ID=7 DENSE_DELTA_SNAPSHOT_ID=13 CUDA_VISIBLE_DEVICES=2,3 buck2 run mode/dev-nosan -c fbcode.nvcc_arch=a100,h100 -c fbcode.enable_gpu_sections=true caffe2/torch/fb/model_transform/fx2trt/packaging:load_net_predictor -- --loadMode=DenseOnly --baseNetFile=/data/users/$USER/models/${MODEL_ENTITY_ID}/${SNAPSHOT_ID}/${MODEL_ENTITY_ID}_${SNAPSHOT_ID}.predictor.disagg.gpu.${MODULE} --moduleName=${MODULE} --predictor_hardware_type 1 --submodToDevice "" --deltaNetFile /data/users/$USER/models/${MODEL_ENTITY_ID}/${SNAPSHOT_ID}/delta_${DENSE_DELTA_SNAPSHOT_ID}/${MODEL_ENTITY_ID}_${SNAPSHOT_ID}.predictor.disagg.gpu.${MODULE} ``` Local replayer: ``` MODEL_TYPE=mtml_ctr_instagram_model MODEL_ENTITY_ID=900234243 SNAPSHOT_ID=7 DENSE_DELTA_SNAPSHOT_ID=13 USE_SERVABLE=0 HARDWARE_TYPE=0 DENSE_DELTA_IDS=${DENSE_DELTA_SNAPSHOT_ID} ENABLE_REALTIME_UPDATE=1 CUDA_VISIBLE_DEVICES=6,7 sh ./sigrid/predictor/scripts/start_gpu_with_gif.sh ${MODEL_ENTITY_ID}_${SNAPSHOT_ID} /data/users/$USER/models/${MODEL_ENTITY_ID}/${SNAPSHOT_ID} 7455 USE_SERVABLE=0 sh sigrid/predictor/scripts/start_gpu_replayer_localhost_with_gif.sh ${MODEL_ENTITY_ID}_${SNAPSHOT_ID} 10 ${MODEL_TYPE} /data/users/$USER/requests/filter_requests_mtml_ctr_instagram_model_500 localhost /data/users/$USER/models/${MODEL_ENTITY_ID}/${SNAPSHOT_ID} true 7455 ``` Rollback Plan: Differential Revision: D76520301 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155872 Approved by: https://github.com/SherlockNoMad	2025-06-18 19:13:22 +00:00
Alexander Zhipa	48de3da253	fix: avoid flamegraph script setup conflicts (#156310 ) Fixes #156309 Instead of any kind of locking and busy waits leaving room for multiple script downloads to happen, while only one `rename` will succeed and others will silently fail, removing any temporary files created during this process. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156310 Approved by: https://github.com/malfet Co-authored-by: Alexander Zhipa <azzhipa@amazon.com>	2025-06-18 19:06:22 +00:00
Luca Wehrstedt	cbafba5794	Allow forcing FSDP2 to always use SUM reductions (#155915 ) NCCL zero-copy support only works for SUM reductions. FSDP2, by default, was prefering AVG reductions or, when using `set_reduce_scatter_divide_factor`, PreMulSum reductions. Moreover, PreMulSum reductions had a few bugs, such as #155903 and #155904. This PR adds a flag to always use SUM reductions, potentially requiring separate pre-/post-scaling kernels, and reworks the `set_reduce_scatter_divide_factor` logic to make it safer (and renaming it to avoid confusion). Differential Revision: [D76895058](https://our.internmc.facebook.com/intern/diff/D76895058) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155915 Approved by: https://github.com/xunnanxu	2025-06-18 18:57:47 +00:00
windsonsea	9944cd0949	Convert to markdown: quantization-accuracy-debugging.rst, quantization-backend-configuration.rst, quantization-support.rst, random.rst (#155520 ) Related to #155032 - ✅ quantization-accuracy-debugging.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/quantization-accuracy-debugging.html) vs [main](https://docs.pytorch.org/docs/main/quantization-accuracy-debugging.html) - ✅ quantization-backend-configuration.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/quantization-backend-configuration.html) vs [main](https://docs.pytorch.org/docs/main/quantization-backend-configuration.html) - ✅ quantization-support.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/quantization-support.html) vs [main](https://docs.pytorch.org/docs/main/quantization-support.html) - ✅ random.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/random.html) vs [main](https://docs.pytorch.org/docs/main/random.html) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155520 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-18 18:46:04 +00:00
Jeff Daily	30d3cf62fb	support CUBLASLT_MATMUL_MATRIX_SCALE_OUTER_VEC_32F (#154680 ) Requires CUDA >= 12.9 and sm_90. hipBLASLt has a similar enum but is not available until ROCm 7.0. Support the new enum early using a cmake test. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154680 Approved by: https://github.com/malfet, https://github.com/atalman	2025-06-18 18:39:01 +00:00
xinan.lin	aee2bfc5ba	[Intel GPU] Update xpu triton commit pin for PyTorch release 2.8. (#154194 ) As title. Thanks @anmyachev for the work on compatibility adaptation. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154194 Approved by: https://github.com/jansel	2025-06-18 18:17:07 +00:00
Ayan Das	2620361d19	Add batching rule for torch.matrix_exp (#155202 ) ## Summary Adds the missing batching rule for `torch.matrix_exp` to enable efficient `vmap` support. Previously, using `vmap` with `matrix_exp` would trigger a performance warning and fall back to a slow loop-based implementation, even though `matrix_exp` natively supports batched inputs. Fixes #115992 ## Details `torch.matrix_exp` is an alias for `torch.linalg.matrix_exp`. This PR adds vmap support by registering `matrix_exp` with `OP_DECOMPOSE`, which reuses the existing CompositeImplicitAutograd decomposition to automatically generate batching behavior from the operation's simpler component operations. ## Testing The existing test suite for vmap and matrix_exp should cover this change. The fix enables: - No performance warning when using `vmap(torch.matrix_exp)` - Efficient native batched execution instead of loop-based fallback Edit: Updated Details section to accurately reflect the implementation approach (decomposition rather than batch rule registration) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155202 Approved by: https://github.com/zou3519	2025-06-18 17:35:35 +00:00
eqy	a5f59cc2ea	[cuDNN][64-bit indexing] update conv depthwise 64bit indexing dispatch condition to match native kernel (#156140 ) The native kernel doesn't support batch splitting so the previous check wasn't aggressive enough in dispatching to cuDNN https://github.com/pytorch/pytorch/issues/155225 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156140 Approved by: https://github.com/ngimel	2025-06-18 17:32:36 +00:00
PyTorch MergeBot	94f8679019	Revert "[PT2][partitioners] raise getitems in partitioners to allow earlier release of buffers (#155809 )" This reverts commit 6d3a4356f61b28a14abd95f641e2615deb186365. Reverted https://github.com/pytorch/pytorch/pull/155809 on behalf of https://github.com/laithsakka due to pr_time_benchmarks ([comment](https://github.com/pytorch/pytorch/pull/155809#issuecomment-2985022572))	2025-06-18 16:52:19 +00:00
Nikita Shulga	36f7a027b5	[MPS] Implement upsample_trilinear as Metal shader (#156263 ) But only forward for now Pull Request resolved: https://github.com/pytorch/pytorch/pull/156263 Approved by: https://github.com/dcci ghstack dependencies: #156256, #156090	2025-06-18 16:10:02 +00:00
charan-ponnada	bf06190e21	Integrated AMD AWS runners into Pytorch CI (#153704 ) Integrated AMD AWS runners into PyTorch CI, including the linux.24xl.amd for performance tests, the linux.8xl.amd with AVX512 support for unit and periodic tests, and the linux.12xl.amd with AVX2 support for unit and periodic tests. Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/153704 Approved by: https://github.com/malfet, https://github.com/jithunnair-amd Co-authored-by: kiriti-pendyala <kiriti.pendyala@amd.com>	2025-06-18 15:58:22 +00:00
PyTorch MergeBot	ce3406817d	Revert "[dynamo] control one_graph behavior additionally through config (#154283 )" This reverts commit fe37db4f1270745d6c523623143332ddf263af55. Reverted https://github.com/pytorch/pytorch/pull/154283 on behalf of https://github.com/atalman due to inductor/test_flex_decoding.py::TestFlexDecodingCUDA::test_do_not_trigger_dynamic_shapes_on_empty_block_mask_cuda GH job link HUD commit link ([comment](https://github.com/pytorch/pytorch/pull/154283#issuecomment-2984795214))	2025-06-18 15:53:32 +00:00
PyTorch MergeBot	c5d3e7a4ff	Revert "[dynamo] add set_fullgraph decorator/context manager (#154289 )" This reverts commit 920f6e681ec70b664ed952255b8c1f97962f5de0. Reverted https://github.com/pytorch/pytorch/pull/154289 on behalf of https://github.com/atalman due to inductor/test_flex_decoding.py::TestFlexDecodingCUDA::test_do_not_trigger_dynamic_shapes_on_empty_block_mask_cuda GH job link HUD commit link ([comment](https://github.com/pytorch/pytorch/pull/154289#issuecomment-2984774814))	2025-06-18 15:51:06 +00:00
PyTorch MergeBot	408d9884b0	Revert "[dynamo] fix set_fullgraph for nested calls (#154782 )" This reverts commit 3c8c48f79344356c58e91b9c8588f85ff806e1c8. Reverted https://github.com/pytorch/pytorch/pull/154782 on behalf of https://github.com/atalman due to inductor/test_flex_decoding.py::TestFlexDecodingCUDA::test_do_not_trigger_dynamic_shapes_on_empty_block_mask_cuda GH job link HUD commit link ([comment](https://github.com/pytorch/pytorch/pull/154782#issuecomment-2984764330))	2025-06-18 15:47:21 +00:00
PyTorch MergeBot	6201981f48	Revert "[dynamo] handle fullgraph toggle using nested torch.compile (#155166 )" This reverts commit 614a41514545cbdd15757ef2586d433d7d34041c. Reverted https://github.com/pytorch/pytorch/pull/155166 on behalf of https://github.com/atalman due to inductor/test_flex_decoding.py::TestFlexDecodingCUDA::test_do_not_trigger_dynamic_shapes_on_empty_block_mask_cuda [GH job link](https://github.com/pytorch/pytorch/actions/runs/15726606697/job/44333233942) [HUD commit link](`a6a3a44144`) ([comment](https://github.com/pytorch/pytorch/pull/155166#issuecomment-2984751600))	2025-06-18 15:43:22 +00:00
Tugsbayasgalan (Tugsuu) Manlaibaatar	d290fe7690	Remove legacy export testing path (#156093 ) Summary: After this diff stack lands, we are pretty much done with the training IR migration. So there is no need to run extensive legacy export test. Test Plan: CI Rollback Plan: Differential Revision: D76734378 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156093 Approved by: https://github.com/desertfire	2025-06-18 15:36:44 +00:00
Jeff Daily	7531bd6491	[ROCm] upgrade to 6.4.1 patch release (#156112 ) Fixes #155292. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156112 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-18 15:21:44 +00:00
Aleksandar Samardžić	830a335a7d	Refine alignment check along dynamic dimension for grouped MMs (#155466 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155466 Approved by: https://github.com/ngimel	2025-06-18 15:15:05 +00:00
Xuan Zhang	6d3a4356f6	[PT2][partitioners] raise getitems in partitioners to allow earlier release of buffers (#155809 ) Problem & Solution: Assume we have something like: ``` x = some_op(...) x0 = x[0] do_something_with_and_is_last_use_of(x0) do_a_bunch_of_other_things() x1 = x[1] ``` In this case, the memory associated with `x0` cannot be released until `x1 = x[1]`. Since `x1 = x[1]` does not use additional memory, it would be beneficial to move and `x1 = x[1]` and all such `getitem` operations to be immediately after `x = some_op(...)` such as ``` x = some_op(...) x0 = x[0] x1 = x[1] do_something_with_and_is_last_use_of(x0) do_a_bunch_of_other_things() ``` Results: For instance, for the `res2net101_26w_4s` model in pytorch benchmark, when running with `aot_eager` backend and with `activation_memory_budget=0.4`, the peak memory are * baseline: 7.73GiB * with the chage: 6.45GiB As a sanity check, for the same setting with `inductor` backend, the peak memory is not regressed. cc and credit to @ShatianWang for noticing this issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155809 Approved by: https://github.com/fmassa, https://github.com/bdhirsh ghstack dependencies: #155943	2025-06-18 14:38:55 +00:00
Pearu Peterson	c177abd217	Disable pinning check when loading sparse tensors (#154638 ) Disables pinning check as unnecessary and to fix https://github.com/pytorch/pytorch/issues/153143 when loading sparse tensor from external storage with sparse tensor invariants check enabled. Fixes https://github.com/pytorch/pytorch/issues/153143 . For FC, to be landed two weeks after https://github.com/pytorch/pytorch/pull/154617, see https://github.com/pytorch/pytorch/pull/154617#issuecomment-2919643612. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154638 Approved by: https://github.com/amjames, https://github.com/ngimel	2025-06-18 14:33:36 +00:00
PyTorch MergeBot	8f02161d10	Revert "[dynamo] raise hard error if error is encountered while tracing resume function prologue (#154564 )" This reverts commit a6a3a441442a96f38d0771c985f753223cea2ba0. Reverted https://github.com/pytorch/pytorch/pull/154564 on behalf of https://github.com/atalman due to inductor/test_flex_decoding.py::TestFlexDecodingCUDA::test_do_not_trigger_dynamic_shapes_on_empty_block_mask_cuda [GH job link](https://github.com/pytorch/pytorch/actions/runs/15726606697/job/44333233942) [HUD commit link](`a6a3a44144`) ([comment](https://github.com/pytorch/pytorch/pull/154564#issuecomment-2984409088))	2025-06-18 14:19:39 +00:00
Luca Wehrstedt	b30e04b3c8	Make the NCCL PG Options and Config copyable and safe to init standalone (#155700 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155700 Approved by: https://github.com/kwen2501	2025-06-18 13:36:27 +00:00
Xia, Weiwen	1bb9b1858b	[CPU][Inductor] Improve A16W4 GEMM template performance by using block_n=32 (#156174 ) Summary We found that using `block_n=32` brings better performance for A16W4 than `block_n=64` because cache locality is better and parallelism is better if N is small and more cores are used. For example, when running Llama-3.1-8B with A16W4 and batch size = 16 on 43 cores, `block_n=32` is faster by >10% E2E for both first and next token. Test plan ``` pytest test/inductor/test_cpu_select_algorithm.py -k test_int4_woq_mm_amx ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156174 Approved by: https://github.com/leslie-fang-intel	2025-06-18 13:17:46 +00:00
frost-intel	d99cac2816	[Kineto][submodule] Update kineto pin for XPU toggle feature (#155488 ) Part of #154898 Update kineto submodule Summary: We add the toggleCollectionDynamic functionality to XPUPTI in Kineto, so profiler can be enabled/disabled dynamically. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155488 Approved by: https://github.com/guangyey, https://github.com/sraikund16	2025-06-18 12:39:58 +00:00
Aleksei Nikiforov	c11888e7a6	Skip more tests on s390x (#155210 ) Make CI for s390x green before fixing and restoring tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155210 Approved by: https://github.com/seemethere	2025-06-18 12:07:17 +00:00
Xuehai Pan	402ae09e41	[BE] fix typos in c10/ (#156078 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156078 Approved by: https://github.com/malfet, https://github.com/cyyever	2025-06-18 10:24:44 +00:00
David Berard	f45f483884	[user triton] AOT Inductor support for new host-side TMA api (#155879 ) This adds support for the host-side TMA api (TensorDescriptor.from_tensor) for AOTI. Note: this should support all the same features as the old (experimental) TMA api, but not some new features of the new TMA, like mxfp4 support. Note: one complexity with the new TMA api is that a single TMA descriptor passed to the python kernel turns into 1 + 2 * N args in the cubin function signature, for a rank-N tensor. What this PR contains: 1) device_op_overrides.py: add a rough copy of fillTMADescriptor from https://github.com/triton-lang/triton/blob/main/third_party/nvidia/backend/driver.c#L283. However, the fillTMADescriptor implementation in Triton is significantly modified, so that much of the computation (about swizzling and data types) is done before the time of the TMA construction. For simplicity, I've moved the computation into the cuda helper kernel (as was the previous strategy with fill2DTMADescriptor); but long term we might want to unify our implementation with the upstream implementation 2) device_op_overrides.py: introduces a struct "StableTMADescriptor" which stores some of the 1 + 2 * N args for the cubin signature (along with the global shape, which is not strictly needed, but this cleans up the call to the triton kernel 3) plumbing through cpp_wrapper_gpu.py. The main thing to note is: the code generated by cpp_wrapper_gpu.py generally refers to the StableTMADescriptor object when it passes around a "tma descriptor" variable. At the very end (in generate_args_decl), the StableTMADescriptor is unwrapped and the individual arguments are passed into the cubin. Tests: test_aot_inductor.py's test_triton_kernel_tma_descriptor_{N}d_dynamic_{D}_tma_version_{V}_cuda: for N in {1, 2} and D in {True, False}, and V = {new, old}, this test passes (or is skipped, if the appropriate TMA API is not available). Tested on H100 for Triton 3.3 and Triton 3.4. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155879 Approved by: https://github.com/desertfire	2025-06-18 09:35:11 +00:00
Junjie Wang (PyTorch)	577baa4116	[c10d] Add a logger for all nccl collectives with its time duration when completed (#156008 ) Summary: We want to build a logging table for tracking the collective time spent on GPU for all internal workloads. Since we have a cudaEventQuery for both the start and end of a collective (We rolled out ECudaEventStart (enableTiming) fully already), we plan to add this logging table inside the watchdog of PyTorch ProcessGroupNCCL so that we get to know the duration of collectives. Test Plan: CI + dry run. Rollback Plan: Differential Revision: D76552340 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156008 Approved by: https://github.com/fegin, https://github.com/eqy	2025-06-18 09:08:42 +00:00
Wang, Chuanqi	c5a4fe9c17	[CI] fix the ci image name for public copy in ghcr (#156169 ) After the PR #152209 landed, the name of ci image public copy in ghcr is not correct. For example, https://github.com/pytorch/pytorch/actions/runs/15698468716/job/44228133522#step:10:8. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156169 Approved by: https://github.com/malfet	2025-06-18 08:16:56 +00:00
William Wen	a6a3a44144	[dynamo] raise hard error if error is encountered while tracing resume function prologue (#154564 ) This should prevent bad resume function prologues from slipping by. In particular, graph breaks in resume function prologues will now hard error. Implementation details: - The resume function prologue is surrounded by `LOAD_CONST arg, STORE_FAST __is_tracing_resume_prologue` instructions. The first sequence has `arg=True` and the second sequence has `arg=False`. - InstructionTranslator will know when it is tracing a resume function prologue when it detects `STORE_FAST __is_tracing_resume_prologue`. The top of stack will be True to mark the start of the prologue, False to mark the end. - When `convert_frame.py` detects that an error occurred while the InstructionTranslator was tracing a resume function prologue, we will wrap the exception and hard error Pull Request resolved: https://github.com/pytorch/pytorch/pull/154564 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289, #154782, #155166	2025-06-18 07:27:20 +00:00
William Wen	614a415145	[dynamo] handle fullgraph toggle using nested torch.compile (#155166 ) See added test for the case that this PR handles. In particular, the semantics for nested torch.compile with toggled fullgraph settings was strange before - `@torch.compile(fullgraph=True)` overrides the existing fullgraph setting, while `@torch.compile(fullgraph=False)` does not. Note that this change will add an extra frame to any inlined torch.compile'd function (which I don't expect to happen frequently). Pull Request resolved: https://github.com/pytorch/pytorch/pull/155166 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289, #154782	2025-06-18 07:27:20 +00:00
William Wen	3c8c48f793	[dynamo] fix set_fullgraph for nested calls (#154782 ) - Make the fullgraph argument of set_fullgraph a positional argument - Fix behavior on nested calls by updating `tracer.error_on_graph_break` in more places. In particular, a tracer's error_on_graph_break is set to the inlined tracer's error_on_graph_break upon the latter's exit. We also track error_on_graph_break in the speculation log now, since if we encounter a nested graph break, we will restart analysis and we need to somehow remember the error_on_graph_break setting after attempting to run the nested function (but we don't actually trace into it in the restart analysis). Pull Request resolved: https://github.com/pytorch/pytorch/pull/154782 Approved by: https://github.com/jansel ghstack dependencies: #154283, #154289	2025-06-18 07:27:09 +00:00
William Wen	920f6e681e	[dynamo] add set_fullgraph decorator/context manager (#154289 ) Implements https://github.com/pytorch/pytorch/issues/144908. Implementation notes: - `set_fullgraph` is implemented using `patch_config`, which changes config correctly during runtime and tracing. - Moved setting `config.error_on_graph_break` from convert_frame.py to eval_frame.py. This is because this should only be done at the top-level decorated function. If we kept this in convert_frame.py, we would be changing `config.error_on_graph_break` on every top-level frame, which causes confusing behavior (see added test for example). - InstructionTranslator reads from `config.error_on_graph_break` every `step()`. This is to determine the value of `config.error_on_graph_break` at the time of the graph break, because tracer cleanup will restore the value of `config.error_on_graph_break` . - `convert_frame.py` determines whether we should abort tracing (fullgraph=True) or continue (fullgraph=False) by reading the value of the tracer's `error_on_graph_break`. If there is no tracer (failed to initialize), then default to reading `config.error_on_graph_break`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154289 Approved by: https://github.com/jansel, https://github.com/zou3519 ghstack dependencies: #154283	2025-06-18 07:27:00 +00:00
William Wen	fe37db4f12	[dynamo] control one_graph behavior additionally through config (#154283 ) `torch.compile` now always goes through `torch._dynamo._optimize`. fullgraph is now implemented in `torch.compile` by looking at `config.error_on_graph_break`. Export still goes through `torch._dynamo._optimize_assert`, which uses `tx.one_graph` instead of `config.error_on_graph_break`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154283 Approved by: https://github.com/jansel, https://github.com/anijain2305	2025-06-18 07:26:52 +00:00
Brian Hirsh	ccc6279b40	flex attention: fix dispatch order for tensor subclasses, avoid hardcoding call to faketensor impl in dynamo (#151719 ) This is enough to get @XilunWu 's stack in a state where his flex_attention DTensor implementations worked E2E for me. It also required these changes on the DTensor side, to properly add a DTensor rule for flex backward: P1789852198 There are two problems: (1) in the normal dispatcher, we have a precedence ordering between modes and subclasses. Modes are dispatched to first, but modes are allowed to return NotImplemented, giving subclasses a chance to run. This normally happens automatically in `FakeTensorMode.__torch_dispatch__` and `FunctionalTensorMode.__torch_dispatch__`. However, since HOPs implement these two modes themselves, HOPs do not get this benefit. For now, I ended up hardcoding this `NotImplemented` logic directly into the functional/fake rules for flex attention. Having to do this for every HOP seems a bit painful. If we could plumb every HOP through `Fake[\|Functional]TensorMode.__torch_dispatch__` then we would get this support. Another option could be to just assume that most HOP <> mode implementations want the same treatment by default, and hardcode this `NotImplemented` logic into `torch/_ops.py`. I'm not sure if we'd need a way for the HOP to opt out of this though. (2) We were hardcoding a call to flex attention's fake implementation in dynamo to run fake prop. This is technically wrong for subclasses, because it doesn't give subclasses the chance to interpose on the op and desugar it before fake prop runs. I tweaked dynamo's logic to call the op, and let the dispatcher handle invoking the fake implementation. Testing Xilun is adding some DTensor tests in his PR that will end up testing this logic. If folks would prefer, though, I can try to add a test that uses another subclass instead that is maybe more basic. This is the tlparse that his DTensor test gnerated for me: https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/hirsheybar/0196c1d3-a9a2-46ea-a46d-aa21618aa060/custom/rank_0/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151719 Approved by: https://github.com/ydwu4 Co-authored-by: drisspg <drisspguessous@gmail.com>	2025-06-18 07:02:04 +00:00
Ruben Rodriguez Buchillon	bdb1553b77	[inductor][cutlass] binary remote cache (#156248 ) Summary: # Why speed up cutlass kernel generation and retrieval # What using the _ManifoldCache, make a KernelBinaryCache that uploads/downloads kernels and their error files. only register the handler internally this is the OSS only part of the change, to facilitate integration Test Plan: ## prove that we can upload successfully ``` buck2 run @mode/opt scripts/coconutruben/torchmm:experiment 2>&1 ``` ``` manifold ls coconutruben-test-01/tree/cutlass_concept_2 673184 cfkykew2fw5572hjr4e7jbog7oix7xjkegtn2ovikyhxe6pr4tcw.so 649776 cpjqda67c6ojj75z3ddnmfbxinpm7yp7rc2q2oxwsrtwsnacklqv.so ``` ## prove that we can download successfully ``` buck2 run @mode/opt scripts/coconutruben/torchmm:experiment 2>&1 ``` ``` I0611 12:48:38.759000 935012 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:65] Successfully downloaded /var/tmp/torchinductor_coconutruben/fk/cfkykew2fw5572hjr4e7jbog7oix7xjkegtn2ovikyhxe6pr4tcw.so I0611 12:48:38.760000 935012 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:65] Successfully downloaded /var/tmp/torchinductor_coconutruben/pj/cpjqda67c6ojj75z3ddnmfbxinpm7yp7rc2q2oxwsrtwsnacklqv.so ``` ## prove that we can upload errors successfully ``` buck2 run @mode/opt scripts/coconutruben/torchmm:experiment 2>&1 ``` ``` manifold ls coconutruben-test-01/tree/cutlass_concept_2 4846 cqiq4vjbvytdofutoxisa3pqjplgpgmt2sh7dtatiw4bqt5rtjgc.so.error 4846 cqymdwsfsirhkqglv7sbjyvqkrt3ryql4mtb45tekt76347ee6sx.so.error ``` ## prove that we can download errors successfully ``` buck2 run @mode/opt scripts/coconutruben/torchmm:experiment 2>&1 ``` ``` I0611 12:56:14.078000 1001022 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:74] Successfully downloaded /var/tmp/torchinductor_coconutruben/qi/cqiq4vjbvytdofutoxisa3pqjplgpgmt2sh7dtatiw4bqt5rtjgc.so.error I0611 12:56:14.079000 1001022 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:74] Successfully downloaded /var/tmp/torchinductor_coconutruben/qy/cqymdwsfsirhkqglv7sbjyvqkrt3ryql4mtb45tekt76347ee6sx.so.error ``` ## showing timing information ``` I0616 11:22:29.169000 2249769 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:71] Successfully downloaded /var/tmp/torchinductor_coconutruben/fk/cfkykew2fw5572hjr4e7jbog7oix7xjkegtn2ovikyhxe6pr4tcw.so (download: 0.842s, write: 0.000s, total: 0.842s) I0616 11:22:29.169000 2249769 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:71] Successfully downloaded /var/tmp/torchinductor_coconutruben/pj/cpjqda67c6ojj75z3ddnmfbxinpm7yp7rc2q2oxwsrtwsnacklqv.so (download: 0.838s, write: 0.001s, total: 0.838s) ``` Reviewed By: henrylhtsang Pull Request resolved: https://github.com/pytorch/pytorch/pull/156248 Approved by: https://github.com/henrylhtsang	2025-06-18 06:51:22 +00:00
PyTorch UpdateBot	96df866410	[audio hash update] update the pinned audio hash (#156259 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156259 Approved by: https://github.com/pytorchbot	2025-06-18 06:02:46 +00:00
Kaichao You	a5df6ffbc2	Improve IPC for Expandable Segments to use fabric handle when possible (#156074 ) Improve upon https://github.com/pytorch/pytorch/pull/130890 , inspired by https://github.com/pytorch/pytorch/pull/130890#issuecomment-2278882984 , we can automatically use the fabric handle for IPC when possible. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156074 Approved by: https://github.com/ngimel, https://github.com/malfet	2025-06-18 05:22:06 +00:00
henrylhtsang	29867b211a	[cutlass backend] Add __init__.py to cutlass_lib_extensions (#156234 ) When using docker with cutlass backend, we can get ``` No module named 'torch._inductor.codegen.cuda.cutlass_lib_extensions' ``` First reported by @nWEIdia in https://github.com/pytorch/pytorch/issues/155888 Evidence that this fixes: https://github.com/pytorch/pytorch/pull/156136 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156234 Approved by: https://github.com/mlazos, https://github.com/Skylion007	2025-06-18 05:03:43 +00:00
Nikita Shulga	c28e74e457	[MPS] Add nearest_3d forward and backward (#156090 ) Introduce generalizable `UpsampleParams` structure in `UpSample.h`, which could be shared between CPU and MPS Delete `upsample_nearest3d` MPS fallback and replace it with proper shader Pull Request resolved: https://github.com/pytorch/pytorch/pull/156090 Approved by: https://github.com/kulinseth, https://github.com/dcci ghstack dependencies: #156256	2025-06-18 04:48:15 +00:00
functionstackx	a82c171bb2	remove skipifrocm from composability tests (#156036 ) Porting over DTensor training codebase to rocm atm and was reading through a 2D unit tests and noticed a couple of the unit tests already work on rocm even though it is being skipped. pipeline parallel tests pass too tested locally <img width="561" alt="image" src="https://github.com/user-attachments/assets/7c40c0f2-2de8-4cf1-8e36-0ba2bba46baa" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156036 Approved by: https://github.com/jeffdaily	2025-06-18 04:24:42 +00:00
Daniel Galvez	9ed0060225	Provide access to the cudaGraph_t underlying a CUDAGraph. (#155164 ) There are a few considerations here: 1. A user might want to modify the cudaGraph_t either during the stream capture or after the stream capture (but before instantiation). This draft implements modification after stream capture only, though support could be added for modification during stream capture by applying https://github.com/pytorch/pytorch/pull/140979/files#diff-d7302d133bb5e0890fc94de9aeea4d9d442555a3b40772c9db10edb5cf36a35cR391-R404 2. Previously, the cudaGraph_t would be destroyed before the end of capture_end() unless the user had previously called enable_debug_mode(). There is no way to implement this correctly without removing this restriction, or forcing the user to always call enable_debug_mode(). However, enable_debug_mode() is a confusing API (despite being an instance method, it would modify a static global variable; thus, putting one CUDAGraph object into debug mode puts all of them into debug mode, which is not acceptable in my opinion). Therefore, I made enable_debug_mode() into a no-op. This means that the CPU memory usage will increase after this change. I think this is likely to be fine. 3. No python bindings yet. These should be easy to add. It is probably worthwhile to take some time to make sure that the returned cudaGraph_t can be converted into the cuda-python cudaGraph_t in a reasonable, hopefully type-safe, manner (but without making cuda-python a dependency of pytorch), since I imagine most users will use the pip cuda-python package to make modifications. 4. There are two foot guns: a. The cudaGraph_t returned by raw_cuda_graph() is not owned by the user, so it will be destroyed once the owning CUDAGraph is destroyed (or calls reset()). b. The following seuquence won't work as intended: ``` g = torch.cuda.CUDAGraph() with torch.cuda.graph(g): foo() g.replay() raw_graph = g.raw_cuda_graph() modify(raw_graph) g.replay() ``` This won't work because the user must call instantiate() again after modifying cudaGraph_t. You could add a "safety" mechanism by traversing the cudaGraph_t to create a hash and seeing if the hash changes between calls to replay(), but this is likely way too expensive. I think these two foot guns are probably okay given that this a bit of an experts' API. Fixes #155106 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155164 Approved by: https://github.com/ngimel	2025-06-18 03:39:28 +00:00
Simon Fan	17b38b850e	[ca] Allow using compiled autograd context managers during backward runtime (#156120 ) Added an invariant that nested compiled autograd context managers must exit before their parent context manager. This allows us to defer the thread check. FIXES https://github.com/pytorch/pytorch/issues/152219 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156120 Approved by: https://github.com/jansel ghstack dependencies: #155521, #155480	2025-06-18 03:01:15 +00:00
CaoE	10d41c7d20	Add SDPA patterns for T5 models (#155455 ) * Add SDPA patterns for T5 models. * Remove the stride check of mask, and do contiguous for mask in flash attention when the stride of last dim != 1 & != 0. This allows more SDPAs with complex mask to be accelerated using flash attention, such as the T5 model, where the generated masks may be not continuous. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155455 Approved by: https://github.com/Valentine233, https://github.com/leslie-fang-intel, https://github.com/jansel	2025-06-18 02:09:55 +00:00
Mark Molinaro	4851863e3f	fix hack to check if register_buffer has been overridden (#155963 ) Followup on https://github.com/pytorch/pytorch/pull/125971 `self.register_buffer` will always be a a bound method on the instance (`self`) while `torch.nn.Module.register_buffer` is an unbound class method. `is`-ing these two things will never yield `True`. Instead, lets check the [original function object](https://docs.python.org/3/reference/datamodel.html#method.__func__). Note that the current logic doesn't break anything because the `else` branch will still do the "right thing" in the case `register_buffer` hasn't been overrridden, but it does mean we do less work! Example demonstration: ```python class Base: def register_buffer(self, buffer): pass class InheritedOk(Base): pass class InheritedOverride(Base): def register_buffer(self, buffer): pass b = Base() ok = InheritedOk() override = InheritedOverride() print(f"b.register_buffer is Base.register_buffer: {b.register_buffer is Base.register_buffer}") # False print(f"ok.register_buffer is Base.register_buffer: {ok.register_buffer is Base.register_buffer}") # False print(f"override.register_buffer is Base.register_buffer: {override.register_buffer is Base.register_buffer}") # False print(f"b.register_buffer.__func__ is Base.register_buffer: {b.register_buffer.__func__ is Base.register_buffer}") # True print(f"ok.register_buffer.__func__ is Base.register_buffer: {ok.register_buffer.__func__ is Base.register_buffer}") # True print(f"override.register_buffer.__func__ is Base.register_buffer: {override.register_buffer.__func__ is Base.register_buffer}") # False ``` (I can make an associated issue if needed, but didnt see it required [in the contributing guidelines](https://github.com/pytorch/pytorch/blob/main/CONTRIBUTING.md#merging-your-change)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155963 Approved by: https://github.com/mikaylagawarecki	2025-06-18 01:50:30 +00:00
nirajkamalk	202d2ae53a	Convert rst to md: rpc.rst, signal.rst, size.rst, special.rst (#155430 ) Fixes #155033 - [x] [rpc.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/rpc.rst) - [x] [signal.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/signal.rst) - [x] [size.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/size.rst) - [sparse.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/sparse.rst) fixed in #155438 due to large size. - [x] [special.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/special.rst) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155430 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-18 01:27:04 +00:00
Nicolas Macchioni	68996dc183	[BE][2/X] Phase out usage of `use_max_autotune()` (#155848 ) See #155847 for context Pull Request resolved: https://github.com/pytorch/pytorch/pull/155848 Approved by: https://github.com/masnesral	2025-06-18 01:18:09 +00:00
Jane Xu	e8bfce9a43	Document how to use stack-based APIs with StableIValue (#155984 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155984 Approved by: https://github.com/albanD, https://github.com/zou3519	2025-06-18 01:10:23 +00:00
Nikita Shulga	541297daae	[Build] Allow metal shaders to include ATen headers (#156256 ) No-op change that will be used later to share structs between CPU and Metal Pull Request resolved: https://github.com/pytorch/pytorch/pull/156256 Approved by: https://github.com/dcci	2025-06-18 01:03:25 +00:00
xinan.lin	3dabc351bb	[Break XPU] Fix XPU UT failures introduced by community. (#156091 ) Fixes #15089, Fixes #156063, Fixes #155689, Fixes #155692, Fixes #156146 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156091 Approved by: https://github.com/jansel	2025-06-17 23:43:37 +00:00
Howard Huang	38e1e5d54c	Add get_pipeline_order() for Gpipe and 1F1B (#155935 ) The [schedule visualizer](https://github.com/pytorch/pytorch/blob/main/torch/distributed/pipelining/_schedule_visualizer.py) relies on `self.pipeline_order` to be populated. The `_PipelineScheduleRuntime` also depends on this to run the IR. The single stage schedules do not implement this so this PR adds that. Also fixes a bug in the schedule visualizer Pull Request resolved: https://github.com/pytorch/pytorch/pull/155935 Approved by: https://github.com/wconstab	2025-06-17 23:39:17 +00:00
bobrenjc93	5435e75399	[ez] rename choice_timings -> choice_timings_fn (#156099 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156099 Approved by: https://github.com/mlazos ghstack dependencies: #155982, #155996, #156053	2025-06-17 23:30:27 +00:00
Manuel Candales	12b02137af	[MPS] Add benchmark for scan operations (#156241 ) Comparison of cumsum performance before and after Metal implementaton: Previous performance (using torch==2.7.1): ```[------------------------------- -------------------------------] \| eager \| compile 1 threads: ------------------------------------------------------- cumsum-dim0-32x32 (torch.float16) \| 131.0 \| 136.9 cumsum-dim0-128x128 (torch.float16) \| 116.9 \| 121.2 cumsum-dim0-512x512 (torch.float16) \| 132.5 \| 151.9 cumsum-dim0-1024x1024 (torch.float16) \| 150.0 \| 163.0 cumsum-dim1-32x32 (torch.float16) \| 125.9 \| 140.9 cumsum-dim1-128x128 (torch.float16) \| 116.4 \| 129.4 cumsum-dim1-512x512 (torch.float16) \| 135.9 \| 150.1 cumsum-dim1-1024x1024 (torch.float16) \| 139.5 \| 154.2 cumsum-1d-100 (torch.float16) \| 119.5 \| 127.1 cumsum-1d-10000 (torch.float16) \| 128.9 \| 142.5 cumsum-1d-1000000 (torch.float16) \| 140.6 \| 145.6 cumsum-dim0-32x32 (torch.float32) \| 115.7 \| 132.5 cumsum-dim0-128x128 (torch.float32) \| 118.0 \| 131.5 cumsum-dim0-512x512 (torch.float32) \| 138.8 \| 151.6 cumsum-dim0-1024x1024 (torch.float32) \| 155.5 \| 164.2 cumsum-dim1-32x32 (torch.float32) \| 127.2 \| 141.7 cumsum-dim1-128x128 (torch.float32) \| 117.7 \| 130.5 cumsum-dim1-512x512 (torch.float32) \| 138.2 \| 152.3 cumsum-dim1-1024x1024 (torch.float32) \| 144.4 \| 158.6 cumsum-1d-100 (torch.float32) \| 118.6 \| 128.0 cumsum-1d-10000 (torch.float32) \| 125.5 \| 141.5 cumsum-1d-1000000 (torch.float32) \| 143.9 \| 158.4 cumsum-dim0-32x32 (torch.bfloat16) \| 106.6 \| 137.6 cumsum-dim0-128x128 (torch.bfloat16) \| 118.1 \| 131.0 cumsum-dim0-512x512 (torch.bfloat16) \| 140.0 \| 154.3 cumsum-dim0-1024x1024 (torch.bfloat16) \| 153.2 \| 164.4 cumsum-dim1-32x32 (torch.bfloat16) \| 127.9 \| 132.6 cumsum-dim1-128x128 (torch.bfloat16) \| 116.5 \| 129.6 cumsum-dim1-512x512 (torch.bfloat16) \| 136.5 \| 151.2 cumsum-dim1-1024x1024 (torch.bfloat16) \| 139.8 \| 144.8 cumsum-1d-100 (torch.bfloat16) \| 115.7 \| 129.4 cumsum-1d-10000 (torch.bfloat16) \| 125.0 \| 143.3 cumsum-1d-1000000 (torch.bfloat16) \| 127.8 \| 143.4 Times are in microseconds (us). ``` Current performance: ``` [-------------------------------- --------------------------------] \| eager \| compile 1 threads: --------------------------------------------------------- cumsum-dim0-32x32 (torch.float16) \| 107.4 \| 123.8 cumsum-dim0-128x128 (torch.float16) \| 134.2 \| 145.8 cumsum-dim0-512x512 (torch.float16) \| 207.3 \| 231.6 cumsum-dim0-1024x1024 (torch.float16) \| 318.9 \| 355.3 cumsum-dim1-32x32 (torch.float16) \| 98.0 \| 114.3 cumsum-dim1-128x128 (torch.float16) \| 110.8 \| 121.6 cumsum-dim1-512x512 (torch.float16) \| 193.0 \| 209.1 cumsum-dim1-1024x1024 (torch.float16) \| 844.7 \| 870.8 cumsum-1d-100 (torch.float16) \| 108.4 \| 125.0 cumsum-1d-10000 (torch.float16) \| 784.7 \| 852.3 cumsum-1d-1000000 (torch.float16) \| 65855.2 \| 66725.9 cumsum-dim0-32x32 (torch.float32) \| 114.7 \| 115.7 cumsum-dim0-128x128 (torch.float32) \| 139.0 \| 151.6 cumsum-dim0-512x512 (torch.float32) \| 197.3 \| 208.0 cumsum-dim0-1024x1024 (torch.float32) \| 312.7 \| 332.9 cumsum-dim1-32x32 (torch.float32) \| 92.0 \| 110.8 cumsum-dim1-128x128 (torch.float32) \| 114.2 \| 125.0 cumsum-dim1-512x512 (torch.float32) \| 186.2 \| 196.1 cumsum-dim1-1024x1024 (torch.float32) \| 752.0 \| 825.0 cumsum-1d-100 (torch.float32) \| 112.4 \| 122.0 cumsum-1d-10000 (torch.float32) \| 793.5 \| 863.5 cumsum-1d-1000000 (torch.float32) \| 66431.8 \| 66040.0 cumsum-dim0-32x32 (torch.bfloat16) \| 111.6 \| 121.6 cumsum-dim0-128x128 (torch.bfloat16) \| 139.0 \| 138.4 cumsum-dim0-512x512 (torch.bfloat16) \| 217.6 \| 230.1 cumsum-dim0-1024x1024 (torch.bfloat16) \| 305.2 \| 325.6 cumsum-dim1-32x32 (torch.bfloat16) \| 100.5 \| 110.9 cumsum-dim1-128x128 (torch.bfloat16) \| 112.8 \| 125.0 cumsum-dim1-512x512 (torch.bfloat16) \| 187.8 \| 208.9 cumsum-dim1-1024x1024 (torch.bfloat16) \| 790.9 \| 864.7 cumsum-1d-100 (torch.bfloat16) \| 111.6 \| 124.6 cumsum-1d-10000 (torch.bfloat16) \| 778.1 \| 844.9 cumsum-1d-1000000 (torch.bfloat16) \| 64654.3 \| 64082.5 Times are in microseconds (us). ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156241 Approved by: https://github.com/malfet	2025-06-17 22:30:22 +00:00
PyTorch MergeBot	fa4f07b5b8	Revert "[Docs] Convert to markdown to fix 155032 (#155520 )" This reverts commit cd66ff80307862ef8e75520054ecd19a5eff9f7e. Reverted https://github.com/pytorch/pytorch/pull/155520 on behalf of https://github.com/atalman due to breaks multiple test_quantization.py::TestQuantizationDocs::test_quantization_ ([comment](https://github.com/pytorch/pytorch/pull/155520#issuecomment-2981996091))	2025-06-17 22:22:50 +00:00
Julian De la Barrera Brandner	54998c2daa	Document padding size limitations in nn.modules.padding (#134840 ) (#155618 ) Fixes #134840 Added documentation to clarify padding size constraints for all padding modes in nn.modules.padding: - Circular padding: size must be less than or equal to the corresponding input dimension - Reflection padding: size must be less than the corresponding input dimension - Replication padding: output dimensions must remain positive These changes help prevent runtime errors when users attempt to use large padding values. ## PR Checklist - [x] The PR title and message follow our [commit guidelines](https://github.com/pytorch/pytorch/blob/main/CONTRIBUTING.md#commit-message-format) - [x] The PR is made against the correct branch - [x] The PR is labeled with `docathon` - [x] The PR is labeled with `module: nn` - [x] The PR is labeled with `documentation` - [x] The PR description includes a reference to the issue being fixed - [x] The PR includes tests if applicable - [x] The PR includes documentation changes - [x] The PR has been tested locally Pull Request resolved: https://github.com/pytorch/pytorch/pull/155618 Approved by: https://github.com/AlannaBurke, https://github.com/malfet	2025-06-17 22:16:48 +00:00
Jane Xu	937529f0b3	Pass by const ref instead of by value in StableIValue from (#156126 ) I realize I was passing stable::Tensors by value (thus making a copy every time) which is not what I want from the `from` function that converts Ts to StableIValues. `from` should not mutate the input and should be read-only. I asked an LLM whether this is API BC breaking (with an intuition that it shouldn't be), and it said no, cuz: 1. "Passing by const reference is more permissive than passing by value. e.g., if T is a type that has a deleted or inaccessible copy constructor (e.g., std::unique_ptr), the original code would have been invalid, while the new code would be valid." Nice. We are good with additive. 2. We didn't modify the original input before (cuz we took a copy) and we don't now (cuz we promise const). Update: The LLM failed to mention primitives, with which we should not pass references around, so we are only changing the signatures of std::optional<T> and stable::Tensor Pull Request resolved: https://github.com/pytorch/pytorch/pull/156126 Approved by: https://github.com/swolchok ghstack dependencies: #155367, #155977	2025-06-17 22:11:30 +00:00
Daniel Galvez	4c0aa37dda	Support stream capture of event record and wait nodes in cuda graphs (#155372 ) These are created by the user passing cudaEventRecordExternal and cudaEventWaitExternal to cudaEventRecordWithFlags() and cudaStreamWaitEvent() respectively. We do this by allowing the user to specify external=True when constructing a torch.cuda.Event(). If external=False, the cudaEventRecord and cudaStreamWaitEvent API's have a different meaning described here: https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#cross-stream-dependencies-and-events In short, they will be used to experess fork and join operations in the graph if external=False. External events can be used for expressing a fine-grained dependency on the outcome of some nodes in a cuda graph (rather than all nodes). They can also be used for timing parts of a cuda graph's execution, rather than timing the entire graph's execution. Finishes #146145 I'm a dummy and don't know how to use ghstack at this time. The first commit is a bug fix for _CudaKernel, which would previously always launch work on the NULL stream, rather than the user-passed stream. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155372 Approved by: https://github.com/ngimel	2025-06-17 21:44:51 +00:00
Oguz Ulgen	8e02cd9c5a	Skip cache related configs for cache config serialization (#156195 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156195 Approved by: https://github.com/masnesral	2025-06-17 21:24:07 +00:00
Junjie Wang (PyTorch)	3106a33e41	[fr] Fix one error in analysis script when subPG world size is smaller than global size (#156156 ) Summary: We run into an interesting case when we see so many mismatches while lot of mismatch turns out to be a fully match. The reason is that we use the dump ranks (which is from 0 to 79) to compare against the local pg ranks (0 to 7) this leads to false positive of mismatches. We can just check whether dump ranks contain all expected ranks or not, that should be sufficient. Test Plan: Test with the failed case with the script and we now see the correct behavior + new unit test case. Rollback Plan: Differential Revision: D76775373 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156156 Approved by: https://github.com/VieEeEw	2025-06-17 21:17:58 +00:00
henrylhtsang	bb462a6237	[cutlass backend] Fix prescreening non-deterministic problem (#156144 ) Differential Revision: [D76642615](https://our.internmc.facebook.com/intern/diff/D76642615/) What do we expect to see when we run two identical matmul back to back? We expect to see the second one spending no time in precompilation, autotuning and prescreening. However, the introduction of prescreening bring some non-deterministics-ness. Basically, we have 1. prescreening of first matmul chooses a set of kernels to advance to autotuning 2. autotuning re-does the autotuning of the winners, potentially changing their timings a bit 3. second prescreening results in a slightly different set of kernels 4. since not all timings are present, an autotune is re-done. With this diff: ``` SingleProcess AUTOTUNE benchmarking takes 3.8633 seconds and 134.7364 seconds precompiling for 32 choices and 24.4472 seconds prescreening SingleProcess AUTOTUNE benchmarking takes 0.0003 seconds and 0.0027 seconds precompiling for 32 choices and 0.0006 seconds prescreening ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156144 Approved by: https://github.com/mlazos	2025-06-17 20:39:06 +00:00
windsonsea	cd66ff8030	[Docs] Convert to markdown to fix 155032 (#155520 ) Fix #155032 - ✅ quantization-accuracy-debugging.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/quantization-accuracy-debugging.html) vs [main](https://docs.pytorch.org/docs/main/quantization-accuracy-debugging.html) - ✅ quantization-backend-configuration.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/quantization-backend-configuration.html) vs [main](https://docs.pytorch.org/docs/main/quantization-backend-configuration.html) - ✅ quantization-support.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/quantization-support.html) vs [main](https://docs.pytorch.org/docs/main/quantization-support.html) - ✅ quantization.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/quantization.html) vs [main](https://docs.pytorch.org/docs/main/quantization.html) - ✅ random.rst: [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155520/random.html) vs [main](https://docs.pytorch.org/docs/main/random.html) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155520 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-17 20:29:45 +00:00
Nicolas Macchioni	50940270ae	[BE][3/X] Phase out usage of `use_max_autotune()` (#155849 ) See #155847 for context Pull Request resolved: https://github.com/pytorch/pytorch/pull/155849 Approved by: https://github.com/masnesral	2025-06-17 20:26:29 +00:00
Xuehai Pan	b020971e78	[BE] fix typos in torchgen/ (#156083 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156083 Approved by: https://github.com/jingsh ghstack dependencies: #156079, #156082	2025-06-17 19:25:50 +00:00
Xuehai Pan	a69785b3ec	[BE] fix typos in tools/ (#156082 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156082 Approved by: https://github.com/soulitzer ghstack dependencies: #156079	2025-06-17 19:25:50 +00:00
Xuehai Pan	ccea6ddac3	[BE] fix typos in cmake/ (#156079 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156079 Approved by: https://github.com/Skylion007	2025-06-17 19:25:43 +00:00
Dmitry Nikolaev	5eb5c3700b	[ROCm] enable batched eigen decomposition (syevD_batched) on ROCm (#154525 ) This PR implements `Batched Eigen Decomposition` (syevD_batched) on ROCm by calling rocSolver directly. cuSolver doesn't support syevD_batched and neither does hipSolver. Direct call to rocSolver is required. `syevD_batched` will be used on ROCm if all the following conditions are met: - `rocSolver version >= 3.26` - input data type is `float` or `double` - batch size >= 2 Otherwise, non-batched `syevD` will be used on ROCm (complex data types, batch size==1, rocSolver <3.26) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154525 Approved by: https://github.com/Mellonta	2025-06-17 19:20:36 +00:00
PyTorch MergeBot	ec08eb8ba2	Revert "[inductor][cutlass] binary remote cache (#156106 )" This reverts commit 9a2c669425379eb264f896390b8fcd8d3f2ce959. Reverted https://github.com/pytorch/pytorch/pull/156106 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/156106#issuecomment-2981533904))	2025-06-17 19:07:49 +00:00
Aidyn-A	4a26bb8a12	[C10][CUDA] Eagerly create context on torch.cuda.set_device(device) call (#155900 ) Fixes #155668 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155900 Approved by: https://github.com/ngimel	2025-06-17 18:59:44 +00:00
Thien Tran	fc177801af	Enable FP8 row-wise scaled-mm for sm12x (#155991 ) ## Update using Cutlass 3.x (2025/06/15) Following @alexsamardzic's advice, I tried out Cutlass 3.x API and it's impressive (rated specs is 419 TFLOPS) M \| N \| K \| TFLOPS ---\|---\|---\|-------- 16\|4096\|4096\|17.56 64\|4096\|4096\|69.63 256\|4096\|4096\|266.57 1024\|4096\|4096\|339.28 4096\|4096\|4096\|388.91 This uses the same SM100 template. The only difference is - Cluster size is fixed to `<1,1,1>` since sm120 does not have multicast feature - ~~Tile size is fixed to `<128,128,128>` due to default kernel schedule does not support `<64,128,128>`. I will work a bit on improve perf for small M.~~ Fixed. Use `KernelTmaWarpSpecializedPingpong` when TileShape.M == 64 Perf for small M is still bad since it seems like Cutlass does not support TileShape.M < 64 for this kernel. It's possible to boost perf a bit by using TileShape `<64,64,128>`. ## Original using SM89 I tried using cutlass FP8 row-wise scaled-mm for sm89 on sm120 (5090) and it works. I guess it makes sense because sm120 matmul uses the standard sm80 PTX instructions (`cp.async`+`mma` and friends). Simple benchmark script ```python import torch from torch._inductor.utils import do_bench_using_profiling N, K = 4096, 4096 for M in [16, 64, 256, 1024, 4096]: A = torch.randn(M, K, device="cuda").to(torch.float8_e4m3fn) B = torch.randn(N, K, device="cuda").to(torch.float8_e4m3fn).T scale_A = torch.ones(M, 1).cuda() scale_B = torch.ones(1, N).cuda() out = torch._scaled_mm(A, B, scale_A, scale_B, out_dtype=torch.bfloat16) out_ref = ((A.float() @ B.float()) * scale_A * scale_B).bfloat16() torch.testing.assert_close(out, out_ref) latency_us = do_bench_using_profiling(lambda: torch._scaled_mm(A, B, scale_A, scale_B, out_dtype=torch.bfloat16)) tflops = (2 * M * N * K) / latency_us / 1e9 print(f"{M=}\t{N=}\t{K=}\t{tflops:.2f} TFLOPS") ``` M \| N \| K \| TFLOPS ---\|---\|---\|--- 16 \| 4096 \| 4096 \| 25.73 TFLOPS 64 \| 4096 \| 4096 \| 71.84 TFLOPS 256 \| 4096 \| 4096 \| 86.40 TFLOPS 1024 \| 4096 \| 4096 \| 112.12 TFLOPS 4096 \| 4096 \| 4096 \| 121.24 TFLOPS Accodring to [RTX Blackwell Whitepaper](https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf), FP8 MMA with FP32 accumulate is 419 TFLOPS. So the result is quite bad here... However, if I change `ThreadblockSwizzle` to `cutlass::gemm::threadblock::GemmIdentityThreadblockSwizzle<1>` M \| N \| K \| TFLOPS ---\|---\|---\|-------- 16\|4096\|4096\|27.13 TFLOPS 64\|4096\|4096\|84.84 TFLOPS 256\|4096\|4096\|96.75 TFLOPS 1024\|4096\|4096\|110.21 TFLOPS 4096\|4096\|4096\|122.98 TFLOPS Small M slightly improves, but large M is still bad. If I further change `ThreadBlockShape=<128, 64, 128>, WarpShape=<64, 32, 128>, NumStages=3` for M>256, which is taken from [cutlass example 58](https://github.com/NVIDIA/cutlass/blob/v3.9.2/examples/58_ada_fp8_gemm/ada_fp8_gemm.cu), I get the following results M \| N \| K \| TFLOPS ---\|---\|---\|-------- 1024\|4096\|4096\|313.28 4096\|4096\|4096\|376.73 Which is much closer to hardware limit. And it also means this kernel is sufficient to get the most perf out of sm120. Only need better tuned configs. To make sure this high perf is only obtainable with `GemmIdentityThreadblockSwizzle<1>` + `ThreadBlockShape=<128, 64, 128>, WarpShape=<64, 32, 128>, NumStages=3`, I also try using `ThreadblockSwizzleStreamK` + `ThreadBlockShape=<128, 64, 128>, WarpShape=<64, 32, 128>, NumStages=3` M \| N \| K \| TFLOPS ---\|---\|---\|-------- 1024\|4096\|4096\|144.03 4096\|4096\|4096\|156.86 A bit better than current configs, but still very far away from hardware limit. @alexsamardzic I noticed you chose this configs in #149978. Do you have any numbers how the current configs perform on sm89? Update: Using triton codegen-ed from inductor `compiled_scaled_mm = torch.compile(torch._scaled_mm, dynamic=False, mode="max-autotune-no-cudagraphs")` M \| N \| K \| TFLOPS ---\|---\|---\|-------- 16\|4096\|4096\|25.60 64\|4096\|4096\|71.74 256\|4096\|4096\|161.64 1024\|4096\|4096\|185.89 4096\|4096\|4096\|215.53 Better than default configs, but still far away from the config above for compute-bound Pull Request resolved: https://github.com/pytorch/pytorch/pull/155991 Approved by: https://github.com/drisspg, https://github.com/eqy	2025-06-17 18:52:44 +00:00
Scott Wolchok	e323d46b61	ELU: compute ELU(0) with the cheaper definition (#155765 ) Both halves of the ELU definition yield 0 when evaluated at 0. Let's choose the half that doesn't require expm1. (I have no particular evidence that the input is often 0 in any case, but this seems like a free win.) Differential Revision: [D76481038](https://our.internmc.facebook.com/intern/diff/D76481038/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155765 Approved by: https://github.com/ezyang	2025-06-17 18:20:22 +00:00
Animesh Jain	8b0e0e4f23	[dynamo] Support tracing of functools.lru_cached method (#156125 ) Fixes https://github.com/pytorch/pytorch/issues/155841 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156125 Approved by: https://github.com/williamwen42	2025-06-17 18:11:32 +00:00
Svetlana Karslioglu	fc5ae12293	Fix issue with right-nav (#156119 ) Enable on page right nav. For autosummary, we need to set `"show_toc_level": 2` so that navigation is enabled. Example: * Main: https://docs.pytorch.org/docs/main/special.html - right nav (under On this page) is empty. * Preview: https://docs-preview.pytorch.org/pytorch/pytorch/156119/special.html - right nav (under On this page) has a all the object listed <img width="1125" alt="Screenshot 2025-06-16 at 2 48 16 PM" src="https://github.com/user-attachments/assets/0790bb72-5997-4542-9847-0a89be4598c0" /> vs <img width="1030" alt="Screenshot 2025-06-16 at 2 48 55 PM" src="https://github.com/user-attachments/assets/4897c49c-044d-4bea-a8cd-490c90cca2b0" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156119 Approved by: https://github.com/albanD	2025-06-17 18:09:51 +00:00
Catherine Lee	32c1611263	[CI][run_test] Fix rerun logic for failing at exit (#155853 ) Sometimes a test file reports success according to pytest, but fails afterwards, and the rerun logic doesn't handle that correctly. The name of the last run test is saved in order to do more efficient reruns (target the last run test for a rerun without rerunning the entire file). This usually correct, ex test fails and pytest catches it -> lastrun = the test that failed, test segfaults (pytest doesn't catch) -> lastrun is the test that segfaulted. But sometimes pytest reports a success, but the process has non zero exit code. The two cases I know of are hangs and double freeing at exit. In this case, its unclear which test caused the failure, so lastrun is set to be the first test that ran in that session, so that during the next session it will start from the beginning in an attempt to replicate the error (an alternate solution would be to just fail and not rerun, which might be the better option). But then it reruns with runsingle, which prevents lastrun from being reset (not sure why, I'm pretty sure there's no difference between resetting and not normally), so lastrun becomes the last test that ran, and its not always true that lastrun is the one that caused it. Then on the next run, it starts from the last test and the process now exits cleanly Short term solution here: ensure the lastrun is always set to the initial value if the session succeeds. This is correct even in the normal path because initial value shouldn't change in that case Things that still need to be fixed: * log says "running single test" which is not true * no xml reports get generated here * also no xml reports get generated on segfault * docs for this I think I have a PR that fixes the above but its old so I need to take another look Testing: This from when I was based on a commit that had a hang for macs, and before I added the skips in inductor array ref: `cc862d2c14` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155853 Approved by: https://github.com/malfet	2025-06-17 17:51:40 +00:00
Nikita Shulga	6629eaf0c6	[CMAKE] Fix torch_cpu relink logic if metal shaders are recompiled (#156193 ) Beforehand, shader recompilation updated `caffe2/aten/src/ATen/metallib_dummy.cpp` but `torch_cpu` were dependent on `aten/src/ATen/metallib_dummy.cpp` Test plan: Run `python3 ../tools/build_with_debinfo.py ../aten/src/ATen/native/mps/kernels/UpSample.metal` and observe that torch_cpu is being relinked Pull Request resolved: https://github.com/pytorch/pytorch/pull/156193 Approved by: https://github.com/manuelcandales	2025-06-17 17:49:33 +00:00
Manuel Candales	a4ea242edc	[MPS] Implement scan metal kernels (#156100 ) Implements metal kernels for scan operations: - Migrates cumsum and cumprod from MPSGraph implementation to Metal. - Fixes #154881 - Adds MPS backend support for cummin and cummax Pull Request resolved: https://github.com/pytorch/pytorch/pull/156100 Approved by: https://github.com/malfet	2025-06-17 17:44:22 +00:00
Jane Xu	9a5c59368d	Replace all RAIIATH with Tensor in libtorch_agnostic test, test some APIs (#155977 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155977 Approved by: https://github.com/albanD ghstack dependencies: #155367	2025-06-17 17:36:31 +00:00
Jane Xu	b115a4c03a	torch::stable::Tensor beginnings, mainly mem mgmt (#155367 ) ``` // The torch::stable::Tensor class is a highlevel C++ header-only wrapper around // the C shim Tensor APIs. We've modeled this class after TensorBase, as custom // op kernels only really need to interact with Tensor metadata (think sizes, // strides, device, dtype). Other functions on Tensor (like empty_like) should // live like the ATen op that they are and exist outside of this struct. // // There are several goals of this class over AtenTensorHandle and // RAIIAtenTensorHandle: // 1. torch::stable::Tensor is a nicer UX much closer to torch::Tensor than the // C APIs with AtenTensorHandle. Under the hood we still call to these C shim // APIs to preserve stability. // 2. RAIIAtenTensorHandle boils down to a uniq_ptr that forces the user to pass // around ownership. This makes it difficult to pass one input into 2 // different functions, e.g., doing something like c = a(t) + b(t) for // stable::Tensor t. Thus, we use a shared_ptr here. ``` This PR: - exemplifies the above - adds test cases in libtorch_agnostic to make sure the file actually works - includes the results of a battle with template specialization Pull Request resolved: https://github.com/pytorch/pytorch/pull/155367 Approved by: https://github.com/albanD	2025-06-17 17:36:31 +00:00
Soumith Chintala	2625c70aec	Update CODEOWNERS (#156182 ) as title says. removing me as codeowner for cpp extensions Pull Request resolved: https://github.com/pytorch/pytorch/pull/156182 Approved by: https://github.com/albanD	2025-06-17 17:15:41 +00:00
rzou	a24afbff3f	Support torch.cuda.*Tensor in Dynamo (#156107 ) Summary: This PR adds support for torch.cuda.FloatTensor and friends in Dynamo. These are indeed legacy APIs, but that doesn't stop us from adding support for them in torch.compile. I add support for these in the same way that we support torch.Tensor: these APIs can be safely put into the Dynamo graph. Fixes #130722 Test Plan: - new test Pull Request resolved: https://github.com/pytorch/pytorch/pull/156107 Approved by: https://github.com/williamwen42	2025-06-17 16:31:10 +00:00
Ruben Rodriguez Buchillon	9a2c669425	[inductor][cutlass] binary remote cache (#156106 ) Summary: # Why speed up cutlass kernel generation and retrieval # What using the _ManifoldCache, make a KernelBinaryCache that uploads/downloads kernels and their error files. only register the handler internally Test Plan: ## prove that we can upload successfully ``` buck2 run mode/opt scripts/coconutruben/torchmm:experiment 2>&1 ``` ``` manifold ls coconutruben-test-01/tree/cutlass_concept_2 673184 cfkykew2fw5572hjr4e7jbog7oix7xjkegtn2ovikyhxe6pr4tcw.so 649776 cpjqda67c6ojj75z3ddnmfbxinpm7yp7rc2q2oxwsrtwsnacklqv.so ``` ## prove that we can download successfully ``` buck2 run mode/opt scripts/coconutruben/torchmm:experiment 2>&1 ``` ``` I0611 12:48:38.759000 935012 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:65] Successfully downloaded /var/tmp/torchinductor_coconutruben/fk/cfkykew2fw5572hjr4e7jbog7oix7xjkegtn2ovikyhxe6pr4tcw.so I0611 12:48:38.760000 935012 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:65] Successfully downloaded /var/tmp/torchinductor_coconutruben/pj/cpjqda67c6ojj75z3ddnmfbxinpm7yp7rc2q2oxwsrtwsnacklqv.so ``` ## prove that we can upload errors successfully ``` buck2 run mode/opt scripts/coconutruben/torchmm:experiment 2>&1 ``` ``` manifold ls coconutruben-test-01/tree/cutlass_concept_2 4846 cqiq4vjbvytdofutoxisa3pqjplgpgmt2sh7dtatiw4bqt5rtjgc.so.error 4846 cqymdwsfsirhkqglv7sbjyvqkrt3ryql4mtb45tekt76347ee6sx.so.error ``` ## prove that we can download errors successfully ``` buck2 run mode/opt scripts/coconutruben/torchmm:experiment 2>&1 ``` ``` I0611 12:56:14.078000 1001022 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:74] Successfully downloaded /var/tmp/torchinductor_coconutruben/qi/cqiq4vjbvytdofutoxisa3pqjplgpgmt2sh7dtatiw4bqt5rtjgc.so.error I0611 12:56:14.079000 1001022 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:74] Successfully downloaded /var/tmp/torchinductor_coconutruben/qy/cqymdwsfsirhkqglv7sbjyvqkrt3ryql4mtb45tekt76347ee6sx.so.error ``` ## showing timing information ``` I0616 11:22:29.169000 2249769 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:71] Successfully downloaded /var/tmp/torchinductor_coconutruben/fk/cfkykew2fw5572hjr4e7jbog7oix7xjkegtn2ovikyhxe6pr4tcw.so (download: 0.842s, write: 0.000s, total: 0.842s) I0616 11:22:29.169000 2249769 /data/users/coconutruben/fbsource/fbcode/caffe2/torch/_inductor/fb/kernel_binary_remote_cache.py:71] Successfully downloaded /var/tmp/torchinductor_coconutruben/pj/cpjqda67c6ojj75z3ddnmfbxinpm7yp7rc2q2oxwsrtwsnacklqv.so (download: 0.838s, write: 0.001s, total: 0.838s) ``` Rollback Plan: Reviewed By: henrylhtsang Differential Revision: D76454741 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156106 Approved by: https://github.com/henrylhtsang Co-authored-by: atalman <atalman@fb.com>	2025-06-17 16:24:10 +00:00
David Berard	d66b4bcc3f	[inductor][triton pin] Support triton builtins after #7054 (#156031 ) Triton's PR 7054 modifies the builtins to take _semantic as a kwarg instead of _builder. To handle this, this PR checks the signature of tl.core.view (to see if it takes _builder or _semantic), and adds a wrapper converting _semantic to _builder if the new _semantic kwarg is being used. (Previously-)failing test: `python test/inductor/test_cooperative_reductions.py -k test_welford_non_power_of_2_rsplit_persistent_True_x_9_r_8000_rsplit_37` Differential Revision: [D76801240](https://our.internmc.facebook.com/intern/diff/D76801240) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156031 Approved by: https://github.com/NikhilAPatel	2025-06-17 16:09:55 +00:00
Yuki Kobayashi	d083841c0e	Fix a small sphinx markup error (#156061 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156061 Approved by: https://github.com/colesbury	2025-06-17 15:36:02 +00:00
Catherine Lee	0079c80b35	[CI] Do not constrain memory for ROCm testing in CI (#156115 ) Fixes ROCm OOMs introduced by https://github.com/pytorch/pytorch/pull/155631 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156115 Approved by: https://github.com/jeffdaily	2025-06-17 15:30:36 +00:00
windsonsea	7fcad0231c	[Docs] Convert to markdown to fix 155025 (#155789 ) Related to #155025 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155789 Approved by: https://github.com/svekars	2025-06-17 15:08:14 +00:00
Nikita Shulga	4886ba64dc	[BE] Refactor functions from optional_submodules (#155954 ) And use `pathlib.Path` instead of `os.path` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155954 Approved by: https://github.com/Skylion007 ghstack dependencies: #155947	2025-06-17 14:41:52 +00:00
Frank Lin	cf90c9f8d1	[Draft][CUDA] Use runtime driver API for cuStreamWriteValue32 (#156097 ) Fixes #154073 Reference: https://github.com/NVIDIA/Fuser/pull/4197 See PR #154097 @nWEIdia is currently out of the office, so I’ve temporarily taken over his work. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156097 Approved by: https://github.com/ngimel, https://github.com/cyyever Co-authored-by: Wei Wang <weiwan@nvidia.com>	2025-06-17 14:15:49 +00:00
Xuehai Pan	42015db6a9	[BE] fix typos in benchmarks/ (#156077 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156077 Approved by: https://github.com/Skylion007, https://github.com/malfet ghstack dependencies: #156069	2025-06-17 13:12:18 +00:00
Luca Wehrstedt	0a0023d984	Enable NCCL zero-copy (user buffer registration) for FSDP2 (#150564 ) In recent versions NCCL introduced support for "user buffer registration", i.e., allowing user-owned memory (such as regular PyTorch tensors) to be "registered" (pinned, page-locked, etc.) with all the various hardware (NVLink, InfiniBand, ...) in order to support zero-copy transfers and thus accelerate communication and reduce resource footprint of NCCL's kernels (which reduces contention). This was already exposed in PyTorch through a custom allocator provided by the NCCL process group. DDP already uses this, via a memory pool to allow caching and reusing. FSDP2 is also particularly suited to leverage user buffer registration because the buffers it passes to NCCL are allocated by FSDP2 itself, since it anyways needs to (de)interleave the parameters to/from these private buffers. This PR adds an extra flag to FSDP2 that tells it to use the ProcessGroup allocator for these private buffers, thus allowing it to leverage NCCL zero-copy (when supported). Pull Request resolved: https://github.com/pytorch/pytorch/pull/150564 Approved by: https://github.com/kwen2501, https://github.com/weifengpy, https://github.com/syed-ahmed	2025-06-17 12:54:58 +00:00
Wang, Chuanqi	11bb1ece50	[CI] Fix triton version split issue (#155670 ) Fix a bug caused by #155313, refer https://github.com/pytorch/pytorch/actions/runs/15576592378/job/43862613039?pr=154194#step:7:652 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155670 Approved by: https://github.com/atalman, https://github.com/EikanWang	2025-06-17 12:42:40 +00:00
Xuehai Pan	1cce73b5f4	[build] Change `--cmake{,-only}` arguments to envvars to support modern Python build frontend (#156045 ) See also: - #156029 - #156027 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156045 Approved by: https://github.com/ezyang ghstack dependencies: #156040, #156041	2025-06-17 11:40:24 +00:00
Xuehai Pan	57084ca846	[BE][setup] allow passing pytorch-specific `setup.py` options from envvars (#156041 ) See also: - #156029 - #156027 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156041 Approved by: https://github.com/ezyang ghstack dependencies: #156040	2025-06-17 11:40:24 +00:00
LuFengqing	092aed1b18	[Intel GPU] Enable GQA and different head_dim of value for SDPA (#150992 ) In OneDNN v3.7, SDPA doesn't support num_head_q != num_head_kv (aka GQA) and head_dim_qk != head_dim_v. In OneDNN v3.8, SDPA supports these two scenarios. Enable them in this PR. SDPA UTs pass in local test. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150992 Approved by: https://github.com/guangyey, https://github.com/drisspg, https://github.com/EikanWang Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-06-17 11:09:51 +00:00
Wei (Will) Feng	4a8f5e752b	[FSDP2] explain user contract for fully_shard (#156070 ) <img width="896" alt="Screenshot 2025-06-16 at 1 36 00 AM" src="https://github.com/user-attachments/assets/7cdea256-2454-49c7-8b32-24549a13134d" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/156070 Approved by: https://github.com/mori360	2025-06-17 10:03:19 +00:00
Xuehai Pan	8d7ee0f4fb	[BE] fix typos in .ci/, .circleci/, .github/ (#156069 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156069 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-06-17 09:43:14 +00:00
Xuehai Pan	2e0e08588e	[BE][PYFMT] migrate PYFMT for `torch/[e-n]*/` to `ruff format` (#144553 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/144553 Approved by: https://github.com/ezyang ghstack dependencies: #144551	2025-06-17 08:18:47 +00:00
cyy	95cb42c45d	Use CMAKE_COLOR_DIAGNOSTICS (#154583 ) `CMAKE_COLOR_DIAGNOSTICS` was introduced in CMake 2.24. Use it to simplify CMake code. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154583 Approved by: https://github.com/ezyang	2025-06-17 04:52:31 +00:00
cyy	d43c0bdf46	[CI] Move ASAN jobs to clang-18 (#149099 ) Use clang-18 for ASAN jobs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/149099 Approved by: https://github.com/ezyang	2025-06-17 04:51:07 +00:00
Animesh Jain	7b0118884e	[invoke_subgraph][inductor] Dont fallback on complex dtype (#155885 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155885 Approved by: https://github.com/jansel, https://github.com/zou3519 ghstack dependencies: #155828	2025-06-17 04:47:12 +00:00
Animesh Jain	ffcc6fea7b	[invoke_subgraph] Ignore redundantly nested invoke_subgraph (#155828 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155828 Approved by: https://github.com/zou3519	2025-06-17 04:47:12 +00:00
Nikita Shulga	b1713c6655	[MPS][Testing][BE] Fix samples for full_like (#156026 ) Now that device is known, one can avoid creating tensors of `torch.double` type Pull Request resolved: https://github.com/pytorch/pytorch/pull/156026 Approved by: https://github.com/dcci ghstack dependencies: #156121	2025-06-17 04:46:26 +00:00
Ke Wen	82672206b7	[SymmMem] Make get_rank_to_global_rank return const ref (#156117 ) Avoiding a copy, not expecting a caller to change its value. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156117 Approved by: https://github.com/fegin ghstack dependencies: #155506, #155835, #155968, #155971, #155975, #156116	2025-06-17 04:13:18 +00:00
Ke Wen	eea3bcb3d1	[SymmMem] Cache rank_to_global_rank exchange (#156116 ) The rank-to-global-rank exchange is a major overhead in `NVSHMEMSymmetricMemory` creation. We should cache its result on per-group basis. Before this change: ``` TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py exchanged_n_times: 18 ``` After this change: ``` TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py exchanged_n_times: 1 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156116 Approved by: https://github.com/fegin, https://github.com/ngimel ghstack dependencies: #155506, #155835, #155968, #155971, #155975	2025-06-17 04:12:37 +00:00
Oguz Ulgen	a2a75be0f8	Rename inductor cache (#156128 ) Requested by Simon on a different PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/156128 Approved by: https://github.com/xmfan	2025-06-17 03:57:18 +00:00
henrylhtsang	45382b284d	[cutlass backend] changes how gpu_kernels_o are handled for cutlass (#155875 ) Currently, we do it a bit hacky: Look at all the .o we have from this session, add them all to AOTI. This for example doesn't work if we do multiple AOTI compilation in one session, without clearing the inductor cache. Also I want to change how cutlass .so are compiled. Hence this change. This change is broken down since @coconutruben is trying to make a change to the same files too. Differential Revision: [D76563003](https://our.internmc.facebook.com/intern/diff/D76563003/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155875 Approved by: https://github.com/ColinPeppler	2025-06-17 02:06:54 +00:00
cyy	64bb6317a5	[Accelerator] Fix Python typing in accelerator (#152394 ) There are some changes: 1. Use keywords for arguments if possible. 2. `__exit__ ` of `device_index` is changed to return None. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152394 Approved by: https://github.com/XuehaiPan, https://github.com/guangyey, https://github.com/ezyang Co-authored-by: Xuehai Pan <XuehaiPan@outlook.com> Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-06-17 01:27:40 +00:00
William Wen	1f0eb79e3e	[dynamo] fix KeyError in LOAD_FAST_CHECK (#155763 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155763 Approved by: https://github.com/StrongerXi, https://github.com/jansel ghstack dependencies: #155761	2025-06-17 00:54:16 +00:00
William Wen	4e833c2005	[dynamo] support tracing weakref callback (#155761 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155761 Approved by: https://github.com/StrongerXi, https://github.com/jansel	2025-06-17 00:54:16 +00:00
xadupre	e6252f62ef	[ONNX] Implements converter for higher order ops scan (#154513 ) Fixes #151327 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154513 Approved by: https://github.com/justinchuby Co-authored-by: Justin Chu <justinchuby@users.noreply.github.com>	2025-06-17 00:54:07 +00:00
Pian Pawakapan	b618817479	[PGO] include ints/floats in suggested whitelist (#155980 ) Made the mistake of dropping these Pull Request resolved: https://github.com/pytorch/pytorch/pull/155980 Approved by: https://github.com/bobrenjc93	2025-06-17 00:41:38 +00:00
Benjamin Glass	4311aea5e7	[AOTInductor] Add class declarations to torch._C._aoti interface file (#155128 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155128 Approved by: https://github.com/desertfire ghstack dependencies: #155149	2025-06-17 00:10:57 +00:00
yifanmao	82fb904140	Add warning for incorrected grad results at world size 1 (#154928 ) Add warning for the issue discussed at https://github.com/pytorch/pytorch/issues/144045 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154928 Approved by: https://github.com/weifengpy	2025-06-17 00:08:04 +00:00
mori360	eb4cf59ecd	Add FSDP2 logging (#155826 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/155826 Approved by: https://github.com/weifengpy	2025-06-16 23:49:58 +00:00
Nikita Shulga	6e2992a998	Remove unused Azure pipeline trigger script (#156134 ) ## Summary - delete `.circleci/scripts/trigger_azure_pipeline.py` ## Testing - `python3 -m pip install flake8` - `python3 -m flake8 .circleci/scripts` ------ https://chatgpt.com/codex/tasks/task_e_6850a55f530c83279036800308fb6871 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156134 Approved by: https://github.com/izaitsevfb	2025-06-16 23:42:52 +00:00
codingwithsurya	4781b0ee60	[SymmMem] Add NVSHMEM GET support to Triton (#155890 ) Adds NVSHMEM GET operation support for Triton kernels: - Add `getmem_block` core.extern wrapper for nvshmemx_getmem_block - Add basic `test_triton_get` for 2-rank GET operation - Add `test_triton_get_ring` for ring topology GET across arbitrary ranks Tests: `$ TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py` `TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_get` ```python @skipIfRocm @requires_triton() def test_triton_get(self) -> None: @triton.jit def get_kernel(dst_ptr, src_ptr, numel: tl.constexpr, peer: tl.constexpr): nvshmem.getmem_block(dst_ptr, src_ptr, numel, peer) # ... setup code ... val = 7 inp = symm_mem.empty(numel, dtype=dtype, device=self.device).fill_( val if rank == 0 else -1 ) out = symm_mem.empty(numel, dtype=dtype, device=self.device).fill_(-1) peer = 1 - rank if rank == 1: # Rank 1 gets data from rank 0 get_kernel[(1, 1, 1)](dst_ptr, src_ptr, numel=numel, peer=peer, extern_libs=nvshmem_lib) dist.barrier() print(f"[Rank {rank}] inp buffer: {inp}") print(f"[Rank {rank}] out buffer: {out}") print(f"[Rank {rank}] got data from peer {peer}") ``` ``` [Rank 0] inp buffer: tensor([7, 7, 7, 7, 7, 7, 7, 7], device='cuda:0', dtype=torch.int8) [Rank 1] inp buffer: tensor([-1, -1, -1, -1, -1, -1, -1, -1], device='cuda:1', dtype=torch.int8) ... [Rank 1] out buffer: tensor([7, 7, 7, 7, 7, 7, 7, 7], device='cuda:1', dtype=torch.int8) ... [Rank 1] got data from peer 0 ---------------------------------------------------------------------- Ran 2 tests in 17.046s OK ``` ```python @skipIfRocm @requires_triton() def test_triton_get_ring(self) -> None: @triton.jit def get_kernel(dst_ptr, src_ptr, numel: tl.constexpr, peer: tl.constexpr): nvshmem.getmem_block(dst_ptr, src_ptr, numel, peer) # ... setup code ... # Ring topology: each rank gets data from the rank to its left peer = (rank - 1) % world_size # All ranks execute the get operation get_kernel[(1, 1, 1)](dst_ptr, src_ptr, numel=numel, peer=peer, extern_libs=nvshmem_lib) dist.barrier() print(f"[Rank {rank}] inp buffer: {inp}") print(f"[Rank {rank}] out buffer: {out}") print(f"[Rank {rank}] got data from peer {peer}") ``` ``` Output (8 GPUs): [Rank 0] inp buffer: tensor([0, 0, 0, 0, 0, 0, 0, 0], device='cuda:0', dtype=torch.int8) [Rank 2] inp buffer: tensor([2, 2, 2, 2, 2, 2, 2, 2], device='cuda:2', dtype=torch.int8) [Rank 5] inp buffer: tensor([5, 5, 5, 5, 5, 5, 5, 5], device='cuda:5', dtype=torch.int8) [Rank 6] inp buffer: tensor([6, 6, 6, 6, 6, 6, 6, 6], device='cuda:6', dtype=torch.int8) [Rank 3] inp buffer: tensor([3, 3, 3, 3, 3, 3, 3, 3], device='cuda:3', dtype=torch.int8) [Rank 1] inp buffer: tensor([1, 1, 1, 1, 1, 1, 1, 1], device='cuda:1', dtype=torch.int8) [Rank 2] out buffer: tensor([1, 1, 1, 1, 1, 1, 1, 1], device='cuda:2', dtype=torch.int8) [Rank 5] out buffer: tensor([4, 4, 4, 4, 4, 4, 4, 4], device='cuda:5', dtype=torch.int8) [Rank 0] out buffer: tensor([7, 7, 7, 7, 7, 7, 7, 7], device='cuda:0', dtype=torch.int8) [Rank 2] got data from peer 1 [Rank 5] got data from peer 4 [Rank 0] got data from peer 7 [Rank 7] inp buffer: tensor([7, 7, 7, 7, 7, 7, 7, 7], device='cuda:7', dtype=torch.int8) [Rank 6] out buffer: tensor([5, 5, 5, 5, 5, 5, 5, 5], device='cuda:6', dtype=torch.int8) [Rank 3] out buffer: tensor([2, 2, 2, 2, 2, 2, 2, 2], device='cuda:3', dtype=torch.int8) [Rank 6] got data from peer 5 [Rank 3] got data from peer 2 [Rank 1] out buffer: tensor([0, 0, 0, 0, 0, 0, 0, 0], device='cuda:1', dtype=torch.int8) [Rank 1] got data from peer 0 [Rank 4] inp buffer: tensor([4, 4, 4, 4, 4, 4, 4, 4], device='cuda:4', dtype=torch.int8) [Rank 7] out buffer: tensor([6, 6, 6, 6, 6, 6, 6, 6], device='cuda:7', dtype=torch.int8) [Rank 7] got data from peer 6 [Rank 4] out buffer: tensor([3, 3, 3, 3, 3, 3, 3, 3], device='cuda:4', dtype=torch.int8) [Rank 4] got data from peer 3 ---------------------------------------------------------------------- Ran 1 test in 41.045s OK ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155890 Approved by: https://github.com/kwen2501, https://github.com/mandroid6	2025-06-16 23:18:15 +00:00
Nikita Shulga	bb1f3d1a55	[MPSInductor] Improve `_default` dtype inference (#156121 ) By just adding 'mps' as one of the backend options and fixing reduction op to actually return tuple of CSEVariable's rather than tuple of strings Test plan: CI Pull Request resolved: https://github.com/pytorch/pytorch/pull/156121 Approved by: https://github.com/dcci	2025-06-16 23:11:53 +00:00
Nicolas Macchioni	508cdc4fc9	[BE][4/X] Phase out usage of `use_max_autotune()` (#155850 ) See #155847 for context Pull Request resolved: https://github.com/pytorch/pytorch/pull/155850 Approved by: https://github.com/masnesral	2025-06-16 23:10:26 +00:00
Yiming Zhou	f2d70898c6	[nativert] Move OpKernel to PyTorch core (#156011 ) Summary: Moves OpKernel base class to PyTorch core. It is an abstract interface representing a kernel, which is responsible for executing a single Node in the graph. Torch Native Runtime RFC: pytorch/rfcs#72 Test Plan: buck2 run mode/dev-nosan caffe2/test/cpp/nativert:op_kernel_test Rollback Plan: Differential Revision: D76525939 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156011 Approved by: https://github.com/zhxchen17	2025-06-16 22:53:10 +00:00
PyTorch MergeBot	35ecd7c2d4	Revert "[Cutlass] Fix buffer missing issues (#155897 )" This reverts commit 9bd42c15707a4b410ee005d5916e882a7db432bb. Reverted https://github.com/pytorch/pytorch/pull/155897 on behalf of https://github.com/atalman due to failing internal tests ([comment](https://github.com/pytorch/pytorch/pull/155897#issuecomment-2978391416))	2025-06-16 22:44:39 +00:00
PyTorch MergeBot	190f76fa31	Revert "Implement guard collectives (#155558 )" This reverts commit 5a5a05a6a3be376130848e235df73b752eef0230. Reverted https://github.com/pytorch/pytorch/pull/155558 on behalf of https://github.com/malfet due to Hmm, may be I'm looking at the wrong metric, but `c92f1075aa/1` shows that test started to pass after PR were reverted ([comment](https://github.com/pytorch/pytorch/pull/155558#issuecomment-2978337152))	2025-06-16 22:26:52 +00:00
Ting Lu	c92f1075aa	Fix if condition for CUDA 12.9 Win build (#156108 ) follow-up for https://github.com/pytorch/pytorch/pull/155799/files Currently the last if condition will be executed for CUDA 12.9, overriding previous CUDA_ARCH_LIST. We should exclude 12.9 from the last if condition to fix this. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156108 Approved by: https://github.com/atalman	2025-06-16 21:57:34 +00:00
Scott Wolchok	cce4d213a6	Remove non-header-only dep from c10_headers target (#155858 ) It depends on /c10/util:base which is not header-only. Differential Revision: [D76552750](https://our.internmc.facebook.com/intern/diff/D76552750/) NOTE FOR REVIEWERS: This PR has internal Meta-specific changes or comments, please review them on [Phabricator](https://our.internmc.facebook.com/intern/diff/D76552750/)! Pull Request resolved: https://github.com/pytorch/pytorch/pull/155858 Approved by: https://github.com/ezyang	2025-06-16 21:41:25 +00:00
bobrenjc93	a24ce67dee	[ez] fix grammar error in comment (#156053 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156053 Approved by: https://github.com/jingsh ghstack dependencies: #155982, #155996	2025-06-16 20:53:07 +00:00
Colin Peppler	247113e03e	Add size_hint_or_throw (#155615 ) ## Summary `TypeError("Cannot convert symbols to int")` is coming up more recently since more unbacked symints are making its way into Inductor. See https://github.com/pytorch/pytorch/issues/155484 - One way to deal with this is to add `size_hint_or_throw` to throw if we try to pull a hint from an unbacked expr. - Then, repurpose `size_hint` to accommodate unbacked symints by setting a default fallback or adding an appropriate fallback for each callsite. This PR adds `size_hint_or_throw` which will throw if unbacked symints exist - use `size_hint_or_throw` -- usually when the callee can try/catch the exception or guards against unbacked symints ------ with Codex https://chatgpt.com/codex/tasks/task_e_684869dfc740832882c88d05534cc8f9 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155615 Approved by: https://github.com/ezyang, https://github.com/laithsakka, https://github.com/jingsh Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-06-16 20:46:51 +00:00
Justin Silver	008345be9d	Fix #155018 (convert distributed rst to markdown) (#155528 ) Used [rst2myst tool](https://rst-to-myst.readthedocs.io/en/latest/) Fixes #155018 Docs comparison (check out the 'new' whenever docs build) 1. distributed.checkpoint ([old](https://docs.pytorch.org/docs/main/distributed.checkpoint.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155528/distributed.checkpoint.html)) 2. distributed.elastic ([old](https://docs.pytorch.org/docs/main/distributed.elastic.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155528/distributed.elastic.html)) 3. distributed.fsdp.fully_shard ([old](https://docs.pytorch.org/docs/main/distributed.fsdp.fully_shard.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155528/distributed.fsdp.fully_shard.html)) 4. distributed.optim ([old](https://docs.pytorch.org/docs/main/distributed.optim.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155528/distributed.optim.html)) 5. distributed.pipelining ([old](https://docs.pytorch.org/docs/main/distributed.pipelining.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155528/distributed.pipelining.html)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155528 Approved by: https://github.com/wz337, https://github.com/svekars	2025-06-16 20:46:09 +00:00
Xuan Zhang	eb2af14f8e	[PT2][partitioners] Add aten.split to view_ops list [relanding #155424 ] (#155943 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155943 Approved by: https://github.com/ShatianWang	2025-06-16 20:42:54 +00:00
PyTorch MergeBot	03488d820c	Revert "[MPS][Testing][BE] Fix samples for full_like (#156026 )" This reverts commit 2d832c9587fd99db295b62d0c9b459d509c19d06. Reverted https://github.com/pytorch/pytorch/pull/156026 on behalf of https://github.com/atalman due to Sorry breaks MPS tests: test_ops.py::TestMathBitsCPU::test_neg_view_full_like_cpu_float64 [GH job link](https://github.com/pytorch/pytorch/actions/runs/15683608879/job/44182730620) [HUD commit link](`2d832c9587`) ([comment](https://github.com/pytorch/pytorch/pull/156026#issuecomment-2977903074))	2025-06-16 19:50:26 +00:00
Pian Pawakapan	6d2155db49	[PGO] no code state update on dynamic=False (#155961 ) Summary: When tensor size changes are detected on `dynamic=False`, overwrites the PGO state with the newest static shapes to reflect the latest frame state, instead of updating automatic dynamic. A longer term solution, if we move to shared PGO state between multiple jobs, would be to update automatic dynamic, but avoid suggesting/logging the whitelist (compiling with `dynamic=False` should already override any dynamic PGO that's read, so we're fine there). This way if any particular job runs with `dynamic=False`, it won't statically overwrite the entire PGO state if it's shared with many other jobs. Test Plan: test/dynamo/test_pgo.py Rollback Plan: Differential Revisi,on: D76630499 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155961 Approved by: https://github.com/bobrenjc93	2025-06-16 19:47:55 +00:00
Edward Z. Yang	5a5a05a6a3	Implement guard collectives (#155558 ) When running a distributed job with compiler collectives enabled, if one rank recompiles while others do not, this leads to a deadlock (as not everyone will rendezvous with the compiler collective from the recompile). Although there aren't any convenient ways to cheaply solve this problem, if you are willing to force everyone to sync when evaluating guards, you can just force everyone to recompile if anyone requires a recompile. So the way guard collectives work is: 1. Perform compiled code lookup (evaluating guards) 2. Run a collective, communicating if you found a compiled code or not 3. If anyone requires recompile, force everyone to recompile One current deficiency in the implementation is we can't conveniently track the time it takes to run this collective. I need to test if we actually successfully are running the collective on a separate stream, or if we have to wait for user collectives to all finish. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/155558 Approved by: https://github.com/Microve	2025-06-16 19:46:16 +00:00
PyTorch MergeBot	61b271e0f3	Revert "Implement guard collectives (#155558 )" This reverts commit 38e5e81e55fc5d85d6cf8a83c96c88578995e3fe. Reverted https://github.com/pytorch/pytorch/pull/155558 on behalf of https://github.com/atalman due to Breaks CI, sorry: [GH job link](https://github.com/pytorch/pytorch/actions/runs/15683161593/job/44181274826) [HUD commit link](`38e5e81e55`) ([comment](https://github.com/pytorch/pytorch/pull/155558#issuecomment-2977871178))	2025-06-16 19:40:46 +00:00
Georgia Phillips	7cf38d2a05	Make benchmark by op for TS model work with sample inputs (#155988 ) Summary: Add pickle input type to allow for running ptvsc2_predictor_bench to get individual node benchmarks for SR Test Plan: ``` buck2 run mode/opt caffe2/caffe2/fb/predictor:ptvsc2_predictor_bench -- --scripted_model=/data/users/georgiaphillips/models/742055223/1/742055223_1.predictor.local --pt_inputs=/data/users/georgiaphillips/models/742055223/0/mix.pt --pt_enable_static_runtime=1 --compare_results=0 --iters=1000 --warmup_iters=100 --num_threads=1 --do_profile=1 --method_name=${MODULE_NAME}.forward --set_compatibility --do_benchmark=1 --pytorch_predictor_default_model_id=${MODEL_ENTITY_ID}_${SNAPSHOT_ID} --input_type=pickle ``` Rollback Plan: Reviewed By: dolpm Differential Revision: D76554920 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155988 Approved by: https://github.com/dolpm	2025-06-16 19:15:07 +00:00
Julian De la Barrera Brandner	2dc1627451	[doc] Add documentation for division by zero behavior in autograd (#155987 ) Fixes #128796 This PR adds documentation about the behavior of division by zero operations in PyTorch's autograd system. The documentation explains: 1. How division by zero produces `inf` values following IEEE-754 floating point arithmetic 2. How autograd handles these cases and why masking after division can lead to `nan` gradients 3. Provides concrete examples showing the issue 4. Recommends two solutions: - Masking before division - Using MaskedTensor (experimental API) The documentation is added to the autograd notes section, making it easily discoverable for users who encounter this common issue. This addresses the original issue #128796 which requested better documentation of this behavior to help users avoid common pitfalls when dealing with division by zero in their models. dditional changes: - Fixed formatting consistency by replacing curly apostrophes with straight apostrophes in the existing documentation Pull Request resolved: https://github.com/pytorch/pytorch/pull/155987 Approved by: https://github.com/soulitzer Co-authored-by: sekyondaMeta <127536312+sekyondaMeta@users.noreply.github.com>	2025-06-16 19:02:12 +00:00
Simon Fan	907d0931cc	[ca] default on in CI, with fallback for tests in test/compiled_autograd_skips/ (#155480 ) For every test that is ran with PYTORCH_TEST_WITH_DYNAMO=1, turn on compiled autograd via config if it is not skipped Pull Request resolved: https://github.com/pytorch/pytorch/pull/155480 Approved by: https://github.com/jansel ghstack dependencies: #155521	2025-06-16 18:45:03 +00:00
Simon Fan	9ff9c28fe8	[ca] Functionalize AccumulateGrad (#155521 ) This PR changes compiled autograd's handling of gradient accumulation, by proxying it as a `call_accumulate_grad`, which does the .grad mutation in python bytecode for dynamo to see. For eager, the only change is the leaf invariant check was moved up. Before: - Compiled Autograd Engine: proxies call to inductor accumulate_grad op - Dynamo: polyfills the inductor accumulate_grad op (not respecting all of the accumulateGrad implementation e.g. sparse, gradient layout contract) ```python new_grad_strided: "f32[s21]" = torch.empty_like(getitem_1); getitem_1 = None copy_: "f32[s21]" = new_grad_strided.copy_(aot3_tangents_1); copy_ = None ``` - AOTAutograd: functionalizes the copy_ After: - Compiled Autograd Engine: proxies call to `call_accumulate_grad`, which calls `torch._dynamo.compiled_autograd.ops.AccumulateGrad`/`AccumulateGrad_apply_functional_no_hooks_ivalue`, similar to other functional autograd implementations, but also sets .grad from python. Hooks are still handled separately from this call. - Dynamo: `torch._dynamo.compiled_autograd.ops.AccumulateGrad` was allow_in_graph'd - AOTAutograd: traces into the op, with FunctionalTensors. While functionalizing the tensors, we insert an autograd Error node to ensure that we don't use the autograd meta from tracing. This clashes with the "leaf variable has been moved into the graph interior" error check, I could not find a way to identify a FunctionalTensor subclass from C++, so I bypass that for Error nodes in the compiled case. In the CI PR, this fixes 19 tests relating to sparse tensors, and more are hidden by an earlier failure in dynamo Pull Request resolved: https://github.com/pytorch/pytorch/pull/155521 Approved by: https://github.com/jansel	2025-06-16 18:45:02 +00:00
Benjamin Glass	42ff6a4a5c	[Inductor] Delay codegen for fallback arguments and improve typing (#154371 ) Delays code generation for arguments to fallback ops. This is inspired by #155642, and likely fixes similar memory leaks. Additionally, prepare for the next PR in the stack by tightening up typing on a `cpp_wrapper` interface that's only used in one (well-typed) place, as well as downstream effects of that change. In particular, this enabled: 1. removing a number of now clearly unnecessary asserts 2. adding a few more targeted asserts to validate the code's current assumptions 3. removing some unneeded control flow in several functions Pull Request resolved: https://github.com/pytorch/pytorch/pull/154371 Approved by: https://github.com/desertfire	2025-06-16 18:00:04 +00:00
Xuehai Pan	4162c0f702	[BE][setup] gracefully handle envvars representing a boolean in `setup.py` (#156040 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156040 Approved by: https://github.com/malfet	2025-06-16 17:56:31 +00:00
angelayi	f48a157660	[aoti] Add more to error message (#155974 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/155974 Approved by: https://github.com/yushangdi	2025-06-16 17:49:52 +00:00
windsonsea	fbd88ae2b5	Convert to markdown: checkpoint.rst (#156009 ) Related to #155014 Use two commits to have a try. ```bash 1800 git mv docs/source/checkpoint.rst docs/source/checkpoint.md 1802 git commit -m "[Docs] Rename checkpoint.rst" 1803 git push origin ckpoint # update the markdown file 1805 git add . 1806 git commit -m "modify checkpoint.md" 1807 git push origin ckpoint ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/156009 Approved by: https://github.com/svekars	2025-06-16 17:48:23 +00:00
windsonsea	a10024d7de	Convert complex_numbers.rst to markdown (#156039 ) Related to #155014 Have a try by following https://github.com/pytorch/pytorch/pull/155899#issuecomment-2974715750 Pull Request resolved: https://github.com/pytorch/pytorch/pull/156039 Approved by: https://github.com/svekars	2025-06-16 17:24:37 +00:00
PyTorch MergeBot	e9fdaf8701	Revert "[Quant][CPU] fix fake_quantize_per_tensor_affine of inf values (#155109 )" This reverts commit e375d21bb9b0ef6fefe7a8af5a054a17de8c63c9. Reverted https://github.com/pytorch/pytorch/pull/155109 on behalf of https://github.com/malfet due to Looks like it broke ROCM tests ([comment](https://github.com/pytorch/pytorch/pull/155109#issuecomment-2977428354))	2025-06-16 17:22:55 +00:00
Justin Chu	45596ec58f	Delete tools/onnx/update_default_opset_version.py (#156055 ) The tool is no longer relevant. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156055 Approved by: https://github.com/titaiwangms	2025-06-16 17:21:36 +00:00
PyTorch MergeBot	365ce465f3	Revert "[C10][CUDA] Eagerly create context on torch.cuda.set_device(device) call (#155900 )" This reverts commit 8142a0286016e63a0e91b5667e1fb1a5e868ffd7. Reverted https://github.com/pytorch/pytorch/pull/155900 on behalf of https://github.com/clee2000 due to causing some sort of hang? in test_distributed_spawn [GH job link](https://github.com/pytorch/pytorch/actions/runs/15678895788/job/44168117193) [HUD commit link](`8142a02860`) note to self: bad TD ([comment](https://github.com/pytorch/pytorch/pull/155900#issuecomment-2977365699))	2025-06-16 16:59:25 +00:00
albanD	2a4e357192	Fix compilation warning with gcc14 (#155934 ) Note that nccl still doesn't work so you have to build with `USE_NCCL=0` @eqy is that something being tracked there? Pull Request resolved: https://github.com/pytorch/pytorch/pull/155934 Approved by: https://github.com/malfet, https://github.com/janeyx99	2025-06-16 16:43:15 +00:00
PyTorch MergeBot	503362d019	Revert "Unify dynamic shapes APIs naming 2 (expect_true and check) (#155776 )" This reverts commit 603a54a9b33e1aabe1407721d7935b881a160968. Reverted https://github.com/pytorch/pytorch/pull/155776 on behalf of https://github.com/atalman due to failing internal build ([comment](https://github.com/pytorch/pytorch/pull/155776#issuecomment-2977041192))	2025-06-16 15:13:53 +00:00
PyTorch MergeBot	b8d96c3f78	Revert "[cuBLASLt][cuBLAS] Support 2D bias and `beta != 1.0` in cuBLASLt (#154170 )" This reverts commit 47c8810b5275179833d6b33ca3d70922f485272c. Reverted https://github.com/pytorch/pytorch/pull/154170 on behalf of https://github.com/atalman due to failing torchrec tests ([comment](https://github.com/pytorch/pytorch/pull/154170#issuecomment-2976990461))	2025-06-16 14:59:01 +00:00
Xuehai Pan	013dfeabb4	[BE] fix typos in top-level files (#156067 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156067 Approved by: https://github.com/malfet ghstack dependencies: #156066	2025-06-16 14:56:07 +00:00
Xuehai Pan	6c493e2b14	[BE] add `codespell` linter (#156066 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156066 Approved by: https://github.com/malfet	2025-06-16 14:56:07 +00:00
Nikita Shulga	2d832c9587	[MPS][Testing][BE] Fix samples for full_like (#156026 ) Now that device is known, one can avoid creating tensors of `torch.double` type Pull Request resolved: https://github.com/pytorch/pytorch/pull/156026 Approved by: https://github.com/dcci	2025-06-16 14:27:42 +00:00
Nikita Shulga	831c9010c7	[BE] Remove non-existing operator from unimplemented list (#156025 ) Never heard of torch.login :) Pull Request resolved: https://github.com/pytorch/pytorch/pull/156025 Approved by: https://github.com/dcci	2025-06-16 14:14:58 +00:00
Edward Z. Yang	38e5e81e55	Implement guard collectives (#155558 ) When running a distributed job with compiler collectives enabled, if one rank recompiles while others do not, this leads to a deadlock (as not everyone will rendezvous with the compiler collective from the recompile). Although there aren't any convenient ways to cheaply solve this problem, if you are willing to force everyone to sync when evaluating guards, you can just force everyone to recompile if anyone requires a recompile. So the way guard collectives work is: 1. Perform compiled code lookup (evaluating guards) 2. Run a collective, communicating if you found a compiled code or not 3. If anyone requires recompile, force everyone to recompile One current deficiency in the implementation is we can't conveniently track the time it takes to run this collective. I need to test if we actually successfully are running the collective on a separate stream, or if we have to wait for user collectives to all finish. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/155558 Approved by: https://github.com/Microve	2025-06-16 14:09:14 +00:00
dependabot[bot]	05faba4028	Bump requests from 2.32.2 to 2.32.4 in /.github (#155491 ) Bumps [requests](https://github.com/psf/requests) from 2.32.2 to 2.32.4. - [Release notes](https://github.com/psf/requests/releases) - [Changelog](https://github.com/psf/requests/blob/main/HISTORY.md) - [Commits](https://github.com/psf/requests/compare/v2.32.2...v2.32.4) --- updated-dependencies: - dependency-name: requests dependency-version: 2.32.4 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2025-06-16 06:48:08 -07:00
PyTorch UpdateBot	d6ee5144ca	[xla hash update] update the pinned xla hash (#156064 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned xla hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156064 Approved by: https://github.com/pytorchbot	2025-06-16 11:11:10 +00:00
Aidyn-A	8142a02860	[C10][CUDA] Eagerly create context on torch.cuda.set_device(device) call (#155900 ) Fixes #155668 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155900 Approved by: https://github.com/ngimel	2025-06-16 10:55:47 +00:00
Anthony Barbier	bf7e290854	Add __main__ guards to jit tests (#154725 ) This PR is part of a series attempting to re-submit https://github.com/pytorch/pytorch/pull/134592 as smaller PRs. In jit tests: - Add and use a common raise_on_run_directly method for when a user runs a test file directly which should not be run this way. Print the file which the user should have run. - Raise a RuntimeError on tests which have been disabled (not run) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154725 Approved by: https://github.com/clee2000	2025-06-16 10:28:45 +00:00
Justin Chu	f810e98143	[ONNX] Update default opset to 18 (#156023 ) Update default opset for the torchscript exporter to 18 to match the dynamo exporter, because support was actaully added and tested in https://github.com/pytorch/pytorch/pull/118828. In the next version we should plan to update to opset 21 or higher. This change also removes the hard limit on the torchscript exporter for more flexibility. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156023 Approved by: https://github.com/Skylion007	2025-06-16 08:40:49 +00:00
Bob Ren	39c605e8b3	remove allow-untyped-defs from context.py (#155622 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155622 Approved by: https://github.com/Skylion007	2025-06-16 07:38:34 +00:00
Xia, Weiwen	d9799a2ee7	Support boolean tensor for torch.fused_moving_avg_obs_fake_quant on CUDA (#153699 ) Fixes #153310 As the title Test plan ``` pytest test/quantization/core/test_workflow_ops.py -k test_fused_obs_fake_quant_moving_avg ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153699 Approved by: https://github.com/mingfeima, https://github.com/jerryzh168	2025-06-16 07:10:06 +00:00
PyTorch UpdateBot	156b28e62a	[audio hash update] update the pinned audio hash (#155648 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155648 Approved by: https://github.com/pytorchbot	2025-06-16 03:57:28 +00:00
Dhia naouali	c620d0b5c7	convert: rst to myst pr2/2 (#155911 ) Fixes #155038 parent [PR](https://github.com/pytorch/pytorch/pull/155375) (made two PRs to pass sanity check) this PR converts the following three .rst files with the mentioned referenced in each file - [torch.compiler_faq](https://github.com/pytorch/pytorch/blob/main/docs/source/torch.compiler_faq.rst) - torch.compiler_troubleshooting - nonsupported_numpy_feats - torchdynamo_fine_grain_tracing - [torch.compiler_fine_grain_apis](https://github.com/pytorch/pytorch/blob/main/docs/source/torch.compiler_fine_grain_apis.rst) - None - [torch.compiler_get_started](https://github.com/pytorch/pytorch/blob/main/docs/source/torch.compiler_get_started.rst) - torch.compiler_overview - torch.compiler_api - torchdynamo_fine_grain_tracing I made the suggested edits by the maintainers as commented in the parent PR (used git mv on all files, yet it still appeared as delete-create action) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155911 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-16 00:44:44 +00:00
David Berard	c83041cac2	[test][triton pin] add device-side TMA tests (AOTI + test_triton_kernels) (#155827 ) Tests added: ``` python test/inductor/test_triton_kernels.py -k test_on_device_tma python test/inductor/test_triton_kernels.py -k test_add_kernel_on_device_tma python test/inductor/test_aot_inductor.py -k test_triton_kernel_on_device_tma ``` These pass on Triton 3.3 but not yet on Triton 3.4 (note: to support tests for both Triton versions, there's two triton kernels - one for old api and one for new api - and a given version of the test will only run if that version of the API is available). Pull Request resolved: https://github.com/pytorch/pytorch/pull/155827 Approved by: https://github.com/FindHao ghstack dependencies: #155777, #155814	2025-06-15 20:24:19 +00:00
David Berard	bc9b8ea230	[user triton] JIT inductor support for new host-side TMA api (#155814 ) This PR adds JIT inductor support for user-defined triton kernels using the new host-side TMA api. * handle TensorDescriptor.from_tensor in ir.py * codegen TensorDescriptor.from_tensor in wrapper.py * generate the right signature for functions that take TensorDescriptor arguments (i.e. in the @triton_heuristics.user_autotune decorator) AOTI support is not implemented yet. Tests: ran test_triton_kernels.py w/ both Triton 3.3 and 3.4 and there were no failures. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155814 Approved by: https://github.com/aakhundov ghstack dependencies: #155777	2025-06-15 20:24:19 +00:00
David Berard	b7c95acc6c	[user triton] triton_kernel_wrap support for new host-side TMA API (#155777 ) This adds support for user-defined triton kernels using TensorDescriptor.from_tensor into triton_kernel_wrap: i.e. storing metadata about the TMA descriptors and doing mutation analysis. Major changes: * TMADescriptorMetadata has changed: previously it was a dict[str, tuple[list[int], list[int], int]]. But now there are two metadata formats: one for experimental API and one for stable API. Now the metadata format is dict[str, tuple[str, tuple[...]]], where tuple[...] is tuple[list[int], list[int], int] for experimental and tuple[list[int],] for stable API. And then most handling of the metadata has to be branched based on whether the metadata represents a stable or experimental TMA descriptor * mutation analysis: unlike experimental TMA (where the mutation analysis / ttir analysis pretends that the TMA descriptor is actually just a tensor), we need to construct an actual TMA descriptor before getting the Triton frontend to create the TTIR (otherwise assertions fail). A TensorDescriptor (i.e. stable TMA API descriptor) passed into a python triton kernel actually turns into 1 + 2N parameters in the TTIR (for a rank-N tensor), so the arg list also needs to be patched for this reason (in generate_ttir) mutation analysis: now we also need to pass tma_descriptor_metadata into the mutation analysis, in order to create the TMA descriptors that are passed into the frontend code (ie. the previous point). This is why all the mutation tests are modified with an extra return value (the tma_descriptor_metadata) Inductor is not modified (Inductor just errors out if you use a stable API tma descriptor). This will be the next PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155777 Approved by: https://github.com/aakhundov	2025-06-15 20:24:19 +00:00
Animesh Jain	54976bca10	[dynamo] Provide helper functions for guard filter hook (#155083 ) Collection of ready-made guard filters. One issue is that they are not composable - `filter1(filter2(guard))`. On the other hand, they are easy to use. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155083 Approved by: https://github.com/zhxchen17, https://github.com/jansel	2025-06-15 17:49:36 +00:00
Yu, Guangye	0935a97d95	[Dynamo] Add torch.accelerator API to trace_rules (#155884 ) # Motivation - Add binding API and non-binding API in torch.accelerator to trace rules. - Add some function in torch.accelerator to const fold functon list for Dynamo capature. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155884 Approved by: https://github.com/jansel, https://github.com/EikanWang ghstack dependencies: #155787, #155788	2025-06-15 17:09:57 +00:00
Yu, Guangye	b51d803785	[Dynamo] Add XPU API to trace_rules (#155788 ) # Motivation - Add binding API and non-bindling API to trace rules for XPU; - Add some XPU API to the const fold function for Dynamo capture. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155788 Approved by: https://github.com/jansel, https://github.com/EikanWang ghstack dependencies: #155787	2025-06-15 17:09:57 +00:00
Yu, Guangye	69acba2b19	[Dynamo] Add generic and XPU-specific Stream&Event in UserDefineClass (#155787 ) # Motivation - Add XPU-specific Stream and Event to in graph calss list for Dynamo capture. - Add generic Stream and Event to i graph class list for Dynamo capture. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155787 Approved by: https://github.com/jansel, https://github.com/EikanWang	2025-06-15 17:09:57 +00:00
nirajkamalk	53cd18f6b3	Update gradient behavior note in torch.amin and torch.amax (#155071 ) Fixes #155048 The behavior of `min` and `max` were changed in #43519. The note about gradient behavior in torch.amin and torch.amax docs are updated to reflect this change: New note: `amax, amin, max(dim), min(dim) evenly distributes gradient between equal values when there are multiple input elements with the same minimum or maximum value.` cc - @spzala @svekars @soulitzer @sekyondaMeta @AlannaBurke @ezyang @gqchen @nikitaved @Varal7 @xmfan Pull Request resolved: https://github.com/pytorch/pytorch/pull/155071 Approved by: https://github.com/soulitzer	2025-06-15 16:09:31 +00:00
PyTorch UpdateBot	655b3b14ff	[executorch hash update] update the pinned executorch hash (#156007 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned executorch hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/156007 Approved by: https://github.com/pytorchbot	2025-06-15 04:51:37 +00:00
Austin Wahle	517d2995e0	Add__int__ and __float__ methods to _sympy.functions.Identity (#155873 ) Fixes #155688 Root Cause: in [`torch/_inductor/index_propagation.py`](`f151b20123/torch/_inductor/index_propagation.py (L57-L68)`) When creating a `TypedExpr` from an `Identity` (a `torch.utils._sympy.functions.Identity`, not a `sympy.matrices.expressions.Identity `) and the inner value of the identity, `Identity.args[0]`, is any torch int type, the `TypedExpr.__post_init__` method tries to cast the Identity object to a python `int`. This is where to `TypeError` from the issue was raised, because Identity does not know how to cast to an `int`. Fix: Define `__int__` method for `torch.utils._sympy.functions.Identity`. wlog for `float` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155873 Approved by: https://github.com/williamwen42	2025-06-15 04:24:40 +00:00
Aaron Gokaslan	6ebe9a4f47	[BE][Ez]: Optimize nvshmem alloc with missing move (#156000 ) Saw this in another PR where there was a missing move on this potentially very hot path with Pull Request resolved: https://github.com/pytorch/pytorch/pull/156000 Approved by: https://github.com/kwen2501, https://github.com/cyyever	2025-06-15 03:04:08 +00:00
Ke Wen	32eee8ed22	[SymmMem] Add nvshmem_free (#155975 ) Stack from [ghstack](https://github.com/ezyang/ghstack) (oldest at bottom): Calling `nvshmem_free` when an `NVSHMEMAllocation` is being destructed. Use a `is_finalizing()` as a guard as done in `CUDASymmetricMemory.cu` to avoid "driver shutting down" error (destruction fiasco). Pull Request resolved: https://github.com/pytorch/pytorch/pull/155975 Approved by: https://github.com/ngimel ghstack dependencies: #155506, #155835, #155968, #155971	2025-06-15 01:23:49 +00:00
fduwjj	b8aee84fb9	[c10d][fr] Shrink the range of mutex lock to avoid deadlock (#155949 ) While looking into a case when FR dump (actual dump not monitoring thread) takes 30 mins, I realized that our global write lock is grabbed too early so the second effort to dump FR without stack trace will fail because of a deadlock because the global write lock is still hold. So we should only grab the lock when we are ready to write so that we are less likely to keep the lock forever. Also I did an audit to the lock within FR as well and found that there is one place we can shrink as well. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155949 Approved by: https://github.com/Skylion007	2025-06-15 00:37:42 +00:00
Howard Huang	3159ee2ad3	Update test_schedule_multiproc to use world_size=2 (#155921 ) The multiproc schedule tests previously ran with world_size=2, and PP tests became flakier due to the longer pipeline execution, this is restoring previously behavior. This will fix the tests (https://github.com/pytorch/pytorch/issues/154373, https://github.com/pytorch/pytorch/issues/154391, https://github.com/pytorch/pytorch/issues/154408, https://github.com/pytorch/pytorch/issues/154443, https://github.com/pytorch/pytorch/issues/154481 In follow up PRs I will refactor the tests and move some tests to use large world sizes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155921 Approved by: https://github.com/fduwjj, https://github.com/Skylion007 ghstack dependencies: #155920	2025-06-15 00:24:18 +00:00
Howard Huang	8e1471bdc9	Allow MultiProcContinuousTest to set world_size (#155920 ) `MultiProcContinuousTest` will automatically set world_size to number of devices. This change allows this attribute to be modified by the derived test class Pull Request resolved: https://github.com/pytorch/pytorch/pull/155920 Approved by: https://github.com/fduwjj	2025-06-15 00:24:17 +00:00
Michael Lazos	9bd42c1570	[Cutlass] Fix buffer missing issues (#155897 ) Handles constants and constant folding with aoti. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155897 Approved by: https://github.com/henrylhtsang	2025-06-15 00:08:50 +00:00
Henry Tsang	a35b3a9b95	[cutlass backend][forward fix] use _cuda_compiler path to check if nvcc exists (#155939 ) Differential Revision: D76571828 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155939 Approved by: https://github.com/Skylion007, https://github.com/masnesral	2025-06-15 00:01:57 +00:00
eqy	47c8810b52	[cuBLASLt][cuBLAS] Support 2D bias and `beta != 1.0` in cuBLASLt (#154170 ) Fixes https://github.com/pytorch/pytorch/issues/153590 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154170 Approved by: https://github.com/malfet	2025-06-14 23:34:31 +00:00
bobrenjc93	0fa361e429	[ez] fix typo in _inductor/scheduler.py (#155996 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155996 Approved by: https://github.com/Skylion007 ghstack dependencies: #155982	2025-06-14 21:21:35 +00:00
Ke Wen	77ac3a0965	[SymmMem] Remove wrappers around nvshmem APIs (#155971 ) `NVSHMEMSymmetricMemory.cu` and `nvshmem_extension.cu` are under the same compilation condition now (i.e. only when `USE_NVSHMEM=True`), see https://github.com/pytorch/pytorch/blob/main/caffe2/CMakeLists.txt#L1013-L1018. Therefore there is no need to build an extra layer to hide dependency. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155971 Approved by: https://github.com/Skylion007 ghstack dependencies: #155506, #155835, #155968	2025-06-14 19:58:09 +00:00
Ke Wen	2c0d94a7de	[SymmMem] Remove unused ptr_to_symm_mem_ (#155968 ) No code enqueues entries to `ptr_to_symm_mem_`, thus it is always empty. This PR removes it and supports relying functionalities via the `allocations_` map. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155968 Approved by: https://github.com/Skylion007 ghstack dependencies: #155506, #155835	2025-06-14 19:57:06 +00:00
Aaron Gokaslan	a317c63d1b	[BE]: Update NCCL to 2.27.3 (#155233 ) Fixes: https://github.com/pytorch/pytorch/issues/155052 and https://github.com/pytorch/pytorch/issues/153517 This upgrade is needed to effectively use those symmetric memory kernels anyway. Also fixes some nasty NCCL bugs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155233 Approved by: https://github.com/nWEIdia, https://github.com/kwen2501, https://github.com/atalman, https://github.com/eqy	2025-06-14 19:20:31 +00:00
Jithun Nair	794ef6c9b8	Enable manywheel build and smoke test on main branch for ROCm (#153287 ) Fixes issue of not discovering breakage of ROCm wheel builds until the nightly job runs e.g. https://github.com/pytorch/pytorch/pull/153253 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153287 Approved by: https://github.com/jeffdaily	2025-06-14 19:14:31 +00:00
albanD	5285d10243	remove duplicated pybind flag in mps code (#155936 ) gcc14 (at least) warns that this is already defined Pull Request resolved: https://github.com/pytorch/pytorch/pull/155936 Approved by: https://github.com/Skylion007	2025-06-14 18:41:12 +00:00
Aaron Orenstein	e95e8eed0a	mypy 1.16.0 (#155821 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155821 Approved by: https://github.com/ezyang, https://github.com/zou3519	2025-06-14 18:18:43 +00:00
Marcin Pioch	ce79056471	Custom FX pass for inductor's backend registration (#154841 ) This PR is related to RFC #153532. It is an extension to Inductor's backend registration interface to allow to register custom FX passes by the backend. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154841 Approved by: https://github.com/jansel Co-authored-by: Jason Ansel <jansel@jansel.net>	2025-06-14 17:29:54 +00:00
Laith Sakka	603a54a9b3	Unify dynamic shapes APIs naming 2 (expect_true and check) (#155776 ) The functions guard_lt, guard_equals, and guard_leq work similarly to torch.check and expect_true, but they operate on SymPy expressions. Notably, guard_equals applies local replacements before comparison, which might be better extracted into a separate function. This pull request standardizes naming conventions to match symbolic_shapes.py. Specifically, - it introduces size_vars.expect_true and size_vars.check. - guard_lt becomes check_lt - guard_leq becomes check_leq - guard_equals becomes check_equals I am also seeing a couple of wrong usages !! that i will fix in the next PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/155776 Approved by: https://github.com/bobrenjc93 ghstack dependencies: #154774	2025-06-14 17:13:53 +00:00
Laith Sakka	c219dbd2fc	avoid gso in has_internal_overlap (#155870 ) existing comment already explains it Pull Request resolved: https://github.com/pytorch/pytorch/pull/155870 Approved by: https://github.com/bobrenjc93	2025-06-14 17:13:20 +00:00
Xuehai Pan	279cae52e7	[BE][PYFMT] migrate PYFMT for `torch/ao/` to `ruff format` (#148185 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/148185 Approved by: https://github.com/ezyang	2025-06-14 16:47:04 +00:00
cyy	c2beeadeb4	[Reland] Use 3.27 as the minimum CMake version (#154783 ) Reland of #153153, which was incidentally closed. Update the minimum CMake version to 3.27 because of it provides more CUDA targets such as CUDA::nvperf_host so that it is possible to remove some of our forked CUDA modules. See https://github.com/pytorch/pytorch/pull/153783. It's also possible to facilitate future third-party updates such as FBGEMM (its current shipped version requires 3.21). Pull Request resolved: https://github.com/pytorch/pytorch/pull/154783 Approved by: https://github.com/ezyang	2025-06-14 16:37:51 +00:00
Tugsbayasgalan (Tugsuu) Manlaibaatar	370fc49dde	Handle aten.to at submodule boundaries (#153972 ) Summary: #buildall Test Plan: CI Differential Revision: D74582970 When we decompose to inference IR, aten.to can sometimes disappear. As a result, export module call graph tree will start containing dead nodes because previous provenance tracking is insufficient. This PR fixes that. The caveat is that this won't work in general for tensor subclass inputs to submodule that user wants to preserve signature because we always desugar the tensor subclass into constituent tensors in inference IR making it impossible to preserve the original calling convention. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153972 Approved by: https://github.com/avikchaudhuri	2025-06-14 16:13:29 +00:00
PyTorch UpdateBot	d42c11819f	[executorch hash update] update the pinned executorch hash (#153436 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned executorch hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153436 Approved by: https://github.com/pytorchbot	2025-06-14 16:09:41 +00:00
bobrenjc93	70b68caf58	Fix logging of failed tensorified ops (#155982 ) Tested via ``` TORCH_LOGS="torch.fx.passes._tensorify_python_scalars" tlp python test/inductor/test_torchinductor_dynamic_shapes.py -k test_unspecialized_float_fallback_symint_specialization I0613 21:50:38.247000 4163366 torch/fx/passes/_tensorify_python_scalars.py:314] [0/1] Failed to tensorify <built-in function pow> I0613 21:50:38.247000 4163366 torch/fx/passes/_tensorify_python_scalars.py:314] [0/1] Failed to tensorify <built-in function floor> ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155982 Approved by: https://github.com/flaviotruzzi	2025-06-14 14:23:54 +00:00
Xia, Weiwen	e375d21bb9	[Quant][CPU] fix fake_quantize_per_tensor_affine of inf values (#155109 ) Fixes #154328 Summary Fail reason: The input value is infinity in float and it has undefined behavior to convert it to int64_t. On X86, it will be converted to the min value of int64_t, which is not expected. Fix: Clamping `(input * inv_scale + zero_point)` to `[quant_min, quant_max]` before converting it to int64_t. Test plan ``` pytest test/quantization/core/test_workflow_ops.py -k test_fake_quantize_per_tensor_affine_inf ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155109 Approved by: https://github.com/leslie-fang-intel, https://github.com/jerryzh168	2025-06-14 14:12:38 +00:00
Xuehai Pan	1a568f4e5d	[BE][Easy] bump `isort` to 6.0.1 (#155919 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155919 Approved by: https://github.com/Skylion007 ghstack dependencies: #155909, #155914	2025-06-14 12:29:01 +00:00
Xuehai Pan	5467765990	[BE][Easy] bump `ruff` to 0.11.13 (#155914 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155914 Approved by: https://github.com/Skylion007 ghstack dependencies: #155909	2025-06-14 12:29:01 +00:00
Xuehai Pan	736a15a81a	[torchgen] Fix `ruff format` for `# fmt: skip` comment for function signature (#155909 ) See also: - astral-sh/ruff#18658 This fix follows the suggestion from: - https://github.com/astral-sh/ruff/issues/18658#issuecomment-2970130276 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155909 Approved by: https://github.com/ezyang	2025-06-14 12:28:55 +00:00
Aaron Gokaslan	d859e65826	[DCP][Ez]: Fix broadcast_object bug in DCP utils (#155912 ) Fixes #152310. Broadcast_object is now symmetric with gather_object and scatter_object. It was likely a typo that wasn't fixed in https://github.com/pytorch/pytorch/pull/147675 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155912 Approved by: https://github.com/ezyang	2025-06-14 12:14:14 +00:00
Xuehai Pan	596b418391	[BE][PYFMT] migrate PYFMT for `{torch,test}/{nn,optim}/**` to `ruff format` (#144548 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/144548 Approved by: https://github.com/ezyang	2025-06-14 11:27:04 +00:00
penknife6153	3e38feb05f	[inductor] Add configuration control for CUTLASS operation selection. (#155770 ) Added a new configuration option `cutlass_enabled_ops` that allows users to control which operations use CUTLASS lowerings. By default, CUTLASS is enabled for all operations (maintaining backward compatibility), but users can now selectively enable it only for specific operations to optimize compilation time. Fixes #155718 ## Usage Examples ```bash # Enable CUTLASS for all operations (default behavior) export TORCHINDUCTOR_CUTLASS_ENABLED_OPS="ALL" # Enable CUTLASS only for matrix multiplication operations export TORCHINDUCTOR_CUTLASS_ENABLED_OPS="mm,addmm" # Enable CUTLASS only for batch operations export TORCHINDUCTOR_CUTLASS_ENABLED_OPS="bmm,baddbmm" # Disable CUTLASS for all operations export TORCHINDUCTOR_CUTLASS_ENABLED_OPS="" ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155770 Approved by: https://github.com/henrylhtsang	2025-06-14 08:19:54 +00:00
FFFrog	1982ec2d22	Add api info for torch._C._nn.pyi (#148405 ) APis involved are as followed: - adaptive_avg_pool2d - adaptive_avg_pool3d - binary_cross_entropy - col2im ISSUE Related: https://github.com/pytorch/pytorch/issues/148404 Pull Request resolved: https://github.com/pytorch/pytorch/pull/148405 Approved by: https://github.com/ezyang	2025-06-14 07:57:07 +00:00
Laith Sakka	7070ab3180	use guard_or_false in checkInBoundsForStorage (#155874 ) this was added in https://github.com/pytorch/pytorch/pull/147354, the comment already justify guard_or_false Pull Request resolved: https://github.com/pytorch/pytorch/pull/155874 Approved by: https://github.com/bobrenjc93	2025-06-14 07:21:26 +00:00
Laith Sakka	d79651571f	assume sparse tensor not coalesced_ gsv -> guard_or_false. (#155869 ) preserve current behavior. Generalize it such that no need for torch._check_is_size to opt into this, and make it work for more complex unbacked sizes with ranges [-inf, inf] Pull Request resolved: https://github.com/pytorch/pytorch/pull/155869 Approved by: https://github.com/bobrenjc93	2025-06-14 07:19:56 +00:00
Xuehai Pan	e7da21806f	[Easy][BE] update recommanded VS Code settings (#152760 ) Changes: - Remove old invalid settings and replace with new settings. - Add commonly used VS Code extensions to support `cmake`, `ruff`, `mypy`, `flake8`, `editorconfig`, and spell checker. Also, add corresponding settings. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152760 Approved by: https://github.com/drisspg	2025-06-14 07:11:10 +00:00
cyy	1393f71e07	Use CUDA language in generated CMakeLists.txt from cpp_builder.py (#155979 ) The CMake CUDA module has been deprecated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155979 Approved by: https://github.com/ezyang	2025-06-14 06:52:51 +00:00
David Berard	c843909d9e	[flex attention][triton pin] use new TMA API (#155771 ) Triton 3.4 will remove the experimental TMA APIs: https://github.com/triton-lang/triton/pull/6488. Ahead of this, we are replacing the experimental TMA API usage with the stable TMA API in flex attention. This means that flex attention TMA will stop working with Triton 3.2 or Triton 3.3/3.3.1 for now (but it should work for Triton 3.4 in the PyTorch 2.8 release, and Meta-internal triton 3.3.1fb, which have the new TMA API). This PR does the following: * replace the experimental TMA APIs with the stable TMA APIs * remove the workspace args. Testing: I ran test/inductor/test_flex_attention.py on a H100 with @mandroid6's PR #153662 patched in to turn on TMA [TODO: confirm results once all the local tests pass, but from the first 100 tests I ran locally, all the failing tests were also failing on #153662 alone] Note: When #153662 lands, turning on TMA support by default, it should be checking specifically for stable TMA API support (commented on PR) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155771 Approved by: https://github.com/mandroid6, https://github.com/nmacchioni	2025-06-14 06:34:16 +00:00
Oguz Ulgen	92b7ed6d07	Add Helion softmax test (#155976 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155976 Approved by: https://github.com/jansel	2025-06-14 05:53:21 +00:00
Phillip Liu	9338d85d45	[ProcessGroupNCCL] Added log when fr dump triggered from pipe (#155754 ) Summary: TSIA Created from CodeHub with https://fburl.com/edit-in-codehub Test Plan: eyes Sandcastle run Differential Revision: D76472617 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155754 Approved by: https://github.com/fduwjj, https://github.com/Skylion007	2025-06-14 04:34:29 +00:00
Ti-Tai Wang	bf897b4cea	[ONNX] Support 0/1 on dynamic dimension (#155717 ) Previous to this PR, the exporter does not support dynamic dim with traced inputs containing 0/1. But after https://github.com/pytorch/pytorch/pull/148696, this is supported by torch.export.export. This PR adds the patch to torch.onnx.export. However, there is still known pitfall existing because the difference between eager and export. Compiler needs to decide the exported shape ahead, and whether the "hidden broadcasting" being applied results in different export. For example, ```python import torch class Model(torch.nn.Module): def forward(self, x, y, z): return torch.cat((x, y), axis=1) + z model = Model() x = torch.randn(2, 3) y = torch.randn(2, 5) z = torch.randn(1, 8) model(x, y, z) DYN = torch.export.Dim.DYNAMIC ds = {0: DYN, 1: DYN} with torch.fx.experimental._config.patch(backed_size_oblivious=True): ep = torch.export.export(model, (x, y, z), dynamic_shapes=(ds, ds, ds)) print(ep) """ ExportedProgram: class GraphModule(torch.nn.Module): def forward(self, x: "f32[s7, s16]", y: "f32[s7, s43]", z: "f32[s7, s16 + s43]"): # sym_size_int: "Sym(s7)" = torch.ops.aten.sym_size.int(x, 0) sym_size_int_1: "Sym(s16)" = torch.ops.aten.sym_size.int(x, 1) sym_size_int_2: "Sym(s7)" = torch.ops.aten.sym_size.int(y, 0) sym_size_int_3: "Sym(s43)" = torch.ops.aten.sym_size.int(y, 1) sym_size_int_4: "Sym(s7)" = torch.ops.aten.sym_size.int(z, 0) sym_size_int_5: "Sym(s16 + s43)" = torch.ops.aten.sym_size.int(z, 1) # File: /home/titaiwang/pytorch/test_export.py:7 in forward, code: return torch.cat((x, y), axis=1) + z cat: "f32[s7, s16 + s43]" = torch.ops.aten.cat.default([x, y], 1); x = y = None # eq: "Sym(True)" = sym_size_int_2 == sym_size_int; sym_size_int_2 = None _assert_scalar_default = torch.ops.aten._assert_scalar.default(eq, "Runtime assertion failed for expression Eq(s58, s35) on node 'eq'"); eq = _assert_scalar_default = None add_1: "Sym(s16 + s43)" = sym_size_int_1 + sym_size_int_3; sym_size_int_1 = sym_size_int_3 = None eq_1: "Sym(True)" = add_1 == sym_size_int_5; add_1 = sym_size_int_5 = None _assert_scalar_default_1 = torch.ops.aten._assert_scalar.default(eq_1, "Runtime assertion failed for expression Eq(s16 + s43, s23) on node 'eq_1'"); eq_1 = _assert_scalar_default_1 = None eq_2: "Sym(True)" = sym_size_int == sym_size_int_4; sym_size_int = sym_size_int_4 = None _assert_scalar_default_2 = torch.ops.aten._assert_scalar.default(eq_2, "Runtime assertion failed for expression Eq(s35, s7) on node 'eq_2'"); eq_2 = _assert_scalar_default_2 = None # File: /home/titaiwang/pytorch/test_export.py:7 in forward, code: return torch.cat((x, y), axis=1) + z add: "f32[s7, s16 + s43]" = torch.ops.aten.add.Tensor(cat, z); cat = z = None return (add,) Graph signature: # inputs x: USER_INPUT y: USER_INPUT z: USER_INPUT # outputs add: USER_OUTPUT Range constraints: {s7: VR[0, int_oo], s16: VR[0, int_oo], s43: VR[0, int_oo], s16 + s43: VR[0, int_oo]} """ ep.module()(x, y, z) """ Traceback (most recent call last): File "/home/titaiwang/pytorch/test_export.py", line 20, in <module> ep.module()(x, y, z) File "/home/titaiwang/pytorch/torch/fx/graph_module.py", line 840, in call_wrapped return self._wrapped_call(self, args, kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/titaiwang/pytorch/torch/fx/graph_module.py", line 416, in __call__ raise e File "/home/titaiwang/pytorch/torch/fx/graph_module.py", line 403, in __call__ return super(self.cls, obj).__call__(args, *kwargs) # type: ignore[misc] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/titaiwang/pytorch/torch/nn/modules/module.py", line 1767, in _wrapped_call_impl return self._call_impl(args, *kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/titaiwang/pytorch/torch/nn/modules/module.py", line 1873, in _call_impl return inner() ^^^^^^^ File "/home/titaiwang/pytorch/torch/nn/modules/module.py", line 1800, in inner args_kwargs_result = hook(self, args, kwargs) # type: ignore[misc] ^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/titaiwang/pytorch/torch/_dynamo/eval_frame.py", line 895, in _fn return fn(args, *kwargs) ^^^^^^^^^^^^^^^^^^^ File "/home/titaiwang/pytorch/torch/export/_unlift.py", line 83, in _check_input_constraints_pre_hook _check_input_constraints_for_graph( File "/home/titaiwang/pytorch/torch/_export/utils.py", line 426, in _check_input_constraints_for_graph _check_symint( File "/home/titaiwang/pytorch/torch/_export/utils.py", line 338, in _check_symint raise RuntimeError( RuntimeError: Expected input at args[2].shape[0] to be equal to 2, but got 1 """ ``` The explanation (from @pianpwk): In the model we have `return torch.cat((x, y), axis=1) + z`. Before this add is executed, the LHS has shape `[s7, s16 + s43]`, while the z has shape, say `[s8, s16 + s43]` (we don't know `s7 == s8` yet). When we execute this add, the compiler is making a decision: does broadcasting apply or not? The choices are: 1) Yes -> then we must specialize `s8` to 1 2) No -> then this element-wise op is only valid if the shapes match up, and we assume `s7 == s8`. Unfortunately export can only follow one of these options, and in avoiding 0/1 specialization (because a dynamic dimension was requested), it assumed case 2). For an operation like a + b, in eager semantics it's possible to have all options (either a == 1 OR b == 1 OR a == b), but with export we need to make a decision on what the output shape of this operation is, and keeping all branches alive requires expressing the output shape with a conditional (e.g. output shape = `a if b == 1 else b`), which is pretty hard for the compiler to reason about. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155717 Approved by: https://github.com/justinchuby	2025-06-14 04:04:47 +00:00
FFFrog	187828dcb4	[OpenReg][5/N] add set_.source_Storage for openreg (#155191 ) Changes: - add set_.source_Storage for openreg to support torch.load & torch.serialization - uncomment some related tests in the test_openreg.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/155191 Approved by: https://github.com/albanD ghstack dependencies: #153947, #154018, #154019, #154106, #154181, #155101	2025-06-14 03:44:32 +00:00
FFFrog	e4fd0bf771	[OpenReg][4/N] Migrate cpp_extensions_open_device_registration to OpenReg (#155101 ) As the title stated. Involved testcases: - test_open_device_storage_pin_memory - test_open_device_serialization Pull Request resolved: https://github.com/pytorch/pytorch/pull/155101 Approved by: https://github.com/albanD ghstack dependencies: #153947, #154018, #154019, #154106, #154181	2025-06-14 03:44:32 +00:00
FFFrog	1e7989cad5	[OpenReg][3/N] Migrate cpp_extensions_open_device_registration to OpenReg (#154181 ) As the title stated. Involved testcases: - test_open_device_quantized - test_open_device_random - test_open_device_tensor - test_open_device_packed_sequence - test_open_device_storage Pull Request resolved: https://github.com/pytorch/pytorch/pull/154181 Approved by: https://github.com/albanD ghstack dependencies: #153947, #154018, #154019, #154106	2025-06-14 03:44:32 +00:00
FFFrog	7e5f29b2de	[OpenReg][2/N] Migrate cpp_extensions_open_device_registration to OpenReg (#154106 ) As the title stated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154106 Approved by: https://github.com/nareshrajkumar866, https://github.com/albanD ghstack dependencies: #153947, #154018, #154019	2025-06-14 03:44:32 +00:00
FFFrog	676abded4b	[OpenReg][1/N] Migrate cpp_extensions_open_device_registration to OpenReg (#154019 ) As the title stated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154019 Approved by: https://github.com/albanD ghstack dependencies: #153947, #154018	2025-06-14 03:44:32 +00:00
FFFrog	d3d469092f	[Openreg] Split TestOpenReg into two parts (#154018 ) ---- - TestPrivateUse1: testing 3rd accelerator integration mechinasm itself - TestOpenReg: testing openreg itself Pull Request resolved: https://github.com/pytorch/pytorch/pull/154018 Approved by: https://github.com/albanD ghstack dependencies: #153947	2025-06-14 03:44:31 +00:00
FFFrog	cafd2344d6	[OpenReg] add manual_seed related capabilities (#153947 ) Changes: - Add manual_seed manual_seed_all initial_seed and so on - Delay execution of self._lazy_init more deeply Pull Request resolved: https://github.com/pytorch/pytorch/pull/153947 Approved by: https://github.com/albanD	2025-06-14 03:44:31 +00:00
Sean McGovern	297805fd8f	Typo fixes for "overridden" in comments and function names (#155944 ) This word appears often in class descriptions and is not consistently spelled. Update comments and some function names to use the correct spelling consistently. Facilitates searching the codebase. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155944 Approved by: https://github.com/Skylion007	2025-06-14 03:37:38 +00:00
Edgar Romo Montiel	ca3cabd24a	Convert to markdown: named_tensor.rst, nested.rst, nn.attention.bias.rst, nn.attention.experimental.rst, nn.attention.flex_attention.rst #155028 (#155696 ) Fixes #155028 This pull request updates the documentation by transitioning from .rst to .md format. It introduces new Markdown files for the documentation of named_tensor, nested, nn.attention.bias, nn.attention.experimental, and nn.attention.flex_attention Pull Request resolved: https://github.com/pytorch/pytorch/pull/155696 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-14 03:32:00 +00:00
dolpm	cdfa33a328	[nativert] move execution frame to torch (#155830 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D76369008 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155830 Approved by: https://github.com/zhxchen17	2025-06-14 03:28:55 +00:00
Nicolas Macchioni	a6084b71ed	[BE][1/X] Phase out usage of `use_max_autotune()` (#155847 ) `use_max_autotune()` is likely not what people expect it to be; Originally, `use_max_autotune()` was setup to decide when we should include Triton templates as choices in GEMM autotuning. As expected, `use_max_autotune()=True` if `max_autotune=True` or `max_autotune_gemm=True`. However, with the addition of the offline GEMM autotuning cache two years back `use_max_autotune()=True` also in the case that `search_autotune_cache=True`; in this case though, `search_autotune_cache=True` should never trigger autotuning. Over time, people have used `use_max_autotune()` likely without realizing that this gives unexpected behavior if `search_autotune_cache=True`. We could rename the method to be more clear, but prefer to phase it out entirely for maximal clarity. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155847 Approved by: https://github.com/jingsh, https://github.com/masnesral	2025-06-14 03:16:20 +00:00
Yiming Zhou	7982b8c703	[BE][AOTI] Remove duplicate schema for ExternKernelNode (#155867 ) Summary: The definition of `ExternKernelNode` and `ExternKernelNodes` schema in `torch/_export/serde/aoti_schema.py` is a complete duplicate of the ones in `torch/_export/serde/schema.py`. Test Plan: CI Rollback Plan: Differential Revision: D76558294 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155867 Approved by: https://github.com/jingsh	2025-06-14 02:03:27 +00:00
Yiming Zhou	8f5f01bf19	[BE][AOTI] Combine DynamicArgType in Proxy Executors (#155871 ) Summary: As title. Move the duplicate definition to the base class header `proxy_executor.h` Test Plan: CI Rollback Plan: Differential Revision: D76559180 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155871 Approved by: https://github.com/yushangdi	2025-06-14 01:52:43 +00:00
PyTorch MergeBot	4574b39aa4	Revert "[BE]: Sync cusparselt 12.9 with static build and other cuda 12 (#155709 )" This reverts commit bbbced94a43cf764ddfe719e7d4c161a3992830c. Reverted https://github.com/pytorch/pytorch/pull/155709 on behalf of https://github.com/clee2000 due to broke lint [GH job link](https://github.com/pytorch/pytorch/actions/runs/15645591737/job/44082402642) [HUD commit link](`bbbced94a4`) landrace with 155819? easy forward fix but its the end of the week so idk when id get a review ([comment](https://github.com/pytorch/pytorch/pull/155709#issuecomment-2972094849))	2025-06-14 01:43:16 +00:00
Nikita Shulga	c10339559d	[BE] Better uv detection in `pip init` (#155972 ) If one has some UV and non-UV environments locally, one shoudl call `uv pip install` only on the UV-enabled ones, which could be detected by checking if `uv/python` path is present in `sys.base_prefix` Fixes https://github.com/pytorch/pytorch/issues/152999 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155972 Approved by: https://github.com/janeyx99	2025-06-14 01:35:50 +00:00
PyTorch MergeBot	d7e3c9ce82	Revert "Enable manywheel build and smoke test on main branch for ROCm (#153287 )" This reverts commit 3b6569b1ef4b9ff25f5b75fe0a216d6d084d573f. Reverted https://github.com/pytorch/pytorch/pull/153287 on behalf of https://github.com/clee2000 due to broke lint [GH job link](https://github.com/pytorch/pytorch/actions/runs/15646152483/job/44083912145) [HUD commit link](`3b6569b1ef`) ([comment](https://github.com/pytorch/pytorch/pull/153287#issuecomment-2972088294))	2025-06-14 01:32:27 +00:00
anwang	c165b36a31	[MTIA Aten Backend] Migrate relu / relu_ (#155927 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate relu / relu_. Note: Pytorch in-tree implementation delegates relu to clamp_min, so no more need to launch relu kernel. https://www.internalfb.com/code/fbsource/[0c9eedb2fc8f99bcca00cb67a5738cfe07e39349]/fbcode/caffe2/aten/src/ATen/native/Activation.cpp?lines=512-520 Let me know if any concern about this Differential Revision: [D75803582](https://our.internmc.facebook.com/intern/diff/D75803582/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155927 Approved by: https://github.com/egienvalue ghstack dependencies: #154632, #154659, #155925, #155926	2025-06-14 01:24:48 +00:00
anwang	50f6431e0a	[MTIA Aten Backend] Migrate sqrt.out / rsqrt.out / sin.out / silu.out (#155926 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate sqrt.out / rsqrt.out / sin.out / silu.out Differential Revision: [D75801847](https://our.internmc.facebook.com/intern/diff/D75801847/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155926 Approved by: https://github.com/egienvalue ghstack dependencies: #154632, #154659, #155925	2025-06-14 01:24:48 +00:00
anwang	7b11cb8c12	[MTIA Aten Backend] Migrate tanh.out and tanh_backward.grad_input (#155925 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate tanh.out and tanh_backward.grad_input Differential Revision: [D75769242](https://our.internmc.facebook.com/intern/diff/D75769242/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155925 Approved by: https://github.com/egienvalue ghstack dependencies: #154632, #154659	2025-06-14 01:24:41 +00:00
anwang	0185d3a5ed	[MTIA Aten Backend] Migrate bitwise_or.Tensor_out (#154659 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate bitwise_or.Tensor_out from out-of-tree to in-tree. Differential Revision: [D75629937](https://our.internmc.facebook.com/intern/diff/D75629937/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154659 Approved by: https://github.com/egienvalue ghstack dependencies: #154632	2025-06-14 01:24:34 +00:00
anwang	163cdaaa3a	[MTIA Aten Backend] Migrate bitwise_not.out (#154632 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff Migrate bitwise_not.out from out-of-tree to in-tree. Differential Revision: [D75610643](https://our.internmc.facebook.com/intern/diff/D75610643/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154632 Approved by: https://github.com/egienvalue	2025-06-14 01:24:27 +00:00
Jinhang Choi	04cf2c9d24	fix tensor print behavior for MAIA (#155609 ) This pull request fixes the tensor print behavior for `MAIA` to account for the absence of double-precision support in its backend. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155609 Approved by: https://github.com/soulitzer	2025-06-14 01:04:12 +00:00
cz2h	dabb55baff	Add resolve in add decomp to enable view (#153945 ) Fixes #148950. During the construction of graph and running the node of add under [interpreter](/github.com/pytorch/pytorch/blob/d68d4d31f4824f1d1e0d1d6899e9879ad19b0754/torch/fx/interpreter.py#L301 ), the functional argument of conj complex tensor gets cloned. This result in always having .is_conj() evaluted to false in decomposition function. Propose a fix of calling resolve_conj() in the decomposition of complex tensor add. Test as below `python test/dynamo/test_repros.py ReproTests.test_add_complex_conj` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153945 Approved by: https://github.com/jansel	2025-06-14 00:41:50 +00:00
Nikita Shulga	fec571cfd4	[BE][CI] Remove hardshrink integer exclusions (#155965 ) As they are not called anyway Pull Request resolved: https://github.com/pytorch/pytorch/pull/155965 Approved by: https://github.com/dcci	2025-06-14 00:32:57 +00:00
Boyuan Feng	38410cf9b5	Fix DDPOptimizer issue on static tensor index (#155746 ) We rely on `_try_get_metadata_from_dynamo()` to get static input indices. When the meta info is missing, it just returns an empty list of static input indices. This wrong list of static input indices lead to repeated cudagraph re-recording, which looks like a hang from the user perspective. `bc3972b80a/torch/_functorch/aot_autograd.py (L1025-L1031)` The root cause is `split_module` in DDP Optimizer loses meta info and gm attributes. This PR fixes the issue by propagating these metadata from original module to submodules. `bc3972b80a/torch/_dynamo/backends/distributed.py (L515-L517)` Fixes #140395 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155746 Approved by: https://github.com/xmfan, https://github.com/bdhirsh	2025-06-14 00:15:58 +00:00
Jithun Nair	3b6569b1ef	Enable manywheel build and smoke test on main branch for ROCm (#153287 ) Fixes issue of not discovering breakage of ROCm wheel builds until the nightly job runs e.g. https://github.com/pytorch/pytorch/pull/153253 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153287 Approved by: https://github.com/jeffdaily	2025-06-14 00:05:57 +00:00
Aaron Gokaslan	bbbced94a4	[BE]: Sync cusparselt 12.9 with static build and other cuda 12 (#155709 ) followup for https://github.com/pytorch/pytorch/pull/154980 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155709 Approved by: https://github.com/tinglvv, https://github.com/atalman, https://github.com/nWEIdia, https://github.com/cyyever	2025-06-13 23:10:01 +00:00
Nikita Shulga	d512584718	[BE] Refactor clamp dtypes check (#155930 ) By introducing `check_for_unsupported_clamp_dtypes` similar to `check_for_unsupported_isin_dtypes` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155930 Approved by: https://github.com/albanD, https://github.com/janeyx99, https://github.com/clee2000 ghstack dependencies: #155470	2025-06-13 23:05:02 +00:00
Nikita Shulga	0cb85c188f	[BE] Move optional submodules checkout to its own module (#155947 ) To expand it to optional eigen checkout later Pull Request resolved: https://github.com/pytorch/pytorch/pull/155947 Approved by: https://github.com/Skylion007	2025-06-13 23:02:38 +00:00
dggaytan	3003c681ef	Converting .rst files to .md files (#155377 ) Fixes #155036 This pull request updates the documentation for several modules by transitioning from .rst to .md format, improving readability and usability. It introduces new Markdown files for the documentation of torch.ao.ns._numeric_suite, torch.ao.ns._numeric_suite_fx, AOTInductor, AOTInductor Minifier, and the torch.compiler API Pull Request resolved: https://github.com/pytorch/pytorch/pull/155377 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-13 22:54:27 +00:00
ggsmith842	799443605b	Convert to markdown: distributed.tensor.parallel.rst, distributed.tensor.rst, distributions.rst, dlpack.rst (#155297 ) Fixes #155019 ## Description Convert to markdown: distributed.tensor.parallel.rst, distributed.tensor.rst, distributions.rst, dlpack.rst ## Checklist - [X] dlpack.rst converted to dlpack.md --> [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155297/dlpack.html) - [X] distributions.rst converted to distributions.md --> [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155297/distributions.html) - [X] distributed.tensor.rst converted to distributed.tensor.md --> [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155297/distributed.tensor.html) - [X] distributed.tensor.parallel.rst converted to distributed.tensor.parallel.md --> [Preview](https://docs-preview.pytorch.org/pytorch/pytorch/155297/distributed.tensor.parallel.html) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155297 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-13 22:08:37 +00:00
Nikita Shulga	764c02b78b	[BE] Raise `NotImplementedError` (#155470 ) When op is unimplemented for a specific dtype Which makes more sense, than a RuntimeError Example ```python >>> import torch >>> torch.nn.Hardshrink()(torch.randint(0, 5, (10,))) NotImplementedError: "hardshrink_cpu" not implemented for 'Long' ``` release notes bc-breaking: After this release `NotImplementedError` exception will be raised when ATen operation is called on the combinaiton of input tensor dtypes it has not been implemented for Mark few more unary ops as unimplemented to satisfy foreach testing error reporting consistency between CPU and CUDA Pull Request resolved: https://github.com/pytorch/pytorch/pull/155470 Approved by: https://github.com/albanD, https://github.com/Skylion007	2025-06-13 22:07:03 +00:00
Catherine Lee	d59ed21d0f	[CI] Reuse old whl: track why failed to use the old whl (#155860 ) As in title Any other things I should track? Pull Request resolved: https://github.com/pytorch/pytorch/pull/155860 Approved by: https://github.com/malfet	2025-06-13 22:01:31 +00:00
Catherine Lee	3596c0c77f	Fix test after revert (#155946 ) ex test_dynamic_shapes.py::TestUbackedOps::test_unbacked_reshape2 [GH job link](https://github.com/pytorch/pytorch/actions/runs/15642199583/job/44073674212) [HUD commit link](`06408dae49`) started after 06408dae49d06b6146fdd9d7a37eb5dde4f5e78d idk what the test does so maybe theres a better way to fix this Pull Request resolved: https://github.com/pytorch/pytorch/pull/155946 Approved by: https://github.com/yangw-dev, https://github.com/huydhn, https://github.com/malfet	2025-06-13 21:52:07 +00:00
Catherine Lee	eef253d9f6	[CI] Keep going display on HUD: upload log when test fails (#155371 ) I guess this is more of an RFC Goal: Enable keep going so that we can get information immediately for failures. We want be aware of failures as soon as possible, especially on the main branch, this is so that reverts can happen quickly. Proposal: A job with `keep-going` will continue through errors in `python run_test.py`. If a test fails, before it runs the next test, it will upload a fake log that should have enough information in it so that viewing the log will be able to tell you what failed and any stack traces/error logs, and should be able to be parsed by log classifier to get a line. I am getting the log by concating the test logs in test/test-reports, which is all the text outputted by pytest (unless someone runs with `ci-verbose-test-logs` label). There are obviously many things this won't catch, ex output outside of run_test.py, some output inside of run_test.py, but it should be enough. After a log finishes, eventually its raw log is uploaded to ossci-raw-job-status s3 bucket and the log classifier will read it to do classification. This means we will have to change log classifier to read from this bucket as well. I'm thinking just add an input parameter to log classifier like https://github.com/pytorch/test-infra/pull/6723/files Also upload the temp results to a temp attribute instead of the real one To overwrite the conclusion on HUD, I'm thinking a lambda that is s3 put trigger on the fake log being put into s3, that does something similar to log classifier where it just mutates the entry `13a990b678/aws/lambda/log-classifier/src/network.rs (L85)` to add a new field like "will_fail": true, and also triggers the log classifier to run Then we change HUD/ClickHouse to point the raw log url to the alternate place, the new "will_fail" field as the conclusion, and the temp log classifier result if needed Why always write to temp attribution/column? I am unsure about overwriting the real results with fake ones Pros: Not many changes outside of HUD/UI Cons: Lots of moving parts, lots of temp fields that will require adjustment for queries, temp fields never really get deleted Pull Request resolved: https://github.com/pytorch/pytorch/pull/155371 Approved by: https://github.com/malfet	2025-06-13 21:21:55 +00:00
mori360	e5ed267f83	Update h100-distributed image (#155861 ) Move non inductor workflows cuda 12.6->cuda 12.8 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155861 Approved by: https://github.com/seemethere	2025-06-13 21:17:05 +00:00
Joona Havukainen	20a74c370b	Add error message with assert to topK if ndims() - dim > 4 (#155475 ) Addressing #154890 Not really a proper fix but at least it's more informative than the current crash. For a more long term solution I'm testing if we can use the TopK API released in MacOS14 as it does not have the same MPSScan op issue that the Sort and ArgSort are hitting. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155475 Approved by: https://github.com/kulinseth	2025-06-13 21:10:06 +00:00
Runtian (Rachel) Li	049dc48d1e	fix code chunk indentation for `jit_language_reference_v2.md` (#155937 ) Fixes https://github.com/pytorch/pytorch/issues/155023 Related PR: #155781 Description: As discussed, this PR is a follow-up update for `jit_language_reference_v2.md` by deleting the code chunk indentation. Checklist: - [x] The issue being fixed is referenced above (Fixes https://github.com/pytorch/pytorch/issues/155023) - [x] Only one issue is addressed in this pull request - [x] Labels from the issue that this PR is fixing are added to this pull request - [x] No unnecessary issues are included into this pull request. @pytorchbot label "topic: docs" @pytorchbot label "topic: not user facing" @pytorchbot label docathon-h1-2025 @pytorchbot label "module: docs" Pull Request resolved: https://github.com/pytorch/pytorch/pull/155937 Approved by: https://github.com/jingsh, https://github.com/svekars	2025-06-13 21:05:23 +00:00
GdoongMathew	731351bb4a	Convert rst to markdown - optim.rst #155031 (#155813 ) Fixes #155031 ![image](https://github.com/user-attachments/assets/36507ca1-eb1e-4358-9e66-ce25ec8a2be1) @pytorchbot label "docathon-h1-2025" "module: docs" "topic: not user facing" "topic: docs" Pull Request resolved: https://github.com/pytorch/pytorch/pull/155813 Approved by: https://github.com/AlannaBurke	2025-06-13 21:03:39 +00:00
Benjamin Glass	92388bb2ab	[export] Remove broken check for multiple cpp files in PT2 package (#155149 ) This check was recently added, but (when fixed to refer to CPP rather than library files) fails with the separate kernel and wrapper build of AOTInductor. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155149 Approved by: https://github.com/angelayi	2025-06-13 21:02:31 +00:00
windsonsea	7d1b3f599d	[Docs] Convert to markdown cond.rst, config_mod.rst (#155653 ) Related to #155014 Only included 2 files in this PR: - cond.rst - config_mod.rst Pull Request resolved: https://github.com/pytorch/pytorch/pull/155653 Approved by: https://github.com/svekars	2025-06-13 20:58:57 +00:00
henrylhtsang	fdf5d97fa8	[cutlass backend][ez] Log timings from prescreening (#155757 ) Differential Revision: [D76474669](https://our.internmc.facebook.com/intern/diff/D76474669/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155757 Approved by: https://github.com/ColinPeppler	2025-06-13 20:44:04 +00:00
Justin Silver	f3e6c8e834	Fix #155016 for Docathon - convert rst to markdown (#155198 ) Used [rst2myst tool](https://rst-to-myst.readthedocs.io/en/latest/) One note is that "Created On" and "Last Updated On" banner doesn't show in the markdown files... I'm not sure if that's just an artifact of my local build though. Fixes #155016 Docs comparison (check out the 'new' whenever docs build) 1. cuda ([old](https://docs.pytorch.org/docs/main/cuda.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155198/cuda.html)) 2. cuda.tunable ([old](https://docs.pytorch.org/docs/main/cuda.tunable.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155198/cuda.tunable.html)) 3. leave cudnn_persistent_rnn.rst as is because it's reused in docstrings 4. cudnn_rnn_determinism.rst as is because it's reused in docstrings. 5. data ([old](https://docs.pytorch.org/docs/main/data.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155198/data.html)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155198 Approved by: https://github.com/albanD, https://github.com/svekars	2025-06-13 20:24:34 +00:00
Ankita George	bf798a2f01	Change _hfstorage to hfstorage (#155837 ) Summary: Change HF classes to not have an underscore, there-by making them public, we will add documentation to them following this Test Plan: ensure existing tests pass Rollback Plan: Differential Revision: D76364024 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155837 Approved by: https://github.com/saumishr	2025-06-13 20:19:51 +00:00
zeshengzong	77f884c2ec	Optimize Tensor.backward type hints (#155656 ) Fixes #81963 ## Test Result ![image](https://github.com/user-attachments/assets/67380fdc-73c4-43d8-b2a5-5e16d63f4fd3) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155656 Approved by: https://github.com/soulitzer	2025-06-13 19:16:48 +00:00
PyTorch MergeBot	06408dae49	Revert "Add view_simple as meta function for view, and avoid calling reshape_view_helper. (#154757 )" This reverts commit 0029259bdfeee627181df2b9f5ff6979f65090ec. Reverted https://github.com/pytorch/pytorch/pull/154757 on behalf of https://github.com/laithsakka due to post land issue ([comment](https://github.com/pytorch/pytorch/pull/154757#issuecomment-2971385787))	2025-06-13 19:11:43 +00:00
Michael Lazos	4628f1b7a9	[Hierarchical-Compile] Track mutations for setitem (#155880 ) This fixes a bug in tensor variable where we would not do things like set the example value on setitem nodes (but these don't typically have users so it doesn't matter) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155880 Approved by: https://github.com/anijain2305	2025-06-13 18:59:31 +00:00
Ting Lu	344731fb25	Add CUDA 12.9.1 sbsa nightly binaries (#155819 ) https://github.com/pytorch/pytorch/issues/155196 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155819 Approved by: https://github.com/atalman	2025-06-13 18:52:41 +00:00
fduwjj	ce44877961	[c10d][PGNCCL] Make watchdog thread a class (#155831 ) By extracting both monitor thread and watchdog thread into a separate class this will help us learn what dependencies we have for each thread and it will kind of simplify the consolidation work for each thread (consolidating from thread per PG instance to per PG class) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155831 Approved by: https://github.com/d4l3k, https://github.com/kwen2501	2025-06-13 18:05:22 +00:00
Dhia-naouali	c5d00e150a	convert: rst to myst pr 1/2 (#155840 ) Fixes #155038 parent [PR](https://github.com/pytorch/pytorch/pull/155375) (made two PRs to pass sanity check) this PR converts the following two .rst files - [torch.compiler_dynamo_overview](https://github.com/pytorch/pytorch/blob/main/docs/source/torch.compiler_dynamo_overview.rst) - [torch.compiler_fake_tensor](https://github.com/pytorch/pytorch/blob/main/docs/source/torch.compiler_fake_tensor.rst) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155840 Approved by: https://github.com/sekyondaMeta	2025-06-13 18:02:28 +00:00
Nikita Shulga	36bf81e363	[BE] Fix minifier when one has multiple Python runtimes (#155918 ) By using `sys.executable` instead of `"python"` Otherwise, it fails on Ubuntu with `python not found` error Pull Request resolved: https://github.com/pytorch/pytorch/pull/155918 Approved by: https://github.com/seemethere, https://github.com/ZainRizvi, https://github.com/zou3519	2025-06-13 17:55:04 +00:00
Runtian (Rachel) Li	093aaccae2	convert `jit_language_reference_v2.rst` to `jit_language_reference_v2.md` (#155781 ) Fixes https://github.com/pytorch/pytorch/issues/155023 Description: converted `jit_language_reference_v2.rst` to `jit_language_reference_v2.md` I indented the code blocks to minimize the file difference to pass the sanity check for no more than 2000 lines of change. I will submit another PR to fix the indentation after this PR is merged. Checklist: - [x] The issue being fixed is referenced above (Fixes https://github.com/pytorch/pytorch/issues/155023) - [x] Only one issue is addressed in this pull request - [x] Labels from the issue that this PR is fixing are added to this pull request - [x] No unnecessary issues are included into this pull request. @pytorchbot label "topic: docs" @pytorchbot label "topic: not user facing" @pytorchbot label docathon-h1-2025 @pytorchbot label module: docs Pull Request resolved: https://github.com/pytorch/pytorch/pull/155781 Approved by: https://github.com/svekars	2025-06-13 17:33:10 +00:00
PyTorch UpdateBot	f0bee87eea	[xla hash update] update the pinned xla hash (#155779 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned xla hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155779 Approved by: https://github.com/pytorchbot	2025-06-13 17:13:37 +00:00
Aidyn-A	1f3cc4875c	[ATen][CUDA][cuSOLVER] Add cusolverDnXsyevBatched for torch.linalg.eigh (#155695 ) This PR add a new API for SYEV operation of cuSOLVER [`cusolverDnXsyevBatched`](https://docs.nvidia.com/cuda/cusolver/index.html#cusolverdnxsyevbatched) which is a new alternative to [`cusolverDn<t>syevjBatched`](https://docs.nvidia.com/cuda/cusolver/index.html#cusolverdn-t-syevjbatched). This API was introduced in cuSOLVER as part of 64-bit API in CUDA Tool Kit 12.6.2. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155695 Approved by: https://github.com/Skylion007, https://github.com/eqy	2025-06-13 17:12:26 +00:00
Nikita Shulga	b6add8c8ba	[MPSInductor] Fix remainder implementation for int types (#155891 ) Introduce `c10:🤘:remainder` and call it from both inductor and eager implementation, with integer specialization, which should make it much faster than before, while still compliant with Python way of rounding up negative numbers. This allows one to remove complex type detection logic from mps codegen and rely on Metal(C++) type system to figure out input and output types. This fixes compilation of something like ```python @torch.compile def f(x, y): return x[y % 5] ``` which beforehand failed to compile with ``` torch._inductor.exc.InductorError: SyntaxError: failed to compile #include <c10/metal/utils.h> kernel void generated_kernel( device float* out_ptr0, constant long* in_ptr0, constant float* in_ptr1, uint xindex [[thread_position_in_grid]] ) { int x0 = xindex; auto tmp0 = in_ptr0[x0]; auto tmp1 = 12; auto tmp2 = static_cast<float>(tmp0) - static_cast<float>(tmp1) * metal::floor(static_cast<float>(tmp0) / static_cast<float>(tmp1)); auto tmp3 = 1024; auto tmp4 = static_cast<long>(tmp3); auto tmp5 = tmp2 + tmp4; auto tmp6 = tmp2 < 0; auto tmp7 = tmp6 ? tmp5 : tmp2; if ((tmp7 < 0) && (tmp7 > 1024)) return; auto tmp9 = in_ptr1[tmp7]; out_ptr0[x0] = static_cast<float>(tmp9); } with program_source:372:28: error: array subscript is not an integer auto tmp9 = in_ptr1[tmp7]; ^~~~~ ``` This fixes fail_to_compile for GPT2ForSequenceClassification Huggingface model using `transformers==4.44.2` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155891 Approved by: https://github.com/manuelcandales	2025-06-13 16:42:56 +00:00
Georgia Phillips	9462106b7e	[nativert] Move graph_passes to nativert (#155411 ) Summary: Move graph_passes to nativert Test Plan: CI Rollback Plan: Differential Revision: D76205048 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155411 Approved by: https://github.com/zhxchen17	2025-06-13 16:41:01 +00:00
Chao Gu	338a8c7853	fix slice w/ dynamic shapes (#153131 ) Summary: guard_size_oblivious has side effects that'll result in invalid strides when slice nodes take negative index on dynamic input shapes. Cause overflow error with a huge number “9223372036854776048” Test Plan: CIs should pass. Differential Revision: D74354663 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153131 Approved by: https://github.com/laithsakka	2025-06-13 15:53:17 +00:00
Chang Pan	a5938ff431	[BE][c10d/Store]add check in pyi (#155855 ) (#155865 ) Summary: "check" is already binded https://fburl.com/code/9lx1zf9o which is also documented in https://docs.pytorch.org/docs/stable/distributed.html add it to pyi for type checking Test Plan: skip Rollback Plan: Differential Revision: D76547457 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155865 Approved by: https://github.com/fduwjj	2025-06-13 15:39:27 +00:00
Stephen Jia	bee93f9f0d	Move glslc to cas to enable remote execution (#155832 ) Meta: `fbsource//xplat/caffe2:gen_torch_vulkan_spv_cpp` takes on average 2 min to build and is one of topmost slow targets in fbandroid. See: https://fb.workplace.com/groups/2840058936242210/posts/4067730240141734 This target hat to run locally because it uses manifold backend for dotslash. This diff moves the `glslc` to cas backend so that it can run on RE. Here are commands executed: ``` % manifold get dotslash_glslc/flat/glslc-linux-x86_64.tar.gz % manifold get dotslash_glslc/flat/glslc-macos-v2024_4.tar.gz % manifold get dotslash_glslc/flat/glslc-windows-v2024_3.tar % ls -rw-r--r-- 1 navidq staff 2.0M Jun 12 10:02 glslc-linux-x86_64.tar.gz -rw-r--r-- 1 navidq staff 4.7M Jun 12 10:03 glslc-macos-v2024_4.tar.gz -rw-r--r-- 1 navidq staff 4.4M Jun 12 10:03 glslc-windows-v2024_3.tar % frecli --use-case dotslash cas upload-blob --skip-find-missing glslc-linux-x86_64.tar.gz ea5d674e0e7e9782be3f5c309e3484732e5b3a331cbe3258f3e929002811627b:2072937 % frecli --use-case dotslash cas upload-blob --skip-find-missing glslc-macos-v2024_4.tar.gz 1331dc691835e4676832b7c21ef669083a3acc8856981583d0698192f466c51a:4898649 % frecli --use-case dotslash cas upload-blob --skip-find-missing glslc-windows-v2024_3.tar 76181fbb1ce5c62d0c905db26df3a64e999d0baff2e93270775921daa91e3a1a:4585984 ``` Differential Revision: [D76513735](https://our.internmc.facebook.com/intern/diff/D76513735/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155832 Approved by: https://github.com/GregoryComer	2025-06-13 14:38:51 +00:00
PyTorch MergeBot	ce6e0523f9	Revert "[BE] Raise `NotImplementedError` (#155470 )" This reverts commit 5ab6a3fb6fd37c542060c606edd4b95c7e3cae82. Reverted https://github.com/pytorch/pytorch/pull/155470 on behalf of https://github.com/malfet due to foreach tests are failing on ROCm because we are not running the same on CUDA ([comment](https://github.com/pytorch/pytorch/pull/155470#issuecomment-2970592124))	2025-06-13 14:32:50 +00:00
James Wu	3819584f12	[precompile] Implement PrecompileContext for recording precompile artifacts, integrate with CompilePackage (#154415 ) This PR implements a basic interface and test for PrecompileContext, a special CacheArtifactManager specifically designed for precompile. The job of a PrecompileContext is to record things precompile needs as torch is compiling, dump it all into bytes, and then stitch it back together into a cache of callables. ## Why use CacheArtifactManager? Precompile needs a way to record various serializable data as torch is compiling. CacheArtifactManager already does this today pretty well, handling a lot of serialization and cache information. So we're reusing a bunch of that infrastructure directly. ## How is it different from CacheArtifactManager? Unlike regular CacheArtifactManager, PrecompileContext needs to be able to take the recorded artifacts and stitch them together after deserialization, to create a single working callable. Since PrecompileContext doesn't need the cache keys, the "key" field of PrecompileArtifacts can be used for metadata relating to how to stitch the individual functions being compiled together into a full callable. For example, on a given dynamo compile, if there are multiple functions (via graph breaks or recompiles) being compiled, MegaCache would organize it like so: ![image](https://github.com/user-attachments/assets/49a0a75b-1e7f-4d96-8d81-6769fe5a53ca) Whereas we'd visualize PrecompileContext's result like so: ![image](https://github.com/user-attachments/assets/fcc0dd4e-dfbf-4b13-9c08-2e99b373180b) For now, we just handle eager mode; in the diff above, I'll hook up the other backend artifacts from PrecompileContext. After this PR, precompile consists of three main interfaces: ### CompilePackage - Everything needed to run one torch.compile'd function (including graph breaks) - `__init__(fn, cache_entry)` Initializes with a DynamoCacheEntry - `install(backends)` load precompile artifacts into function's dynamo state with a dictionary of backends - `cache_entry()` return a serializable cache entry to save ### DynamoStore - Responsible for tracking CompilePackages on disk (and/or in memory) - `load_package(path)`: load a package given a torch compiled function and a path to the cache artifact - `save_package(package, path): Save a CompiledPackage to a path. Calls PrecompileContext to grab backend data - `record_package(package)`: Record a package to PrecompileContext (for global serialization/deserialization) ### PrecompileContext - Overarching context for serializing and deserializing precompile artifacts. Supports global and local setups. - `serialize()`: (Global) serializes all artifacts in PrecompileContext into bytes - `populate_caches(bytes)`: (Global) takes serialized bytes and puts them into DynamoStore (TODO) - `serialize_artifact_by_key(key)`: (Local) serialize a single artifact by its cache key <img width="1455" alt="image" src="https://github.com/user-attachments/assets/99b61330-7607-4763-bdbc-85b366e82cdd" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/154415 Approved by: https://github.com/zhxchen17 ghstack dependencies: #155118	2025-06-13 14:11:24 +00:00
James Wu	b2fc9cfea1	[precompile] Add CompilePackage to serialize dynamo states. (#155118 ) Adding a per torch.compile() object CompilePackage which tracks dynamo artifact. CompilePackage is considered a low level component and should not be directly exposed to end users. It has the following interface: 1. `CompilePackage.__init__()` which optionally takes previously serialized dynamo states. a. when `dynamo` argument is None, it will contruct a brand new CompilePackage object. b. when `dynamo` argument is not None, it will load a pre-compiled dynamo state. 2. `package.save()` which dumps the dynamo states into _DynamoCacheEntry. 3. `package.install(backends)` which will handle all the side-effectful global scope updates with compiled functions and resume functions. This diff focus on making the low level mechanism for precompile. It will be left to upper level interface to use these API to build more user-facing frontend. Differential Revision: [D75956538](https://our.internmc.facebook.com/intern/diff/D75956538/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155118 Approved by: https://github.com/jamesjwu Co-authored-by: James Wu <jjwu@meta.com>	2025-06-13 13:54:10 +00:00
xinan.lin	670dab6c63	[AOTI] Enable OP `test__weight_int4pack_mm_with_scales_and_zeros` in AOTI. (#155780 ) The op test__weight_int4pack_mm_with_scales_and_zeros is for Intel GPU. It is functionally equivalent to the CUDA/CPU op test__weight_int4pack_mm (with the constraint that oneDNN only supports integer zero points, which is why we need this API). Since test__weight_int4pack_mm is already included in AOTI's fallback list, this PR adds support for XPU. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155780 Approved by: https://github.com/jansel	2025-06-13 11:12:13 +00:00
Avik Chaudhuri	463fe36532	fix error message on specialization with Dim.DYNAMIC (#155738 ) Previously specialization error messages would render sources that were pretty far from source-code names. E.g., given args named `x, y, zs`, the source for `y.size()[0]` would be rendered as `args[0][1].size()[0]`. This is because we created artificial local names following `(args, kwargs)` structure instead of reusing signatures. This PR fixes that situation. Basically we map prefixes of key paths that correspond to original arg names to root sources corresponding to those names; the rest of the key paths hang from these root sources. Differential Revision: [D76461391](https://our.internmc.facebook.com/intern/diff/D76461391/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155738 Approved by: https://github.com/bobrenjc93	2025-06-13 10:33:46 +00:00
anwang	6abe450a6f	[pytorch Aten] Delete unused duplicate clamp_stub, to avoid compile error (#154631 ) I found the `clamp_stub` in `UnaryOps.h` is not used. And it's a duplicate of the `clamp_stub` in `TensorCompare.cpp`: https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/native/TensorCompare.cpp#L313-L314 This diff/PR deletes it as this duplicate caused build failure for me: ``` ATen/native/UnaryOps.h:109:1: error: redefinition of 'clamp_stub_DECLARE_DISPATCH_type' ``` Differential Revision: [D75612521](https://our.internmc.facebook.com/intern/diff/D75612521/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154631 Approved by: https://github.com/Skylion007, https://github.com/cyyever, https://github.com/nautsimon ghstack dependencies: #154589, #154591	2025-06-13 10:01:51 +00:00
anwang	1cc31b213d	[MTIA Aten Backend] Migrate bitwise_and.Tensor_out (#154591 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff - Migrate where.self and where.self_out - Add tests for dtype casting and shape broadcasting Differential Revision: [D75578498](https://our.internmc.facebook.com/intern/diff/D75578498/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154591 Approved by: https://github.com/malfet ghstack dependencies: #154589	2025-06-13 10:01:51 +00:00
fengqing.lu@intel.com	65b9c13cce	[Intel GPU] Enable safe softmax for XPU SDPA (#151999 ) Fix https://github.com/intel/torch-xpu-ops/issues/1432#event-16899653975 When one row of Q*K attention score is masked with `-inf`, `softmax(score)` would output `NaN` for whole row which would cause model corruption. With this new flag, it would output `0` for whole row which is aligned with Pytorch CPU/CUDA's behavior. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151999 Approved by: https://github.com/guangyey, https://github.com/EikanWang, https://github.com/drisspg Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-06-13 08:53:47 +00:00
anwang	56b03df6ac	[MTIA Aten Backend] Migrate where.self and where.self_out (#154589 ) # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This diff - Migrate where.self and where.self_out - Add tests for dtype casting and shape broadcasting Differential Revision: [D75577304](https://our.internmc.facebook.com/intern/diff/D75577304/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154589 Approved by: https://github.com/malfet	2025-06-13 08:25:13 +00:00
ZhaoqiongZ	3d595fd559	update get start xpu (#151886 ) update link and product name add print to print ```torch.xpu.is_available()``` result in code snippet for user not using command python Pull Request resolved: https://github.com/pytorch/pytorch/pull/151886 Approved by: https://github.com/guangyey, https://github.com/AlannaBurke Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-06-13 07:46:13 +00:00
Isuru Fernando	53d06e18d9	[dynamo] add missing algorithm header (#154754 ) Needed for `std::max(<initializer-list>)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154754 Approved by: https://github.com/Skylion007, https://github.com/anijain2305	2025-06-13 06:56:11 +00:00
Bob Ren	6020440683	remove allow-untyped-defs from adaround_fake_quantize.py (#155621 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155621 Approved by: https://github.com/Skylion007	2025-06-13 06:14:22 +00:00
Ke Wen	99e99d5bfe	[a2av] Test must allocate tensors symmetrically (#155835 ) This is a requirement of most SHMEM backends. Otherwise, allocations may misalign across ranks. In this PR, we make the (total) input size and output size a constant number, even though the split sizes are created random. (Previously we sum the splits up as input size, which creates misalignment in SHMEM heap across ranks). Pull Request resolved: https://github.com/pytorch/pytorch/pull/155835 Approved by: https://github.com/fduwjj, https://github.com/fegin, https://github.com/Skylion007 ghstack dependencies: #155506	2025-06-13 06:05:38 +00:00
angelayi	0860606729	[export] Add meta[val] to getattr nodes (#154934 ) Fixes [P1830293318](https://www.internalfb.com/intern/paste/P1830293318/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154934 Approved by: https://github.com/yushangdi, https://github.com/muchulee8	2025-06-13 05:48:21 +00:00
Nikita Shulga	25717da8c8	[BE] Don't run the same tests on 2xlarge and 4xlarge (#155859 ) Also, speedup builds by moving them to 4xlarge instances Pull Request resolved: https://github.com/pytorch/pytorch/pull/155859 Approved by: https://github.com/ZainRizvi, https://github.com/wdvr	2025-06-13 05:40:20 +00:00
fduwjj	a87dfc7480	[symm_mem] Update CMakeList to reflect code moving a dedicated folder (#155823 ) We moved all symm_mem code into a folder ([CudaDMAConnectivity](https://github.com/pytorch/pytorch/pull/155573)) but somehow forgot update for CudaDMAConnectivity in the CMakeList. Users see errors: RuntimeError: DMA connectivity detector for cuda over nvlink is not available while torch.distributed.init_process_group(backend=backend). So this PR should fix it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155823 Approved by: https://github.com/Skylion007	2025-06-13 05:27:59 +00:00
Justin Silver	70bb34929a	Convert to .md: draft_export.rst, export.ir_spec.rst, fft.rst (#155567 ) Used [rst2myst tool](https://rst-to-myst.readthedocs.io/en/latest/) Fixes #155020. This PR is split into 3 to pass sanity check. Docs comparison (check out the 'new' whenever docs build) 1. draft_export ([old](https://docs.pytorch.org/docs/main/draft_export.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155567/draft_export.html)) 2. export.ir_spec ([old](https://docs.pytorch.org/docs/main/export.ir_spec.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155567/export.ir_spec.html)) 3. fft ([old](https://docs.pytorch.org/docs/main/fft.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155567/fft.html)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155567 Approved by: https://github.com/svekars	2025-06-13 05:19:43 +00:00
Henry Tsang	b878ca0c91	[cutlass backend] add fp8 to cutlass benchmark script (#155507 ) Summary: Add fp8. Right now FP8 only allows fast_accum. Test Plan: ``` Experiment group: _scaled_mm (8192x8192, 8192x8192) torch.float8_e4m3fn +-----------------------+--------------------+--------------------+----------------------+--------------------+ \| name \| forward_time (us) \| teraflops (TFLOPS) \| compilation_time (s) \| perf_over_aten (%) \| +-----------------------+--------------------+--------------------+----------------------+--------------------+ \| aten \| 967.1226739883423 \| 1136.8895149998868 \| 1.219131228979677 \| NA \| \| triton \| 1764.6185159683228 \| 623.08743664783 \| 20.373826419003308 \| 82.46067054670186 \| \| triton_persistent_tma \| 1769.0335512161255 \| 621.5323768280928 \| 20.48663099599071 \| 82.91718297956578 \| \| cutlass_lvl_default \| 790.5075550079346 \| 1390.8932568835019 \| 13.788519630907103 \| -18.26191482535096 \| \| cutlass_lvl_3332 \| 803.7384748458862 \| 1367.996757884245 \| 226.81587297911756 \| -16.89384434227684 \| +-----------------------+--------------------+--------------------+----------------------+--------------------+ ``` Rollback Plan: Differential Revision: D76310809 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155507 Approved by: https://github.com/ColinPeppler	2025-06-13 05:11:15 +00:00
GdoongMathew	2ba930d4ce	Convert rst to markdown - profiler.rst #155031 (#155559 ) Fixes https://github.com/pytorch/pytorch/issues/155031 * [profiler.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/profiler.rst) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155559 Approved by: https://github.com/svekars	2025-06-13 05:02:54 +00:00
Runtian (Rachel) Li	e8b3dfa7c0	convert jit_language_reference.rst to jit_language_reference.md (#155633 ) Part of changes https://github.com/pytorch/pytorch/issues/155023 (parent PR https://github.com/pytorch/pytorch/pull/155429) - converted jit_language_reference.rst to jit_language_reference.md @pytorchbot label "topic: docs" @pytorchbot label "topic: not user facing" @pytorchbot label docathon-h1-2025 @pytorchbot label module: docs Pull Request resolved: https://github.com/pytorch/pytorch/pull/155633 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-13 04:58:28 +00:00
Runtian (Rachel) Li	3f65e38b73	Convert hub.rst to hub.md (#155483 ) Part of changes https://github.com/pytorch/pytorch/issues/155023 (parent PR https://github.com/pytorch/pytorch/pull/155429) @pytorchbot label "topic: docs" @pytorchbot label "topic: not user facing" @pytorchbot label docathon-h1-2025 @pytorchbot label module: docs Pull Request resolved: https://github.com/pytorch/pytorch/pull/155483 Approved by: https://github.com/svekars	2025-06-13 04:39:55 +00:00
Will Constable	0a6b66c881	Inductor comms reorder logs to tlparse (#155737 ) Hacked test_inductor_collectives test to demonstrate this works: https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/whc/de50ff33-f460-406b-bfa9-457e6e17395b/custom/-_0_0_0/reorder_communication_preserving_peak_memory_9.txt?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 Follow up: it would be nice to move the logging out of this pass and into the broader comms pass loop, where the before/after each pass visualization could be logged into the same tlparse file. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155737 Approved by: https://github.com/bdhirsh	2025-06-13 02:59:42 +00:00
Bin Bao	f151b20123	[AOTI] Remove the emit_current_arch_binary option (#155768 ) Summary: Remove the option as generating fatbin with PTX only doesn't work on H100, so switch to always include one PTX and one SASS for fatbin. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155768 Approved by: https://github.com/angelayi	2025-06-13 02:06:07 +00:00
FFFrog	020da74437	[Easy] Remove empty file (#155796 ) As the title stated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155796 Approved by: https://github.com/malfet ghstack dependencies: #155772	2025-06-13 01:42:11 +00:00
zeshengzong	905b194a2e	Replace device check of TORCH_INTERNAL_ASSERT with TORCH_CHECK (#155318 ) Fixes #136849 ## Test Result ```python >>> import torch >>> device = torch.cuda.device_count() + 1 >>> torch.cuda.current_stream(device) # INTERNAL ASSERT FAILED Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/home/zong/code/pytorch/torch/cuda/__init__.py", line 1083, in current_stream streamdata = torch._C._cuda_getCurrentStream( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ RuntimeError: Device index value 3 is out of index range [0, 2) >>> torch.cuda.default_stream(device) # INTERNAL ASSERT FAILED Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/home/zong/code/pytorch/torch/cuda/__init__.py", line 1101, in default_stream streamdata = torch._C._cuda_getDefaultStream( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ RuntimeError: Device index value 3 is out of index range [0, 2) >>> torch.cuda.set_per_process_memory_fraction(0.5, device) # INTERNAL ASSERT FAILED Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/home/zong/code/pytorch/torch/cuda/memory.py", line 193, in set_per_process_memory_fraction torch._C._cuda_setMemoryFraction(fraction, device) RuntimeError: Allocator not initialized for device : did you call init? ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155318 Approved by: https://github.com/albanD	2025-06-13 01:20:19 +00:00
Laith Sakka	d7e657da35	pyfmt lint more torch/utils files (#155812 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155812 Approved by: https://github.com/Skylion007 ghstack dependencies: #155782, #155783	2025-06-12 23:51:42 +00:00
angelayi	4d3ecefda5	[aoti][mps] Use cpp sym-expr printer (#155646 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155646 Approved by: https://github.com/desertfire ghstack dependencies: #155752, #154287, #155582, #155583	2025-06-12 23:33:28 +00:00
angelayi	2e65d72e1e	[aoti][mps] Fix int/symint kernel args (#155583 ) Integer arguments to mps kernels need to go through a different function, since `aoti_torch_mps_set_arg` only takes a Tensor. So I added a `aoti_torch_mps_set_arg_int`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155583 Approved by: https://github.com/desertfire ghstack dependencies: #155752, #154287, #155582	2025-06-12 23:33:28 +00:00
angelayi	ffbda61fbe	[aoti][mps] Fix dynamic dispatch size (#155582 ) In the case where we pass in a symint to the `dispatch` call, the compiler errors, so we need to cast the input to int64_t. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155582 Approved by: https://github.com/malfet ghstack dependencies: #155752, #154287	2025-06-12 23:33:15 +00:00
angelayi	a4ab392251	[aoti][mps] mps constants support (#154287 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154287 Approved by: https://github.com/malfet ghstack dependencies: #155752	2025-06-12 23:33:07 +00:00
angelayi	8821a9dc4e	[BE][aoti][mps] Fix tests to use common function (#155752 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155752 Approved by: https://github.com/desertfire, https://github.com/malfet	2025-06-12 23:32:59 +00:00
Nikita Shulga	5ab6a3fb6f	[BE] Raise `NotImplementedError` (#155470 ) When op is unimplemented for a specific dtype Which makes more sense, than a RuntimeError Example ```python >>> import torch >>> torch.nn.Hardshrink()(torch.randint(0, 5, (10,))) NotImplementedError: "hardshrink_cpu" not implemented for 'Long' ``` release notes bc-breaking: After this release `NotImplementedError` exception will be raised when ATen operation is called on the combinaiton of input tensor dtypes it has not been implemented for Pull Request resolved: https://github.com/pytorch/pytorch/pull/155470 Approved by: https://github.com/albanD, https://github.com/Skylion007	2025-06-12 23:19:12 +00:00
Natalia Gimelshein	d9b8369f39	fix warning spam for list indexing (#155815 ) Per title, #154806 incorrectly placed a warning Pull Request resolved: https://github.com/pytorch/pytorch/pull/155815 Approved by: https://github.com/Skylion007, https://github.com/albanD	2025-06-12 23:07:24 +00:00
Laith Sakka	2903e5ad3c	pyfmt lint more export files (#155783 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155783 Approved by: https://github.com/Skylion007 ghstack dependencies: #155782	2025-06-12 23:04:11 +00:00
Laith Sakka	86b1116f22	pyfmt lint torch/_custom_op/* (#155782 ) file torch/_custom_op/functional.py does not exisits file torch/_custom_op/__init__.py is empty. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155782 Approved by: https://github.com/Skylion007	2025-06-12 23:04:11 +00:00
Jithun Nair	4cdbdcdbcf	Switch to miniconda for ROCm CI (#155239 ) Related to https://github.com/pytorch/pytorch/issues/148335 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155239 Approved by: https://github.com/jeffdaily	2025-06-12 22:55:47 +00:00
Randolf Scholz	f04fd4dc4e	typing: allow integer in bitwise operations (#155704 ) Fixes #155701 (false positives) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155704 Approved by: https://github.com/Skylion007, https://github.com/aorenste	2025-06-12 22:40:17 +00:00
angelayi	938515fa75	[aoti] Update cshim for all backends (#155604 ) Fixes https://github.com/pytorch/pytorch/issues/155349 `python torchgen/gen.py --update-aoti-c-shim` will now update all cpu/cuda/mps/xpu shims -- I verified this using `aten._print.default`, but didn't commit the changes since I'm not sure if we actually want to add this. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155604 Approved by: https://github.com/desertfire, https://github.com/janeyx99	2025-06-12 22:10:58 +00:00
Mikayla Gawarecki	38bfd462b8	Use swap_tensors path in nn.Module.to for FakeTensor (#152539 ) Fixes https://github.com/pytorch/pytorch/issues/148977 Differential Revision: [D76458023](https://our.internmc.facebook.com/intern/diff/D76458023) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152539 Approved by: https://github.com/albanD	2025-06-12 22:08:21 +00:00
Frost Mitchell	db01f1032f	Support XPU in memory tracker (#150703 ) This PR adds support for XPU devices to the distributed MemoryTracker tool, including unit test for XPU. Specifically, this code adds tracking for a few alloc-related statistics for XPUCachingAllocator. It also adapts the existing memory tracker tool to be device agnostic, by getting the device module and recording the necessary memory stats. (I get the device module instead of using `torch.accelerator` methods, as that API is still in-progress.) Pull Request resolved: https://github.com/pytorch/pytorch/pull/150703 Approved by: https://github.com/EikanWang, https://github.com/guangyey, https://github.com/gujinghui, https://github.com/d4l3k	2025-06-12 21:33:52 +00:00
Brian Hirsh	154a39bfbd	basic compile support for grouped_mm (#153384 ) grouped_mm is used in torchtitan, this adds just enough support in compile to allow inductor to lower it as a fallback kernel. I imagine that at some point in the future it may be valuable to get inductor to support templating grouped_mm, although this PR just provides basic support. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @ipiszy @chenyang78 @kadeng @muchulee8 @amjames @chauhang @aakhundov @ngimel @eellison Pull Request resolved: https://github.com/pytorch/pytorch/pull/153384 Approved by: https://github.com/eellison	2025-06-12 21:24:51 +00:00
Prachi Gupta	f2b44424a1	[ROCm] Skip _stress_cuda and test_ddp_apply_optim_in_backward (#155724 ) These tests are flaky on ROCm and have been skipped via Github issues, but the bot keeps closing the issues after not observing the failures for these tests in the rerun_disabled_tests runs (not sure why they don't fail there), and we have to keep reopening them. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155724 Approved by: https://github.com/jeffdaily Co-authored-by: Jithun Nair <37884920+jithunnair-amd@users.noreply.github.com>	2025-06-12 21:18:04 +00:00
lzhang2	590fe4d2d7	Skip updating the default device distributed backend if already registered (#155320 ) Motivation: PyTorch maintain a `default_device_backend_map` https://github.com/pytorch/pytorch/blob/main/torch/distributed/distributed_c10d.py#L269 , which indicates the default distributed backend if no backend name is specified in user frontend (like `init_process_group`). Currently, `"xpu": XCCL` is also in this `default_device_backend_map`. However, if another process group name is registered as XPU distributed backend, it immediately replaces XCCL in this default map, which is not what we want. Therefore, we would like to skip updating the default distributed backend if one is already registered in the map. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155320 Approved by: https://github.com/guangyey, https://github.com/d4l3k	2025-06-12 21:17:06 +00:00
Catherine Lee	29391c7cf9	[ez] Mark linalg svd memory allocation test as serial b/c OOMing on cu128 (#155811 ) `9df2e8020f/1` `8e8d4b13b0 (43980565863-box)` started OOMing after switching to cuda 12.8 Maybe b/c I made some changes fix the per process memory fraction so each proc has fewer memory ``` 2025-06-12T15:29:50.4998758Z FAILED [0.0124s] test_linalg.py::TestLinalgCUDA::test_svd_memory_allocation_cuda_complex128 - torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.10 GiB. GPU 0 has a total capacity of 7.43 GiB of which 6.85 GiB is free. Process 80272 has 68.75 MiB memory in use. Process 83346 has 68.75 MiB memory in use. Process 83365 has 374.75 MiB memory in use. Process 83384 has 70.75 MiB memory in use. 2.90 GiB allowed; Of the allocated memory 240.00 MiB is allocated by PyTorch, and 2.00 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155811 Approved by: https://github.com/huydhn, https://github.com/malfet, https://github.com/atalman, https://github.com/eqy	2025-06-12 21:05:32 +00:00
Parag Ekbote	093fd47dbe	Add a Additional Example that showcases the usage of `torch.autograd.functional.jacobian` (#155683 ) Fixes #132140 As described in the issue, I've added an example that showcases the use of higher-dimensional inputs and outputs, batched inputs, and the vectorize=True with `torch.autograd.functional.jacobian`. Could you please review? Pull Request resolved: https://github.com/pytorch/pytorch/pull/155683 Approved by: https://github.com/soulitzer	2025-06-12 19:46:55 +00:00
Ankita George	e6d71f3789	Support re-sharding for safetensors checkpoints (#154519 ) This change will add the ability to support re-sharding for hf safetensors checkpoints. This is done by adding more metadata when saving each file. This metadata captures the size and offset of the saved shard. This can be used to re-shard on load by using this information to create the chunks belonging to TensorStorageMetadata class. Differential Revision: [D75226344](https://our.internmc.facebook.com/intern/diff/D75226344/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154519 Approved by: https://github.com/saumishr	2025-06-12 19:38:29 +00:00
Amandeep Chhabra	d3da03d6fa	[2/n]passing event log handler to record function calls (#155457 ) Summary: This diff modifies the elastic agent's API to pass the event log handler to the record function calls. This change enables the elastic agent to log events to a specific destination, improving the monitoring and debugging capabilities of the distributed training process. Test Plan: unit tests ran an e2e training job. Differential Revision: D75194115 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155457 Approved by: https://github.com/d4l3k	2025-06-12 19:35:08 +00:00
Justin Silver	e085012335	Fix #155020 - rst2markdown for export.rst (split PR) (#155753 ) Used [rst2myst tool](https://rst-to-myst.readthedocs.io/en/latest/) Fixes #155020. This PR is split into 3 to pass sanity check. This is the 3rd one. Docs comparison (check out the 'new' whenever docs build) 1. export ([old](https://docs.pytorch.org/docs/main/export.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155567/export.html)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155753 Approved by: https://github.com/sekyondaMeta	2025-06-12 19:30:52 +00:00
Laith Sakka	4bb936d8b7	refresh expected results (#155817 ) some changes landed when the test is recently unstable with out updating the results. <img width="564" alt="Screenshot 2025-06-12 at 9 26 32 AM" src="https://github.com/user-attachments/assets/9a83f18b-f2a8-485d-a58e-67d8c161eb18" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/155817 Approved by: https://github.com/yushangdi	2025-06-12 19:14:21 +00:00
Grant Smith	7986c0dba6	rename distributed.rst to md (#155767 ) Fixes #155019 For sanity checks, split PR to have this one only include distributed.rst -> distributed.md Preview -> [distributed.md](https://docs-preview.pytorch.org/pytorch/pytorch/155767/distributed.html) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155767 Approved by: https://github.com/sekyondaMeta	2025-06-12 18:42:15 +00:00
Nikita Shulga	bcad962550	[BE][Testing] Delete some unused code (#155760 ) - Fix typo in class name `OpenRgistration`->`OpenRegistration` - Use existing `common` alias of `torch.testing._internal.common_utils`, i.e. `s/torch.testing._internal.common_utils.markDynamoStrictTest/common.markDynamoStrictTest/` - Remove unused `TEST_CUDA` and `TEST_ROCM` are unused in that file Pull Request resolved: https://github.com/pytorch/pytorch/pull/155760 Approved by: https://github.com/albanD	2025-06-12 18:41:53 +00:00
Yidi Wu	fac0cc16ef	[scan] fix doc of scan and list the restrctions. (#155577 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155577 Approved by: https://github.com/zou3519	2025-06-12 18:22:28 +00:00
Mu-Chu Lee	a1257446f8	[AOTInductor] Memory leak fix for Fallback Kernels (#155642 ) Summary: We generate AtenTensorHandles for Fallback kernels regardless of the arg type. If we indeed "fallback", we will regenerate the AtenTensorHandles that will cause the first handle being generated not recycled, thus a memory leak would occur. Test Plan: python test/inductor/test_aot_inductor.py -k test_fallback_mem_leak Reviewers: Subscribers: Tasks: Tags: Pull Request resolved: https://github.com/pytorch/pytorch/pull/155642 Approved by: https://github.com/jingsh, https://github.com/desertfire	2025-06-12 17:42:56 +00:00
atalman	0d3d84d866	[CD] Windows Magma build 12.9 and cuda scripts (#155799 ) Scripts needed to build Magma and CUDA on windows Same as https://github.com/pytorch/pytorch/pull/146653 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155799 Approved by: https://github.com/jeanschmidt	2025-06-12 17:41:24 +00:00
Chris Sidebottom	430cc1c636	Run tests on Amazon EC2 M8g Instances (#153940 ) Requires machines configured here: https://github.com/pytorch/test-infra/pull/6642 This adds additional test runs against AWS Graviton4 processors, alongside existing runs against AWS Graviton3 and AWS Graviton2 processors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153940 Approved by: https://github.com/fadara01, https://github.com/malfet	2025-06-12 17:33:08 +00:00
Shangdi Yu	522a18bd6c	Fix provenance unit test (#155747 ) Summary: Fix the test to adapt added provenance tracking in D75837494 Test Plan: ``` buck2 run @//mode/dev-nosan fbcode//caffe2/test:fx -- -r test_graph_provenance ``` Rollback Plan: Differential Revision: D76466778 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155747 Approved by: https://github.com/YUNQIUGUO	2025-06-12 17:26:43 +00:00
zpcore	50d8168c8b	[DTensor] Support in gradient placement for local_map() (#155181 ) Support `in_grad_placements` argument in torch.distributed.tensor.experimental.local_map(). The argument helps enforce placement of gradient of the input Dtensor. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155181 Approved by: https://github.com/wanchaol	2025-06-12 17:07:04 +00:00
henrylhtsang	6c0b42fd2f	[inductor][cutlass backend] Log prescreening elpase (#155508 ) Differential Revision: [D76311352](https://our.internmc.facebook.com/intern/diff/D76311352/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155508 Approved by: https://github.com/jingsh	2025-06-12 16:48:52 +00:00
Simon Mahns	c1ae768baa	Basic MTIA ATen CMake (#155477 ) Summary: Basic ATen CMake Differential Revision: D75203592 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155477 Approved by: https://github.com/andyanwang, https://github.com/cyyever	2025-06-12 16:29:32 +00:00
Laith Sakka	f4376cac54	unify symbolic_shapes and sizevars dynamic shapes APIs naming 1 (#154774 ) Inductor have a set of APIs that allows performing symbolic evaluations similar to that of symbolic shapes but it operates on sympy expressions instead of symnodes. Namings are not consistent making them consistent in this stack. Step 1 : unify statically_know_true naming! for consistent experience. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154774 Approved by: https://github.com/drisspg, https://github.com/bobrenjc93, https://github.com/eellison	2025-06-12 16:11:55 +00:00
Runtian (Rachel) Li	9df2e8020f	fix code indentation for fx.md (#155764 ) Fixes https://github.com/pytorch/pytorch/issues/155023 Related PR: #155482 Description: As discussed here https://github.com/pytorch/pytorch/pull/155482#pullrequestreview-2918032289, I removed indentation for python code blocks as a follow-up modification for fx.md Checklist: - [x] The issue being fixed is referenced above (Fixes https://github.com/pytorch/pytorch/issues/155023) - [x] Only one issue is addressed in this pull request - [x] Labels from the issue that this PR is fixing are added to this pull request - [x] No unnecessary issues are included into this pull request. @pytorchbot label "topic: docs" @pytorchbot label "topic: not user facing" @pytorchbot label docathon-h1-2025 @pytorchbot label module: docs Pull Request resolved: https://github.com/pytorch/pytorch/pull/155764 Approved by: https://github.com/svekars	2025-06-12 16:02:33 +00:00
Pian Pawakapan	75824035d3	[dynamic shapes] skip fused linear path if not definitely contiguous (#155051 ) Falls back to non-fused linear -> add bias path for non-contiguous tensors with unbacked sizes Pull Request resolved: https://github.com/pytorch/pytorch/pull/155051 Approved by: https://github.com/laithsakka	2025-06-12 15:55:21 +00:00
Catherine Lee	51560797ce	[CI] Reuse old whl: switch default to always (#155572 ) Switch default to always reuse old whl I have a few worries about API rate limits Pull Request resolved: https://github.com/pytorch/pytorch/pull/155572 Approved by: https://github.com/huydhn, https://github.com/malfet, https://github.com/seemethere, https://github.com/atalman	2025-06-12 15:43:29 +00:00
Aleksandar Samardžić	62fa3f5aeb	Support tuning of _grouped_mm (#153953 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153953 Approved by: https://github.com/ngimel	2025-06-12 15:39:35 +00:00
Henry Tsang	6b3eef6d31	[cutlass backend] Only consider to use re worker if nvcc doesn't exist (#155745 ) Differential Revision: D76463340 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155745 Approved by: https://github.com/masnesral	2025-06-12 15:23:52 +00:00
Manuel Candales	851a6fa82d	[MPS] Migrate softshrink (forward and backward) to Metal kernel (#155586 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155586 Approved by: https://github.com/malfet ghstack dependencies: #155304, #155316, #155462, #155479, #155571	2025-06-12 15:02:43 +00:00
PyTorch MergeBot	2a3b41cbd0	Revert "[CI] Use `setup-python` from for Mac tests (#155698 )" This reverts commit 2b9d638e3333e6e9ae324e1486774e83292e1883. Reverted https://github.com/pytorch/pytorch/pull/155698 on behalf of https://github.com/malfet due to It causes weird flaky failures in MPS and do not upload usage logs anymore ([comment](https://github.com/pytorch/pytorch/pull/155698#issuecomment-2967120676))	2025-06-12 14:42:32 +00:00
Zhengxu Chen	0fd711df19	[export] Allow user frame to be None when symbolic shape tries to get stacktrace. (#155744 ) Summary: Fixing https://github.com/pytorch/pytorch/issues/155605 Test Plan: CI Rollback Plan: Differential Revision: D76463358 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155744 Approved by: https://github.com/angelayi	2025-06-12 14:36:29 +00:00
Bin Bao	dd1b6621bc	Remove C10_DEPRECATED references in c10 (#151058 ) Summary: Revive https://github.com/pytorch/pytorch/pull/138406. Only limit the scope to files in c10. Summary from the original PR, ``` Looking in the code I see // NB: __cplusplus doesn't work for MSVC, so for now MSVC always uses // the "__declspec(deprecated)" implementation and not the C++14 // "[[deprecated]]" attribute. We tried enabling "[[deprecated]]" for C++14 on // MSVC, but ran into issues with some older MSVC versions. But looking at the MSVC C++ support table I see that the [[deprecated]] attribute is supported as of MSVC 2015 and that the vast majority of C++17 features became supported in MSVC 2015 or later. Since PyTorch is C++17 now, I infer that PyTorch must not support versions of MSVC earlier than MSVC 2015, so the versions of MSVC supported by PyTorch must support [[deprecated]]. Therefore, since we are finished deprecating old MSVCs we can deprecate C10_DEPRECATED. ``` Test Plan: CI Differential Revision: D72762767 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151058 Approved by: https://github.com/r-barnes	2025-06-12 13:38:03 +00:00
FFFrog	d632cf2cc9	[Easy][Code Clean] Remove the unused and undefined function in pickler (#155772 ) As the title stated. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155772 Approved by: https://github.com/malfet	2025-06-12 13:03:36 +00:00
Wang, Chuanqi	8e8d4b13b0	[XPU] Simplify XPU `make triton` by install from PyTorch source (#155675 ) Remove install from source code build Pull Request resolved: https://github.com/pytorch/pytorch/pull/155675 Approved by: https://github.com/atalman	2025-06-12 13:02:23 +00:00
David Berard	132babe7e0	[user triton] dynamo support for new host-side TMA API (#155662 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155662 Approved by: https://github.com/aakhundov ghstack dependencies: #155510	2025-06-12 12:56:23 +00:00
Aaron Gokaslan	9cced33c7c	[BE]: Update cudnn to 9.10.2.21 (#155576 ) Update to CUDNN 9.10.2.21 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155576 Approved by: https://github.com/eqy, https://github.com/atalman	2025-06-12 12:50:36 +00:00
atalman	c199a4d0fd	Move non inductor workflows cuda 12.6->cuda 12.8 (#155234 ) Move non inductor workflows cuda 12.6->cuda 12.8 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155234 Approved by: https://github.com/Skylion007, https://github.com/zxiiro, https://github.com/cyyever, https://github.com/malfet	2025-06-12 12:42:34 +00:00
leeeizhang	eecaa0bbc6	[Multiprocesing] Fix `_release_ipc_counter` missing in rebuilding cuda ipc tensor with UntypedStorage (#155312 ) Fixes https://github.com/pytorch/pytorch/issues/155311 To avoid `torch.multiprocessing.reductions::rebuild_cuda_tensor` failed on untyped storage, this FIX PR adds the `_release_ipc_counter` into UntypedStorage like the previous legacy typed storage. `e2d141dbde/torch/storage.py (L1466-L1469)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155312 Approved by: https://github.com/mikaylagawarecki	2025-06-12 10:41:58 +00:00
Laith Sakka	0029259bdf	Add view_simple as meta function for view, and avoid calling reshape_view_helper. (#154757 ) address https://github.com/pytorch/pytorch/issues/153303 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154757 Approved by: https://github.com/bobrenjc93, https://github.com/leslie-fang-intel	2025-06-12 09:58:15 +00:00
Michael Lazos	d3d655ad14	[Hierarchical-Compile] Hash int args in addition to input shapes (#155655 ) Fixes Swsl_resnext101_32x16d in TIMM Pull Request resolved: https://github.com/pytorch/pytorch/pull/155655 Approved by: https://github.com/anijain2305	2025-06-12 06:35:12 +00:00
David Berard	c3ecabf059	[inductor][triton pin] add support for new TMA API for mm.py templates (#155723 ) Triton 3.4 will remove the experimental TMA APIs: https://github.com/triton-lang/triton/pull/6488 For mm.py templates, this PR adds support for using the new APIs when they are available (and otherwise falls back to the experimental APIs). For flex_attention, we'll remove TMA support for Triton 3.2 and 3.3 (versions of triton that don't have the new API). For mm_scaled_grouped.py, https://github.com/pytorch/pytorch/pull/150944 will remove TMA support for Triton 3.2. Note: we attempted this earlier with https://github.com/pytorch/pytorch/pull/154858, but this broke TMA usage in Triton 3.2. Differential Revision: [D76444471](https://our.internmc.facebook.com/intern/diff/D76444471) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155723 Approved by: https://github.com/NikhilAPatel	2025-06-12 06:25:47 +00:00
Nikita Shulga	2b9d638e33	[CI] Use `setup-python` from for Mac tests (#155698 ) Instead of `setup-miniconda` - Remove `CONDA_RUN` macro... - Hack the search path in `macos-test.sh` to put both python and python3 aliases first in the path (not sure what other action are messing with path environment variable) - Skip `TestMultiprocessing.test_fs_sharing` as even though it completes, it hangs on the shutdown both in CI and in all local setups I have - Skip `TestCppExtensionOpenRgistration.test_base_device_registration` as it hangs on the shutdown as well Pull Request resolved: https://github.com/pytorch/pytorch/pull/155698 Approved by: https://github.com/atalman ghstack dependencies: #155476, #155493, #155601, #155515, #155697	2025-06-12 04:58:00 +00:00
Yiming Zhou	57e4d7b5cc	[nativert] Move DelegateExecutor to PyTorch core (#155581 ) Summary: Moves DelegateExecutor base class to PyTorch core. It provides the extension point of backend delegation for NativeRT. Torch Native Runtime RFC: pytorch/rfcs#72 Test Plan: This is only a virtual base class. So relying on internal CI is sufficient. Rollback Plan: Differential Revision: D76351984 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155581 Approved by: https://github.com/zhxchen17	2025-06-12 04:33:31 +00:00
Animesh Jain	a9d5157e25	[dynamo] Use BINARY_SUBSCR for pre-graph bytecode for regular dict accesses (#155727 ) vLLM profiler sets with_stack=True that shows the dict_getitem on the profiler, both inflating the numbers and confusing compile users. This PR keeps BINARY_SUBSCR for regular dicts, while using `dict.__getitem__` only for dict subclasses. Using binary_subscr is little bit faster, but not enough to make any major latency improvements. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155727 Approved by: https://github.com/zou3519, https://github.com/StrongerXi, https://github.com/jansel	2025-06-12 04:02:29 +00:00
Animesh Jain	c9e9a0c823	[inductor][invoke_subgraph] Mark invoke_subgraph outputs as user_visible to constrain output strides (#155395 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155395 Approved by: https://github.com/zou3519	2025-06-12 03:58:16 +00:00
Shangdi Yu	9f5153b1a4	Preserve GrpahModule node stack trace after torch package deserializaion re-tracing (#155638 ) Summary: urrently the node.meta["stack_trace"] is not preserved when we torch package/load GraphModule, which means the original stack trace is lost. When we re-trace the packaged graph module, we just get a stack trace like fx-generated._0...... Adding the node.meta["stack_trace"] to torch packaged graph module Test Plan: ``` buck2 run @//mode/dev-nosan fbcode//caffe2/test:package -- -r TestPackageFX ``` Rollback Plan: Differential Revision: D76379692 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155638 Approved by: https://github.com/angelayi	2025-06-12 03:48:27 +00:00
Nikita Shulga	ce9ba071fd	[BE] Fix warning in open_registration_extension.cpp (#155755 ) Namely ``` /Users/nshulga/git/pytorch/pytorch/test/cpp_extensions/open_registration_extension.cpp:306:33: warning: left operand of comma operator has no effect [-Wunused-value] 306 \| at::Tensor first = at::empty((2,3)).to(at::DeviceType::PrivateUse1); ``` Or switching between Python and C++ is hard In Python `(2, 3)` creates a tuple, in C/C++ it's just a integral literal 3 P.S. I could have vibe-coded the fix with Claude: https://claude.ai/share/82479e88-84cb-4299-aa2f-dafd28ee2d55 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155755 Approved by: https://github.com/huydhn, https://github.com/atalman	2025-06-12 03:01:30 +00:00
Yiming Zhou	d96dec8415	[export] Fix serialization for call_torchbind hop with as_none argument (#155647 ) Summary: As title. D75251816 broke one internal test. This diff fixes it. Test Plan: Internal CI Differential Revision: D76383202 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155647 Approved by: https://github.com/ydwu4	2025-06-12 02:59:03 +00:00
Kazuaki Ishizaki	b00b641ff1	[Docs] Convert to markdown: accelerator.rst, amp.rst, autograd.rst, backends.rst, benchmark_utils.rst (#155762 ) Fixes #155013 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155762 Approved by: https://github.com/svekars	2025-06-12 02:55:06 +00:00
Xia, Weiwen	b6f84b3b0f	[Inductor][CPU] Use AMX-based microkernels when M > 4 for GEMM template for INT4 weight (#155444 ) Summary GEMM templates for INT4 weights are used for lowering `aten._weight_int4pack_mm_for_cpu` with Inductor when max-autotune is on. Currently, AMX-based microkernels are used only when M >= 16 if input tensor has shape [M, K]. However, we find that AMX kernel brings performance benefit when 4 < M < 16. For example, on a 6th gen of Intel(R) Xeon(R) platform, E2E latency can be improved by up to > 20% when running Llama-3.1-8B on 32 cores for M = 8. So, this PR changes the threshold so that AMX is used when M > 4. Test plan ``` pytest test/inductor/test_cpu_select_algorithm.py -k test_int4_woq_mm ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155444 Approved by: https://github.com/sanchitintel, https://github.com/leslie-fang-intel	2025-06-12 02:28:48 +00:00
Simon Fan	212575f994	[ca] Annotate AccumulateGrad branching and add polyfill tests (#155289 ) Annotates AccumulateGrad and tracks the semantics for AccumulateGrad's polyfill , except for Scenario 1.4: Cloning MKLDNN new_grad and Scenario 2.2: Vmap-incompatible. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155289 Approved by: https://github.com/jansel, https://github.com/albanD	2025-06-12 02:10:52 +00:00
Yu, Guangye	d84efde3f0	Move _storage_Use_Count to be gerneric (#155451 ) # Motivation `torch._C._storage_Use_Count` should be a generic API that is not aware of device type. It is also used in `337cd7c53d/torchtune/training/_activation_offloading.py (L323)` to do some memory optimization. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155451 Approved by: https://github.com/albanD	2025-06-12 01:39:04 +00:00
PyTorch MergeBot	8372d0986a	Revert "[PT2][partitioners] Add aten.split to view_ops list (#155424 )" This reverts commit e1db10e05aa720aef1989773adcf48f311bcf920. Reverted https://github.com/pytorch/pytorch/pull/155424 on behalf of https://github.com/clee2000 due to I think this broke inductor/test_cpu_repro.py::CPUReproTests::test_transpose_with_norm [GH job link](https://github.com/pytorch/pytorch/actions/runs/15596830833/job/43931044625) [HUD commit link](`e1db10e05a`) but idk how, reverting to see if it fixes the problem ([comment](https://github.com/pytorch/pytorch/pull/155424#issuecomment-2964717706))	2025-06-12 01:38:34 +00:00
Catherine Lee	9b122aab5d	Fix set per proc memory fraction when running tests (#155631 ) env setting needs to happen before pool creation for it to take effect In theory this should fix some OOMs and also cause some OOMs, but this PR is green so idk alt options: use initializer? Pull Request resolved: https://github.com/pytorch/pytorch/pull/155631 Approved by: https://github.com/huydhn, https://github.com/malfet, https://github.com/seemethere, https://github.com/atalman	2025-06-12 01:28:08 +00:00
Pian Pawakapan	8ad6197b46	[draft export] avoid storing intermediate real tensors in proxies (#154630 ) Handles GC for non-strict draft export; GPU memory usage shouldn't be much more than eager mode + input tensors now. While trying to do draft export CPU offloading, I found out GC is feasible, because in non-strict, there's 2 places holding references to a `.real_tensor` attribute: 1) the FakeTensors in fake tensor prop, but these are held by the actual variables in the model's forward call, and so the real tensor gets gc-ed along with the fake one when the variable goes out of scope. 2) A clone of the fake tensor in 1) stored in `proxy.node.meta["val"]`, which was added in https://github.com/pytorch/pytorch/pull/150948. But we didn't actually need to store them on intermediate values; the placeholders are enough for retracing/lowering. Avoiding storing the intermediate values in 2), the values in 1) should be naturally GC-ed, and the real-tensor memory usage for non-strict should be pretty similar to eager computation? Strict still OOMs; dynamo still holds these in variable tracking, and not sure how to GC those. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154630 Approved by: https://github.com/angelayi, https://github.com/yushangdi	2025-06-12 01:18:57 +00:00
Shangdi Yu	4e19477196	[nativert] Move Pytree (#155136 ) Summary: fbcode/sigmoid/core/common -> fbcode/caffe2/torch/nativert/common Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 Test Plan: ``` buck run fbcode//mode/dev-nosan //caffe2/test/cpp/nativert:pytree_test ``` OSS CI Rollback Plan: Differential Revision: D75965059 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155136 Approved by: https://github.com/zhxchen17, https://github.com/XuehaiPan, https://github.com/zou3519	2025-06-12 01:10:34 +00:00
Wanchao Liang	ee5c2908cb	[dtensor] refactor PlacementStrategy -> OpSpec, move utils to OpSchema (#155592 ) as titled. It's sometimes confusing to use PlacementStrategy as a name, as we also have OpStrategy and TupleStrategy, the latter two contain the former, so it is better to make the naming clearer. Renaming PlacementStrategy -> OpSpec as it is an operator spec that contains output_spec + input_specs. Also found some utils can be merged to OpSchema so included together in this PR Pull Request resolved: https://github.com/pytorch/pytorch/pull/155592 Approved by: https://github.com/awgu	2025-06-12 00:51:36 +00:00
Huy Do	7485ef078f	Run torch.compile benchmark more frequently on H100 (#155719 ) We have more capacity now with 20+ `linux.aws.h100` runners, half of them are idle. Running benchmark more frequently would utilize these runner better and provide early signals multiple times per day. Running every 8 hours to start with. The workflow usually finishes within 5 hours https://github.com/pytorch/pytorch/actions/runs/15578331612/job/43878878434 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155719 Approved by: https://github.com/atalman	2025-06-12 00:24:21 +00:00
Ke Wen	9e9484d022	[SymmMem] Enable NVSHMEM for Triton (#155506 ) (This is an Experimental feature) Allow Triton kernels to invoke NVSHMEM device functions. ### Example Triton program Key parts: - Call `nvshmem.enable_triton()` to initialize; - Call `nvshmem.putmem_block` in Triton kernel; - Add `extern_libs` kwarg at kernel invocation. ``` import torch.distributed._symmetric_memory._nvshmem_triton as nvshmem @triton.jit def put_kernel( dst_ptr, src_ptr, numel: tl.constexpr, peer: tl.constexpr, BLOCK_SIZE: tl.constexpr, ): nvshmem.putmem_block(dst_ptr, src_ptr, numel, peer) if __name__ == "__main__": # Enable NVSHMEM for Triton nvshmem_lib = nvshmem.enable_triton() # Use torch Symmetric Memory to allocate Symmetric tensors ... peer = 1 - rank if rank == 0: kernel = put_kernel[(1, 1, 1)]( dst_ptr, src_ptr, numel=numel, peer=peer, BLOCK_SIZE=BLOCK_SIZE, extern_libs=nvshmem_lib, ) dist.barrier() if rank == 1: print(f"Rank {rank}: received {out=}") ``` ### Test output: ``` $ TORCH_SYMMMEM=NVSHMEM python test/distributed/test_nvshmem.py -k test_triton_put Rank 0: writing value 5 to Peer 1 Rank 1: received out=tensor([5, 5, 5, 5, 5, 5, 5, 5], device='cuda:1', dtype=torch.int8) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155506 Approved by: https://github.com/ngimel, https://github.com/fegin, https://github.com/fduwjj	2025-06-12 00:22:49 +00:00
Justin Silver	cf9878d7a2	Fix #155022 rst to markdown conversion (#155540 ) Used [rst2myst tool](https://rst-to-myst.readthedocs.io/en/latest/) Fixes #155022 Docs comparison (check out the 'new' whenever docs build) 1. func.ux_limitations ([old](https://docs.pytorch.org/docs/main/func.ux_limitations.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/func.ux_limitations.html)) 2. func.whirlwind_tour ([old](https://docs.pytorch.org/docs/main/func.whirlwind_tour.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/func.whirlwind_tour.html)) 3. future_mod ([old](https://docs.pytorch.org/docs/main/future_mod.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/future_mod.html)) 4. futures ([old](https://docs.pytorch.org/docs/main/futures.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/futures.html)) 5. fx.experimental ([old](https://docs.pytorch.org/docs/main/fx.experimental.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155540/fx.experimental.html)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155540 Approved by: https://github.com/AlannaBurke, https://github.com/svekars	2025-06-12 00:21:22 +00:00
Sidharth	7918978653	[dynamo] uploaded full json file of all unimplemented_v2() calls currently in repository (#155758 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155758 Approved by: https://github.com/williamwen42	2025-06-12 00:17:28 +00:00
Tsung-Hsien Lee	a6210fd07b	[c10d] Enhance `get_process_group_ranks()` to accept `group=None` (#154902 ) Summary: This diff enhances the `get_process_group_ranks()` function to accept `group=None` as an optional argument. This allows the function to return all ranks associated with the default process group if no group is specified. Test Plan: contbuild & OSS CI Rollback Plan: Differential Revision: D75817800 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154902 Approved by: https://github.com/wz337	2025-06-11 23:41:03 +00:00
eqy	bd3c32916c	[cuDNN] Enabled dilation for deterministic convolutions in cuDNN (#154292 ) Provides order-of-magnitude speedup over fallback impl. https://github.com/pytorch/pytorch/issues/28777 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154292 Approved by: https://github.com/Skylion007 Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-06-11 23:35:52 +00:00
Ankita George	c13e725edd	Updates to HFStorageReader to use TensorStorageMetadata instead of BytesStorageMetadata (#154518 ) As we prepare to support re-sharding, the current approach of using BytesStorageMetadata to read safetenstors won't work anymore. Before, we didn't need to read the metadata of the safetensors file from its header because we were just loading the contents of the file directly into tensors with safetensor.load() that would handle the metadata and deserialization. But now, in preparation of handling re-sharding, we need to read the metadata directly from the header of the safetensors file and store it directly in TensorStorageMetadata objects so that we can perform re-sharding. Re-sharding won't currently work, as we need extra metadata to be stored on each save, so that will be added in a subsequent PR. In addition this PR adds an integration test in addition to the unit tests. It also removes the HfFileSystem import because that's only needed if users are using HfFileSystem, but we want to support any backend. Differential Revision: [D74891998](https://our.internmc.facebook.com/intern/diff/D74891998/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154518 Approved by: https://github.com/saumishr	2025-06-11 23:35:05 +00:00
jafraustro	1b032384b1	Convert rst files to md (#155369 ) Fixes #155021 Fixes #155158 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155369 Approved by: https://github.com/svekars, https://github.com/malfet	2025-06-11 23:00:52 +00:00
Nikita Shulga	48921721d8	[MPS] Fix binary builds (#155733 ) Introduced by https://github.com/pytorch/pytorch/pull/155611 All functions in those headers must be static and inline Pull Request resolved: https://github.com/pytorch/pytorch/pull/155733 Approved by: https://github.com/seemethere, https://github.com/atalman	2025-06-11 22:55:33 +00:00
Colin Peppler	c1446e1e9d	[easy] revert unintended changes from #152579 (#155614 ) Summary: I accidentally removed a test and a small change in my pr: https://github.com/pytorch/pytorch/pull/152579 - `test_load_package_multiple_gpus` from https://github.com/pytorch/pytorch/pull/152093 Rollback Plan: Differential Revision: D76370555 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155614 Approved by: https://github.com/jingsh	2025-06-11 22:54:58 +00:00
Yidi Wu	4a954fc185	[refactor] make do_auto_functionalize_v2 take HopInstance (#154192 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154192 Approved by: https://github.com/zou3519 ghstack dependencies: #155261, #154072, #154191	2025-06-11 22:52:37 +00:00
Yidi Wu	d6be87648f	[hop schema] add schema.tree_spec to support pytree inputs (#154191 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154191 Approved by: https://github.com/zou3519 ghstack dependencies: #155261, #154072	2025-06-11 22:52:37 +00:00
Yidi Wu	6ded656aee	[hop] auto functionalize invoke_subgraph (#154072 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154072 Approved by: https://github.com/zou3519 ghstack dependencies: #155261	2025-06-11 22:52:28 +00:00
Yidi Wu	20fb8f5d1f	[refactor] make check input alias and mutation easier to use (#155261 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155261 Approved by: https://github.com/zou3519	2025-06-11 22:52:21 +00:00
Shunting Zhang	61e13782dd	[inductor] handle -1 for pointless view pairs (#155295 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155295 Approved by: https://github.com/laithsakka, https://github.com/jansel	2025-06-11 22:20:36 +00:00
loganthomas	458cc7213b	DOC: Convert to markdown: mobile_optimizer.rst, model_zoo.rst, module_tracker.rst, monitor.rst, mps_environment_variables.rst (#155702 ) Fixes #155026 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155702 Approved by: https://github.com/sekyondaMeta, https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-11 22:16:04 +00:00
Shatian Wang	e1db10e05a	[PT2][partitioners] Add aten.split to view_ops list (#155424 ) Summary: Add `aten.split` to view_ops list in partitioners.py Test Plan: na Rollback Plan: Differential Revision: D76011951 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155424 Approved by: https://github.com/xuanzhang816	2025-06-11 22:12:13 +00:00
PyTorch MergeBot	f59c76b549	Revert "[BE]: Update cudnn to 9.10.2.21 (#155576 )" This reverts commit 2d3615f577894c7a117a55e85bb8371bb598ec50. Reverted https://github.com/pytorch/pytorch/pull/155576 on behalf of https://github.com/malfet due to breaks the same test again (I remember there were a version that adjusted tolerances), see `bc3972b80a/1` ([comment](https://github.com/pytorch/pytorch/pull/155576#issuecomment-2964404710))	2025-06-11 22:03:45 +00:00
Shangdi Yu	bc3972b80a	[reland] Add stack_trace on make_fx (#155486 ) Summary: Previosuly, we only add stack trace in class _ModuleStackTracer(PythonKeyTracer) for non-strict export. I moved this stack trace logic to the parent class PythonKeyTracer, this way the graph traced from Module using make_fx will have stack_trace as well. Motivation: we've observed some uses cases where users first use make_fx on the Module, and then run export on the resulting graph. If the result of make_fx doesn't have stack trace, the stack trace information is lost. User needs to turn this on by passing in `stack_trace=True` to make_fx. We don't make this the default option since this might increase inductor compilation time (`make_fx` is used in inductor to trace graph patterns for pattern matching). It's also turned on if `_inductor.config.trace.enabled` is True. preserving stack trace is on by default for ModuleStackTracer, which is used for non-strict export. Test Plan: ``` buck run test:test_export -- -r test_stack_trace buck run fbcode//caffe2/test/dynamo:test_dynamo -- -k test_autocast_ordering ``` Rollback Plan: Differential Revision: D76298692 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155486 Approved by: https://github.com/angelayi, https://github.com/zou3519	2025-06-11 21:27:43 +00:00
Pian Pawakapan	9bd0830ed8	[dynamic shapes] guard_or_false for cat, repeat (#155290 ) Summary: assumes: - specified repeats are non-negative - 1d cat arguments like [u0] aren't non-zero sized (replaces existing size-oblivious) Test Plan: test_export Rollback Plan: Differential Revision: D76092011 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155290 Approved by: https://github.com/laithsakka	2025-06-11 21:03:32 +00:00
Manuel Candales	4609699bfd	[MPS] Migrate leaky_relu (forward and backward) to Metal kernel (#155571 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155571 Approved by: https://github.com/malfet ghstack dependencies: #155304, #155316, #155462, #155479	2025-06-11 20:58:46 +00:00
Manuel Candales	f8d93b3783	[MPS] Migrate hardswish (forward and backward) to Metal kernel (#155479 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155479 Approved by: https://github.com/kulinseth, https://github.com/malfet ghstack dependencies: #155304, #155316, #155462	2025-06-11 20:58:46 +00:00
Zhihan Fang	db5970c1a6	[coreml-backend-tool] fix pytorch-backended issue on new coremltools (#155543 ) Summary: the new coreml tool is export mlpakage instead mlmodel in default option. when we use new 8.0 coreml tool to convert to backend, the error is ``` Exception: MLModel of type mlProgram cannot be loaded just from the model spec object. It also needs the path to the weights file. Please provide that as well, using the 'weights_dir' argument. ``` Test Plan: tested with internal workflow Rollback Plan: Differential Revision: D76325462 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155543 Approved by: https://github.com/shoumikhin	2025-06-11 20:52:26 +00:00
Laith Sakka	cec264c8c6	remove single remaining gso from compute_stride (#155635 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155635 Approved by: https://github.com/ColinPeppler	2025-06-11 20:36:21 +00:00
Laith Sakka	cc09d3a5ba	remove float args benchmark (#155674 ) This benchmark very sensitive. removing it for now until we make it better . <img width="755" alt="Screenshot 2025-06-11 at 12 01 25 AM" src="https://github.com/user-attachments/assets/01a45ae5-2028-42a2-b819-c30d4db3b5d4" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/155674 Approved by: https://github.com/bdhirsh, https://github.com/bobrenjc93	2025-06-11 20:34:58 +00:00
Aaron Gokaslan	2d3615f577	[BE]: Update cudnn to 9.10.2.21 (#155576 ) Update to CUDNN 9.10.2.21 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155576 Approved by: https://github.com/eqy, https://github.com/atalman	2025-06-11 20:32:07 +00:00
Catherine Lee	94ae615337	[trymerge] Error on ghstack commit with multiple PRs (#154941 ) see https://github.com/pytorch/pytorch/issues/154427#issuecomment-2932941343 for context Errors if do not find 1 match in ghstack commit Pull Request resolved: https://github.com/pytorch/pytorch/pull/154941 Approved by: https://github.com/malfet, https://github.com/seemethere, https://github.com/atalman	2025-06-11 20:26:50 +00:00
Justin Silver	b7a73a2cdb	Convert to markdown: export.programming_model.rst (#155659 ) Converts only export.programming_model.rst to markdown Used [rst2myst tool](https://rst-to-myst.readthedocs.io/en/latest/) Fixes #155020, but split into a second PR to pass sanity check Docs comparison (check out the 'new' whenever docs build) 1. export.programming_model ([old](https://docs.pytorch.org/docs/main/export.programming_model.html) vs. [new](https://docs-preview.pytorch.org/pytorch/pytorch/155659/export.programming_model.html)) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155659 Approved by: https://github.com/sekyondaMeta	2025-06-11 20:23:46 +00:00
Shuqi Yang	1b6772a90f	A small fix in do_bench_using_profiling (#155500 ) Summary: Results: https://docs.google.com/document/d/1B_4rtiDFPH_jV3VpnqLPnInwDMpF7yX29G82UoJTcu8/edit?tab=t.0 Test Plan: ``` buck2 run mode/opt -c fbcode.enable_gpu_sections=true ai_acceleration/float8/benchmarks/bench:bench_fp8_shapes_eval 2>&1 \| tee output44.txt ``` Rollback Plan: Differential Revision: D76298690 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155500 Approved by: https://github.com/yoyoyocmu, https://github.com/nmacchioni	2025-06-11 20:06:19 +00:00
Scott Wolchok	1dd0b1d12b	Unbreak torch.is_vulkan_available() on Mac (re-send of #154675 , please stamp) (#155595 ) This is a new PR duplicating #154675 due to merge issues with that PR coming from my old (now updated) version of ghstack. I am a Vulkan noob, but this extension and flag seem to be necessary. See "Encounted VK_ERROR_INCOMPATIBLE_DRIVER" at https://vulkan-tutorial.com/Drawing_a_triangle/Setup/Instance . (For anyone trying to repro at home, I have the following homebrew packages installed, not all of which may be necessary: molten-vk, vulkan-headers, vulkan-loader, vulkan-tools, vulkan-utility-libraries. I also have VK_ICD_FILENAMES set to /opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.json, and I built PyTorch with USE_VULKAN=1. Making sure vkcube works helped me debug this setup.) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155595 Approved by: https://github.com/malfet	2025-06-11 19:51:35 +00:00
Oguz Ulgen	d1947a8707	Migrate from lru_cache to cache (#155613 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155613 Approved by: https://github.com/ezyang ghstack dependencies: #155612	2025-06-11 19:44:18 +00:00
PyTorch MergeBot	f80a61adf5	Revert "[dynamo] added github_cli to detect unimplemented_v2 calls (#155610 )" This reverts commit 5dd07c70e53a86b73f49711b8186d86dc4f1b32a. Reverted https://github.com/pytorch/pytorch/pull/155610 on behalf of https://github.com/malfet due to Looks like it fails on every pull request, based on https://github.com/pytorch/pytorch/actions/workflows/check-unimplemented-calls.yml, but it does not run on trunk ([comment](https://github.com/pytorch/pytorch/pull/155610#issuecomment-2963929765))	2025-06-11 19:31:55 +00:00
Ti-Tai Wang	1e373d02d5	[ONNX] Change deprecation message from 2.8 to 2.9 (#155580 ) ~~The PR: https://github.com/pytorch/pytorch/pull/152478 did not respect the release policy that the deprecation should happen after the deprecation message has been set for 2 releases. This PR postpone 2.8 to the rightful version 2.10.~~ ~~NOTE: "as early as" 2.10 shall give ONNX users more time to adapt and provide feedback.~~ To follow the upcoming torchscript deprecation, `torch.onnx.export` expects to switch dynamo=True (also turn on fallback=True for bc) on torch 2.9. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155580 Approved by: https://github.com/justinchuby, https://github.com/tugsbayasgalan	2025-06-11 19:31:29 +00:00
Pian Pawakapan	3f29642ecf	Update XLA pin (#155471 ) Update pin after XLA PR https://github.com/pytorch/xla/pull/9312 landed Pull Request resolved: https://github.com/pytorch/pytorch/pull/155471 Approved by: https://github.com/laithsakka	2025-06-11 19:16:52 +00:00
Aleksandar Samardžić	f8baec8984	Update auto-tuning support for _scaled_grouped_mm (#150944 ) 1. Enable strided inputs 2. Implement "2d/2d", "3d/2d" and "3d/3d" combinations of inputs 3. Fix non-TMA load variant 4. Replace experimental_device_tensormap_create2d with _experimental_make_tensor_descriptor 5. Fix cases when group size along K dimension is not multiple of block size along K 6. Updated meta registration 7. Update synthetic offsets creation Pull Request resolved: https://github.com/pytorch/pytorch/pull/150944 Approved by: https://github.com/ngimel, https://github.com/davidberard98	2025-06-11 19:12:52 +00:00
Simon Fan	6dfada220e	[ca] better error message for subclasses not supported by FakeTensor (#155481 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155481 Approved by: https://github.com/jansel ghstack dependencies: #155473, #155570	2025-06-11 19:09:29 +00:00
Simon Fan	5dcc718a77	[dynamo][ci] update PYTORCH_TEST_WITH_DYNAMO xfail/skips script for 3.13 (#155570 ) No more 311 runners, tested by generating the files for the next PRs Pull Request resolved: https://github.com/pytorch/pytorch/pull/155570 Approved by: https://github.com/zou3519 ghstack dependencies: #155473	2025-06-11 19:09:29 +00:00
Simon Fan	87b002b6fb	[ca] make torch.compile API respect ambient disable contexts (#155473 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155473 Approved by: https://github.com/jansel	2025-06-11 19:09:29 +00:00
Manuel Candales	be124a61a4	[MPS] Migrate hardsigmoid (forward and backward) to Metal kernel (#155462 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155462 Approved by: https://github.com/malfet ghstack dependencies: #155304, #155316	2025-06-11 19:09:23 +00:00
Oguz Ulgen	c04a4e7094	Add types to torch/utils/_triton.py (#155612 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155612 Approved by: https://github.com/jamesjwu	2025-06-11 19:04:10 +00:00
Kazuaki Ishizaki	2002e3a311	[Docs] Convert to markdown: torch.compiler_transformations.rst, torch.compiler.config.rst (#155347 ) Part of changes #155040 (parent PR #155120) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155347 Approved by: https://github.com/svekars	2025-06-11 18:55:30 +00:00
Runtian (Rachel) Li	925fbfca27	Convert fx.rst to fx.md (#155482 ) Part of changes #155023 (parent PR #155429) @pytorchbot label "topic: docs" @pytorchbot label "topic: not user facing" @pytorchbot label docathon-h1-2025 @pytorchbot label module: docs Pull Request resolved: https://github.com/pytorch/pytorch/pull/155482 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-11 18:46:35 +00:00
Eddie Yan	4d9d884c3f	[NCCL] Expose new `ncclConfig_t` flags in 2.27 (#155379 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155379 Approved by: https://github.com/Skylion007	2025-06-11 18:26:55 +00:00
Pian Pawakapan	247f83e0a4	[dynamic shapes] guard individual terms in sym_and; user-code-friendly sym_and/sym_or (#154737 ) Previously when processing `sym_and(a, b, c)`, symbolic shapes wouldn't individually process a, b, and c and store their implications. This would lead us to data-dependent error on individual checks, e.g. we stored `u0 >= 0 & u0 <= 10`, but then couldn't figure out `u0 <= 10`. This handles that, and also makes `sym_and/or` user-code friendly, for testing. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154737 Approved by: https://github.com/laithsakka	2025-06-11 18:08:06 +00:00
Nikita Shulga	c1cbaca7fd	[CI] Move setuptools requirements from conda to pip (#155697 ) Needed for `import z5` to work without warning, otherwise `LoggingTests.test_logs_out` will fail Pull Request resolved: https://github.com/pytorch/pytorch/pull/155697 Approved by: https://github.com/atalman ghstack dependencies: #155476, #155493, #155601, #155515	2025-06-11 18:03:18 +00:00
PyTorch MergeBot	3a43dba21f	Revert "[cuBLASLt][cuBLAS] Support 2D bias and `beta != 1.0` in cuBLASLt (#154170 )" This reverts commit dc5e8f7999cccb51efcf0f5fe197a740a918c73d. Reverted https://github.com/pytorch/pytorch/pull/154170 on behalf of https://github.com/malfet due to It broke ROCM, see `c75c732481/1` ([comment](https://github.com/pytorch/pytorch/pull/154170#issuecomment-2963708109))	2025-06-11 18:01:08 +00:00
Nikita Shulga	c75c732481	[CI] Disable ET tests (#155708 ) I'm tired of seeing red on PRs and it has been consistently broken since May 30th per `59eb61b2d1/10` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155708 Approved by: https://github.com/clee2000, https://github.com/atalman	2025-06-11 17:56:52 +00:00
penknife6153	59eb61b2d1	[inductor] Improve GEMM logging to display batch size for batched operations (#155544 ) Improves the GEMM overview logging in PyTorch Inductor to properly display batch size information for batched matrix operations like `torch.bmm` and `torch.baddbmm`. Fixes #155307 ## Problem The current GEMM logging for `torch.bmm` shows: ```python # Repro import os os.environ["TORCH_LOGS"] = "inductor" import torch M, N, K = 1024, 1024, 1024 dtype = torch.bfloat16 A = torch.randn(10, M, K, device="cuda", dtype=dtype) B = torch.randn(10, K, N, device="cuda", dtype=dtype) compiled_model = torch.compile(torch.bmm, fullgraph=True) _ = compiled_model(A, B) ``` Before: ``` Name \| M \| N \| K \| Count ---------------------------------------------------------------------------------------------------- aten.bmm \| 1024 \| 1024 \| 1024 \| 1 ---------------------------------------------------------------------------------------------------- ``` The batch size (10) is missing from the logs, making it unclear what the actual operation dimensions were. ## Solution After: ``` Name \| B \| M \| N \| K \| Count ---------------------------------------------------------------------------------------------------------------------------------- aten.bmm \| 10 \| 1024 \| 1024 \| 1024 \| 1 aten.mm \| - \| 1024 \| 1024 \| 1024 \| 2 ---------------------------------------------------------------------------------------------------------------------------------- ``` ## Changes Made ### 1. Enhanced Parsing Logic in compile_fx.py - Detects batched operations by checking if operation name ends with `'bmm'` or `'baddbmm'` - For batched operations: takes last 4 parts as `batch, m, n, k` - For non-batched operations: takes last 3 parts as `m, n, k` - Dedicated "B" column: Added separate column for batch size instead of embedding in operation name - Shows batch size for batched operations, shows "-" for non-batched operations ### 2. Updated All MM Operations for Consistency - bmm.py: - Extract batch size from `mat1.get_size()[0]` for both `tuned_bmm` and `tuned_baddbmm` - Use positional counter keys: `aten.bmm_{batch_size}_{m}_{n}_{k}` - Enhanced log messages to include batch size information - mm.py: Updated counter keys for consistency: - `aten.mm_{m}_{n}_{k}` (no batch dimension) - `aten.addmm_{m}_{n}_{k}` (no batch dimension) - `aten._int_mm_{m}_{n}_{k}` (no batch dimension) - `aten._scaled_mm.default_{m}_{n}_{k}` (no batch dimension) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155544 Approved by: https://github.com/jansel, https://github.com/BoyuanFeng	2025-06-11 16:57:40 +00:00
Colin Peppler	7b7cd56f5e	[export] support linear & layer_norm unbacked (#155260 ) ## What - use `definitely_contiguous_for_memory_format` instead of `is_contiguous` when the non-contiguous case is fine if we encounter a DDE. - use ref's contiguous over Aten's contiguous because Aten's version will DDE and stop tracing. ref's version will use `definitely_contiguous_for_memory_format` and clone if there's a DDE. ## Example DDEs - Fixed with `definitely_contiguous_for_memory_format` in `fast_binary_impl` ``` torch._dynamo.exc.UserError: Could not guard on data-dependent expression Eq((u0//387), 0) (unhinted: Eq((u0//387), 0)). (Size-like symbols: u0) Caused by: layer_norm = self.layer_norm(linear) # caffe2/test/export/test_export.py:4566 in forward (_subclasses/fake_impls.py:1022 in fast_binary_impl) ``` - Fixed with `refs.contiguous` instead of calling aten's contiguous (that'd require a bigger re-write in Aten) ``` File "c10/core/TensorImpl.h", line 825, in torch::autograd::THPVariable_contiguous(_object, _object, _object) File "c10/core/SymbolicShapeMeta.h", line 87, in c10::TensorImpl::is_contiguous_default(c10::MemoryFormat) const File "c10/core/SymbolicShapeMeta.cpp", line 250, in c10::SymbolicShapeMeta::init_is_contiguous() const torch.fx.experimental.symbolic_shapes.GuardOnDataDependentSymNode: Could not guard on data-dependent expression Eq(128((u0//387)), 0) (unhinted: Eq(128((u0//387)), 0)). (Size-like symbols: u0) Caused by: (_refs/__init__.py:3302 in native_layer_norm) ``` - Fixed with `definitely_contiguous_for_memory_format` in ref's contiguous ``` torch.fx.experimental.symbolic_shapes.GuardOnDataDependentSymNode: Could not guard on data-dependent expression 387((u0//387)) < 2 (unhinted: 387*((u0//387)) < 2). (Size-like symbols: u0) Caused by: (_prims_common/__init__.py:279 in is_contiguous) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155260 Approved by: https://github.com/laithsakka ghstack dependencies: #155499	2025-06-11 16:47:34 +00:00
Jing Shan	b49edc0d6c	[Export] Fix some typos in docstring (#155485 ) Summary: nit change, fix the doc string Test Plan: CI Rollback Plan: Differential Revision: D76297740 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155485 Approved by: https://github.com/ColinPeppler	2025-06-11 16:44:38 +00:00
zeshengzong	18bf6addc4	set_grad_enabled add str and repr for prints (#155681 ) Fixes #86718 ## Test Result ```python >>> import torch >>> torch.set_grad_enabled(False) torch.autograd.grad_mode.set_grad_enabled(mode=False) >>> print(torch.set_grad_enabled(False)) torch.autograd.grad_mode.set_grad_enabled(mode=False) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155681 Approved by: https://github.com/soulitzer	2025-06-11 16:01:03 +00:00
Eddie Yan	dc5e8f7999	[cuBLASLt][cuBLAS] Support 2D bias and `beta != 1.0` in cuBLASLt (#154170 ) Fixes https://github.com/pytorch/pytorch/issues/153590 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154170 Approved by: https://github.com/malfet	2025-06-11 15:20:48 +00:00
PyTorch MergeBot	45c5a23237	Revert "Add Intel GPU info collection to the collect env script (#137846 )" This reverts commit 5264f8cd8d08272003298cdefe6bd60b1b8c80b4. Reverted https://github.com/pytorch/pytorch/pull/137846 on behalf of https://github.com/malfet due to Just testing if it will fix PR time benchmarks signal ([comment](https://github.com/pytorch/pytorch/pull/137846#issuecomment-2963232606))	2025-06-11 15:18:47 +00:00
Nikita Shulga	359e8f5d69	[CI] Use `setup-python` from test-infra to do MacOS builds (#155515 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155515 Approved by: https://github.com/cyyever, https://github.com/Skylion007, https://github.com/atalman ghstack dependencies: #155476, #155493, #155601	2025-06-11 15:11:38 +00:00
David Berard	9328a7fb58	[triton pin][tests] refactor test_triton_kernel.py tests to test new & old API (#155510 ) This splits out the tests so we can independently test both the new and old API. Note: the new API doesn't work yet - we still need to fix those tests. Differential Revision: [D76318840](https://our.internmc.facebook.com/intern/diff/D76318840) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155510 Approved by: https://github.com/oulgen	2025-06-11 13:52:15 +00:00
Ting Lu	4c3da611c2	Add CUDA 12.9.1 x86 nightly binaries (#154980 ) Adding CUDA 12.9.1 to nightly binaries matrix for linux (x86) builds. Add sbsa and libtorch build docker images, builds addition will be follow-up PRs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154980 Approved by: https://github.com/eqy, https://github.com/atalman	2025-06-11 13:43:17 +00:00
Kurt Mohler	013cf1e330	[MPS] Move expm1 op to Metal (#155611 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155611 Approved by: https://github.com/malfet	2025-06-11 13:06:14 +00:00
Bin Bao	44df7cf28d	[AOTI] Fix embed_kernel_binary error when max_autotune is ON (#155569 ) Summary: Stop removing cubin files so that it won't be missing when max_autotune is ON. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155569 Approved by: https://github.com/angelayi, https://github.com/yushangdi	2025-06-11 12:27:36 +00:00
Boyuan Feng	f34ab1628b	[Graph Partition] move cpu scalar tensor to gpu (#154464 ) cudagraph does not support cpu tensors. In this PR, we update the graph by explicitly moving cpu tensors to gpu when profitable, relying on graph partition to split off this data copy, and cudagraphifying the remaining gpu ops. This PR unblocked the graph partition + cudagraph on speech_transformer, leading to 39.5% speedup on inference [P1830602200](https://www.internalfb.com/phabricator/paste/view/P1830602200), 85% speedup on training [P1831115315](https://www.internalfb.com/phabricator/paste/view/P1831115315). Close: #119241 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154464 Approved by: https://github.com/eellison, https://github.com/mlazos	2025-06-11 10:22:45 +00:00
Wang, Chuanqi	eaceb243df	[BE] Update the XPU support package to 2025.1.3 (#154346 ) Fixes #153632 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154346 Approved by: https://github.com/EikanWang, https://github.com/atalman	2025-06-11 09:46:18 +00:00
Laith Sakka	2585960b47	remove redundent type_id (#155539 ) Those were added in https://github.com/pytorch/pytorch/pull/92229 to prevent confusion of overloads. but the variants that accepts SymBool are all removed in https://github.com/pytorch/pytorch/pull/112890 with the introduction of SymbolicShapeMeta. Hence that dummy arg is not needed anymore. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155539 Approved by: https://github.com/ezyang	2025-06-11 08:46:56 +00:00
David Berard	717a099d42	Revert "[flex attention][triton pin] triton_helpers shim for TMA apis (#154858 )" (#155640 ) This reverts commit ea7b233015ff00098df687884be4e2efbf7a55fa. It fails internal tests in fbcode, but they weren't running so we didn't get signal Reverting w/ a PR/diff because the PR has been landed for ~1 week - too old to revert directly from internal. Differential Revision: [D76380887](https://our.internmc.facebook.com/intern/diff/D76380887) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155640 Approved by: https://github.com/nmacchioni, https://github.com/danzimm	2025-06-11 07:37:47 +00:00
Oguz Ulgen	0e2013a12d	Add helion x pt2 test (#155513 ) This kinda just worked out of the box, shocking. PT2 traced into helion and emitted it as a user defined triton kernel: P1836496774 In the long run, we do not actually want this, but rather to create a helion HOP so we can do fusions etc. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155513 Approved by: https://github.com/zou3519, https://github.com/jansel	2025-06-11 07:08:06 +00:00
bobrenjc93	5b9db4335e	Include c++ stack traces when we hit constraint violation (#155603 ) Example new error message ``` torch.fx.experimental.symbolic_shapes.ConstraintViolationError: Constraints violated (L['x'].size()[0])! For more information, run with TORCH_LOGS="+dynamic". - You marked L['x'].size()[0] as dynamic but your code specialized it to be a constant (5). Either remove the mark_dynamic or use a less strict API such as maybe_mark_dynamic or Dim.AUTO. Framework stack: File "??", line 0, in _start File "", line 0, in __libc_start_main_alias_2 File "??", line 0, in __libc_start_call_main File "/usr/local/src/conda/python-3.10.16/Modules/main.c", line 1094, in Py_BytesMain File "/usr/local/src/conda/python-3.10.16/Modules/main.c", line 357, in pymain_run_file_obj File "/usr/local/src/conda/python-3.10.16/Python/pythonrun.c", line 90, in _PyRun_AnyFileObject File "/usr/local/src/conda/python-3.10.16/Python/pythonrun.c", line 456, in _PyRun_SimpleFileObject File "/usr/local/src/conda/python-3.10.16/Python/pythonrun.c", line 1208, in pyrun_file File "/usr/local/src/conda/python-3.10.16/Python/pythonrun.c", line 1312, in run_mod File "/usr/local/src/conda/python-3.10.16/Python/pythonrun.c", line 1291, in run_eval_code_obj File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 1134, in PyEval_EvalCode File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/scratch/repro.py", line 9, in <module> foo(x) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/eval_frame.py", line 699, in compile_wrapper return fn(args, kwargs) File "offloadstuff.c", line 0, in dynamo__custom_eval_frame File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 305, in _PyObject_Call File "/usr/local/src/conda/python-3.10.16/Objects/typeobject.c", line 7494, in slot_tp_call File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 431, in _PyObject_Call_Prepend File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/convert_frame.py", line 1469, in __call__ return self._torchdynamo_orig_callable( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 112, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 215, in _PyObject_MakeTpCall File "/usr/local/src/conda/python-3.10.16/Objects/typeobject.c", line 7494, in slot_tp_call File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 431, in _PyObject_Call_Prepend File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 153, in _PyObject_FastCallDictTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/convert_frame.py", line 1248, in __call__ result = self._inner_convert( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 112, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 215, in _PyObject_MakeTpCall File "/usr/local/src/conda/python-3.10.16/Objects/typeobject.c", line 7494, in slot_tp_call File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 431, in _PyObject_Call_Prepend File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 153, in _PyObject_FastCallDictTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/convert_frame.py", line 625, in __call__ return _compile( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/convert_frame.py", line 1092, in _compile guarded_code = compile_inner(code, one_graph, hooks, transform) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_utils_internal.py", line 97, in wrapper_function return function(args, *kwargs) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/convert_frame.py", line 779, in compile_inner return _compile_inner(code, one_graph, hooks, transform) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/convert_frame.py", line 818, in _compile_inner out_code = transform_code_object(code, transform) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/bytecode_transformation.py", line 1424, in transform_code_object transformations(instructions, code_options) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/convert_frame.py", line 265, in _fn return fn(args, kwargs) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/convert_frame.py", line 743, in transform tracer.run() File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/symbolic_convert.py", line 3531, in run super().run() File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/symbolic_convert.py", line 1359, in run while self.step(): File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/symbolic_convert.py", line 1263, in step self.dispatch_table[inst.opcode](self, inst) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/symbolic_convert.py", line 422, in impl self.push(fn_var.call_function(self, self.popn(nargs), {})) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/variables/builtin.py", line 1160, in call_function return handler(tx, args, kwargs) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/variables/builtin.py", line 792, in <lambda> return lambda tx, args, kwargs: obj.call_function( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/variables/builtin.py", line 1160, in call_function return handler(tx, args, kwargs) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/variables/builtin.py", line 1120, in _handle_insert_op_in_graph return wrap_fx_proxy(tx, proxy) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/variables/builder.py", line 2500, in wrap_fx_proxy return wrap_fx_proxy_cls(target_cls=TensorVariable, kwargs) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 267, in PyVectorcall_Call File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/variables/builder.py", line 2566, in wrap_fx_proxy_cls return _wrap_fx_proxy( File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/variables/builder.py", line 2664, in _wrap_fx_proxy example_value = get_fake_value(proxy.node, tx, allow_non_graph_fake=True) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/utils.py", line 3205, in get_fake_value ret_val = wrap_fake_exception( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/utils.py", line 2705, in wrap_fake_exception return fn() File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/utils.py", line 3206, in <lambda> lambda: run_node(tx.output, node, args, kwargs, nnmodule) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_dynamo/utils.py", line 3373, in run_node return node.target(args, kwargs) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5917, in do_call_core File "/usr/local/src/conda/python-3.10.16/Objects/methodobject.c", line 430, in cfunction_vectorcall_FASTCALL File "/usr/local/src/conda/python-3.10.16/Objects/abstract.c", line 891, in binary_op1 File "/usr/local/src/conda/python-3.10.16/Objects/typeobject.c", line 7284, in slot_nb_multiply File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Objects/descrobject.c", line 344, in method_vectorcall_VARARGS_KEYWORDS File "python_variable_methods.cpp", line 0, in _object torch::autograd::TypeError_to_NotImplemented_<&torch::autograd::THPVariable_mul>(_object, _object, _object) File "python_variable_methods.cpp", line 0, in torch::autograd::THPVariable_mul(_object, _object, _object) File "??", line 0, in at::_ops::mul_Tensor::call(at::Tensor const&, at::Tensor const&) File "offloadstuff.c", line 0, in c10::impl::BoxedKernelWrapper<at::Tensor (at::Tensor const&, at::Tensor const&), void>::call(c10::BoxedKernel const&, c10::OperatorHandle const&, c10::DispatchKeySet, at::Tensor const&, at::Tensor const&) File "PyInterpreter.cpp", line 0, in torch::detail::(anonymous namespace)::ConcretePyInterpreterVTable::python_dispatcher(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >) const File "offloadstuff.c", line 0, in c10::OperatorHandle::callBoxedForDispatchKey(c10::DispatchKey, std::vector<c10::IValue, std::allocator<c10::IValue> >&) const File "PythonFallbackKernel.cpp", line 0, in void c10::BoxedKernel::make_boxed_function<&(anonymous namespace)::pythonTLSSnapshotFallback>(c10::OperatorKernel, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >) File "PyInterpreter.cpp", line 0, in torch::detail::(anonymous namespace)::ConcretePyInterpreterVTable::python_dispatcher(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >) const File "offloadstuff.c", line 0, in c10::OperatorHandle::callBoxedForDispatchKey(c10::DispatchKey, std::vector<c10::IValue, std::allocator<c10::IValue> >&) const File "VariableType_0.cpp", line 0, in c10::impl::make_boxed_from_unboxed_functor<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (c10::DispatchKeySet, at::Tensor const&, at::Tensor const&), &torch::autograd::VariableType::(anonymous namespace)::mul_Tensor>, at::Tensor, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, at::Tensor const&> >, false>::call(c10::OperatorKernel, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >) File "VariableType_0.cpp", line 0, in torch::autograd::VariableType::(anonymous namespace)::mul_Tensor(c10::DispatchKeySet, at::Tensor const&, at::Tensor const&) File "??", line 0, in at::_ops::mul_Tensor::redispatch(c10::DispatchKeySet, at::Tensor const&, at::Tensor const&) File "offloadstuff.c", line 0, in c10::impl::BoxedKernelWrapper<at::Tensor (at::Tensor const&, at::Tensor const&), void>::call(c10::BoxedKernel const&, c10::OperatorHandle const&, c10::DispatchKeySet, at::Tensor const&, at::Tensor const&) File "PyInterpreter.cpp", line 0, in torch::detail::(anonymous namespace)::ConcretePyInterpreterVTable::python_dispatcher(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >) const File "offloadstuff.c", line 0, in c10::OperatorHandle::callBoxedForDispatchKey(c10::DispatchKey, std::vector<c10::IValue, std::allocator<c10::IValue> >&) const File "PythonFallbackKernel.cpp", line 0, in (anonymous namespace)::pythonFallback(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >) File "PyInterpreter.cpp", line 0, in torch::detail::(anonymous namespace)::ConcretePyInterpreterVTable::dispatch(c10::OperatorHandle const&, std::vector<c10::IValue, std::allocator<c10::IValue> >) const File "??", line 0, in torch::handle_torch_function_no_python_arg_parser(c10::ArrayRef<_object>, _object, _object, char const, _object, char const, torch::TorchFunctionName) File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 577, in PyObject_CallMethod File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/utils/_stats.py", line 27, in wrapper return fn(args, *kwargs) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_subclasses/fake_tensor.py", line 1346, in __torch_dispatch__ return self.dispatch(func, types, args, kwargs) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_subclasses/fake_tensor.py", line 2029, in dispatch return self._cached_dispatch_impl(func, types, args, kwargs) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_subclasses/fake_tensor.py", line 1442, in _cached_dispatch_impl return self._dispatch_impl(func, types, args, kwargs) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_subclasses/fake_tensor.py", line 2552, in _dispatch_impl return maybe_propagate_real_tensors(fast_impl(self, args, *kwargs)) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_subclasses/fake_impls.py", line 956, in fast_binary_impl final_shape = infer_size(final_shape, shape) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/_subclasses/fake_impls.py", line 916, in infer_size torch._check( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/__init__.py", line 1669, in _check _check_with(RuntimeError, cond, message) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/__init__.py", line 1632, in _check_with if expect_true(cond): File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/symbolic_shapes.py", line 1686, in expect_true return a.node.expect_true( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/sym_node.py", line 552, in expect_true return self.guard_bool(file, line) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/sym_node.py", line 536, in guard_bool r = self.evaluate() File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/sym_node.py", line 510, in evaluate return self.shape_env.evaluate_sym_node(self, size_oblivious) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/symbolic_shapes.py", line 7113, in evaluate_sym_node return self.evaluate_expr( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 112, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 215, in _PyObject_MakeTpCall File "/usr/local/src/conda/python-3.10.16/Modules/_functoolsmodule.c", line 1020, in bounded_lru_cache_wrapper File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 267, in PyVectorcall_Call File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/recording.py", line 272, in wrapper return retlog(fn(args, *kwargs)) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 267, in PyVectorcall_Call File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/symbolic_shapes.py", line 7215, in evaluate_expr return self._inner_evaluate_expr( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 112, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 215, in _PyObject_MakeTpCall File "/usr/local/src/conda/python-3.10.16/Modules/_functoolsmodule.c", line 1020, in bounded_lru_cache_wrapper File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/recording.py", line 272, in wrapper return retlog(fn(args, *kwargs)) File "/usr/local/src/conda/python-3.10.16/Python/ceval.c", line 5945, in do_call_core File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/symbolic_shapes.py", line 7238, in _inner_evaluate_expr return self._evaluate_expr( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/symbolic_shapes.py", line 7505, in _evaluate_expr self._maybe_guard_rel(g) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 112, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 215, in _PyObject_MakeTpCall File "/usr/local/src/conda/python-3.10.16/Modules/_functoolsmodule.c", line 1020, in bounded_lru_cache_wrapper File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/symbolic_shapes.py", line 6758, in _maybe_guard_rel self._refine_ranges(expr) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/symbolic_shapes.py", line 7709, in _refine_ranges self._set_replacement( File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/fx/experimental/symbolic_shapes.py", line 6667, in _set_replacement self.framework_specialization_stacks[source] = CapturedTraceback.extract(cpp=True) File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 114, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Include/internal/pycore_ceval.h", line 46, in _PyEval_EvalFrame File "/home/bobren/local/a/pytorch/torch/utils/_traceback.py", line 207, in extract torch._C._profiler.gather_traceback(python=True, script=script, cpp=cpp), File "/usr/local/src/conda/python-3.10.16/Include/cpython/abstract.h", line 112, in _PyObject_VectorcallTstate File "/usr/local/src/conda/python-3.10.16/Objects/call.c", line 215, in _PyObject_MakeTpCall File "/usr/local/src/conda/python-3.10.16/Objects/methodobject.c", line 543, in cfunction_call File "offloadstuff.c", line 0, in pybind11::cpp_function::dispatcher(_object, _object, _object) File "offloadstuff.c", line 0, in pybind11::cpp_function::initialize<std::shared_ptr<torch::CapturedTraceback> (&)(bool, bool, bool), std::shared_ptr<torch::CapturedTraceback>, bool, bool, bool, pybind11::name, pybind11::scope, pybind11::sibling, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v>(std::shared_ptr<torch::CapturedTraceback> (&)(bool, bool, bool), std::shared_ptr<torch::CapturedTraceback> ()(bool, bool, bool), pybind11::name const&, pybind11::scope const&, pybind11::sibling const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&)::{lambda(pybind11::detail::function_call&)#3}::_FUN(pybind11::detail::function_call&) File "??", line 0, in torch::CapturedTraceback::gather(bool, bool, bool) File "??", line 0, in torch::unwind::unwind() User stack: File "/home/bobren/local/a/pytorch/scratch/repro.py", line 5, in foo return torch.randn(5) x ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155603 Approved by: https://github.com/zou3519, https://github.com/cyyever ghstack dependencies: #155133	2025-06-11 05:00:36 +00:00
Yiming Zhou	84c14361c2	[ez][AOTI] Add test for std::nullopt return in custom op (#155636 ) Summary: As title. Follow up of https://github.com/pytorch/pytorch/pull/154286 Test Plan: buck2 run mode/dev-nosan caffe2/test/inductor:test_aot_inductor_custom_ops -- -r test_fn_with_optional_tensor_nullopt_output Rollback Plan: Differential Revision: D76378892 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155636 Approved by: https://github.com/zou3519, https://github.com/cyyever	2025-06-11 03:52:31 +00:00
zeshengzong	1e690b6c41	Replace TORCH_INTERNAL_ASSERT with TORCH_CHECK in set_history (#155453 ) Fixes #154357 ## Test Result ```bash >>> import torch >>> >>> x = torch.tensor(1, device=torch.device('cpu')) >>> y = torch.tensor([1.0, 2.0, 3.0], requires_grad=True) >>> z0 = (x.abs() * y).prod(dtype=torch.int16) Traceback (most recent call last): File "<stdin>", line 1, in <module> RuntimeError: Autograd not support dtype: Short ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155453 Approved by: https://github.com/albanD, https://github.com/soulitzer	2025-06-11 03:46:48 +00:00
Arsh Zahed	110ae0f433	Custom Op handle 1-element tuples (#155447 ) Fixes #150472 Modification of [PR 151408](https://github.com/pytorch/pytorch/pull/151408). This PR modifies the return parsing in `infer_schema` to handle the case of a Tuple with a single element. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155447 Approved by: https://github.com/bdhirsh, https://github.com/zou3519	2025-06-11 03:43:40 +00:00
Brian Hirsh	a2b0b2698d	inductor codecache: include private inductor configs in cache key (#153672 ) Fixes https://github.com/pytorch/torchtitan/issues/1185 It looks like inductor's logic to include inductor configs in the cache key skips configs with a leading underscore by default. This came up in torchtitan - there's an asyncTP pipelining pass in inductor gated by a private config, and by not caching on the config we were attempting to use asyncTP when we shouldn't be. I'm not sure how worried we should be on the blast radius of this change. On the one hand: (1) it technically fixes any silent correctness issues in the cache around any other private inductor configs (it looks like there are a few) (2) there is some risk that there are some "harmless" configs that we are now including in the key, which may increase false negatives. I do see that there is an explicit list for "configs we want to ignore for caching" (`_save_config_ignore`), so my hope is that all harmless configs are already encapsulated there. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153672 Approved by: https://github.com/oulgen	2025-06-11 01:33:24 +00:00
Jing Xu	5264f8cd8d	Add Intel GPU info collection to the collect env script (#137846 ) As title, add Intel GPU info collection to the collect env script Output examples: 1. CPU on Windows ``` C:\Users\user\miniforge3\envs\py310\lib\site-packages\torch\_subclasses\functional_tensor.py:279: UserWarning: Failed to initialize NumPy: No module named 'numpy' (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\pytorch\torch\csrc\utils\tensor_numpy.cpp:81.) cpu = _conversion_method_template(device=torch.device("cpu")) Collecting environment information... PyTorch version: 2.8.0.dev20250528+cpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Microsoft Windows 11 Enterprise (10.0.22631 64-bit) GCC version: Could not collect Clang version: Could not collect CMake version: Could not collect Libc version: N/A Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:06:35) [MSC v.1943 64 bit (AMD64)] (64-bit runtime) Python platform: Windows-10-10.0.22631-SP0 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: False HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Name: 12th Gen Intel(R) Core(TM) i7-1270P Manufacturer: GenuineIntel Family: 198 Architecture: 9 ProcessorType: 3 DeviceID: CPU0 CurrentClockSpeed: 1711 MaxClockSpeed: 2200 L2CacheSize: 9216 L2CacheSpeed: None Revision: None Versions of relevant libraries: [pip3] torch==2.8.0.dev20250528+cpu [conda] torch 2.8.0.dev20250528+cpu pypi_0 pypi ``` 2. XPU on Windows ``` Collecting environment information... PyTorch version: 2.8.0a0+gitef6306e Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Microsoft Windows 10 Pro (10.0.19045 64-bit) GCC version: (GCC) 13.1.0 Clang version: Could not collect CMake version: version 3.29.3 Libc version: N/A Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:06:35) [MSC v.1943 64 bit (AMD64)] (64-bit runtime) Python platform: Windows-10-10.0.19045-SP0 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: True XPU used to build PyTorch: 20250101 Intel GPU driver version: * 32.0.101.6795 (20250520000000.****+) Intel GPU models onboard: Intel(R) Arc(TM) A770 Graphics Intel GPU models detected: * [0] _XpuDeviceProperties(name='Intel(R) Arc(TM) A770 Graphics', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.33184', total_memory=15915MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=128, sub_group_sizes=[8 16 32], has_fp16=1, has_fp64=0, has_atomic64=1) HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: ---------------------- Name: Intel(R) Xeon(R) Platinum 8260 CPU @ 2.40GHz Manufacturer: GenuineIntel Family: 179 Architecture: 9 ProcessorType: 3 DeviceID: CPU0 CurrentClockSpeed: 2401 MaxClockSpeed: 2401 L2CacheSize: 24576 L2CacheSpeed: None Revision: 21767 ---------------------- Name: Intel(R) Xeon(R) Platinum 8260 CPU @ 2.40GHz Manufacturer: GenuineIntel Family: 179 Architecture: 9 ProcessorType: 3 DeviceID: CPU1 CurrentClockSpeed: 2200 MaxClockSpeed: 2401 L2CacheSize: 24576 L2CacheSpeed: None Revision: 21767 Versions of relevant libraries: [pip3] intel_extension_for_pytorch==2.8.10+gitb3ea3a1 [pip3] numpy==2.1.2 [pip3] optree==0.13.1 [pip3] pytorch-triton-xpu==3.3.1+gitb0e26b73 [pip3] torch==2.8.0a0+gitef6306e [conda] intel-extension-for-pytorch 2.8.10+gitb3ea3a1 pypi_0 pypi [conda] mkl 2025.1.0 pypi_0 pypi [conda] mkl-dpcpp 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-blas 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-datafitting 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-dft 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-lapack 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-rng 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-sparse 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-stats 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-vm 2025.1.0 pypi_0 pypi [conda] pytorch-triton-xpu 3.3.1+gitb0e26b73 pypi_0 pypi [conda] torch 2.8.0a0+gitef6306e pypi_0 pypi ``` 3. CPU on Linux ``` /opt/python/cp312-cp312/lib/python3.12/site-packages/torch/_subclasses/functional_tensor.py:279: UserWarning: Failed to initialize NumPy: No module named 'numpy' (Triggered internally at /pytorch/torch/csrc/utils/tensor_numpy.cpp:81.) cpu = _conversion_method_template(device=torch.device("cpu")) Collecting environment information... PyTorch version: 2.8.0.dev20250528+cpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: AlmaLinux 8.10 (Cerulean Leopard) (x86_64) GCC version: (GCC) 14.2.1 20250110 (Red Hat 14.2.1-7) Clang version: Could not collect CMake version: version 4.0.0 Libc version: glibc-2.28 Python version: 3.12.10 (main, Apr 19 2025, 05:03:56) [GCC 14.2.1 20250110 (Red Hat 14.2.1-7)] (64-bit runtime) Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.28 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: False HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Byte Order: Little Endian CPU(s): 88 On-line CPU(s) list: 0-87 Thread(s) per core: 2 Core(s) per socket: 22 Socket(s): 2 NUMA node(s): 2 Vendor ID: GenuineIntel CPU family: 6 Model: 85 Model name: Intel(R) Xeon(R) Gold 6238M CPU @ 2.10GHz Stepping: 7 CPU MHz: 1000.000 CPU max MHz: 3700.0000 CPU min MHz: 1000.0000 BogoMIPS: 4200.00 Virtualization: VT-x L1d cache: 32K L1i cache: 32K L2 cache: 1024K L3 cache: 30976K NUMA node0 CPU(s): 0-21,44-65 NUMA node1 CPU(s): 22-43,66-87 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 intel_ppin ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts vnmi pku ospke avx512_vnni md_clear flush_l1d arch_capabilities Versions of relevant libraries: [pip3] torch==2.8.0.dev20250528+cpu [conda] Could not collect ``` 5. XPU on Linux ``` Collecting environment information... PyTorch version: 2.8.0.dev20250516+xpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Ubuntu 22.04.4 LTS (x86_64) GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0 Clang version: Could not collect CMake version: version 3.31.6 Libc version: glibc-2.35 Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:19:12) [GCC 13.3.0] (64-bit runtime) Python platform: Linux-5.15.50-051550-generic-x86_64-with-glibc2.35 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: True XPU used to build PyTorch: 20250101 Intel GPU driver version: * intel_opencl: 24.39.31294.21-1032~22.04 * level_zero: 1.17.44.0-1022~22.04 Intel GPU models onboard: * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 Intel GPU models detected: * [0] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [1] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [2] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [3] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [4] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [5] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [6] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [7] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 52 bits physical, 57 bits virtual Byte Order: Little Endian CPU(s): 224 On-line CPU(s) list: 0-223 Vendor ID: GenuineIntel Model name: Intel(R) Xeon(R) Platinum 8480+ CPU family: 6 Model: 143 Thread(s) per core: 2 Core(s) per socket: 56 Socket(s): 2 Stepping: 6 CPU max MHz: 3800.0000 CPU min MHz: 800.0000 BogoMIPS: 4000.00 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cat_l2 cdp_l3 invpcid_single intel_ppin cdp_l2 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow vnmi flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm rdt_a avx512f avx512dq rdseed adx smap avx512ifma clflushopt clwb intel_pt avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local split_lock_detect avx_vnni avx512_bf16 wbnoinvd dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req avx512vbmi umip pku ospke waitpkg avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg tme avx512_vpopcntdq la57 rdpid bus_lock_detect cldemote movdiri movdir64b enqcmd fsrm md_clear serialize tsxldtrk pconfig arch_lbr avx512_fp16 flush_l1d arch_capabilities Virtualization: VT-x L1d cache: 5.3 MiB (112 instances) L1i cache: 3.5 MiB (112 instances) L2 cache: 224 MiB (112 instances) L3 cache: 210 MiB (2 instances) NUMA node(s): 2 NUMA node0 CPU(s): 0-55,112-167 NUMA node1 CPU(s): 56-111,168-223 Vulnerability Itlb multihit: Not affected Vulnerability L1tf: Not affected Vulnerability Mds: Not affected Vulnerability Meltdown: Not affected Vulnerability Mmio stale data: Not affected Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl and seccomp Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Vulnerability Spectre v2: Mitigation; Enhanced IBRS, IBPB conditional, RSB filling Vulnerability Srbds: Not affected Vulnerability Tsx async abort: Not affected Versions of relevant libraries: [pip3] numpy==2.2.5 [pip3] pytorch-triton-xpu==3.3.0+git0bcc8265 [pip3] torch==2.8.0.dev20250516+xpu [conda] mkl 2025.1.0 pypi_0 pypi [conda] numpy 2.2.5 pypi_0 pypi [conda] onemkl-sycl-blas 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-dft 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-lapack 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-rng 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-sparse 2025.1.0 pypi_0 pypi [conda] pytorch-triton-xpu 3.3.0+git0bcc8265 pypi_0 pypi [conda] torch 2.8.0.dev20250516+xpu pypi_0 pypi ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/137846 Approved by: https://github.com/guangyey, https://github.com/malfet Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com> Co-authored-by: Nikita Shulga <2453524+malfet@users.noreply.github.com>	2025-06-11 01:22:06 +00:00
Michael Lazos	3040ca6d0f	[Cutlass] Include fp8 headers in aoti cpp wrapper (#155173 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155173 Approved by: https://github.com/desertfire ghstack dependencies: #154829, #154835, #155195	2025-06-11 01:21:16 +00:00
soulitzer	1ed243f01c	Add missing attr access check for legacy autograd.Function (#155055 ) Fixes https://github.com/pytorch/pytorch/issues/154981 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155055 Approved by: https://github.com/albanD ghstack dependencies: #154509, #154852	2025-06-11 01:03:49 +00:00
Sidharth	5dd07c70e5	[dynamo] added github_cli to detect unimplemented_v2 calls (#155610 ) This PR adds the workflow of whenever a dev makes a PR that contains files under torch/_dynamo, we check for any unimplemented_v2() callsites and if any of them have been modified in some sort of way, the workflow fails and lists them exactly which callsites and let's them know what the command lines are to update the registry. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155610 Approved by: https://github.com/williamwen42	2025-06-11 00:40:56 +00:00
soulitzer	3580b8dde4	[BE] Mention debug=True in AC error messages (#155593 ) See https://github.com/pytorch/pytorch/issues/155171#issuecomment-2949415407 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155593 Approved by: https://github.com/janeyx99	2025-06-11 00:32:41 +00:00
Ankita George	dbec08bc1c	Changes to HFStorageWriter to support saving shards of tensors (#154742 ) (#155566 ) Summary: As we move towards supporting saving partial tensors natively with HFStorageWriter, there are some simple changes that need to be made to make this happen. - The current approach for distributed writes is that every rank has full tensors, but we split up the writing of these full tensors across all available ranks. We're removing this logic that was in the HFSavePlanner and instead assuming that every rank has a shard and saving every rank's local state - as a result we can probably remove the HFSavePlanner, but keeping it as a placeholder for now - the current naming of files doesn't support shards as its in the format "model-00001-of-00004.safetensors", but if every rank is writing the same file names they will overwrite eachother, so this adds a shard-00001 prefix, so that the rank files don't overwrite eachother - don't save the metadata file models.safetensors.index.json if sharding is enabled. This file expects a 1 to 1 ratio between tensor and filename, but this doesn't make sense in the sharded saving approach, so we can just get rid of this file - make the "fqn_to_file_index" map optional. This is to describe which files to save which tensors in, but if users don't want to provide this, we can just save all the tensors to one file. If they run into issues, they can choose how to split up their tensors to be more friendly with 5GB HF remote storage file size soft limit. Test Plan: test_hf_storage.py Reviewed By: saumishr Differential Revision: D75099862 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155566 Approved by: https://github.com/saumishr	2025-06-10 23:37:47 +00:00
Siddharth Kotapati	2161be8497	Move unary trig ops to metal kernels (#154465 ) Move inverse trig unary ops, sinh, & cosh to metal kernel Pull Request resolved: https://github.com/pytorch/pytorch/pull/154465 Approved by: https://github.com/malfet Co-authored-by: Nikita Shulga <2453524+malfet@users.noreply.github.com>	2025-06-10 22:56:59 +00:00
Joel Schlosser	c4b93e6579	Replace frame_traced_fn hook with get_traced_code() util (#155249 ) #153622 introduced a hook for getting the relevant code objects after frame tracing. The idea is to have vLLM use this instead of monkey-patching `inline_call_()` to determine the source code files to hash. Unfortunately, the hook runs too late; the vLLM backend needs access to the set of source code filenames while it's running. This PR replaces the newly-added hook with a utility function that a backend can call to get this information. I've made the change in vLLM and can verify that this allows the information to be queried at the right time. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155249 Approved by: https://github.com/zou3519	2025-06-10 22:40:58 +00:00
dolpm	8892b782a8	[nativert] move execution planner to torch (#155374 ) Summary: att Test Plan: ci Rollback Plan: Differential Revidsion: D76167093 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155374 Approved by: https://github.com/zhxchen17	2025-06-10 22:36:06 +00:00
fduwjj	ffc6cbfaf7	[symm_mem] Move all symm mem code into a dedicated folder (#155573 ) We arrive at a point when so many files are related to symmetric memory and files are scattered around in the cpp side. Let's first put all related code (symmetric memory related) into a separate folder. We can do further refactoring later if needed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155573 Approved by: https://github.com/fegin, https://github.com/d4l3k	2025-06-10 22:30:11 +00:00
Nikita Shulga	3e131f7779	[CI] Move tlparse to requirements files (#155601 ) Not sure why we had it that way to begin with Pull Request resolved: https://github.com/pytorch/pytorch/pull/155601 Approved by: https://github.com/seemethere ghstack dependencies: #155476, #155493	2025-06-10 22:25:47 +00:00
Jane Xu	94da4523ec	Disable foreach tests that depend on profiler for CUDA 12.6 (#155596 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155596 Approved by: https://github.com/clee2000, https://github.com/malfet	2025-06-10 22:21:06 +00:00
atalman	672ac2ec86	Reapply "Cleanup VS 2019 refs in pytorch (#145863 )" (#152613 ) (#155478 ) This reverts commit e4f2282. I believe fix PR was landed https://github.com/pytorch/pytorch/pull/153480 that triggered the revert. Hence this is reland. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155478 Approved by: https://github.com/malfet	2025-06-10 22:20:14 +00:00
PyTorch MergeBot	3b7c5e6fa5	Revert "[inductor][triton pin] TMA shim refactor & mm, mm_scaled_grouped support (#155182 )" This reverts commit b07725a9516028a485153c4b5356b3e33b990f81. Reverted https://github.com/pytorch/pytorch/pull/155182 on behalf of https://github.com/davidberard98 due to fails on triton 3.2 (internally) ([comment](https://github.com/pytorch/pytorch/pull/155182#issuecomment-2960664845))	2025-06-10 21:53:01 +00:00
Oleksandr Stashuk	d2f06d2b06	[logs] Change autotune data into separate items (#155525 ) Summary: Split the autotune data into multiple keys and items : this is better for storage of the data and easier querying. Test Plan: ``` TORCHINDUCTOR_MAX_AUTOTUNE=1 tlp buck run (sample) ``` Rollback Plan: Differential Revision: D76303514 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155525 Approved by: https://github.com/jamesjwu, https://github.com/masnesral	2025-06-10 21:47:07 +00:00
GdoongMathew	14f3639e09	Convert to .md: onnx_verification.rst, onnx.rst, package.rst, (#155556 ) Fixes https://github.com/pytorch/pytorch/issues/155031 * [onnx_verification.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/onnx_verification.rst) * [onnx.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/onnx.rst) * [package.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/package.rst) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155556 Approved by: https://github.com/AlannaBurke, https://github.com/sekyondaMeta	2025-06-10 21:40:40 +00:00
nirajkamalk	ae0f1f8984	Convert to markdown onnx rst (#155228 ) Fixes #155030 Converts the following files to MyST markdown and ensure that the doc tests are green: - [x] [onnx_dynamo_onnxruntime_backend.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/onnx_dynamo_onnxruntime_backend.rst) - [x] [onnx_dynamo.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/onnx_dynamo.rst) - [x] [onnx_ops.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/onnx_ops.rst) - [onnx_torchscript_supported_aten_ops.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/onnx_torchscript_supported_aten_ops.rst) - not changed as it is autogenerated - [onnx_torchscript.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/onnx_torchscript.rst) - fixed in #155390 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155228 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-10 21:33:07 +00:00
atalman	7a03b0d2ca	[BE] Remove CUDA 11 artifacts. Fix Check Binary workflow (#155555 ) Please see: https://github.com/pytorch/pytorch/issues/147383 1. Remove CUDA 11 build and test artifacts. One place CUDA 12.4 2. Fix Check Binary Workflow to use Stable Cuda version variable rather then hardcoded one Pull Request resolved: https://github.com/pytorch/pytorch/pull/155555 Approved by: https://github.com/malfet, https://github.com/Skylion007	2025-06-10 21:32:08 +00:00
PyTorch MergeBot	40fefe2871	Revert "[BE] Update cudnn to 9.10.1.4 (#155122 )" This reverts commit 73220d52fd67b5f4f5b15e0e0433e09733c93f31. Reverted https://github.com/pytorch/pytorch/pull/155122 on behalf of https://github.com/atalman due to wrong pr description ([comment](https://github.com/pytorch/pytorch/pull/155122#issuecomment-2960592004))	2025-06-10 21:13:18 +00:00
Alberto A. Gallegos	8a396c5635	DOC: Convert to markdown: torch.compiler_best_practices_for_backends.rst, torch.compiler_cudagraph_trees.rst, torch.compiler_custom_backends.rst, torch.compiler_dynamic_shapes.rst, torch.compiler_dynamo_deepdive.rst (#155137 ) Fixes #155037 [torch.compiler_best_practices_for_backends.rst](https://github.com/pytorch/pytorch/tree/main/docs/source/torch.compiler_best_practices_for_backends.rst) shows error 404 cc @svekars @sekyondaMeta @AlannaBurke Pull Request resolved: https://github.com/pytorch/pytorch/pull/155137 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-10 20:51:05 +00:00
jafraustro	01b8f5e685	Convert to markdown: testing.rst, threading_environment_variables.rst, torch_cuda_memory.rst, torch_environment_variables.rst, torch_nccl_environment_variables.rst (#155523 ) Fixes #155035 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155523 Approved by: https://github.com/AlannaBurke, https://github.com/svekars	2025-06-10 20:38:36 +00:00
Yidi Wu	545fbd58dc	[export] inline jit.scripted function in export (#155180 ) When we export a scripted function, we inline the original callable stored in "_torchdynamo_inline", this is the same strategy as torch.compile path. We do the same thing for script method, where a "\_\_wrapped\_\_" attribute points to the original callable in most cases. There are some corner cases we identified: top-level jit.scripted modules' method doesn't have a \_\_wrapped\_\_. In this case, we fall back to the original scripted approach. Maybe there're more such cases but need verification. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155180 Approved by: https://github.com/zou3519	2025-06-10 20:34:12 +00:00
PyTorch UpdateBot	a666cf3b38	[xla hash update] update the pinned xla hash (#154348 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned xla hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154348 Approved by: https://github.com/pytorchbot	2025-06-10 20:33:31 +00:00
Colin Peppler	c9404faacb	[refactor] is_known_channels_last_contiguous* -> definitely_channels_last_contiguous* (#155499 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155499 Approved by: https://github.com/laithsakka	2025-06-10 20:29:46 +00:00
Max Podkorytov	94763f5ca7	[ROCm][Inductor][CK] add kBatch as runtime parameter to CK-tile gemms (#155389 ) Similar to old-CK gemms ### Testing Rely on existing coverage in `test/inductor/test_ck_backend.py` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155389 Approved by: https://github.com/chenyang78	2025-06-10 20:25:02 +00:00
Catherine Lee	ab51a93737	[CI] Set PATH during build to include location of sccache wrapped nvcc (#155464 ) Sccache wasn't working for nvcc on jammy, so manually set the path to include where nvcc is I had problems with always making nvcc a wrapper in some inductor tests where I got ``` sccache: encountered fatal error sccache: error: PCH not supported by nvcc sccache: caused by: PCH not supported by nvcc ``` and I also got an error (only on clang) when trying to set CMAKE_CUDA_COMPILER_LAUNCHER to /opt/cache/bin/sccache or sccache ``` ccache: error: failed to execute compile sccache: caused by: Compiler not supported: "nvcc warning : Support for offline compilation for architectures prior to \'<compute/sm/lto>_75\' will be removed in a future release (Use -Wno-deprecated-gpu-targets to suppress warning).\nnvcc fatal : Failed to preprocess host compiler properties.\n" ``` Non jammy cuda jobs' docker images used a different dockerfile, which set CMAKE_CUDA_COMPILER_LAUNCHER `e895e9689c/.ci/docker/ubuntu-cuda/Dockerfile (L110)` Alt solution: Given that I only get the error on clang, I could set CMAKE_CUDA_COMPILER_LAUNCHER=sccache only when not using clang Setting CUDA_NVCC_EXECUTABLE doesn't fail but also doesn't result in cache hits/misses Pull Request resolved: https://github.com/pytorch/pytorch/pull/155464 Approved by: https://github.com/malfet, https://github.com/huydhn	2025-06-10 20:23:33 +00:00
Eddie Yan	35e8f2593c	[CUDA] Fix missing bounds check in `Softmax.cu` (#154778 ) Uncovered by @ngimel, same as change in #144009 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154778 Approved by: https://github.com/ngimel, https://github.com/cyyever, https://github.com/malfet	2025-06-10 20:03:54 +00:00
Aidyn-A	0ca2a79f5b	[TEST] Modernize test_sort_large (#155546 ) Since its introduction ~4 years ago, the test `test_sort_large` has always been deselected because it requires 200GB of CUDA memory. Now, as we do have GPUs this big, it gets selected, but fails with `var_mean` not being a member if `torch.Tensor` and `var_mean` accepting only floating point tensors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155546 Approved by: https://github.com/eqy	2025-06-10 19:59:12 +00:00
Catherine Lee	ea23eb4b98	[ez][CI] Reuse old whl: turn off on releases, add docs files to ok list (#155346 ) Add docs/*/.md and docs/*/.rst to files that are ok to reuse old whls Prevent using old whls on release branches Move check for changed files earlier to reduce api usage? Pull Request resolved: https://github.com/pytorch/pytorch/pull/155346 Approved by: https://github.com/malfet, https://github.com/huydhn	2025-06-10 19:57:40 +00:00
redwrasse	8a22551300	Fixes OpInfo gradient checks for ctc_loss (#154590 ) Fixes #67462 Re-enables `OpInfo` gradient checks for the restricted scenarios where the current `ctc_loss` implementation is accurate and consistent. The desired `ctc_loss` gradient behavior appears to be an ongoing discussion, see https://github.com/pytorch/pytorch/issues/52241. The `OpInfo` gradient checks can be updated if/as the underlying implementation advances. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154590 Approved by: https://github.com/soulitzer	2025-06-10 19:56:39 +00:00
Anthony Barbier	954ce94950	Add __main__ guards to quantization tests (#154728 ) This PR is part of a series attempting to re-submit https://github.com/pytorch/pytorch/pull/134592 as smaller PRs. In quantization tests: - Add and use a common raise_on_run_directly method for when a user runs a test file directly which should not be run this way. Print the file which the user should have run. - Raise a RuntimeError on tests which have been disabled (not run) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154728 Approved by: https://github.com/ezyang	2025-06-10 19:46:07 +00:00
Ryan Guo	07eb374e7e	[dynamo] Avoid unncessary caching source codegen (#155376 ) We only need to cache a source (e.g., `x.y.z`) into a temporary local if it's used multiple times in the codegen, otherwise we'd just be creating redundant `DUP` and `STORE_FAST tmp_...` instructions, which might degrade perf and definitely makes generated bytecode harder to read. Example: ```python import torch @torch.compile(backend="eager") def fn(x, y): return x + y fn(torch.ones(2), torch.ones(1)) ``` Original bytecode: ```verbatim [0/0] [__bytecode] 3 0 RESUME 0 [0/0] [__bytecode] [0/0] [__bytecode] 5 2 LOAD_FAST 0 (x) [0/0] [__bytecode] 4 LOAD_FAST 1 (y) [0/0] [__bytecode] 6 BINARY_OP 0 (+) [0/0] [__bytecode] 10 RETURN_VALUE ``` Modified bytecode (before this patch): ```verbatim [__bytecode] 3 0 RESUME 0 [__bytecode] 2 LOAD_GLOBAL 1 (NULL + __compiled_fn_1_578c8d9a_2a9b_4d15_bac7_267591cdee32) [__bytecode] 14 LOAD_FAST 0 (x) [__bytecode] 16 COPY 1 [__bytecode] 18 STORE_FAST 3 (tmp_1) [__bytecode] 20 LOAD_FAST 1 (y) [__bytecode] 22 COPY 1 [__bytecode] 24 STORE_FAST 4 (tmp_2) [__bytecode] 26 PRECALL 2 [__bytecode] 30 CALL 2 [__bytecode] 40 STORE_FAST 2 (graph_out_0) [__bytecode] 42 LOAD_FAST 2 (graph_out_0) [__bytecode] 44 LOAD_CONST 1 (0) [__bytecode] 46 BINARY_SUBSCR [__bytecode] 56 DELETE_FAST 2 (graph_out_0) [__bytecode] 58 RETURN_VALUE ``` Modified bytecode (after this patch): ```verbatim [__bytecode] 3 0 RESUME 0 [__bytecode] 2 LOAD_GLOBAL 1 (NULL + __compiled_fn_1_2c498af2_ce5c_49cb_abba_a0c7489b09ce) [__bytecode] 14 LOAD_FAST 0 (x) [__bytecode] 16 LOAD_FAST 1 (y) [__bytecode] 18 PRECALL 2 [__bytecode] 22 CALL 2 [__bytecode] 32 STORE_FAST 2 (graph_out_0) [__bytecode] 34 LOAD_FAST 2 (graph_out_0) [__bytecode] 36 LOAD_CONST 1 (0) [__bytecode] 38 BINARY_SUBSCR [__bytecode] 48 DELETE_FAST 2 (graph_out_0) [__bytecode] 50 RETURN_VALUE ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155376 Approved by: https://github.com/williamwen42	2025-06-10 19:38:15 +00:00
Manuel Candales	91ee9ee82d	[MPS][BE] Refactor round_decimals shader code to leverage new macro (#155316 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155316 Approved by: https://github.com/Skylion007, https://github.com/malfet ghstack dependencies: #155304	2025-06-10 19:29:57 +00:00
Anthony Barbier	b1b8e57cda	Add __main__ guards to ao tests (#154612 ) This is the first PR of a series in an attempt to get the content of #134592 merged as smaller PRs (Given that the original one was closed due to a lack of reviewers). This specific PR contains: - Add and use a common raise_on_run_directly method for when a user runs a test file directly which should not be run this way. Print the file which the user should have run. - Update ao tests. There will be follow up PRs to update the other test suites but I don't have permissions to create branches directly on pytorch/pytorch so I can't create a stack and therefore will have to create them one at the time. Cc @jerryzh168 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154612 Approved by: https://github.com/jcaip	2025-06-10 18:33:09 +00:00
Shunting Zhang	0b677560e6	[inductor] use int64 for large index (#154575 ) Split reduction may need add an extra mask to avoid invalid index. Previously we always uses torch.int32 dtype. That causes problem when the tensor numel exceeds 2^31. Fix https://github.com/pytorch/pytorch/issues/154168 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154575 Approved by: https://github.com/ngimel, https://github.com/jansel	2025-06-10 18:30:43 +00:00
Manuel Candales	0f47e76937	[MPS] Implement hardshrink metal kernel (#155304 ) Implements the forward and backward hardshrink operators as Metal kernels. In order to support the lambda parameter, we extend the `exec_unary_kernel` and `exec_binary_kernel` methods. Now they take an optional Scalar and an optional ScalarType argument. When the optional ScalarType is provided, it overrides the type of the Scalar. We add a new `REGISTER_UNARY_ALPHA_OP` macro, and modify the existing `REGISTER_BINARY_ALPHA_OP` to support the new feature. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155304 Approved by: https://github.com/malfet	2025-06-10 18:20:27 +00:00
PyTorch MergeBot	8347268edc	Revert "Make open device registration tests standalone (#153855 )" This reverts commit 8823138e47a3200c313f6bf2d21eb689d8150f39. Reverted https://github.com/pytorch/pytorch/pull/153855 on behalf of https://github.com/clee2000 due to causing some linux aarch64 tests to fail [GH job link](https://github.com/pytorch/pytorch/actions/runs/15566289293/job/43832373302) [HUD commit link](`8823138e47`), should be easy fix, rename in places where its mentioned, there might be more than just aarch64 though ([comment](https://github.com/pytorch/pytorch/pull/153855#issuecomment-2960191503))	2025-06-10 18:11:24 +00:00
Han, Chao1	cb9b479f4f	XPU enable XCCL by default (#154963 ) Enable USE_XCCL=ON by default when building PyTorch XPU binary, which is on par with NCCL for PyTorch CUDA binary build. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154963 Approved by: https://github.com/cyyever, https://github.com/guangyey, https://github.com/chuanqi129, https://github.com/EikanWang, https://github.com/malfet Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-06-10 17:56:13 +00:00
Sidharth	0b6c0898e6	[dynamo] added additional_info functionality (#155526 ) There is now functionality for the developer to add a --additional-info arg to the add and update dev terminal command to include any additional info the dev might want to remark about the graph break. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155526 Approved by: https://github.com/williamwen42	2025-06-10 17:40:50 +00:00
Joel Schlosser	8823138e47	Make open device registration tests standalone (#153855 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153855 Approved by: https://github.com/janeyx99	2025-06-10 17:33:26 +00:00
NikhilAPatel	c88e87f355	[Inductor] Set Triton Allocator Function For Use with New TMA API in Inductor (#155373 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155373 Approved by: https://github.com/davidberard98	2025-06-10 17:09:04 +00:00
Aaron Gokaslan	73220d52fd	[BE] Update cudnn to 9.10.1.4 (#155122 ) Follow up to #152782 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155122 Approved by: https://github.com/malfet, https://github.com/atalman, https://github.com/eqy	2025-06-10 16:59:00 +00:00
zhxchen17	38c4d05535	[precompile] Ensure @disable()-ed function won't trigger recompile from precompile bytecode. (#155363 ) In a precompiled bytecode, it looks like the following: ``` pre-graph bytecode ... compiled graph code ... post-graph bytecode ``` In pre-graph bytecode we have calls into helper functions like torch._dynamo.utils.call_size which will invoke @disable inside the bytecode. Normally torch.compile() will handle these frames fine, but for precompile we will load bytecode from a clean state of dynamo and we want a way to assert recompile never happen, so the current way to ensure this is by doing set_stance("fail_on_recompile") (open to any other idea to test this, but IMO this is the closest thing we have today). This approach doesn't work when util functions like call_size() is involved and this PR fixes a bunch of places to make sure "fail_on_recompile" can skip through the functions meant to be skipped during compilation. Differential Revision: [D76156867](https://our.internmc.facebook.com/intern/diff/D76156867/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155363 Approved by: https://github.com/jamesjwu, https://github.com/jansel ghstack dependencies: #155329	2025-06-10 16:13:38 +00:00
zhxchen17	ddee927f31	[precompile] Add low level C API to load precompiled dynamo code on functions. (#155329 ) While loading deserialized dynamo states back from disk, precompile will need a direct way to access ExtraState and populate guarded bytecode as cache entries. This diff adds two API at code level to load precompiled guard + bytecode entries. 1. _load_precompile_entry() will append an entry to a precompile entry list per code object. This precompile entry will be looked up before normal compiled entries. 2. _reset_precompile_entries() will clean up all the installed existing entries. This is useful to prevent a case where user call loading multiple times and explode the number of entries on the list. Differential Revision: [D76083247](https://our.internmc.facebook.com/intern/diff/D76083247/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155329 Approved by: https://github.com/jamesjwu, https://github.com/jansel	2025-06-10 16:13:38 +00:00
Nichols A. Romero	e8d29c45e0	[ROCm][TunableOp] Unit test to verify that there is only one kernel launch per PyTorch API invocation. (#155077 ) TunableOp UT covers breakage that was fixed in this PR: https://github.com/pytorch/pytorch/pull/153764 After tuning is complete, verify that there is only one kernel launch. for each PyTorch API invocation Pull Request resolved: https://github.com/pytorch/pytorch/pull/155077 Approved by: https://github.com/jeffdaily	2025-06-10 16:11:43 +00:00
Kazuaki Ishizaki	08d15d3ec1	[Docs] Convert to markdown: torch.compiler_troubleshooting.rst (#155514 ) Part of changes #155040 (parent PR #155120) Follow-up of #155351. I split the changes of `torch.compiler_troubleshooting.rst ` into #155351 and this PR due to the 2000-line limit in one PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155514 Approved by: https://github.com/svekars	2025-06-10 15:41:31 +00:00
PyTorch MergeBot	eb152ab1dd	Revert "Inductor logging + analysis of torch.profile (#149697 )" This reverts commit 060838c2312ad207c7afe2c86f8a484afea5f328. Reverted https://github.com/pytorch/pytorch/pull/149697 on behalf of https://github.com/clee2000 due to broke a bunch of tests internally D76299454, probably also broke rocm inductor/test_analysis.py::TestAnalysisCUDA::test_augment_trace_against_flop_counter_maxat0_cuda_float16 [GH job link](https://github.com/pytorch/pytorch/actions/runs/15545277599/job/43766911025) [HUD commit link](`060838c231`) ([comment](https://github.com/pytorch/pytorch/pull/149697#issuecomment-2959747153))	2025-06-10 15:38:40 +00:00
wengshiy	b44306d368	Add dont constant fold flag (#154945 ) For support https://github.com/pytorch/ao/issues/2228 > What we want to do now is to enable FP8 quantization in PyTorch. And similar as INT8 quantization, we need to insert quantize and dequantize ops into the graph. > > However we met problems with these q/dq ops both in the PyTorch core and Torchao. > > PyTorch core: > > The quantize_per_tensor op does not support FP8. We want to fix it via https://github.com/pytorch/pytorch/pull/153601. And as you commented, the op is deprecated. > Torchao: > > In the fusion pass in Inductor, we want to match the pattern fp8_weight -> torchao.dequantize_affine_float8 -> fp32_op and fuse it as fp8_weight -> weight_pack -> fp8_op. We have done so for INT8 PT2E quantization. However, the pattern matching pass is applied after a constant folding pass in Inductor: > `100ec0b34a/torch/_inductor/fx_passes/freezing_patterns.py (L69C1-L74C1)` > After constant_fold(gm), the pattern will be folded as fp32_weight -> fp32_op. Then the original pattern cannot be found any more and the FP8 semantics is lost since the pattern is entirely in fp32 now. > For INT8, the int8_weight -> quantized_decomposed.dequantize_per_channel -> fp32_op pattern won't be folded because we mark quantized_decomposed.dequantize_per_channel impure so that it won't be folded: `100ec0b34a/torch/_inductor/constant_folding.py (L139C1-L149C1)` . But for the torchao.dequantize_affine_float8, we cannot do this because > It is an op from Torchao, which is unknown to the constant folder > It is decomposed to smaller ops, so we cannot put it in the list as a single op. > So, we think an easy and short-term solution is to modify the ops in PyTorch core via https://github.com/pytorch/pytorch/pull/153601. > However, if we want to resolve the issue with Torchao, we need to > Add a method in the constant folder in Inductor to allow registration of impure ops Based on [Jansel‘s reply](https://github.com/pytorch/ao/issues/2228#issuecomment-2914560340), add dont constant fold flag on this patch Pull Request resolved: https://github.com/pytorch/pytorch/pull/154945 Approved by: https://github.com/jansel Co-authored-by: Jason Ansel <jansel@jansel.net>	2025-06-10 14:52:26 +00:00
Nikita Shulga	68f36683f0	[Testing] Add more models to MPSInductor tests (#155494 ) Enable all 46 HuggingFace models, only GPT2ForSequenceClassification fails to compile with a rather strange `array subscript is not an integer` error Pull Request resolved: https://github.com/pytorch/pytorch/pull/155494 Approved by: https://github.com/dcci ghstack dependencies: #155476, #155493	2025-06-10 14:40:59 +00:00
Xuanteng Huang	c8d39a1045	[docs] Add docstring indicating UB for converting inf to int (#154781 ) Fixes #154724 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154781 Approved by: https://github.com/malfet	2025-06-10 14:04:50 +00:00
PyTorch MergeBot	805297981a	Revert "[Testing] Add more models to MPSInductor tests (#155494 )" This reverts commit f154f9b3040369a7979d5de7acb6fe21433eda83. Reverted https://github.com/pytorch/pytorch/pull/155494 on behalf of https://github.com/malfet due to I'm blind ([comment](https://github.com/pytorch/pytorch/pull/155494#issuecomment-2959319787))	2025-06-10 13:45:32 +00:00
Aby Mathew C	e53ddaf1f6	Adapt dtensor tests to be device agnostic (#154840 ) ##MOTIVATION This PR includes minor changes to skip some unsupported tests on Intel Gaudi devices as well as to make some of the tests more device agnostic. Please refer to this RFC as well: https://github.com/pytorch/rfcs/pull/66 ##CHANGES - test_dtensor_compile.py : Make some of the tests device agnostic . ( Replace "cuda" hard codings with self.device_type) - test_dtensor.py and test_comm_mode_features.py: Skip some tests which are unsupported on Intel Gaudi devices. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154840 Approved by: https://github.com/EikanWang, https://github.com/guangyey, https://github.com/albanD	2025-06-10 12:43:16 +00:00
Nikita Shulga	f154f9b304	[Testing] Add more models to MPSInductor tests (#155494 ) Enable all hugging face models Pull Request resolved: https://github.com/pytorch/pytorch/pull/155494 Approved by: https://github.com/dcci ghstack dependencies: #155476, #155493	2025-06-10 12:30:38 +00:00
Vivek Nayak	75f258dd1f	Fix spelling mistake (#155495 ) Summary: Change "primtivies" to "primitives". Test Plan: n/a Rollback Plan: Differential Revision: D76229938 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155495 Approved by: https://github.com/angelayi, https://github.com/cyyever	2025-06-10 09:06:58 +00:00
Laith Sakka	a205e8fd73	Apply all replacements on backward graph args during inductor codegen. (#155469 ) Summary: temporary mitigation for https://github.com/pytorch/pytorch/issues/155468 Test Plan: NA Rollback Plan: Differential Revision: D76096355 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155469 Approved by: https://github.com/bobrenjc93	2025-06-10 08:56:18 +00:00
Wang, Chuanqi	5116293f7e	[XPU] Split triton version as 2 files to decouple triton version bump (#155313 ) Triton XPU shares its version file with the community one. When the community updates Triton version, it will temporarily break the XPU CI/CD because they use different repositories and commits. To decouple Triton version bumps between the community and XPU, we propose splitting the version into two separate files. Refer the latest community triton version bump PR #153117 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155313 Approved by: https://github.com/etaf, https://github.com/EikanWang, https://github.com/atalman	2025-06-10 08:49:03 +00:00
Xuehai Pan	2cdcd16e83	[Easy] update pip sources for CUDA in nightly pull tool (#149143 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/149143 Approved by: https://github.com/ezyang, https://github.com/cyyever ghstack dependencies: #145685	2025-06-10 08:07:30 +00:00
Xuehai Pan	0319044e92	[Easy] update pip sources for ROCm in nightly pull tool (#145685 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/145685 Approved by: https://github.com/ezyang	2025-06-10 08:07:30 +00:00
Wang, Chuanqi	9d2d227003	[CI] Fix XPU runner setup status issue (#155443 ) Flow with PR #155194, fix the timeout exit code issue refer https://github.com/pytorch/pytorch/actions/runs/15526078422/job/43706927778?pr=154962#step:3:74 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155443 Approved by: https://github.com/etaf, https://github.com/atalman, https://github.com/EikanWang	2025-06-10 08:06:37 +00:00
Michael Lazos	5dfe1787b5	[Inductor] Limit fusions to a node distance of 64 (#154688 ) fix for https://github.com/pytorch/pytorch/issues/154652 and https://fb.workplace.com/groups/1075192433118967/permalink/1484799079148049/ [window 128 dashboard run here w/ no regressions](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Sun%2C%2001%20Jun%202025%2006%3A38%3A41%20GMT&stopTime=Sun%2C%2008%20Jun%202025%2006%3A38%3A41%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cuda%20(a100)&lBranch=mlazos/fuse-window&lCommit=8576f00ebfa53567d7bddc89d9882df9eb990561&rBranch=main&rCommit=9d59b516e9b3026948918e3ff8c2ef55a33d13ad) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154688 Approved by: https://github.com/eellison, https://github.com/Raymo111	2025-06-10 07:32:23 +00:00
Edward Z. Yang	8b8684466a	Add a stub AGENTS.md for Codex (#155459 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/155459 Approved by: https://github.com/albanD, https://github.com/malfet	2025-06-10 07:20:21 +00:00
David Berard	b07725a951	[inductor][triton pin] TMA shim refactor & mm, mm_scaled_grouped support (#155182 ) Follow-up to #154858. Triton 3.4 will provide a different API for TMA compared to Triton 3.3; the TMA shim in triton_helpers dispatches to the correct API. First, this refactors the TMA shim to drop args that aren't supported from Triton 3.2 to Triton 3.4: in particular, strides (Triton 3.2 version doesn't accept non-contiguous inputs, so we just infer contiguous strides in Triton 3.4) and element_ty (Triton 3.4 doesn't support this arg, so in Triton 3.2 we just infer it from base_ptr). Second, this updates mm.py & mm_scaled_grouped.py to use the TMA shim. Differential Revision: [D76318784](https://our.internmc.facebook.com/intern/diff/D76318784) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155182 Approved by: https://github.com/drisspg	2025-06-10 06:48:42 +00:00
atalman	8153340d10	[CI/CD] Remove CUDA 11.8 builds (#155509 ) This removes CUDA 11.8 from CI/CD Please see: https://github.com/pytorch/pytorch/issues/147383 TODO: Will followup of cleaning CUDA 11.8 config from scripts Pull Request resolved: https://github.com/pytorch/pytorch/pull/155509 Approved by: https://github.com/cyyever, https://github.com/huydhn, https://github.com/malfet	2025-06-10 05:16:41 +00:00
Mikayla Gawarecki	671a9d175b	Add warning for module full backward hook when no input requires gradient (#155339 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155339 Approved by: https://github.com/Skylion007	2025-06-10 04:42:06 +00:00
Animesh Jain	e25ce0f928	[invoke_subgraph] Use eager input vals to constrain input strides (#155291 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155291 Approved by: https://github.com/ezyang, https://github.com/zou3519	2025-06-10 04:06:09 +00:00
PyTorch MergeBot	660695f11d	Revert "Move non inductor workflows cuda 12.6->cuda 12.8 (#155234 )" This reverts commit ede6ead8cd8e925cb093f2b3016342e645bd728d. Reverted https://github.com/pytorch/pytorch/pull/155234 on behalf of https://github.com/clee2000 due to causing a bunch of tests to fail? ex test_nn.py::TestNNDeviceTypeCUDA::test_variable_sequence_cuda_float32 [GH job link](https://github.com/pytorch/pytorch/actions/runs/15545607752/job/43773157441) [HUD commit link](`ede6ead8cd`), some of the failures attributed to broken trunk on friday seem real? ([comment](https://github.com/pytorch/pytorch/pull/155234#issuecomment-2957578769))	2025-06-10 03:34:36 +00:00
CaoE	76644c9ff5	Make require_contiguous require exact strides instead of stride order (#148424 ) Make `require_contiguous` require exact strides instead of stride order. Pull Request resolved: https://github.com/pytorch/pytorch/pull/148424 Approved by: https://github.com/eellison Co-authored-by: eellison <elias.ellison@gmail.com>	2025-06-10 02:36:18 +00:00
Nikita Shulga	b916d8a583	[Testing] Shard MacOS inductor perf tests (#155493 ) One dtype per shard, as current job takes 2+ hours to finish Pull Request resolved: https://github.com/pytorch/pytorch/pull/155493 Approved by: https://github.com/dcci ghstack dependencies: #155476	2025-06-10 02:26:22 +00:00
Edward Z. Yang	52edfb2cbc	Updates to README about CUDA install dir and conda not required (#155458 ) Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/155458 Approved by: https://github.com/malfet	2025-06-10 01:30:34 +00:00
Sahdev Zala	f34335bf33	Convert compiler rst files to markdown (#155335 ) Convert following compiler rst files to md file. torch.compiler_inductor_profiling.rst torch.compiler_ir.rst torch.compiler_nn_module.rst torch.compiler_performance_dashboard.rst torch.compiler_profiling_torch_compile.rst Fixes #155039 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155335 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-10 01:12:11 +00:00
Yiming Zhou	1851f50866	[AOTI] Add int return type support for custom op in proxy executor (#155465 ) Summary: When a custom op has int return type in its schema. The returned value will be specialized and such behaviour is different from a symint return type. This diff only added support for int return type. As the returned int will be specialized and fused into downstream kernels (if being used), we can simply skip the int return type in the proxy executor. Note that in the eager run, the returned int will be specialized to the value defined in the real impl of the custom op. In exported program or in AOTI, the returned int will be specialized to the value defined in the fake impl of the custom op. So the definitions of the return value should be consistent across real and fake impl of the custom op. Otherwise the eager run and AOTI run will have different results. Test Plan: ``` buck2 run mode/dev-nosan caffe2/test/inductor:test_aot_inductor_custom_ops -- -r test_fn_with_int_output ``` Rollback Plan: Differential Revision: D76159406 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155465 Approved by: https://github.com/angelayi	2025-06-10 01:07:15 +00:00
angelayi	da50835bde	[aoti] Support c10 calls (#155256 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155256 Approved by: https://github.com/malfet	2025-06-10 00:45:59 +00:00
Ting Lu	07e340e29c	Build magma-cuda 129 (#155496 ) followup for https://github.com/pytorch/pytorch/pull/155340 https://github.com/pytorch/pytorch/issues/155196 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155496 Approved by: https://github.com/atalman	2025-06-10 00:32:24 +00:00
Kurt Mohler	e7698ff5cf	[MPS] Move abs op to Metal (#155474 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155474 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-06-10 00:23:59 +00:00
PyTorch MergeBot	7a48cc6990	Revert "[cuBLASLt][cuBLAS] Support 2D bias and `beta != 1.0` in cuBLASLt (#154170 )" This reverts commit b8bc2c2660e84034ff15232e2161e3ef9a6656d0. Reverted https://github.com/pytorch/pytorch/pull/154170 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it starts failing on ROCm ([comment](https://github.com/pytorch/pytorch/pull/154170#issuecomment-2957346976))	2025-06-10 00:18:23 +00:00
David Berard	a9a0501ec4	[user triton] mutation analysis for on-device TMA (#155380 ) Previously, the user-defined triton kernel mutation analysis would not detect mutation caused by TMA store, if the TMA descriptor was created via on-device TMA creation. This PR adds partial support for mutation analysis on programs that do stores via on-device TMA. On-device TMA works like this: ``` @triton.jit def kernel(A_ptr, workspace_ptr, ...): tl.extra.cuda.experimental_device_tensormap_create2d(workspace_ptr, A_ptr, ...) tl._experimental_descriptor_store(workspace_ptr, data, ...) ``` The first call (tensormap_create2d) mutates the contents of workspace_ptr to contain a data (including the fact that this TMA descriptor points to A_ptr). The second call (experimental_descriptor_store) writes to the location specified by the data in workspace_ptr: A_ptr, in this case. The approach here is to do a first pass to identify all the experimental_descriptor_stores (and collect the associated descriptor values); and then during mutation analysis, any tma creation on a mutated descriptor value (e.g. on `workspace_ptr` in the above example) will actually register as a mutation to the associated data pointer (e.g. `data` in the above example). Consider this example, which I'll used to describe the pros/cons of this approach. ``` @triton.jit def create_tma(global_ptr, workspace_ptr): tl.extra.cuda.experimental_device_tensormap_create2d(workspace_ptr, global_ptr, ...) @triton.jit def kernel(A, B, workspace_ptr): create_tma(A, workspace_ptr) workspace_B = workspace_ptr + 128 create_tma(B, workspace_B) data = tl._experimental_descriptor_load(workspace_ptr, ...) tl._experimental_descriptor_store(workspace_B, data, ...) ``` An alternative approach could be to simply modify the `tl.extra.cuda.experimental_device_tensormap_create2d` so that it returns a descriptor, and to use that descriptor in subsequent uses (i.e. to "functionalize" the uses of the tma creation API). However, this would (a) require "functionalization" through any function calls (e.g. to `create_tma`), and (b) would lead to both `A` and `B` being marked as mutated (i.e. mutation to `workspace_B` -> mutation to `workspace_ptr` -> mutation to `A`). A downside of the current approach is that it doesn't understand offsets into workspaces. e.g. if one were to recompute workspace_B instead of reusing the variable, the analysis pass would not understand that these values point to the same descriptor. Differential Revision: [D76175117](https://our.internmc.facebook.com/intern/diff/D76175117) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155380 Approved by: https://github.com/oulgen	2025-06-10 00:07:18 +00:00
Anthony Barbier	2578796e23	Fix sqlite3 in x86 Docker container. (#155211 ) Some core modules for versions of python installed in /opt depend on libraries in /usr/local but those libraries are not copied over from the base container. For example: /opt/python/cp312-cp312/bin/python3 -c "import sqlite3" ImportError: libsqlite3.so: cannot open shared object file: No such file or directory Pull Request resolved: https://github.com/pytorch/pytorch/pull/155211 Approved by: https://github.com/huydhn	2025-06-09 23:42:02 +00:00
Kazuaki Ishizaki	5df3bf13ec	[Docs] Convert to markdown: torch.compiler_troubleshooting.rst (#155351 ) Part of changes #155040 (parent PR #155120) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155351 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-09 23:18:31 +00:00
PyTorch MergeBot	e12597090c	Revert "Update auto-tuning support for _scaled_grouped_mm (#150944 )" This reverts commit 09328eb02f5412d2211b5fd638ce82d0e03b9c1f. Reverted https://github.com/pytorch/pytorch/pull/150944 on behalf of https://github.com/davidberard98 due to breaks internal usage & complicates triton pin update - more details in https://github.com/pytorch/pytorch/pull/150944#issuecomment-2957246463 ([comment](https://github.com/pytorch/pytorch/pull/150944#issuecomment-2957248841))	2025-06-09 23:12:56 +00:00
Michael Lazos	40d02eb481	[Cutlass] Allow filtering by fast_accum for scaled_mm (#155195 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155195 Approved by: https://github.com/drisspg ghstack dependencies: #154829, #154835	2025-06-09 22:46:18 +00:00
PyTorch MergeBot	2c1a93a0ae	Revert "[Graph Partition] move cpu scalar tensor to gpu (#154464 )" This reverts commit c1f531f0b0e6faf443d90f8de2936e866c8c27c2. Reverted https://github.com/pytorch/pytorch/pull/154464 on behalf of https://github.com/clee2000 due to some of the newly added tests are failing internally, along with some other tests, D75913292 ([comment](https://github.com/pytorch/pytorch/pull/154464#issuecomment-2957201054))	2025-06-09 22:43:20 +00:00
Qasim Khan	82e6475d92	Add doc for missing functions for torch.special module (#155074 ) Fixes #132178 Added all the missing functions that had a docstring but were not present in the documentation Pull Request resolved: https://github.com/pytorch/pytorch/pull/155074 Approved by: https://github.com/albanD	2025-06-09 22:28:26 +00:00
albanD	bdbf2792a8	Fix docs build (#155129 ) Not sure why the online doc build passes but it fails locally with these broken strings... ~Also pinning numpy version even though it is technically optional to ensure users have the right version as most users have numpy in their environment anyways.~ Pull Request resolved: https://github.com/pytorch/pytorch/pull/155129 Approved by: https://github.com/janeyx99, https://github.com/svekars	2025-06-09 22:25:20 +00:00
Nikita Shulga	034a7f6437	[BE] Raise better exception in `torch.[con]cat[enate]` (#155460 ) By replacing `TORCH_CHECK` with `TORCH_CHECK_VALUE` Also make redispatching from aliases an even simpler, by just calling respective original class Addresses feedback raised in https://github.com/pytorch/pytorch/pull/155383/files#r2133952368 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155460 Approved by: https://github.com/Skylion007, https://github.com/albanD	2025-06-09 22:18:00 +00:00
Ting Lu	398fca9dcf	Add almalinux CUDA 12.9 docker build, required for magma build (#155340 ) https://github.com/pytorch/pytorch/issues/155196 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155340 Approved by: https://github.com/cyyever, https://github.com/atalman	2025-06-09 22:10:24 +00:00
atalman	ede6ead8cd	Move non inductor workflows cuda 12.6->cuda 12.8 (#155234 ) Move non inductor workflows cuda 12.6->cuda 12.8 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155234 Approved by: https://github.com/Skylion007, https://github.com/zxiiro, https://github.com/cyyever, https://github.com/malfet	2025-06-09 22:04:19 +00:00
Gabriel Ferns	060838c231	Inductor logging + analysis of torch.profile (#149697 ) Prereqs: - https://github.com/pytorch/pytorch/pull/152708 Features: 1. Adds inductor's estimate of flops and bandwidth to the json trace events that perfetto uses. 1. Only use the tflops estimation from triton if we don't have the info from the datasheet because Triton's estimates are inaccurate. I have a backlog item to fix triton flops estimation upstream. New `DeviceInfo` class, and new function `get_device_tflops`. 1. New helpers `countable_fx` and `count_flops_fx` helps get the flops of an `fx.Node`. 1. Extends Triton `torch.profiler` logging to `DebugAutotuner`. 1. New script `profile_analysis.py`: `--augment_trace` adds perf estimates to any perfetto json trace, `--analyze` creates a summary table of these perf estimates, and `--diff` will compare two traces side by side: ```python Device(NVIDIA H100, 0): Kernel Name \| resnet Kernel Count \| resnet FLOPS \| resnet bw gbps \| resnet Dur (ms) \| resnet Achieved FLOPS % \| resnet Achieved Bandwidth % \| newresnet Kernel Count \| newresnet FLOPS \| newresnet bw gbps \| newresnet Dur (ms) \| newresnet Achieved FLOPS % \| newresnet Achieved Bandwidth % --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- triton_poi_fused__native_batch_norm_legi \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 triton_red_fused__native_batch_norm_legi \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 triton_poi_fused__native_batch_norm_legi \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 void cutlass::Kernel2<cutlass_80_tensoro \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 triton_red_fused__native_batch_norm_legi \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 void cutlass::Kernel2<cutlass_80_tensoro \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 triton_poi_fused__native_batch_norm_legi \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 triton_red_fused__native_batch_norm_legi \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 triton_per_fused__native_batch_norm_legi \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 triton_poi_fused__native_batch_norm_legi \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 triton_per_fused__native_batch_norm_legi \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 triton_poi_fused__native_batch_norm_legi \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 void cublasLt::splitKreduce_kernel<32, 1 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 triton_per_fused__native_batch_norm_legi \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 triton_poi_fused__native_batch_norm_legi \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 void cask_plugin_cudnn::xmma_cudnn::init \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 triton_per_fused__native_batch_norm_legi \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 triton_per_fused__native_batch_norm_legi \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 void cutlass::Kernel2<cutlass_80_tensoro \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 triton_poi_fused__native_batch_norm_legi \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 triton_poi_fused__native_batch_norm_legi \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 triton_poi_fused__native_batch_norm_legi \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 triton_per_fused__native_batch_norm_legi \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 triton_poi_fused__native_batch_norm_legi \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 triton_per_fused__native_batch_norm_legi \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 triton_poi_fused__native_batch_norm_legi \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 triton_per_fused__native_batch_norm_legi \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 std::enable_if<!(false), void>::type int \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 triton_poi_fused_add_copy__38 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 triton_poi_fused_convolution_0 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 triton_poi_fused_convolution_1 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 void convolve_common_engine_float_NHWC<f \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 triton_per_fused__native_batch_norm_legi \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 triton_per_fused__native_batch_norm_legi \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 triton_poi_fused__native_batch_norm_legi \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 void at::native::(anonymous namespace):: \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 void at::native::vectorized_elementwise_ \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/149697 Approved by: https://github.com/eellison, https://github.com/shunting314	2025-06-09 21:43:21 +00:00
Roy Hvaara	b95dadd717	[MPS] Enable RProp test for non-contiguous (#155439 ) I believe this issue has already been fixed, but I don't know the hero PR. I'm relying on ci signals to verify it's fixed across macOS versions. Fixes #118117 xref #115350 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155439 Approved by: https://github.com/Skylion007	2025-06-09 21:29:09 +00:00
Roy Hvaara	3490a4f906	[MPS] Enable optimizer tests affected by addcdiv (#155437 ) Tracked in #118115. Fixed in #124442. This PR unskips the tests. xref #115350 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155437 Approved by: https://github.com/Skylion007	2025-06-09 21:27:37 +00:00
Eddie Yan	b8bc2c2660	[cuBLASLt][cuBLAS] Support 2D bias and `beta != 1.0` in cuBLASLt (#154170 ) Fixes https://github.com/pytorch/pytorch/issues/153590 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154170 Approved by: https://github.com/malfet	2025-06-09 21:23:32 +00:00
Yiming Zhou	eba5fc91ac	[nativert] Move serialization to PyTorch core (#155229 ) Summary: Serialization contains utilities to deserialize a graph saved on disk in json format as defined in `torch/csrc/utils/generated_serialization_types.h` to the in-memory representation as defined in `torch/nativert/graph/Graph.h` Test Plan: buck2 run @mode/dev-nosan caffe2/test/cpp/nativert:serialization_test Rollback Plan: Differential Revision: D76012641 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155229 Approved by: https://github.com/zhxchen17	2025-06-09 21:12:30 +00:00
Max Podkorytov	1e6a653234	[ROCm][Inductor][CK] Split ck and ck-tile inductor backend(s) (#155294 ) ... and fix ck-tile instances not being generated due to incorrect caching ### Testing Added test cases for CKTILE instances ``` pytest test/inductor/test_ck_backend.py -k gemm_backends_CKTILE ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155294 Approved by: https://github.com/coconutruben	2025-06-09 20:40:26 +00:00
PyTorch MergeBot	620415e018	Revert "Add stack_trace on make_fx (#155155 )" This reverts commit d4d0ede6bacb4b3b33c0e4aa4cb0e79d34e697ec. Reverted https://github.com/pytorch/pytorch/pull/155155 on behalf of https://github.com/malfet due to Not sure why it was merged, it indeed breaks those tests in CI ([comment](https://github.com/pytorch/pytorch/pull/155155#issuecomment-2956973633))	2025-06-09 20:40:13 +00:00
Nikita Shulga	abbdf9f363	[BE][Testing] Unskip `ones_like`/`zeros_like` testing on MPS (#155476 ) But skip `double` dtype form OpInfo variants for this test Pull Request resolved: https://github.com/pytorch/pytorch/pull/155476 Approved by: https://github.com/Skylion007, https://github.com/dcci	2025-06-09 20:37:44 +00:00
eellison	ea37f72099	enable test (#155342 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155342 Approved by: https://github.com/Skylion007, https://github.com/bdhirsh ghstack dependencies: #154768	2025-06-09 19:26:05 +00:00
Shangdi Yu	d4d0ede6ba	Add stack_trace on make_fx (#155155 ) Summary: Previosuly, we only add stack trace in `class _ModuleStackTracer(PythonKeyTracer)` for non-strict export. I moved this stack trace logic to the parent class `PythonKeyTracer`, this way the graph traced from Module using make_fx will have stack_trace as well. Motivation: we've observed some uses cases where users first use `make_fx` on the Module, and then run `export` on the resulting graph. If the result of `make_fx` doesn't have stack trace, the stack trace information is lost. Test Plan: ``` buck run test:test_export -- -r test_stack_trace ``` Rollback Plan: Differential Revision: D75985427 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155155 Approved by: https://github.com/angelayi, https://github.com/zou3519	2025-06-09 18:31:57 +00:00
Justin Silver	2aade5ee9f	Fix weight tensor documentation #134896 (#155093 ) Fixes #134896 ## Description Remove line about 'weight' tensor needing to be of floating point type. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155093 Approved by: https://github.com/AlannaBurke	2025-06-09 18:07:21 +00:00
Aaron Gokaslan	3863bbb55b	[BE]: Update cusparselt to 0.7.1 (#155232 ) Needed to support sparse operations on Blackwell, and implements new features for the library. Also optimizes library sizes vs 0.7 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155232 Approved by: https://github.com/nWEIdia, https://github.com/malfet	2025-06-09 18:01:23 +00:00
PyTorch MergeBot	79bdafe5b6	Revert "Custom FX pass for inductor's backend registration (#154841 )" This reverts commit e694280d1215caf70f41575f2611bfa26c69ebdb. Reverted https://github.com/pytorch/pytorch/pull/154841 on behalf of https://github.com/clee2000 due to failing some tests internally D76135706 ([comment](https://github.com/pytorch/pytorch/pull/154841#issuecomment-2956357711))	2025-06-09 16:56:45 +00:00
IvanKobzarev	0083032e75	[aotd] Support mutations in reordering_to_mimic_autograd_engine (#155353 ) Original issue: https://github.com/pytorch/pytorch/issues/154820 Dedicated sub-issue: https://github.com/pytorch/pytorch/issues/155242 Backward graph is reordered by partitioners.py: reordering_to_mimic_autograd_engine Which only records in the backward graph compute that starts from tangents. Mutation of primals(inputs) in backward can be disconnected from backward. Handling this copy_ specifically, as we add this mutation in framework and this is the only mutation that exist. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155353 Approved by: https://github.com/bdhirsh, https://github.com/zou3519	2025-06-09 16:39:47 +00:00
Brian Hirsh	6c05f2fca0	[test] use JK to force graph break on slow aliasing/mutation/dynamic_shape behavior (#155257 ) Summary: test to unblock shampoo, needs cleanup Test Plan: CI Rollback Plan: steps: - jk.update: jk: pytorch/compiler:aliased_inputs_with_mutation_and_dyn_shapes_killswitch constant_bool: null consistent_pass_rate: null fractional_host_rollout: null sampling_rate: null - manual.note: content: Set it to false. Reviewed By: c00w Differential Revision: D76051868 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155257 Approved by: https://github.com/c00w	2025-06-09 16:21:59 +00:00
Cui, Yifeng	4a4cac0cef	Update torch-xpu-ops commit pin (#154962 ) Update the torch-xpu-ops commit to [intel/torch-xpu-ops@`a3a196`](`a3a196ccdb`) includes: - Enhanced Adaptive Average Pooling 2D Backward Kernel for performance and code simplification - Group Norm Backward Optimization with vectorization and parallel reduction - Support CL path for MaxUnpooling2d and MaxUnpooling3d - Rename USE_ONEMKL as USE_ONEMKL_XPU and set it as default ON - Refactor USE_XCCL & USE_C10D_XCCL option Pull Request resolved: https://github.com/pytorch/pytorch/pull/154962 Approved by: https://github.com/EikanWang	2025-06-09 15:54:13 +00:00
Sheng Fu	b9b84d8011	Generate unique id for tensor storage object by observing the week pointer of tensor storage object (#154859 ) Summary: PyTorch execution trace records tensor storage data in the trace. The tensor storage data includes storage id, offset, number of elements, and number of byte for each element. PARAM et-replay uses this information to allocate/free the tensors. However, the current implementation of generating tensor storage id does not guarantee it is unique. ExecutionTraceObserver maintains a lookup table to map the memory address of the tensor storage object to an unique id. If a new memory address is found, it will be put into that hash table and associate it to a new id. This implementation does not guarantee the storage object is unique since the memory that the address points to may be released and then re-allocated to a different tensor storage object. Test Plan: buck2 run mode/opt caffe2/test:test_profiler_cuda -- profiler.test_execution_trace.TestExecutionTraceCUDA Differential Revision: D75749065 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154859 Approved by: https://github.com/eellison, https://github.com/ngimel	2025-06-09 15:46:27 +00:00
Justin Chu	79aef14169	[ONNX] Set the name of the producing node using the value name (#155413 ) When comparing two graphs exported using different opset versions, even though the value names are the same in both graphs, the node names did not match, causing model-explorer to not be able to sync the two graphs. This change updates the names of the nodes that directly produce the output values, for better correspondence across exported graphs. ![image](https://github.com/user-attachments/assets/3c00ca18-221f-4add-8429-4bcf12069036) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155413 Approved by: https://github.com/cyyever, https://github.com/xadupre	2025-06-09 13:03:58 +00:00
Amandeep Chhabra	e15848669f	[1/n]adding torch.distributed.run option to provide destination for event logging (#154644 ) (#155268 ) Summary: Problem Statement Currently, torch distributed elastic does not support to an option specify destination for event logging from torch.distributed.run. recording events to default destination: https://fburl.com/code/7f9b0993 The default destination is "null". *Solution* adding option in torch.destributed.run to specify event_logging_destination. The default value will be "null" which is current default so it won;t affect users unless the specify it via command line. Test Plan: https://www.internalfb.com/mlhub/pipelines/runs/mast/f738408681-TrainingApplication_torch_distributed_run_3?job_attempt=0&version=0&tab=execution_details&env=PRODUCTION Rollback Plan: Reviewed By: kiukchung Differential Revision: D75183591 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155268 Approved by: https://github.com/d4l3k	2025-06-09 10:43:52 +00:00
Yuanhao Ji	9968c854b6	[Dynamo] Replace `unimplemented` with `unimplemented_v2` in `torch/_dynamo/variables/tensor.py` (#153146 ) Part of #147913 Replace `unimplemented` with`unimplemented_v2` in `torch/_dynamo/variables/tensor.py` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153146 Approved by: https://github.com/williamwen42 Co-authored-by: William Wen <william.wen42@gmail.com>	2025-06-09 06:27:50 +00:00
Yiming Zhou	9b4a748e29	[nativert] Move Weights to PyTorch core (#155156 ) Summary: Moves Weights class to PyTorch core Torch Native Runtime RFC: pytorch/rfcs#72 README: https://github.com/pytorch/pytorch/blob/main/torch/nativert/OVERVIEW.md Test Plan: buck2 run mode/dev-nosan caffe2/test/cpp/nativert:weights_test Differential Revision: D75973156 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155156 Approved by: https://github.com/zhxchen17	2025-06-09 05:49:32 +00:00
PyTorch MergeBot	6fb6293159	Revert "Add Intel GPU info collection to the collect env script (#137846 )" This reverts commit c6b4f98625bb6b22bb9a60112a6d58e684a97e1b. Reverted https://github.com/pytorch/pytorch/pull/137846 on behalf of https://github.com/etaf due to This is breaking tests on xpu, detail log: https://hud.pytorch.org/pr/pytorch/pytorch/154962#43700962849 ([comment](https://github.com/pytorch/pytorch/pull/137846#issuecomment-2954517883))	2025-06-09 03:13:27 +00:00
James Wu	be2ad70cfa	Fix dynamo tracing into AOTAutogradCache results in cpu tensors (#155251 ) On this line, we see that the bw_compiler that dynamo uses for AotAutograd automatically disables the backward runnable: `05dd638ee9/torch/_dynamo/backends/common.py (L76)` This disables dynamo in the bw_compiler but also disables the runnable the compiler returns. On a AOTAutogradCache hit, however, we never call the bw_compiler! So we don't disable dynamo properly. This only has an effect on certain cases of cpu tensors' backwards, where the backward is being done in python land, and dynamo unnecessarily tries to trace through the inductor generated code. It also only matters if the backward is being accessed outside of dynamo itself (say, in a graph break in eager mode), since dynamo properly disables the forward function already. ``` I0605 09:58:40.135000 3981970 torch/_dynamo/eval_frame.py:517] TorchDynamo attempted to trace the following frames: [ I0605 09:58:40.135000 3981970 torch/_dynamo/eval_frame.py:517] * fn /home/jjwu/test.py:9 I0605 09:58:40.135000 3981970 torch/_dynamo/eval_frame.py:517] * cast /data/users/jjwu/a/pytorch-env/lib/python3.10/typing.py:1737 I0605 09:58:40.135000 3981970 torch/_dynamo/eval_frame.py:517] * call /tmp/torchinductor_jjwu/rq/crq327nhoyjzog5n3qlchauucdrunrtutwmmoh7ipoe2ngnson5s.py:35 I0605 09:58:40.135000 3981970 torch/_dynamo/eval_frame.py:517] * fn /home/jjwu/test.py:9 I0605 09:58:40.135000 3981970 torch/_dynamo/eval_frame.py:517] * cast /data/users/jjwu/a/pytorch-env/lib/python3.10/typing.py:1737 I0605 09:58:40.135000 3981970 torch/_dynamo/eval_frame.py:517] * call /tmp/torchinductor_jjwu/rq/crq327nhoyjzog5n3qlchauucdrunrtutwmmoh7ipoe2ngnson5s.py:35 I0605 09:58:40.135000 3981970 torch/_dynamo/eval_frame.py:517] ] ``` This PR fixes the issue and adds a unit test showing that with or without cache hit, the frames dynamo is tracing is identical. Fixes https://github.com/pytorch/pytorch/issues/154536 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155251 Approved by: https://github.com/bdhirsh, https://github.com/anijain2305	2025-06-09 02:06:16 +00:00
Parag Ekbote	2908c10259	Document the default garbage_collection_threshold value and improve the organization of cuda docs (#155341 ) Fixes #150917 As mentioned in the issue, I've updated the documentation of `garbage_collection_threshold`and improved the organization. Could you please review? Pull Request resolved: https://github.com/pytorch/pytorch/pull/155341 Approved by: https://github.com/AlannaBurke, https://github.com/ngimel	2025-06-08 22:09:35 +00:00
Abhinav Tharamel	d41f62b7a0	Fix/issue #155027 (#155252 ) Fixes #155027 Converted RST files to Markdown Pull Request resolved: https://github.com/pytorch/pytorch/pull/155252 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-08 21:17:31 +00:00
Nikita Shulga	3d82a1dfb5	Add checks for empty tensor list (#155383 ) Vibe-coded with Codex, after collecting a backtrace, see https://chatgpt.com/s/cd_68438be8a1248191adbfa0a5f000e60b Even though, check for empty tensor list exists in `at::cat` crash might happens while resolving named dimension to position, by calling `dimname_to_position(tensors[0], dim)`, see backtrace below ``` (lldb) up frame #1: 0x00000001101146dc libtorch_cpu.dylib`at::TensorBase::has_names(this=0x0000000000000000) const at TensorBase.h:559:10 556 bool has_names() const { 557 // If a user is using unnamed tensors, then we can short-circuit right here. 558 // Otherwise, impl::has_names attempts to retrieve names. -> 559 if (!impl_->has_named_tensor_meta()) { 560 return false; 561 } 562 return impl::has_names(unsafeGetTensorImpl()); (lldb) up frame #2: 0x00000001101144c4 libtorch_cpu.dylib`at::dimname_to_position(tensor=0x0000000000000000, dim=Dimname @ 0x000000016fdfe348) at NamedTensorUtils.cpp:23:3 20 int64_t dimname_to_position(const Tensor& tensor, Dimname dim) { 21 TORCH_CHECK(dim.type() != NameType::WILDCARD, 22 "Please look up dimensions by name, got: name = None."); -> 23 TORCH_CHECK(tensor.has_names(), 24 "Name ", dim, " not found in ", toDimnameRepr(tensor), "."); 25 const auto names = tensor.names(); 26 ``` TODOs: - May be move test from `test_tensor_creation.py` to OpInfo (not sure which one is more readable) - Replace `TORCH_CHECK` with `TORCH_CHECK_VALUE` and adjust unit tests Fixes https://github.com/pytorch/pytorch/issues/155306 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155383 Approved by: https://github.com/cyyever, https://github.com/ezyang ghstack dependencies: #155382	2025-06-08 18:53:19 +00:00
PyTorch MergeBot	95448b2ce6	Revert "[Inductor] Improve typing, and prepare for ABI-compatible AOTI C-shim dispatching (#154371 )" This reverts commit 65b1aedd09e98fcafcdd893ca4924f4fa598fd18. Reverted https://github.com/pytorch/pytorch/pull/154371 on behalf of https://github.com/clee2000 due to see henry's comment above. This was reverted internally because it causes a memory leak and OOMs on AMD? ([comment](https://github.com/pytorch/pytorch/pull/154371#issuecomment-2954192879))	2025-06-08 17:37:29 +00:00
Narek Malkhasyan	30293b8b5e	Preserve Enum types during torch.export serialization and deserialization (#154821 ) Fixes #154674 Addresses an issue where `torch.export` does not correctly preserve Python `Enum` types during the save/load round-trip. Previously, Enum inputs were serialized by value only, causing their type to be lost after deserialization. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154821 Approved by: https://github.com/XuehaiPan, https://github.com/Skylion007, https://github.com/yushangdi, https://github.com/angelayi	2025-06-08 17:30:31 +00:00
PyTorch MergeBot	27df0c56b7	Revert "[inductor] use int64 for large index (#154575 )" This reverts commit 2596e3d0617852469241be8777cf46db5c83928c. Reverted https://github.com/pytorch/pytorch/pull/154575 on behalf of https://github.com/clee2000 due to broke inductor/test_op_dtype_prop.py::TestCaseCUDA::test_op_dtype_propagation_add_cuda_int32 [GH job link](https://github.com/pytorch/pytorch/actions/runs/15510656657/job/43673763835) [HUD commit link](`2596e3d061`), note for self: bad TD ([comment](https://github.com/pytorch/pytorch/pull/154575#issuecomment-2954175761))	2025-06-08 16:58:59 +00:00
Xuehai Pan	49888e6be0	[BE] Polish `Makefile` (#155425 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155425 Approved by: https://github.com/ezyang	2025-06-08 16:37:12 +00:00
Bob Ren	b981fb6744	Add docblock to torch/_dynamo/variables/builtin.py (#155402 ) Add comprehensive module docstring explaining built-in function and type variable tracking, including handling of Python built-ins, type constructors, operators, and special constructs during symbolic execution. Originally generated by claude but reviewed and edited by me. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155402 Approved by: https://github.com/Skylion007 ghstack dependencies: #155403	2025-06-08 15:24:29 +00:00
Aleksandar Samardžić	09328eb02f	Update auto-tuning support for _scaled_grouped_mm (#150944 ) 1. Enable strided inputs 2. Implement "2d/2d", "3d/2d" and "3d/3d" combinations of inputs 3. Fix non-TMA load variant 4. Replace experimental_device_tensormap_create2d with _experimental_make_tensor_descriptor 5. Fix cases when group size along K dimension is not multiple of block size along K 6. Updated meta registration 7. Update synthetic offsets creation Pull Request resolved: https://github.com/pytorch/pytorch/pull/150944 Approved by: https://github.com/ngimel	2025-06-08 10:18:13 +00:00
Bob Ren	1339e88105	Add docblock to torch/_dynamo/side_effects.py (#155403 ) Add comprehensive module docstring explaining side effect tracking and management, including mutation tracking, context changes, aliasing, and state preservation during symbolic execution. Originally generated by claude but reviewed and edited by me. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155403 Approved by: https://github.com/williamwen42	2025-06-08 07:02:30 +00:00
Bob Ren	0756ebcd48	Add docblock to torch/_dynamo/trace_rules.py (#155401 ) Add comprehensive module docstring explaining the tracing rules and policies that govern TorchDynamo's compilation decisions, including skip rules, inlining policies, and library-specific handling. Originally generated by claude but reviewed and edited by me. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155401 Approved by: https://github.com/williamwen42	2025-06-08 04:30:03 +00:00
Shivam Raikundalia	abf4da0d24	[Profiler] Induce Inductor Import before Profiling (#155243 ) Fixes #151829 Summary: Currently, inductor has a lazy init which causes certain aten ops to run during a profiling run. This ends up cluttering the function events especially for smaller traces. One of the attempts to fix this was to simply remove that import from the profiler entirely but it looks like the import happens somewhere downstream anyways and the event still flood our profile. To fix this, we induce the inductor import during prepare trace if the inductor is present. This way regardless of how the workload imports the inductor the actual init process will be done before tracing starts, resulting in more accurate tracing. Test Plan: Added test, also ran N7316820 manually and went from getting many events on the first run to the following output (only difference is Runtime Triggered Module Loading which is CUPTI overhead event): ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg Self CUDA Self CUDA % CUDA total CUDA time avg # of Calls aten::mul_ 1.40% 340.638us 99.92% 24.390ms 24.390ms 1.535us 100.00% 4.605us 4.605us 1 cudaLaunchKernel 0.60% 146.533us 98.52% 24.049ms 24.049ms 0.000us 0.00% 3.070us 3.070us 1 Runtime Triggered Module Loading 6.14% 1.500ms 6.14% 1.500ms 1.500ms 1.535us 100.00% 1.535us 1.535us 1 Runtime Triggered Module Loading 91.78% 22.403ms 91.78% 22.403ms 22.403ms 1.535us 100.00% 1.535us 1.535us 1 void at::native::vectorized_elementwise_kernel<4, at... 0.00% 0.000us 0.00% 0.000us 0.000us 1.535us 100.00% 1.535us 1.535us 1 cudaDeviceSynchronize 0.08% 20.031us 0.08% 20.031us 20.031us 0.000us 0.00% 0.000us 0.000us 1 ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg Self CUDA Self CUDA % CUDA total CUDA time avg # of Calls aten::mul_ 82.81% 484.396us 94.26% 551.378us 551.378us 1.440us 100.00% 1.440us 1.440us 1 cudaLaunchKernel 11.45% 66.982us 11.45% 66.982us 66.982us 0.000us 0.00% 0.000us 0.000us 1 void at::native::vectorized_elementwise_kernel<4, at... 0.00% 0.000us 0.00% 0.000us 0.000us 1.440us 100.00% 1.440us 1.440us 1 cudaDeviceSynchronize 5.74% 33.581us 5.74% 33.581us 33.581us 0.000us 0.00% 0.000us 0.000us 1 ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ Rollback Plan: Differential Revision: D76056511 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155243 Approved by: https://github.com/ngimel	2025-06-07 23:58:50 +00:00
Catherine Lee	f1f49e56b0	[CI] remove xfail sm89 job (#155244 ) No need to collect more data Pull Request resolved: https://github.com/pytorch/pytorch/pull/155244 Approved by: https://github.com/janeyx99, https://github.com/huydhn, https://github.com/Skylion007	2025-06-07 21:04:57 +00:00
Yuki Kobayashi	11bc29856d	Fix some incorrect reST markups in the document (#154831 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154831 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-06-07 19:09:46 +00:00
Shunting Zhang	2596e3d061	[inductor] use int64 for large index (#154575 ) Split reduction may need add an extra mask to avoid invalid index. Previously we always uses torch.int32 dtype. That causes problem when the tensor numel exceeds 2^31. Fix https://github.com/pytorch/pytorch/issues/154168 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154575 Approved by: https://github.com/ngimel, https://github.com/jansel	2025-06-07 18:41:46 +00:00
cyy	f6e18bc105	Fix CUDA 12.8 docker tag (#155087 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/155087 Approved by: https://github.com/nWEIdia, https://github.com/Skylion007	2025-06-07 16:39:42 +00:00
Jeff Daily	783a4c1f50	[ROCm] fix nightly wheel, second attempt (#155388 ) Fixes #155207. hipsparselt logic was still broken, but smoke test didn't catch Pull Request resolved: https://github.com/pytorch/pytorch/pull/155388 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-07 15:57:55 +00:00
Wei Wang	ab56e5add9	[CUDA][BUILD] Add back the capability to use env TORCH_CUDA_ARCH_LIST (#155314 ) Add back the capability to use env TORCH_CUDA_ARCH_LIST to control how downstream projects (which uses find_package (torch)) build. Follow up to: https://github.com/pytorch/pytorch/pull/152715 Before this PR, On a CPU only machine, building a downstream project would ignore the TORCH_CUDA_ARCH_LIST setting (if set) and go straight to the auto GPU detection mode, in which case there would be no GPU detected and an excessive list of cuda architectures may be used. This also means that there is no way to build a binary that would be targeting a different SM on the current machine a developer is using. After this PR, TORCH_CUDA_ARCH_LIST is effective for developers to control explicitly which SM architectures to build. p.s. I think this PR might have been the original intent of https://github.com/pytorch/pytorch/pull/152715 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155314 Approved by: https://github.com/janeyx99, https://github.com/eqy, https://github.com/atalman	2025-06-07 15:52:39 +00:00
bobrenjc93	456f40cb09	Add docblock for autotune_cache.py (#155133 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155133 Approved by: https://github.com/aorenste	2025-06-07 14:50:09 +00:00
xinan.lin	29e6033ff3	[Break XPU] Fix failed test cases which are introduced by community for XPU. (#155317 ) Fixes #155186, Fixes #154701 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155317 Approved by: https://github.com/jansel	2025-06-07 14:46:30 +00:00
Kshiteej K	694028f502	update get_default_device to also respect torch.device ctx manager (#148621 ) Fixes https://github.com/pytorch/pytorch/issues/131328 Pull Request resolved: https://github.com/pytorch/pytorch/pull/148621 Approved by: https://github.com/ezyang	2025-06-07 14:26:17 +00:00
Animesh Jain	db491825e0	[invoke_subgraph] Add logging (#155284 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155284 Approved by: https://github.com/zou3519 ghstack dependencies: #155270	2025-06-07 11:31:53 +00:00
Animesh Jain	0f3f59784d	[invoke_subgraph] Throw assertion on uncaptured speculate_subgraph (#155270 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155270 Approved by: https://github.com/zou3519	2025-06-07 11:31:53 +00:00
Boyuan Feng	c1f531f0b0	[Graph Partition] move cpu scalar tensor to gpu (#154464 ) cudagraph does not support cpu tensors. In this PR, we update the graph by explicitly moving cpu tensors to gpu when profitable, relying on graph partition to split off this data copy, and cudagraphifying the remaining gpu ops. This PR unblocked the graph partition + cudagraph on speech_transformer, leading to 39.5% speedup on inference [P1830602200](https://www.internalfb.com/phabricator/paste/view/P1830602200), 85% speedup on training [P1831115315](https://www.internalfb.com/phabricator/paste/view/P1831115315). Close: #119241 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154464 Approved by: https://github.com/eellison	2025-06-07 06:59:39 +00:00
Mengwei Liu	386aa72003	[BE] Cleanup old ExecuTorch codegen and runtime code (#154165 ) Summary: These files are added to pytorch/pytorch before ExecuTorch is opensourced. Now is a good time to remove it from pytorch/pytorch, since the code is moved to pytorch/executorch already. Test Plan: Rely on CI jobs. Differential Revision: [D75985423](https://our.internmc.facebook.com/intern/diff/D75985423) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154165 Approved by: https://github.com/kimishpatel, https://github.com/Skylion007, https://github.com/cyyever	2025-06-07 06:54:12 +00:00
dolpm	da1f8980df	[nativert] move function schema to torch (#154948 ) Summary: att Test Plan: ci Rollback Plan: Differential Revision: D75826905 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154948 Approved by: https://github.com/zhxchen17	2025-06-07 05:45:30 +00:00
Xiaodong Wang	5fbaa041e7	SDPA support gfx950 (#155103 ) Summary: Seems to run, just not the optimal performance. e.g. ck_tile doesn't have those gfx942 optimizations it seems https://github.com/ROCm/composable_kernel/blob/develop/include/ck_tile/ops/fmha/block/variants.hpp#L27 Test Plan: ``` +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| Batch Size \| Sequence Length \| Heads \| Head Dim \| Flash Time (µs) \| Mem Eff Time (µs) \| Math Time (µs) \| Flex Time (µs) \| xformers Time (µs) \| Flash TFlops \| Mem Eff TFlops \| Math TFlops \| Flex TFlops \| xformers TFlops \| Speedup (Flash/Math) \| Speedup (MemEff/Math) \| Speedup (Flex/Math) \| Speedup (xformers/Math) \| xformers trace_url \| Flash trace_url \| +==============+===================+=========+============+===================+=====================+==================+==================+======================+================+==================+===============+===============+===================+========================+=========================+=======================+===========================+======================+===================+ \| 1 \| 4096 \| 16 \| 64 \| 179.737 \| 182.874 \| 3106.6 \| 359.662 \| 205.506 \| 382.334 \| 375.776 \| 22.1205 \| 191.067 \| 334.392 \| 17.2841 \| 16.9877 \| 8.63754 \| 15.1169 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 4096 \| 32 \| 128 \| 617.271 \| 623.38 \| 7169.73 \| 998.961 \| 654.534 \| 445.312 \| 440.947 \| 38.3387 \| 275.164 \| 419.96 \| 11.6152 \| 11.5014 \| 7.17719 \| 10.9539 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 8192 \| 16 \| 64 \| 667.032 \| 670.118 \| 13031.8 \| 1383.42 \| 768.452 \| 412.091 \| 410.193 \| 21.0928 \| 198.694 \| 357.703 \| 19.5371 \| 19.4471 \| 9.42 \| 16.9586 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 8192 \| 32 \| 128 \| 2074.64 \| 2214.81 \| 29186.9 \| 3916.35 \| 2404.29 \| 529.978 \| 496.437 \| 37.6714 \| 280.749 \| 457.313 \| 14.0684 \| 13.1781 \| 7.45257 \| 12.1395 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 16384 \| 16 \| 64 \| 2456.6 \| 2472.38 \| 51095.8 \| 5647.01 \| 3008.09 \| 447.574 \| 444.718 \| 21.5186 \| 194.707 \| 365.518 \| 20.7994 \| 20.6666 \| 9.0483 \| 16.9861 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 16384 \| 32 \| 128 \| 8048.8 \| 8070.96 \| 113478 \| 15580.8 \| 9768.71 \| 546.423 \| 544.922 \| 38.7569 \| 282.274 \| 450.218 \| 14.0987 \| 14.06 \| 7.2832 \| 11.6165 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| Batch Size \| Sequence Length \| Heads \| Head Dim \| Flash Time (µs) \| Mem Eff Time (µs) \| Math Time (µs) \| Flex Time (µs) \| xformers Time (µs) \| Flash TFlops \| Mem Eff TFlops \| Math TFlops \| Flex TFlops \| xformers TFlops \| Speedup (Flash/Math) \| Speedup (MemEff/Math) \| Speedup (Flex/Math) \| Speedup (xformers/Math) \| xformers trace_url \| Flash trace_url \| +==============+===================+=========+============+===================+=====================+==================+==================+======================+================+==================+===============+===============+===================+========================+=========================+=======================+===========================+======================+===================+ \| 1 \| 4096 \| 16 \| 64 \| 692.323 \| 697.649 \| 4241.81 \| 1562.26 \| 906.441 \| 248.148 \| 246.254 \| 40.5012 \| 109.968 \| 189.531 \| 6.12693 \| 6.08015 \| 2.71518 \| 4.67963 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 4096 \| 32 \| 128 \| 2263.22 \| 2267.38 \| 9482.64 \| 7003.8 \| 2765.5 \| 303.636 \| 303.079 \| 72.4687 \| 98.1174 \| 248.489 \| 4.1899 \| 4.18221 \| 1.35393 \| 3.42891 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 8192 \| 16 \| 64 \| 2553.94 \| 2572.68 \| 15909.8 \| 5697.16 \| 3284.77 \| 269.073 \| 267.112 \| 43.193 \| 120.621 \| 209.206 \| 6.22953 \| 6.18415 \| 2.79259 \| 4.84352 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 8192 \| 32 \| 128 \| 8187.67 \| 8201.71 \| 35449.2 \| 26424.3 \| 10364.5 \| 335.722 \| 335.147 \| 77.5413 \| 104.025 \| 265.21 \| 4.32959 \| 4.32218 \| 1.34154 \| 3.42025 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 16384 \| 16 \| 64 \| 9948.15 \| 9815.47 \| 62815.1 \| 23741.9 \| 12710 \| 276.31 \| 280.046 \| 43.7598 \| 115.778 \| 216.269 \| 6.31425 \| 6.39961 \| 2.64575 \| 4.94217 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 16384 \| 32 \| 128 \| 32187.6 \| 32035.6 \| 137832 \| 102075 \| 40623.4 \| 341.595 \| 343.216 \| 79.7716 \| 107.716 \| 270.66 \| 4.28216 \| 4.30248 \| 1.35031 \| 3.39293 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ ``` Rollback Plan: Differential Rev,ision: D75934358 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155103 Approved by: https://github.com/jeffdaily, https://github.com/malfet	2025-06-07 03:47:29 +00:00
Stella Laurenzo	30387ab2e4	[ROCm] Adds initialization support for PyTorch when built from ROCm wheels. (#155285 ) AMD is beginning to roll out ROCm distribution via Python wheels. This patch adds the `__init__.py` hook that is necessary to bootstrap ROCm correctly on Linux and Windows when built from these wheels. See draft, developer documentation describing the mechanism here: https://github.com/ROCm/TheRock/blob/main/docs/packaging/python_packaging.md This operates to similar effect as how Torch can depend on CUDA wheels, with some differences: * ROCm libraries and checks are delegated to helpers in the `rocm_sdk` module, which knows how to find and configure access to the installed libraries. This limits the amount of plumbing and path machinations that must match up between the framework and ROCm. * When building torch against ROCm, no ROCm system install is needed: instead the proper SDK development wheel is installed and the `CMAKE_PREFIX_PATH` is obtained via `rocm-sdk path --cmake`. * It is expected that whoever produces such a build will also place a generated `_rocm_init.py` in the `torch` module with initialization logic to preload libraries, check versions, verify GPU compatibility, etc. * See [build_prod_wheels.py](https://github.com/ROCm/TheRock/blob/main/external-builds/pytorch/build_prod_wheels.py) for an example build script that is being used to generate nightlies in this configuration. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155285 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-07 02:59:03 +00:00
Nikita Shulga	f140fac8dc	[MPS] Implement erfc (#155382 ) And migrate `erf` to Metal kernel Use `erf` approximations from https://github.com/ml-explore/mlx/blob/main/mlx/backend/metal/kernels/erf.h as previous approximation did not match the CPU implementation After that, `erfc(x) := 1.0 - erf(x)` Fixes https://github.com/pytorch/pytorch/issues/155337 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155382 Approved by: https://github.com/manuelcandales, https://github.com/dcci	2025-06-07 02:35:12 +00:00
Oleksandr Stashuk	400f439670	[pt][easy] Rename metadata column (#155365 ) Summary: Fixing typo: our logging requires autotuning_data instead of autotune_data, making it consistent Test Plan: Run benchmark, observe in perfetto trace proper name Rollback Plan: Differential Revision: D76159393 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155365 Approved by: https://github.com/masnesral, https://github.com/Skylion007	2025-06-07 02:25:55 +00:00
William Wen	81b0b308ca	[dynamo] constant fold torch.cuda.is_initialized (#155300 ) Fixes https://github.com/pytorch/pytorch/issues/129659 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155300 Approved by: https://github.com/StrongerXi, https://github.com/jansel	2025-06-07 02:21:11 +00:00
Stella Laurenzo	10cd1de518	[ROCm] Make optional features in LoadHIP better conditioned. (#155305 ) * The `rocm-core` CMake package only started appearing in ROCm 6.4, so rework the version probing to work if it is not present. Also collapses the unneeded operating system conditioning in favor of feature probing. * Make `hipsparselt` optional: it only started appearing in ROCm 6.4 and it is not in all recent distribution channels yet. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155305 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-07 02:20:55 +00:00
Nikita Shulga	5596cefba6	Fix segfault during NumPy string tensor conversion (#155364 ) By checking dtype first, but add elemnt_size check as well Fixes https://github.com/pytorch/pytorch/issues/155328 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155364 Approved by: https://github.com/Skylion007	2025-06-07 01:55:00 +00:00
Abhishek Kumar	be2e43264d	[CI]Update windows runner to windows-2022 (#154368 ) As per info in : actions/runner-images#12045 We need to change window runner. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154368 Approved by: https://github.com/cyyever, https://github.com/atalman	2025-06-07 01:39:19 +00:00
Aaron Gokaslan	83d22256f8	[BE][Ez]: Improve typing in torch._logging (#155345 ) Add a few missing returns in torch._logging and use ruff to infer the obvious ones. LazyStr now properly checks the return type of the Callable and the args and kwargs passed to it Pull Request resolved: https://github.com/pytorch/pytorch/pull/155345 Approved by: https://github.com/ezyang	2025-06-07 00:04:39 +00:00
Jane Xu	9b4db093cb	Add C shim for at::pad and fix some typos (#155226 ) As stated, we would like a pad shim to support custom ops wanting to build in an ABI stable manner. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155226 Approved by: https://github.com/desertfire	2025-06-06 23:08:39 +00:00
Francisco R Castro Garcia	cd82096973	DOC: Convert to markdown: ddp_comm_hooks.rst, debugging_environment_variables.rst, deploy.rst, deterministic.rst, distributed.algorithms.join.rst (#155298 ) Fixes #155017 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155298 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-06 22:44:50 +00:00
Aaron Gokaslan	457dd79927	[BE][Ez]: Remove unnecessary accesses of dim vector (#155334 ) It's better because you return less date, encapsulate more, and no longer need special handling of symvec vs nonsym vec dim(). Also removes a few casts and fixes a few potential edge cases relating to unsigned comparisons Pull Request resolved: https://github.com/pytorch/pytorch/pull/155334 Approved by: https://github.com/ezyang	2025-06-06 21:28:25 +00:00
Kazuaki Ishizaki	c95705dac2	[Docs] Convert to markdown: torch.compiler_troubleshooting_old.rst, torch.compiler.rst (#155348 ) Part of changes #155040 (parent PR #155120) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155348 Approved by: https://github.com/svekars	2025-06-06 21:26:24 +00:00
eellison	d2a2bfcb58	Turn on new tiling by default (#154768 ) Turning on in fbcode to come. Also updates `max_tiles` to have a default value of None. The existing tiling logic doesn't really handle max_tiles=3 well, but we do in the new tiling logic, so we default to 3 in the new logic and 2 elsewhere unless max_tiles has been explicitly set. TB runners have been very unstable recently (do we need to bump batch size ?) but e.g. for a [recent torchbench](https://hud.pytorch.org/benchmark/torchbench/inductor_with_cudagraphs?dashboard=torchinductor&startTime=Tue,%2027%20May%202025%2015:38:26%20GMT&stopTime=Tue,%2003%20Jun%202025%2015:38:26%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cuda%20(a100)&lBranch=gh/eellison/803/head&lCommit=8480c220db4eb3c9e2b58d85a698d0a7113a6e37&rBranch=main&rCommit=0cd18ba1ca35d87916723d445c06664615dcae12) inference run we had 15 models with a lower execution time (i.g. green) and 2 models with higher (i.e.. red) I am doing another run and will update here. Dynamic shapes is not yet turned on because there are a lot of fixes to be done in splitting that don't work yet.. See: ``` (Pdb) p expr ((s25s85)//32) (Pdb) p FloorDiv(expr, expr) ((s25s85)//(32(((s25s85)//32)))) ``` and also - unbacked shape is not multiple of itself. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154768 Approved by: https://github.com/jansel	2025-06-06 21:19:35 +00:00
Animesh Jain	bc5a11b581	[easy][invoke_subgraph] Remove skip from already fixed test (#155286 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155286 Approved by: https://github.com/zou3519	2025-06-06 21:16:22 +00:00
Wei Feng	0d8c029584	[FSDP2] keep root unsharded when not specifying reshard_after_forward (#155319 ) for `fully_shard(model)` without explicitly setting `reshard_after_forward=True/False`, we keep root unsharded. When user explicitly set `reshard_after_forward`, we respect it Pull Request resolved: https://github.com/pytorch/pytorch/pull/155319 Approved by: https://github.com/mori360	2025-06-06 20:29:31 +00:00
loganthomas	4f5b34427b	DOC: Convert to markdown: torch.overrides.rst, type_info.rst, utils.rst, xpu.rst (#155088 ) Fixes #155041 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155088 Approved by: https://github.com/svekars Co-authored-by: Svetlana Karslioglu <svekars@meta.com>	2025-06-06 20:16:13 +00:00
Animesh Jain	067fd0b3ab	[dynamo][cleanup] Simplify disabling of the helper functions on tensor properties (#155259 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155259 Approved by: https://github.com/zhxchen17	2025-06-06 19:44:40 +00:00
Ke Wen	749757ac1b	[a2av] Align length of major dimension in output of 2D a2av (#155172 ) Downstream consumer of the 2D all-to-all-v is often a group GEMM. Today the GEMM often have an alignment requirement on the chunk sizes within grouped sequence, where each chunk carries the tokens headed for an expert. For example, `torch._group_mm` requires an alignment of 8. This PR adds that alignment capability, when user passes in a `major_align` argument, so that no extra padding step is needed. The key in supporting that is making the output offsets aligned to such value. (Output offsets are returned to the users in the 3rd row of `in_out_splits`, on device. The 2nd row, output splits, are unaffected by this alignment value -- i.e. reflecting true number of tokens for an expert.) The algorithm is as follows. ![502413288_678786854922438_530852083153996358_n](https://github.com/user-attachments/assets/557624a3-150e-4ab6-ba8b-1dbaa5ac01ac) In detailed implementation, we use warp scan to calculate prefix sum on the "block" illustrated above. As a result, the "block" size, i.e. `npes` is currently limited to warp size 32. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155172 Approved by: https://github.com/ngimel ghstack dependencies: #153653, #153677, #155058	2025-06-06 19:39:44 +00:00
Jovian Anthony Jaison	1ccc57e428	Log backward no-op to tlparse and pt2 compile events. (#154544 ) Summary: Log backward no-op to tlparse and pt2 compile events. Test Plan: $ rm -rf /tmp/r && TORCH_TRACE=/tmp/r buck2 run //scripts/jovian:backward_noop_repro_compile Used print statements to verify we enter the logging code region. Differential Revision: D75231665 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154544 Approved by: https://github.com/c00w	2025-06-06 18:08:19 +00:00
Blaine Burton Rister	2e2ea7290a	[Inductor] Support autotuning in the FX backend. (#155049 ) # Feature If `config.triton.autotune_at_compile_time` is set to `True`, autotune Triton kernels during FX conversion. Else, stick with the existing behavior of using the first precompiled config. # Test plan Added CI tests verifying that the tuner is called iff this flag is set, with and without dynamic shapes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155049 Approved by: https://github.com/jansel	2025-06-06 17:44:14 +00:00
Ke Wen	453bc9fbdf	[a2av] 2D all-to-all-vdev (#155058 ) A 2D AllToAllv shuffle is illustrated below: (`world_size` = 2, `ne` = 2, where `ne` is number of experts per rank) ``` Source: \| Rank 0 \| Rank 1 \| \| c0 \| c1 \| c2 \| c3 \| d0 \| d1 \| d2 \| d3 \| Dest : \| Rank 0 \| Rank 1 \| \| c0 \| d0 \| c1 \| d1 \| c2 \| d2 \| c3 \| d3 \| ``` where each `c_i` / `d_i` are slices of the `input` tensor, targeting expert `i`, with length indicated by input splits (in `in_out_splits[0]`). That is, the 2D AllToAllv shuffle achieves a transpose from rank-major order at input to expert-major order at output. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155058 Approved by: https://github.com/ngimel ghstack dependencies: #153653, #153677	2025-06-06 17:35:39 +00:00
Oleksandr Stashuk	64436c38c9	[logs] Add autotuning data (#154771 ) Summary: Add autotuning logging data to scuba/chrome trace. Test Plan: ``` TORCHINDUCTOR_MAX_AUTOTUNE=1 tlp buck run //scripts/sashko:compilation_sample ``` Open https://interncache-all.fbcdn.net/manifold/perfetto-artifacts/tree/ui/index.html#!/viewer?local_cache_key=00000000-0000-0000-92db-f23383ebf5b5, search for template_autotuning, see in metadata strides (see screenshot) Differential Revision: D75457770 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154771 Approved by: https://github.com/masnesral, https://github.com/PaulZhang12	2025-06-06 17:12:55 +00:00
Natalia Gimelshein	706bc41c4c	pass mempool arg through emptyCache (#155315 ) Fixing typo in a previous PR #154746 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155315 Approved by: https://github.com/Skylion007	2025-06-06 16:14:26 +00:00
Aleksei Nikiforov	7ae7c14143	Reduce scope of s390x CI (#155208 ) The purpose of this change is to reduce scope of s390x CI to stop it potentially blocking usual workflows for other users while still keeping nightly builds and tests for me to look at. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155208 Approved by: https://github.com/malfet	2025-06-06 16:07:34 +00:00
bobrenjc93	fc77269262	Add randint_like tensor overload for high (#154899 ) Fixes #135664 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154899 Approved by: https://github.com/StrongerXi	2025-06-06 15:48:00 +00:00
PyTorch MergeBot	7e4c097b07	Revert "[inductor] Add typing to _inductor/ir.py (#149958 )" This reverts commit 529e0357c6c4e74f8cd32c29198c5f1c9f6e329d. Reverted https://github.com/pytorch/pytorch/pull/149958 on behalf of https://github.com/malfet due to Looks like it broke inductor_torchbind tests, due to more graphbreaks, see `b0fbbef136/1` ([comment](https://github.com/pytorch/pytorch/pull/149958#issuecomment-2949583209))	2025-06-06 15:19:16 +00:00
PyTorch MergeBot	b0fbbef136	Revert "Turn on new tiling by default (#154768 )" This reverts commit 7dcc77e422dcf97ce35991a138ab635a5cb88731. Reverted https://github.com/pytorch/pytorch/pull/154768 on behalf of https://github.com/malfet due to Looks like it broke inductor CPU, see `231eb9902b/1` ([comment](https://github.com/pytorch/pytorch/pull/154768#issuecomment-2949468396))	2025-06-06 14:40:03 +00:00
Nikita Shulga	231eb9902b	[MPS][BE] Extend ndim_and_dtypes to 4 elements (#155272 ) Metal arguments must be 8 bytes aliged (or may be 16 bytes), so running any strided (or typecasted) binary op with MTL_DEBUG_LAYER leads to exception ``` % MTL_DEBUG_LAYER=1 python3 ../test/test_mps.py -v -k test_output_match_add 2025-06-05 15:41:34.201 Python[86653:16826825] Metal API Validation Enabled test_output_match_add_mps_bfloat16 (__main__.TestConsistencyMPS.test_output_match_add_mps_bfloat16) ... validateComputeFunctionArguments:1083: failed assertion `Compute Function(add_strided_bfloat_bfloat): argument ndim[0] from buffer(7) with offset(0) and length(12) has space for 12 bytes, but argument has a length(16).' zsh: abort MTL_DEBUG_LAYER=1 python3 ../test/test_mps.py -v -k test_output_match_add ``` Extend it to 4 elements and pass output dtype, which will be used by binary_op later on anyway Test plan: Run abovementioned command with `MTL_DEBUG_LAYER=1` and make sure everything passes Pull Request resolved: https://github.com/pytorch/pytorch/pull/155272 Approved by: https://github.com/angelayi, https://github.com/dcci, https://github.com/cyyever	2025-06-06 14:20:21 +00:00
Tom Ritchford	529e0357c6	[inductor] Add typing to _inductor/ir.py (#149958 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/149958 Approved by: https://github.com/Skylion007	2025-06-06 14:15:01 +00:00
Edward Z. Yang	348fd45065	Support detached checkout in tools/nightly.py (#154314 ) Prompt for Sonnet 3.7 in Claude Code: Only inspect tools/nightly.py, all other files are irrelevant to your task. Do not use any shell commands. Task: Add a --detach argument to this script which instead of making a new branch just directly checks out the correct commit in detached mode. With two interventions: - Branch and detach are mutually exclusive. So you should consolidate them into a single argument. Why don't we take over the 'None' option? - Do you know that nightly_version is guaranteed to be a commit hash? It seems it would be safer to explicitly pass --detach I tested by running `python tools/nightly.py checkout` and observing that my worktree was detached at this point. Signed-off-by: Edward Z. Yang <ezyang@meta.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/154314 Approved by: https://github.com/XuehaiPan, https://github.com/malfet	2025-06-06 13:28:29 +00:00
drisspg	907aea032d	Add claude local md files (#155299 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155299 Approved by: https://github.com/ezyang	2025-06-06 13:28:26 +00:00
Aaron Gokaslan	6b1211df29	[BE]: Backport runtime_checkable perf improvements/behavior from 3.12 (#155130 ) Backports some behavior changes and performance improvements with runtime_checkable in 3.12 to older versions of Python. Should be free performance improvement on typing checking protocols since everything works on Python 3.12. The difference between the two versions of runtime_checkable is [these lines](`40e22ebb2c/src/typing_extensions.py (L800-L823)`). Pull Request resolved: https://github.com/pytorch/pytorch/pull/155130 Approved by: https://github.com/rec, https://github.com/aorenste	2025-06-06 13:28:05 +00:00
Yu, Guangye	10cef1e25d	Remove torch XPU ABI=0 build logic for old compiler (#150095 ) # Motivation Follow https://github.com/pytorch/pytorch/pull/149888, this PR intends to remove ABI=0 build logic for PyTorch XPU build with old compiler (< 2025.0). For newer compilers >= 2025.0, the ABI is neutral by default without requiring additional compilation options (`-fpreview-breaking-changes`). # Additional Context This PR depends on XPU CI pass, which will be fixed by https://github.com/pytorch/pytorch/pull/149843 and https://github.com/intel/torch-xpu-ops/pull/1515 Pull Request resolved: https://github.com/pytorch/pytorch/pull/150095 Approved by: https://github.com/EikanWang, https://github.com/malfet	2025-06-06 13:13:19 +00:00
Nikita Shulga	58e5d20c57	[BE] Delete IS_SPMM_AVAILABLE() logic (#155296 ) As it's been available on all currently supported platforms Pull Request resolved: https://github.com/pytorch/pytorch/pull/155296 Approved by: https://github.com/clee2000	2025-06-06 13:12:35 +00:00
Animesh Jain	271ca679a8	[reland][dynamo] Record the pre-graph bytecode using fast record function event (#154974 ) reland of https://github.com/pytorch/pytorch/pull/154769 @diff-train-skip-merge Pull Request resolved: https://github.com/pytorch/pytorch/pull/154974 Approved by: https://github.com/Lucaskabela, https://github.com/jansel	2025-06-06 13:11:03 +00:00
PyTorch MergeBot	9656251bb1	Revert "[BE] Update cudnn to 9.10.1.4 (#155122 )" This reverts commit a14f427db68e54500ef4cd9ed34cb9537263bb74. Reverted https://github.com/pytorch/pytorch/pull/155122 on behalf of https://github.com/malfet due to Looks like it breaks a bunch of tests, see `36a722e20d/1` ([comment](https://github.com/pytorch/pytorch/pull/155122#issuecomment-2949209801))	2025-06-06 13:03:49 +00:00
中野博文	36a722e20d	[typo] Fix 'intialize' -> 'initialize' in proxy_tensor.py (#155301 ) ## Description Fixes a typo in the comment of `torch/fx/experimental/proxy_tensor.py`, changing "intialize" to "initialize". ## Issue None ## Type of change - [x] Typo fix ## Checklist - [x] My code follows the style guidelines of this project - [x] I have performed a self-review of my own code - [x] My changes generate no new warnings Pull Request resolved: https://github.com/pytorch/pytorch/pull/155301 Approved by: https://github.com/jingsh, https://github.com/ezyang, https://github.com/cyyever	2025-06-06 10:43:44 +00:00
zeshengzong	9d59b516e9	Make device check throw specific error (#155085 ) Fixes #122757 The fix is lost after revert and rebase previous PR https://github.com/pytorch/pytorch/pull/150750 (only change of tests are merged). ## Test Result ```python >>> import torch >>> >>> model_output = torch.randn(10, 5).cuda() >>> labels = torch.randint(0, 5, (10,)).cuda() >>> weights = torch.randn(5) >>> >>> loss_fn = torch.nn.CrossEntropyLoss(weight=weights) >>> loss = loss_fn(input=model_output, target=labels) Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/home/zong/code/pytorch/torch/nn/modules/module.py", line 1767, in _wrapped_call_impl return self._call_impl(args, kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/zong/code/pytorch/torch/nn/modules/module.py", line 1778, in _call_impl return forward_call(args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/zong/code/pytorch/torch/nn/modules/loss.py", line 1297, in forward return F.cross_entropy( ^^^^^^^^^^^^^^^^ File "/home/zong/code/pytorch/torch/nn/functional.py", line 3476, in cross_entropy return torch._C._nn.cross_entropy_loss( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ RuntimeError: Expected all tensors to be on the same device, but got weight is on cpu, different from other tensors on cuda:0 (when checking argument in method wrapper_CUDA_nll_loss_forward) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155085 Approved by: https://github.com/mikaylagawarecki	2025-06-06 07:00:04 +00:00
Wang, Chuanqi	07da8a469b	[CI] fix xpu-smi hang issue on some xpu runners (#155194 ) To workaround xpu-smi hang issue on some XPU runners, refer https://github.com/pytorch/pytorch/actions/runs/15431583674/job/43431289026?pr=154962 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155194 Approved by: https://github.com/EikanWang, https://github.com/malfet	2025-06-06 06:51:26 +00:00
Marcin Pioch	e694280d12	Custom FX pass for inductor's backend registration (#154841 ) This PR is related to RFC #153532. It is an extension to Inductor's backend registration interface to allow to register custom FX passes by the backend. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154841 Approved by: https://github.com/jansel Co-authored-by: Jason Ansel <jansel@jansel.net>	2025-06-06 06:49:44 +00:00
Jing Xu	c6b4f98625	Add Intel GPU info collection to the collect env script (#137846 ) As title, add Intel GPU info collection to the collect env script Output examples: 1. CPU on Windows ``` C:\Users\user\miniforge3\envs\py310\lib\site-packages\torch\_subclasses\functional_tensor.py:279: UserWarning: Failed to initialize NumPy: No module named 'numpy' (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\pytorch\torch\csrc\utils\tensor_numpy.cpp:81.) cpu = _conversion_method_template(device=torch.device("cpu")) Collecting environment information... PyTorch version: 2.8.0.dev20250528+cpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Microsoft Windows 11 Enterprise (10.0.22631 64-bit) GCC version: Could not collect Clang version: Could not collect CMake version: Could not collect Libc version: N/A Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:06:35) [MSC v.1943 64 bit (AMD64)] (64-bit runtime) Python platform: Windows-10-10.0.22631-SP0 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: False HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Name: 12th Gen Intel(R) Core(TM) i7-1270P Manufacturer: GenuineIntel Family: 198 Architecture: 9 ProcessorType: 3 DeviceID: CPU0 CurrentClockSpeed: 1711 MaxClockSpeed: 2200 L2CacheSize: 9216 L2CacheSpeed: None Revision: None Versions of relevant libraries: [pip3] torch==2.8.0.dev20250528+cpu [conda] torch 2.8.0.dev20250528+cpu pypi_0 pypi ``` 2. XPU on Windows ``` Collecting environment information... PyTorch version: 2.8.0a0+gitef6306e Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Microsoft Windows 10 Pro (10.0.19045 64-bit) GCC version: (GCC) 13.1.0 Clang version: Could not collect CMake version: version 3.29.3 Libc version: N/A Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:06:35) [MSC v.1943 64 bit (AMD64)] (64-bit runtime) Python platform: Windows-10-10.0.19045-SP0 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: True XPU used to build PyTorch: 20250101 Intel GPU driver version: * 32.0.101.6795 (20250520000000.****+) Intel GPU models onboard: Intel(R) Arc(TM) A770 Graphics Intel GPU models detected: * [0] _XpuDeviceProperties(name='Intel(R) Arc(TM) A770 Graphics', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.33184', total_memory=15915MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=128, sub_group_sizes=[8 16 32], has_fp16=1, has_fp64=0, has_atomic64=1) HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: ---------------------- Name: Intel(R) Xeon(R) Platinum 8260 CPU @ 2.40GHz Manufacturer: GenuineIntel Family: 179 Architecture: 9 ProcessorType: 3 DeviceID: CPU0 CurrentClockSpeed: 2401 MaxClockSpeed: 2401 L2CacheSize: 24576 L2CacheSpeed: None Revision: 21767 ---------------------- Name: Intel(R) Xeon(R) Platinum 8260 CPU @ 2.40GHz Manufacturer: GenuineIntel Family: 179 Architecture: 9 ProcessorType: 3 DeviceID: CPU1 CurrentClockSpeed: 2200 MaxClockSpeed: 2401 L2CacheSize: 24576 L2CacheSpeed: None Revision: 21767 Versions of relevant libraries: [pip3] intel_extension_for_pytorch==2.8.10+gitb3ea3a1 [pip3] numpy==2.1.2 [pip3] optree==0.13.1 [pip3] pytorch-triton-xpu==3.3.1+gitb0e26b73 [pip3] torch==2.8.0a0+gitef6306e [conda] intel-extension-for-pytorch 2.8.10+gitb3ea3a1 pypi_0 pypi [conda] mkl 2025.1.0 pypi_0 pypi [conda] mkl-dpcpp 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-blas 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-datafitting 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-dft 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-lapack 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-rng 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-sparse 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-stats 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-vm 2025.1.0 pypi_0 pypi [conda] pytorch-triton-xpu 3.3.1+gitb0e26b73 pypi_0 pypi [conda] torch 2.8.0a0+gitef6306e pypi_0 pypi ``` 3. CPU on Linux ``` /opt/python/cp312-cp312/lib/python3.12/site-packages/torch/_subclasses/functional_tensor.py:279: UserWarning: Failed to initialize NumPy: No module named 'numpy' (Triggered internally at /pytorch/torch/csrc/utils/tensor_numpy.cpp:81.) cpu = _conversion_method_template(device=torch.device("cpu")) Collecting environment information... PyTorch version: 2.8.0.dev20250528+cpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: AlmaLinux 8.10 (Cerulean Leopard) (x86_64) GCC version: (GCC) 14.2.1 20250110 (Red Hat 14.2.1-7) Clang version: Could not collect CMake version: version 4.0.0 Libc version: glibc-2.28 Python version: 3.12.10 (main, Apr 19 2025, 05:03:56) [GCC 14.2.1 20250110 (Red Hat 14.2.1-7)] (64-bit runtime) Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.28 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: False HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Byte Order: Little Endian CPU(s): 88 On-line CPU(s) list: 0-87 Thread(s) per core: 2 Core(s) per socket: 22 Socket(s): 2 NUMA node(s): 2 Vendor ID: GenuineIntel CPU family: 6 Model: 85 Model name: Intel(R) Xeon(R) Gold 6238M CPU @ 2.10GHz Stepping: 7 CPU MHz: 1000.000 CPU max MHz: 3700.0000 CPU min MHz: 1000.0000 BogoMIPS: 4200.00 Virtualization: VT-x L1d cache: 32K L1i cache: 32K L2 cache: 1024K L3 cache: 30976K NUMA node0 CPU(s): 0-21,44-65 NUMA node1 CPU(s): 22-43,66-87 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 intel_ppin ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts vnmi pku ospke avx512_vnni md_clear flush_l1d arch_capabilities Versions of relevant libraries: [pip3] torch==2.8.0.dev20250528+cpu [conda] Could not collect ``` 5. XPU on Linux ``` Collecting environment information... PyTorch version: 2.8.0.dev20250516+xpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Ubuntu 22.04.4 LTS (x86_64) GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0 Clang version: Could not collect CMake version: version 3.31.6 Libc version: glibc-2.35 Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:19:12) [GCC 13.3.0] (64-bit runtime) Python platform: Linux-5.15.50-051550-generic-x86_64-with-glibc2.35 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: True XPU used to build PyTorch: 20250101 Intel GPU driver version: * intel_opencl: 24.39.31294.21-1032~22.04 * level_zero: 1.17.44.0-1022~22.04 Intel GPU models onboard: * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 Intel GPU models detected: * [0] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [1] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [2] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [3] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [4] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [5] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [6] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [7] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 52 bits physical, 57 bits virtual Byte Order: Little Endian CPU(s): 224 On-line CPU(s) list: 0-223 Vendor ID: GenuineIntel Model name: Intel(R) Xeon(R) Platinum 8480+ CPU family: 6 Model: 143 Thread(s) per core: 2 Core(s) per socket: 56 Socket(s): 2 Stepping: 6 CPU max MHz: 3800.0000 CPU min MHz: 800.0000 BogoMIPS: 4000.00 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cat_l2 cdp_l3 invpcid_single intel_ppin cdp_l2 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow vnmi flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm rdt_a avx512f avx512dq rdseed adx smap avx512ifma clflushopt clwb intel_pt avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local split_lock_detect avx_vnni avx512_bf16 wbnoinvd dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req avx512vbmi umip pku ospke waitpkg avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg tme avx512_vpopcntdq la57 rdpid bus_lock_detect cldemote movdiri movdir64b enqcmd fsrm md_clear serialize tsxldtrk pconfig arch_lbr avx512_fp16 flush_l1d arch_capabilities Virtualization: VT-x L1d cache: 5.3 MiB (112 instances) L1i cache: 3.5 MiB (112 instances) L2 cache: 224 MiB (112 instances) L3 cache: 210 MiB (2 instances) NUMA node(s): 2 NUMA node0 CPU(s): 0-55,112-167 NUMA node1 CPU(s): 56-111,168-223 Vulnerability Itlb multihit: Not affected Vulnerability L1tf: Not affected Vulnerability Mds: Not affected Vulnerability Meltdown: Not affected Vulnerability Mmio stale data: Not affected Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl and seccomp Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Vulnerability Spectre v2: Mitigation; Enhanced IBRS, IBPB conditional, RSB filling Vulnerability Srbds: Not affected Vulnerability Tsx async abort: Not affected Versions of relevant libraries: [pip3] numpy==2.2.5 [pip3] pytorch-triton-xpu==3.3.0+git0bcc8265 [pip3] torch==2.8.0.dev20250516+xpu [conda] mkl 2025.1.0 pypi_0 pypi [conda] numpy 2.2.5 pypi_0 pypi [conda] onemkl-sycl-blas 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-dft 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-lapack 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-rng 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-sparse 2025.1.0 pypi_0 pypi [conda] pytorch-triton-xpu 3.3.0+git0bcc8265 pypi_0 pypi [conda] torch 2.8.0.dev20250516+xpu pypi_0 pypi ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/137846 Approved by: https://github.com/guangyey, https://github.com/malfet Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com> Co-authored-by: Nikita Shulga <2453524+malfet@users.noreply.github.com>	2025-06-06 05:53:24 +00:00
PyTorch MergeBot	d3d64c6db0	Revert "Add pinned numpy and fix build (#155129 )" This reverts commit a3098a74d494020dbb906c05ef047013e1921662. Reverted https://github.com/pytorch/pytorch/pull/155129 on behalf of https://github.com/malfet due to Broke test_spectral_op, looks like missing xfail, see `0db3e0cf29/1` ([comment](https://github.com/pytorch/pytorch/pull/155129#issuecomment-2947951632))	2025-06-06 03:14:47 +00:00
PyTorch MergeBot	0db3e0cf29	Revert "Add Intel GPU info collection to the collect env script (#137846 )" This reverts commit e1180c7228ba8c8b16cabf78706d4a67ca189a6b. Reverted https://github.com/pytorch/pytorch/pull/137846 on behalf of https://github.com/malfet due to Breaks doc test, but should be easily fixable ([comment](https://github.com/pytorch/pytorch/pull/137846#issuecomment-2947935940))	2025-06-06 03:08:48 +00:00
Simon Fan	28796f71d0	Redo D75092426: [internal] Expose additional metadata to compilation callbacks (#155063 ) Originally https://github.com/pytorch/pytorch/pull/153596 --------------- Summary: via reverting D75708685 gate the ROCm failure Test Plan: Unit tests in OSS, sandcastle Rollback Plan: Bifferential Revision: D75894349 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155063 Approved by: https://github.com/masnesral	2025-06-05 23:40:31 +00:00
Xuan Zhang	72453a6676	[PT2][comms] put `visualize_overlap` in a try-except block (#155222 ) Summary: For simple FSDP, this `visualize_overlap` function is throwing errors. Seems to be a mistake here since `visualize_overlap` is called twice here and one is in try-except and one is not, so doing the same for both places. Test Plan: :) Rollback Plan: Reviewed By: Microve Bifferential Revision: D75985733 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155222 Approved by: https://github.com/yf225	2025-06-05 23:39:48 +00:00
Brian Coutinho	9bae2fcf99	[profiler] Enable all configured activities in CUPTI Range profiler mode (#154749 ) Summary: Updates the pytorch range profiler mode (metrics mode) to support all trace activitity types. Reviewed By: sraikund16 Bifferential Revision: D75568693 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154749 Approved by: https://github.com/sraikund16	2025-06-05 23:38:54 +00:00
Shangdi Yu	26f066bb61	Add AOTI model name config (#154129 ) Summary: If a model name is specified in aoti config, the generated files will use that model name as file stem. Test Plan: ``` buck2 run mode/dev-nosan caffe2/test/inductor:test_aot_inductor -- -r test_using_model_name_for_files ``` Bifferential Revision: D75102034 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154129 Approved by: https://github.com/desertfire	2025-06-05 23:38:11 +00:00
Nicolas Macchioni	fa705f7912	[BE] minor refactor + some comments on behavior (#154695 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154695 Approved by: https://github.com/masnesral, https://github.com/eellison	2025-06-05 23:00:46 +00:00
Jeff Daily	9e88d6c857	[ROCm] manywheel missing hipsparselt deps (#155254 ) Bundle libhipsparselt.so and auxiliary files into wheel. Dependency added by hipsparselt integration #150578. Fixes #155207. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155254 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-05 22:45:36 +00:00
Jing Xu	e1180c7228	Add Intel GPU info collection to the collect env script (#137846 ) As title, add Intel GPU info collection to the collect env script Output examples: 1. CPU on Windows ``` C:\Users\user\miniforge3\envs\py310\lib\site-packages\torch\_subclasses\functional_tensor.py:279: UserWarning: Failed to initialize NumPy: No module named 'numpy' (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\pytorch\torch\csrc\utils\tensor_numpy.cpp:81.) cpu = _conversion_method_template(device=torch.device("cpu")) Collecting environment information... PyTorch version: 2.8.0.dev20250528+cpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Microsoft Windows 11 Enterprise (10.0.22631 64-bit) GCC version: Could not collect Clang version: Could not collect CMake version: Could not collect Libc version: N/A Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:06:35) [MSC v.1943 64 bit (AMD64)] (64-bit runtime) Python platform: Windows-10-10.0.22631-SP0 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: False HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Name: 12th Gen Intel(R) Core(TM) i7-1270P Manufacturer: GenuineIntel Family: 198 Architecture: 9 ProcessorType: 3 DeviceID: CPU0 CurrentClockSpeed: 1711 MaxClockSpeed: 2200 L2CacheSize: 9216 L2CacheSpeed: None Revision: None Versions of relevant libraries: [pip3] torch==2.8.0.dev20250528+cpu [conda] torch 2.8.0.dev20250528+cpu pypi_0 pypi ``` 2. XPU on Windows ``` Collecting environment information... PyTorch version: 2.8.0a0+gitef6306e Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Microsoft Windows 10 Pro (10.0.19045 64-bit) GCC version: (GCC) 13.1.0 Clang version: Could not collect CMake version: version 3.29.3 Libc version: N/A Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:06:35) [MSC v.1943 64 bit (AMD64)] (64-bit runtime) Python platform: Windows-10-10.0.19045-SP0 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: True XPU used to build PyTorch: 20250101 Intel GPU driver version: * 32.0.101.6795 (20250520000000.****+) Intel GPU models onboard: Intel(R) Arc(TM) A770 Graphics Intel GPU models detected: * [0] _XpuDeviceProperties(name='Intel(R) Arc(TM) A770 Graphics', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.33184', total_memory=15915MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=128, sub_group_sizes=[8 16 32], has_fp16=1, has_fp64=0, has_atomic64=1) HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: ---------------------- Name: Intel(R) Xeon(R) Platinum 8260 CPU @ 2.40GHz Manufacturer: GenuineIntel Family: 179 Architecture: 9 ProcessorType: 3 DeviceID: CPU0 CurrentClockSpeed: 2401 MaxClockSpeed: 2401 L2CacheSize: 24576 L2CacheSpeed: None Revision: 21767 ---------------------- Name: Intel(R) Xeon(R) Platinum 8260 CPU @ 2.40GHz Manufacturer: GenuineIntel Family: 179 Architecture: 9 ProcessorType: 3 DeviceID: CPU1 CurrentClockSpeed: 2200 MaxClockSpeed: 2401 L2CacheSize: 24576 L2CacheSpeed: None Revision: 21767 Versions of relevant libraries: [pip3] intel_extension_for_pytorch==2.8.10+gitb3ea3a1 [pip3] numpy==2.1.2 [pip3] optree==0.13.1 [pip3] pytorch-triton-xpu==3.3.1+gitb0e26b73 [pip3] torch==2.8.0a0+gitef6306e [conda] intel-extension-for-pytorch 2.8.10+gitb3ea3a1 pypi_0 pypi [conda] mkl 2025.1.0 pypi_0 pypi [conda] mkl-dpcpp 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-blas 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-datafitting 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-dft 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-lapack 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-rng 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-sparse 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-stats 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-vm 2025.1.0 pypi_0 pypi [conda] pytorch-triton-xpu 3.3.1+gitb0e26b73 pypi_0 pypi [conda] torch 2.8.0a0+gitef6306e pypi_0 pypi ``` 3. CPU on Linux ``` /opt/python/cp312-cp312/lib/python3.12/site-packages/torch/_subclasses/functional_tensor.py:279: UserWarning: Failed to initialize NumPy: No module named 'numpy' (Triggered internally at /pytorch/torch/csrc/utils/tensor_numpy.cpp:81.) cpu = _conversion_method_template(device=torch.device("cpu")) Collecting environment information... PyTorch version: 2.8.0.dev20250528+cpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: AlmaLinux 8.10 (Cerulean Leopard) (x86_64) GCC version: (GCC) 14.2.1 20250110 (Red Hat 14.2.1-7) Clang version: Could not collect CMake version: version 4.0.0 Libc version: glibc-2.28 Python version: 3.12.10 (main, Apr 19 2025, 05:03:56) [GCC 14.2.1 20250110 (Red Hat 14.2.1-7)] (64-bit runtime) Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.28 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: False HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Byte Order: Little Endian CPU(s): 88 On-line CPU(s) list: 0-87 Thread(s) per core: 2 Core(s) per socket: 22 Socket(s): 2 NUMA node(s): 2 Vendor ID: GenuineIntel CPU family: 6 Model: 85 Model name: Intel(R) Xeon(R) Gold 6238M CPU @ 2.10GHz Stepping: 7 CPU MHz: 1000.000 CPU max MHz: 3700.0000 CPU min MHz: 1000.0000 BogoMIPS: 4200.00 Virtualization: VT-x L1d cache: 32K L1i cache: 32K L2 cache: 1024K L3 cache: 30976K NUMA node0 CPU(s): 0-21,44-65 NUMA node1 CPU(s): 22-43,66-87 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 intel_ppin ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts vnmi pku ospke avx512_vnni md_clear flush_l1d arch_capabilities Versions of relevant libraries: [pip3] torch==2.8.0.dev20250528+cpu [conda] Could not collect ``` 5. XPU on Linux ``` Collecting environment information... PyTorch version: 2.8.0.dev20250516+xpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: Ubuntu 22.04.4 LTS (x86_64) GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0 Clang version: Could not collect CMake version: version 3.31.6 Libc version: glibc-2.35 Python version: 3.10.17 \| packaged by conda-forge \| (main, Apr 10 2025, 22:19:12) [GCC 13.3.0] (64-bit runtime) Python platform: Linux-5.15.50-051550-generic-x86_64-with-glibc2.35 Is CUDA available: False CUDA runtime version: No CUDA CUDA_MODULE_LOADING set to: N/A GPU models and configuration: No CUDA Nvidia driver version: No CUDA cuDNN version: No CUDA Is XPU available: True XPU used to build PyTorch: 20250101 Intel GPU driver version: * intel_opencl: 24.39.31294.21-1032~22.04 * level_zero: 1.17.44.0-1022~22.04 Intel GPU models onboard: * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 * Intel(R) Data Center GPU Max 1550 Intel GPU models detected: * [0] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [1] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [2] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [3] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [4] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [5] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [6] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) * [7] _XpuDeviceProperties(name='Intel(R) Data Center GPU Max 1550', platform_name='Intel(R) oneAPI Unified Runtime over Level-Zero', type='gpu', driver_version='1.6.31294+21', total_memory=65536MB, max_compute_units=512, gpu_eu_count=512, gpu_subslice_count=64, max_work_group_size=1024, max_num_sub_groups=64, sub_group_sizes=[16 32], has_fp16=1, has_fp64=1, has_atomic64=1) HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 52 bits physical, 57 bits virtual Byte Order: Little Endian CPU(s): 224 On-line CPU(s) list: 0-223 Vendor ID: GenuineIntel Model name: Intel(R) Xeon(R) Platinum 8480+ CPU family: 6 Model: 143 Thread(s) per core: 2 Core(s) per socket: 56 Socket(s): 2 Stepping: 6 CPU max MHz: 3800.0000 CPU min MHz: 800.0000 BogoMIPS: 4000.00 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cat_l2 cdp_l3 invpcid_single intel_ppin cdp_l2 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow vnmi flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm rdt_a avx512f avx512dq rdseed adx smap avx512ifma clflushopt clwb intel_pt avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local split_lock_detect avx_vnni avx512_bf16 wbnoinvd dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req avx512vbmi umip pku ospke waitpkg avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg tme avx512_vpopcntdq la57 rdpid bus_lock_detect cldemote movdiri movdir64b enqcmd fsrm md_clear serialize tsxldtrk pconfig arch_lbr avx512_fp16 flush_l1d arch_capabilities Virtualization: VT-x L1d cache: 5.3 MiB (112 instances) L1i cache: 3.5 MiB (112 instances) L2 cache: 224 MiB (112 instances) L3 cache: 210 MiB (2 instances) NUMA node(s): 2 NUMA node0 CPU(s): 0-55,112-167 NUMA node1 CPU(s): 56-111,168-223 Vulnerability Itlb multihit: Not affected Vulnerability L1tf: Not affected Vulnerability Mds: Not affected Vulnerability Meltdown: Not affected Vulnerability Mmio stale data: Not affected Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl and seccomp Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Vulnerability Spectre v2: Mitigation; Enhanced IBRS, IBPB conditional, RSB filling Vulnerability Srbds: Not affected Vulnerability Tsx async abort: Not affected Versions of relevant libraries: [pip3] numpy==2.2.5 [pip3] pytorch-triton-xpu==3.3.0+git0bcc8265 [pip3] torch==2.8.0.dev20250516+xpu [conda] mkl 2025.1.0 pypi_0 pypi [conda] numpy 2.2.5 pypi_0 pypi [conda] onemkl-sycl-blas 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-dft 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-lapack 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-rng 2025.1.0 pypi_0 pypi [conda] onemkl-sycl-sparse 2025.1.0 pypi_0 pypi [conda] pytorch-triton-xpu 3.3.0+git0bcc8265 pypi_0 pypi [conda] torch 2.8.0.dev20250516+xpu pypi_0 pypi ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/137846 Approved by: https://github.com/guangyey, https://github.com/malfet Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com> Co-authored-by: Nikita Shulga <2453524+malfet@users.noreply.github.com>	2025-06-05 22:35:04 +00:00
Murray Steele	0a092c7de6	Enable CPP Extension Open Registration tests on Arm (#144774 ) Enables most tests under CPP Extension Open Registration as they pass on Arm now. Pull Request resolved: https://github.com/pytorch/pytorch/pull/144774 Approved by: https://github.com/aditew01, https://github.com/fadara01, https://github.com/malfet	2025-06-05 22:32:28 +00:00
eellison	0827464002	Replace runtime type parameterization (#155221 ) See: ``` >>> import timeit; print(f"OrderedSet[str](): {timeit.timeit('OrderedSet[str]()', setup='from torch.utils._ordered_set import OrderedSet', number=1000000):.6f}s, OrderedSet(): {timeit.timeit('OrderedSet()', setup='from torch.utils._ordered_set import OrderedSet', number=1000000):.6f}s") ``` > `OrderedSet[str]()`: 0.354622s, OrderedSet(): 0.095376s Type parameterization should be on type hint, not in runtime. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155221 Approved by: https://github.com/Skylion007, https://github.com/jansel	2025-06-05 21:43:54 +00:00
eellison	7dcc77e422	Turn on new tiling by default (#154768 ) Turning on in fbcode to come. Also updates `max_tiles` to have a default value of None. The existing tiling logic doesn't really handle max_tiles=3 well, but we do in the new tiling logic, so we default to 3 in the new logic and 2 elsewhere unless max_tiles has been explicitly set. TB runners have been very unstable recently (do we need to bump batch size ?) but e.g. for a [recent torchbench](https://hud.pytorch.org/benchmark/torchbench/inductor_with_cudagraphs?dashboard=torchinductor&startTime=Tue,%2027%20May%202025%2015:38:26%20GMT&stopTime=Tue,%2003%20Jun%202025%2015:38:26%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cuda%20(a100)&lBranch=gh/eellison/803/head&lCommit=8480c220db4eb3c9e2b58d85a698d0a7113a6e37&rBranch=main&rCommit=0cd18ba1ca35d87916723d445c06664615dcae12) inference run we had 15 models with a lower execution time (i.g. green) and 2 models with higher (i.e.. red) I am doing another run and will update here. Dynamic shapes is not yet turned on because there are a lot of fixes to be done in splitting that don't work yet.. See: ``` (Pdb) p expr ((s25s85)//32) (Pdb) p FloorDiv(expr, expr) ((s25s85)//(32(((s25s85)//32)))) ``` and also - unbacked shape is not multiple of itself. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154768 Approved by: https://github.com/jansel	2025-06-05 21:34:09 +00:00
tvukovic-amd	a85ad55525	[ROCm][Windows] Fix offload gpu arch list in tests (#155212 ) Added fix to get ROCM_PROPERTY_ARCH_LIST value in set_target_properties in c10/cuda and caffe2 tests Pull Request resolved: https://github.com/pytorch/pytorch/pull/155212 Approved by: https://github.com/malfet	2025-06-05 20:30:28 +00:00
Michael Lazos	9a42f01586	[Cutlass] EVT dynamic shapes support (#154835 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154835 Approved by: https://github.com/henrylhtsang ghstack dependencies: #154829	2025-06-05 20:17:01 +00:00
Michael Lazos	5911f870c0	[Cutlass] fp8 dynamic shapes test (#154829 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154829 Approved by: https://github.com/henrylhtsang, https://github.com/eellison	2025-06-05 20:17:01 +00:00
Shangdi Yu	606d73bde4	Adding from_node for nodes in gm.module() (#155053 ) Summary: Adding "from_node" information that indicates which nodes are unlifted in `.module()` call. The lifted nodes will have "ExportedProgram.module().unlift()" passname in the last entry of from_node. Test Plan: ``` buck run fbcode//caffe2/test:test_export -- -r test_from_node_metadata_export ``` Rollback Plan: Reviewed By: angelayi Differential Revision: D75837494 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155053 Approved by: https://github.com/angelayi	2025-06-05 20:11:56 +00:00
Yidi Wu	c8c892b4a5	[scan] disable functionalization key in backward tracing (#154343 ) Previously, we didn't disable functionalization key when materializing backward graph. This causes the torch.zeros_like call for the case where grad is None to return a functional tensor that's not tracked by the proxy tensor mode. This PR fixes it by putting the tracing code under disable functionalization ctx manager. Fixes https://github.com/pytorch/pytorch/issues/153437. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154343 Approved by: https://github.com/zou3519	2025-06-05 20:06:33 +00:00
Joel Schlosser	5e93abe3c0	Address docs for clip_grad functions (#155125 ) This PR takes the opinionated stance that `torch.nn.utils.<func>` should be the preferred API over `torch.nn.utils.clip_grad.<func>`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155125 Approved by: https://github.com/albanD, https://github.com/mikaylagawarecki, https://github.com/janeyx99	2025-06-05 19:22:09 +00:00
Nikita Shulga	dd41a3907c	[MPS] Fix unary/binary ops for 2**32+ elem tensors (#155183 ) By using `TensorIterator::with_32bit_indexing()` primitive Add `bind_tensors` helper function that correctly sets up MPS tensors originating from TensorIterator TODO: Add comments to bind_tensors as well asunit test, based on ``` python -c "import torch;print((torch.rand(1, 1024, 1024, dtype=torch.bfloat16, device='mps') + torch.rand(5000, 1, 1, dtype=torch.bfloat16, device='mps')).sin())" ``` Fixes https://github.com/pytorch/pytorch/issues/154828 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155183 Approved by: https://github.com/cyyever, https://github.com/dcci, https://github.com/Skylion007 ghstack dependencies: #155150, #155178, #155184	2025-06-05 18:57:14 +00:00
PyTorch MergeBot	05dd638ee9	Revert "Add dont constant fold flag (#154945 )" This reverts commit 196c95d463367f15999c0cddc9eb89031e9988ab. Reverted https://github.com/pytorch/pytorch/pull/154945 on behalf of https://github.com/malfet due to This broke halide test sanity, see `a3098a74d4/1` ([comment](https://github.com/pytorch/pytorch/pull/154945#issuecomment-2945598901))	2025-06-05 18:25:59 +00:00
albanD	a3098a74d4	Add pinned numpy and fix build (#155129 ) Not sure why the online doc build passes but it fails locally with these broken strings... Also pinning numpy version even though it is technically optional to ensure users have the right version as most users have numpy in their environment anyways. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155129 Approved by: https://github.com/janeyx99, https://github.com/svekars	2025-06-05 17:44:18 +00:00
henrylhtsang	2481c4b2ea	[cutlass backend] add teraflops and increase rep for benchmark script (#154944 ) Differential Revision: [D75840023](https://our.internmc.facebook.com/intern/diff/D75840023/) I think I will continue to use do_bench for now. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154944 Approved by: https://github.com/mlazos	2025-06-05 17:20:29 +00:00
David Berard	be2ab96347	Inductor unit tests: cuda 12.6 -> 12.8 (#155056 ) Fixes #154938 When we update the Triton version in CI, we'll require cuda >= 12.8 for certain AOTI tests to pass: these AOTI tests try to run nvcc on the triton-generated PTX, and triton-generated PTX is PTX 8.7, which requires CUDA 12.8 Regarding the revert & reland: * This PR causes the python 3.13 version to be bumped from 3.13.2 to 3.13.3. test_deopt_from_append_list starts unexpectedly passing on 3.13.3, so I originally modified the test in https://github.com/pytorch/pytorch/pull/155167 to xfail only for <=3.13.2 * However there was a land race with https://github.com/pytorch/pytorch/pull/150796, which introduced another test that passes only for >=3.13.3. Resolution: * @guilhermeleobas reverted https://github.com/pytorch/pytorch/pull/150796 so I will reland this (and I've merged the test_deopt_from_append_list change into this PR. And based on Guilherme's feedback, I'm just skipping the test instead of selectively failing/passing the test. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155056 Approved by: https://github.com/atalman, https://github.com/nWEIdia	2025-06-05 17:17:27 +00:00
Animesh Jain	cadcb5d368	[inductor] disable compiler on the compiled_module_main (#155169 ) Fixes https://github.com/pytorch/pytorch/issues/154536 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155169 Approved by: https://github.com/jamesjwu, https://github.com/bdhirsh	2025-06-05 16:37:45 +00:00
Animesh Jain	13ea0f2c0a	[dynamo][dynamic] Recompilation hint for nn module integer attributes (#154867 ) For program like this ``` class Mod(torch.nn.Module): def __init__(self): super().__init__() self.c = 0 def forward(self, x): self.c += 1 return x * self.c ``` You can check the recompile reasons at https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmpzv9z6Q/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 ![image](https://github.com/user-attachments/assets/856a95fd-0533-4abc-a213-1f73ae2cb766) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154867 Approved by: https://github.com/zou3519	2025-06-05 16:37:22 +00:00
Aaron Gokaslan	a14f427db6	[BE] Update cudnn to 9.10.1.4 (#155122 ) Follow up to #152782 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155122 Approved by: https://github.com/malfet, https://github.com/atalman	2025-06-05 16:07:25 +00:00
atalman	cd361fc247	[CI] Migrate focal (ubuntu 20.04) images to jammy (ubuntu 22.04) (#154437 ) Fixes https://github.com/pytorch/pytorch/issues/154157 Inductor Workflows where moved from focal to jammy here: https://github.com/pytorch/pytorch/pull/154153 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154437 Approved by: https://github.com/Skylion007, https://github.com/cyyever, https://github.com/davidberard98, https://github.com/huydhn	2025-06-05 15:24:07 +00:00
Jane Xu	e895e9689c	Update docs build to specify <3.13 in CONTRIBUTING (#155140 ) Python 3.13 removed the deprecated imghdr module, so our docs build does not compile with 3.13+. Mention it in our contributing guide so people know before committing to the wrong version oop. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155140 Approved by: https://github.com/drisspg, https://github.com/cyyever ghstack dependencies: #155126	2025-06-05 15:16:48 +00:00
Jane Xu	2f3f8339ec	[BE] Document device memory apis in correct module (#155126 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155126 Approved by: https://github.com/msaroufim, https://github.com/Skylion007	2025-06-05 15:16:48 +00:00
Narek Malkhasyan	7999735d23	[CUDA][MPS] Fix torch.arange bound validation for large float inputs (#154320 ) Fixes #153133 Fixes an inconsistency in torch.arange on CUDA and MPS backends when using float32 and large input values. Previously, invalid ranges (e.g., start > end with a positive step) could silently return empty tensors due to precision loss in validation logic. The fix introduces double precision validation for checking whether the step sign is consistent with the range direction. This ensures torch.arange behaves consistently with CPU for large float32 inputs, and raises an appropriate error when the range is invalid. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154320 Approved by: https://github.com/malfet	2025-06-05 14:51:25 +00:00
Nikita Shulga	ed661a5f11	[MPS] Fix complex scalar binding to Metal tensors (#155184 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155184 Approved by: https://github.com/dcci ghstack dependencies: #155150, #155178	2025-06-05 14:34:57 +00:00
kiersten-stokes	9bf6593e96	Fix docstring for `torch.UntypedStorage.from_file` (#155067 ) Fixes #130629 Happy to revert the second commit if we think it's making the test too fragile for the future Pull Request resolved: https://github.com/pytorch/pytorch/pull/155067 Approved by: https://github.com/malfet	2025-06-05 14:30:49 +00:00
PyTorch MergeBot	a1057cda31	Revert "Add CPython generator/contextlib tests (#150796 )" This reverts commit d5f642211f14593c8c78af98a1fb7cfb63039ce5. Reverted https://github.com/pytorch/pytorch/pull/150796 on behalf of https://github.com/guilhermeleobas due to This is breaking tests on trunk. https://hud.pytorch.org/hud/pytorch/pytorch/main/1?per_page=50&name_filter=3.13&mergeEphemeralLF=true ([comment](https://github.com/pytorch/pytorch/pull/150796#issuecomment-2944469866))	2025-06-05 13:51:54 +00:00
wengshiy	196c95d463	Add dont constant fold flag (#154945 ) For support https://github.com/pytorch/ao/issues/2228 > What we want to do now is to enable FP8 quantization in PyTorch. And similar as INT8 quantization, we need to insert quantize and dequantize ops into the graph. > > However we met problems with these q/dq ops both in the PyTorch core and Torchao. > > PyTorch core: > > The quantize_per_tensor op does not support FP8. We want to fix it via https://github.com/pytorch/pytorch/pull/153601. And as you commented, the op is deprecated. > Torchao: > > In the fusion pass in Inductor, we want to match the pattern fp8_weight -> torchao.dequantize_affine_float8 -> fp32_op and fuse it as fp8_weight -> weight_pack -> fp8_op. We have done so for INT8 PT2E quantization. However, the pattern matching pass is applied after a constant folding pass in Inductor: > `100ec0b34a/torch/_inductor/fx_passes/freezing_patterns.py (L69C1-L74C1)` > After constant_fold(gm), the pattern will be folded as fp32_weight -> fp32_op. Then the original pattern cannot be found any more and the FP8 semantics is lost since the pattern is entirely in fp32 now. > For INT8, the int8_weight -> quantized_decomposed.dequantize_per_channel -> fp32_op pattern won't be folded because we mark quantized_decomposed.dequantize_per_channel impure so that it won't be folded: `100ec0b34a/torch/_inductor/constant_folding.py (L139C1-L149C1)` . But for the torchao.dequantize_affine_float8, we cannot do this because > It is an op from Torchao, which is unknown to the constant folder > It is decomposed to smaller ops, so we cannot put it in the list as a single op. > So, we think an easy and short-term solution is to modify the ops in PyTorch core via https://github.com/pytorch/pytorch/pull/153601. > However, if we want to resolve the issue with Torchao, we need to > Add a method in the constant folder in Inductor to allow registration of impure ops Based on [Jansel‘s reply](https://github.com/pytorch/ao/issues/2228#issuecomment-2914560340), add dont constant fold flag on this patch Pull Request resolved: https://github.com/pytorch/pytorch/pull/154945 Approved by: https://github.com/leslie-fang-intel, https://github.com/jansel Co-authored-by: Jason Ansel <jansel@jansel.net>	2025-06-05 13:42:44 +00:00
PyTorch MergeBot	e01fde8213	Revert "[reland][dynamo] Record the pre-graph bytecode using fast record function event (#154974 )" This reverts commit bee9c70c5d4b681ec1f2adf92eca1205b372634a. Reverted https://github.com/pytorch/pytorch/pull/154974 on behalf of https://github.com/malfet due to Broke inductor tests, see `3c72b9fd8f/1` ([comment](https://github.com/pytorch/pytorch/pull/154974#issuecomment-2944370617))	2025-06-05 13:36:21 +00:00
PyTorch MergeBot	3c72b9fd8f	Revert "SDPA support gfx950 (#155103 )" This reverts commit b9312c56bf5f277e341c0185da748e3475d0807f. Reverted https://github.com/pytorch/pytorch/pull/155103 on behalf of https://github.com/malfet due to looks like it broke mi300 tests, see `9a4c08ddfc/1` ([comment](https://github.com/pytorch/pytorch/pull/155103#issuecomment-2944331460))	2025-06-05 13:33:17 +00:00
PyTorch MergeBot	523b637cbe	Revert "[test][dynamo] skip test_deopt_from_append_list on python>=3.13.3 (#155167 )" This reverts commit 1c828786c28b8cd2a6be2397cc2af65e3266c5fa. Reverted https://github.com/pytorch/pytorch/pull/155167 on behalf of https://github.com/malfet due to This broke a bunch of 3.13 tests, see `fa3c38c7ae/1` ([comment](https://github.com/pytorch/pytorch/pull/155167#issuecomment-2944318067))	2025-06-05 13:27:40 +00:00
PyTorch MergeBot	f60b2712dd	Revert "Inductor unit tests: cuda 12.6 -> 12.8 (#155056 )" This reverts commit bb43ced6e2c9e1cdc17923826aaf58466c2ffd4b. Reverted https://github.com/pytorch/pytorch/pull/155056 on behalf of https://github.com/malfet due to This broke a bunch of 3.13 tests, see `fa3c38c7ae/1` ([comment](https://github.com/pytorch/pytorch/pull/155167#issuecomment-2944318067))	2025-06-05 13:27:40 +00:00
Roy Hvaara	9a4c08ddfc	[MPS] Parametrize `test_scaled_dot_product_attention_autocast` (#155005 ) Also moving comments inside the function scope for some of my previous regression tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155005 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-06-05 13:24:53 +00:00
zeshengzong	fa3c38c7ae	Add tensor overlap check for `cross` (#154999 ) Fixes #132031 ## Test Result ```python In [1]: import torch ...: torch.manual_seed(0) ...: torch.cuda.manual_seed(0) ...: a = torch.randn(3, 4) ...: b = torch.randn(3, 4) ...: torch.cross(a, b, out=a) --------------------------------------------------------------------------- RuntimeError Traceback (most recent call last) Cell In[1], line 6 4 a = torch.randn(3, 4) 5 b = torch.randn(3, 4) ----> 6 torch.cross(a, b, out=a) RuntimeError: unsupported operation: some elements of the input tensor and the written-to tensor refer to a single memory location. Please clone() the tensor before performing the operation. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154999 Approved by: https://github.com/lezcano	2025-06-05 10:00:01 +00:00
Ivan Zaitsev	5b65628906	Workflow to tag trunk commits with `trunk/{commit-sha}` tags (#155170 ) This PR adds workflow to automate tagging commits on the `main` branch. The workflow includes validation and retry with exponential backoff. The rationale for this is to work around the github limitation on using workflow_dispatch (requires branch or tag). We want to use workflow_dispatch to rerun CI workflows with parameters (trunk, pull, etc). --- ### Testing Tested using almost identical workflow in a personal repo (the difference is in repository_owner check and backoff settings). * successful tag push: https://github.com/izaitsevfb/deleteme/actions/runs/15454729765/job/43504630765 * validation: PR commit (fails) https://github.com/izaitsevfb/deleteme/actions/runs/15454743572/job/43504669720 * tagging of the old commit on main: https://github.com/izaitsevfb/deleteme/actions/runs/15453805748/job/43501885903 * tag already exists: https://github.com/izaitsevfb/deleteme/actions/runs/15454756077/job/43504706980 * invalid sha on workflow dispatch: https://github.com/izaitsevfb/deleteme/actions/runs/15454611077/job/43504286858 * retry with exponential backoff on failure (via tag rule blocklist): https://github.com/izaitsevfb/deleteme/actions/runs/15454768346/job/43504743486 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155170 Approved by: https://github.com/huydhn	2025-06-05 09:50:58 +00:00
Animesh Jain	bee9c70c5d	[reland][dynamo] Record the pre-graph bytecode using fast record function event (#154974 ) reland of https://github.com/pytorch/pytorch/pull/154769 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154974 Approved by: https://github.com/Lucaskabela, https://github.com/jansel	2025-06-05 07:25:04 +00:00
Boyuan Feng	be16f21ca6	[Graph Partition] add symints to get_graph_inputs (#154679 ) During `codegen_inputs`, we check whether there are undefined symbols: `65b1aedd09/torch/_inductor/codegen/wrapper.py (L1668-L1674)` Previously, for graph partition inputs, we do not explicitly add symints. `65b1aedd09/torch/_inductor/codegen/wrapper.py (L3265-L3272)` We relied on sizes/strides of TensorBox for codegen symint inputs. For example, a tensor with shape `[s0, 2]` will implicitly codegen `s0` as an input here. This works fine in most cases since backed symint has to come from some tensor shapes. `65b1aedd09/torch/_inductor/codegen/wrapper.py (L1624-L1632)` In rare cases, this does not work. One example is saved tensors for backward where a tensor may have shape `[2s0, 2]`. Since `2s0` is an expression but not a symbol, `codegen_input_symbol_assignment` would not handle `s0` and later there would be an error when `_verify_input_symbol_assignment`. The fix is add symints to `get_graph_inputs`. An alternative way is to update `codegen_input_symbol_assignment` but I want to minimize the change to graph partition only. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154679 Approved by: https://github.com/eellison	2025-06-05 06:46:28 +00:00
PyTorch MergeBot	d3c8f36ba0	Revert "[Intel GPU] Make SDPA output has the same stride as Query. (#154340 )" This reverts commit 0f10df71a66cb1b0c3659381b7db8e06d95f0d67. Reverted https://github.com/pytorch/pytorch/pull/154340 on behalf of https://github.com/etaf due to This PR breaks hugging face E2E run on XPU. ([comment](https://github.com/pytorch/pytorch/pull/154340#issuecomment-2942954192))	2025-06-05 06:46:24 +00:00
David Berard	bb43ced6e2	Inductor unit tests: cuda 12.6 -> 12.8 (#155056 ) Fixes #154938 When we update the Triton version in CI, we'll require cuda >= 12.8 for certain AOTI tests to pass: these AOTI tests try to run nvcc on the triton-generated PTX, and triton-generated PTX is PTX 8.7, which requires CUDA 12.8 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155056 Approved by: https://github.com/atalman, https://github.com/nWEIdia ghstack dependencies: #155167	2025-06-05 05:59:06 +00:00
David Berard	1c828786c2	[test][dynamo] skip test_deopt_from_append_list on python>=3.13.3 (#155167 ) Not sure why, apparently this test starts passing on python 3.13.3 (while it fails on python <=3.13.2) and it's causing unexpected passes on xfail-ed tests when newer versions of python are used, e.g. in #155056. Verified locally in a python 3.13.1 vs. python 3.13.3 conda env. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155167 Approved by: https://github.com/williamwen42	2025-06-05 05:59:06 +00:00
PyTorch MergeBot	93012d2290	Revert "[forward fix] add support for MemoryFormat after type tightening (#154658 )" This reverts commit 0fdd568b785812da86e69d65632de77d2ee945c7. Reverted https://github.com/pytorch/pytorch/pull/154658 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/154658#issuecomment-2942752048))	2025-06-05 05:01:40 +00:00
PyTorch MergeBot	5130ac64f4	Revert "Add randint_like tensor overload for high (#154899 )" This reverts commit 72fe1d5f42aa9bffa876932a3b4fcae052b99168. Reverted https://github.com/pytorch/pytorch/pull/154899 on behalf of https://github.com/seemethere due to Failing internal tests see https://fburl.com/diff/bai044ob ([comment](https://github.com/pytorch/pytorch/pull/154899#issuecomment-2942740661))	2025-06-05 04:54:05 +00:00
drisspg	80703ca332	[FlexAttention] Allow dispatch to SAC for flex (#150080 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/150080 Approved by: https://github.com/zou3519	2025-06-05 04:34:27 +00:00
James Wu	fa63de0866	Handle empty linemaps in PyCodeCache (#155064 ) Some functions have empty linemaps, and if you call `PyCodeCache.stack_frames_for_code` on code in the wrong order, you'll end up triggering a too many values to unpack issue: https://github.com/pytorch/pytorch/issues/154536 Specifically, if you populate PyCodeCache's linemap via caching, and then request the stack frames of a inductor generated output file that has an empty linemap, this function will try to unpack too many arguments. Test plan: ``` import os os.environ["TORCHINDUCTOR_FX_GRAPH_CACHE"] = "1" os.environ["TORCHINDUCTOR_AUTOGRAD_CACHE"] = "1" import torch @torch.compile def fn(x: torch.Tensor): (x_grad,) = torch.autograd.grad(x.sum(), x) return x_grad x = torch.randn(10, 10, requires_grad=True) result = fn(x) ``` Run this twice and see that everything works as expected. It's hard to exactly pinpoint a good unit test for this: it requires a whole lot of moving parts to get the issue to trigger because: - The callsite in question in dynamo, without caching, will always run before generating the code, so cls.linemaps[path] will be None most of the time - The inductor generated output needs to call back into dynamo via `assert_size_stride` - In our test case, the CompiledBackward needs to not have linemaps, and also be called in the middle of a graph break while compiling a different cached function. Caching switches the order the PyCodeCache.linemap is populated (i.e. either before or after the graph break is evaluated), which causes the issue. All these things need to interact together to create the bug, so it's a bit difficult to write a simple unit test. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155064 Approved by: https://github.com/bdhirsh	2025-06-05 03:54:35 +00:00
fduwjj	450180fbcd	[c10d][fr] Add the log of thread name and thread id into fr (#155142 ) There is an ask from internal head users to have thread id and thread name inside fr. This would be useful to users when it comes to cases when we launches collectives not just on main thread as well. Differential Revision: [D75973919](https://our.internmc.facebook.com/intern/diff/D75973919) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155142 Approved by: https://github.com/kwen2501	2025-06-05 03:33:01 +00:00
Xiaodong Wang	b9312c56bf	SDPA support gfx950 (#155103 ) Summary: Seems to run, just not the optimal performance. e.g. ck_tile doesn't have those gfx942 optimizations it seems https://github.com/ROCm/composable_kernel/blob/develop/include/ck_tile/ops/fmha/block/variants.hpp#L27 Test Plan: ``` +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| Batch Size \| Sequence Length \| Heads \| Head Dim \| Flash Time (µs) \| Mem Eff Time (µs) \| Math Time (µs) \| Flex Time (µs) \| xformers Time (µs) \| Flash TFlops \| Mem Eff TFlops \| Math TFlops \| Flex TFlops \| xformers TFlops \| Speedup (Flash/Math) \| Speedup (MemEff/Math) \| Speedup (Flex/Math) \| Speedup (xformers/Math) \| xformers trace_url \| Flash trace_url \| +==============+===================+=========+============+===================+=====================+==================+==================+======================+================+==================+===============+===============+===================+========================+=========================+=======================+===========================+======================+===================+ \| 1 \| 4096 \| 16 \| 64 \| 179.737 \| 182.874 \| 3106.6 \| 359.662 \| 205.506 \| 382.334 \| 375.776 \| 22.1205 \| 191.067 \| 334.392 \| 17.2841 \| 16.9877 \| 8.63754 \| 15.1169 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 4096 \| 32 \| 128 \| 617.271 \| 623.38 \| 7169.73 \| 998.961 \| 654.534 \| 445.312 \| 440.947 \| 38.3387 \| 275.164 \| 419.96 \| 11.6152 \| 11.5014 \| 7.17719 \| 10.9539 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 8192 \| 16 \| 64 \| 667.032 \| 670.118 \| 13031.8 \| 1383.42 \| 768.452 \| 412.091 \| 410.193 \| 21.0928 \| 198.694 \| 357.703 \| 19.5371 \| 19.4471 \| 9.42 \| 16.9586 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 8192 \| 32 \| 128 \| 2074.64 \| 2214.81 \| 29186.9 \| 3916.35 \| 2404.29 \| 529.978 \| 496.437 \| 37.6714 \| 280.749 \| 457.313 \| 14.0684 \| 13.1781 \| 7.45257 \| 12.1395 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 16384 \| 16 \| 64 \| 2456.6 \| 2472.38 \| 51095.8 \| 5647.01 \| 3008.09 \| 447.574 \| 444.718 \| 21.5186 \| 194.707 \| 365.518 \| 20.7994 \| 20.6666 \| 9.0483 \| 16.9861 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 16384 \| 32 \| 128 \| 8048.8 \| 8070.96 \| 113478 \| 15580.8 \| 9768.71 \| 546.423 \| 544.922 \| 38.7569 \| 282.274 \| 450.218 \| 14.0987 \| 14.06 \| 7.2832 \| 11.6165 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| Batch Size \| Sequence Length \| Heads \| Head Dim \| Flash Time (µs) \| Mem Eff Time (µs) \| Math Time (µs) \| Flex Time (µs) \| xformers Time (µs) \| Flash TFlops \| Mem Eff TFlops \| Math TFlops \| Flex TFlops \| xformers TFlops \| Speedup (Flash/Math) \| Speedup (MemEff/Math) \| Speedup (Flex/Math) \| Speedup (xformers/Math) \| xformers trace_url \| Flash trace_url \| +==============+===================+=========+============+===================+=====================+==================+==================+======================+================+==================+===============+===============+===================+========================+=========================+=======================+===========================+======================+===================+ \| 1 \| 4096 \| 16 \| 64 \| 692.323 \| 697.649 \| 4241.81 \| 1562.26 \| 906.441 \| 248.148 \| 246.254 \| 40.5012 \| 109.968 \| 189.531 \| 6.12693 \| 6.08015 \| 2.71518 \| 4.67963 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 4096 \| 32 \| 128 \| 2263.22 \| 2267.38 \| 9482.64 \| 7003.8 \| 2765.5 \| 303.636 \| 303.079 \| 72.4687 \| 98.1174 \| 248.489 \| 4.1899 \| 4.18221 \| 1.35393 \| 3.42891 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 8192 \| 16 \| 64 \| 2553.94 \| 2572.68 \| 15909.8 \| 5697.16 \| 3284.77 \| 269.073 \| 267.112 \| 43.193 \| 120.621 \| 209.206 \| 6.22953 \| 6.18415 \| 2.79259 \| 4.84352 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 8192 \| 32 \| 128 \| 8187.67 \| 8201.71 \| 35449.2 \| 26424.3 \| 10364.5 \| 335.722 \| 335.147 \| 77.5413 \| 104.025 \| 265.21 \| 4.32959 \| 4.32218 \| 1.34154 \| 3.42025 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 16384 \| 16 \| 64 \| 9948.15 \| 9815.47 \| 62815.1 \| 23741.9 \| 12710 \| 276.31 \| 280.046 \| 43.7598 \| 115.778 \| 216.269 \| 6.31425 \| 6.39961 \| 2.64575 \| 4.94217 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ \| 1 \| 16384 \| 32 \| 128 \| 32187.6 \| 32035.6 \| 137832 \| 102075 \| 40623.4 \| 341.595 \| 343.216 \| 79.7716 \| 107.716 \| 270.66 \| 4.28216 \| 4.30248 \| 1.35031 \| 3.39293 \| \| \| +--------------+-------------------+---------+------------+-------------------+---------------------+------------------+------------------+----------------------+----------------+------------------+---------------+---------------+-------------------+------------------------+-------------------------+-----------------------+---------------------------+----------------------+-------------------+ ``` Rollback Plan: Differential Revision: D75934358 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155103 Approved by: https://github.com/yoyoyocmu	2025-06-05 03:26:38 +00:00
Wei Wang	a01bb9da14	[CI][CUDA] Re-enable the test-nan-assert on CUDA12 (#154448 ) We need to reenable this test because there are recent changes that could be relevant to test_nan_assert. I've already tested that there would be hang if we don't remove the "pg._allgather_base(output, nan_tensor)" in between the "backend._set_enable_nan_check" calls. Why was it "working" previously? Because previously only cu118 distributed was running and this "backend._set_enable_nan_check" change was not tested in the merge process (skip logic is if "not CUDA 12 and above", skip). Workaround #153479 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154448 Approved by: https://github.com/kwen2501	2025-06-05 02:09:31 +00:00
PyTorch MergeBot	5e03433443	Revert "Inductor logging + analysis of torch.profile (#149697 )" This reverts commit e5afbe31245287a92fe328c404b3557e5c5eca73. Reverted https://github.com/pytorch/pytorch/pull/149697 on behalf of https://github.com/malfet due to Broke rocm, see `642687af29/1` ([comment](https://github.com/pytorch/pytorch/pull/149697#issuecomment-2942415600))	2025-06-05 01:38:13 +00:00
Nikita Shulga	642687af29	[MPS][BE] Some refactor in preparation for 64-bit iterators (#155178 ) set input/output tensors only once Get rid of `is_storage_dense` predicate, as `iter.is_contiguous` serves the same purpose Pull Request resolved: https://github.com/pytorch/pytorch/pull/155178 Approved by: https://github.com/dcci, https://github.com/cyyever ghstack dependencies: #155150	2025-06-05 01:24:31 +00:00
Laith Sakka	3398d1d459	support bmm and mm_plus_mm in generated templates cache (#154904 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154904 Approved by: https://github.com/drisspg, https://github.com/eellison, https://github.com/jansel ghstack dependencies: #154891, #154892	2025-06-05 00:36:01 +00:00
Guilherme Leobas	21f45f7afb	Add CPython int/float tests (#150795 ) Tests: * test_int.py * test_int_literal.py * test_float.py Minor changes were made to each test to run them inside Dynamo One can reproduce the changes by downloading the tests from CPython and applying the diff: ```bash for f in "test_int" "test_int_literal" "test_float"; do wget -O "test/dynamo/cpython/3_13/${f}.py" "https://raw.githubusercontent.com/python/cpython/refs/heads/3.13/Lib/test/${f}.py" git apply "test/dynamo/cpython/3_13/${f}.diff" done ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150795 Approved by: https://github.com/williamwen42	2025-06-05 00:28:53 +00:00
Guilherme Leobas	d5f642211f	Add CPython generator/contextlib tests (#150796 ) Tests: * test_generator.py * test_generator_stop.py * test_contextlib.py Minor changes were made to each test to run them inside Dynamo. We intentionally didn't copy the binary files stored in `python/Lib/test/archivetestdata` for security reasons. There's a single test that requires a binary file and it is skipped because of that. The tests were downloaded from CPython 3.13 and the diff was generated using `git diff` to apply the changes: ```bash for f in "test_contextlib" "test_generators" "test_generator_stop"; do wget -O "test/dynamo/cpython/3_13/${f}.py" "https://raw.githubusercontent.com/python/cpython/refs/heads/3.13/Lib/test/${f}.py" git apply "test/dynamo/cpython/3_13/${f}.diff" done ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150796 Approved by: https://github.com/williamwen42	2025-06-05 00:18:29 +00:00
Thomas Bohnstingl	fb5a787a8f	[HOP] Added clone for outputs of create_bw_fn that are aliasing the inputs (#153932 ) This PR fixes an issue with the new way of creating the bw graph introduced for cond. In particular, there is an issue if the bw function simply aliases the inputs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153932 Approved by: https://github.com/ydwu4	2025-06-04 23:52:52 +00:00
Laith Sakka	b0a2ca65ef	support more prologue functions in generated templates cache (#154892 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154892 Approved by: https://github.com/jansel, https://github.com/eellison ghstack dependencies: #154891	2025-06-04 23:45:36 +00:00
Laith Sakka	51b4c51973	add missing check for caching triton template caching (#154891 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154891 Approved by: https://github.com/eellison	2025-06-04 23:45:36 +00:00
Shivam Raikundalia	1083bc749d	[Memory Snapshot] Add Flag to Toggle Global and Local Callbacks for Annotations (#154932 ) Summary: There are some cases where we want only local annotations for memory snapshot such as executing inside the cudastream callback, which cannot execute CUDA operators. Thus the cuda errors happen: Exception in RecordFunction callback: CUDA error: operation not permitted However, we need to have an option to turn on the globally so that on-demand snapshot can get annotations. Additionally, there may be some cases in which auto-trace will also want annotations using record functions so we expose the flag to the auto-trace as well. Test Plan: Run MVAI executable and see that the errors go away Rollback Plan: Differential Revision: D75831687 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154932 Approved by: https://github.com/mzzchy, https://github.com/sanrise	2025-06-04 23:15:19 +00:00
Tushar Jain	7cf5b36ec2	Release GIL in PG destructor (#154976 ) Summary: Gloo PG doesn't release GIL, which results in python code hanging until the destructor completes. The destructor waits for all work on the PG to complete which can take a long time. Test Plan: Ran ``` $ pytest --log-cli-level=INFO -vs torchft/local_sgd_integ_test.py ``` with a large timeout on the async work. Call to `gil_scoped_release` doesn't show up in the gdb stack trace. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154976 Approved by: https://github.com/d4l3k, https://github.com/dcci, https://github.com/fduwjj	2025-06-04 23:10:55 +00:00
Animesh Jain	c881f2ddf3	[reland][dynamo] Mark a vt unspecialized nn module variable source earlier (#155099 ) Reland of https://github.com/pytorch/pytorch/pull/154780 Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/155099 Approved by: https://github.com/williamwen42	2025-06-04 23:05:36 +00:00
Nikita Shulga	992be94dab	[MPS][BE] Better error messages (#155150 ) "Can't be indexed using 32-bit iterator" is not really helpful error This PR distinguishes between error from old indexing helper function as well as to binaryTensorIterator Adds the same warning to unary op, otherwise it just runs and returns incorrect value Test plan (manual, don't have machine with enough RAM to run it reliable in CI): ``` % python -c "import torch;print(torch.rand(1, 1024, 1024, dtype=torch.bfloat16, device='mps') + torch.rand(5000, 1, 1, dtype=torch.bfloat16, device='mps'))" RuntimeError: add can't be indexed using 32-bit iterator for shape [1048576, 5000] ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155150 Approved by: https://github.com/Skylion007, https://github.com/dcci	2025-06-04 22:53:51 +00:00
Mu-Chu Lee	f5e2e4c4f1	[Inductor] Include math and torch in launcher scope (#154673 ) Summary: For grid computation, if we have sympy, it is possible we have math and torch used. We include the math and torch module in the launcher scope to make sure those grid get computed correctly. Test Plan: Check phabricator for internal cmd. Differential Revision: D75642931 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154673 Approved by: https://github.com/Skylion007, https://github.com/davidberard98	2025-06-04 22:32:19 +00:00
Mikayla Gawarecki	671553bd23	Update documentation wording for transformer-related layers (#155123 ) <img width="947" alt="Screenshot 2025-06-04 at 1 33 53 PM" src="https://github.com/user-attachments/assets/4dbb66b3-43f4-4d04-afb5-dc80cec0f2cd" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/155123 Approved by: https://github.com/albanD, https://github.com/jbschlosser	2025-06-04 22:20:32 +00:00
Sidharth	6f23ca53bb	[dynamo] sample gb_registry json file for website testing purposes (#155160 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/155160 Approved by: https://github.com/StrongerXi, https://github.com/williamwen42	2025-06-04 22:14:48 +00:00
angelayi	c8566a0b98	[export] Use patching in test (#155132 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/155132 Approved by: https://github.com/pianpwk	2025-06-04 21:41:26 +00:00
ANotFox	65a5eb8d27	Fix for ambiguity in linalg.norm()'s ord argument of +2 & -2 (#155148 ) Fixes #136453 ### Description --- Fixed the ambiguity by referencing a hyperlink to wikipedia's SVD/Singular Values section as per past discussion (by other contributors) on the above thread. In the ord argument, for values `+2` and `-2`, the `singular value` now points to [this section of singular values on the wiki SVD page](https://en.wikipedia.org/wiki/Singular_value_decomposition#Singular_values,_singular_vectors,_and_their_relation_to_the_SVD). ### Why not mention SVD --- For conciseness (expanding 'largest singular value' -> 'largest singular value of a SVD' is too much, i think, wrt rest of the table) --- I hope this is satisfactory. Please let me know if I have missed anything essential; cheers. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155148 Approved by: https://github.com/Skylion007, https://github.com/lezcano	2025-06-04 21:15:20 +00:00
Thomas Bohnstingl	b084e1b81c	[HOP] Rework Autograd DispatchKey for scan and map (#153336 ) This PR introduces the `py_autograd_impl` instead of the `DispatchKey.Autograd` for some HOPs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153336 Approved by: https://github.com/ydwu4	2025-06-04 20:54:02 +00:00
Sidharth	0404785f3b	[dynamo] [3/3] added cmd_update_gb_type which supports updating an existing gb_type properties and optional arg to change gb_type name (#154985 ) The user can now use the terminal to update the registry whenever they update an existing gb_type's properties. Additionally, if the user changes the gb_type description itself, they can update the registry as well. Terminal command template for updating existing gb_type: python [path to gb_id_mapping.py] update "existing_gb_type" [path to file where user added callsite] Terminal command template for updating existing gb_type name (can also be used if the user changed the other properties as well including the gb_type name): python [path to gb_id_mapping.py] update "existing_gb_type" [path to file where user added callsite] --new_gb_type "new_name_for_existing_gb_type" Pull Request resolved: https://github.com/pytorch/pytorch/pull/154985 Approved by: https://github.com/williamwen42	2025-06-04 20:10:02 +00:00
Gabriel Ferns	e5afbe3124	Inductor logging + analysis of torch.profile (#149697 ) Prereqs: - https://github.com/pytorch/pytorch/pull/152708 Features: 1. Adds inductor's estimate of flops and bandwidth to the json trace events that perfetto uses. 1. Only use the tflops estimation from triton if we don't have the info from the datasheet because Triton's estimates are inaccurate. I have a backlog item to fix triton flops estimation upstream. New `DeviceInfo` class, and new function `get_device_tflops`. 1. New helpers `countable_fx` and `count_flops_fx` helps get the flops of an `fx.Node`. 1. Extends Triton `torch.profiler` logging to `DebugAutotuner`. 1. New script `profile_analysis.py`: `--augment_trace` adds perf estimates to any perfetto json trace, `--analyze` creates a summary table of these perf estimates, and `--diff` will compare two traces side by side: ```python Device(NVIDIA H100, 0): Kernel Name \| resnet Kernel Count \| resnet FLOPS \| resnet bw gbps \| resnet Dur (ms) \| resnet Achieved FLOPS % \| resnet Achieved Bandwidth % \| newresnet Kernel Count \| newresnet FLOPS \| newresnet bw gbps \| newresnet Dur (ms) \| newresnet Achieved FLOPS % \| newresnet Achieved Bandwidth % --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- triton_poi_fused__native_batch_norm_legi \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 \| 24 \| 0 \| 0.11395268248131513 \| 2.5919166666666666 \| 0 \| 0.003401572611382541 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 \| 142 \| 16932673552.422373 \| 0.2585007824198784 \| 12.441619718309857 \| 0.08683422334575583 \| 0.007716441266265022 triton_red_fused__native_batch_norm_legi \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 \| 39 \| 0 \| 0.13990024992108846 \| 5.752589743589743 \| 0 \| 0.004176126863316074 triton_poi_fused__native_batch_norm_legi \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 \| 25 \| 0 \| 0.31824055917536503 \| 2.5291999999999994 \| 0 \| 0.009499718184339253 void cutlass::Kernel2<cutlass_80_tensoro \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 \| 98 \| 16211056473.596165 \| 0.42972434051025826 \| 7.130408163265306 \| 0.08313362294151874 \| 0.012827592254037562 triton_red_fused__native_batch_norm_legi \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 \| 73 \| 0 \| 0.3225381327611705 \| 9.987068493150682 \| 0 \| 0.009628003963020014 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 \| 15 \| 0 \| 1.4491211346487216 \| 4.439333333333333 \| 0 \| 0.043257347302946926 void cutlass::Kernel2<cutlass_80_tensoro \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 \| 186 \| 14501701145.337954 \| 0.2667131401910989 \| 7.873865591397849 \| 0.07436769818122027 \| 0.007961586274361157 triton_poi_fused__native_batch_norm_legi \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 \| 33 \| 0 \| 1.4924556538193923 \| 4.3101515151515155 \| 0 \| 0.044550915039384846 triton_red_fused__native_batch_norm_legi \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 \| 29 \| 0 \| 0.25562590522631107 \| 6.296275862068965 \| 0 \| 0.007630624036606301 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 \| 13 \| 0 \| 0.5870562174192726 \| 2.7397692307692307 \| 0 \| 0.01752406619162008 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 \| 34 \| 0 \| 0.41409928846284 \| 2.853588235294117 \| 0 \| 0.012361172789935523 triton_per_fused__native_batch_norm_legi \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 \| 34 \| 0 \| 0.11705315007018151 \| 3.460647058823529 \| 0 \| 0.0034941238826919864 triton_poi_fused__native_batch_norm_legi \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 \| 16 \| 0 \| 0.17207853197124584 \| 2.3459375000000002 \| 0 \| 0.005136672596156592 triton_per_fused__native_batch_norm_legi \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 \| 30 \| 0 \| 0.2639714322022256 \| 6.131199999999999 \| 0 \| 0.007879744244842555 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 \| 100 \| 11875430356.891787 \| 0.19494470869421385 \| 16.36534 \| 0.06089964285585531 \| 0.005819245035648175 triton_poi_fused__native_batch_norm_legi \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 \| 8 \| 0 \| 0.9854096626224687 \| 3.2757500000000004 \| 0 \| 0.029415213809625928 void cublasLt::splitKreduce_kernel<32, 1 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 \| 56 \| 34377923395.147064 \| 0.8310300045762317 \| 3.4199999999999986 \| 0.17629704305203628 \| 0.024806865808245714 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 \| 23 \| 0 \| 0.9944002965861103 \| 3.2431304347826084 \| 0 \| 0.02968359094286896 triton_per_fused__native_batch_norm_legi \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 \| 10 \| 0 \| 0.1826801058931057 \| 4.428800000000001 \| 0 \| 0.00545313748934644 triton_poi_fused__native_batch_norm_legi \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 \| 10 \| 0 \| 0.3168973585366449 \| 2.5471999999999997 \| 0 \| 0.009459622642884923 triton_poi_fused__native_batch_norm_legi \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 \| 34 \| 0 \| 1.1463614897015777 \| 4.124323529411764 \| 0 \| 0.03421974596124114 void cask_plugin_cudnn::xmma_cudnn::init \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 \| 44 \| 44045510816.64277 \| 2.0661232850348643 \| 3.6887499999999993 \| 0.22587441444432194 \| 0.06167532194133924 sm90_xmma_fprop_implicit_gemm_f32f32_tf3 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 \| 95 \| 7876855400.165316 \| 0.4694941555946739 \| 18.224315789473682 \| 0.04039413025725802 \| 0.014014750913273854 triton_per_fused__native_batch_norm_legi \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 \| 41 \| 0 \| 0.06825669875995298 \| 3.0384146341463416 \| 0 \| 0.002037513395819492 triton_poi_fused__native_batch_norm_legi \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 \| 23 \| 0 \| 0.08808154712430301 \| 2.3275652173913044 \| 0 \| 0.0026292999141582997 triton_per_fused__native_batch_norm_legi \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 \| 40 \| 0 \| 0.18179321034952417 \| 4.556825 \| 0 \| 0.005426662995508183 triton_poi_fused__native_batch_norm_legi \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 \| 15 \| 0 \| 0.5887415155454232 \| 2.783866666666667 \| 0 \| 0.017574373598370836 void cutlass::Kernel2<cutlass_80_tensoro \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 \| 38 \| 14242013806.264643 \| 0.256592404353939 \| 7.217631578947369 \| 0.0730359682372546 \| 0.007659474756834 triton_poi_fused__native_batch_norm_legi \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 \| 21 \| 0 \| 0.5842860973430516 \| 2.7779047619047623 \| 0 \| 0.017441376040091088 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 \| 16 \| 0 \| 0.11509365173486417 \| 3.5959375000000002 \| 0 \| 0.0034356313950705724 triton_poi_fused__native_batch_norm_legi \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 \| 14 \| 0 \| 0.1704672000243914 \| 2.4044285714285714 \| 0 \| 0.00508857313505646 triton_poi_fused__native_batch_norm_legi \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 \| 58 \| 0 \| 2.307520779930795 \| 8.190706896551722 \| 0 \| 0.06888121731136704 triton_per_fused__native_batch_norm_legi \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 \| 29 \| 0 \| 0.037243248971881276 \| 3.0277586206896556 \| 0 \| 0.001111738775280038 triton_poi_fused__native_batch_norm_legi \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 \| 20 \| 0 \| 0.04741699795428918 \| 2.2911500000000005 \| 0 \| 0.0014154327747549007 triton_per_fused__native_batch_norm_legi \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 \| 25 \| 0 \| 0.13357016893727824 \| 3.37536 \| 0 \| 0.003987169222008305 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 \| 13 \| 0 \| 0.3089862268300253 \| 2.8111538461538457 \| 0 \| 0.009223469457612694 triton_poi_fused__native_batch_norm_legi \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 \| 17 \| 0 \| 0.3129385387909844 \| 2.673 \| 0 \| 0.009341448919133863 triton_per_fused__native_batch_norm_legi \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 \| 19 \| 0 \| 0.2215568162533158 \| 3.8837368421052636 \| 0 \| 0.0066136363060691275 std::enable_if<!(false), void>::type int \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 \| 23 \| 504916805.19297093 \| 1.0118296096314707 \| 8.113913043478261 \| 0.0025893169497075447 \| 0.030203868944223014 triton_poi_fused_add_copy__38 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 \| 56 \| 0 \| 0 \| 2.132482142857143 \| 0 \| 0 triton_poi_fused_convolution_0 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 \| 18 \| 0 \| 0.43458610794936897 \| 2.773333333333334 \| 0 \| 0.012972719640279667 triton_poi_fused_convolution_1 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 \| 17 \| 0 \| 0.028816312469162712 \| 2.6145882352941174 \| 0 \| 0.0008601884319153051 void convolve_common_engine_float_NHWC<f \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 \| 44 \| 8641868995.31118 \| 0.024730540008465626 \| 25.87327272727273 \| 0.04431727689903169 \| 0.0007382250748795709 triton_per_fused__native_batch_norm_legi \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 \| 12 \| 0 \| 0.6809930918986744 \| 4.82675 \| 0 \| 0.020328151996975356 triton_per_fused__native_batch_norm_legi \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 \| 14 \| 0 \| 0.02883030597936608 \| 2.6651428571428575 \| 0 \| 0.0008606061486377935 triton_per_fused__native_batch_norm_legi \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 \| 16 \| 0 \| 0.0014658988233201874 \| 2.098 \| 0 \| 4.375817383045335e-05 triton_poi_fused__native_batch_norm_legi \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 \| 13 \| 0 \| 0.9926297180284697 \| 3.2367692307692306 \| 0 \| 0.02963073785159611 triton_poi_fused__native_batch_norm_legi \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 \| 9 \| 0 \| 1.3008817095666507 \| 3.0863333333333336 \| 0 \| 0.03883228983781048 void at::native::(anonymous namespace):: \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 \| 98 \| 0 \| 0.09174335613709389 \| 4.408520408163265 \| 0 \| 0.0027386076458833994 void at::native::vectorized_elementwise_ \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 \| 7 \| 0 \| 0 \| 1.7278571428571428 \| 0 \| 0 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/149697 Approved by: https://github.com/eellison, https://github.com/shunting314	2025-06-04 20:03:46 +00:00
Ben Koopman	4d576442e9	Fix incorrect get_default_qat_qconfig in prepare_qat_fx docs. (#155100 ) Fixes #144522 ## Description FX QAT docs for prepare_qat_fx incorrectly used get_default_qat_qconfig when it should use get_default_qat_qconfig_mapping for a qconfig_mapping. Previous example code incorrectly used `get_default_qat_qconfig`, resulting in a qconfig being incorrectly passed to `prepare_qat_fx`. `prepare_qat_fx` requires a `qconfig_mapping`, not a single `qconfig`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/155100 Approved by: https://github.com/jerryzh168	2025-06-04 18:51:40 +00:00
Sidharth	6c8241c089	[dynamo] [2/3] added add_new_gb_type functionality (#154886 ) The user can now use the terminal to update the registry whenever they create a new unimplemented_v2() callsite. Terminal command template: python [path to gb_id_mapping.py] add "new_gb_type" [path to file where user added callsite] Before the user added a new gb_type: <img width="619" alt="Screenshot 2025-06-02 at 1 33 54 PM" src="https://github.com/user-attachments/assets/7258cab1-a184-4200-9d56-7b21d243d6d8" /> After the user added a new gb_type: <img width="366" alt="Screenshot 2025-06-02 at 1 34 47 PM" src="https://github.com/user-attachments/assets/5c383e94-268c-4f6d-9111-7b18c856222e" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/154886 Approved by: https://github.com/williamwen42 ghstack dependencies: #154738	2025-06-04 18:44:37 +00:00
Sidharth	681a8189d7	[dynamo] [1/3] updated gbid mapping for initial registry creation (#154738 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154738 Approved by: https://github.com/williamwen42	2025-06-04 18:44:37 +00:00
Bin Bao	197080337b	[AOTI] Extend torchgen to generate C shim with version number (#147745 ) Summary: While it is ok to add a new arg with defaul value to a fallback op in Python, it will be BC-breaking for the C shim. This PR adds an automatic approach to update C shim files when specifying a version number with a list of new args for the modified op. See https://github.com/pytorch/pytorch/pull/154848 as an example on how to do that. Pull Request resolved: https://github.com/pytorch/pytorch/pull/147745 Approved by: https://github.com/yushangdi	2025-06-04 18:40:34 +00:00
Shangdi Yu	1d67849e43	[AOTInductor] Activate CPU test for package and update weights (#155078 ) Summary: looks like CPU is enabled for update_constant_buffer in D71177509 enable these tests as well. Test Plan: ``` buck2 test @//mode/dev-nosan //caffe2/test/inductor:aot_inductor_package -- -r "test_package_without_weight" -v buck2 test @//mode/dev-nosan //caffe2/test/inductor:aot_inductor_package -- -r "test_package_user_managed_weight" -v buck2 test @//mode/dev-nosan //caffe2/test/inductor:aot_inductor_package -- -r "test_update_weights" -v ``` Rollback Plan: Differential Revision: D75908993 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155078 Approved by: https://github.com/angelayi	2025-06-04 17:57:20 +00:00
fduwjj	956716880f	[c10d][gloo] Enable using c10::Half for gloo (#153862 ) Testing with https://github.com/pytorch/gloo/pull/446 and we see that the numerical issues reported in https://github.com/pytorch/pytorch/issues/152300 is indeed resolved and we added a unit test for it. Also update submodule gloo to reflect the change on the gloo side. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153862 Approved by: https://github.com/d4l3k, https://github.com/clee2000, https://github.com/malfet	2025-06-04 17:53:08 +00:00
Xuan Zhang	9eb7e67727	[PT2][memory] correct wait tensor output size (#153569 ) This PR correctly handles the output buffer size of wait tensor nodes. ![image](https://github.com/user-attachments/assets/fdcc5eb7-58cf-42a2-84b2-ce949cb9db92) See [this doc](https://docs.google.com/document/d/1lkKulwIb-fYL_p8jn1SD6Lh1PoAKBgpBsU5sAH80leI/edit?tab=t.0#bookmark=id.w3n4k1y4rdz8) with testing details [internal only] Pull Request resolved: https://github.com/pytorch/pytorch/pull/153569 Approved by: https://github.com/eellison	2025-06-04 17:49:25 +00:00
Ke Wen	34c6371d24	Add NVSHMEM to PYTORCH_EXTRA_INSTALL_REQUIREMENTS (#154568 ) NVSHMEM 3.2.5 (released Mar 2025) have both cu11 and cu12 builds. See: https://pypi.nvidia.com/nvidia-nvshmem-cu12/ https://pypi.nvidia.com/nvidia-nvshmem-cu11/ Pull Request resolved: https://github.com/pytorch/pytorch/pull/154568 Approved by: https://github.com/atalman ghstack dependencies: #154538	2025-06-04 17:43:24 +00:00
James Wu	b3e666ae17	[easy] Bump STATIC_CUDA_LAUNCHER_VERSION=1 (#154861 ) This turns on STATIC_CUDA_LAUNCHER internally for a some low risk entitlements. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154861 Approved by: https://github.com/Skylion007, https://github.com/eellison	2025-06-04 17:38:06 +00:00
sandishkumarhn	e9c31fb86d	[torch.compile] handle a custom __delattr__ method correctly (#150899 ) Fixes #150765 - handle a custom __delattr__ method correctly Test: ``` import torch class MyObject: def __init__(self, val): self.val = val # Flag to track deletion attempts instead of using print self.deletion_attempted = False def __delattr__(self, attr): if attr == "val": # Set flag instead of printing self.deletion_attempted = True else: super().__delattr__(attr) @torch.compile(fullgraph=True, backend="eager") def test(input_tensor): instance_a = MyObject(1) instance_b = MyObject(2) del instance_a.val del instance_b.val exists_a = hasattr(instance_a, 'val') exists_b = hasattr(instance_b, 'val') deletion_attempted_a = instance_a.deletion_attempted deletion_attempted_b = instance_b.deletion_attempted return input_tensor + 1, exists_a, exists_b, deletion_attempted_a, deletion_attempted_b # Run the test result = test(torch.ones(1)) print(f"Result tensor: {result[0]}") print(f"val attribute still exists on instance_a: {result[1]}") print(f"val attribute still exists on instance_b: {result[2]}") print(f"Deletion was attempted on instance_a: {result[3]}") print(f"Deletion was attempted on instance_b: {result[4]}") ``` output: ``` (base) sany@sandishs-Laptop pytorch % python3 test_delattr_fix.py Result tensor: tensor([2.]) val attribute still exists on instance_a: True val attribute still exists on instance_b: True Deletion was attempted on instance_a: True Deletion was attempted on instance_b: True ``` ``` (pytorch-dev) sany@sandishs-Laptop pytorch % python3 -m pytest test/dynamo/test_repros.py::ReproTests::test_delattr_return -v ========================================================= test session starts ========================================================= platform darwin -- Python 3.12.5, pytest-8.3.5, pluggy-1.5.0 -- /Library/Frameworks/Python.framework/Versions/3.12/bin/python3 cachedir: .pytest_cache rootdir: /Users/sany/git/pytorch configfile: pytest.ini plugins: typeguard-4.3.0 collected 1 item Running 1 items in this shard test/dynamo/test_repros.py::ReproTests::test_delattr_return PASSED [0.0659s] [100%] ========================================================== 1 passed in 1.71s ========================================================== (pytorch-dev) sany@sandishs-Laptop pytorch % ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150899 Approved by: https://github.com/jansel, https://github.com/StrongerXi	2025-06-04 17:27:20 +00:00
PyTorch MergeBot	4405dc1487	Revert "Always set CPU affinity for benchmark jobs (#154569 )" This reverts commit 629fca295e1257c2c54d1b6316ed4fa00e6044d6. Reverted https://github.com/pytorch/pytorch/pull/154569 on behalf of https://github.com/anijain2305 due to potentially causing compile time regressions, unsure ([comment](https://github.com/pytorch/pytorch/pull/154569#issuecomment-2940737778))	2025-06-04 16:52:15 +00:00
dependabot[bot]	8f08f90b61	Bump pillow from 10.0.1 to 10.3.0 in /.github/requirements (#154416 ) Bumps [pillow](https://github.com/python-pillow/Pillow) from 10.0.1 to 10.3.0. - [Release notes](https://github.com/python-pillow/Pillow/releases) - [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst) - [Commits](https://github.com/python-pillow/Pillow/compare/10.0.1...10.3.0) --- updated-dependencies: - dependency-name: pillow dependency-version: 10.3.0 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2025-06-04 09:37:13 -07:00
Nikita Shulga	aed938f3a8	Enable check_gomp for Ubuntu OSes (#155119 ) And ARM platform Pull Request resolved: https://github.com/pytorch/pytorch/pull/155119 Approved by: https://github.com/atalman	2025-06-04 15:57:08 +00:00
PyTorch MergeBot	20912673a6	Revert "Add __main__ guards to jit tests (#154725 )" This reverts commit 1a55fb0ee87eaa8b376aaa82d95d213fe0fbe64b. Reverted https://github.com/pytorch/pytorch/pull/154725 on behalf of https://github.com/malfet due to This added 2nd copy of raise_on_run to common_utils.py which caused lint failures, see https://github.com/pytorch/pytorch/actions/runs/15445374980/job/43473457466 ([comment](https://github.com/pytorch/pytorch/pull/154725#issuecomment-2940503905))	2025-06-04 15:42:52 +00:00
PyTorch MergeBot	6f93ce3c86	Revert "[Cutlass] fp8 dynamic shapes test (#154829 )" This reverts commit 36596ad2a009a0906848fa264954d4b200efc50e. Reverted https://github.com/pytorch/pytorch/pull/154829 on behalf of https://github.com/seemethere due to This is failing internal tests see, [fburl.com/diff/3gomp7i3](https://fburl.com/diff/3gomp7i3). Please re-land this as a co-dev diff ([comment](https://github.com/pytorch/pytorch/pull/154829#issuecomment-2940494361))	2025-06-04 15:36:27 +00:00
PyTorch MergeBot	3fa3dbdb1f	Revert "[Cutlass] EVT dynamic shapes support (#154835 )" This reverts commit 4224a7df01a9607830da771fd4884c8eba150630. Reverted https://github.com/pytorch/pytorch/pull/154835 on behalf of https://github.com/seemethere due to This is part of a stack that is failing internal tests see, [fburl.com/diff/3gomp7i3](https://fburl.com/diff/3gomp7i3). Please re-land this as a co-dev diff ([comment](https://github.com/pytorch/pytorch/pull/154835#issuecomment-2940463211))	2025-06-04 15:33:09 +00:00
Jeff Daily	3ce5102927	[ROCm] fix CI failures from inductor periodic (#154896 ) Similar idea as https://github.com/pytorch/pytorch/pull/154497, but for ROCm. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154896 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-06-04 15:28:43 +00:00
PyTorch MergeBot	a99a01a677	Revert "[dynamo] Mark a vt unspecialized nn module variable source earlier (#154780 )" This reverts commit cc96febb979da16b0a0b758020b330a49c72b7e7. Reverted https://github.com/pytorch/pytorch/pull/154780 on behalf of https://github.com/seemethere due to This fails internal testing see, https://fburl.com/diff/b0yuxk4w ([comment](https://github.com/pytorch/pytorch/pull/154780#issuecomment-2940381691))	2025-06-04 15:03:34 +00:00
PyTorch MergeBot	a0f2544502	Revert "[dynamo][dynamic] Recompilation hint for nn module integer attributes (#154867 )" This reverts commit 6c2f941e250ba34a920f476c8a9ee30e6153fc15. Reverted https://github.com/pytorch/pytorch/pull/154867 on behalf of https://github.com/seemethere due to This fails internal testing see, https://fburl.com/diff/b0yuxk4w ([comment](https://github.com/pytorch/pytorch/pull/154780#issuecomment-2940381691))	2025-06-04 15:03:34 +00:00
Anthony Barbier	1a55fb0ee8	Add __main__ guards to jit tests (#154725 ) This PR is part of a series attempting to re-submit https://github.com/pytorch/pytorch/pull/134592 as smaller PRs. In jit tests: - Add and use a common raise_on_run_directly method for when a user runs a test file directly which should not be run this way. Print the file which the user should have run. - Raise a RuntimeError on tests which have been disabled (not run) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154725 Approved by: https://github.com/Skylion007	2025-06-04 14:44:08 +00:00
Anthony Barbier	3f34d26040	Add __main__ guards to distributed tests (#154628 ) This is the first PR of a series in an attempt to re-submit #134592 as smaller PRs. In distributed tests: - Ensure all files which should call run_tests do call run_tests. - Raise a RuntimeError on tests which have been disabled (not run) - Remove any remaining uses of "unittest.main()"" Cc @wconstab @clee2000 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154628 Approved by: https://github.com/Skylion007	2025-06-04 14:39:57 +00:00
Anthony Barbier	c8d44a2296	Add __main__ guards to fx tests (#154715 ) This PR is part of a series attempting to re-submit #134592 as smaller PRs. In fx tests: - Add and use a common raise_on_run_directly method for when a user runs a test file directly which should not be run this way. Print the file which the user should have run. - Raise a RuntimeError on tests which have been disabled (not run) - Remove any remaining uses of "unittest.main()"" Pull Request resolved: https://github.com/pytorch/pytorch/pull/154715 Approved by: https://github.com/Skylion007	2025-06-04 14:38:50 +00:00
Anthony Barbier	cf9cad31df	Add __main__ guards to tests (#154716 ) This PR is part of a series attempting to re-submit https://github.com/pytorch/pytorch/pull/134592 as smaller PRs. Add missing `if __name__ == "__main__":` guards to some tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154716 Approved by: https://github.com/Skylion007	2025-06-04 14:38:13 +00:00
Kshitij Khode	ca0c2985d3	[ONNX] Allow exporter to export SDPA to Attention onnx operator (#154596 ) Fixes [#149662](https://github.com/pytorch/pytorch/issues/149662) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154596 Approved by: https://github.com/justinchuby, https://github.com/titaiwangms Co-authored-by: Justin Chu <justinchuby@users.noreply.github.com>	2025-06-04 14:29:44 +00:00
zeshengzong	31d12b3955	Fix avg_pool2d param kernel_size descripthon (#154353 ) Fixes part of #153149 ## Test Result ![image](https://github.com/user-attachments/assets/216ffd2b-dd2b-4cf6-9fca-aeed075be5e7) ![image](https://github.com/user-attachments/assets/820cd184-1f8e-4a7a-b64e-15dfb9c7dad2) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154353 Approved by: https://github.com/colesbury	2025-06-04 11:55:01 +00:00
soulitzer	2af78d368f	Skip another test file that doesn't run gradcheck for slow gradcheck (#154852 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154852 Approved by: https://github.com/albanD	2025-06-04 07:47:09 +00:00
fengqing.lu@intel.com	0f10df71a6	[Intel GPU] Make SDPA output has the same stride as Query. (#154340 ) Fixes [#153903](https://github.com/pytorch/pytorch/issues/153903). Currently the output tensor of SDPA XPU is always defined as contiguous stride, while CPU/CUDA flash_attention and cudnn_attention allocate output tensor with stride the same as Query. This PR aligns XPU's behavior with CUDA/CPU to make XPU compatible to CPU/CUDA's modeling code. The function `alloc_with_matching_layout` is copied from cudnn `8c16d0e404/aten/src/ATen/native/cudnn/MHA.cpp (L874)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154340 Approved by: https://github.com/Skylion007, https://github.com/EikanWang, https://github.com/guangyey	2025-06-04 07:16:56 +00:00
Yiming Zhou	1e20745532	[ez][AOTI] Fix index offset for Optional Tensor Return (#155073 ) Summary: As title. See added test for more context. Test Plan: buck2 run mode/dev-nosan caffe2/test/inductor:test_aot_inductor_custom_ops -- -r test_fn_with_optional_tensor_output_2 Rollback Plan: Differential Revision: D75900658 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155073 Approved by: https://github.com/angelayi	2025-06-04 06:22:46 +00:00
angelayi	d2bfd97d71	[export] Refactor pt2 save/load (#152495 ) Refactor the pt2 archive saving to consolidate the format of torch.export.save and torch._inductor.package.package_aoti. This PR adds the following functions, which torch.export.save and AOTI packaging calls into: ```python package_pt2( f: FileLike, *, exported_programs: Optional[Union[ExportedProgram, dict[str, ExportedProgram]]] = None, aoti_files: Optional[Union[list[str], dict[str, list[str]]]] = None, extra_files: Optional[dict[str, Any]] = None, ) -> FileLike @dataclass class PT2ArchiveContents: exported_programs: dict[str, ExportedProgram] aoti_runners: dict[str, AOTICompiledModel] extra_files: dict[str, Any] load_pt2(f: FileLike) -> PT2ArchiveContents ``` Power users directly call into these APIs if they want to bundle multiple exported programs, aoti files, or extra metadata. This is how the pt2 archive looks like ([spec](https://docs.google.com/document/d/1RQ4cmywilnFUT1VE-4oTGxwXdc8vowCSZsrRgo3wFA8/edit?tab=t.0)): ``` ├── archive_format ├── version ├── .data ├── data │ ├── aotinductor │ │ └── model1 │ │ ├── model1.cpp │ │ ├── model1.so # currently AOTI automatically moves weights in here, TODO to move it out │ │ ├── cg7domx3woam3nnliwud7yvtcencqctxkvvcafuriladwxw4nfiv.cubin │ │ └── cubaaxppb6xmuqdm4bej55h2pftbce3bjyyvljxbtdfuolmv45ex.cubin │ ├── weights │ │ ├── model1.pt # TODO to dedup weights between model1/model2 │ │ └── model2.pt │ └── constants │ │ ├── model1.pt # TODO to dedup weights between model1/model2 │ │ └── model2.pt │ └── sample_inputs │ ├── model1.pt # TODO to dedup weights between model1/model2 │ └── model2.pt ├── extra │ └── user_metadata.txt └── models ├── model1.json └── model2.json ``` Future todos: - unbundle the weights -- instead of .pt, we can use bin files, which will also allow us to dedup weights if we store multiple models - update aoti_compile_and_package to also save the exported program - integrate TNR with this packaging flow Pull Request resolved: https://github.com/pytorch/pytorch/pull/152495 Approved by: https://github.com/yushangdi	2025-06-04 06:04:29 +00:00
Yuanhao Ji	75b24c273b	Export `torch::utils::tensor_to_numpy` (#154178 ) Fixes #154105 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154178 Approved by: https://github.com/albanD, https://github.com/Skylion007, https://github.com/youkaichao	2025-06-04 05:48:27 +00:00
fengqing.lu@intel.com	7b074346e0	[Intel GPU] Support f32 intermediate dtype, headdim size <=576 and f32 causal mask for SDPA (#152091 ) In OneDNN v3.7, SDPA has below defects: 1. The dtype of intermediate value is the same as QKV, while Pytorch uses FP32 dtype for intermediate value to make sure better accuracy. 2. Only support headdim size <= 256. 3. Don't support implict causal mask when QKV is FP32. We need to build an attention mask explicitly with aten ops. In OneDNN v3.8, they have update for these defects. Since these are tiny changes, I decided to put them in single PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152091 Approved by: https://github.com/EikanWang, https://github.com/guangyey, https://github.com/drisspg	2025-06-04 05:18:36 +00:00
fduwjj	4d93985d13	[c10d] Separate monitoring thread into a class in PGNCCL (#153977 ) This is the start of a series of efforts to consolidating auxiliary threads in PGNCCL, aka watchdog and heartbeat_monitoring threads. Right now we launch these two threads per PG instances, i.e., if users create hundred or thousand instances of PG or subPGs, we will end up with that twice many side threads which is not efficient. We have a RFC to consolidate them (https://github.com/pytorch/pytorch/issues/146956). Right now both threads are assigned with so many functionalities so it is hard to do the consolidations in one shot, we will try to split it into at least two steps (PRs) to make it easier to test and review. We did our first attemp in https://github.com/pytorch/pytorch/pull/153668 but we also want to try to see if we can make monitoring thread a class. This PR is doing the first step to make monitoring thread a class. The next step to also extract watchdog to be a separate class so that we know its dependency. What we did in this PR: 1. Move all related variables and methods into a class named `HeartbeatMonitor`. 2. Correct some errors in the original logics inside monitoring thread loop. 3. Move the error propagation check to watchdog thread which is more relevant. This is totally fine since we rolled out EventCache out fully so watchdog hang is rare now. Today there are two major functions inside heartbeat monitoring thread today: 1. Check the heartbeat of watchdog thread every 8 minutes. If no heartbeat detected and we are sure monitoring thread has not been stopped, we will kill the program by SIG_ABORT. 2. We check TCPStore every 30 sec to see if any watchdog timeout happens on other ranks, if so we will initiate a dump signal on the current rank as well. (We do this only in the default PG) Differential Revision: [D75799278](https://our.internmc.facebook.com/intern/diff/D75799278) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153977 Approved by: https://github.com/kwen2501, https://github.com/d4l3k	2025-06-04 04:07:07 +00:00
tvukovic-amd	ec35a36820	[ROCm][Windows] Fix building tests for multiple architectures (#154979 ) Fixing building C10_CUDA_ALL_TEST_FILES and Caffe2_HIP_TEST_SRCS for multiple architectures Pull Request resolved: https://github.com/pytorch/pytorch/pull/154979 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-06-04 03:53:21 +00:00
bobrenjc93	72fe1d5f42	Add randint_like tensor overload for high (#154899 ) Fixes #135664 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154899 Approved by: https://github.com/StrongerXi ghstack dependencies: #154863	2025-06-04 03:37:09 +00:00
Nikita Shulga	6b0c6f2856	[BE] Delete pre-CUDA-10.1 code from SparseCUDABlas (#155079 ) As latest PyTorch is no longer buildable against it CUDA-10, so this is essentially a dead code Made small change to hipify script to rename `cusparseGetErrorString` to `hipsparseGetErrorString` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155079 Approved by: https://github.com/atalman, https://github.com/cyyever	2025-06-04 03:29:24 +00:00
Nikita Shulga	9f39028629	[MPS][BE] Move sigmoid op to Metal (#155080 ) Fixes https://github.com/pytorch/pytorch/issues/154895 Pull Request resolved: https://github.com/pytorch/pytorch/pull/155080 Approved by: https://github.com/dcci, https://github.com/cyyever ghstack dependencies: #154936, #155002, #155081	2025-06-04 03:28:11 +00:00
Blaine Burton Rister	437df54cc8	[Inductor] Fix a few FX conversion bugs. (#154958 ) # Feature This PR fixes two bugs with Inductor's FX backend. 1. When extracting offsets from `ReinterpretView`'s, we accidentally took the offset of the parent layout instead of the view's layout. This case is triggered when multiple kernels write into the same buffer due to `torch.cat`. 2. In certain rare cases, `V.graph.graph_inputs` can contain a constant input value. In case this happens, create a new `sympy.Symbol` for the input, for compatibility with the existing `SymbolBuffer` abstraction mapping to an FX placeholder. This case is triggered when calling `torch._inductor.compile` on certain modules coming from `torch.export`. # Test plan Added a couple of tests exposing these bugs. 1. Concat with multiple kernels writing to the same buffer. 3. `Export` -> `torch._inductor.compile` with a constant input. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154958 Approved by: https://github.com/jansel	2025-06-04 03:09:44 +00:00
Justin Chu	3e57de1251	[ONNX] Create support for rotary embeddings (#154745 ) This PR registers the RotaryEmbedding op in the `torch.ops.onnx` name spaces and allows the exporter to recognize and export onnx operators. ## Design ONNX operators of their respective opset version is implemented in torch/onnx/ops/_impl.py, and are registered in the torch.ops.onnx namespace following the following rule: `OpType-version => torch.ops.onnx.OpType.opset{version}` For example, `RotaryEmbedding-23` becomes `torch.ops.onnx.RotaryEmbedding.opset23` This name is parsed by the exporter to create an onnx node in the graph without having to go through translation. When users use the ops in the model, we provide more convenient, unversioned functions under `torch.onnx.ops` that will dispatch to the implementations based on user input (type and provided attributes). For example, users can directly call `torch.onnx.ops.rotary_embedding()` to use the op natively in their pytorch models. I chose snake case naming to make the functions more pythonic and aligned with other torch apis. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154745 Approved by: https://github.com/titaiwangms	2025-06-04 03:07:43 +00:00
mori360	37e6bf8adf	Switch to _apply_to_tensors for dataclass input (#154897 ) Fixes https://github.com/pytorch/pytorch/issues/153077 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154897 Approved by: https://github.com/weifengpy	2025-06-04 02:19:52 +00:00
Natalia Gimelshein	34e3930401	fix numpy compatibility for 2d small list indices (#154806 ) Will fix #119548 and linked issues once we switch from warning to the new behavior, but for now, given how much this syntax was used in our test suite, we suspect a silent change will be disruptive. We will change the behavior after 2.8 branch is cut. Numpy behavior was changed at least in numpy 1.24 (more than 2 years ago) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154806 Approved by: https://github.com/cyyever, https://github.com/Skylion007, https://github.com/albanD	2025-06-04 01:58:52 +00:00
Feng Tian	e2760544fa	[PT] expose FlightRecord API for building (#154866 ) Summary: as titled Test Plan: CI Rollback Plan: Differential Revision: D75803611 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154866 Approved by: https://github.com/fduwjj, https://github.com/d4l3k	2025-06-04 01:25:52 +00:00
Nikita Shulga	d8e4c1c363	[BE] Define `REGISTER_UNARY_TI_DISPATCH` (#155081 ) That creates _kernel_mps function that takes iterator and calls stub for it Pull Request resolved: https://github.com/pytorch/pytorch/pull/155081 Approved by: https://github.com/dcci ghstack dependencies: #154936, #155002	2025-06-04 01:15:37 +00:00
PyTorch MergeBot	50de6ae253	Revert "[BE][Ez]: Fully type nn.utils.clip_grad (#154801 )" This reverts commit 9ce2732b685da527308dc2dc4b2eeb4e252f57d1. Reverted https://github.com/pytorch/pytorch/pull/154801 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/154801#issuecomment-2937886337))	2025-06-04 00:41:27 +00:00
eellison	40a8770154	Incorporate coalesce analysis in codegen (#153751 ) This pr uses the coalescing information in generating a tiling. The previous tiling heuristic would have each dependency generate a tiling. Then, we sum up the score for each generated tiling, preferring any 2d tiling over the default. The new tiling heuristics scores each tiling by its global coalesced memory. This gives both a potentially better tiling (especially for more complicated, 3d patterns) as well as information we can use in generating block sizes. In triton heuristics, for generating 3d tiled reductions, we take the same total block size that the 2d reduction would use, then distribute the block according to whichever block coalesces the most memory. The motivating kernel is in https://github.com/pytorch/pytorch/issues/149982 which is a 32 element reduction. A smaller version of it is [here](https://gist.github.com/eellison/0fa9396f5479eb4dba09756e3bf6ff2a). We need to run this kernel once in the forward per linear layer on a contiguous tensor, and once in the backward on a transposed tensor. While the contiguous kernel has coalesced accesses, and is performant on master, the transposed version accesses uncoalesced memory on main and is ~2.8x slower. See, this [full log](https://gist.github.com/eellison/fa644bfd9d0ae11dadb62e17a5d48a83) from the above repro. Now, with this PR, it is only ~1.15x slower. See the [updated log](https://gist.github.com/eellison/0b2b653309494d28cf7b48929a022075). Pull Request resolved: https://github.com/pytorch/pytorch/pull/153751 Approved by: https://github.com/jansel ghstack dependencies: #153723, #153730, #153748	2025-06-04 00:22:57 +00:00
Animesh Jain	6c2f941e25	[dynamo][dynamic] Recompilation hint for nn module integer attributes (#154867 ) For program like this ``` class Mod(torch.nn.Module): def __init__(self): super().__init__() self.c = 0 def forward(self, x): self.c += 1 return x * self.c ``` You can check the recompile reasons at https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/.tmpzv9z6Q/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000 ![image](https://github.com/user-attachments/assets/856a95fd-0533-4abc-a213-1f73ae2cb766) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154867 Approved by: https://github.com/zou3519 ghstack dependencies: #154780	2025-06-04 00:05:53 +00:00
xinan.lin	cbdacd32fe	[AOTI][Intel GPU] Support multi_arch_kernel_binary option for XPU. (#154514 ) Following the design of #154413, this PR add XPU support for generating kernel binary files that support multiple archs. Fixes #154682, Fixes #154683, Fixes 154689, Fixes #154685 , Fixes #154690, Fixes #154681 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154514 Approved by: https://github.com/desertfire, https://github.com/EikanWang	2025-06-03 23:02:00 +00:00
xinan.lin	8f0e3f446d	[Inductor UT] Reuse test_fused_attention.py for Intel GPU. (#154110 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154110 Approved by: https://github.com/eellison, https://github.com/jansel, https://github.com/EikanWang ghstack dependencies: #154091	2025-06-03 23:01:05 +00:00
xinan.lin	6c40e6606f	[Inductor] Add attention pattern for model DistilBert in transformers==4.44.2. (#154091 ) This PR add a attention fusion pattern that match the attention of DistilDistilBert in transformers==4.44.2 at `953196a43d/src/transformers/models/distilbert/modeling_distilbert.py (L212)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154091 Approved by: https://github.com/jansel, https://github.com/eellison	2025-06-03 23:01:05 +00:00
Michael Lazos	4224a7df01	[Cutlass] EVT dynamic shapes support (#154835 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154835 Approved by: https://github.com/henrylhtsang ghstack dependencies: #154775, #154761, #154829	2025-06-03 22:20:34 +00:00
Michael Lazos	36596ad2a0	[Cutlass] fp8 dynamic shapes test (#154829 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154829 Approved by: https://github.com/henrylhtsang, https://github.com/eellison ghstack dependencies: #154775, #154761	2025-06-03 22:20:33 +00:00
Michael Lazos	1c2b9cecd2	[Cutlass] Support bias arg for fp8 GEMM (#154761 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154761 Approved by: https://github.com/drisspg ghstack dependencies: #154775	2025-06-03 22:20:27 +00:00
Michael Lazos	5735729597	[Cutlass] Cleanup gemm_template evt handling (#154775 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154775 Approved by: https://github.com/henrylhtsang, https://github.com/eellison	2025-06-03 22:20:18 +00:00
Yiming Zhou	71499fee6b	[3/3] Add build rule and test for Graph in nativert (#154532 ) We split the large PR for adding Graph.h and Graph.cpp to nativert into 3 smaller PRs: 1. Add header file 2. Add source file 3. Add test and build rules Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 4 classes have been introduced: `Graph`, `Node`, `Value`, `Type` - `Type` represents the kind of a `Value` - `Value` represents a single symbolic value, it could be any kind that exists in `Type`. Values are inputs and outputs of a `Node`. - `Node` represents a single unit of execution, typically a PyTorch op. - `Graph` represents a model's computation graph, which is designed to facilitate transformation/analysis. Differential Revision: [D75495273](https://our.internmc.facebook.com/intern/diff/D75495273/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154532 Approved by: https://github.com/SherlockNoMad ghstack dependencies: #154530, #154531	2025-06-03 21:52:05 +00:00
Yiming Zhou	b4c399d445	[2/3] Add source file for Graph in nativert (#154531 ) We split the large PR for adding Graph.h and Graph.cpp to nativert into 3 smaller PRs: 1. Add header file 2. Add source file 3. Add test and build rules. Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 4 classes have been introduced: `Graph`, `Node`, `Value`, `Type` - `Type` represents the kind of a `Value` - `Value` represents a single symbolic value, it could be any kind that exists in `Type`. Values are inputs and outputs of a `Node`. - `Node` represents a single unit of execution, typically a PyTorch op. - `Graph` represents a model's computation graph, which is designed to facilitate transformation/analysis. Differential Revision: [D75492405](https://our.internmc.facebook.com/intern/diff/D75492405/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154531 Approved by: https://github.com/SherlockNoMad ghstack dependencies: #154530	2025-06-03 21:51:52 +00:00
Yiming Zhou	55873dcb0d	[1/3] Add header file for Graph in nativert (#154530 ) We split the large PR for adding Graph.h and Graph.cpp to `nativert` into 3 smaller PRs: 1. Add header file 2. Add source file 3. Add test and build rules. Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 4 classes have been introduced: `Graph`, `Node`, `Value`, `Type` - `Type` represents the kind of a `Value` - `Value` represents a single symbolic value, it could be any kind that exists in `Type`. Values are inputs and outputs of a `Node`. - `Node` represents a single unit of execution, typically a PyTorch op. - `Graph` represents a model's computation graph, which is designed to facilitate transformation/analysis. Differential Revision: [D75491860](https://our.internmc.facebook.com/intern/diff/D75491860/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154530 Approved by: https://github.com/SherlockNoMad	2025-06-03 21:51:47 +00:00
LifengWang	69a57d9486	add JSON output support for operator benchmark (#154410 ) To better support the integration of operator benchmark performance data into the OSS benchmark database for the dashboard, I’ve added a JSON output format that meets the required specifications: https://github.com/pytorch/pytorch/wiki/How-to-integrate-with-PyTorch-OSS-benchmark-database#output-format Since the current operator benchmark already has a flag `--output-json` to support saving the results into a JSON file, I add a new flag `--output-json-for-dashboard` for this feature. At the same time, I renamed the `--output-dir` to `--output-csv` for a clearer and more intuitive expression. An example of the JSON output of the operator benchmark. ``` [ { "benchmark": { "name": "PyTorch operator benchmark - add_M1_N1_K1_cpu", "mode": "inference", "dtype": "float32", "extra_info": { "input_config": "M: 1, N: 1, K: 1, device: cpu" } }, "model": { "name": "add_M1_N1_K1_cpu", "type": "micro-benchmark", "origins": [ "pytorch" ] }, "metric": { "name": "latency", "unit": "us", "benchmark_values": [ 2.074 ], "target_value": null } }, { "benchmark": { "name": "PyTorch operator benchmark - add_M64_N64_K64_cpu", "mode": "inference", "dtype": "float32", "extra_info": { "input_config": "M: 64, N: 64, K: 64, device: cpu" } }, "model": { "name": "add_M64_N64_K64_cpu", "type": "micro-benchmark", "origins": [ "pytorch" ] }, "metric": { "name": "latency", "unit": "us", "benchmark_values": [ 9.973 ], "target_value": null } }, ] ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154410 Approved by: https://github.com/huydhn	2025-06-03 21:29:24 +00:00
Scott Wolchok	8e1474d3c6	[inductor] small cleanups in torch/_inductor/codegen/mps.py (#154921 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154921 Approved by: https://github.com/jansel, https://github.com/Skylion007	2025-06-03 20:57:25 +00:00
cyy	debd095149	Avoid index integer overflow in gemm_notrans_ (#154809 ) Use uint64_t index types to avoid ``` torch_np/numpy_tests/core/test_einsum.py::TestEinsum::test_einsum_broadcast /var/lib/jenkins/workspace/aten/src/ATen/native/cpu/BlasKernel.cpp:132:24: runtime error: signed integer overflow: 9223365439786057728 + 13194139533312 cannot be represented in type 'long' #0 0x7f30d26166ba in std::enable_if<std::is_same_v<long, long>, void>::type at::native::cpublas::(anonymous namespace)::gemm_notrans_<long, long, long>(long, long, long, long, long const, long, long const, long, long, long, long) /var/lib/jenkins/workspace/aten/src/ATen/native/cpu/BlasKernel.cpp:132:24 #1 0x7f30d26166ba in void at::native::cpublas::(anonymous namespace)::gemm_core_<long, long, long>(at::native::TransposeType, at::native::TransposeType, long, long, long, long, long const, long, long const, long, long, long, long) /var/lib/jenkins/workspace/aten/src/ATen/native/cpu/BlasKernel.cpp:451:12 #2 0x7f30d25fba1b in at::native::cpublas::(anonymous namespace)::cpublas_gemm_impl(c10::ScalarType, at::native::TransposeType, at::native::TransposeType, long, long, long, c10::Scalar const&, void const, long, void const, long, c10::Scalar const&, void, long)::$_2::operator()() const::'lambda2'()::operator()() const /var/lib/jenkins/workspace/aten/src/ATen/native/cpu/BlasKernel.cpp:485:3 #3 0x7f30d25fba1b in at::native::cpublas::(anonymous namespace)::cpublas_gemm_impl(c10::ScalarType, at::native::TransposeType, at::native::TransposeType, long, long, long, c10::Scalar const&, void const, long, void const, long, c10::Scalar const&, void, long)::$_2::operator()() const /var/lib/jenkins/workspace/aten/src/ATen/native/cpu/BlasKernel.cpp:485:3 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154809 Approved by: https://github.com/soulitzer	2025-06-03 19:28:34 +00:00
karthickai	10c3e6ec43	[inductor][dynamo] Include operator name in size/stride/alignment assertion (#152353 ) Fixes #151930 This PR updates the `assert_size_stride` and `assert_alignment` functions in [guards.cpp](https://github.com/pytorch/pytorch/blob/main/torch/csrc/dynamo/guards.cpp) to accept an optional `op_name` argument and includes it in the error messages. The corresponding type stubs in [guards.pyi](https://github.com/pytorch/pytorch/blob/main/torch/_C/_dynamo/guards.pyi) are updated to match the new function arg. In [inductor/ir.py](https://github.com/pytorch/pytorch/blob/main/torch/_inductor/ir.py) extracts the operator name from the FX graph and passes it into the `codegen_size_asserts` and `codegen_alignment_asserts` functions, so that generated assertions in Triton code include the op name for better debugging. Added unit tests inside [test_torchinductor.py](https://github.com/pytorch/pytorch/blob/main/test/inductor/test_torchinductor.py). - Verified both successful and failing assertion cases include the operator name. - Verified that generated Triton code contains the op name inside the asserts. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152353 Approved by: https://github.com/jansel, https://github.com/shunting314	2025-06-03 19:21:15 +00:00
Animesh Jain	cc96febb97	[dynamo] Mark a vt unspecialized nn module variable source earlier (#154780 ) I am working on providing some skip guard helper functions to allow users to reduce guard overhead. This is a refactor to allow that. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154780 Approved by: https://github.com/StrongerXi, https://github.com/jansel	2025-06-03 19:19:47 +00:00
David Berard	ea7b233015	[flex attention][triton pin] triton_helpers shim for TMA apis (#154858 ) Triton 3.4 will remove the experimental TMA apis: https://github.com/triton-lang/triton/pull/6488 To allow compatibility across different triton versions, we implement a shim layer which calls the new API if available, and otherwise falls back to the experimental API. Test: `python test/inductor/test_flex_attention.py TestFlexAttentionCUDA.test_GQA_causal_mask_cuda` which previously fails w/ triton-lang/tritoncda4229558c5dca7f7c4734bedd3e596ebcae0b8, but now passes. Note: we'll need to apply this for other things in inductor, this just does it for flex attention. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154858 Approved by: https://github.com/NikhilAPatel, https://github.com/drisspg	2025-06-03 19:15:48 +00:00
atalman	85fb13d0d1	[BE] Cleanup cuda 12.4 artifacts from scripts and workflows (#154893 ) Remove artifacts. CUDA 12.4 was deprecated. hence no need to keep this code around Pull Request resolved: https://github.com/pytorch/pytorch/pull/154893 Approved by: https://github.com/nWEIdia, https://github.com/malfet, https://github.com/tinglvv	2025-06-03 18:43:40 +00:00
David Berard	c014e9d7cd	[inductor][test] test_padding.py: use inductor TestCase instead of dynamo TestCase (#154935 ) test_pad_3d_tensor fails if you run it multiple times in a row, because the cache is populated and inductor skips the logic that increments the counter. To fix this, switch these tests to use inductor's TestCase / run_tests instead of dynamo's - this way, a fresh inductor cache is used. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154935 Approved by: https://github.com/Skylion007	2025-06-03 18:36:44 +00:00
Jane Xu	e8183f8d3d	add #pragma once to stable/library.h (#154920 ) This shoulda been there and it was an oversight that it was not! We do not want the same translation unit to process this header multiple times. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154920 Approved by: https://github.com/albanD, https://github.com/Skylion007	2025-06-03 18:34:53 +00:00
Ryan Guo	6f7694f18f	[dynamo] Reconstruct defaultdict properly (#154931 ) `DefaultDictVariable` inherited `ConstDictVariable.reconstruct`, causing dynamo to reconstruct a `DefaultDictVariable` into a dict rather than defaultdict. This patch fixes that. Fixes #138412. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154931 Approved by: https://github.com/williamwen42, https://github.com/zou3519 ghstack dependencies: #154930	2025-06-03 18:18:40 +00:00
Ryan Guo	467235027c	[AOTDispatch] Use the proper meta function for `_amp_foreach_non_finite_check_and_unscale_` (#154930 ) As title, this fixes part of #138412. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154930 Approved by: https://github.com/zou3519	2025-06-03 18:18:40 +00:00
Svetlana Karslioglu	462579af11	Update merge_rules.yaml (#155008 ) - add new docs reviewers Pull Request resolved: https://github.com/pytorch/pytorch/pull/155008 Approved by: https://github.com/malfet	2025-06-03 18:09:23 +00:00
Nikita Shulga	f714599c57	[MPS][BE] Extend torch.special. to integer dtypes (#155002 ) By changing the functor to looks as follows ```metal struct xlog1py_functor { template <typename T, enable_if_t<is_floating_point_v<T>, bool> = true> inline T operator()(const T a, const T b) { return static_cast<T>(c10:🤘:xlog1py(a, b)); } template <typename T, enable_if_t<is_integral_v<T>, bool> = true> inline float operator()(const T a, const T b) { return c10:🤘:xlog1py(float(a), float(b)); } }; ``` Repeat the same for `zeta`, `chebyshev_polynomial_[tuvw]_functor` and `hermite_polynomial_h[e]_functor` Pull Request resolved: https://github.com/pytorch/pytorch/pull/155002 Approved by: https://github.com/Skylion007, https://github.com/dcci ghstack dependencies: #154936	2025-06-03 17:52:41 +00:00
bubuss	31405a69fb	[typing] Add missing type annotations to torch.nn.init module (#154504 ) ## Summary Adds missing type annotations to `torch.nn.init` and removes `# mypy: allow-untyped-defs` since all functions are now properly typed. ## Changes - Added missing type annotations to initialization functions in the module. - Added missing typing imports: `Any`, `Callable`, `Union` - Removed `# mypy: allow-untyped-defs` comment - Create Literal types for kaiming initialization mode and nonlinearity. - Created `__all__` ## Why Better IDE support, catches type errors earlier, and brings the module up to PyTorch's typing standards. No runtime changes - purely additive typing improvements. Tested with existing test suite and lintrunner. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154504 Approved by: https://github.com/Skylion007	2025-06-03 17:33:32 +00:00
cora-codes	40142978d7	Add type annotation to orthogonal_ (#154927 ) Trivial charge, but I want pyright to stop yelling at me Pull Request resolved: https://github.com/pytorch/pytorch/pull/154927 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-06-03 17:00:02 +00:00
albanD	1f131fe56b	Update bug-report.yml (#154857 ) Update issue template for binary data and numerical notes. Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/154857 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-06-03 16:13:07 +00:00
fduwjj	ff92b42fc3	[c10d][gloo] Integrate vendor generic FR into gloo (#152614 ) This is a first quick prototyping for FR integration for gloo. Few features gaps: - Input/Output numels for each collective - Whether to use c10::Event or where to use it. - Where to dump the FR traces. (The dump api is provided in this PR) Differential Revision: [D75803601](https://our.internmc.facebook.com/intern/diff/D75803601) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152614 Approved by: https://github.com/d4l3k ghstack dependencies: #154929	2025-06-03 16:12:54 +00:00
Howard Huang	283f876ab6	[PP] Fix disabled flaky tests (#154856 ) Fix https://github.com/pytorch/pytorch/issues/154373, https://github.com/pytorch/pytorch/issues/154391, https://github.com/pytorch/pytorch/issues/154408, https://github.com/pytorch/pytorch/issues/154443, https://github.com/pytorch/pytorch/issues/154481 Because MultiProcContinousTest [now executes the tests with 8 GPUs instead of 2](https://github.com/pytorch/pytorch/pull/153653), our PP tests comparing gradients have become flakier due to the longer pipeline. The gradients are still close but we need to relax the tolerance. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154856 Approved by: https://github.com/Skylion007	2025-06-03 15:55:29 +00:00
Alanna Burke	250e9af4da	Removing per torch.compile audit. (#154572 ) Removing https://pytorch.org/docs/stable/torch.compiler_best_practices_for_backends.html per torch.compile audit Pull Request resolved: https://github.com/pytorch/pytorch/pull/154572 Approved by: https://github.com/williamwen42, https://github.com/svekars	2025-06-03 15:41:52 +00:00
Ke Wen	3685b10170	Turn on compile with NVSHMEM (#154538 ) Before: `USE_NVSHMEM=1` need to be explicit set in build environment. After: `USE_NVSHMEM=1` is the default for CUDA/Rocm on Linux. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154538 Approved by: https://github.com/ngimel	2025-06-03 15:24:24 +00:00
Ruisi Zhang	a1a268aff5	[dtensor] fix simplefsdp mixed-precision training bugs (#154975 ) This is a follow-up on the previous dtensor redistribute PR: https://github.com/pytorch/pytorch/pull/150740, which enables SimpleFSDP's mixed-precision training. In the most recent integration in TorchTitan: https://github.com/pytorch/torchtitan/pull/1250, we found some discrepancies between SimpleFSDP's `fully_shard` and `replicate` modes when MPT is enabled. After debugging, I found the problem is in dtensor redistribute --`local_tensor` is taken out again from the original `input`. Thus, the dtensor used for communication has its original precision instead of using `forward_dtype`. This PR fixes this issue and corrects previously added test cases. After fixing the bug, the loss curves of `fully_shard` and `replicate` mode match perfectly. ![loss](https://github.com/user-attachments/assets/a8faddae-a476-48c0-a411-3fe04d2233bd) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154975 Approved by: https://github.com/tianyu-l	2025-06-03 14:47:36 +00:00
eellison	2608927cfb	Solve for tilings (#153748 ) Find variables that coalesce the reads and writes and score the total size. If uncoalesced memory expressions are found, look for additional tiling of variables which will coalesce memory accesses. For instance - for the following expression: `(32*p0) // 2048`, tiling p0 by 64 will make this expression coalesced. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153748 Approved by: https://github.com/jansel ghstack dependencies: #153723, #153730	2025-06-03 14:37:30 +00:00
David Svantesson	812deecaab	Add option to define OpenBLAS version for manylinux Dockerfile_2_28_aarch64 (#150106 ) Adds optional variable OPENBLAS_VERSION to `.ci/docker/common/install_openblas.sh` used to define which version of OpenBLAS to install. Adds argument to `Dockerfile_2_28_aarch64` image. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150106 Approved by: https://github.com/aditew01, https://github.com/fadara01, https://github.com/malfet Co-authored-by: Fadi Arafeh <115173828+fadara01@users.noreply.github.com>	2025-06-03 14:35:54 +00:00
eellison	0adbde4d35	Analyze coalesced mem (#153730 ) Analyze memory expressions to see if they contain a coalescing symbol. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153730 Approved by: https://github.com/jansel ghstack dependencies: #153723	2025-06-03 14:29:06 +00:00
Nikita Shulga	e9266f807a	[BE] Use vendored packaging for testing (#154946 ) As the rest of the torch uses it, test should rely on it as well Pull Request resolved: https://github.com/pytorch/pytorch/pull/154946 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-06-03 14:22:53 +00:00
Nikita Shulga	9cdce682a1	[MPS][BE] Reimplement log1p as Metal shader (#154936 ) That should make it faster than MPSGraph implementation, but also improves accuracy for small inputs, by using the algorithm described in [What Every Computer Scientist Should Know About Floating-Point Arithmetic](https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html#1202), i.e. $log(1+x) = \frac{x * log(1+x)}{(1 + x) - 1}$ if $1 +x \neq 1$ else just $x$ Also tried using first 3 elements of Taylor series in Horner's form which also seems to work fine, i.e. $log(1+x) \approx x * (1 -x (\frac{1}{2} - \frac{x}{3}))$ Replaced less accurate log1p implementation in `c10/metal/special_math.h` with generic one. Parametrize and modify regression test to check for accuracy of small values TODOs: - Do proper implementation for complex values as well, perhaps using `0408ba0a76/mlx/backend/metal/kernels/utils.h (L339)` - May be implement it using Remez-like algorithm documented here `207f3b2b25/lib/msun/src/s_log1pf.c (L37)` - Or use llvm's implementation from `f393986b53/libclc/clc/lib/generic/math/clc_log1p.inc (L22)` - Benchmark which algorithm is faster and delivers better accuracy Pull Request resolved: https://github.com/pytorch/pytorch/pull/154936 Approved by: https://github.com/dcci, https://github.com/Skylion007	2025-06-03 14:10:13 +00:00
eellison	00dfd3891e	[Tiling rewrite pt1] Normalize reads and writes to common iter space (#153723 ) In order to take the globally best tiling, we need to normalize all the node read and writes to a common iteration space. This first pr finds a common split among nodes in a fused scheduler node, and then normalizes reads and writes to the common split. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153723 Approved by: https://github.com/jansel	2025-06-03 14:04:34 +00:00
Animesh Jain	635b73e697	[dynamo][guards] Flush cache to more accurately measure guard overhead (#154764 ) We observed that guard overhead at runtime using profiler traces was higher than reported in this profiling function at the compile time. After investigation, we found that f_locals are already in cache and that was causing the guard overhead to be way smaller while profiling during the compilation. To be more realistic, we flush the cache here. Profiling the guard overhead during compilation (in addition to at runtime) allows faster iteration time, and logging in tlparse and internal databases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154764 Approved by: https://github.com/zou3519, https://github.com/jansel, https://github.com/StrongerXi	2025-06-03 11:50:57 +00:00
Aidyn-A	71a0af8a14	[TEST][Quantization] Skip test_learnable due to hypothesis (#152819 ) As per comment in https://github.com/pytorch/pytorch/issues/111471#issuecomment-1866933243 the tests are failing due to hypothesis. This PR adds a skip to those tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152819 Approved by: https://github.com/eqy	2025-06-03 11:23:15 +00:00
bobrenjc93	ea5b9eca74	Combine sticky pgo key with job id (#154863 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154863 Approved by: https://github.com/Mingming-Ding	2025-06-03 07:58:38 +00:00
Boyuan Feng	a4da1d4a47	[Graph Partition] support standalone_compile (#154698 ) For graph partition, `write_get_raw_stream_header_once` is done once so the autotune code may not have the header. This PR additionally calls `write_get_raw_stream_header` in `codegen_device_guard_enter` before `get_raw_stream` is used. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154698 Approved by: https://github.com/oulgen	2025-06-03 07:40:42 +00:00
fduwjj	d91c85babb	[c10d][fr] Split cuda and non-cuda fr logic into two cpp file (#154929 ) During the integration fr with gloo I found that put all logic inside one cpp with both build Macro does not work in the current linkage set up in the bazil file. If we put the cpp in the libtorch_cpu, then cuda side build will fail, if we put both we get complaint about ld.lld: error: duplicate symbol: typeinfo for c10d::DebugInfoWriter. To fix this, we need to move the common logic into another header file and we use different cpp file for cpu and cuda so that fr can be used in both cases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154929 Approved by: https://github.com/kwen2501	2025-06-03 07:00:14 +00:00
Bin Bao	13044b2b04	Move c10/macros/Export.h to torch/standalone (#154850 ) Summary: The goal of this PR and future follow-up PRs is to group a set of header files required by AOTInductor Standalone in a separate directory, ensuring they are implemented in a header-only manner. Test Plan: CI Bifferential Revision: D75756619 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154850 Approved by: https://github.com/janeyx99	2025-06-03 06:18:59 +00:00
PyTorch MergeBot	a7e496a896	Revert "[dynamo] Record the pre-graph bytecode using fast record function event (#154769 )" This reverts commit 409c396a48584de1ab14e1be6957663d548ad89e. Reverted https://github.com/pytorch/pytorch/pull/154769 on behalf of https://github.com/seemethere due to This fails internal tests see [fburl.com/diff/67gyp7gp](https://fburl.com/diff/67gyp7gp) ([comment](https://github.com/pytorch/pytorch/pull/154769#issuecomment-2933629894))	2025-06-03 06:13:49 +00:00
PyTorch MergeBot	b86aaaae0b	Revert "[dynamo][guards] Flush cache to more accurately measure guard overhead (#154764 )" This reverts commit 7dee89913072f1499c5265d8e92d23c30fc6a7f1. Reverted https://github.com/pytorch/pytorch/pull/154764 on behalf of https://github.com/seemethere due to This fails internal tests see [fburl.com/diff/67gyp7gp](https://fburl.com/diff/67gyp7gp) ([comment](https://github.com/pytorch/pytorch/pull/154769#issuecomment-2933629894))	2025-06-03 06:13:49 +00:00
henrylhtsang	d375e64279	[cutlass backend][forward fix] hex the cutlass key instead of decode (#154885 ) This is mainly following how it is done for torch_key. Error was: ``` UnicodeDecodeError: 'utf-8' codec can't decode bytes in position 0-1: invalid continuation byte ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154885 Approved by: https://github.com/jingsh, https://github.com/mlazos	2025-06-03 06:00:16 +00:00
Yuanhao Ji	8af447224e	Improve error message for `torch.fft.ihfft2` when input's dtype is complex (#149692 ) Fixes #149625 For the case mentioned in the issue, will get: ``` RuntimeError: Only supports floating-point dtypes, but found: ComplexDouble ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/149692 Approved by: https://github.com/malfet	2025-06-03 05:54:56 +00:00
henrylhtsang	295ea202f6	[inductor] Add kernel_hash_key to ChoiceCaller (#154470 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154470 Approved by: https://github.com/mlazos	2025-06-03 04:01:49 +00:00
cyy	388912dd94	Remove AttributeError constructor (#154808 ) It is a private API and uses C vsnprintf, which is not type safe. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154808 Approved by: https://github.com/Skylion007 Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-06-03 03:49:09 +00:00
PyTorch MergeBot	ef92653022	Revert "Remove AttributeError constructor (#154808 )" This reverts commit 3239da0c732c4ad736df7081ea44c1cd79c01145. Reverted https://github.com/pytorch/pytorch/pull/154808 on behalf of https://github.com/cyyever due to Need format code ([comment](https://github.com/pytorch/pytorch/pull/154808#issuecomment-2933286113))	2025-06-03 03:40:41 +00:00
Wei Feng	b3cb0e83de	[FSDP2] respect reshard_after_forward=True for root model (#154704 ) resolve https://github.com/pytorch/pytorch/issues/154655 `fully_shard(root, reshard_after_forward=True)` didn't really reshard parameters after forward, because we assumed root model will be used in backward immeidately. The assumption becomes invalid in 2 cases * we have 3 roots for CLIP, T5, FLUX. we should reshard parameters are CLIP and T5 immeidately after their forward for recommendation model, we may have mutiple root for dense part Change default beahvior to always respect `reshard_after_forward=True` Differential Revision: [D75663200](https://our.internmc.facebook.com/intern/diff/D75663200) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154704 Approved by: https://github.com/mori360	2025-06-03 03:12:45 +00:00
Anatoly Myachev	ff35c0cdfd	[inductor] Change `_constexpr_to_value` -> `_unwrap_if_constexpr` (#154905 ) To adapt to the changes from: `f480e2f697` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154905 Approved by: https://github.com/davidberard98	2025-06-03 03:10:56 +00:00
cyy	e3cf73ee49	Move remaining CI jobs to VS 2022 (#154811 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/154811 Approved by: https://github.com/huydhn	2025-06-03 02:21:24 +00:00
Yuanyuan Chen	3239da0c73	Remove AttributeError constructor (#154808 ) It is a private API and uses C vsnprintf, which is not type safe. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154808 Approved by: https://github.com/Skylion007 Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com>	2025-06-03 02:18:51 +00:00
David Berard	28cb3c0fe5	[test][inductor] attempt to fix duplicate registration issue (#154865 ) Fixes #154216 In #154216, there's a duplicate registration error thrown from registering `test::foo` twice. I expect that this is caused by having two tests that both register a `test::foo` op in the same test file. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154865 Approved by: https://github.com/NikhilAPatel, https://github.com/jingsh	2025-06-03 01:11:47 +00:00
David Berard	6cb6da6ea2	[triton pin][test] relax codecache test checks for number of triton artifacts (#154879 ) Triton has added another artifact that gets generated (triton-lang/triton#6992), so `test_cache_load_function` started failing as there are now 8 (instead of 7) artifacts. Instead of figuring out a way to check exactly which set of artifacts will get generated, I instead modified the test to just check that there are _at least_ 6 artifacts, to account for different platforms (intel/amd/nvidia) and different triton versions (which may or may not have a `.source` artifact) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154879 Approved by: https://github.com/oulgen, https://github.com/masnesral	2025-06-03 00:52:54 +00:00
Isuru Fernando	7f44b589be	[dynamo] fix pruning locals with ShapeEnvSource (#154752 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154752 Approved by: https://github.com/zhxchen17	2025-06-03 00:35:11 +00:00
David Berard	47a142c3c2	[triton pin][tests] update inductor/profiler launch_(enter\|exit)_hooks tests (#154894 ) Fixes #154223 Triton has updated launch_(enter\|exit)_hooks so that they are now in `knobs`. @danzimm already fixed this in #152457 - this just updates the test. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154894 Approved by: https://github.com/jingsh, https://github.com/NikhilAPatel	2025-06-03 00:14:14 +00:00
Catherine Lee	731acbfb0b	[CI] Reuse old whl on PRs (#154662 ) Turn off main branch only gating for reusing old whls Pull Request resolved: https://github.com/pytorch/pytorch/pull/154662 Approved by: https://github.com/huydhn	2025-06-03 00:10:39 +00:00
Georgia Phillips	af9f18e87e	[nativert] Free stale execution frames (#154636 ) Summary: This was implemented in SR due to caching of runtime instances building up and causing some memory usage spikes after some large amount traffic went through the model, and then once traffic went down, SR was still caching all the previous usage. We need something similar on the Sigmoid side to make sure the static dispatch modules aren't hogging memory. Currently, all ExecutionFrame objects are being cached, and never freed if stale. Test Plan: Added extra execution frames in tmp commit D75257998 and ran local replayer test to confirm extra execution frames get cleaned up down to min size, which is set at 8 {F1978532047} Also tested by modifying load_net_predictor (modifications also in D75257998) to run benchmarkNumIterations twice - once with benchmarkNumThreads, and once with only one thread. Also set clearing interval at one second. Verified that execution frames get cleared when we drop down to one thread. {F1978558984} ``` buck2 test 'mode/dev-nosan' fbcode//sigmoid/inference/test_gpu:model_runner_test -- ModelRunnerTest.Basic_InterpreterCuda_Multithread_Cleanup --run-disabled --print-passing-details ``` Bifferential Revision: D75257992 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154636 Approved by: https://github.com/zhxchen17, https://github.com/dolpm	2025-06-02 23:44:12 +00:00
PyTorch MergeBot	37eb909c94	Revert "[Inductor] Add attention pattern for model DistilBert in transformers==4.44.2. (#154091 )" This reverts commit 7b25ff7cf2e6096c103da0068e417216a41be7a9. Reverted https://github.com/pytorch/pytorch/pull/154091 on behalf of https://github.com/seemethere due to I root caused this PR to some failures, I tried to resolve with https://github.com/pytorch/pytorch/pull/154923 but it looks like there are more failures with my fix ([comment](https://github.com/pytorch/pytorch/pull/154091#issuecomment-2932848880))	2025-06-02 23:22:43 +00:00
PyTorch MergeBot	ac65e94f45	Revert "[Inductor UT] Reuse test_fused_attention.py for Intel GPU. (#154110 )" This reverts commit 2dfc0e33273fe50dcbb3d363da02c8cc485b4adc. Reverted https://github.com/pytorch/pytorch/pull/154110 on behalf of https://github.com/seemethere due to This is part of a stack with failures internally, I tried to resolve with https://github.com/pytorch/pytorch/pull/154923 but it looks like there are more failures ([comment](https://github.com/pytorch/pytorch/pull/154110#issuecomment-2932845168))	2025-06-02 23:20:11 +00:00
PyTorch MergeBot	e3af628b0d	Revert "Add CPython exception tests (#150789 )" This reverts commit 67fb9b7cc3f7d2ebbb104296f2b11776f4adbb22. Reverted https://github.com/pytorch/pytorch/pull/150789 on behalf of https://github.com/seemethere due to This is failing upstream in trunk, see `67fb9b7cc3` ([comment](https://github.com/pytorch/pytorch/pull/150789#issuecomment-2932823586))	2025-06-02 23:12:15 +00:00
Animesh Jain	7dee899130	[dynamo][guards] Flush cache to more accurately measure guard overhead (#154764 ) We observed that guard overhead at runtime using profiler traces was higher than reported in this profiling function at the compile time. After investigation, we found that f_locals are already in cache and that was causing the guard overhead to be way smaller while profiling during the compilation. To be more realistic, we flush the cache here. Profiling the guard overhead during compilation (in addition to at runtime) allows faster iteration time, and logging in tlparse and internal databases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154764 Approved by: https://github.com/zou3519, https://github.com/jansel, https://github.com/StrongerXi ghstack dependencies: #154769	2025-06-02 23:01:58 +00:00
Animesh Jain	409c396a48	[dynamo] Record the pre-graph bytecode using fast record function event (#154769 ) ![image](https://github.com/user-attachments/assets/1d06618b-1c14-4ed5-ab7b-dcfecbb4d632) Adds another event in the profiler traces. This can help us find models where pre-graph bytecode is very expensive. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154769 Approved by: https://github.com/zou3519, https://github.com/williamwen42, https://github.com/StrongerXi, https://github.com/jansel	2025-06-02 22:33:27 +00:00
eellison	f6b83d4cc6	sort iteration over index vars (#154846 ) Fix for https://github.com/pytorch/pytorch/issues/154741 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154846 Approved by: https://github.com/Skylion007, https://github.com/bdhirsh	2025-06-02 22:06:00 +00:00
Catherine Lee	d6420d4f85	[CI] Reuse old whl: replace the version (#154773 ) Replace the git version, so whl name goes from `torch-something+git<old commit>` to `torch-something+git<new commit>` Renamed a bunch of variables to hopefully be more clear Tested on `ef210ad54b` * Removed gating that prevents it from running on PRs (which is going to be merged soon) * Removed gating that checks for which files can be changed (since this PR has stuff outside of the acceptable list) * The above two allow the whl to be reused, and I added assert 1 == 2 in common_utils and checked that jobs failed (meaning they were using updated code despite not building) Checked that the whl in the docker image has the right commit sha, didn't check torch.__version__ though Pull Request resolved: https://github.com/pytorch/pytorch/pull/154773 Approved by: https://github.com/malfet	2025-06-02 22:02:41 +00:00
Catherine Lee	e1644e40a7	[ez][TD] Fix TD indexer workflow (#154868 ) Update docker image, and fix gpu flag env var Example failure: https://github.com/pytorch/pytorch/actions/runs/15381170311/job/43272174443 Tested on `9cb28f03e5` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154868 Approved by: https://github.com/Skylion007	2025-06-02 21:33:19 +00:00
Henry Tsang	104c31598f	[cutlass backend][ez] Make load config from local more resilient (#154740 ) Differential Revision: D75693211 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154740 Approved by: https://github.com/ColinPeppler	2025-06-02 21:12:12 +00:00
Guilherme Leobas	731e635c95	Add CPython math/cmath tests (#150794 ) Tests: * test_math.py * test_cmath.py Minor changes were made to each test to run them inside Dynamo One can reproduce the changes by downloading the tests from CPython and applying the diff: ```bash for f in "test_math" "test_cmath"; do wget -O "test/dynamo/cpython/3_13/${f}.py" "https://raw.githubusercontent.com/python/cpython/refs/heads/3.13/Lib/test/${f}.py" git apply "test/dynamo/cpython/3_13/${f}.diff" done ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150794 Approved by: https://github.com/zou3519	2025-06-02 20:49:44 +00:00
Guilherme Leobas	67fb9b7cc3	Add CPython exception tests (#150789 ) ---- * test_baseexception.py * test_exceptions.py * test_exception_variations.py * test_raise.py * test_sys.py Minor changes were made to each test to run them inside Dynamo One can reproduce the changes by downloading the tests from CPython and applying the diff: ```bash for f in "test_raise" "test_sys" "test_exceptions" "test_baseexception" "test_exception_variations"; do wget -O "test/dynamo/cpython/3_13/${f}.py" "https://raw.githubusercontent.com/python/cpython/refs/heads/3.13/Lib/test/${f}.py" git apply "test/dynamo/cpython/3_13/${f}.diff" done ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150789 Approved by: https://github.com/zou3519	2025-06-02 20:44:41 +00:00
Wei Wang	48807d568e	[CI][CUDA] Migrate remaining cu118 jobs to cu128 (#154169 ) Contributing to the fix of #147383 and #154119 Additional steps required: `3218b1b684/.github/workflows/lint.yml` cu118 needs to be updated. Make install_cuda.sh accept both 12.8 and 12.8.* as CUDA_VERSION argument. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154169 Approved by: https://github.com/eqy, https://github.com/malfet, https://github.com/atalman, https://github.com/tinglvv	2025-06-02 20:22:14 +00:00
Ryan Guo	9d3ad82ca7	[dynamo] Remove all `skipIfTorchDynamo` in `test_tensor_creation_ops.py` (#154693 ) Looks like they are no longer needed. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154693 Approved by: https://github.com/Skylion007, https://github.com/zou3519	2025-06-02 20:14:35 +00:00
bobrenjc93	984b1a80e3	[ez] add docs for *eager_then_compile stances (#154818 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154818 Approved by: https://github.com/williamwen42 ghstack dependencies: #154802, #154826, #154822, #154823, #154805	2025-06-02 19:04:35 +00:00
bobrenjc93	28f27886eb	Vary batch size when running dynamic shapes benchmarks (#154805 ) This better measures the actual runtime performance of dynamic shapes where we aren't guaranteed to have similar shapes as the original hint. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154805 Approved by: https://github.com/Skylion007 ghstack dependencies: #154802, #154826, #154822, #154823	2025-06-02 18:56:18 +00:00
bobrenjc93	33f2d0ff45	add reference to stances from dynamic shapes doc (#154823 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154823 Approved by: https://github.com/Skylion007, https://github.com/williamwen42 ghstack dependencies: #154802, #154826, #154822	2025-06-02 18:47:19 +00:00
bobrenjc93	d99e9568ec	Add docs for how to mark as unbacked (#154822 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154822 Approved by: https://github.com/Skylion007 ghstack dependencies: #154802, #154826	2025-06-02 18:30:57 +00:00
Animesh Jain	1258aac1c2	[dynamo] Upcast torch.Size + tuple to be of size torch.Size (#154830 ) Fixes https://github.com/pytorch/pytorch/issues/154432 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154830 Approved by: https://github.com/StrongerXi, https://github.com/Skylion007, https://github.com/williamwen42	2025-06-02 17:57:23 +00:00
bobrenjc93	9fe1b40d17	[ez] add dynamic sources docs (#154826 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154826 Approved by: https://github.com/Skylion007 ghstack dependencies: #154802	2025-06-02 17:53:30 +00:00
PyTorch MergeBot	69e22301da	Revert "[inductor] Add kernel_hash_key to ChoiceCaller (#154470 )" This reverts commit 7a79de1c0f31200f95a48a9e69fbd2df2a3c735d. Reverted https://github.com/pytorch/pytorch/pull/154470 on behalf of https://github.com/seemethere due to Failing internal inductor tests, author is aware and suggested revert. D75767762 ([comment](https://github.com/pytorch/pytorch/pull/154470#issuecomment-2931717432))	2025-06-02 17:43:23 +00:00
Oguz Ulgen	113224b530	Enable non blocking remote cache write (#154837 ) Test Plan: Ran ``` buck2 run mode/opt //scripts/oulgen:runner ``` twice and got https://fburl.com/scuba/pt2_remote_cache/u7u1uqh1 Differential Revision: D75770423 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154837 Approved by: https://github.com/jamesjwu	2025-06-02 17:36:43 +00:00
PyTorch MergeBot	67067512a1	Revert "[BE] Cleanup old ExecuTorch codegen and runtime code (#154165 )" This reverts commit 515c19a3856e953c0fe23a0ed4fa844f8eea34d8. Reverted https://github.com/pytorch/pytorch/pull/154165 on behalf of https://github.com/seemethere due to This is failing when attempting to test against executorch main internally, author has acknowledged that this should be reverted ([comment](https://github.com/pytorch/pytorch/pull/154165#issuecomment-2931489616))	2025-06-02 16:28:46 +00:00
Joona Havukainen	981bdb39ca	Enable ConvTranspose3D for FP32 and Complex64 (#154696 ) Fixes #154615 Enables using ConvTranspose3D since it seems support exists both on MacOS 14 and 15. For the half dtypes the discrepancy of CPU and GPU implementations is too large to conclude whether there is a bug in the implementation or not without a more rigorous study on what bounds are there to the expected error. So they are left unsupported for now and an assert is added to notify the user if the op is called with fp16 or bf16 inputs. Tests for ConvTranspose3D were enabled for the supported data types. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154696 Approved by: https://github.com/malfet	2025-06-02 16:24:03 +00:00
angelayi	77d85a4629	Symintify baddbmm (#154656 ) Previously we would specialize on the shape in this if-statement Pull Request resolved: https://github.com/pytorch/pytorch/pull/154656 Approved by: https://github.com/pianpwk	2025-06-02 15:23:14 +00:00
angelayi	e22be781b7	Symintify repeat_interleave (#154660 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/154660 Approved by: https://github.com/pianpwk	2025-06-02 15:19:39 +00:00
cyy	f6275bf0fe	Bump pocketfft submodule to the latest (#154845 ) Fixes #154843 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154845 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-06-02 14:54:13 +00:00
Anthony Shoumikhin	dfd6849e77	Update lint_urls.sh (#154838 ) Do not match empty urls pieces like "https://" Add headers for better handling urls like "https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/data-sheets/amd-instinct-mi300x-data-sheet.pdf" Pull Request resolved: https://github.com/pytorch/pytorch/pull/154838 Approved by: https://github.com/Skylion007	2025-06-02 14:50:34 +00:00
PyTorch UpdateBot	c65e9ad77a	Update slow tests (#154347 ) This PR is auto-generated weekly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/weekly.yml). Update the list of slow tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154347 Approved by: https://github.com/pytorchbot	2025-06-02 11:30:56 +00:00
Pearu Peterson	ff4515fde5	Add optional check_pinning argument to _validate_sparse_compressed_tensor/coo_args (#154759 ) As in the title. A prerequisite to https://github.com/pytorch/pytorch/pull/154638 . Pull Request resolved: https://github.com/pytorch/pytorch/pull/154759 Approved by: https://github.com/amjames, https://github.com/ngimel ghstack dependencies: #154610	2025-06-02 10:17:07 +00:00
Pearu Peterson	3f3c1f419f	User-controlled sparse tensor validation when loading data from external storage (#154610 ) This PR lets users to control sparse tensor invariants validation (that can be expensive, especially, for sparse tensors with many indices) when loading data from external sources. By default, the validation of sparse tensor invariants is disabled. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154610 Approved by: https://github.com/amjames, https://github.com/ngimel	2025-06-02 10:17:07 +00:00
PyTorch UpdateBot	9258cfc227	[audio hash update] update the pinned audio hash (#154776 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154776 Approved by: https://github.com/pytorchbot	2025-06-02 05:36:13 +00:00
Wei Wang	16d05e130c	[CI][CUDA][UCC] Update test_c10d_ucc.py - remove xfailIfLinux because it now succeeds (#150979 ) pytest -v test/distributed/test_c10d_ucc.py -k test_save_load ============================================================================================== test session starts ============================================================================================== platform linux -- Python 3.12.3, pytest-8.1.1, pluggy-1.5.0 -- /usr/bin/python cachedir: .pytest_cache hypothesis profile 'default' -> database=DirectoryBasedExampleDatabase(PosixPath('/opt/pytorch/pytorch/.hypothesis/examples')) rootdir: /opt/pytorch/pytorch configfile: pytest.ini plugins: anyio-4.9.0, hypothesis-6.130.13, flakefinder-1.1.0, rerunfailures-15.0, xdist-3.6.1, xdoctest-1.0.2, typeguard-4.3.0 collected 63 items / 62 deselected / 1 selected Running 1 items in this shard test/distributed/test_c10d_ucc.py::DistributedDataParallelTest::test_save_load_checkpoint PASSED [65.2581s] [100%] ================================================================================== 1 passed, 62 deselected in 68.78s (0:01:08) @ptrblck @eqy @tinglvv @atalman @malfet Pull Request resolved: https://github.com/pytorch/pytorch/pull/150979 Approved by: https://github.com/eqy	2025-06-02 03:24:35 +00:00
Zeming Lin	cd3d2b75b3	Update README.md - James has the wrong github link. (#151473 ) Unless I'm wrong, the James on the pytorch paper is not the account linked to in the README.md. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151473 Approved by: https://github.com/albanD	2025-06-02 01:53:44 +00:00
Mengwei Liu	515c19a385	[BE] Cleanup old ExecuTorch codegen and runtime code (#154165 ) Summary: These files are added to pytorch/pytorch before ExecuTorch is opensourced. Now is a good time to remove it from pytorch/pytorch, since the code is moved to pytorch/executorch already. Test Plan: Rely on CI jobs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154165 Approved by: https://github.com/kimishpatel, https://github.com/Skylion007, https://github.com/cyyever	2025-06-02 01:47:02 +00:00
angelayi	0d0058d90d	Fix flaky test in test_custom_ops (#152484 ) Hopefully fixes https://github.com/pytorch/pytorch/issues/151301, https://github.com/pytorch/pytorch/issues/151281 by making the ops have different names Pull Request resolved: https://github.com/pytorch/pytorch/pull/152484 Approved by: https://github.com/zou3519	2025-06-02 01:45:28 +00:00
Aaron Gokaslan	80af98c6c3	[BE]: Update nlohmann submodule to 3.12.0 (#154817 ) This is mostly compiler fixes, C++20 fixes, and clang-tidy fixes. Should be entirely backwards compatible with our current version Pull Request resolved: https://github.com/pytorch/pytorch/pull/154817 Approved by: https://github.com/jansel, https://github.com/malfet	2025-06-02 01:29:58 +00:00
Aaron Gokaslan	2b2245d5db	[BE]: Replace printf with fmtlib call (#154814 ) Safer, faster, more concise, and better type checking. Also add a few misc changes in the file. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154814 Approved by: https://github.com/jansel	2025-06-01 22:27:08 +00:00
Aaron Gokaslan	206e9d5160	[BE]: Update cpp-httplib submodule to 0.20.1 (#154825 ) Updates cpp-httplib to 0.20.1. This mostly updates OSS with a bunch of CMake, CXX compiler errors, and bugfixes from upstream. It's a header only library so should be pretty straightforward to upgrade Pull Request resolved: https://github.com/pytorch/pytorch/pull/154825 Approved by: https://github.com/malfet	2025-06-01 21:44:23 +00:00
Aaron Gokaslan	064bb3cebc	[BE]: Replace a couple of call sites with fmtlib printf (#154533 ) This is faster, and memory safe implementation of printf functions coming from fmtlib. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154533 Approved by: https://github.com/cyyever, https://github.com/jansel	2025-06-01 21:16:34 +00:00
Nikita Shulga	0350c7e72c	[BE] Introduce torch.AcceleratorError (#152023 ) Which inherits from `RuntimeError` and contains `error_code`, which in case of CUDA should contain error returned by `cudaGetLastError` `torch::detail::_new_accelerator_error_object(c10::AcceleratorError&)` follows the pattern of CPython's [`PyErr_SetString`](`cb8a72b301/Python/errors.c (L282)`), namely - Convert cstr into Python string with `PyUnicode_FromString` - Create new exception object using `PyObject_CallOneArg` just like it's done in [`_PyErr_CreateException`](`cb8a72b301/Python/errors.c (L32)`) - Set `error_code` property using `PyObject_SetAttrString` - decref all temporary references Test that it works and captures CPP backtrace (in addition to CI) by running ```python import os os.environ['TORCH_SHOW_CPP_STACKTRACES'] = '1' import torch x = torch.rand(10, device="cuda") y = torch.arange(20, device="cuda") try: x[y] = 2 print(x) except torch.AcceleratorError as e: print("Exception was raised", e.args[0]) print("Captured error code is ", e.error_code) ``` which produces following output ``` Exception was raised CUDA error: device-side assert triggered CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1 Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions. Exception raised from c10_cuda_check_implementation at /home/ubuntu/pytorch/c10/cuda/CUDAException.cpp:41 (most recent call first): C++ CapturedTraceback: #4 std::_Function_handler<std::shared_ptr<c10::LazyValue<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > const> (), c10::SetStackTraceFetcher(std::function<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > ()>)::{lambda()#1}>::_M_invoke(std::_Any_data const&) from Logging.cpp:0 #5 c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) from ??:0 #6 c10::cuda::c10_cuda_check_implementation(int, char const, char const, int, bool) [clone .cold] from CUDAException.cpp:0 #7 void at::native::gpu_kernel_impl<at::native::AbsFunctor<float> >(at::TensorIteratorBase&, at::native::AbsFunctor<float> const&) [clone .isra.0] from tmpxft_000191fc_00000000-6_AbsKernel.cudafe1.cpp:0 #8 at::native::abs_kernel_cuda(at::TensorIteratorBase&) from ??:0 #9 at::Tensor& at::native::unary_op_impl_with_complex_to_float_out<at::native::abs_stub_DECLARE_DISPATCH_type>(at::Tensor&, at::Tensor const&, at::native::abs_stub_DECLARE_DISPATCH_type&, bool) [clone .constprop.0] from UnaryOps.cpp:0 #10 at::(anonymous namespace)::(anonymous namespace)::wrapper_CUDA_out_abs_out(at::Tensor const&, at::Tensor&) from RegisterCUDA_0.cpp:0 #11 at::_ops::abs_out::call(at::Tensor const&, at::Tensor&) from ??:0 #12 at::native::abs(at::Tensor const&) from ??:0 #13 c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (at::Tensor const&), &at::(anonymous namespace)::(anonymous namespace)::wrapper_CompositeExplicitAutograd__abs>, at::Tensor, c10::guts::typelist::typelist<at::Tensor const&> >, at::Tensor (at::Tensor const&)>::call(c10::OperatorKernel, c10::DispatchKeySet, at::Tensor const&) from RegisterCompositeExplicitAutograd_0.cpp:0 #14 at::_ops::abs::redispatch(c10::DispatchKeySet, at::Tensor const&) from ??:0 #15 torch::autograd::VariableType::(anonymous namespace)::abs(c10::DispatchKeySet, at::Tensor const&) from VariableType_1.cpp:0 #16 c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (c10::DispatchKeySet, at::Tensor const&), &torch::autograd::VariableType::(anonymous namespace)::abs>, at::Tensor, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&> >, at::Tensor (c10::DispatchKeySet, at::Tensor const&)>::call(c10::OperatorKernel, c10::DispatchKeySet, at::Tensor const&) from VariableType_1.cpp:0 #17 at::_ops::abs::call(at::Tensor const&) from ??:0 #18 at::native::isfinite(at::Tensor const&) from ??:0 #19 c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (at::Tensor const&), &at::(anonymous namespace)::(anonymous namespace)::wrapper_CompositeImplicitAutograd__isfinite>, at::Tensor, c10::guts::typelist::typelist<at::Tensor const&> >, at::Tensor (at::Tensor const&)>::call(c10::OperatorKernel, c10::DispatchKeySet, at::Tensor const&) from RegisterCompositeImplicitAutograd_0.cpp:0 #20 at::_ops::isfinite::call(at::Tensor const&) from ??:0 #21 torch::autograd::THPVariable_isfinite(_object, _object, _object) from python_torch_functions_2.cpp:0 #22 PyObject_CallFunctionObjArgs from ??:0 #23 _PyObject_MakeTpCall from ??:0 #24 _PyEval_EvalFrameDefault from ??:0 #25 _PyObject_FastCallDictTstate from ??:0 #26 _PyStack_AsDict from ??:0 #27 _PyObject_MakeTpCall from ??:0 #28 _PyEval_EvalFrameDefault from ??:0 #29 _PyFunction_Vectorcall from ??:0 #30 _PyEval_EvalFrameDefault from ??:0 #31 _PyFunction_Vectorcall from ??:0 #32 _PyEval_EvalFrameDefault from ??:0 #33 _PyFunction_Vectorcall from ??:0 #34 _PyEval_EvalFrameDefault from ??:0 #35 PyFrame_GetCode from ??:0 #36 PyNumber_Xor from ??:0 #37 PyObject_Str from ??:0 #38 PyFile_WriteObject from ??:0 #39 _PyWideStringList_AsList from ??:0 #40 _PyDict_NewPresized from ??:0 #41 _PyEval_EvalFrameDefault from ??:0 #42 PyEval_EvalCode from ??:0 #43 PyEval_EvalCode from ??:0 #44 PyUnicode_Tailmatch from ??:0 #45 PyInit__collections from ??:0 #46 PyUnicode_Tailmatch from ??:0 #47 _PyRun_SimpleFileObject from ??:0 #48 _PyRun_AnyFileObject from ??:0 #49 Py_RunMain from ??:0 #50 Py_BytesMain from ??:0 #51 __libc_init_first from ??:0 #52 __libc_start_main from ??:0 #53 _start from ??:0 Captured error code is 710 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/152023 Approved by: https://github.com/eqy, https://github.com/mradmila, https://github.com/ngimel ghstack dependencies: #154436	2025-06-01 21:02:43 +00:00
Nikita Shulga	f7c09f864a	[Docs] Reformat sparse example (#154785 ) Not sure why, but rst fails to colorize multiline inputs, but works fine for single line commands Test plan: \| [Before](https://docs.pytorch.org/docs/main/sparse.html#construction) \| [After](https://docs-preview.pytorch.org/pytorch/pytorch/154785/sparse.html#construction) \| \| ------------- \| ------------- \| \| <img width="466" alt="image" src="https://github.com/user-attachments/assets/96a5c52a-1804-4d05-a5cf-c10221aaddf6" /> \| <img width="477" alt="image" src="https://github.com/user-attachments/assets/99565288-5c0b-4e8e-bd60-f016ebc207b5" /> \| Fixes https://github.com/pytorch/pytorch/issues/154779 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154785 Approved by: https://github.com/janeyx99, https://github.com/Skylion007	2025-06-01 20:56:14 +00:00
JungHoyoun	c2e9115757	Fix typo in dcp module (#154815 ) Fixed the docstring in `validate_checkpoint_id` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154815 Approved by: https://github.com/Skylion007	2025-06-01 18:18:45 +00:00
bobrenjc93	b90fc2ec27	[ez] delete code that died a long time ago (#154802 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154802 Approved by: https://github.com/Skylion007	2025-06-01 14:57:03 +00:00
Aaron Gokaslan	0cd18ba1ca	[BE][Ez] Update deprecated pybind11 functions (#154798 ) * getType() is deprecated, replace it with new/proper static method. These are backwards compatible with old pybind11 versions we support. So break this off before we upgrade to pybind11 3.0 where these methods are dropped in #154115 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154798 Approved by: https://github.com/jansel, https://github.com/cyyever	2025-06-01 06:17:50 +00:00
Aaron Gokaslan	bfae151269	[BE][Ez]: Remove unneeded mypy suppressions (#154800 ) Improvements in typing have made this suppression unnecessary Pull Request resolved: https://github.com/pytorch/pytorch/pull/154800 Approved by: https://github.com/cyyever, https://github.com/jansel	2025-06-01 06:10:41 +00:00
Natalia Gimelshein	9cbbc2593b	test for 146431 (#154786 ) Adds test for #146431 that was fixed by #154746 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154786 Approved by: https://github.com/Skylion007, https://github.com/galv Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>	2025-06-01 04:17:54 +00:00
cyy	5616fa4a68	[Submodule] Bump flatbuffers to v24.12.23 (#143964 ) This sub-module has not been updated for a long time. Pull Request resolved: https://github.com/pytorch/pytorch/pull/143964 Approved by: https://github.com/Skylion007	2025-06-01 02:25:57 +00:00
Aaron Gokaslan	c33fc9dae3	[BE][Ez]: Update VulkanMemoryAllocator to 3.3.0 (#154796 ) Last update to this submodule was 3 years ago, and the API is pretty stable and this is a minor version release update. Part of a bunch of PRs to eradicate low CMake required versions Pull Request resolved: https://github.com/pytorch/pytorch/pull/154796 Approved by: https://github.com/jansel	2025-06-01 00:30:56 +00:00
Aaron Gokaslan	9ce2732b68	[BE][Ez]: Fully type nn.utils.clip_grad (#154801 ) Full types clip_grad and exposed typing annotations that were hidden by a bad decorator Pull Request resolved: https://github.com/pytorch/pytorch/pull/154801 Approved by: https://github.com/jansel	2025-05-31 23:06:45 +00:00
Aaron Gokaslan	dbad6d71c7	[BE][Ez]: Unskip conv1d MPS test (#154795 ) Fixes issue I noticed where conv1d test is skipped for complex types unconditionally Pull Request resolved: https://github.com/pytorch/pytorch/pull/154795 Approved by: https://github.com/jansel	2025-05-31 23:01:19 +00:00
Aaron Gokaslan	b85c460749	[BE][Ez]: Update NVTX submodule to 3.2.1 (#154797 ) Update NVTX3 submodule to 3.2.1. * Mostly improved compiler support, Python support, and better CMake and C++ support. * Also has a few new APIs to support fancy new features. * This is header only library so should be an easy non-invasive change. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154797 Approved by: https://github.com/jansel	2025-05-31 23:01:13 +00:00
Pearu Peterson	6a781619bf	Temporarily disable sparse tensor validation when loading from external storage. (#154758 ) As in the title per https://github.com/pytorch/pytorch/issues/153143#issuecomment-2917793067 . The plan is to workout a solution that will allow (1) disabling pinned memory check to fix the original issue and (2) switching off the sparse tensor validation for maximal performance in loading sparse tensors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154758 Approved by: https://github.com/amjames, https://github.com/ngimel	2025-05-31 19:45:44 +00:00
FindHao	c99e91b1d7	[BE]Enhance _get_clean_triton.py to auto-generate launch_params if missing (#154666 ) Previously, @Chillee wrote a script https://github.com/pytorch/pytorch/pull/125811 to remove inductor dependency for inductor compiled triton kernels. We'd like to automate the process of obtaining the launch parameters. Added functionality to the torch/utils/_get_clean_triton.py to automatically generate the launch_params file if it does not exist and the auto_generate_params flag is set to True. This includes running the input file in a subprocess with the appropriate environment variable. Updated the get_clean_triton function and the main script to support this new feature, allowing users to disable auto-generation via a command-line argument. # Test Plan test embedding op in TritonBench ``` # generate inductor compiled triton kernels TORCH_COMPILE_DEBUG=1 TORCHINDUCTOR_FX_GRAPH_CACHE=0 python run.py --op embedding --mode fwd --precision fp32 --metrics nsys_rep --only inductor_embedding --num-inputs 1 --input-id 11 # run the script to get rid of inductor dependency. By default, triton_only_repro.py is the output file name. python ~/pytorch/torch/utils/_get_clean_triton.py ~/tritonbench/torch_compile_debug/run_2025_05_29_14_47_50_497790-pid_849274/torchinductor/model__0_forward_1.0/output_code.py ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154666 Approved by: https://github.com/davidberard98	2025-05-31 19:27:56 +00:00
xu-shawn	c014e4bcaa	Fix typo in vec256 interleave2 (#154784 ) Fix a typo where the elements in a vector are mislabeled Pull Request resolved: https://github.com/pytorch/pytorch/pull/154784 Approved by: https://github.com/malfet, https://github.com/Skylion007	2025-05-31 14:17:10 +00:00
fan.mo	daff263062	[Functorch] Support Functorch for PrivateUse1 backend (#154700 ) This PR enable that functorch to be used in 3rd party backends. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154700 Approved by: https://github.com/zou3519	2025-05-31 07:28:45 +00:00
Nicolas Macchioni	15e9119a69	[BE] install_triton_wheel.sh update for internal dev (#154637 ) internal devgpu gets mad at `pip install ...` but `python3 -m pip install ...` is fine Pull Request resolved: https://github.com/pytorch/pytorch/pull/154637 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-05-31 06:57:56 +00:00
Animesh Jain	7368eeba5e	[dynamo][guards] Prevent LENGTH guard on nn modules (#154763 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154763 Approved by: https://github.com/williamwen42	2025-05-31 05:32:31 +00:00
henrylhtsang	7a79de1c0f	[inductor] Add kernel_hash_key to ChoiceCaller (#154470 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154470 Approved by: https://github.com/mlazos	2025-05-31 03:09:37 +00:00
PyTorch MergeBot	bd10ea4e6c	Revert "Use 3.27 as the minimum CMake version (#153153 )" This reverts commit ad26ec6abe51d528124bc5fbbacaa87aef077ab8. Reverted https://github.com/pytorch/pytorch/pull/153153 on behalf of https://github.com/cyyever due to It still breaks windows debug builds ([comment](https://github.com/pytorch/pytorch/pull/153153#issuecomment-2923997777))	2025-05-31 02:14:24 +00:00
Peter Y. Yeh	43390d8b13	ROCm Sparsity through HipSparseLT (#150578 ) TLDR: - This pull request introduces support for hipSPARSELt in ROCm, current usage would be semi-structure sparsity. - Require ROCm 6.4 && gfx942/gfx950. - The average performance uplift (compare to dense operation) is ~ 20% in ROCm 6.4 but expect further performance lift along the way. ### Dense vs. Sparse Performance Comparison #### NT (Row-major) Average Uplift: `1.20` \| M \| N \| K \| hipsparselt-bench (us) \| hipblaslt-bench get all (us) \| Uplift \| \|-------\|--------\|--------\|-------------------------\|-------------------------------\|--------\| \| 14336 \| 8 \| 4096 \| 20.05 \| 25.3 \| 1.26 \| \| 4096 \| 8 \| 14336 \| 21.07 \| 25.28 \| 1.20 \| \| 3072 \| 3072 \| 10240 \| 299.05 \| 351.82 \| 1.18 \| \| 3072 \| 1536 \| 768 \| 18.56 \| 20.05 \| 1.08 \| \| 3072 \| 17664 \| 768 \| 163.13 \| 173.91 \| 1.07 \| \| 3072 \| 196608 \| 768 \| 1717.30 \| 1949.63 \| 1.14 \| \| 3072 \| 24576 \| 768 \| 206.84 \| 242.98 \| 1.17 \| \| 3072 \| 6144 \| 768 \| 53.90 \| 56.88 \| 1.06 \| \| 3072 \| 98304 \| 768 \| 833.77 \| 962.28 \| 1.15 \| \| 768 \| 1536 \| 768 \| 8.53 \| 19.65 \| 2.30 \| \| 768 \| 17664 \| 768 \| 46.02 \| 46.84 \| 1.02 \| \| 768 \| 196608 \| 768 \| 463.15 \| 540.46 \| 1.17 \| \| 768 \| 24576 \| 768 \| 54.32 \| 59.55 \| 1.10 \| \| 768 \| 6144 \| 768 \| 19.47 \| 20.15 \| 1.03 \| \| 768 \| 98304 \| 768 \| 231.88 \| 258.73 \| 1.12 \| --- #### NN (Row-major) Average Uplift: `1.13` \| M \| N \| K \| hipsparselt-bench (us) \| hipblaslt-bench get all (us) \| Uplift \| \|-----\|--------\|-------\|-------------------------\|-------------------------------\|--------\| \| 768 \| 1536 \| 3072 \| 27.50 \| 28.78 \| 1.05 \| \| 768 \| 17664 \| 3072 \| 125.06 \| 158.94 \| 1.27 \| \| 768 \| 196608 \| 3072 \| 1568.38 \| 1767.12 \| 1.13 \| \| 768 \| 24576 \| 3072 \| 171.05 \| 203.49 \| 1.19 \| \| 768 \| 6144 \| 3072 \| 58.72 \| 60.39 \| 1.03 \| \| 768 \| 98304 \| 3072 \| 787.15 \| 887.60 \| 1.13 \| ------------------------- This pull request introduces support for hipSPARSELt in ROCm, alongside various updates and improvements to the codebase and test suite. The changes primarily involve adding configuration flags, updating conditional checks, and ensuring compatibility with hipSPARSELt. ### ROCm and hipSPARSELt Support: * [`BUILD.bazel`](diffhunk://#diff-7fc57714ef13c3325ce2a1130202edced92fcccc0c6db34a72f7b57f60d552a3R292): Added `@AT_HIPSPARSELT_ENABLED@` substitution to enable hipSPARSELt support. * [`aten/CMakeLists.txt`](diffhunk://#diff-0604597797bb21d7c39150f9429d6b2ace10b79ab308514ad03f76153ae8249bR104-R110): Introduced a conditional flag to enable hipSPARSELt support based on ROCm version. * [`aten/src/ATen/CMakeLists.txt`](diffhunk://#diff-ce80f3115ab2f6be5142f0678a1fc92c6b2d7727766ce44f48726c99e720f777R37): Added `AT_HIPSPARSELT_ENABLED` configuration. * [`aten/src/ATen/cuda/CUDAConfig.h.in`](diffhunk://#diff-8bb82da825ca87c28233abacffa1b0566c73a54990b7a77f3f5108d3718fea15R11): Defined `AT_HIPSPARSELT_ENABLED` macro. * `caffe2/CMakeLists.txt`, `cmake/Dependencies.cmake`, `cmake/public/LoadHIP.cmake`: Included hipSPARSELt in the ROCm dependencies. [[1]](diffhunk://#diff-c5ee05f1e918772792ff6f2a3f579fc2f182e57b1709fd786ef6dc711fd68b27R1380) [[2]](diffhunk://#diff-12e8125164bbfc7556b1781a8ed516e333cc0bf058acb7197f7415be44606c72L1084-R1084) [[3]](diffhunk://#diff-b98e27b9a5f196a6965a99ee5a7bb15b3fc633d6375b767635b1b04ccb2fd3d5R153) ### Codebase Updates: * [`aten/src/ATen/native/sparse/cuda/cuSPARSELtOps.cpp`](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3R1-R6): Added hipSPARSELt support checks and initialization functions. Updated various methods to conditionally handle hipSPARSELt. [[1]](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3R1-R6) [[2]](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3R22-R67) [[3]](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3R78-R85) [[4]](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3R97-R109) [[5]](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3R183-R188) [[6]](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3L134-R200) [[7]](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3R213-R222) [[8]](diffhunk://#diff-ae921dd1584ab98fdd9c25a3521047795de702223f5b65fdaa45a5bd92b4d1f3L217-R285) ### Test Suite Updates: * [`test/test_sparse_semi_structured.py`](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR50-R65): Added checks for hipSPARSELt availability and updated test conditions to skip tests not supported on ROCm. [[1]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR50-R65) [[2]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR228) [[3]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR239) [[4]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR250) [[5]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR579) [[6]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR624) [[7]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR661) [[8]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR695) [[9]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR730) [[10]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR755) [[11]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR771) [[12]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR809) [[13]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR844) [[14]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cL840-R854) [[15]](diffhunk://#diff-b7b57bc1e34145ef89c7929751d5d26aeecc8edfb37da9c60e9d3f0a1335133cR1005) Pull Request resolved: https://github.com/pytorch/pytorch/pull/150578 Approved by: https://github.com/jeffdaily	2025-05-31 02:03:40 +00:00
cyy	ad26ec6abe	Use 3.27 as the minimum CMake version (#153153 ) Update the minimum CMake version to 3.27 because of it provides more CUDA targets such as `CUDA::nvperf_host` so that it is possible to remove some of our forked CUDA modules. See https://github.com/pytorch/pytorch/pull/153783. It's also possible to facilitate future third-party updates such as FBGEMM (its current shipped version requires 3.21). Pull Request resolved: https://github.com/pytorch/pytorch/pull/153153 Approved by: https://github.com/malfet	2025-05-31 01:54:35 +00:00
PyTorch MergeBot	3e71016459	Revert "Aten vector default constructors set to 0, add fnmadd and fnmsub (#154298 )" This reverts commit 489afa829a248ca64c4b2dffe2e6d601b8816cf9. Reverted https://github.com/pytorch/pytorch/pull/154298 on behalf of https://github.com/izaitsevfb due to breaks linux-jammy-aarch64-py3.10 / build ([comment](https://github.com/pytorch/pytorch/pull/154298#issuecomment-2923966688))	2025-05-31 01:51:59 +00:00
Robert Burke	489afa829a	Aten vector default constructors set to 0, add fnmadd and fnmsub (#154298 ) Test Plan: The only functional change is zero-initialization instead of undefined-initialization. If tests pass, I think it should be fine. Differential Revision: D75345074 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154298 Approved by: https://github.com/swolchok	2025-05-31 01:32:45 +00:00
dolpm	472773c7f9	[nativert] move OpKernelKind enum to torch (#154756 ) Summary: att Test Plan: ci Differential Revision: D75703996 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154756 Approved by: https://github.com/zhxchen17, https://github.com/cyyever	2025-05-31 01:31:29 +00:00
Natalia Gimelshein	f01e628e3b	Resubmit Remove MemPoolContext (#154042 ) (#154746 ) Summary: Per title Test Plan: Added tests + existing tests Differential Revision: D75695030 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154746 Approved by: https://github.com/malfet	2025-05-31 01:21:54 +00:00
Siddharth Kotapati	932733e0e6	Fix memory leaks in mps_linear_nograph (#154765 ) Fixes some memory leaks which were identified as part of the investigation of https://github.com/pytorch/pytorch/issues/154329. This doesn't appear to be the whole solution but wanted to merge this anyway since it's a quick fix In my tests I see roughly 3MB of unexpected memory growth before this change, and after this change I see 2.2MB of memory growth Pull Request resolved: https://github.com/pytorch/pytorch/pull/154765 Approved by: https://github.com/malfet	2025-05-31 00:46:12 +00:00
PyTorch MergeBot	108422ac26	Revert "Use 3.27 as the minimum CMake version (#153153 )" This reverts commit 78624679a876a21acb14bf075ba6beccff21b9a0. Reverted https://github.com/pytorch/pytorch/pull/153153 on behalf of https://github.com/cyyever due to It still breaks windows debug builds ([comment](https://github.com/pytorch/pytorch/pull/153153#issuecomment-2923785799))	2025-05-31 00:28:03 +00:00
mori360	da4aacabac	Add h100_distributed label (#154562 ) Add h100_distributed label, testing distributed 3D composability tests on 8*H100 GPU node. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154562 Approved by: https://github.com/seemethere	2025-05-31 00:17:43 +00:00
David Berard	9b5308cd58	[upstream triton] support build with setup.py in ./python/ or in ./ (#154635 ) Upstream triton has moved setup.py from python/ to ./. This PR allows versions to be buildable by checking the location of setup.py and choosing the cwd of the build commands based on the location. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154635 Approved by: https://github.com/atalman	2025-05-31 00:15:43 +00:00
Catherine Lee	b019a33f8f	[ez][CI] Reuse old whl: remove old zip/whl (#154770 ) Forgot that unzip doesn't get rid of the zip so the old one is still there Unrelated: figure out how to update the git version Pull Request resolved: https://github.com/pytorch/pytorch/pull/154770 Approved by: https://github.com/ZainRizvi, https://github.com/malfet	2025-05-31 00:13:24 +00:00
PyTorch MergeBot	0fab32290a	Revert "[draft export] avoid storing intermediate real tensors in proxies (#154630 )" This reverts commit 5acb8d50801e6d110790993464611314dd1bd54b. Reverted https://github.com/pytorch/pytorch/pull/154630 on behalf of https://github.com/malfet due to This still ooms, at least occasionally see `78624679a8/1` ([comment](https://github.com/pytorch/pytorch/pull/154630#issuecomment-2923759745))	2025-05-31 00:07:56 +00:00
Yidi Wu	faf973da5e	[refactor] move materialize_as_graph to _higher_order_ops/utils.py (#154070 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154070 Approved by: https://github.com/zou3519	2025-05-31 00:06:44 +00:00
cyy	78624679a8	Use 3.27 as the minimum CMake version (#153153 ) Update the minimum CMake version to 3.27 because of it provides more CUDA targets such as `CUDA::nvperf_host` so that it is possible to remove some of our forked CUDA modules. See https://github.com/pytorch/pytorch/pull/153783. It's also possible to facilitate future third-party updates such as FBGEMM (its current shipped version requires 3.21). Pull Request resolved: https://github.com/pytorch/pytorch/pull/153153 Approved by: https://github.com/malfet	2025-05-31 00:01:52 +00:00
Pian Pawakapan	5f1c3c67b2	[pgo] log dynamic whitelist in PT2 Compile Events (#154747 ) Summary: logs the whitelist to PT2 Compile Events Test Plan: loggercli codegen GeneratedPt2CompileEventsLoggerConfig Reviewed By: bobrenjc93 Differential Revision: D75617963 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154747 Approved by: https://github.com/angelayi	2025-05-30 23:54:24 +00:00
Aaron Gokaslan	bbda22e648	[BE][Ez]: Optimize unnecessary lambda with operator (#154722 ) Automated edits performed by FURB118. Operator is implemented in C and way faster when passed to another C method like sorted, max etc as a `key=` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154722 Approved by: https://github.com/jansel	2025-05-30 23:47:10 +00:00
Catherine Lee	0f3db20132	[ez][CI] Do not reuse old whl if deleting files (#154731 ) Thankfully very few commits actually delete files so I don't think has affected anything Pull Request resolved: https://github.com/pytorch/pytorch/pull/154731 Approved by: https://github.com/Skylion007	2025-05-30 22:35:13 +00:00
David Berard	eb93c0adb1	[inductor][AMD] support special kwargs in AMD triton configs (#154605 ) Context: AMD triton kernels can be launched with special kwargs, like `waves_per_eu`. Triton configs with these kwargs look like this: ``` triton.Config({ "BLOCK_SIZE": 64, "waves_per_eu": 2, }) ``` in comparison, nvidia's special kwargs are explicit parameters on the config, e.g. num_warps: ``` triton.Config( {"BLOCK_SIZE": 64}, num_warps=4, ) ``` Problem: this causes custom triton kernels w/ PT2 to error out, because there's a kwarg in the triton.Config that doesn't appear in the kernel signature. Solution: When splicing in the constexpr values into the arg list, ignore any values in the config kwargs list if they don't appear in the function signature. Differential Revision: [D75599629](https://our.internmc.facebook.com/intern/diff/D75599629/) NOTE FOR REVIEWERS: This PR has internal Meta-specific changes or comments, please review them on [Phabricator](https://our.internmc.facebook.com/intern/diff/D75599629/)! Pull Request resolved: https://github.com/pytorch/pytorch/pull/154605 Approved by: https://github.com/njriasan	2025-05-30 22:24:32 +00:00
PyTorch MergeBot	1193bf0855	Revert "convert inductor codecache to use getArtifactLogger (#153766 )" This reverts commit 5b6fd277f954b789649501e21e9689a42d565e13. Reverted https://github.com/pytorch/pytorch/pull/153766 on behalf of https://github.com/malfet due to I want to revert this change as I'm 90+% certain it somehow broke testing ([comment](https://github.com/pytorch/pytorch/pull/153766#issuecomment-2923620806))	2025-05-30 22:20:07 +00:00
Justin Chu	26aa8dcf27	[ONNX] Simplify onnx test dependencies (#154732 ) Simplify onnx test dependencies and bump onnxscript to 0.3 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154732 Approved by: https://github.com/Skylion007	2025-05-30 21:58:04 +00:00
Pian Pawakapan	5acb8d5080	[draft export] avoid storing intermediate real tensors in proxies (#154630 ) Handles GC for non-strict draft export; GPU memory usage shouldn't be much more than eager mode + input tensors now. While trying to do draft export CPU offloading, I found out GC is feasible, because in non-strict, there's 2 places holding references to a `.real_tensor` attribute: 1) the FakeTensors in fake tensor prop, but these are held by the actual variables in the model's forward call, and so the real tensor gets gc-ed along with the fake one when the variable goes out of scope. 2) A clone of the fake tensor in 1) stored in `proxy.node.meta["val"]`, which was added in https://github.com/pytorch/pytorch/pull/150948. But we didn't actually need to store them on intermediate values; the placeholders are enough for retracing/lowering. Avoiding storing the intermediate values in 2), the values in 1) should be naturally GC-ed, and the real-tensor memory usage for non-strict should be pretty similar to eager computation? Strict still OOMs; dynamo still holds these in variable tracking, and not sure how to GC those. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154630 Approved by: https://github.com/angelayi, https://github.com/yushangdi	2025-05-30 21:06:55 +00:00
Feny Patel	abc2264e8f	remove another instance of mtia_workloadd from pytorch (#154739 ) Summary: ^ Test Plan: CIs Differential Revision: D75692171 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154739 Approved by: https://github.com/sraikund16	2025-05-30 20:50:46 +00:00
Paul Zhang	22a4cabd19	[Inductor] Add NaN assert to returned values from generated code (#154455 ) Summary: It is possible to have `reinterpret_tensor` in the output of inductor codegen, e.g. `reinterpret_tensor(buf366, (1024, ), (1, ), 0)` in the return tuple. This adds assertions to all return values from inductor codegen to prevent nans from slipping through and being hard to trace. Test Plan: NaN asserts properly generated in example gemm script: vars = (buf1, primals_2, buf2, primals_1, ) for var in vars: if isinstance(var, torch.Tensor): assert not var.isnan().any().item() assert not var.isinf().any().item() Pull Request resolved: https://github.com/pytorch/pytorch/pull/154455 Approved by: https://github.com/eellison	2025-05-30 20:32:56 +00:00
Aaron Gokaslan	ed1ff7d0fb	[BE][Ez]: Update mimalloc submodule to 2.2.3 (#154720 ) Updating minor version of mimalloc. The old version is more than 2 years old, and the newer release has performance fixes and compiler fixes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154720 Approved by: https://github.com/jansel	2025-05-30 20:17:13 +00:00
Aaron Gokaslan	2f03673ebf	[BE][Ez]: Enable ClangFormat aten/src/core/Formatting.cpp (#154719 ) Follow up to #152830 . Noticed the file was excluded from fromatting, opt in to clang-format since it's really close anyway. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154719 Approved by: https://github.com/jansel	2025-05-30 19:52:43 +00:00
Alessandro Sangiorgi	f57754e815	[Inductor] Record Triton’s Base32 Cache Key in .best_config for Debugging (#154618 ) This is a follow-up PR of the reverted one https://github.com/pytorch/pytorch/pull/148981 re-opening for visibility : Modified TorchInductor’s autotuning flow so that each best_config JSON file also includes the Triton “base32” (or base64) cache key. Motivation Debugging & Analysis: With this change, we can quickly identify which compiled binary and IRs belongs to a given best config. The impact is minimal since it is only an extra field in .best_config. It can help advanced performance tuning or kernel-level debugging. Also, since Triton already stores cubin/hsaco in its cache, developers/researchers can avoid to set store_cubin = True since they can get the cubin/hsaco in the Triton cache and with the code provided in this PR, they can easily match the best_config with the right Triton cache directory for the "best" kernel. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154618 Approved by: https://github.com/jansel	2025-05-30 19:30:25 +00:00
Isalia20	d6edefefbf	[CUDA] Fixes for backwards in memefficient attn for large tensors (#154663 ) followup to #154029. @ngimel Backwards had the same problem as well so this PR fixes it and adds support for logsumexp computation in the forward pass. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154663 Approved by: https://github.com/ngimel	2025-05-30 19:30:07 +00:00
Alexander Grund	d89d213118	Fix test_tensorboard when started w/o tensorboard package (#154709 ) If `TEST_TENSORBOARD == False` then `DataType` is not defined or imported. However it is used unconditionally when defining the test with `parametrize` which leads to an NameError crashing the test execution on start. Provide a Dummy to make it syntactially correct. Tests will be skipped on start. ``` File "/dev/shm/build/pytorch-v2.2.1/test/test_tensorboard.py", line 885, in <module> class TestTensorProtoSummary(BaseTestCase): File "/dev/shm/build/pytorch-v2.2.1/test/test_tensorboard.py", line 889, in TestTensorProtoSummary (torch.float16, DataType.DT_HALF), ^^^^^^^^ NameError: name 'DataType' is not defined Got exit code 1, retrying... test_tensorboard 1/1 failed! [Errno 2] No such file or directory: '/dev/shm/build/pytorch-v2.2.1/.pytest_cache/v/cache/stepcurrent/test_tensorboard_0_0dba8bc00bbe233f' ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154709 Approved by: https://github.com/Skylion007	2025-05-30 19:18:43 +00:00
atalman	22641f42b6	[Binary-builds]Use System NCCL by default in CI/CD. (#152835 ) Use System NCCl by default. The correct nccl version is already built into the Manylinux docker image. Will followup with PR on detecting if user has NCCL installed and enabling USE_SYSTEM_NCCL by default in this case. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152835 Approved by: https://github.com/malfet	2025-05-30 18:51:48 +00:00
Ryan Guo	967937872f	[dynamo] Remove dead code path for `torch.Tensor.view(*shape)` (#154646 ) This was introduced in early days of Dynamo, and looks like it's been fixed since -- the regression test `test_transpose_for_scores` passes. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154646 Approved by: https://github.com/Skylion007, https://github.com/zou3519 ghstack dependencies: #154645	2025-05-30 18:50:58 +00:00
Ryan Guo	f9dc20c7a3	[dynamo] Fix syntax error in aot graph from kwarg-less `torch.Tensor.[random_\|uniform_]` calls (#154645 ) As title, fixes #151432, see more context in the issue discussion. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154645 Approved by: https://github.com/zou3519	2025-05-30 18:50:58 +00:00
PyTorch MergeBot	fb67fa9968	Revert "[Inductor] Add NaN assert to returned values from generated code (#154455 )" This reverts commit aec3ef100844631cb7c4ce2725157984eb9cebfe. Reverted https://github.com/pytorch/pytorch/pull/154455 on behalf of https://github.com/malfet due to Looks like it broke inductor/test_compile_subprocess.py::CpuTests::test_AllenaiLongformerBase, see `35fc5c49b4/1`(default%2C%20&mergeEphemeralLF=true ([comment](https://github.com/pytorch/pytorch/pull/154455#issuecomment-2923154249))	2025-05-30 18:45:01 +00:00
PyTorch MergeBot	35fc5c49b4	Revert "[internal] Expose additional metadata to compilation callbacks (#153596 )" This reverts commit f889dea97dad3cc506d43e379a469334417040c8. Reverted https://github.com/pytorch/pytorch/pull/153596 on behalf of https://github.com/izaitsevfb due to introduces bunch of callback-related failures on rocm ([comment](https://github.com/pytorch/pytorch/pull/153596#issuecomment-2923139061))	2025-05-30 18:39:27 +00:00
Aaron Gokaslan	b6b9311f4f	[BE][Ez]: Fix typo in dynamo utils #154639 (#154748 ) Fixes a typo in #154639 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154748 Approved by: https://github.com/ngimel	2025-05-30 18:39:01 +00:00
Guilherme Leobas	bbdf469f0e	Add CPython dict tests (#150791 ) Tests: * test_dict.py * test_ordered_dict.py * test_userdict.py Minor changes were made to each test to run them inside Dynamo One can reproduce the changes by downloading the tests from CPython and applying the diff: ```bash for f in "test_dict" "test_ordered_dict" "test_userdict"; do wget -O "test/dynamo/cpython/3_13/${f}.py" "https://raw.githubusercontent.com/python/cpython/refs/heads/3.13/Lib/test/${f}.py" git apply "test/dynamo/cpython/3_13/${f}.diff" done ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150791 Approved by: https://github.com/zou3519	2025-05-30 18:17:09 +00:00
Aaron Gokaslan	2120eeb8de	[BE][Ez]: Improve dynamo utils typing with TypeIs and TypeGuard (#154639 ) Adds some additional TypeIs and TypeGuard to some _dynamo utils for additional type narrowing Pull Request resolved: https://github.com/pytorch/pytorch/pull/154639 Approved by: https://github.com/jansel	2025-05-30 18:09:50 +00:00
zeshengzong	1b569e5490	Fix load_state_dict description (#154599 ) Fixes #141364 Fix missing description in `assign` param ## Test Result ### Before ![image](https://github.com/user-attachments/assets/5928c691-4e31-463b-aa0a-86eb8bb452e5) ### After ![image](https://github.com/user-attachments/assets/036631a2-0f20-4a71-95c3-2c0fd732293e) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154599 Approved by: https://github.com/colesbury, https://github.com/mikaylagawarecki	2025-05-30 18:08:59 +00:00
Shivam Raikundalia	30ac7f4d4e	[EZ/Memory Snapshot] Remove Handle even if compile_context not set (#154664 ) Summary: When setting the memory snapshot callback we register and unregister callbacks for performance reasons. For ease of use, it makes sense to just remove all callbacks regardless of which flags are enabled. The enable stays behind a feature flag, this just changes the disable to ignore the flag itself. Test Plan: Ran without any flags and saw all callbacks removed. Differential Revision: D75636035 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154664 Approved by: https://github.com/sanrise, https://github.com/aaronenyeshi	2025-05-30 18:08:37 +00:00
dolpm	65d8dba735	[nativert] move layout planner settings to torch (#154668 ) Summary: att Test Plan: ci Differential Revision: D75633031 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154668 Approved by: https://github.com/zhxchen17	2025-05-30 17:33:27 +00:00
Sidharth	3bdceab124	[dynamo] fix: added star operator for graph_break_hints (#154713 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154713 Approved by: https://github.com/zou3519, https://github.com/williamwen42	2025-05-30 17:31:03 +00:00
Henry Hu	802ffd06c8	[Export] Add math module for deserialization (#154643 ) Summary: As title Test Plan: ci Differential Revision: D75580646 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154643 Approved by: https://github.com/yushangdi	2025-05-30 17:29:25 +00:00
Aaron Orenstein	fc0135ca11	Re-enable FakeTensor caching for SymInts (#152662 ) Summary: This backs out D60320595 which itself turned off FakeTensor caching when a SymInt was present. There has been a lot of dynamic shape fixes done this year and tests pass so I'm assuming some of that work fixed what was breaking previously. Test Plan: Reran the tests listed in T196779132 and they pass. ## Perf ### Instruction Counter Benchmark: - 26% win on add_loop_eager_dynamic - 13% win on add_loop_inductor_dynamic_gpu ### Perf Dashboard Compilation Latency wins across the board but especially strong on the dynamic tests (like cudagraphs_dynamic) - for example MobileBertForMaskedLM went from 66s -> 50s. Differential Revision: [D75467694](https://our.internmc.facebook.com/intern/diff/D75467694) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152662 Approved by: https://github.com/anijain2305	2025-05-30 17:23:36 +00:00
Pian Pawakapan	3027051590	[export] avoid float/bool specialization for scalar tensor construction (#154661 ) Fixes #153411 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154661 Approved by: https://github.com/angelayi	2025-05-30 17:18:21 +00:00
bobrenjc93	e7bf72c908	[multigraph] fix composabilty with aotautograd cache (#153526 ) AOTAutogradCache uses FXGraphCache which uses the tracing context to get the ShapeEnv. Although the TracingContext global_context is cleared by the time we get around to reusing it, we don't actually need it. We just need the ShapeEnv in the TracingContext, which isn't cleared at the end of dynamo and does persist. This PR adds the tracing context manager around the specialized compile to ensure our caching infrastructure can get access to the ShapeEnv. A test was also added to prove correctness. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153526 Approved by: https://github.com/jamesjwu, https://github.com/zou3519 ghstack dependencies: #153433, #153449	2025-05-30 16:56:17 +00:00
Ryan Guo	7183f52675	[dynamo] Support namedtuple subclass (#153982 ) Fixes #133762. This involves 1. support tuple subclass constructed inside compile region. 2. handle the "fake" global scope associated with NamedTuple-generated `__new__`. 3. handle `namedtuple._tuplegetter` more faithfully. Differential Revision: [D75488091](https://our.internmc.facebook.com/intern/diff/D75488091) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153982 Approved by: https://github.com/jansel ghstack dependencies: #154176	2025-05-30 16:14:37 +00:00
Ryan Guo	8002d22ce3	[dynamo] Trace into descriptor with `__set__` (#154176 ) As title, this patch basically implements https://github.com/python/cpython/blob/3.11/Objects/object.c#L1371-L1452, and make the `__get__` handling more robust. I ran into this while fixing #133762. Differential Revision: [D75488090](https://our.internmc.facebook.com/intern/diff/D75488090) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154176 Approved by: https://github.com/jansel	2025-05-30 16:14:37 +00:00
PyTorch MergeBot	31f95b5d2e	Revert "inductor codecache: include private inductor configs in cache key (#153672 )" This reverts commit 2c1cb38d9516e10474b4f12a2e839046648a71a8. Reverted https://github.com/pytorch/pytorch/pull/153672 on behalf of https://github.com/malfet due to Looks like it regressed pr_time_benchmarks, see `ba3f91af97/1` ([comment](https://github.com/pytorch/pytorch/pull/153672#issuecomment-2922759739))	2025-05-30 15:54:14 +00:00
Guilherme Leobas	4b1f047a33	Add CPython list/tuple tests (#150790 ) Tests: * test_list.py * test_tuple.py * test_userlist.py Minor changes were made to each test to run them inside Dynamo One can reproduce the changes by downloading the tests from CPython and applying the diff: ```bash for f in "test_raise" "test_list" "test_tuple" "test_userlist"; do wget -O "test/dynamo/cpython/3_13/${f}.py" "https://raw.githubusercontent.com/python/cpython/refs/heads/3.13/Lib/test/${f}.py" git apply "test/dynamo/cpython/3_13/${f}.diff" done ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/150790 Approved by: https://github.com/williamwen42	2025-05-30 15:53:38 +00:00
Randolf Scholz	ba3f91af97	Type hints for distributions/utils (#154712 ) Fixes #144196 Part of #144219 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154712 Approved by: https://github.com/Skylion007	2025-05-30 15:50:31 +00:00
Bin Bao	0f81c7a28d	[CI] Pin the torchao version used when testing torchbench (#154723 ) Summary: To fix a recent CI breakage. As a follow-up, the torchao pin in .github/ci_commit_pins/torchao.txt is 6-month old. We should bump up that once we verify this fix works. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154723 Approved by: https://github.com/eellison	2025-05-30 15:04:26 +00:00
PyTorch MergeBot	7e8532077f	Revert "Use 3.27 as the minimum CMake version (#153153 )" This reverts commit 1ece53b157db4425ad12cae31fb570c591dc19e7. Reverted https://github.com/pytorch/pytorch/pull/153153 on behalf of https://github.com/cyyever due to It still breaks windows debug builds ([comment](https://github.com/pytorch/pytorch/pull/153153#issuecomment-2922369830))	2025-05-30 13:16:33 +00:00
cyy	1ece53b157	Use 3.27 as the minimum CMake version (#153153 ) Update the minimum CMake version to 3.27 because of it provides more CUDA targets such as `CUDA::nvperf_host` so that it is possible to remove some of our forked CUDA modules. See https://github.com/pytorch/pytorch/pull/153783. It's also possible to facilitate future third-party updates such as FBGEMM (its current shipped version requires 3.21). Pull Request resolved: https://github.com/pytorch/pytorch/pull/153153 Approved by: https://github.com/malfet	2025-05-30 11:25:30 +00:00
Laith Sakka	9d6f0d5991	avoid sym_max on nested int in is_contiguous. (#154633 ) calling is_contiguous will fail due to sym_max not being supported for nested int, this address in a way consistent with make_contiguous_strides_for Pull Request resolved: https://github.com/pytorch/pytorch/pull/154633 Approved by: https://github.com/bobrenjc93	2025-05-30 09:59:33 +00:00
Zhang, Jianyi	3c05167489	[Intel GPU] fix matmul accuracy when offset > 0 (#154495 ) This pr will make matmul tensors contiguous if they are not 64 byte alignment. oneDNN requires a minimal alignment of 64 https://uxlfoundation.github.io/oneDNN/dev_guide_c_and_cpp_apis.html#intel-r-processor-graphics-and-xe-architecture-graphics Fixes https://github.com/intel/torch-xpu-ops/issues/1656 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154495 Approved by: https://github.com/liangan1, https://github.com/guangyey, https://github.com/EikanWang	2025-05-30 09:53:51 +00:00
Paul Zhang	aec3ef1008	[Inductor] Add NaN assert to returned values from generated code (#154455 ) Summary: It is possible to have `reinterpret_tensor` in the output of inductor codegen, e.g. `reinterpret_tensor(buf366, (1024, ), (1, ), 0)` in the return tuple. This adds assertions to all return values from inductor codegen to prevent nans from slipping through and being hard to trace. Test Plan: NaN asserts properly generated in example gemm script: vars = (buf1, primals_2, buf2, primals_1, ) for var in vars: if isinstance(var, torch.Tensor): assert not var.isnan().any().item() assert not var.isinf().any().item() Pull Request resolved: https://github.com/pytorch/pytorch/pull/154455 Approved by: https://github.com/eellison	2025-05-30 08:53:24 +00:00
Bob Ren	dc82e911e7	remove allow-untyped-defs from torch/utils/data/datapipes/iter/filelister.py (#154624 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154624 Approved by: https://github.com/Skylion007	2025-05-30 08:38:05 +00:00
PyTorch MergeBot	639f459cb6	Revert "[Inductor] Add NaN assert to returned values from generated code (#154455 )" This reverts commit c3de2c7c6bc865b9fabd2db8f2af6383936aa653. Reverted https://github.com/pytorch/pytorch/pull/154455 on behalf of https://github.com/huydhn due to Sorry for reverting your change, I am trying to see if it help fix the broken trunk below. It it does not help, I will reland the PR ([comment](https://github.com/pytorch/pytorch/pull/154455#issuecomment-2921562089))	2025-05-30 08:11:22 +00:00
Simon Fan	f889dea97d	[internal] Expose additional metadata to compilation callbacks (#153596 ) These hooks are used by internal stuck job detection to associate compilation events with the compile lease. Previously, we only had events for Dynamo and Inductor compilation. And recently, the callback handler was updated to ignore nested events. So the Inductor event was only really used by lazy backward. Here, I remove the inductor event, and add an explicit lazy backward one. Additionally, I add other runtime compilation events: autotuning and cudagraphs. I also expose the CompileId as a string to avoid imports, this will let internal UIs track each graph's contribution to the timeout. ```python class CallbackTrigger(enum.Enum): # most common case, dynamo attempts to trace a new frame DYNAMO = 1 # backward compilation can be deferred to runtime LAZY_BACKWARD = 2 # some backends autotune at runtime TRITON_AUTOTUNING = 3 # cudagraphs record at runtime CUDAGRAPH_RECORDING = 4 ``` Differential Revision: [D75092426](https://our.internmc.facebook.com/intern/diff/D75092426) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153596 Approved by: https://github.com/masnesral	2025-05-30 08:07:04 +00:00
drisspg	208965a9d6	Fix unbackend symint error (#154672 ) ## Summary Me and @laithsakka spoke offline about this one, TLDR is that we wanted this ![image](https://github.com/user-attachments/assets/2e537612-3261-4fbe-a6b9-f8ff92ba3c37) to also be true for Inductor. In that vein we added two new apis to size-vars which is `guard_or_false`, or `guard_or_true` with the semantics: guard_or_false, guard_or_true: Those APIs may add guards, but will never fail with data-dependent errors; They will try to evaluate the expression with the possibility of adding guards, if that fails due to data dependency, instead of hard failing. False or True are returned. When to use this? Performance optimizations that warrant a recompilation. Take the general path and add a runtime check. ``` # Consider this branching. if x==0: return 1 else return 10 # To make data dependent friendly, it can be written as the following: if guard_or_false(x==0): return 1 else torch.check(x!=0) # runtime check return 10 ``` However there is still 1 more api to add to make this example work which is the torch.check which works with expressions, I will leave that to the @laithsakka Pull Request resolved: https://github.com/pytorch/pytorch/pull/154672 Approved by: https://github.com/laithsakka	2025-05-30 07:45:01 +00:00
Bob Ren	5a7442b91f	remove allow-untyped-defs from torch/distributed/checkpoint/resharding.py (#154626 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154626 Approved by: https://github.com/Skylion007	2025-05-30 07:43:04 +00:00
Bob Ren	d66a55def0	remove allow-untyped-defs from torch/distributed/elastic/utils/logging.py (#154625 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154625 Approved by: https://github.com/Skylion007	2025-05-30 07:37:56 +00:00
Bob Ren	382b38ed1b	remove allow-untyped-defs from torch/nn/utils/_expanded_weights/conv_expanded_weights.py (#154623 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154623 Approved by: https://github.com/Skylion007	2025-05-30 07:32:57 +00:00
baodii	bcbd2a22b2	[Intel GPU] OneDNN primitive cache support for Int4 WOQ gemm on XPU (#147693 ) * add onednn primitive cache for int4 gemm for xpu Pull Request resolved: https://github.com/pytorch/pytorch/pull/147693 Approved by: https://github.com/EikanWang, https://github.com/liangan1, https://github.com/guangyey, https://github.com/ZhiweiYan-96 Co-authored-by: Yan, Zhiwei <zhiwei.yan@intel.com> Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-05-30 07:26:36 +00:00
Bob Ren	0df96e3921	remove allow-untyped-defs from torch/ao/quantization/stubs.py (#154622 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154622 Approved by: https://github.com/Skylion007	2025-05-30 07:26:09 +00:00
Xuanteng Huang	30f7079c93	[FSDP2] allow different dtypes for no grad model params (#154103 ) Fixes #154082 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154103 Approved by: https://github.com/weifengpy	2025-05-30 07:00:54 +00:00
PyTorch MergeBot	d173ba5a75	Revert "Remove MemPoolContext (#154042 )" This reverts commit 3b38989b5f8f918cf1ad38bdade059608544af4b. Reverted https://github.com/pytorch/pytorch/pull/154042 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/154042#issuecomment-2921401100))	2025-05-30 06:53:37 +00:00
Ivan Zaitsev	0fdd568b78	[forward fix] add support for MemoryFormat after type tightening (#154658 ) Summary: fixes error: ``` raise AssertionError(f"Unexpected type in c_type_for_prim_type: {type_=}") AssertionError: Unexpected type in c_type_for_prim_type: type_=MemoryFormat ``` after https://github.com/pytorch/pytorch/pull/154371 \| D75568111 Test Plan: ``` buck test 'fbcode//mode/opt' fbcode//deeplearning/aot_inductor/test:test_custom_ops -- --exact 'deeplearning/aot_inductor/test:test_custom_ops - test_export_extern_fallback_nodes (deeplearning.aot_inductor.test.test_custom_ops.TestAOTInductorProxyExecutor)' ``` Differential Revision: D75617432 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154658 Approved by: https://github.com/Camyll, https://github.com/atalman, https://github.com/malfet	2025-05-30 06:53:25 +00:00
Henry Tsang	a4b0023f3b	[cutlass backend] Cache config generation locally and remotely (#154686 ) Summary: Trying to cache the json list of configs. There are probably some more work: * preset * filelock (?) * for cases where we generate from scratch, save it to local as well (?) Test Plan: tested offline Reviewed By: coconutruben Differential Revision: D75334439 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154686 Approved by: https://github.com/coconutruben, https://github.com/ColinPeppler	2025-05-30 05:40:46 +00:00
PyTorch MergeBot	ba51f4876d	Revert "Enable C++ dynamic shape guards by default (#140756 )" This reverts commit dc0f09a4785349fc3b4e4d3dc3c02b018e5a0534. Reverted https://github.com/pytorch/pytorch/pull/140756 on behalf of https://github.com/izaitsevfb due to seem to break dynamo tests ([comment](https://github.com/pytorch/pytorch/pull/140756#issuecomment-2921151663))	2025-05-30 03:52:02 +00:00
PyTorch MergeBot	852b99eba0	Revert "[c10d] Separate monitoring thread into a class in PGNCCL (#153977 )" This reverts commit 0db9c64d68dcdf25210357c4f7a41618441091d4. Reverted https://github.com/pytorch/pytorch/pull/153977 on behalf of https://github.com/izaitsevfb due to breaks lots of jobs internally, safer to revert, see D75628917 ([comment](https://github.com/pytorch/pytorch/pull/153977#issuecomment-2921146129))	2025-05-30 03:46:43 +00:00
Bob Ren	20ee5f9044	remove allow-untyped-defs from elastic_distributed_sampler.py (#154620 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154620 Approved by: https://github.com/Skylion007	2025-05-30 03:29:45 +00:00
bobrenjc93	9c06dff1ce	[multigraph] use specializations in compile_and_call_fx_graph (#153449 ) The goal of this multigraph work is to enable a compiled region that has a single dynamo trace but multiple backend specializations. This work was inspired by vLLM which does this in a somewhat hacky way where they use a custom backend to capture a dynamo graph and then manually invoke compile_fx multiple times to get specialized graphs. There's really two parts of this work: The frontend changes: 1) we introduce an optional kwarg `specialize_on` to mark_{dynamic,unbacked} that takes in a list of specializations. I debated other methods including specifying specializations via decorators, but ultimately decided this approach was more harmonious. The big issue with decorators is the difficulty of composing well with the rest of the torch.compile ecosystem including graph breaks, lazy initialization of variable trackers and symbolic variables, etc. The backend changes (this PR): 1) We capture the backend_specialization specified in the mark_{dynamic,unbacked} API into a SymbolicContext. See changes in `/_dynamo/variables/builder.py` 2) After we are done dynamo tracing, we will lazily (more on this later) invoke `call_user_compiler` up to N + 1 times for N specializations and 1 generic graph. Under the hood this will call compile_fx, which composes nicely with both Async Compile and AOTAutogradCache. We do this by using a context manager to patch in specialization specific axioms into the ShapeEnv before invoking the user compiler. 3) When we have specializations, we install a lazy specialized dispatch function that checks each specialization and dispatches to the first one that matches. Instead of doing all of the specialization compiles up front, we do the compiles lazily. The first time a specialization is invoked, we will do the compilation and save it in a cache so subsequent invocations are fast. If none of the specializations match, we dispatch to the generic graph. I decided to do this over returning N different GuardedCodes since 1) it doesn't pollute the dynamo cache (eg. if you have 8 specializations, you would hit the cache limit) 2) it naturally incorporates the hierarchical lattice structure of the guards since the specializations are always necessarily stricter than the generic region's guards. I benchmarked this PR stack with #152596 and found around a 50% reduction when dispatching to the specialized regions: ![495269647_576053105510082_9189856138964956774_n](https://github.com/user-attachments/assets/66030fed-d62e-4d87-940f-aa13c99b1a73) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153449 Approved by: https://github.com/zou3519 ghstack dependencies: #153433	2025-05-30 03:19:49 +00:00
Paul Zhang	c3de2c7c6b	[Inductor] Add NaN assert to returned values from generated code (#154455 ) Summary: It is possible to have `reinterpret_tensor` in the output of inductor codegen, e.g. `reinterpret_tensor(buf366, (1024, ), (1, ), 0)` in the return tuple. This adds assertions to all return values from inductor codegen to prevent nans from slipping through and being hard to trace. Test Plan: NaN asserts properly generated in example gemm script: vars = (buf1, primals_2, buf2, primals_1, ) for var in vars: if isinstance(var, torch.Tensor): assert not var.isnan().any().item() assert not var.isinf().any().item() Differential Revision: D74691131 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154455 Approved by: https://github.com/eellison	2025-05-30 03:09:37 +00:00
Dylan Maloy	4a302b5731	NativeRT readme (#154581 ) Summary: att Test Plan: ci Differential Revision: D75557667 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154581 Approved by: https://github.com/Skylion007, https://github.com/zhxchen17, https://github.com/yiming0416	2025-05-30 02:50:53 +00:00
Yu, Guangye	adfd5b293a	Enhance UT on elapsed_time for XPUEvent (#154494 ) # Motivation UT enhancement to avoid the incorrect elapsed time return by xpu's Event. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154494 Approved by: https://github.com/EikanWang	2025-05-30 02:00:02 +00:00
Yiming Zhou	0289313551	[AOTI] Support OptionalTensor return type in AOTI proxy executor (#154286 ) Summary: When a C++ custom op returns an uninitialized tensor, it will be marked as None in Python. For this scenario, the user should mark the possibly uninitialized return as Tensor? in the custom op schema. This diff adds `as_optional_tensor` type to export schema and the support for optional tensor in AOTI proxy executor. Test Plan: ``` buck2 run mode/dev-nosan caffe2/test/inductor:test_aot_inductor_custom_ops -- -r test_fn_with_optional_tensor_output ``` Differential Revision: D75262529 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154286 Approved by: https://github.com/desertfire	2025-05-30 01:53:00 +00:00
Pian Pawakapan	58ead04ee9	[dynamic shapes] unbacked safe unsqueeze (#154087 ) Also ran into this working on https://github.com/SWivid/F5-TTS Pull Request resolved: https://github.com/pytorch/pytorch/pull/154087 Approved by: https://github.com/laithsakka	2025-05-30 01:41:57 +00:00
bobrenjc93	172015fc11	[multigraph] add specialize_on kwarg to mark_{dynamic,unbacked} (#153433 ) The goal of this multigraph work is to enable a compiled region that has a single dynamo trace but multiple backend specializations. This work was inspired by vLLM which does this in a somewhat hacky way where they use a custom backend to capture a dynamo graph and then manually invoke compile_fx multiple times to get specialized graphs. There's really two parts of this work: The frontend changes (this PR): 1) we introduce an optional kwarg `specialize_on` to mark_{dynamic,unbacked} that takes in a list of specializations. I debated other methods including specifying specializations via decorators, but ultimately decided this approach was more harmonious. The big issue with decorators is the difficulty of composing well with the rest of the torch.compile ecosystem including graph breaks, lazy initialization of variable trackers and symbolic variables, etc. The backend changes: 1) We capture the backend_specialization specified in the mark_{dynamic,unbacked} API into a SymbolicContext. See changes in `/_dynamo/variables/builder.py` 2) After we are done dynamo tracing, we will lazily (more on this later) invoke `call_user_compiler` up to N + 1 times for N specializations and 1 generic graph. Under the hood this will call compile_fx, which composes nicely with both Async Compile and AOTAutogradCache. We do this by using a context manager to patch in specialization specific axioms into the ShapeEnv before invoking the user compiler. 3) When we have specializations, we install a lazy specialized dispatch function that checks each specialization and dispatches to the first one that matches. Instead of doing all of the specialization compiles up front, we do the compiles lazily. The first time a specialization is invoked, we will do the compilation and save it in a cache so subsequent invocations are fast. If none of the specializations match, we dispatch to the generic graph. I decided to do this over returning N different GuardedCodes since 1) it doesn't pollute the dynamo cache (eg. if you have 8 specializations, you would hit the cache limit) 2) it naturally incorporates the hierarchical lattice structure of the guards since the specializations are always necessarily stricter than the generic region's guards. I benchmarked this PR stack with #152596 and found around a 50% reduction when dispatching to the specialized regions: ![495269647_576053105510082_9189856138964956774_n](https://github.com/user-attachments/assets/66030fed-d62e-4d87-940f-aa13c99b1a73) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153433 Approved by: https://github.com/zou3519	2025-05-30 01:08:15 +00:00
Chen Lai	9371491529	[Reland][pytorch] Patch the _is_conv_node function (#154473 ) Summary: Add the conv padding ops in pytorch, the corresponding pr in torch ao is https://github.com/pytorch/ao/pull/2257 Test Plan: ``` buck test 'fbcode//mode/opt' fbcode//caffe2/test:quantization_pt2e -- --exact 'caffe2/test:quantization_pt2e - test_conv_padding_bn_relu (quantization.pt2e.test_quantize_pt2e.TestQuantizePT2E)' ``` Differential Revision: D75494468 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154473 Approved by: https://github.com/Skylion007	2025-05-30 00:41:03 +00:00
Nikita Shulga	d6cb0fe576	[MPS] Extend index_copy support to complex dtypes (#154671 ) Should have noticed it during the review Pull Request resolved: https://github.com/pytorch/pytorch/pull/154671 Approved by: https://github.com/dcci ghstack dependencies: #154670	2025-05-30 00:28:13 +00:00
Nikita Shulga	0134150ebb	[MPS][BE] Do not copy sizes/strides unnecesserily (#154670 ) Just pass them as args to `mtl_setArgs`, metaprogramming should deal with the rest Also use `mtl_dispatch1DJob` instead of computing max threadgroup size by nand Pull Request resolved: https://github.com/pytorch/pytorch/pull/154670 Approved by: https://github.com/dcci	2025-05-30 00:28:13 +00:00
Ke Wen	61bfb3df9f	[a2av] Improve tuning for 4 GPUs (#154580 ) ### Problem Running `nvshmem_all_to_all_vdev` on 4 x H100s (fully connected with NVSwitch). Before: ``` Bytes: MiB, Time: us, BusBw: GB/s 0 32.29 16.23 1 33.01 31.76 2 33.01 63.54 4 33.83 123.97 8 49.83 168.34 16 80.82 207.59 32 178.66 187.82 64 335.79 199.86 128 646.72 207.54 256 1268.77 211.57 512 2511.14 213.80 1024 4998.31 214.82 2048 9964.49 215.51 4096 19892.34 215.91 ``` 215 GB/s does not reach the SOL of NV18 (350-400 GB/s). ### Change If the number of peers decreases (say 8 to 4), we do not reduce the number of CTAs; instead, we shift more CTAs towards the data parallel dimension. After: ``` Bytes: MiB, Time: us, BusBw: GB/s 0 25.01 20.96 1 25.70 40.80 2 25.76 81.42 4 28.87 145.26 8 40.79 205.64 16 61.46 272.97 32 111.82 300.06 64 202.40 331.57 128 382.56 350.84 256 739.11 363.19 512 1450.79 370.05 1024 2873.13 373.72 2048 5719.50 375.47 4096 11395.65 376.90 ``` If we look at MoE related region, say 32 MB, we can see a 187 -> 300 GB/s improvement. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154580 Approved by: https://github.com/ngimel	2025-05-30 00:26:13 +00:00
Brian Hirsh	2c1cb38d95	inductor codecache: include private inductor configs in cache key (#153672 ) Fixes https://github.com/pytorch/torchtitan/issues/1185 It looks like inductor's logic to include inductor configs in the cache key skips configs with a leading underscore by default. This came up in torchtitan - there's an asyncTP pipelining pass in inductor gated by a private config, and by not caching on the config we were attempting to use asyncTP when we shouldn't be. I'm not sure how worried we should be on the blast radius of this change. On the one hand: (1) it technically fixes any silent correctness issues in the cache around any other private inductor configs (it looks like there are a few) (2) there is some risk that there are some "harmless" configs that we are now including in the key, which may increase false negatives. I do see that there is an explicit list for "configs we want to ignore for caching" (`_save_config_ignore`), so my hope is that all harmless configs are already encapsulated there. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153672 Approved by: https://github.com/oulgen ghstack dependencies: #153766	2025-05-30 00:24:29 +00:00
Brian Hirsh	5b6fd277f9	convert inductor codecache to use getArtifactLogger (#153766 ) I'm not entirely sure of the background for why inductor codecache code uses default python logging instead of the new TORCH_LOGS-based artifact logging, but switching it over to artifact logging makes it easier to use nice testing utils in the next PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153766 Approved by: https://github.com/oulgen, https://github.com/Skylion007	2025-05-30 00:24:29 +00:00
eqy	818f76a745	[cuDNN] Allow cudnn attention or flash attention in `test_export.py` regex (#154458 ) Analogous to #153272 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154458 Approved by: https://github.com/drisspg	2025-05-29 23:51:09 +00:00
Isuru Fernando	dc0f09a478	Enable C++ dynamic shape guards by default (#140756 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/140756 Approved by: https://github.com/anijain2305, https://github.com/laithsakka ghstack dependencies: #151225	2025-05-29 23:44:43 +00:00
Paul Zhang	0c6c7780d9	[Inductor] Add envvar to disable decomposeK (#154421 ) Summary: Add envvar to Inductor config to disable decomposeK autotuning choice Test Plan: `buck test 'fbcode//mode/opt' fbcode//caffe2/test/inductor:max_autotune -- --exact 'caffe2/test/inductor:max_autotune - test_max_autotune_decompose_k_dynamic_False_sizes2 (caffe2.test.inductor.test_max_autotune.TestMaxAutotune)' --run-disabled` Reviewed By: eellison Differential Revision: D75174823 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154421 Approved by: https://github.com/eellison	2025-05-29 23:34:41 +00:00
Isuru Fernando	9ba67e99bb	[dynamo] keep C++ symbolic shape guards disabled for benchmarks (#151225 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151225 Approved by: https://github.com/anijain2305	2025-05-29 23:29:39 +00:00
Jerry Mannil	d5e0704247	[ROCm] Update maxpool launch config (#154619 ) * Better perf on MI300 with updated launch configs Pull Request resolved: https://github.com/pytorch/pytorch/pull/154619 Approved by: https://github.com/jeffdaily	2025-05-29 23:28:07 +00:00
Joel Schlosser	43b18d098b	Forward fix for test_frame_traced_hook in internal testing (#154641 ) Summary: Fixes the newly-added dynamo test test_frame_traced_hook so it can run internally Test Plan: This is a test change Differential Revision: D75616787 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154641 Approved by: https://github.com/Skylion007	2025-05-29 23:02:01 +00:00
soulitzer	b040d63ce4	Prevent SAC cache from being kept alive by reference cycle (#154651 ) Fixes https://github.com/pytorch/pytorch/issues/154642 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154651 Approved by: https://github.com/xmfan	2025-05-29 22:27:35 +00:00
Aaron Gokaslan	7d17253af8	[BE]: Improve aten formatter with fmtlib (#152830 ) Replaces stateful ostream output with stateless fmtlib, which is signficantly faster and more contained. It is especially faster for the type of complex double formatting found here since it uses the newer [DragonBox algorithm](https://github.com/jk-jeon/dragonbox) for faster floating point formatting (which is the main bottleneck here). This also enables some static time checking of the formatting strings test plan: all tests pass Pull Request resolved: https://github.com/pytorch/pytorch/pull/152830 Approved by: https://github.com/cyyever, https://github.com/malfet, https://github.com/atalman	2025-05-29 22:11:30 +00:00
Paul Zhang	fdbf314278	[Inductor] Cache subgraph autotuning choices properly (#154067 ) Differential Revision: D75170507 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154067 Approved by: https://github.com/eellison	2025-05-29 22:01:44 +00:00
Gabriel Ferns	c7e8e8ee19	Add torch.profile benchmarking function to feedback_fns (#153579 ) Summary: Updates some benchmarking code to have the option to use torch.profile, and passes in a thunk to benchmark_fns to get this information (this will be a different result from `timings`, which are already passed into those functions). Test Plan: Existing unit tests. Differential Revision: D74444990 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153579 Approved by: https://github.com/coconutruben, https://github.com/masnesral, https://github.com/nmacchioni	2025-05-29 21:43:45 +00:00
Arash Pakbin	1237f271aa	[ROCm] MIOpen: Get current device from Torch rather than HIP in handle creation (#154549 ) Get current device from Torch rather than HIP in MIOpen handle creation. The device may have already been set from torch side, otherwise device is set to 0 for handle. Additional audits of cudnn vs miopen Handle.cpp file. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154549 Approved by: https://github.com/jeffdaily, https://github.com/cyyever Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-05-29 21:12:12 +00:00
Arash Pakbin	08fdc64c86	[ROCm] Exposing Some MIOpen Symbols (#2176 ) (#154545 ) This PR exposes some MIOpen symbols, namely: 1. `miopenDataType_t getMiopenDataType(const at::Tensor& tensor)` 2. `miopenHandle_t getMiopenHandle()` 3. `class TensorDescriptor` 4. `class Descriptor` 5. `class FilterDescriptor` 6. `struct ConvolutionDescriptor` 7. `struct DropoutDescriptor` 8. `struct RNNDescriptor` to enable adding extensions that make use of them. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154545 Approved by: https://github.com/jeffdaily, https://github.com/Skylion007 Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-05-29 21:10:45 +00:00
Shivam Raikundalia	83a0e4e6f9	[Visualizer] Start at index with most events (#154571 ) Summary: Oftentimes a single snapshot will contain multiple GPU traces in it based on what the process can see. In this case lets just start with the gpu trace with the highest amount of activity Test Plan: Ran od with: https://www.35929.od.internalfb.com/pytorch_memory_visualizer/mvai_gpu_traces/tree/gpu_snapshot/fire-chujiechen-f701302011/1/rank-1_itrn-3.Mar_01_06_10_09.3747.snapshot.pickle And it started at index 1 instead of 0 Differential Revision: D75555558 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154571 Approved by: https://github.com/aaronenyeshi	2025-05-29 20:49:33 +00:00
Feny Patel	2bc8fec744	deprecate MTIA_WORKLOADD from pytorch (#154627 ) Differential Revision: D75612179 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154627 Approved by: https://github.com/sraikund16	2025-05-29 20:30:40 +00:00
Joaquin	cb56df55dc	[Inductor]Cleanup autotune_fallback_to_aten post-deprecation (#154331 ) Fixes #153298 This PR is the 3rd and final step of #147479 All references to autotune_fallback_to_aten have been removed, and the feature is now deprecated. All calls to should_fallback_to_aten() were also removed, as they were deemed unnecessary. [henrylhtsang](https://github.com/henrylhtsang) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154331 Approved by: https://github.com/henrylhtsang	2025-05-29 20:29:58 +00:00
Huy Do	629fca295e	Always set CPU affinity for benchmark jobs (#154569 ) Because metrics like compilation time requires CPU. I want to see if this help fix https://github.com/pytorch/pytorch/issues/152566 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154569 Approved by: https://github.com/malfet, https://github.com/desertfire	2025-05-29 20:11:47 +00:00
atalman	3afbab66f7	[BE] Remove unused release scripts. Add clarifications for the branch cut process (#154649 ) Scripts in ``scripts/release/promote/`` are not used for a while. We use the ones in test-infra [here](https://github.com/pytorch/test-infra/blob/main/release/) . Hence this small cleanup. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154649 Approved by: https://github.com/Skylion007, https://github.com/huydhn	2025-05-29 19:49:37 +00:00
Shiyan Deng	e8f5c24d17	[rocm]add device guard when initialize single stream (#154433 ) Summary: AMD streams are lazily initialized and sometimes (e.g. when we just want to do event recording on the stream) we might not be setting the device guard while it's initializing which would lead to invalid configuration error. Differential Revision: D75456460 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154433 Approved by: https://github.com/jeffdaily	2025-05-29 19:42:12 +00:00
Jeddie Ji	20ec61a02f	[BE] fix lint errors caused by const SROpFunctor fn (#154552 ) Summary: Remove const quaiflier from SR suggsted from CLANGTIDY. Test Plan: arc lint -a -e extra --take CLANGTIDY caffe2/torch/fb/sparsenn/cpu_operators/to_dense_representation_cpu.cpp Reviewed By: henryoier Differential Revision: D75534056 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154552 Approved by: https://github.com/Skylion007	2025-05-29 19:40:08 +00:00
Bin Bao	5a21d6f982	[AOTI][reland] Support multi-arch when using package_cpp_only (#154608 ) Summary: Reland https://github.com/pytorch/pytorch/pull/154414 Add support of multi_arch_kernel_binary in the package_cpp_only mode. More specifically, generate specific cmake targets to compile .ptx to .fatbin and embed them in the final shared library or binary. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154608 Approved by: https://github.com/yushangdi	2025-05-29 19:32:33 +00:00
fduwjj	0db9c64d68	[c10d] Separate monitoring thread into a class in PGNCCL (#153977 ) This is the start of a series of efforts to consolidating auxiliary threads in PGNCCL, aka watchdog and heartbeat_monitoring threads. Right now we launch these two threads per PG instances, i.e., if users create hundred or thousand instances of PG or subPGs, we will end up with that twice many side threads which is not efficient. We have a RFC to consolidate them (https://github.com/pytorch/pytorch/issues/146956). Right now both threads are assigned with so many functionalities so it is hard to do the consolidations in one shot, we will try to split it into at least two steps (PRs) to make it easier to test and review. We did our first attemp in https://github.com/pytorch/pytorch/pull/153668 but we also want to try to see if we can make monitoring thread a class. This PR is doing the first step to make monitoring thread a class. The next step to also extract watchdog to be a separate class so that we know its dependency. What we did in this PR: 1. Move all related variables and methods into a class named `HeartbeatMonitor`. 2. Correct some errors in the original logics inside monitoring thread loop. 3. Move the error propagation check to watchdog thread which is more relevant. This is totally fine since we rolled out EventCache out fully so watchdog hang is rare now. Today there are two major functions inside heartbeat monitoring thread today: 1. Check the heartbeat of watchdog thread every 8 minutes. If no heartbeat detected and we are sure monitoring thread has not been stopped, we will kill the program by SIG_ABORT. 2. We check TCPStore every 30 sec to see if any watchdog timeout happens on other ranks, if so we will initiate a dump signal on the current rank as well. (We do this only in the default PG) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153977 Approved by: https://github.com/kwen2501, https://github.com/d4l3k	2025-05-29 17:45:04 +00:00
Nicolas Macchioni	6f992e1b3f	[BE][AT] cleanup my old todo (#154542 ) Summary: this todo is very old, and probably not needed anymore. let's have CI figure out if removing this breaks anything Test Plan: CI Differential Revision: D75491068 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154542 Approved by: https://github.com/Skylion007	2025-05-29 17:22:01 +00:00
Nikita Shulga	634ce22601	[MPSInductor] Fix codegen for nested multistage reductions (#154578 ) Yet to write a unittest for it, but this fixes codegen for ``` python3 benchmarks/dynamo/torchbench.py --performance --only hf_T5 --backend inductor --inference --devices mps --float16 ``` By correctly closing triple nested loop Pull Request resolved: https://github.com/pytorch/pytorch/pull/154578 Approved by: https://github.com/jansel, https://github.com/dcci	2025-05-29 17:09:25 +00:00
henrylhtsang	8883e494b3	[cutlass backend][ez] remove indent for cutlass config serialization (#154573 ) Differential Revision: [D75566642](https://our.internmc.facebook.com/intern/diff/D75566642) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154573 Approved by: https://github.com/ColinPeppler	2025-05-29 17:00:52 +00:00
Isalia20	41092cb86c	[MPS] index copy impl (#154326 ) Second most requested op according to #154052 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154326 Approved by: https://github.com/malfet	2025-05-29 16:57:43 +00:00
soulitzer	733e684b11	Skip test file that doesn't run gradcheck for slow gradcheck (#154509 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154509 Approved by: https://github.com/malfet	2025-05-29 16:32:26 +00:00
Bo Li	2c6f24c62d	[ROCm] Updated default workspace for gfx95 (#153988 ) Fixes test_cuda.py::test_cublas_workspace_explicit_allocation on gfx95 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153988 Approved by: https://github.com/jeffdaily	2025-05-29 16:22:17 +00:00
PyTorch MergeBot	53b0f6f543	Revert "Use 3.27 as the minimum CMake version (#153153 )" This reverts commit 4613081b729273a9273185e9ef7470ce76e22da2. Reverted https://github.com/pytorch/pytorch/pull/153153 on behalf of https://github.com/malfet due to It broke windows debug builds, see `ef1d45b12d/1` ([comment](https://github.com/pytorch/pytorch/pull/153153#issuecomment-2919897160))	2025-05-29 16:14:28 +00:00
eellison	ef1d45b12d	Cleanup parent fallback logic (#154006 ) The `parent` in fallback_node_due_to_unsupported_type is a duplication of `unsupported_output_tensor` logic. remove it. tested that the tests in test_add_complex give same codegen. this fixes an issue in mx that @drisspg was running into. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154006 Approved by: https://github.com/drisspg	2025-05-29 13:40:36 +00:00
eellison	d6e29bf875	Reflect back mutation if we clone misaligned tensors (#154442 ) Fix for https://github.com/pytorch/pytorch/issues/152425 inductor specializes whether or not a tensor is 16-bit aligned on the first invocation. then, on subsequent invocations, if we inferred alignment but are passed a non-aligned tensor we clone the tensor. If we infer alignment, then run with unaligned, and mutate the input, we need to reflect back the mutation to the input. This pr adds back that mutation. We could have also been less aggressive about inferring alignment for mutated tensors, but that has a pretty perf hit.See the following benchmark: ``` import torch t = torch.rand(4096 * 4096, device="cuda", dtype=torch.float16) @torch.compile(dynamic=False) def foo(x): return x.add_(1) import triton print(triton.testing.do_bench(lambda: foo(t[:-1]))) torch._dynamo.reset() print(triton.testing.do_bench(lambda: foo(t[1:]))) ``` gives ``` 0.04063070610165596 0.07613472988113162 ``` So almost twice as slow for non-aligned tensors. Tensors changing alignment is a relatively rare case. In the future, we could considering a multi-kernel approach, or codegening a triton kernel that does most of the loads with aligned instructions, and a prologue/epilogue of un-alignment. But, it's yet to be seen this is a huge issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154442 Approved by: https://github.com/bobrenjc93, https://github.com/bdhirsh	2025-05-29 13:36:48 +00:00
Yu, Guangye	3c74a72ea0	Keep XPU compatible with toolchain 2025.2 (#154359 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154359 Approved by: https://github.com/EikanWang, https://github.com/cyyever	2025-05-29 11:12:07 +00:00
Laith Sakka	cd9ff41282	check fallback_value first. (#154493 ) This is just a refactor, not a fix for any issue. we do check fallback_value first and early exit instead of checking it not set over and over. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154493 Approved by: https://github.com/bobrenjc93	2025-05-29 09:06:43 +00:00
Janet Yang	447b481c79	[AOTI] Save data sizes to constants_info (#154534 ) Differential Revision: D75223179 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154534 Approved by: https://github.com/muchulee8	2025-05-29 06:39:13 +00:00
Rachel Guo	9c7ed3e46e	[debug_printer][BE] Fix float8 type printing for min/max value printing (#154466 ) Summary: ATT GH Issue: https://github.com/pytorch/pytorch/issues/149008 Previous: Failed to use debug printing for float8 types due to the limitation of "min_all_cuda" implementation from aten native: `4b39832412/aten/src/ATen/native/cuda/ReduceMinValuesKernel.cu (L51)` Error: Min value: Error: "min_all_cuda" not implemented for 'Float8_e4m3fn' Now: Example output paste: P1824621233 Unblocked float8 type tensor debug printing. Suggest to print the whole value if numel <= threshold. Test Plan: ``` AOT_INDUCTOR_DEBUG_INTERMEDIATE_VALUE_PRINTER=2 TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 TORCHINDUCTOR_ABI_COMPATIBLE=1 TORCH_C OMPILE_DEBUG=1 TORCH_LOGS="+inductor, +schedule, output_code" buck2 run -c fbcode.enable_gpu_sections=true -c fbcode.nvcc_arch=h100 @//mode/opt fbcode//caffe2/test/inductor:test_aot_inductor -- -r test_aoti_debug_printer_float8_dtype_cuda ``` ``` AOT_INDUCTOR_DEBUG_INTERMEDIATE_VALUE_PRINTER=2 TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 TORCHINDUCTOR_ABI_COMPATIBLE=1 TORCH_COMPILE_DEBUG=1 TORCH_LOGS="+inductor, +schedule, output_code" buck2 run -c fbcode.enable_gpu_sections=true -c fbcode.nvcc_arch=h100 @//mode/opt fbcode//caffe2/test/inductor:test_aot_inductor -- -r test_fp8_cuda 2>&1 \| tee fp8_example_printing.txt ``` Differential Revision: D74847967 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154466 Approved by: https://github.com/jingsh, https://github.com/henrylhtsang	2025-05-29 05:48:02 +00:00
henrylhtsang	07343efc15	[cutlass backend] small refactor to flatten the ops to avoid nested for loops (#154576 ) Differential Revision: [D75565429](https://our.internmc.facebook.com/intern/diff/D75565429) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154576 Approved by: https://github.com/ColinPeppler	2025-05-29 04:42:58 +00:00
jianan-gu	b394c6e89c	[Inductor][CPP] Add block sparse for FlexAttention CPU (#147196 ) ## Overview This PR is to optimize FlexAttention CPP template with block sparse. Block sparse is natively supported in FlexAttention block mask structures, thus following logic of the kv blocks from `kv_indice ` and `full_kv_indice ` is the strightforward way to add this optimization. Pull Request resolved: https://github.com/pytorch/pytorch/pull/147196 Approved by: https://github.com/drisspg, https://github.com/leslie-fang-intel	2025-05-29 02:57:02 +00:00
eellison	c0864bb389	Add a (t * 0) pattern (#153161 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153161 Approved by: https://github.com/danielvegamyhre	2025-05-29 02:19:36 +00:00
Aaron Gokaslan	316e7a9293	[BE][Ez]: Denote common types as TypeAlias (#154527 ) Denotes common_types as TypeAlias. This triggered a Ruff rule since we named our TypeAlias off standards so I added a file wide ruff suppression Pull Request resolved: https://github.com/pytorch/pytorch/pull/154527 Approved by: https://github.com/benjaminglass1, https://github.com/aorenste	2025-05-29 02:00:13 +00:00
Jerry Mannil	2d932a2e01	[ROCm] Fix 3D tensor perf degradation with NHWC format (#154522 ) Co-author: @doru1004 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154522 Approved by: https://github.com/jeffdaily	2025-05-29 01:33:49 +00:00
cyy	4613081b72	Use 3.27 as the minimum CMake version (#153153 ) Update the minimum CMake version to 3.27 because of it provides more CUDA targets such as `CUDA::nvperf_host` so that it is possible to remove some of our forked CUDA modules. See https://github.com/pytorch/pytorch/pull/153783. It's also possible to facilitate future third-party updates such as FBGEMM (its current shipped version requires 3.21). Pull Request resolved: https://github.com/pytorch/pytorch/pull/153153 Approved by: https://github.com/malfet	2025-05-29 00:52:44 +00:00
Aaron Orenstein	946a4c2bdc	BE: Type previously untyped decorators (#154515 ) Summary: Cloned #153726 from Skylion007 and fixed internal typing issues. Test Plan: Unit tests pass Differential Revision: D75477355 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154515 Approved by: https://github.com/Skylion007	2025-05-29 00:36:34 +00:00
Menglu Yu	ba0a91b3ea	[4/n][Optimus][Auto-AC] Expose the config to skip the dynamo gaurds to avoid recompile (#154152 ) Summary: context: https://fb.workplace.com/groups/1075192433118967/permalink/1673720956599442/ Thanks Microve for raising the existing dynamo skip API in D75196435 The dynamic shape triggers recompilation, introducing compilation time increase, we expose config that users can skip the dynamo guards to avoid the recompile. Note that it may quantize unnessarily nodes, which can impact NE, QPS and memory saving, needs verification. Differential Revision: D75248430 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154152 Approved by: https://github.com/bobrenjc93	2025-05-29 00:35:37 +00:00
Natalia Gimelshein	22a1b3b5d0	use 4 elements per thread in no-cast elementwise kernel (#154558 ) Reduce elems per thread to 4 in vectorized function also (only for unaligned inputs where there's no vectorization anyway). This slightly reduces binary size (by 4MB) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154558 Approved by: https://github.com/malfet	2025-05-29 00:32:44 +00:00
nirajkamalk	40abb2b403	Fix deprecated amp APIs in docs (#154553 ) Update usage of deprecated amp APIs. Fixes https://github.com/pytorch/tutorials/issues/3331 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154553 Approved by: https://github.com/Skylion007	2025-05-29 00:05:59 +00:00
Michael Lazos	2b3ac17aa2	[Cutlass] Remove spammy log for gemm extensions (#154548 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154548 Approved by: https://github.com/henrylhtsang	2025-05-28 23:55:36 +00:00
William Wen	81b7c96697	[dynamo, nested graph breaks] add skip_frame debugging function (#153773 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153773 Approved by: https://github.com/jansel ghstack dependencies: #151056, #153510, #153772	2025-05-28 23:29:37 +00:00
William Wen	6cda280483	[dynamo, nested graph breaks] remove block stack graph break in output_graph (#153772 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153772 Approved by: https://github.com/jansel ghstack dependencies: #151056, #153510	2025-05-28 23:29:37 +00:00
William Wen	bbd45f1f1f	[dynamo, nested graph breaks] refactor codegen to minimize NULL codegen'ing (#153510 ) Stop codegening NULLs that we need to pop later. Some output_graph.py changes to prepare for nested graph break support. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153510 Approved by: https://github.com/jansel ghstack dependencies: #151056	2025-05-28 23:29:37 +00:00
William Wen	0f0d5749a0	[dynamo, nested graph breaks] small fixes to resume function generation (#151056 ) Old: ~pack resume function stack + locals into a list: we need to be able to pass frame stack+locals in lists to hand off to nested functions in the future, so we implement this part first.~ We are no longer doing this right now since GraphModule/guard variable naming gets messed up. Going forward, our approach will be to keep the top frame unpacked, but pack the rest of the contents of other frames in a list. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151056 Approved by: https://github.com/jansel	2025-05-28 23:29:37 +00:00
Benjamin Glass	65b1aedd09	[Inductor] Improve typing, and prepare for ABI-compatible AOTI C-shim dispatching (#154371 ) Prepares for the next PR in the stack by tightening up typing on a `cpp_wrapper` interface that's only used in one (well-typed) place, as well as downstream effects of that change. In particular, this enabled: 1. removing a number of now clearly unnecessary asserts 2. adding a few more targeted asserts to validate the code's current assumptions 3. removing some unneeded control flow in several functions As far as I can tell, this PR should be functionally neutral. One argument was removed from a `cpp_wrapper` public API, but that argument was unused, and only had a single callsite. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154371 Approved by: https://github.com/desertfire	2025-05-28 23:25:17 +00:00
Shangdi Yu	3e05a48927	Fix clamp type promotion in inductor decomposition (#154471 ) Summary: as title, the clamp type promotion should take min/max arg into consideration as well. Test Plan: ``` buck run fbcode//caffe2/test/inductor:test_aot_inductor -- -r test_clamp_decomposition_cpu python test/inductor/test_torchinductor.py -k test_clamp -v ``` Differential Revision: D75490124 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154471 Approved by: https://github.com/desertfire, https://github.com/chenyang78	2025-05-28 23:24:25 +00:00
bobrenjc93	d865b784e4	Support unbacked whitelist (#154295 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154295 Approved by: https://github.com/angelayi	2025-05-28 23:01:22 +00:00
mieshkiwrk	ef4d57329b	[CAG] Support for call_module at copy paste aot bwd graph (#153827 ) Support for `call_module` in `copy_paste_aot_backward_graph` added recently with PT2.7 Problem is being observed with HPU backend in example repro due to creating fused modules. ``` import torch device = 'cpu' #'hpu' backend = 'inductor' #'hpu_backend' def fn(t1): t1 = t1 * 1 t1_grad = torch.ones_like(t1, device=device) t1.backward(t1_grad, retain_graph=True) return t1 t1 = torch.ones(1, requires_grad=True, device=device) #.squeeze() compiled_fn = torch.compile(fn, backend=backend) result = compiled_fn(t1) with torch._dynamo.compiled_autograd._enable(torch.compile(backend=backend)): result_grad = torch.ones_like(result, device=device) result.backward(result_grad) print(f'{result_grad=}') print(f'{t1.grad=}') ``` With this change I'm getting same results like on CPU, however I'm facing below problem when running with scalar (t1 tensor after squeeze): `torch._dynamo.exc.TorchRuntimeError: Dynamo failed to run FX node with fake tensors: call_function <built-in function getitem>((FakeTensor(..., device='hpu:0', size=()), 0), *{}): got IndexError('invalid index of a 0-dim tensor. Use `tensor.item()` in Python or `tensor.item<T>()` in C++ to convert a 0-dim tensor to a number')` While on CPU there's following warning and None returned: `repro.py:23: UserWarning: The .grad attribute of a Tensor that is not a leaf Tensor is being accessed. Its .grad attribute won't be populated during autograd.backward(). If you indeed want the .grad field to be populated for a non-leaf Tensor, use .retain_grad() on the non-leaf Tensor. If you access the non-leaf Tensor by mistake, make sure you access the leaf Tensor instead. See github.com/pytorch/pytorch/pull/30531 for more informations. (Triggered internally at pytorch/build/aten/src/ATen/core/TensorBody.h:489.) print(f'{t1.grad=}') t1.grad=None` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153827 Approved by: https://github.com/xmfan	2025-05-28 22:52:40 +00:00
bobrenjc93	d62a33c002	[ez] add docblock for _expandsums (#154397 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154397 Approved by: https://github.com/laithsakka ghstack dependencies: #154400, #154398, #154396, #154399	2025-05-28 22:43:26 +00:00
bobrenjc93	0c00e32632	[ez] add docblock for _eval_is_non_overlapping_and_dense (#154399 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154399 Approved by: https://github.com/laithsakka ghstack dependencies: #154400, #154398, #154396	2025-05-28 22:40:03 +00:00
Zhengxu Chen	0f56318152	[precompile] Add Exception type PackageError for unsupported precompile features. (#154430 ) Summary: Today when guard serialization fails, dynamo will raise an internal error like: ``` torch._dynamo.exc.InternalTorchDynamoError: RuntimeError: CLOSURE_MATCH guard cannot be serialized. ``` Adding a dedicated PackageError type to surface the error more clearly. Test Plan: CI Differential Revision: D75452124 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154430 Approved by: https://github.com/jamesjwu, https://github.com/jansel	2025-05-28 22:34:51 +00:00
Serhat Gundem	11129d9317	Add new ops in fallback ops (#154251 ) Fixes #ISSUE_NUMBER ## Background Task: [T222738229](https://www.internalfb.com/intern/tasks/?t=222738229) It's the first starter task on the project _Enabling TorchNative Standalone on Whisper_. We are using cshim to create a layer of abstraction between _libtorch_ and _AOTInductor generated artifacts_. So we needed to add an entry in the cshim for every API surface in libtorch. And we only care about operators that AOTInductor does not handle. And for this task, we only wanted to add it for the following ops. ## What I've done? 4 new fallback ops are added that show up in the Whisper model. (torchgen/aoti/fallback_ops.py) - aten.permute (default) - aten.squueze (dim) - aten.abs (default) - aten.hann_window (default) Then I ran the below command to generate new header C shim header files. As it says [here](`7e86a7c015/torchgen/gen.py (L2424-L2436%20for%20details)`) `python torchgen/gen.py --update-aoti-c-shim` Then, `python setup.py develop` to rebuild PyTorch ## Testing Also 4 new tests have been added on test/inductor/test_aot_inductor.py - test_proxy_executor_permute - test_proxy_executor_abs - test_proxy_executor_squeeze - test_proxy_executor_hann I ran these commands to test it (inside local pytorch root folder): `python test/inductor/test_aot_inductor.py -k test_proxy_executor_permute` `python test/inductor/test_aot_inductor.py -k test_proxy_executor_abs` `python test/inductor/test_aot_inductor.py -k test_proxy_executor_squeeze` `python test/inductor/test_aot_inductor.py -k test_proxy_executor_hann` ## NOTE: I didn't see any order between the tests inside _test/inductor/test_aot_inductor.py_. That's why, I added new tests just after the test given in the example. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154251 Approved by: https://github.com/angelayi	2025-05-28 22:11:07 +00:00
Simon Fan	d2f506cae8	[ca] disable ca for functorch grad and run all HOO tests (#154147 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154147 Approved by: https://github.com/zou3519 ghstack dependencies: #154133	2025-05-28 22:06:13 +00:00
Simon Fan	857f21631d	[ca] fix hop_db tests (#154133 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154133 Approved by: https://github.com/zou3519	2025-05-28 22:06:13 +00:00
bobrenjc93	ed348e7026	Add docblock for TrackedFake (#154396 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154396 Approved by: https://github.com/laithsakka ghstack dependencies: #154400, #154398	2025-05-28 21:19:49 +00:00
bobrenjc93	d311b79c12	add docblock for _fast_expand (#154398 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154398 Approved by: https://github.com/laithsakka ghstack dependencies: #154400	2025-05-28 21:16:47 +00:00
bobrenjc93	e7318b863d	[ez] add docblock to cast_symbool_to_symint_guardless (#154400 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154400 Approved by: https://github.com/laithsakka	2025-05-28 21:11:53 +00:00
Zizeng Meng	f6dcc45c44	[Kineto x Insight] Add device to activity type map in pytorch (#154253 ) Summary: Update the device to ActivityType Map in pytorch. Need to be exported to github Test Plan: Run the ondemand e2e test and insight profiler is triggered during profiling P1819539581: https://www.internalfb.com/intern/paste/P1819539581/ {F1978519960} Insight profiler is not enabled when mtia_insight not specifying in config {F1978527200} Reviewed By: fenypatel99 Differential Revision: D75246621 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154253 Approved by: https://github.com/Skylion007	2025-05-28 20:36:19 +00:00
Ke Wen	e25074d462	[c10d][CI] Change expected return code in Sandcastle for Nan tests (#154441 ) Fixing internal error caused by #153167. `skip_but_pass_in_sandcastle_if` returns exit code 0. But `test_nan_assert` expects exit code -6. So we'd need to set expected return code conditional on `IS_SANDCASTLE`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154441 Approved by: https://github.com/fduwjj, https://github.com/nWEIdia ghstack dependencies: #153167	2025-05-28 20:35:52 +00:00
Huy Do	c381103fd7	Fix the logic of set_cpu_affinity (#154503 ) While investigating https://github.com/pytorch/pytorch/issues/152566, I found two issues with how the cpu affinity is set in benchmark job: * The current logic doesn't work with cgroups slice, the mechanism behind multi-tenant runner: * Using `lscpu` returns all CPUs and not the available ones from cgroups. On the other hand, `nproc` works correctly. For example, on H100, `lscpu` returns 192 CPUs while `nproc` returns 24 (192 / 8) * Setting `taskset -c 0-N` blindly is wrong because CPU 0 is only available to the the first tenant, aka alice. For example, running `taskset -c 0 ls` on any other tenants will fail. To fix this, the ID of available CPUs can be fetched by calling `os.sched_getaffinity(0)`. * The last bug is `taskset` works with logical CPUs https://www.man7.org/linux/man-pages/man1/taskset.1.html, so using the result from `test_inductor_get_core_number` is also wrong because that function returns the number of physical CPUs. ### Testing CPU benchmark jobs look ok * [aarch64 torch.compile benchmark](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Wed%2C%2021%20May%202025%2016%3A40%3A28%20GMT&stopTime=Wed%2C%2028%20May%202025%2016%3A40%3A28%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cpu%20(aarch64)&lBranch=fix-cpu-affinity-cgroups&lCommit=9a6288e083d650c470623f5fe136b1060824021c&rBranch=main&rCommit=dec5ab8d984b8a608140911351d877b9ddb141c2) * [x86 micro benchmark](https://hud.pytorch.org/benchmark/llms?startTime=Wed%2C%2021%20May%202025%2016%3A41%3A26%20GMT&stopTime=Wed%2C%2028%20May%202025%2016%3A41%3A26%20GMT&granularity=day&lBranch=main&lCommit=c1b7dbc52aaa49f4cd147bbe5935110a4a10e3e3&rBranch=refs/tags/ciflow/inductor-micro-benchmark-cpu-x86/154503&rCommit=9a6288e083d650c470623f5fe136b1060824021c&repoName=pytorch%2Fpytorch&benchmarkName=&modelName=All%20Models&backendName=All%20Backends&modeName=All%20Modes&dtypeName=All%20DType&deviceName=cpu%20(x86_64)&archName=All%20Platforms) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154503 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-05-28 19:38:20 +00:00
dolpm	66f53889d5	[nativert] port semaphore to c10 util (#153504 ) Summary: nativert RFC: https://github.com/zhxchen17/rfcs/blob/master/RFC-0043-torch-native-runtime.md To land the runtime into PyTorch core, we will gradually land logical parts of the code into the Github issue and get each piece properly reviewed. This diff adds a simple semaphore interface into c10 until c++20 where we get counting_semaphore gonna need a oss build export to take a look at this... Test Plan: CI Differential Revision: D73882656 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153504 Approved by: https://github.com/zhxchen17	2025-05-28 19:17:30 +00:00
Jithun Nair	24980d2641	[ROCm][CI] Update build-environment for mi300 workflows (#153134 ) so their test times are tracked separately in https://raw.githubusercontent.com/pytorch/test-infra/generated-stats/stats/test-times.json. Currently, both MI200 and MI300 test times get combined into the same key `linux-focal-rocm-py3.10` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153134 Approved by: https://github.com/huydhn	2025-05-28 19:04:53 +00:00
PyTorch MergeBot	d4ab8e74f3	Revert "Fix the Problems About Defining Static Variable in Inline Function (#147095 )" This reverts commit c6fc11af760d4ad1f01cc699a3c6488ab5f41770. Reverted https://github.com/pytorch/pytorch/pull/147095 on behalf of https://github.com/izaitsevfb due to still fails to link internally at meta ([comment](https://github.com/pytorch/pytorch/pull/147095#issuecomment-2917221575))	2025-05-28 18:22:39 +00:00
henrylhtsang	1c7a70b483	[AOTI][cutlass backend] Do not remove the cutlass kernel .o file after packaging (#154155 ) Differential Revision: [D75253009](https://our.internmc.facebook.com/intern/diff/D75253009/) In general, we want to cache the cutlass kernels. Also saw an error saying .o not found. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154155 Approved by: https://github.com/chenyang78	2025-05-28 17:35:19 +00:00
Laith Sakka	66ac724b56	pyfmt lint torch/_export/passes/replace_view_ops_with_view_copy_ops_pass.py (#154488 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154488 Approved by: https://github.com/Skylion007 ghstack dependencies: #154483, #154484, #154485, #154487	2025-05-28 17:07:15 +00:00
Laith Sakka	dfe0f48123	pyfmt lint torch/_export/serde/schema.py (#154487 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154487 Approved by: https://github.com/Skylion007 ghstack dependencies: #154483, #154484, #154485	2025-05-28 17:07:15 +00:00
Laith Sakka	92cebed1bd	pyfmt lint torch/_export/serde/serialize.py (#154485 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154485 Approved by: https://github.com/Skylion007 ghstack dependencies: #154483, #154484	2025-05-28 17:07:07 +00:00
Laith Sakka	b4fe5ca58a	pymft lint torch/utils/weak.py (#154484 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154484 Approved by: https://github.com/Skylion007 ghstack dependencies: #154483	2025-05-28 17:06:58 +00:00
Laith Sakka	4de1b25df7	Remove empty files from execlude lint rule (#154483 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154483 Approved by: https://github.com/Skylion007	2025-05-28 17:06:50 +00:00
Sidharth	70539308ac	[dynamo] updating gb_type names for uniqueness (#154452 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154452 Approved by: https://github.com/williamwen42	2025-05-28 16:54:10 +00:00
Isalia20	e313152a33	SDPA fix memory efficient attention for large batch dim (#154029 ) Fixes #146704 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154029 Approved by: https://github.com/ngimel	2025-05-28 16:53:53 +00:00
Natalia Gimelshein	3b38989b5f	Remove MemPoolContext (#154042 ) Removes MemPoolContext from custom user mempools. The ground truth for which pool should be used is in graph_pools active pool, and MemPoolContext just introduced an opportunity for the pool pointed to by MemPoolContext and active pool in graph_pools to go out of sync (see all the asserts in the code to make sure that happens, and yet it still could happen in a multithread scenario, see my recent PRs (#153990). Pull Request resolved: https://github.com/pytorch/pytorch/pull/154042 Approved by: https://github.com/albanD, https://github.com/syed-ahmed	2025-05-28 16:35:48 +00:00
Jerry Zhang	d23aa7e182	Add deprecation warning for `torch.ao.quantization` (#153892 ) Summary: att Test Plan: (ao) $ PYTHONWARNINGS='default' python Python 3.10.14 \| packaged by conda-forge \| (main, Mar 20 2024, 12:45:18) [GCC 12.3.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> from torch.ao.quantization.quantizer.xnnpack_quantizer import XNNPACKQuantizer printing warning /anaconda3/envs/ao/lib/python3.10/site-packages/torch/ao/quantization/__init__.py:36: DeprecationWarning: torch.ao.quantization is deprecated. Plan is to 1. Remove eager mode quantization (torch.ao.quantization.quantize, torch.ao.quantization.quantize_dynamic), please migrate to use torchao eager mode quantize_ API instead 2. Remove fx graph mode quantization (torch.ao.quantization.quantize_fx.prepare_fx, torch.ao.quantization.quantize_fx.convert_fx, please migrate to use torchao pt2e quantization API instead (prepare_pt2e, convert_pt2e) 3. pt2e quantization has been migrated to torchao (https://github.com/pytorch/ao/tree/main/torchao/quantization/pt2e) see https://dev-discuss.pytorch.org/t/torch-ao-quantization-migration-plan/2810 for more details warnings.warn( >>> a = XNNPACKQuantizer() /anaconda3/envs/ao/lib/python3.10/site-packages/torch/ao/quantization/quantizer/xnnpack_quantizer.py:281: DeprecationWarning: XNNPACKQuantizer is deprecated! Please use xnnpack quantizer in ExecuTorch (https://github.com/pytorch/executorch/tree/main/backends/xnnpack/quantizer) instead warnings.warn(f"{self.__class__.__name__} is deprecated! Please use xnnpack quantizer in ExecuTorch (https://github.com/pytorch/executorch/tree/main/backends/xnnpack/quantizer) instead", DeprecationWarning) >>> Reviewers: Subscribers: Tasks: Tags: Pull Request resolved: https://github.com/pytorch/pytorch/pull/153892 Approved by: https://github.com/Skylion007	2025-05-28 16:25:30 +00:00
Zhengxu Chen	5bf74753f6	[precompile] Prune local scope variables for guard serialization. (#154431 ) Summary: Prune unused local objects from serialized local scope if they are not used in guard reconstruction. This is helpful when a user program takes things like local callable functions or the function call is recursive. Test Plan: test/dynamo/test_guard_serialization.py -k test_function_locals Before pruning locals: ``` state = GuardsState(output_graph=OutputGraphGuardsState(local_scope={'x': tensor([ 0.0461, 0.4024, -1.0115]), 'g': <function ...aints=None, _guards=<torch._guards.GuardsSet object at 0x7fbccc7e9fc0>, _aotautograd_guards=[]), shape_code_parts=None) def pickle_guards_state(state: GuardsState) -> bytes: buf = io.BytesIO() pickler = GuardsStatePickler(buf) try: pickler.dump(state) except AttributeError as e: > raise torch._dynamo.exc.PackageError(str(e)) from e E torch._dynamo.exc.PackageError: Can't pickle local object 'TestGuardSerialization.test_function_locals.<locals>.foo' ``` After the diff ``` Tests finished: Pass 1. Fail 0. Fatal 0. Skip 0. Build failure 0 ``` Differential Revision: D75452123 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154431 Approved by: https://github.com/jansel	2025-05-28 16:03:02 +00:00
Joel Schlosser	9db7bcb3fe	[Dynamo] Introduce hook receiving list of traced code objects (#153622 ) This PR: * Expands `Hooks` with a new, optional `frame_traced_fn` field. It should be a callable receiving the list of traced code objects * Maintains a list of `traced_code` objects in the `TracingContext` of an `OutputGraph` * Whenever an `inline_call()` is encountered, the corresponding code object is added to this set * `OutputGraph`'s associated `f_code` is added to the list just before the hook is called I believe use of this hook should enable the source code hashing that vLLM does in a better way than monkey-patching `inline_call()`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153622 Approved by: https://github.com/jansel	2025-05-28 15:40:09 +00:00
bobrenjc93	476e0a643a	[ez] add docblock for ShapeGuardPythonPrinter (#154403 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154403 Approved by: https://github.com/jingsh ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377, #154378, #154379, #154380, #154381, #154383, #154384, #154385, #154402	2025-05-28 14:17:17 +00:00
bobrenjc93	473a93eb58	[ez] add docblock for _ShapeGuardPrinter (#154402 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154402 Approved by: https://github.com/jingsh ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377, #154378, #154379, #154380, #154381, #154383, #154384, #154385	2025-05-28 14:13:22 +00:00
bobrenjc93	35a473e364	[ez] add docblock for guard_scalar (#154385 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154385 Approved by: https://github.com/jingsh ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377, #154378, #154379, #154380, #154381, #154383, #154384	2025-05-28 14:10:07 +00:00
bobrenjc93	ee4f433963	[ez] add docblock for _guard_or (#154384 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154384 Approved by: https://github.com/pianpwk ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377, #154378, #154379, #154380, #154381, #154383	2025-05-28 14:06:29 +00:00
bobrenjc93	e9b97d19b1	[ez] Make SymNodeImpl comments less misleading (#154480 ) As discussed in DS workchat, it's easy for users to get confused by guarding for these supposedly non-guarding methods. The TL;DR is in the case of non pythonic compilers like XLA, we actually do guard. I've updated the comments accordingly to reduce confusion. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154480 Approved by: https://github.com/pianpwk, https://github.com/Skylion007	2025-05-28 14:04:32 +00:00
PyTorch MergeBot	a75e3a02be	Revert "[dynamo, nested graph breaks] small fixes to resume function generation (#151056 )" This reverts commit 28e7aa21c522e92ea01a62dfdc5e3b74e398d8f0. Reverted https://github.com/pytorch/pytorch/pull/151056 on behalf of https://github.com/malfet due to Not sure which one, but it broke test_error_messages, see `203b0efd63/1` ([comment](https://github.com/pytorch/pytorch/pull/151056#issuecomment-2916437433))	2025-05-28 13:53:50 +00:00
PyTorch MergeBot	9603d6382d	Revert "[dynamo, nested graph breaks] refactor codegen to minimize NULL codegen'ing (#153510 )" This reverts commit 1fe98429222a8ba5e16dd9381f50a8fb90edcf0e. Reverted https://github.com/pytorch/pytorch/pull/153510 on behalf of https://github.com/malfet due to Not sure which one, but it broke test_error_messages, see `203b0efd63/1` ([comment](https://github.com/pytorch/pytorch/pull/151056#issuecomment-2916437433))	2025-05-28 13:53:50 +00:00
PyTorch MergeBot	5fd7004dc9	Revert "[dynamo, nested graph breaks] remove block stack graph break in output_graph (#153772 )" This reverts commit 9a66c30bdc563c62375e5030c4103b67515b8dac. Reverted https://github.com/pytorch/pytorch/pull/153772 on behalf of https://github.com/malfet due to Not sure which one, but it broke test_error_messages, see `203b0efd63/1` ([comment](https://github.com/pytorch/pytorch/pull/151056#issuecomment-2916437433))	2025-05-28 13:53:50 +00:00
PyTorch MergeBot	e86439ed5b	Revert "[dynamo, nested graph breaks] add skip_frame debugging function (#153773 )" This reverts commit aadf9eae63c4793e1107a3b21ede30e5289eeaca. Reverted https://github.com/pytorch/pytorch/pull/153773 on behalf of https://github.com/malfet due to Not sure which one, but it broke test_error_messages, see `203b0efd63/1` ([comment](https://github.com/pytorch/pytorch/pull/151056#issuecomment-2916437433))	2025-05-28 13:53:50 +00:00
Howard Huang	203b0efd63	[PP] Allow unused kwargs in ZB path (#153498 ) This is a fix when an unused kwarg is in the PP stage forward, we try to call `torch.autograd.grad()` and update its gradients when it shouldn't have gradients. Leading to this error: ``` [rank3]:[rank3]: File "/data/users/howardhuang/pytorch/torch/distributed/pipelining/stage.py", line 613, in [rank3]:[rank3]: return lambda: stage_backward_input( [rank3]:[rank3]: File "/data/users/howardhuang/pytorch/torch/distributed/pipelining/_backward.py", line 199, in stage_backward_input [rank3]:[rank3]: dinputs = torch.autograd.grad( [rank3]:[rank3]: File "/data/users/howardhuang/pytorch/torch/autograd/init.py", line 503, in grad [rank3]:[rank3]: result = _engine_run_backward( [rank3]:[rank3]: File "/data/users/howardhuang/pytorch/torch/autograd/graph.py", line 824, in _engine_run_backward [rank3]:[rank3]: return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass [rank3]:[rank3]: RuntimeError: One of the differentiated Tensors does not require grad ``` related issues: https://github.com/pytorch/torchtitan/issues/1188 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153498 Approved by: https://github.com/kwen2501	2025-05-28 13:34:04 +00:00
ILCSFNO	cf7451f279	Fix signature of torch.sparse_coo_tensor() (#152681 ) Fixes #145371 @pearu Searched all and find these codes, wondering whether is the root cause of the issue, could you have a review? Thanks a lot! Pull Request resolved: https://github.com/pytorch/pytorch/pull/152681 Approved by: https://github.com/Skylion007, https://github.com/pearu, https://github.com/nikitaved	2025-05-28 13:16:41 +00:00
Yuanhao Ji	f58143b945	[Typing] Refactor `torch.types.Device` in `torch/cuda/__init__.py` (#153447 ) Part of: #152952 Follow up: #153027 Here is the definition of `torch.types.Device`: `ab997d9ff5/torch/types.py (L74)` So `Optional[Union[Device, int]]` is equivalent to `torch.types.Device`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153447 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-05-28 10:09:31 +00:00
PyTorch MergeBot	fdc339003b	Revert "[AOTI] Support multi-arch when using package_cpp_only (#154414 )" This reverts commit a84d8c4a1cc515db274366537afd0b1492800c2d. Reverted https://github.com/pytorch/pytorch/pull/154414 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it is failing ROCm trunk job ([comment](https://github.com/pytorch/pytorch/pull/154414#issuecomment-2915597821))	2025-05-28 09:23:31 +00:00
Laith Sakka	853958f82c	Fix: Replacements can cause runtime assertions to disappear and can cause invalid inductor code. (#153661 ) Lets explore firs a couple of problem related to replacements and runtime assertions. #### example problem 1 if we have a runtime assertions that u0==s0, u0 is an input coming from mark_unbacked. A replacement u0=s0 will be added, the function f(u0, s0) will become f(s0, s0), this leads to the assert not being inserted during insert_deferred_runtime_asserts. The reason is that insert_deferred_runtime_asserts logic insert each assertion once all its inputs are seen, but u0 will never be seen. Same thing can happen when we defer assertion on backed i.e: s0==s2 ..etc. #### example problem 2 Consider u0==s0, where u0 is coming from a call to .item() Imagine later on that a specialization happens to s0 to become 2. In that case s0 as input wont be seen during insert_deferred_runtime_asserts and the assertion won't be inserted in the graph. Worse, Inductor will generate some code that refers to s0 in the cpp wrapper while it does not exist, causing a failure. internal xref: https://fb.workplace.com/groups/1075192433118967/permalink/1669766396994898/ ## The solution : Runtime assertions insertion loops depend on detecting that the symbols that are used in the runtime assertions are seen, note that those symbols are either graph inputs or generated in the graph from data dependent ops like .item(). The issues above happen when symbols are graph inputs, in order to force the symbols to exist in the graph and to be seen by the runtime assertions we do not do replacements on placeholders expressions during codegen and during runtime assertions insertion. This should not have performance overhead, since we already optimized the graph with replacements, the only effect is not mistakenly dropping graph inputs that are used in runtime assertions. I added extended testing. A solo unrelated follow up that I noticed, is that we might want to rename unbacked symbols in runtime assertions when we do unbacked renaming, but that's a different issue. Other approaches that did not work : #### ban replacements on unbacked. 1. does not work when we defer runtime assertions on backed ex: s0==s1. we could also ban such replacements but problem 2 becomes more problematic. 2. Problem two, it affects the quality of reasoning ! in a bad way. #### Apply specialization on runtime assertions before codegen . 1. Can fix some issues, but may lead also to runtime assertions becoming NOPs. 2. Does not fix the issue if not inserting runtime assertions during insert_deferred_runtime_asserts due to input not being detected. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153661 Approved by: https://github.com/jansel	2025-05-28 09:08:05 +00:00
William Wen	aadf9eae63	[dynamo, nested graph breaks] add skip_frame debugging function (#153773 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153773 Approved by: https://github.com/jansel ghstack dependencies: #151056, #153510, #153772	2025-05-28 08:54:09 +00:00
William Wen	9a66c30bdc	[dynamo, nested graph breaks] remove block stack graph break in output_graph (#153772 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153772 Approved by: https://github.com/jansel ghstack dependencies: #151056, #153510	2025-05-28 08:54:09 +00:00
William Wen	1fe9842922	[dynamo, nested graph breaks] refactor codegen to minimize NULL codegen'ing (#153510 ) Stop codegening NULLs that we need to pop later. Some output_graph.py changes to prepare for nested graph break support. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153510 Approved by: https://github.com/jansel ghstack dependencies: #151056	2025-05-28 08:54:09 +00:00
William Wen	28e7aa21c5	[dynamo, nested graph breaks] small fixes to resume function generation (#151056 ) Old: ~pack resume function stack + locals into a list: we need to be able to pass frame stack+locals in lists to hand off to nested functions in the future, so we implement this part first.~ We are no longer doing this right now since GraphModule/guard variable naming gets messed up. Going forward, our approach will be to keep the top frame unpacked, but pack the rest of the contents of other frames in a list. Pull Request resolved: https://github.com/pytorch/pytorch/pull/151056 Approved by: https://github.com/jansel	2025-05-28 08:54:09 +00:00
cyy	9d04c0f352	Remove outdated CUDA 11 conditions (#154313 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/154313 Approved by: https://github.com/eqy	2025-05-28 08:44:58 +00:00
Pian Pawakapan	1d9b7dd2d1	[PGO] suggest dynamic whitelist for recompilations (#154189 ) suggests `TORCH_COMPILE_DYNAMIC_SOURCES` based off tensor size changes in PGO code state, including parameters. Closing #153442 which took the dynamo guards approach. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154189 Approved by: https://github.com/bobrenjc93	2025-05-28 07:11:43 +00:00
bobrenjc93	fe760b6636	[ez] add docblock for _free_unbacked_symbols_with_path (#154383 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154383 Approved by: https://github.com/pianpwk ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377, #154378, #154379, #154380, #154381	2025-05-28 05:53:50 +00:00
bobrenjc93	8e25ba6963	[ez] add docblock for find_symbol_binding_fx_nodes (#154381 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154381 Approved by: https://github.com/pianpwk ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377, #154378, #154379, #154380	2025-05-28 05:44:26 +00:00
bobrenjc93	08c29deb5f	[ez] add docblock to is_symbol_binding_fx_node (#154380 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154380 Approved by: https://github.com/pianpwk ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377, #154378, #154379	2025-05-28 05:41:19 +00:00
bobrenjc93	07405a6cff	[ez] add docblock for free_unbacked_symbols (#154379 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154379 Approved by: https://github.com/pianpwk ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377, #154378	2025-05-28 05:37:25 +00:00
bobrenjc93	dcdaef5206	[ez] add docblock for free_symbols (#154378 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154378 Approved by: https://github.com/pianpwk ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405, #154377	2025-05-28 05:34:25 +00:00
bobrenjc93	abc3fdc7ac	[ez] add docblock for _iterate_exprs (#154377 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154377 Approved by: https://github.com/pianpwk ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404, #154405	2025-05-28 05:28:58 +00:00
bobrenjc93	ab6cb85cb0	[ez] add docblock for _remove_effect_token_unbacked_bindings (#154405 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154405 Approved by: https://github.com/Skylion007, https://github.com/pianpwk ghstack dependencies: #154374, #154375, #154376, #154386, #154401, #154404	2025-05-28 05:16:14 +00:00
bobrenjc93	fde8f6a8b8	[ez] add docblock for _suggest_torch_checks (#154404 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154404 Approved by: https://github.com/Skylion007 ghstack dependencies: #154374, #154375, #154376, #154386, #154401	2025-05-28 04:45:55 +00:00
bobrenjc93	b82fb57b67	[ez] add docblock for RuntimeAssert (#154401 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154401 Approved by: https://github.com/Skylion007 ghstack dependencies: #154374, #154375, #154376, #154386	2025-05-28 04:43:22 +00:00
bobrenjc93	d64b4a91dd	[ez] remove unused function _constrain_symbol_range (#154386 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154386 Approved by: https://github.com/Skylion007 ghstack dependencies: #154374, #154375, #154376	2025-05-28 04:41:00 +00:00
Laith Sakka	ef90cc18d7	use definitely_contiguous for _prim_elementwise_meta short circuit (#153441 ) * This verifies that the check short circuit is not material. https://github.com/pytorch/pytorch/pull/153431 ``` import torch from torch.export import Dim, export class MyModel(torch.nn.Module): def forward(self, x, ranks): first_k = ranks.max().item() torch._check_is_size(first_k) narrow = x.narrow(dim = 1, start = 0, length = first_k) lt = narrow < narrow.size(1) return lt inps = ( torch.randn((8, 16), device="cuda"), torch.arange(8, device="cuda", dtype=torch.int8) ) spec = { "x": (Dim.AUTO, Dim.AUTO), "ranks": (Dim.AUTO,), } traced = export(MyModel(), inps, dynamic_shapes=spec, strict=True).run_decompositions({}) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153441 Approved by: https://github.com/jansel ghstack dependencies: #153432	2025-05-28 03:41:26 +00:00
Laith Sakka	39df901b2a	introduce definitely_contiguous and use it for reshape and tensor meta data computation. (#153432 ) when a tensor has unbacked symbols it can be general enough to represent both contiguous and non contiguous tensors. in that case we cant really evaluate is_contiguous. In many places in the code base, we check for is_contiguous to take a fast path. but the general path usually works for both contiguous and not contiguous in that case we probably want to use definitely _contiguous API. This is appleid for reshape in this PR and also to tensor meta data computation, the meta data now will have an attribute that says that its contiguous when its always contiguous. We would store that only if definitely _contiguous is true now. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153432 Approved by: https://github.com/bobrenjc93	2025-05-28 03:41:26 +00:00
Sidharth	54f1f29fed	[dynamo] dynamic gb_type -> static gb_type (#154435 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154435 Approved by: https://github.com/williamwen42	2025-05-28 03:14:26 +00:00
ZhiweiYan-96	f12ce4e36b	[Intel GPU] convolution fusion at XPU backend (#154202 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154202 Approved by: https://github.com/EikanWang, https://github.com/guangyey, https://github.com/etaf ghstack dependencies: #140365	2025-05-28 03:14:18 +00:00
FFFrog	c6fc11af76	Fix the Problems About Defining Static Variable in Inline Function (#147095 ) Refer to https://github.com/pytorch/pytorch/issues/125465 for more informations - Remove unused header files - Move the inline function that defines the static variable to .cc Pull Request resolved: https://github.com/pytorch/pytorch/pull/147095 Approved by: https://github.com/cyyever, https://github.com/albanD	2025-05-28 02:47:16 +00:00
bobrenjc93	855eff8e8e	Don't CSE unbacked nodes (#154387 ) * #154440 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154387 Approved by: https://github.com/TroyGarden ghstack dependencies: #154440	2025-05-28 02:21:56 +00:00
bobrenjc93	919a1a17e3	[ez] Replace misleading implementations with NYI (#154440 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154440 Approved by: https://github.com/Skylion007, https://github.com/pianpwk	2025-05-28 02:21:56 +00:00
Bin Bao	a84d8c4a1c	[AOTI] Support multi-arch when using package_cpp_only (#154414 ) Summary: Add support of multi_arch_kernel_binary in the package_cpp_only mode. More specifically, generate specific cmake targets to compile .ptx to .fatbin and embed them in the final shared library or binary. Differential Revision: [D75452096](https://our.internmc.facebook.com/intern/diff/D75452096) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154414 Approved by: https://github.com/angelayi ghstack dependencies: #154412, #154413	2025-05-28 01:20:38 +00:00
Bin Bao	cde82d25b7	[AOTI] Add a multi_arch_kernel_binary option (#154413 ) Summary: CUDA can support multi-arch with the fatbin format. Add this multi_arch_kernel_binary option, so the compiled model binary can run across different GPU archs. Differential Revision: [D75452094](https://our.internmc.facebook.com/intern/diff/D75452094) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154413 Approved by: https://github.com/angelayi ghstack dependencies: #154412	2025-05-28 01:20:38 +00:00
Bin Bao	4d8f3d537a	[AOTI][refactor] Rename embed_cubin to embed_kernel_binary (#154412 ) Summary: Rename as it is not CUDA specific. Differential Revision: [D75452095](https://our.internmc.facebook.com/intern/diff/D75452095) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154412 Approved by: https://github.com/angelayi	2025-05-28 01:20:28 +00:00
bobrenjc93	e79790e14b	[ez] add docblock for _sympy_from_args (#154376 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154376 Approved by: https://github.com/Skylion007 ghstack dependencies: #154374, #154375	2025-05-27 23:43:13 +00:00
atalman	fe082c5ffe	Move inductor workflows focal (ubuntu 20.04) -> jammy (ubuntu 22.04) (#154153 ) Trying to fix: https://github.com/pytorch/pytorch/issues/154157 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154153 Approved by: https://github.com/Skylion007, https://github.com/huydhn, https://github.com/nike4949, https://github.com/cyyever	2025-05-27 23:16:21 +00:00
iupaikov-amd	3f10c9d8af	Fixed an issue with XPU skip so the test_decompose_mem_bound_mm.py suite can be ran correctly (#153245 ) Fixes #153239 Replaced custom decorator with the common one. Although the better way to skip the whole suite would be to add it to skip list in run_test.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/153245 Approved by: https://github.com/jeffdaily	2025-05-27 23:10:25 +00:00
atalman	4b39832412	[CI] Update torchbench pin (#154453 ) Related to https://github.com/pytorch/pytorch/issues/154446 Pins torchbench repo to a https://github.com/pytorch/benchmark/pull/2620 which pins opacus to ``1.5.3`` version Pull Request resolved: https://github.com/pytorch/pytorch/pull/154453 Approved by: https://github.com/wdvr, https://github.com/malfet	2025-05-27 23:08:42 +00:00
atalman	247ea229ba	Create issue template: Release highlight for proposed Feature (#154125 ) Authors: @anitakat @atalman This is related to: https://github.com/pytorch/pytorch/issues/152134 . Adding RFC template for feature submissions Pull Request resolved: https://github.com/pytorch/pytorch/pull/154125 Approved by: https://github.com/anitakat, https://github.com/ZainRizvi, https://github.com/albanD	2025-05-27 22:45:21 +00:00
anwang	53affa273b	[MTIA Aten Backend][1.3/n] Migrate remaining view ops, which all need explicit register in `native_functions.yaml` (#154337 ) See context in D75266206. This diff/PR migrates all the remaining view ops, which all need changes in `native_functions.yaml` and thus need to be exported to PR. Ops covered by this diff: - _reshape_alias - unfold internal: Also delete the entire aten_mtia_view_ops.cpp file, and update corresponding build config. Differential Revision: [D75385411](https://our.internmc.facebook.com/intern/diff/D75385411/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154337 Approved by: https://github.com/nautsimon ghstack dependencies: #154336	2025-05-27 22:18:12 +00:00
Shangdi Yu	eaf355cb11	[BE] Clean up unused parameter input in AOTIModel (#154276 ) Summary: As title Test Plan: CI Differential Revision: D74691763 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154276 Approved by: https://github.com/Skylion007	2025-05-27 22:17:32 +00:00
PyTorch MergeBot	241f8dc84d	Revert "Remove outdated CUDA 11 conditions (#154313 )" This reverts commit 3936e6141c09dab94f21e4fdab7bea4bddf62ac2. Reverted https://github.com/pytorch/pytorch/pull/154313 on behalf of https://github.com/izaitsevfb due to breaks internal builds ([comment](https://github.com/pytorch/pytorch/pull/154313#issuecomment-2914230005))	2025-05-27 21:54:41 +00:00
Jerry Mannil	6be829535f	[ROCm] Improve vectorized elementwise kernel performance in MI300X (#153634 ) * Use non-temporal loads to improve the vectorized elementwise kernel performance on MI300 * Use thread_work_size of 8 or 16 for vectorized elementwise kernel Co-author: @amd-hhashemi Pull Request resolved: https://github.com/pytorch/pytorch/pull/153634 Approved by: https://github.com/jeffdaily	2025-05-27 20:49:32 +00:00
PyTorch MergeBot	555fc05868	Revert "[Inductor] Improve typing, and prepare for ABI-compatible AOTI C-shim dispatching (#154371 )" This reverts commit 6169ca0b65bcb382faa1a2287278b3717c18f127. Reverted https://github.com/pytorch/pytorch/pull/154371 on behalf of https://github.com/benjaminglass1 due to Appears to have broken main ([comment](https://github.com/pytorch/pytorch/pull/154371#issuecomment-2913975736))	2025-05-27 20:39:09 +00:00
Guilherme Leobas	7359705232	Add CPython tests for unittest (#150788 ) Tests: * test_assertions.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/150788 Approved by: https://github.com/williamwen42	2025-05-27 20:26:17 +00:00
Guilherme Leobas	12fc06d267	Add CPython complex tests (#152015 ) Tests: * test_complex.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/152015 Approved by: https://github.com/williamwen42	2025-05-27 20:24:28 +00:00
Guilherme Leobas	3b218e56dc	Add CPython tests for iter/sort (#150797 ) Tests: * test_iter.py * test_sort.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/150797 Approved by: https://github.com/williamwen42	2025-05-27 20:22:34 +00:00
bobrenjc93	4fd8a54a41	[ez] add docblock for is_accessor_node (#154375 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154375 Approved by: https://github.com/Skylion007, https://github.com/pianpwk ghstack dependencies: #154374	2025-05-27 19:47:32 +00:00
tvukovic-amd	b367e5f6a6	[ROCm][Windows] Fix building torch 2.8 wheel with ROCm (added hipblasLt and rocblas directories) (#153144 ) Since rocblas.dll and hipblaslt.dll are copied to torch/lib, rocblas and hipblaslt directories are needed to be stored there too (otherwise we have an error after wheel installation while searching for files in rocblas/library and hipblaslt/library which doesn't exist). This PR fixes this issue. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153144 Approved by: https://github.com/jeffdaily Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2025-05-27 19:40:28 +00:00
PyTorch MergeBot	fa6ca59079	Revert "Move inductor workflows focal (ubuntu 20.04) -> jammy (ubuntu 22.04) (#154153 )" This reverts commit 2bd95f3a1f07132aa00f5c438c5228866d7dd1f8. Reverted https://github.com/pytorch/pytorch/pull/154153 on behalf of https://github.com/malfet due to Broke inductor tests, see `b8452e55bc/1` ([comment](https://github.com/pytorch/pytorch/pull/154153#issuecomment-2913738047))	2025-05-27 19:23:28 +00:00
Benjamin Glass	6169ca0b65	[Inductor] Improve typing, and prepare for ABI-compatible AOTI C-shim dispatching (#154371 ) Prepares for the next PR in the stack by tightening up typing on a `cpp_wrapper` interface that's only used in one (well-typed) place, as well as downstream effects of that change. In particular, this enabled: 1. removing a number of now clearly unnecessary asserts 2. adding a few more targeted asserts to validate the code's current assumptions 3. removing some unneeded control flow in several functions As far as I can tell, this PR should be functionally neutral. One argument was removed from a `cpp_wrapper` public API, but that argument was unused, and only had a single callsite. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154371 Approved by: https://github.com/desertfire	2025-05-27 19:17:41 +00:00
Ryan Guo	75bbd4989c	[dynamo] Support using symint from dispatcher-style tensor subclass (#154130 ) Fixes #146932. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154130 Approved by: https://github.com/laithsakka	2025-05-27 19:05:46 +00:00
PyTorch MergeBot	8c0f07f944	Revert "[ROCm] Improve vectorized elementwise kernel performance in MI300X (#153634 )" This reverts commit 0d4de7872ac019abbd6e87b3391b2276d9d05bd4. Reverted https://github.com/pytorch/pytorch/pull/153634 on behalf of https://github.com/malfet due to Broke inductor jobs, see `b8452e55bc/1` ([comment](https://github.com/pytorch/pytorch/pull/153634#issuecomment-2913619071))	2025-05-27 19:02:59 +00:00
Zizeng Meng	b8452e55bc	[Kineto x Insight] Update Kineto submodule (#154426 ) Summary: We add a new ActivityType::MTIA_INSIGHT in `20f652846f` Test Plan: CI Differential Revision: D75454945 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154426 Approved by: https://github.com/Skylion007	2025-05-27 18:29:29 +00:00
Nikita Shulga	5075df6fee	Make torch importable if compiled without TensorPipe (#154382 ) By delaying the import/hiding it behind `torch.distributed.rpc.is_tensorpipe_avaiable()` check Fixes https://github.com/pytorch/pytorch/issues/154300 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154382 Approved by: https://github.com/Skylion007 ghstack dependencies: #154325	2025-05-27 18:13:38 +00:00
Nikita Shulga	f472ea63bb	[BE] Fix typos in SyntaxError description (#154436 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154436 Approved by: https://github.com/seemethere, https://github.com/wdvr, https://github.com/ZainRizvi	2025-05-27 18:08:58 +00:00
Cyrus Daruwala	cfbd99fdfd	[Pytorch] Add option to CPU Blas GEMM to avoid output downcast (#154012 ) Summary: Dot product for a single output element consists of 3 steps (both input vectors have elements of type scalar_t): 1. elementwise vector multiply (scalar_t x scalar_t -> opmath_t) 2. vector reduction to a scalar value (opmath_t -> opmath_t) 3. optional downcast if opmath_t != out_t The current blas kernel performs steps 1 and 2 correctly, but for step 3, it will always downcast to scalar_t even when opmath_t == output_t (and then do an upcast back to output_t), which results in precision loss. This diff fixes the precision loss in the BlasKernel Test Plan: Attention CI passes Differential Revision: D75023858 topic: not user facing Pull Request resolved: https://github.com/pytorch/pytorch/pull/154012 Approved by: https://github.com/Valentine233, https://github.com/aditew01, https://github.com/CaoE, https://github.com/drisspg	2025-05-27 17:43:21 +00:00
bobrenjc93	1ca082d9a1	[ez] Rewrite comment to be more friendly to non haskellers (#151421 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151421 Approved by: https://github.com/aorenste	2025-05-27 17:32:34 +00:00
bobrenjc93	70fbd5e08c	[ez] Add docblock for resolve_unbacked_bindings (#154374 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154374 Approved by: https://github.com/Skylion007, https://github.com/pianpwk	2025-05-27 17:05:49 +00:00
bobrenjc93	2560c1f3f0	add sticky cache pgo (#154418 ) It's a reland of https://github.com/pytorch/pytorch/pull/154394 that hit some mergebot bug Pull Request resolved: https://github.com/pytorch/pytorch/pull/154418 Approved by: https://github.com/malfet	2025-05-27 16:40:18 +00:00
Boyuan Feng	514409d032	update torchvision pin (#154255 ) Fixes #153985 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154255 Approved by: https://github.com/desertfire	2025-05-27 16:15:25 +00:00
ZhiweiYan-96	0ddfd1ed43	[Intel GPU] Enable mkdnn._linear_pointwise at XPU backend (#140365 ) # Motivation This PR is intended to add post-op fusion support fo Linear. The liner-pointwise fusion is expected to be used in graph mode like torch.compile. The FusionUtils.cpp file defines a utilization APIs for generating primitive attribute. This APIs would also be used for conv-pointwise fusion, which is in #140372. # Validation ```bash python test/xpu/test_fusion.py ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/140365 Approved by: https://github.com/etaf, https://github.com/guangyey, https://github.com/EikanWang	2025-05-27 15:57:15 +00:00
Jerry Mannil	0d4de7872a	[ROCm] Improve vectorized elementwise kernel performance in MI300X (#153634 ) * Use non-temporal loads to improve the vectorized elementwise kernel performance on MI300 * Use thread_work_size of 8 or 16 for vectorized elementwise kernel Co-author: @amd-hhashemi Pull Request resolved: https://github.com/pytorch/pytorch/pull/153634 Approved by: https://github.com/jeffdaily	2025-05-27 15:38:43 +00:00
Xuehai Pan	7ae204c3b6	[BE][CI][Easy] Run `lintrunner` on generated `.pyi` stub files (#150732 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/150732 Approved by: https://github.com/malfet, https://github.com/cyyever, https://github.com/aorenste	2025-05-27 14:58:02 +00:00
Yuanhao Ji	0a7eef140b	Add `torch.Tensor._make_wrapper_subclass` to `torch/_C/__init__.pyi` (#154022 ) Fixes #153790 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154022 Approved by: https://github.com/Skylion007	2025-05-27 14:10:00 +00:00
Nikita Shulga	d88699308f	[CI][MacOS] Move more dependencies to pypi (#154309 ) Hopefully last step before all Mac build/tests could be switched away from conda - Update cmake version from 3.22 to 3.25 as 3.22 from pipy seems to be unusable with python-3.12 - Add `--plat-name macosx_11_0_arm64` to setup.py command - Remove `codesign` for cmake workaround (that was probably never really necessary - Install `libpng` and `jpeg-turbo` when building torchbench and build torchaudio without OpenMP (to be fixed) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154309 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-05-27 13:49:40 +00:00
PyTorch MergeBot	11a51a11af	Revert "introduce definitely_contiguous and use it for reshape and tensor meta data computation. (#153432 )" This reverts commit 5c6d7caaaa08f134c3b17ce032cb014527b53417. Reverted https://github.com/pytorch/pytorch/pull/153432 on behalf of https://github.com/malfet due to Looks like it broke flex attention tests, see https://hud.pytorch.org/hud/pytorch/pytorch/main/1?per_page=50&name_filter=g6.4xlarge&mergeEphemeralLF=true ([comment](https://github.com/pytorch/pytorch/pull/153432#issuecomment-2912562570))	2025-05-27 13:42:34 +00:00
Emmanuel Menage	c52a002a22	Add getDeviceProperties api to torch mtia device (#153577 ) topic: not user facing Test Plan: Internal benchmark. Differential Revision: D74256550 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153577 Approved by: https://github.com/nautsimon	2025-05-27 11:55:58 +00:00
atalman	2bd95f3a1f	Move inductor workflows focal (ubuntu 20.04) -> jammy (ubuntu 22.04) (#154153 ) Trying to fix: https://github.com/pytorch/pytorch/issues/154157 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154153 Approved by: https://github.com/Skylion007, https://github.com/huydhn, https://github.com/nike4949, https://github.com/cyyever	2025-05-27 11:53:47 +00:00
Tom Ritchford	6f86c1ce1d	Add pyrefly.toml (#154144 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154144 Approved by: https://github.com/Skylion007	2025-05-27 10:16:30 +00:00
Laith Sakka	5c6d7caaaa	introduce definitely_contiguous and use it for reshape and tensor meta data computation. (#153432 ) when a tensor has unbacked symbols it can be general enough to represent both contiguous and non contiguous tensors. in that case we cant really evaluate is_contiguous. In many places in the code base, we check for is_contiguous to take a fast path. but the general path usually works for both contiguous and not contiguous in that case we probably want to use definitely _contiguous API. This is appleid for reshape in this PR and also to tensor meta data computation, the meta data now will have an attribute that says that its contiguous when its always contiguous. We would store that only if definitely _contiguous is true now. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153432 Approved by: https://github.com/bobrenjc93	2025-05-27 08:54:31 +00:00
anwang	dec5ab8d98	[MTIA Aten Backend][1.2/n] Migrate as_strided to in-tree, and add unit tests (#154336 ) See context in PR https://github.com/pytorch/pytorch/pull/153670 This diff migrate as_strided to in-tree. I found it's not covered by `test_kernel_eager_ci` so also adding unit tests. Differential Revision: [D75385404](https://our.internmc.facebook.com/intern/diff/D75385404/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154336 Approved by: https://github.com/nautsimon	2025-05-27 06:32:38 +00:00
PyTorch MergeBot	ef6306e1c6	Revert "[executorch hash update] update the pinned executorch hash (#153436 )" This reverts commit 8d6139b8d8a75aab5ead4262ff59d48615ebee31. Reverted https://github.com/pytorch/pytorch/pull/153436 on behalf of https://github.com/malfet due to Broke ET sanity ([comment](https://github.com/pytorch/pytorch/pull/153436#issuecomment-2911206795))	2025-05-27 06:02:14 +00:00
Yu, Guangye	870133b2a0	Use get_device_context in aoti runtime for XPU directly (#154360 ) # Motivation Reuse [c10::xpu::get_device_context](`1bebe0424e/c10/xpu/XPUFunctions.h (L27)`) directly to reduce overhead, as it returns a cached `sycl::context` managed by PyTorch. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154360 Approved by: https://github.com/EikanWang	2025-05-27 05:55:59 +00:00
Jing Xu	8d89cdceb6	fix a compilation issue when TORCH_XPU_ARCH_LIST is an empty string (#153604 ) When `XPU_ARCH_FLAGS` is an empty string, compilation will fail on `C10_STRINGIZE(XPU_ARCH_FLAGS)` in file `torch/csrc/xpu/Module.cpp` on Windows. This PR fixes this issue by setting `TORCH_XPU_ARCH_LIST` to `""` to avoid an empty string conversion in `C10_STRINGIZE()` when compiling without an AOT. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153604 Approved by: https://github.com/guangyey, https://github.com/EikanWang Co-authored-by: Aaron Gokaslan <aaronGokaslan@gmail.com> Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-05-27 05:26:46 +00:00
PyTorch UpdateBot	8d6139b8d8	[executorch hash update] update the pinned executorch hash (#153436 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned executorch hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153436 Approved by: https://github.com/pytorchbot	2025-05-27 04:54:46 +00:00
Boyuan Feng	912af9b2c2	update torchbench pin (#154256 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154256 Approved by: https://github.com/huydhn	2025-05-27 04:40:54 +00:00
Xia, Weiwen	8d319607a7	[CPU][Brgemm] add s8s8 GEMM microkernel API (#154358 ) As the title. `u8s8` and `u8u8` have already been supported. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154358 Approved by: https://github.com/leslie-fang-intel, https://github.com/Skylion007, https://github.com/Valentine233	2025-05-27 03:47:56 +00:00
Georgia Phillips	f8010e7b93	[nativert] Move file_util to pytorch core (#153162 ) Summary: fbcode//sigmoid/core/common -> fbcode//caffe2/torch/nativert/common Test Plan: Github CI Differential Revision: D74328089 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153162 Approved by: https://github.com/zhxchen17	2025-05-27 03:42:47 +00:00
Pat Vignola	70d12ccc3f	[Torch] Fix error message formatting in fp8 comparison logic (#153647 ) Summary: Using `\` includes all the tabs from the next line in the error message. Test Plan: Nothing, simply error message fixing Reviewed By: exclamaforte Differential Revision: D74539234 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153647 Approved by: https://github.com/exclamaforte	2025-05-27 02:51:05 +00:00
Will Feng	100ec0b34a	[Inductor] Allow passing in custom lowering dict to register_lowering() (#154344 ) This PR adds support for passing in custom lowering dict to `register_lowering()`, which allows systems (e.g. Helion, https://github.com/pytorch-labs/helion/pull/80) that uses Inductor to maintain their own lowering dict instead of using the Inductor global `lowerings` dict. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154344 Approved by: https://github.com/jansel	2025-05-27 01:35:26 +00:00
cyy	3936e6141c	Remove outdated CUDA 11 conditions (#154313 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/154313 Approved by: https://github.com/eqy	2025-05-27 00:30:14 +00:00
atalman	6006352ed3	[BE] Refactor manywheel build scripts (#154372 ) 1. Remove `CentOS Linux` cases, since its deprecated 2. Remove logic for old CUDA versions 3. Remove logic for `CUDA_VERSION=12.4` since we deprecated CUDA 12.4 support 4. Simplify setting `USE_CUFILE=1` - only supported on CUDA 12.6 and 12.8 builds Pull Request resolved: https://github.com/pytorch/pytorch/pull/154372 Approved by: https://github.com/malfet, https://github.com/huydhn	2025-05-26 23:17:23 +00:00
PyTorch MergeBot	b643076e4e	Revert "[executorch hash update] update the pinned executorch hash (#153436 )" This reverts commit b6868f290e4882f9c895b1c9476327974288eaba. Reverted https://github.com/pytorch/pytorch/pull/153436 on behalf of https://github.com/malfet due to Broke ET sanity ([comment](https://github.com/pytorch/pytorch/pull/153436#issuecomment-2910692163))	2025-05-26 22:09:16 +00:00
Laith Sakka	aaf5cc13d9	[EASY] use guard_or_false instead of gso in Meta converter (#154234 ) this was added in https://github.com/pytorch/pytorch/pull/141659, the current change keep the same intention "i do not want to fail here if i cant tell if the size is zero or not" i am not familiar enough in the code to know if we need here a runtime check, but looking at current impl it seems that guard_or_false is appropriate to match current behaviour and have the same effect of guard_size_oblivious here. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154234 Approved by: https://github.com/bobrenjc93 ghstack dependencies: #154154, #154164, #154167, #154172	2025-05-26 21:59:52 +00:00
Laith Sakka	e33feddb72	used guard_or_false instead of guard_size_oblivious inside maybe_reduce (#154172 ) This was added in https://github.com/pytorch/pytorch/pull/119562 the idea in this loop seems to be the following. ``` if (TORCH_GUARD_SIZE_OBLIVIOUS(size.sym_eq(1))) { // NB: we could short circuit this once needs_reduce is true but there's // no point since the reduction function will guard on this anyway if (!c10::guard_or_false(size.sym_eq(target), __FILE__, __LINE__)) { needs_reduce = true; } } else { if (!size.sym_eq(target).expect_true(__FILE__, __LINE__)) { fail(); } } ``` 1. if we know size ==1 1.1 : if we know for sure size == target --> no reduce needed. 1.2 : we know for sure that size != target --> we do reduction. 1.3: we could not tell if size == target or not --> we do reduction. 2. if we do now know if size ==1 or not we add a runtime assertions that size ==target and we fail at runtime if size is not equal to target. We could have simplified 1.1 and always do reduction under 1.1, since doing 1.3 without runtime checks implies that it is safe, but i feel the reason could be perf here? idk. anyway using TORCH_GUARD_OR_FALSE instead of TORCH_GUARD_SIZE_OBLIVIOUS here is appropriate. there is really no clear reason for size oblivious reasoning. or for this logic not to apply when size is not size like size is always >=0 anyway. but bad reasoning can make us not able to infer that although we know its true here. python test/dynamo/test_misc.py -k test_validate_outputs_unbacked Pull Request resolved: https://github.com/pytorch/pytorch/pull/154172 Approved by: https://github.com/bobrenjc93 ghstack dependencies: #154154, #154164, #154167	2025-05-26 21:59:52 +00:00
Laith Sakka	ab5137b048	used guard_or_false instead of guard_size_oblivious in is_int_or_symint (#154167 ) This is a short circuit, that we should not fail on. Before this PR we would not fail on u0, u0+u1, only if they are size like. but we will fail on u0-u1.. etc for no need. guard_or_false seems appropriate for that reason. This was added in https://github.com/pytorch/pytorch/pull/122145 there was no unit tests for me to verify why it was added, i could not repo using the associated issue , the example does not work. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154167 Approved by: https://github.com/bobrenjc93 ghstack dependencies: #154154, #154164	2025-05-26 21:59:45 +00:00
Laith Sakka	1da2cc52bc	[EASY] remove guard_size_oblivious from is_nonzero proxy call check (#154164 ) This was added in https://github.com/pytorch/pytorch/pull/149637, torch._check can handle unbacked there is no need for size oblivious reasoning here. Note this does not make is_nonzero unbacked friendly. but that is a different story. I ran the test added in https://github.com/pytorch/pytorch/pull/149637 for veirfication. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154164 Approved by: https://github.com/aorenste, https://github.com/bobrenjc93 ghstack dependencies: #154154	2025-05-26 21:59:29 +00:00
Laith Sakka	f8a2998832	[EASY] used guard_or_false instead of guard_sizes_oblivious in pointless_view (#154154 ) The change is direct and clear, the optimizations removes pointless_view iff it all sizes are the same if not we want to return false, there is no need for size oblivious reasoning. this was added in https://github.com/pytorch/pytorch/pull/139136, run existing tests that are added in that PR. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154154 Approved by: https://github.com/bobrenjc93	2025-05-26 21:59:21 +00:00
atalman	e89ee1e217	Pin almalinux version to 8.10-20250519 (#154367 ) This PR pins Almalinux version to latest supported 8.10 This is related to: https://github.com/pytorch/pytorch/pull/154364 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154367 Approved by: https://github.com/jeanschmidt, https://github.com/wdvr, https://github.com/malfet, https://github.com/huydhn	2025-05-26 20:08:20 +00:00
Randolf Scholz	839c9c6156	Use property instead of ClassVar for `Uniform.arg_constraints` and `Wishart.arg_constraints` (#154361 ) Fixes #154355 For these two distributions, the constraints depend on the actual values, and so `arg_constraints` cannot be a `ClassVar`. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154361 Approved by: https://github.com/Skylion007	2025-05-26 17:48:28 +00:00
PyTorch MergeBot	3f64502c98	Revert "Re-enable FakeTensor caching for SymInts (#152662 )" This reverts commit 7d11c61c26c596076613aa0111892f7cbccae32e. Reverted https://github.com/pytorch/pytorch/pull/152662 on behalf of https://github.com/malfet due to Looks like it broke bunch of inductor tests, see `187d38185e/1` ([comment](https://github.com/pytorch/pytorch/pull/152662#issuecomment-2910293593))	2025-05-26 17:13:22 +00:00
Henry Tsang	187d38185e	[cutlass backend] Do not raise hard error when re worker has cuda compilation error (#154173 ) fbcode specific Differential Revision: D75262641 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154173 Approved by: https://github.com/bertmaher	2025-05-26 17:10:36 +00:00
Yuki Kobayashi	f55f2f42a7	Add missing docstring for `sym_ite` (#154201 ) `sym_ite` is listed in [the reference page](https://docs.pytorch.org/docs/stable/torch.html) and has no document. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154201 Approved by: https://github.com/Skylion007	2025-05-26 15:59:21 +00:00
atalman	02445ec8f0	Almalinux image, install glibc-langpack-en (#154364 ) After update to: https://hub.docker.com/layers/amd64/almalinux/8/images/sha256-4f63eb966695df3c993deeacec7c73d87728e2ea66d3b48fed4b40cb547fa7c2 Started seeing warning: bash: warning: setlocale: LC_ALL: cannot change locale (en_US.UTF-8) and random Segfaults when using python like: https://github.com/pytorch/test-infra/actions/runs/15216565225/job/42901732536 ``` +++ python -c 'import torch' ./check_binary.sh: line 258: 2276 Segmentation fault (core dumped) python -c 'import torch' ``` Installing langpack does resolve these issues: https://github.com/pytorch/test-infra/actions/runs/15256338815/job/42904808826#step:15:2311 Almalinux Docker build without setlocale warning: https://github.com/pytorch/pytorch/actions/runs/15030284546/job/42240978131 Almalinux Docker build with setlocale warning: https://github.com/pytorch/pytorch/actions/runs/15246391200/job/42873875745#step:3:7180 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154364 Approved by: https://github.com/Skylion007, https://github.com/jeanschmidt	2025-05-26 15:56:42 +00:00
Nikita Shulga	4b0ee3f4f2	[BE] Do not templetize unnnecessarily (#154305 ) `${{ os.runner }}` would always evaluate to macOS for those files And architecutre is always ARM64 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154305 Approved by: https://github.com/atalman	2025-05-26 15:00:48 +00:00
Aleksei Nikiforov	7ab4fae62a	Fix s390x vectorization compilation in inductor (#153946 ) Fix s390x vectorization compilation in inductor. One of failing tests is inductor/test_aot_inductor.py::AOTInductorTestABICompatibleCpu::test_add_complex_cpu but it is still disabled. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153946 Approved by: https://github.com/malfet, https://github.com/jgong5	2025-05-26 12:54:25 +00:00
Ben Olson	1bebe0424e	Fix platform detection in MKLDNN CMake file (#142067 ) When building PyTorch with `USE_XPU=True` and Clang, the user sees misleading errors related to incorrect platform detection that assumes that all users that are not using the GNU compilers are on Windows. We can fix this by simply using CMake's builtin platform detection variables. Pull Request resolved: https://github.com/pytorch/pytorch/pull/142067 Approved by: https://github.com/EikanWang, https://github.com/min-jean-cho, https://github.com/guangyey	2025-05-26 06:09:37 +00:00
Narek Malkhasyan	21e42c5d62	More descriptive error message for torch.nanmean() with complex dtypes (#153252 ) Fixes #153132 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153252 Approved by: https://github.com/colesbury	2025-05-26 05:42:57 +00:00
PyTorch UpdateBot	b6868f290e	[executorch hash update] update the pinned executorch hash (#153436 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned executorch hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153436 Approved by: https://github.com/pytorchbot	2025-05-26 04:43:10 +00:00
Aaron Orenstein	7d11c61c26	Re-enable FakeTensor caching for SymInts (#152662 ) Summary: This backs out D60320595 which itself turned off FakeTensor caching when a SymInt was present. There has been a lot of dynamic shape fixes done this year and tests pass so I'm assuming some of that work fixed what was breaking previously. Test Plan: Reran the tests listed in T196779132 and they pass. ## Perf ### Instruction Counter Benchmark: - 26% win on add_loop_eager_dynamic - 13% win on add_loop_inductor_dynamic_gpu ### Perf Dashboard Compilation Latency wins across the board but especially strong on the dynamic tests (like cudagraphs_dynamic) - for example MobileBertForMaskedLM went from 66s -> 50s. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152662 Approved by: https://github.com/anijain2305	2025-05-26 04:17:56 +00:00
Ke Wen	062387fb53	[SymmMem] Speed up tests (#153677 ) Use `MultiProcContinousTest` to avoid re-create ProcessGroup in each test instance. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153677 Approved by: https://github.com/fegin, https://github.com/Skylion007, https://github.com/ngimel ghstack dependencies: #153653	2025-05-26 03:39:11 +00:00
Ke Wen	8c16d0e404	[c10d] Add support for testing SIGABRT return (#153167 ) `SIGABRT` is a common return by negative distributed tests, which checks for effectiveness of NaN assert, watchdog throw, etc. These errors are not detectable by traditional statements like `with self.assertRaises(RuntimeError)`. Instead, we'd need to check for the process's return code, e.g. `SIGABRT(6)` would have a return code of -6. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153167 Approved by: https://github.com/fduwjj	2025-05-26 00:56:05 +00:00
Natalia Gimelshein	b04852e404	Fix deterministic indexing with broadcast (#154296 ) Fixes #79987, now for real. Also removed thrust sort path that was needed for cuda <=11.2 because we no longer support it. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154296 Approved by: https://github.com/soumith	2025-05-25 21:14:50 +00:00
Justin Chu	c3100067ae	[ONNX] Update onnx to 1.18 (#153746 ) Update onnx python package to 1.18. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153746 Approved by: https://github.com/titaiwangms, https://github.com/cyyever, https://github.com/malfet	2025-05-25 20:58:47 +00:00
Laith Sakka	43b2716e89	PYFMT lint grandfathered files 1 (#154261 ) lint: - test/test_fake_tensor.py - test/test_flop_counter.py - torch/_export/verifier.py with same rules as other files, it was a night mare for me to update tests in one of the skipped files with not being able to lint them locally like other files with lintrunner -a. note that those file do have active dev and not old not touched files. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154261 Approved by: https://github.com/angelayi, https://github.com/Skylion007	2025-05-25 17:36:14 +00:00
Nikita Shulga	5677ab9aab	[BE] Correctly pass exceptions raised from `rpc_init` to CPython (#154325 ) By decorating function body with `HANDLE_TH_ERRORS` Partially addresses https://github.com/pytorch/pytorch/issues/154300 I.e. after that change, importing torch no longer crashes but returns a readable (and actionable exception) ``` >>> import torch Traceback (most recent call last): File "<python-input-0>", line 1, in <module> import torch File "/Users/malfet/git/pytorch/pytorch/torch/__init__.py", line 2134, in <module> from torch import _VF as _VF, functional as functional # usort: skip ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/malfet/git/pytorch/pytorch/torch/functional.py", line 8, in <module> import torch.nn.functional as F File "/Users/malfet/git/pytorch/pytorch/torch/nn/__init__.py", line 8, in <module> from torch.nn.modules import * # usort: skip # noqa: F403 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/malfet/git/pytorch/pytorch/torch/nn/modules/__init__.py", line 2, in <module> from .linear import Bilinear, Identity, LazyLinear, Linear # usort: skip ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/malfet/git/pytorch/pytorch/torch/nn/modules/linear.py", line 7, in <module> from torch.nn import functional as F, init File "/Users/malfet/git/pytorch/pytorch/torch/nn/functional.py", line 11, in <module> from torch._jit_internal import ( ...<5 lines>... ) File "/Users/malfet/git/pytorch/pytorch/torch/_jit_internal.py", line 42, in <module> import torch.distributed.rpc File "/Users/malfet/git/pytorch/pytorch/torch/distributed/rpc/__init__.py", line 37, in <module> from torch._C._distributed_rpc import ( # noqa: F401 ...<33 lines>... ) ImportError: cannot import name '_DEFAULT_NUM_WORKER_THREADS' from 'torch._C._distributed_rpc' (unknown location) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154325 Approved by: https://github.com/Skylion007	2025-05-25 17:01:45 +00:00
Nikita Shulga	31ae07b5e7	[CI] Do not install libuv on MacOS (#154307 ) It's tensorpipe submodule and is build from source Same for `dataclasses` as it's needed only for python-3.6 And get rid of `nidia-ml-py` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154307 Approved by: https://github.com/cyyever, https://github.com/Skylion007 ghstack dependencies: #154304	2025-05-25 15:30:38 +00:00
Nikita Shulga	6968386385	[BE] Sort requirements files alphabetically (#154304 ) Using `sort` tool Pull Request resolved: https://github.com/pytorch/pytorch/pull/154304 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-05-25 15:30:38 +00:00
dependabot[bot]	ed27ee8355	Bump setuptools from 70.0.0 to 78.1.1 in /tools/build/bazel (#154075 ) Bumps [setuptools](https://github.com/pypa/setuptools) from 70.0.0 to 78.1.1. <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/pypa/setuptools/blob/main/NEWS.rst">setuptools's changelog</a>.</em></p> <blockquote> <h1>v78.1.1</h1> <h2>Bugfixes</h2> <ul> <li>More fully sanitized the filename in PackageIndex._download. (<a href="https://redirect.github.com/pypa/setuptools/issues/4946">#4946</a>)</li> </ul> <h1>v78.1.0</h1> <h2>Features</h2> <ul> <li>Restore access to _get_vc_env with a warning. (<a href="https://redirect.github.com/pypa/setuptools/issues/4874">#4874</a>)</li> </ul> <h1>v78.0.2</h1> <h2>Bugfixes</h2> <ul> <li>Postponed removals of deprecated dash-separated and uppercase fields in <code>setup.cfg</code>. All packages with deprecated configurations are advised to move before 2026. (<a href="https://redirect.github.com/pypa/setuptools/issues/4911">#4911</a>)</li> </ul> <h1>v78.0.1</h1> <h2>Misc</h2> <ul> <li><a href="https://redirect.github.com/pypa/setuptools/issues/4909">#4909</a></li> </ul> <h1>v78.0.0</h1> <h2>Bugfixes</h2> <ul> <li>Reverted distutils changes that broke the monkey patching of command classes. (<a href="https://redirect.github.com/pypa/setuptools/issues/4902">#4902</a>)</li> </ul> <h2>Deprecations and Removals</h2> <ul> <li>Setuptools no longer accepts options containing uppercase or dash characters in <code>setup.cfg</code>.</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="`8e4868a036`"><code>8e4868a</code></a> Bump version: 78.1.0 → 78.1.1</li> <li><a href="`100e9a61ad`"><code>100e9a6</code></a> Merge pull request <a href="https://redirect.github.com/pypa/setuptools/issues/4951">#4951</a></li> <li><a href="`8faf1d7e0c`"><code>8faf1d7</code></a> Add news fragment.</li> <li><a href="`2ca4a9fe47`"><code>2ca4a9f</code></a> Rely on re.sub to perform the decision in one expression.</li> <li><a href="`e409e80029`"><code>e409e80</code></a> Extract _sanitize method for sanitizing the filename.</li> <li><a href="`250a6d1797`"><code>250a6d1</code></a> Add a check to ensure the name resolves relative to the tmpdir.</li> <li><a href="`d8390feaa9`"><code>d8390fe</code></a> Extract _resolve_download_filename with test.</li> <li><a href="`4e1e89392d`"><code>4e1e893</code></a> Merge <a href="https://github.com/jaraco/skeleton">https://github.com/jaraco/skeleton</a></li> <li><a href="`3a3144f0d2`"><code>3a3144f</code></a> Fix typo: <code>pyproject.license</code> -> <code>project.license</code> (<a href="https://redirect.github.com/pypa/setuptools/issues/4931">#4931</a>)</li> <li><a href="`d751068fd2`"><code>d751068</code></a> Fix typo: pyproject.license -> project.license</li> <li>Additional commits viewable in <a href="https://github.com/pypa/setuptools/compare/v70.0.0...v78.1.1">compare view</a></li> </ul> </details> <br /> [![Dependabot compatibility score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=setuptools&package-manager=pip&previous-version=70.0.0&new-version=78.1.1)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot merge` will merge this PR after your CI passes on it - `@dependabot squash and merge` will squash and merge this PR after your CI passes on it - `@dependabot cancel merge` will cancel a previously requested merge and block automerging - `@dependabot reopen` will reopen this PR if it is closed - `@dependabot close` will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) You can disable automated security fix PRs for this repo from the [Security Alerts page](https://github.com/pytorch/pytorch/network/alerts). </details> Pull Request resolved: https://github.com/pytorch/pytorch/pull/154075 Approved by: https://github.com/Skylion007 Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2025-05-25 15:13:03 +00:00
Nikita Shulga	c113cf5a8f	[BE] Remove unused `conda-env-Linux-X64` (#154303 ) According to https://github.com/search?type=code&q=conda-env-++repo%3Apytorch%2Fpytorch it's not referenced anywhere and has been replaced with `conda-env-ci` a while ago Pull Request resolved: https://github.com/pytorch/pytorch/pull/154303 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-05-25 14:24:28 +00:00
Aaron Gokaslan	d8aed0703e	[BE][Ez]: Enable ruff rule PLW1507. os.environ is not copied (#154120 ) Enables a RUFF rule check against copying os.environ since its' actually a proxy object, not a dict so a shallow copy will be a noop which is rarely desired behavior. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154120 Approved by: https://github.com/malfet	2025-05-25 14:22:57 +00:00
PyTorch MergeBot	54932d865e	Revert "[c10d] Add support for testing SIGABRT return (#153167 )" This reverts commit 03e102dbe8cbffc2e42a3122b262d02f03571de7. Reverted https://github.com/pytorch/pytorch/pull/153167 on behalf of https://github.com/malfet due to It broke lint ([comment](https://github.com/pytorch/pytorch/pull/153167#issuecomment-2907820789))	2025-05-25 13:17:27 +00:00
Keith	c4ef4090c5	Fix segfault on exit in CachingHostAllocator by signaling background thread to exit (#154117 ) Fixes #152008 This PR fixes a segmentation fault that occurred when exiting the program due to improper background thread management in CachingHostAllocator. Previously, the background thread continued running and called process_events() even after the allocator object was destroyed, leading to a crash on exit. `f12d8d60b1/aten/src/ATen/core/CachingHostAllocator.h (L218)` ```cpp // Launch the background thread and process events in a loop. static bool background_thread_flag [[maybe_unused]] = [this] { getBackgroundThreadPool()->run([&]() { while (true) { process_events(); // <-- This line may cause segfault on exit std::this_thread::sleep_for(std::chrono::microseconds(100)); } }); return true; }(); ``` The fix adds a mechanism to signal the background thread to exit before the object is destructed, ensuring the thread stops safely. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154117 Approved by: https://github.com/ngimel, https://github.com/cyyever	2025-05-25 07:46:12 +00:00
Ke Wen	9d922b55ef	[Distributed][CI] Rework continuous TestCase (#153653 ) 1. Reworked `MultiProcContinousTest` to spawn processes during `setUpClass` instead of `main` (so that we can support multiple TestClass'es in one file). 2. The child processes are now an infinite loop, monitoring test IDs passed from main process via a task queue. Reciprocally, the child processes inform the main process completion of a test via a completion queue. 3. Added a test template. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153653 Approved by: https://github.com/d4l3k, https://github.com/fegin, https://github.com/fduwjj	2025-05-25 03:49:29 +00:00
Ke Wen	03e102dbe8	[c10d] Add support for testing SIGABRT return (#153167 ) `SIGABRT` is a common return by negative distributed tests, which checks for effectiveness of NaN assert, watchdog throw, etc. These errors are not detectable by traditional statements like `with self.assertRaises(RuntimeError)`. Instead, we'd need to check for the process's return code, e.g. `SIGABRT(6)` would have a return code of -6. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153167 Approved by: https://github.com/fduwjj	2025-05-25 03:48:34 +00:00
Justin Chu	10c51b11ff	Bump protobuf version and refactor tensorboard tests (#154244 ) In preparation for https://github.com/pytorch/pytorch/pull/153746, I am bumping protobuf to 5.29.4 and fixing the tensorboard tests first. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154244 Approved by: https://github.com/malfet, https://github.com/cyyever	2025-05-25 00:50:07 +00:00
bobrenjc93	53ecb8159a	Introduce statically_known_false (#154291 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154291 Approved by: https://github.com/mengluy0125	2025-05-24 14:23:55 +00:00
xinan.lin	2dfc0e3327	[Inductor UT] Reuse test_fused_attention.py for Intel GPU. (#154110 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154110 Approved by: https://github.com/eellison, https://github.com/jansel, https://github.com/EikanWang	2025-05-24 09:51:33 +00:00
cyy	8fe7ec6721	Add /Zc:preprocessor for torch libraries in MSVC builds (#147825 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/147825 Approved by: https://github.com/janeyx99	2025-05-24 06:57:46 +00:00
Aaron Orenstein	6503b4a96e	Update to using mypy 1.15 (#154054 ) The BC break isn't real - mypy decided to start complaining about the way we were typing that function. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154054 Approved by: https://github.com/Skylion007	2025-05-24 04:30:57 +00:00
Eddie Yan	76ed9db468	[cuBLAS][cuBLASLt] Use cuBLAS default workspace size in Lt (#153556 ) Also enables unified workspaces by default for non-FBCODE use cases. Default Lt workspace size is also updated to match cuBLAS logic for default, including for Blackwell (SM 10.0) and GeForce Blackwell (SM 12.0). Recommended defaults are documented here: https://docs.nvidia.com/cuda/cublas/#cublassetworkspace Pull Request resolved: https://github.com/pytorch/pytorch/pull/153556 Approved by: https://github.com/Skylion007, https://github.com/ngimel	2025-05-24 03:43:35 +00:00
Svetlana Karslioglu	1ab2993345	Add a link to transformer_building_blocks tutorial (#154281 ) Cross-link to https://docs.pytorch.org/tutorials/intermediate/transformer_building_blocks.html Pull Request resolved: https://github.com/pytorch/pytorch/pull/154281 Approved by: https://github.com/mikaylagawarecki	2025-05-24 02:50:24 +00:00
Yu, Guangye	e904d01c16	Make inductor UT to be generic (#154196 ) # Motivation https://github.com/pytorch/pytorch/pull/151773 introduces UT `test_triton_template_generated_code_caching` failed on XPU; https://github.com/pytorch/pytorch/pull/153895 introduces UT `test_mutation_rename` failed on XPU; fix https://github.com/pytorch/pytorch/issues/154218 # Additional Context With this PR, both failed UTs passed on local machine. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154196 Approved by: https://github.com/jansel	2025-05-24 02:47:46 +00:00
Pian Pawakapan	a19f2cdf29	[draft export] skip when no LOC found (#154190 ) Couldn't repro error, but verified fix with @ColinPeppler Pull Request resolved: https://github.com/pytorch/pytorch/pull/154190 Approved by: https://github.com/ColinPeppler	2025-05-24 02:29:34 +00:00
Nikita Shulga	975bbc63db	[MPS][BE] Move fmod/remainder to Metal ops (#154280 ) This accomplishes following: - Fixes correctness problem with large integer types (though probably makes it slower, but this could not be avoided if one wants to compute accurate answer) - Makes op faster for floating point types (as Metal kernel invocation is faster than creating MPSGraph) - Eliminates need for several correctness workarounds Fixes https://github.com/pytorch/pytorch/issues/154171 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154280 Approved by: https://github.com/dcci ghstack dependencies: #154275, #154290	2025-05-24 01:45:33 +00:00
Nikita Shulga	8f08bdb7f2	[MPS][BE] Code dedup (#154290 ) Eliminate some copy-pasta by introducing `REGISTER_FLOAT_BINARY_OP` and `REGISTER_INTEGER_BINARY_OP` macros Use `_METAL_310_PLUS` to guard bfloat dtype use Pull Request resolved: https://github.com/pytorch/pytorch/pull/154290 Approved by: https://github.com/yangw-dev, https://github.com/wdvr ghstack dependencies: #154275	2025-05-24 01:41:31 +00:00
Nikita Shulga	e5f63f4f66	[CI] Move Mac testing to 3.12 (#154177 ) Prep step to completely move away from Conda during the builds.. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154177 Approved by: https://github.com/huydhn, https://github.com/cyyever, https://github.com/atalman ghstack dependencies: #154237, #154268, #154271, #154269, #154270	2025-05-24 01:41:20 +00:00
Catherine Lee	11a490f32f	[CI] Reuse old whl on more workflows (#154285 ) Still only on main branch, not PRs, so that we can monitor Pull Request resolved: https://github.com/pytorch/pytorch/pull/154285 Approved by: https://github.com/malfet	2025-05-24 01:25:35 +00:00
Zhengxu Chen	308beeeb56	[dynamo] Use UUID for compiled function variable names. (#154148 ) Summary: We previously assign each compiled function variable a name based on in-process global counter. This works fine within the same process but when we're trying to serialize the states with precompile, we need a way to load back these compiled functions without causing collision to the existing global scope. Changing the counter to a true global uuid seems to resolve this issue. For example, the new variable name will look like: ``` __compiled_fn_0_7ce7d872_4fe8_4174_b8fd_2496b09b8b43 ``` Test Plan: CI Differential Revision: D75244901 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154148 Approved by: https://github.com/jansel	2025-05-24 01:08:42 +00:00
leslie-fang-intel	7ba6fb69e6	[Inductor][CPP] Enable vectorized fp8 E5M2 quant dequant (#153365 ) Summary This PR enables the vectorization codegen with Inductor CPP backend for `FP8_E5M2` `quant` from `float32` and `dequant` to `float32`. Test Plan ``` python test/inductor/test_cpu_repro.py -k test_dequant_quant_lowering_fp8_e5m2 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153365 Approved by: https://github.com/jansel, https://github.com/jgong5 ghstack dependencies: #152417, #152418, #153364	2025-05-23 23:20:02 +00:00
leslie-fang-intel	84b657d0b5	Add Vectorized FP8 E5M2 (#153364 ) Summary This PR mainly adding the `Vectorized<Float8_e5m2>` class to support the vectorization of `FP8 E5M2` with methods: - Convert to/from `Vectorized<float>` - Common vectorized methods like: `mul`, `abs`, `eq` and etc. Test Plan ``` ./build/bin/vec_test_all_types_AVX512 --gtest_filter=FP8E5M2Test.* ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153364 Approved by: https://github.com/jgong5, https://github.com/CaoE, https://github.com/vkuzo ghstack dependencies: #152417, #152418	2025-05-23 23:11:25 +00:00
leslie-fang-intel	b77a6504fa	[Inductor][CPP] Enable vectorized fp8 quant dequant (#152418 ) Summary This PR enables the vectorization codegen with Inductor CPP backend for `FP8_E4M3` `quant` from `float32` and `dequant` to `float32`. Test Plan ``` python test/inductor/test_cpu_repro.py -k test_dequant_quant_lowering_fp8_e4m3 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/152418 Approved by: https://github.com/jansel, https://github.com/jgong5, https://github.com/CaoE ghstack dependencies: #152417	2025-05-23 23:05:17 +00:00
leslie-fang-intel	080b74ce67	Add Vectorized FP8 E4M3 (#152417 ) Summary This PR mainly adding the `Vectorized<Float8_e4m3fn>` class to support the vectorization of `FP8 E4M3` with methods: - Convert to/from `Vectorized<float>` - Common vectorized methods like: `mul`, `abs`, `eq` and etc. Test Plan ``` ./build/bin/vec_test_all_types_AVX512 --gtest_filter=FP8E4M3Test.* ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/152417 Approved by: https://github.com/mingfeima, https://github.com/CaoE, https://github.com/yanbing-j, https://github.com/jgong5, https://github.com/vkuzo	2025-05-23 22:56:56 +00:00
Ting Lu	bab59d3c28	Upgrade to CUDA 12.8.1 for nightly binaries (#152923 ) Upgrade current CUDA 12.8 builds to 12.8.1 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152923 Approved by: https://github.com/atalman	2025-05-23 22:37:05 +00:00
Scott Wolchok	f0b2706914	remove sleef_arm target (#154166 ) Summary: X-link: https://github.com/pytorch/executorch/pull/11082 We shouldn't need an ARM-specific variant; we have select() where we should need it. Test Plan: CI Reviewed By: nlutsenko Differential Revision: D74356413 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154166 Approved by: https://github.com/kimishpatel, https://github.com/malfet, https://github.com/Skylion007	2025-05-23 22:16:01 +00:00
Zain Rizvi	86a160353e	[BE] Don't run windows builds in pull.yml (#154264 ) We already run windows builds and tests [during trunk.yml](`c13eeaa718/.github/workflows/trunk.yml (L115-L130)`). Spot checking for failures of this job in pull.yml shows that the most of the times this job fails, the failure correlates with other build jobs failing as well, so it's not offering much unique signal. Given that we'll run this job before merging the PR as part of trunk.yml anyways, the trade off of extra signal from getting a windows build signal a little earlier doesn't seem worth the infra investment. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154264 Approved by: https://github.com/malfet	2025-05-23 22:03:19 +00:00
Catherine Lee	65f0cf3df5	[mergebot] Do not block on autoformat workflow (#154236 ) Helps with https://github.com/pytorch/pytorch/issues/154084 Merge sometimes fails due to autoformat failing. I believe it's because author doesn't have write perms/workflow running perms -> needs approval for workflows. On merge, the bot adds the merge label -> triggers autoformat workflow -> needs approval (even though it will end up getting get skipped because the label doesn't match) -> merge sees and fails So I put an ugly exception for the workflow in mergebot Some restrictions to keep in mind: * Need to checkout the PRs code changes to run lint/format on them -> possible security issue if someone modifies a linter/formatter * The (third party) reusable action used in the autoformat workflow requires the trigger to be pull_request Pull Request resolved: https://github.com/pytorch/pytorch/pull/154236 Approved by: https://github.com/malfet	2025-05-23 22:00:34 +00:00
James Wu	bb17f9c98b	[AOTAutogradCache] Fix CHROMIUM_EVENT_LOG being none (#154258 ) It turns out if you import something that's None at import time in python, and later update the value, the one you imported stays none: ``` import torch from torch._dynamo.utils import CHROMIUM_EVENT_LOG class Foo: pass torch._dynamo.utils.CHROMIUM_EVENT_LOG = Foo() print(CHROMIUM_EVENT_LOG) # None ``` This fixes teh bug so we get AOTAUtogradCache instant events again Differential Revision: [D75305770](https://our.internmc.facebook.com/intern/diff/D75305770/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154258 Approved by: https://github.com/oulgen	2025-05-23 21:53:31 +00:00
Nikita Shulga	0e4f1b8a06	[CI] Update MacOS conda requirmenets (#154270 ) Pick package versions which are compatible with both 3.9 and 3.12 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154270 Approved by: https://github.com/clee2000, https://github.com/atalman ghstack dependencies: #154237, #154268, #154271, #154269	2025-05-23 21:44:50 +00:00
Nikita Shulga	5db1503846	[CI] Update MacOS numba and scipy versions (#154269 ) Pick versions that supported by both 3.9 and 3.12 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154269 Approved by: https://github.com/clee2000, https://github.com/atalman ghstack dependencies: #154237, #154268, #154271	2025-05-23 21:44:49 +00:00
Howard Huang	aa3eab2ce6	Fix tcp init when using port 0 (#154156 ) I hit this in tests when calling `init_process_group(init_method="tcp://localhost:0", ...)`. You can't use port 0 due to the bug in the conditional and will get error `ValueError: Error initializing torch.distributed using tcp:// rendezvous: port number missing` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154156 Approved by: https://github.com/d4l3k, https://github.com/Skylion007	2025-05-23 21:41:58 +00:00
Anthony Shoumikhin	3c0b93afc5	Re-enable link linter (#153280 ) And make URL linter always succeed for now. I'll monitor the logs manually and experiment with it futher. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153280 Approved by: https://github.com/albanD	2025-05-23 20:56:25 +00:00
Nikita Shulga	6f34d141ab	[MPS][BE] Delete `complex_div` (#154275 ) An absolute no-op: delete `complex_div` from `UnaryKernel.metal` and use identical one from `c10/metal/utils.h` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154275 Approved by: https://github.com/dcci	2025-05-23 20:53:50 +00:00
Nikita Shulga	dec6a47996	[BE] Delete unused pip-requirements-iOS.txt (#154271 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154271 Approved by: https://github.com/clee2000 ghstack dependencies: #154237, #154268	2025-05-23 20:08:19 +00:00
Nikita Shulga	acd0873d3b	[CI] Fix `TestDynamoTimed.test_ir_count` for 3.12 (#154268 ) Python-3.12 emits the same bytecode as 3.13 for code in question Pull Request resolved: https://github.com/pytorch/pytorch/pull/154268 Approved by: https://github.com/clee2000, https://github.com/atalman ghstack dependencies: #154237	2025-05-23 20:08:19 +00:00
PyTorch MergeBot	28af44285b	Revert "[c10d] Add support for testing SIGABRT return (#153167 )" This reverts commit 499a76b844bbcbc5465cb76c617b3076c1b0fd65. Reverted https://github.com/pytorch/pytorch/pull/153167 on behalf of https://github.com/malfet due to Broke lint, see `fe784c5a2c/1` ([comment](https://github.com/pytorch/pytorch/pull/153167#issuecomment-2905623868))	2025-05-23 19:44:08 +00:00
Shangdi Yu	fe784c5a2c	Fix torchbind path in AOTI package loader (#154265 ) Summary: as title, fix the path in package loader and fix the test to take the additional dir into consideration. Test Plan: ``` buck run 'fbcode//mode/dev-nosan' fbcode//caffe2/test/inductor:torchbind ``` Reviewed By: angelayi Differential Revision: D75308904 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154265 Approved by: https://github.com/clee2000, https://github.com/malfet	2025-05-23 19:32:53 +00:00
PyTorch MergeBot	90855835ff	Revert "[AOTI][cutlass backend] Do not remove the cutlass kernel .o file after packaging (#154155 )" This reverts commit 269fa8028f68b29176e21886108634f48b1eced7. Reverted https://github.com/pytorch/pytorch/pull/154155 on behalf of https://github.com/henrylhtsang due to mistake in PR ([comment](https://github.com/pytorch/pytorch/pull/154155#issuecomment-2905514934))	2025-05-23 19:08:40 +00:00
Angela Yi	3b21d79225	[export] Move PT2ArchiveWriter/Reader to torch/export (#153795 ) Summary: Before: `from sigmoid.core.package.pt2_archive import PT2ArchiveWriter, PT2ArchiveReader, is_sigmoid_package` After: `from torch.export.pt2_archive import PT2ArchiveWriter, PT2ArchiveReader, is_pt2_package` By merging the two PT2ArchiveReader/Writers, into using the native PytorchFileReader/Writer, the open source PT2 archive also changed to have an additional folder. However this PR still maintains support for loading an old PT2 archive which does not have the additional folder. Before: ``` ├── archive_format ├── byteorder ├── .data │ ├── serialization_id │ └── version ├── data │ ├── aotinductor ``` After: ``` ├── tmp │ ├── archive_format │ ├── byteorder │ ├── .data │ │ ├── serialization_id │ │ └── version │ ├── data │ │ ├── aotinductor ``` Test Plan: `buck2 test //sigmoid/...` https://www.internalfb.com/intern/testinfra/testrun/5348024839248187 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153795 Approved by: https://github.com/zhxchen17	2025-05-23 19:04:36 +00:00
Ke Wen	499a76b844	[c10d] Add support for testing SIGABRT return (#153167 ) `SIGABRT` is a common return by negative distributed tests, which checks for effectiveness of NaN assert, watchdog throw, etc. These errors are not detectable by traditional statements like `with self.assertRaises(RuntimeError)`. Instead, we'd need to check for the process's return code, e.g. `SIGABRT(6)` would have a return code of -6. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153167 Approved by: https://github.com/fduwjj	2025-05-23 19:04:28 +00:00
PyTorch MergeBot	561a11aa68	Revert "Patch the _is_conv_node function (#153749 )" This reverts commit c985cec5b2545d46af682d486b18866eee5dffd5. Reverted https://github.com/pytorch/pytorch/pull/153749 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/153749#issuecomment-2905504697))	2025-05-23 19:04:20 +00:00
PyTorch MergeBot	4ff19ecf66	Revert "[export] Move PT2ArchiveWriter/Reader to torch/export (#153795 )" This reverts commit 7e80f23516a86e18ae5bc5579d3005c1e7610102. Reverted https://github.com/pytorch/pytorch/pull/153795 on behalf of https://github.com/malfet due to Looks like it broke lots of tests, see `ec368a1903/1` ([comment](https://github.com/pytorch/pytorch/pull/153795#issuecomment-2905415496))	2025-05-23 18:29:08 +00:00
Svetlana Karslioglu	ec368a1903	Add sitemap (#154158 ) Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/154158 Approved by: https://github.com/albanD	2025-05-23 18:01:00 +00:00
Andy (An) Wang	0d62fd5c3c	[MTIA Aten Backend][2/n] Migrate clamp ops(clamp.out/clamp_min.out/clamp_max.out) from out-of-tree to in-tree (#154015 ) Summary: # Context See the first PR https://github.com/pytorch/pytorch/pull/153670 # This PR 1. Migrate 3 clamp ops from out-of-tree to in-tree(had to migrate the 3 ops altogether, because clamp.out calls all 3 stubs, which are also called by the other 2 ops): - clamp.out - clamp_min.out - clamp_max.out 2. Also enabled structured kernel codegen for MTIA, which is needed by clamp 3. Also introduced the `--mtia` flag to torchgen to prevent OSS from gencoding MTIA code.(Otherwise we got such link error `lib/libtorch_cpu.so: undefined reference to at::detail::empty_mtia`) Differential Revision: D74674418 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154015 Approved by: https://github.com/albanD, https://github.com/nautsimon	2025-05-23 17:59:47 +00:00
Nikita Shulga	bcb2125f0a	[BE][CI] Update expecttest version to 0.3.0 (#154237 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154237 Approved by: https://github.com/Skylion007, https://github.com/albanD, https://github.com/atalman	2025-05-23 17:27:41 +00:00
Tsung-Hsien Lee	cae25ef4e5	[c10d] Enhance Error Logging in `new_subgroups()` for Non-Divisible World Sizes (#154124 ) Summary: The error caused by the world size not being divisible by `group_size` is a common issue encountered by end-users when utilizing applications built on top of `new_subgroups()`. However, these applications may employ different variable names, such as `num_trainers_per_group`, which can make the current error messages less effective despite being correct. To address this, we have improved the error messages to display the actual numbers involved, thereby enhancing their clarity and usefulness. Test Plan: contbuild & OSS CI Differential Revision: D75226925 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154124 Approved by: https://github.com/wz337	2025-05-23 17:12:43 +00:00
henrylhtsang	e927ba6dbd	[inductor][cutlass backend] Add 2 stage autotuning aka prescreening (#153335 ) Motivation: By default, we are tuning the cutlass backend kernels on 3 swizzles. There are runtime params, so they share the same underlying kernel, which saves a lot of compilation time. However, autotuning all combinations of {configs} x {swizzles} is still expensive. Observations: Winner of the {configs} x {swizzles} autotuning is the same as if we do a greedy search: first find the top X winners of {configs} with swizzle 2 (hardcoded), then autotune on the {top X winner configs} x {swizzles}. In other words, we can use a Greedy algorithm to reduce autotuning time. I attach the logs below. This somewhat depends on what X is, but a number like 5-10 works pretty well from empirical observations. Logs: Baseline: https://gist.github.com/henrylhtsang/9a604f150a270dc19524f72a5d4dfac2 ``` AUTOTUNE mm(2048x2048, 2048x2048) strides: [2048, 1], [1, 2048] dtypes: torch.bfloat16, torch.bfloat16 cuda_cutlass_gemm_1776 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1777 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1778 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1800 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1801 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1802 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_9012 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_9013 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_9014 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8940 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8941 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8942 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8934 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8935 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8936 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_2001 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_2002 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_2003 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1848 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1849 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1850 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8964 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8965 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8966 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8958 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8959 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8960 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1929 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1930 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1931 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1770 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1771 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1772 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1953 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1954 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1955 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1995 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1996 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1997 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1794 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1795 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1796 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1842 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1843 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1844 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_9006 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_9007 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_9008 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1923 0.0306 ms 95.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 ``` with prescreening: ``` AUTOTUNE mm(147456x6144, 6144x2048) strides: [6144, 1], [2048, 1] dtypes: torch.bfloat16, torch.bfloat16 cutlass_1a5e81af 4.5469 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.6328 ms 98.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.6836 ms 97.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_161b8b81 4.7224 ms 96.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_161b8b81 4.7234 ms 96.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_161b8b81 4.7274 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_853b6347 4.7369 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.7404 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_161b8b81 4.7711 ms 95.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_8bc6fbda 4.8148 ms 94.4% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_8bc6fbda 4.8159 ms 94.4% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_8bc6fbda 4.8214 ms 94.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_8bc6fbda 4.8302 ms 94.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_0a1c55af 4.8487 ms 93.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_0a1c55af 4.8527 ms 93.7% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_02780d72 4.8617 ms 93.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_0a1c55af 4.8737 ms 93.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_0a1c55af 4.8738 ms 93.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_02780d72 4.9348 ms 92.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_02780d72 4.9763 ms 91.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_853b6347 4.9805 ms 91.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_1a5e81af 5.0225 ms 90.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_853b6347 5.0271 ms 90.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_02780d72 5.0595 ms 89.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_853b6347 5.1434 ms 88.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.1574 ms 88.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_1a5e81af 5.1916 ms 87.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.2018 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.2019 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_c1ffa14b 5.2037 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_1a5e81af 5.5329 ms 82.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_aa6f899c 11.5046 ms 39.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 SingleProcess AUTOTUNE benchmarking takes 1.9526 seconds and 0.0352 seconds precompiling for 32 choices ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153335 Approved by: https://github.com/eellison	2025-05-23 17:12:25 +00:00
Shangdi Yu	04a6fe7914	Update provenance tracking doc (#154062 ) Summary: Update the doc to reflect the changes in https://github.com/pytorch/pytorch/pull/153584/files#diff-e0cdb58c0f84f56f20c5433339b6d83c470dcde47847e2328effea6bedd4cd27 and https://github.com/pytorch/tlparse/pull/110 Test Plan: CI Differential Revision: D75155981 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154062 Approved by: https://github.com/svekars, https://github.com/desertfire	2025-05-23 17:09:52 +00:00
Aleksei Nikiforov	7d8ea5db69	Disable cache and utilization stats uploading steps on s390x (#150297 ) There are no AWS credentials available on s390x runners. These steps are failing anyway due to that. Pull Request resolved: https://github.com/pytorch/pytorch/pull/150297 Approved by: https://github.com/seemethere	2025-05-23 16:49:38 +00:00
Angela Yi	7e80f23516	[export] Move PT2ArchiveWriter/Reader to torch/export (#153795 ) Summary: Before: `from sigmoid.core.package.pt2_archive import PT2ArchiveWriter, PT2ArchiveReader, is_sigmoid_package` After: `from torch.export.pt2_archive import PT2ArchiveWriter, PT2ArchiveReader, is_pt2_package` By merging the two PT2ArchiveReader/Writers, into using the native PytorchFileReader/Writer, the open source PT2 archive also changed to have an additional folder. However this PR still maintains support for loading an old PT2 archive which does not have the additional folder. Before: ``` ├── archive_format ├── byteorder ├── .data │ ├── serialization_id │ └── version ├── data │ ├── aotinductor ``` After: ``` ├── tmp │ ├── archive_format │ ├── byteorder │ ├── .data │ │ ├── serialization_id │ │ └── version │ ├── data │ │ ├── aotinductor ``` Test Plan: `buck2 test //sigmoid/...` https://www.internalfb.com/intern/testinfra/testrun/5348024839248187 Differential Revision: D74616598 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153795 Approved by: https://github.com/zhxchen17	2025-05-23 15:40:25 +00:00
Nikita Shulga	214e4cef9f	Fix RMSNorm doc rendering (#154205 ) By removing `::func::` decorator which adds unneeded parenthesis Test plan: Check https://docs-preview.pytorch.org/pytorch/pytorch/154205/generated/torch.nn.RMSNorm.html#rmsnorm that now renders as <img width="704" alt="image" src="https://github.com/user-attachments/assets/443f605d-75a6-41ef-8971-21e7dc8ef9f6" /> Fixes https://github.com/pytorch/pytorch/issues/154184 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154205 Approved by: https://github.com/mikaylagawarecki	2025-05-23 15:39:29 +00:00
Laith Sakka	9e089bb5b6	change guard_or impl for better perf and simplicity (#153674 ) PR time benchmarks has been showing regressions as we move to guard_or_false, reason is that prev implementation do not cache. This new approach will propagate the fallback value to eval and return it. allowing eval to cache and reducing scamming logs and complexity. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153674 Approved by: https://github.com/bobrenjc93	2025-05-23 15:24:28 +00:00
Aaron Orenstein	4b7abce6a4	Fix fake tensor caching when output has unbacked (#153034 ) We handle fake tensor caching in two ways: 1. If the inputs have no symbols (SymInt, etc) then we cache on the FakeTensorMode. 2. If the inputs have symbols then we cache on the ShapeEnv. This way the symbols in the inputs and outputs are associated with the guards in place at the time of the call. However - it's possible to have an op where there are no symbols in the inputs but there is an unbacked symbol in the output. In this case we shouldn't cache at all because what would that really mean? So this PR changes the caching behavior so that if there's a symbol in the output which doesn't come in some way from the input then we refuse to cache that op. Added a test which checks for this case. While in there I also did a couple other related changes: 1. Added negative caching - if we see that an (op, args) failed to cache previously we don't even bother trying to cache it again. 2. Reworked the inner behavior of _cached_dispatch_impl a little to make it more clear which bits we expect to be able to throw _BypassDispatchCache and add some comments. The latest version of this also: 1. Addresses the problem that caused #153891. The issue was that with caching ops are required to support `__eq__`. Unfortunately _RecordFunction is minimalistic and doesn't support that - so in the off-chance that two keys hash to the same value the `__eq__` check would raise an exception. Apparently this was much more common on MacOS where memory patterns end up with more reuse (so the object IDs are the same and give you the same hash value for objects that use pointer hash). Tested locally on MacOS where running ``` python test/inductor/test_torchinductor.py GPUTests ``` was pretty much guaranteed to fail (at least for me) somewhere around test 100-200 and passed all 800 tests after this change. Another way to test this is to run the inductor tests with `torch._subclasses.fake_tensor._DispatchCacheKey.__hash__` monkey-patched to return a constant (causing all values to hash-collide) but this can't really be checked-in since it causes the cache lookup to turn into an O(n) lookup which takes a crazy long time to run through all the tests... 2. Folds in #153780 to ensure that exceptions raised from the op don't include the context from the cache key bypass. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153034 Approved by: https://github.com/masnesral, https://github.com/tugsbayasgalan	2025-05-23 15:03:31 +00:00
PyTorch MergeBot	866142ff16	Revert "Update the heuristic for AArch64 bmm/baddbmm (#149122 )" This reverts commit d759a517af3e6b2337bf8f8e0d1734e64e470f1b. Reverted https://github.com/pytorch/pytorch/pull/149122 on behalf of https://github.com/jeanschmidt due to breaking internal models, @malfet may you help merge this? ([comment](https://github.com/pytorch/pytorch/pull/149122#issuecomment-2904703075))	2025-05-23 14:54:54 +00:00
Nikita Shulga	5859582ee4	[BE][MPS] Delete unused `complex_mul_out` (#154175 ) It's no longer called, after `mul` has been migrated to binary op Pull Request resolved: https://github.com/pytorch/pytorch/pull/154175 Approved by: https://github.com/dcci, https://github.com/Skylion007	2025-05-23 13:44:24 +00:00
Jonathan Deakin	2225231a14	Enable AArch64 CI scripts to be used for local dev (#143190 ) - Allow user to specify custom ComputeLibrary directory, which is then built rather than checking out a clean copy - Remove `setup.py clean` in build. The CI environment should be clean already, removing this enables incremental rebuilds - Use all cores for building ComputeLibrary Mostly a port of https://github.com/pytorch/builder/pull/2028 with the conda part removed, because aarch64_ci_setup.sh has changed and can now handle being called twice. Pull Request resolved: https://github.com/pytorch/pytorch/pull/143190 Approved by: https://github.com/aditew01, https://github.com/fadara01, https://github.com/malfet Co-authored-by: David Svantesson-Yeung <David.Svantesson-Yeung@arm.com>	2025-05-23 12:09:59 +00:00
Ke Wen	25149cd173	[c10d] Add more tests to prevent extra context (#154174 ) Stack from [ghstack](https://github.com/ezyang/ghstack) (oldest at bottom): Loop a bunch of sync ops and see if any of them creates extra context. Requires nvml to check number of processes resident on a device. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154174 Approved by: https://github.com/atalman	2025-05-23 09:54:01 +00:00
wengshiy	ba5d45d22e	Add assertion to align with cuda (#153233 ) Fixes #153137 Aligned batch_norm_cpu_out assertion to [batch_norm_cuda_out](`a7ea115494/aten/src/ATen/native/cuda/Normalization.cu (L436)`). Pull Request resolved: https://github.com/pytorch/pytorch/pull/153233 Approved by: https://github.com/malfet	2025-05-23 07:32:43 +00:00
Autin Mitra	5623d30228	[Minimizer] Gracefully exit when there is no discrepancy in block mode (#154076 ) Summary: Previously, when there is no discrepancy in results for block mode, net_min_base will throw an OOB error. This occurs due to the block _block_traverse_impl returning an OOB after exhausting subgraphs all the way down to a single node There is also an issue where we may get an unsound subgraph (i.e. mark an earlier node as the "end" even if the correct end is later). This is due to an incorrect check (start_idx == mid) where there can possibly be two values left before the program pre-maturely returns Test Plan: Buck UI: https://www.internalfb.com/buck2/52524c26-ace5-4593-8a4b-843a54eb206a Test UI: https://www.internalfb.com/intern/testinfra/testrun/3096224973363310 Network: Up: 0B Down: 15MiB (reSessionID-cd404e97-395f-49fc-8381-373e90a1378f) Executing actions. Remaining 0/1 Command: test. Time elapsed: 53.7s Tests finished: Pass 7. Fail 0. Fatal 0. Skip 0. Build failure 0 Differential Revision: D75143242 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154076 Approved by: https://github.com/jfix71	2025-05-23 06:42:07 +00:00
Filip Jankovic	8342b9371e	[ROCm] Prefer hipblaslt for gfx1200, gfx1201 (#153610 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153610 Approved by: https://github.com/jeffdaily, https://github.com/atalman	2025-05-23 06:01:53 +00:00
angelayi	26471fc203	[aoti] Initial Metal support (#153959 ) An example generated file: P1816629015 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153959 Approved by: https://github.com/malfet, https://github.com/desertfire ghstack dependencies: #153964	2025-05-23 05:45:35 +00:00
angelayi	b33b7d5c8c	[aoti] Add MPS runner and shim (#153964 ) Added AOTIModelContainerRunnerMps and a shim for mps fallback ops. I also added a mps-specific shim which contains one operator, which will be used to set arguments being passed to the Metal kernel: ``` AOTI_TORCH_EXPORT AOTITorchError aoti_torch_mps_set_arg( AOTIMetalKernelFunctionHandle func, unsigned idx, AtenTensorHandle tensor); ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153964 Approved by: https://github.com/malfet, https://github.com/desertfire	2025-05-23 05:45:35 +00:00
henrylhtsang	269fa8028f	[AOTI][cutlass backend] Do not remove the cutlass kernel .o file after packaging (#154155 ) Differential Revision: [D75253009](https://our.internmc.facebook.com/intern/diff/D75253009/) In general, we want to cache the cutlass kernels. Also saw an error saying .o not found. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154155 Approved by: https://github.com/chenyang78	2025-05-23 04:51:36 +00:00
William Wen	5bb156a7fd	[dynamo] raise observed exception for module attribute errors (#153659 ) Fixes https://github.com/pytorch/pytorch/issues/153605 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153659 Approved by: https://github.com/StrongerXi	2025-05-23 03:56:26 +00:00
PyTorch UpdateBot	db1f33147b	[audio hash update] update the pinned audio hash (#154001 ) This PR is auto-generated nightly by [this action](https://github.com/pytorch/pytorch/blob/main/.github/workflows/nightly.yml). Update the pinned audio hash. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154001 Approved by: https://github.com/pytorchbot	2025-05-23 03:51:21 +00:00
Laith Sakka	c1055f41a6	Data dependent free reshape. (#153198 ) #### change 1: if compute_strides stride fail for reshape just clone. Lets consider the most general case, if torch compile is asked to reshape [u0, u1][u3, u4] -> [u5, u6] what shall it do? The shape is general enough to represent both contiguous and non contiguous tensors, tensors where a clone free reshape can happen and other where a clone free cant happen. The current algorithm will fail due to data dependent errors. The general idea is if its impossible to tell if the reshape can happen in place, (because for some concrete inputs it will and other not) then its ok to take the general path and clone, instead of failing or asking the user to give hints. Because the user want a single graph (single compilations) and this is the only way it can be done. Had this been a view? then the user is explicitly asking for a copy-free reshape, we would fail asking for more information (hints in torch.checks form). with this change reshape works as the following: 1. if we know the input is contiguous we will convert the reshape to view. 2. if compute_strides succeed we will use view. (compute_strides was changed to not fail when when unbacked presented instead it will just return nullptr if it cant compute the strides meaning we shall use a clone). 3. if neither 1, 2 works clone and use a view. Side note: having a view does not mean that inductor will not clone, for inductor there is a pass that converts all views back to reshapes and inductor has its logic dealing with those. #### change 2 : skip _reshape_view_helper and fall back to simpler logic if it fail. We trace _reshape_view_helper when doing fake tensor tracing , but not during proxy tracing. hence such tracing wont effect the graph (only compute output shapes of several operations). We should not fail there, because it should always be possible for us to pass it in case of reshape. i.e. when reshape_symint was called we would have either cloned, or compute_strides succeeded so the view should pass. What I did is the following: we run _reshape_view_helper, if we fail due to unbacked we call _view_simple which will succeed always for reshapes, (might fail for views when its impossible to do the view, in such case we throw the dde that was thrown by the original algorithm). Ideally I would want to register _view_simple as the meta for view and avoid calling _reshape_view_helper completely but I am running some issues with the dispatcher with subclasses and I do not have time to debug it. Namely one test would end up calling some c++ view function that does not support symints during meta dispatch when i register a python meta decompositions ```python test/dynamo/test_subclasses.py SubclassTests.test_subclass_views_dynamic_True ``` https://github.com/pytorch/pytorch/issues/153303.I will follow up with that change in a separate PR. cc @H-Huang @awgu @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @ipiszy @chenyang78 @kadeng @muchulee8 @amjames @chauhang @aakhundov @bdhirsh Two other alternatives for registering _view_simple as meta and the try catch approach in this PR is: 1. call _view_simple if any input is dynamic see #153521 2. if we make is_compiling works for framework code tracing (does not work rn) we can call _view_simple is if is_compiling. #### Note: Reshape can still fail when is_contiguous is called, Next PR will handle that by calling is_known_contiguous. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153198 Approved by: https://github.com/etaf, https://github.com/bobrenjc93	2025-05-23 01:45:16 +00:00
Ruisi Zhang	f74842d665	[DTensor] enable SimpleFSDP's composability with Tensor Parallel (#152286 ) This PR adds support for SimpleFSDP's composability with Tensor Parallel + torch.compile. `_StridedShard` is used in SimpleFSDP/FSDP2 to support correct distributed checkpointing when FSDP+TP is applied. Previously, `_StridedShard` is not guarded by torch.compile. This PR adds `_StridedShard` as an additional placement type to be guarded by torch.compile. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152286 Approved by: https://github.com/bdhirsh	2025-05-23 01:40:38 +00:00
Huy Do	7509b150af	Don't upload compiler benchmark debug info to the benchmark database (#153769 ) During our debug session, @wdvr and I found out that the benchmark database is growing much faster than we expect. After taking a closer look, the majority of them coming from TorchInductor benchmark and the top 3 are all debug information not used by any dashboard atm. In the period of 7 days, there are close to 6 millions records ([query](https://paste.sh/GUVCBa0v#UzszFCZaWQxh7oSVsZtfZdVE)) ``` Benchmark,Metric,Count "TorchInductor","user_stack","1926014" "TorchInductor","reason","1926014" "TorchInductor","model","1926014" ``` Let's skip uploading them to avoid bloating the database. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153769 Approved by: https://github.com/malfet	2025-05-23 01:18:26 +00:00
Benjamin Glass	768cb734ec	cpp_wrapper: build non-performance-sensitive code at O1 (#148773 ) Builds on #148212, applying the same improvements to `cpp_wrapper` mode. Benchmark results: * [A100 Benchmarks](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Wed%2C%2014%20May%202025%2015%3A10%3A05%20GMT&stopTime=Wed%2C%2021%20May%202025%2015%3A10%3A05%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cuda%20(a100)&lBranch=gh/benjaminglass1/77/orig&lCommit=ca7d0a3f16e3c511534d2cd03d695be8524570d3&rBranch=main&rCommit=1075bb37d34e483763a09c7810790d5491441e13) * [x86 Benchmarks](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Wed%2C%2014%20May%202025%2015%3A10%3A05%20GMT&stopTime=Wed%2C%2021%20May%202025%2015%3A10%3A05%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cpu%20(x86)&lBranch=gh/benjaminglass1/77/orig&lCommit=ca7d0a3f16e3c511534d2cd03d695be8524570d3&rBranch=main&rCommit=1075bb37d34e483763a09c7810790d5491441e13) Pull Request resolved: https://github.com/pytorch/pytorch/pull/148773 Approved by: https://github.com/desertfire	2025-05-23 00:51:20 +00:00
Svetlana Karslioglu	3c0cbf4b44	Update GH action to use the correct label (#154126 ) Update GH action to use the correct label for the docathon Pull Request resolved: https://github.com/pytorch/pytorch/pull/154126 Approved by: https://github.com/AlannaBurke, https://github.com/clee2000	2025-05-23 00:29:43 +00:00
Aaron Gokaslan	31f3ee0966	[BE][Ez]: Enable PT014 check for duplicate parameterize test cases (#154118 ) Ruff rule which checks for an error [PT014](https://docs.astral.sh/ruff/rules/pytest-duplicate-parametrize-test-cases/) where a user might specify two duplicate test cases in pytest.parameterize, which is likely an error since it tests the same thing twice. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154118 Approved by: https://github.com/malfet	2025-05-23 00:00:53 +00:00
xinan.lin	7b25ff7cf2	[Inductor] Add attention pattern for model DistilBert in transformers==4.44.2. (#154091 ) This PR add a attention fusion pattern that match the attention of DistilDistilBert in transformers==4.44.2 at `953196a43d/src/transformers/models/distilbert/modeling_distilbert.py (L212)` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154091 Approved by: https://github.com/jansel, https://github.com/eellison	2025-05-22 23:37:03 +00:00
PyTorch MergeBot	59c5fff2aa	Revert "[DDP] rebuilt bucket order when find_unused_parameters=true (#153404 )" This reverts commit a79e621c1c11bcef5f816b9770b751237b84f620. Reverted https://github.com/pytorch/pytorch/pull/153404 on behalf of https://github.com/facebook-github-bot due to Diff reverted internally ([comment](https://github.com/pytorch/pytorch/pull/153404#issuecomment-2902741300))	2025-05-22 22:26:59 +00:00
Yuxuan Chen	f2cce45657	[libc++ readiness][caffe2] No reason to check for "ext/stdio_filebuf.h" (#154080 ) Summary: There should be no reason to check for existence of this GNU C++ header here in this file. It doesn't include it. Removing this condition to make it build under libc++. Differential Revision: D75179136 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154080 Approved by: https://github.com/soumith	2025-05-22 22:23:39 +00:00
Chen Lai	c985cec5b2	Patch the _is_conv_node function (#153749 ) Summary: torch.ops.aten.conv2d.padding is also conv2d node Differential Revision: D74898941 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153749 Approved by: https://github.com/andrewor14, https://github.com/Skylion007	2025-05-22 22:17:02 +00:00
bobrenjc93	413664b3c5	catch CSE recursion depth errors (#154039 ) Fixes #153777 CSE is an optimization and shouldn't block a compile if it hits recursion depth limits. Unfortunately we can't write this iteratively due to a dependency on `ast.unparse` which necessarily needs to do recursion. This PR catches opts out of CSE when we hit recursion depth errors. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154039 Approved by: https://github.com/Microve	2025-05-22 20:17:19 +00:00
Rachel Guo	cad0727fe1	Rename the provenance tracing artifact name for kernel <-> post_grad nodes mapping (#154046 ) Summary: Context: Recently we've added a couple more kernel types support other than inductor generated triton kernels, such as cpu cpp kernels, extern kernels. The name appeared in tlparse chrome link can be confusing to users. Rename from `inductor_triton_kernel_to_post_grad_nodes.json` to `inductor_generated_kernel_to_post_grad_nodes.json` Test Plan: CI Differential Revision: D75159042 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154046 Approved by: https://github.com/yushangdi	2025-05-22 19:20:56 +00:00
atalman	4277907d02	[binary builds] Linux aarch64 CUDA builds. Make sure tag is set correctly (#154045 ) 1. This should set the Manylinux 2.28 tag correctly for CUDA Aarch builds. I believe we used to have something similar in the old script: https://github.com/pytorch/pytorch/blob/main/.ci/aarch64_linux/build_aarch64_wheel.py#L811 ``Tag: cp311-cp311-linux_aarch64 ``-> ``Tag: cp311-cp311-manylinux_2_28_aarch64`` 2. Remove section for CUDA 12.6, since we no longer building CUDA 12.6 aarch64 builds Pull Request resolved: https://github.com/pytorch/pytorch/pull/154045 Approved by: https://github.com/Camyll, https://github.com/malfet	2025-05-22 18:36:13 +00:00
Menglu Yu	788d9cb2d7	[3/n][Optimus][Auto-AC][reland] Support any fp8 quantization type and set scaling as the default" (#154057 ) Summary: This is a reland of D74910193. We change the dtype to torch.float8_e5m2 in unit test since it is not supported. Test Plan: ``` buck2 test 'fbcode//mode/dev-nosan' fbcode//caffe2/test/inductor:quantization ``` Differential Revision: D75169792 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154057 Approved by: https://github.com/Mingming-Ding	2025-05-22 18:26:34 +00:00
skishore	c2660d29a5	[ROCm] Added unit test to test the cuda_pluggable allocator (#154041 ) Added unit test to include the cuda_pluggable allocator and replicate the apex setup.py to build nccl_allocator extension This test to check if this commit https://github.com/pytorch/pytorch/pull/152179 helps to build the cuda pluggable allocator in Rocm/Apex Pull Request resolved: https://github.com/pytorch/pytorch/pull/154041 Approved by: https://github.com/atalman, https://github.com/jeffdaily Co-authored-by: Jithun Nair <jithun.nair@amd.com>	2025-05-22 18:22:15 +00:00
Menglu Yu	5b8f422561	[PT2][Optimus] Fix a typo in decompose_mm (#154048 ) Summary: As titled Differential Revision: D75160513 Pull Request resolved: https://github.com/pytorch/pytorch/pull/154048 Approved by: https://github.com/Mingming-Ding	2025-05-22 18:11:40 +00:00
Nikita Shulga	633ed01145	[MPS] Add support for two more isin variants (#154010 ) `isin_Tensor_Scalar_out` is just a redispatch to eq/neq `isin_Scalar_Tensor_out` redispatches back to generic `isin` op, but needs a small tweak to handle float scalars Make sure that `out` is resized to an expected value in `isin_Tensor_Tensor_out_mps` Add unittests to validate that, but skip them on MacOS-13, where MPS op just returns garbage Before this change both of those failed ```python >>> import torch >>> t = torch.tensor([0, 1, 2], device='mps') >>> torch.isin(t, 1) Traceback (most recent call last): File "<stdin>", line 1, in <module> NotImplementedError: The operator 'aten::isin.Tensor_Scalar_out' is not currently implemented for the MPS device. If you want this op to be considered for addition please comment on https://github.com/pytorch/pytorch/issues/141287 and mention use-case, that resulted in missing op as well as commit hash 3b875c25ea6d8802a0c53af9eb961ddf2f058188. As a temporary fix, you can set the environment variable `PYTORCH_ENABLE_MPS_FALLBACK=1` to use the CPU as a fallback for this op. WARNING: this will be slower than running natively on MPS. >>> torch.isin(1, t) Traceback (most recent call last): File "<stdin>", line 1, in <module> NotImplementedError: The operator 'aten::isin.Scalar_Tensor_out' is not currently implemented for the MPS device. If you want this op to be considered for addition please comment on https://github.com/pytorch/pytorch/issues/141287 and mention use-case, that resulted in missing op as well as commit hash 3b875c25ea6d8802a0c53af9eb961ddf2f058188. As a temporary fix, you can set the environment variable `PYTORCH_ENABLE_MPS_FALLBACK=1` to use the CPU as a fallback for this op. WARNING: this will be slower than running natively on MPS. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154010 Approved by: https://github.com/Skylion007, https://github.com/dcci, https://github.com/manuelcandales ghstack dependencies: #153970, #153971, #153997	2025-05-22 17:59:35 +00:00
Xu Han	7421c21b5e	remove unused code. (#153979 ) Remove the unused cmake code. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153979 Approved by: https://github.com/albanD	2025-05-22 17:50:11 +00:00
Yidi Wu	fc859077a0	[export][cond] support merging constant ints as unbacked symint (#152742 ) @pianpwk points out that this will be helpful to address several data dependent issues in huggingface [models](`e23705e557/src/diffusers/schedulers/scheduling_euler_ancestral_discrete.py (L332)`) with the following pattern: ```python idx = return 0 if u0 else return 1 return x[idx] ``` We could preserve the conditional with a cond. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152742 Approved by: https://github.com/zou3519	2025-05-22 17:25:38 +00:00
PyTorch MergeBot	025c5cc048	Revert "[inductor][cutlass backend] Add 2 stage autotuning aka prescreening (#153335 )" This reverts commit d23762974eae105aad837188d5d2254ea9783b37. Reverted https://github.com/pytorch/pytorch/pull/153335 on behalf of https://github.com/yangw-dev due to sorry the pr is failed internally [D75155648](https://www.internalfb.com/diff/D75155648) ([comment](https://github.com/pytorch/pytorch/pull/153335#issuecomment-2901916364))	2025-05-22 16:52:04 +00:00
PyTorch MergeBot	7d3dab6b90	Revert "[BE]: Type previously untyped decorators (#153726 )" This reverts commit b7d08defe9cfe1595ff680f845b39f5e03a89555. Reverted https://github.com/pytorch/pytorch/pull/153726 on behalf of https://github.com/yangw-dev due to sorry, it seems like your pr failed typecheck error internally, [D75155486](https://www.internalfb.com/diff/D75155486) ([comment](https://github.com/pytorch/pytorch/pull/153726#issuecomment-2901911114))	2025-05-22 16:49:08 +00:00
Michael Lazos	a15550b776	[Cutlass] Use env var for EVT flag (#154099 ) Swaps out hard flag for environment variable in inductor config. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154099 Approved by: https://github.com/eellison	2025-05-22 16:36:57 +00:00
PyTorch MergeBot	a82c8891d5	Revert "[aoti] Add MPS runner and shim (#153964 )" This reverts commit 918ae5d36188f419a47f3b1315f9fb373035ed66. Reverted https://github.com/pytorch/pytorch/pull/153964 on behalf of https://github.com/angelayi due to broke frl build ([comment](https://github.com/pytorch/pytorch/pull/153964#issuecomment-2901876832))	2025-05-22 16:35:59 +00:00
PyTorch MergeBot	47a01f3efb	Revert "[aoti] Initial Metal support (#153959 )" This reverts commit 28bcd9eb30336b370298dbe9677b95019882f2a8. Reverted https://github.com/pytorch/pytorch/pull/153959 on behalf of https://github.com/angelayi due to previous PR broke frl build ([comment](https://github.com/pytorch/pytorch/pull/153959#issuecomment-2901825315))	2025-05-22 16:17:07 +00:00
Isuru Fernando	f419373dd3	[inductor] lowering for fractional_max_pool3d (#148630 ) also a lowering with a reduction for large window_sizes for fractional_max_pool2d Pull Request resolved: https://github.com/pytorch/pytorch/pull/148630 Approved by: https://github.com/eellison	2025-05-22 16:06:29 +00:00
Tom Ritchford	9a8c42ff94	Get rid of unused code in linters (#154043 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154043 Approved by: https://github.com/XuehaiPan, https://github.com/Skylion007	2025-05-22 15:24:54 +00:00
eellison	35ddad284d	update mutation renames (#153895 ) Thanks to @PaulZhang12 for original find. When we finalize a multi template buffer, we need to reflect mutation renaming in dependencies. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153895 Approved by: https://github.com/PaulZhang12	2025-05-22 14:54:39 +00:00
Huy Do	6cd9d66b7f	Allow higher fp16 tolerance for phlippe_resnet on CUDA 12.8 (#154109 ) After https://github.com/pytorch/pytorch/pull/154004, one of the model `phlippe_resnet` needs higher tolerance for fp16 on CUDA 12.8. I can reproduce it locally with: ``` python benchmarks/dynamo/torchbench.py --accuracy --timing --explain --print-compilation-time --inductor --device cuda --training --amp --only phlippe_resnet E0522 02:47:12.392000 2130213 site-packages/torch/_dynamo/utils.py:2949] RMSE (res-fp64): 0.00144, (ref-fp64): 0.00036 and shape=torch.Size([]). res.dtype: torch.float32, multiplier: 3.000000, tol: 0.001000, use_larger_multiplier_for_smaller_tensor: 0 ``` I'm not sure what exactly happens behind the scene, but this should help fix the CI failure. Also remove some left over expected accuracy results for CUDA 12.4 which we are not using anymore on CI for benchmark jobs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154109 Approved by: https://github.com/Skylion007, https://github.com/malfet	2025-05-22 14:25:12 +00:00
IvanKobzarev	4439255148	[aotd] Support saved tensors hooks in aot_autograd (#150032 ) https://github.com/pytorch/pytorch/issues/148222 Goal: At the moment autograd saved tensors hooks are run in eager after compiled forward. They are executed at the same time for all saved tensors. Hooks can be used to reduce amout of memory used for saved tensors, doing quantization or offloading to cpu. This is suboptimal for optimization of peak memory. Better solution will be to put the hooks in the graph, as close as possible to the last usage of the tensor. To get user specified autograd saved tensors hooks in the graph. Logic: UX: If user specifies with torch.autograd.graph.saved_tensors_hooks(pack_gm, unpack_gm). Where pack_gm and unpack_gm are torch.fx.GraphModule. Then AotAutograd will retrace those graph modules, doing decompositions and functionalization in aot_autograd, inlining the result graphs in forward epilogue and backward prologue. User may want to use control logic in the hooks, for example applying quantization only for specific dtypes and sizes. This is also possible, user can put it into torch.fx.wrap function and use symbolic trace to make a GraphModule. In that case AotAutograd cahing will work only in case when user explicitly set to the torch.fx.wrap call_function node "user_cache_hash" metadata. If this metadata set - then aot_autograd cache can use saved cache artifact. If metadata is not set - then cache is bypassed. Dynamo: Dynamo traces pack and unpack hooks and installs them as subgraph and explicitly adds to the output_graph. (As those subgraphs are not used and will not be copied in the result by default). The complexity here is that at this moment we do not have example of inputs for the hooks. We trace pack_hook with some Tensor from the inputs. The result subgraphs are added to the hashing of AotAutograd Cache. In AotAutograd we retrace the graph with the true saved tensors coming from partitioner. Backwards Compatibility: As current hooks are executed in eager mode and not all of them will be traceable - we only try to put in the graph hooks, explicitly marked by user with annotation (@_inlineable_saved_tensors_hooks). For other hooks or if compiled autograd is enabled - keep the same logic. Recompilations: Hooks are guarded with lambda guard matching function id to cause recompilation if user reruns compiled function. Aot_autograd: After partitioner prepared forward and backward module - we trace prepared at Dynamo graphs for pack and unpack hooks and inline them in epilogue of forward and prologue of backward. Forward outputs and backward inputs are changed, transparently for user. We do not try to put it close the last usage etc., relying on inductor to do this optimization. ``` INFO: TRACED GRAPH ===== Forward graph pre saved_tensors_hooks inlining 3 ===== /data/users/ivankobzarev/a/pytorch/torch/fx/_lazy_graph_module.py class GraphModule(torch.nn.Module): def forward(self, primals_1: "Sym(s0)", primals_2: "Sym(s1)", primals_3: "f32[s0, s1][s1, 1]cuda:0"): # File: /data/users/ivankobzarev/a/pytorch/test/functorch/test_aotdispatch.py:6660 in simple_fn, code: x = x + 1 add: "f32[s0, s1][s1, 1]cuda:0" = torch.ops.aten.add.Tensor(primals_3, 1); primals_3 = None # File: /data/users/ivankobzarev/a/pytorch/test/functorch/test_aotdispatch.py:6661 in simple_fn, code: x = SAF.apply(x) view: "f32[s0, s1][s1, 1]cuda:0" = torch.ops.aten.view.default(add, [primals_1, primals_2]) return (view, add, primals_1, primals_2) INFO: TRACED GRAPH ===== Backward graph pre saved_tensors_hooks inlining 3 ===== /data/users/ivankobzarev/a/pytorch/torch/fx/_lazy_graph_module.py class GraphModule(torch.nn.Module): def forward(self, primals_1: "Sym(s0)", primals_2: "Sym(s1)", primals_3: "f32[s0, s1][s1, 1]cuda:0"): # File: /data/users/ivankobzarev/a/pytorch/test/functorch/test_aotdispatch.py:6660 in simple_fn, code: x = x + 1 add: "f32[s0, s1][s1, 1]cuda:0" = torch.ops.aten.add.Tensor(primals_3, 1); primals_3 = None # File: /data/users/ivankobzarev/a/pytorch/test/functorch/test_aotdispatch.py:6661 in simple_fn, code: x = SAF.apply(x) view: "f32[s0, s1][s1, 1]cuda:0" = torch.ops.aten.view.default(add, [primals_1, primals_2]) return (view, add, primals_1, primals_2) INFO: TRACED GRAPH ===== saved_tensors_pack_hook add 3 ===== /data/users/ivankobzarev/a/pytorch/torch/fx/_lazy_graph_module.py class pack_float8(torch.nn.Module): def forward(self, x_1: "f32[s0, s1][s1, 1]cuda:0"): # No stacktrace found for following nodes _to_copy: "f8e4m3fn[s0, s1][s1, 1]cuda:0" = torch.ops.aten._to_copy.default(x_1, dtype = torch.float8_e4m3fn); x_1 = None return (torch.float32, _to_copy) INFO: TRACED GRAPH ===== saved_tensors_unpack_hook add 3 ===== <eval_with_key>.22 from /data/users/ivankobzarev/a/pytorch/torch/fx/experimental/proxy_tensor.py:1225 in wrapped class pack_float8(torch.nn.Module): def forward(self, x_1: "f32[s0, s1][s1, 1]cuda:0"): # No stacktrace found for following nodes _to_copy: "f8e4m3fn[s0, s1][s1, 1]cuda:0" = torch.ops.aten._to_copy.default(x_1, dtype = torch.float8_e4m3fn); x_1 = None return (torch.float32, _to_copy) INFO: TRACED GRAPH ===== Forward graph 3 ===== /data/users/ivankobzarev/a/pytorch/torch/fx/_lazy_graph_module.py class GraphModule(torch.nn.Module): def forward(self, primals_1: "Sym(s0)", primals_2: "Sym(s1)", primals_3: "f32[s0, s1][s1, 1]cuda:0"): # File: /data/users/ivankobzarev/a/pytorch/test/functorch/test_aotdispatch.py:6660 in simple_fn, code: x = x + 1 add: "f32[s0, s1][s1, 1]cuda:0" = torch.ops.aten.add.Tensor(primals_3, 1); primals_3 = None # No stacktrace found for following nodes _to_copy: "f8e4m3fn[s0, s1][s1, 1]cuda:0" = torch.ops.aten._to_copy.default(add, dtype = torch.float8_e4m3fn) # File: /data/users/ivankobzarev/a/pytorch/test/functorch/test_aotdispatch.py:6661 in simple_fn, code: x = SAF.apply(x) view: "f32[s0, s1][s1, 1]cuda:0" = torch.ops.aten.view.default(add, [primals_1, primals_2]); add = None return (view, _to_copy, primals_1, primals_2) INFO: TRACED GRAPH ===== Backward graph 3 ===== <eval_with_key>.21 class GraphModule(torch.nn.Module): def forward(self, primals_1: "Sym(s0)", primals_2: "Sym(s1)", add_packed_2: "f8e4m3fn[s0, s1][s1, 1]cuda:0", tangents_1: "f32[s0, s1][s1, 1]cuda:0"): # No stacktrace found for following nodes _to_copy: "f32[s0, s1][s1, 1]cuda:0" = torch.ops.aten._to_copy.default(add_packed_2, dtype = torch.float32); add_packed_2 = None # File: /data/users/ivankobzarev/a/pytorch/test/functorch/test_aotdispatch.py:6661 in simple_fn, code: x = SAF.apply(x) add_7: "f32[s0, s1][s1, 1]cuda:0" = torch.ops.aten.add.Tensor(tangents_1, _to_copy); tangents_1 = _to_copy = None return (None, None, add_7) ``` Differential Revision: [D72187044](https://our.internmc.facebook.com/intern/diff/D72187044) Pull Request resolved: https://github.com/pytorch/pytorch/pull/150032 Approved by: https://github.com/bdhirsh	2025-05-22 14:09:38 +00:00
zeshengzong	f12d8d60b1	Add hint message when parameters is empty in clip_grad_norm_ (#151529 ) Fixes #148259 ## Changes - Add print warning message when `parameters` generator exhausted ## Test Result ### print warning ```python import torch import torch.nn as nn import torch.optim as optim class SimpleModel(nn.Module): def __init__(self): super(SimpleModel, self).__init__() self.fc = nn.Linear(10, 1) def forward(self, x): return self.fc(x) model = SimpleModel() criterion = nn.MSELoss() optimizer = optim.SGD(model.parameters(), lr=0.01) inputs = torch.randn(16, 10) targets = torch.randn(16, 1) outputs = model(inputs) loss = criterion(outputs, targets) optimizer.zero_grad() loss.backward() params_to_clip = model.parameters() for p in params_to_clip: print(p.shape) max_norm = 1.0 norm_type = 2.0 total_norm = nn.utils.clip_grad_norm_(params_to_clip, max_norm, norm_type) print(f"total_norm: {total_norm}") ``` ```bash /home/zong/code/pytorch/torch/nn/utils/clip_grad.py:222: UserWarning: `parameters` is an empty generator, no gradient clipping will occur. warnings.warn( total_norm: 0.0 ``` ### UT ```bash pytest test/test_nn.py -k test_clip_grad_norm ``` ![image](https://github.com/user-attachments/assets/0aa0f06c-e0a5-43cf-9a97-d7c2747c9180) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151529 Approved by: https://github.com/jbschlosser	2025-05-22 11:23:39 +00:00
leslie-fang-intel	40e6ca24ef	Update CPU Inductor merge rules by adding more CPP Template (#152086 ) Summary Add more CPP Template into the CPU Inductor merge rules. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152086 Approved by: https://github.com/atalman	2025-05-22 09:46:26 +00:00
Aleksei Nikiforov	2f57ee579d	S390x update docker image (#153619 ) Add ninja-build for pytorch tests. Switch to gcc 14 due to fix for precompiled headers and s390x vectorization interaction. Disable -Werror when building onnxruntime. Pin onnx version. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153619 Approved by: https://github.com/huydhn	2025-05-22 09:34:46 +00:00
zeshengzong	d7a83ab67b	Fix `lr_scheduler` unexpectedly calls `step()` when init argument last_epoch is larger than -1 (#149312 ) Fixes #102261 ## Changes - Use flag `_is_initial` to replace `self.last_epoch == 0` condition to judge whether `lr` should be initial value - Add test for `ExponentialLR` checkpoint usecase ## Test Result ```python pytest -s test/optim/test_lrscheduler.py -vv ``` ![image](https://github.com/user-attachments/assets/6fd32bcc-b4fb-4421-b891-620bd4900dc1) Pull Request resolved: https://github.com/pytorch/pytorch/pull/149312 Approved by: https://github.com/janeyx99 Co-authored-by: Jane (Yuan) Xu <31798555+janeyx99@users.noreply.github.com>	2025-05-22 08:42:37 +00:00
Michael Lazos	423fc671e9	[Cutlass] Support float8_e4m3fn GEMM (#153890 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153890 Approved by: https://github.com/drisspg, https://github.com/eellison	2025-05-22 08:37:33 +00:00
Sidharth	c1b7dbc52a	[dynamo] unimplemented -> unimplemented_v2 in variables/dict.py (#154040 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154040 Approved by: https://github.com/williamwen42, https://github.com/StrongerXi	2025-05-22 06:46:10 +00:00
Yu, Guangye	a664cfdf95	Add C10_NODEPRECATED check for xpu (#153935 ) # Motivation Add `C10_NODEPRECATED` check for XPU. This doesn't allow xpu codebase to use `c10::optional`. What's the change about torch-xpu-ops commit update? Deprecate `c10::optional`, `c10::nullopt`, `c10::make_option`, use the counterpart in std instead. # Additional Context This PR depends on https://github.com/intel/torch-xpu-ops/pull/1683 https://github.com/intel/torch-xpu-ops/pull/1690 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153935 Approved by: https://github.com/Skylion007, https://github.com/cyyever	2025-05-22 06:44:04 +00:00
Malaika	482e5b6660	[inductor] Added precompilation_timeout_seconds into a config instead of hardcoded (#153788 ) Fixes #153392 - Updated config.py to add the timeout as a config var to be tuned dynamically (default is 3600s). - Passed the var as a kwarg during call on instance. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153788 Approved by: https://github.com/henrylhtsang	2025-05-22 06:44:02 +00:00
Wei Wang	7128b50a65	[CI][CUDA][Distributed] Move cuda 11.8 distributed pull jobs to cuda 12.6 (#151594 ) This PR moves distributed cuda CI job from cuda 11.8 to cuda 12.6. In doing so, a few unit test failures were exposed, some if not all of which would take a while to root-cause and fix, so temporarily skip them after creating the issues. https://github.com/pytorch/pytorch/issues/153479 test_nan_assert tricky behavior (e.g. skip_but_pass_in_sandcastle, ubuntu 20.04 does not work, ubuntu 22.04 works, Amazon Linux 2023 skip - what is Sandcastle OS?) https://github.com/pytorch/pytorch/issues/153122 CUDA context related https://github.com/pytorch/pytorch/issues/153517 NCCL regression, future NCCL may fix it https://github.com/pytorch/pytorch/issues/154073 skip test_symmetric_memory for cuda 12.6 before it is fixed See: https://github.com/pytorch/pytorch/issues/147383 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151594 Approved by: https://github.com/eqy, https://github.com/atalman, https://github.com/cyyever, https://github.com/huydhn, https://github.com/kwen2501	2025-05-22 06:33:29 +00:00
Laith Sakka	4bcff4af99	Move prologue_supported_inputs computations to def_kernal (#150869 ) This avoid replaying load_input on a cache hit on the generate_code_cache. the idea is that if a template have prologue_loads_all_inputs = True, it means that all all inputs are loaded and hence no need to replay Effect on the current benchmark on a local run on dev server. 18549985383 -> 15072230073 25697270062 -> 20738613297 Pull Request resolved: https://github.com/pytorch/pytorch/pull/150869 Approved by: https://github.com/eellison	2025-05-22 06:24:44 +00:00
Colin L Reliability Rice	4421aee558	torch.compile: Supress stdout / stderr output from subprocesses when local (#153837 ) Summary: This output is extremely noisy - i.e. on a 96 core machine, with 8 ranks, you can get ~700 duplicate set of logs from each worker. Differential Revision: D74907920 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153837 Approved by: https://github.com/aorenste, https://github.com/masnesral	2025-05-22 05:49:43 +00:00
soulitzer	f2af30fee5	Add a HOP to bypass tracing of a wrapper function while tracing the wrapped function (#153487 ) Usage: ```python from torch._higher_order_ops.wrap import dynamo_bypassing_wrapper # Your ordinary function wrapper def my_hop_fn_impl(fn, args, k=1, kwargs): def wrapper(args, *kwargs): out = fn(args, *kwargs) if isinstance(out, tuple): return (out[0] + k,) return out + k return wrapper # Calling `my_hop_fn` instead of the impl directly captures a HOP into the dynamo graph def my_hop_fn(fn, args, k=1, *kwargs): return dynamo_bypassing_wrapper( functools.partial(my_hop_fn_impl, k=k), fn, args, **kwargs ) ``` Notes: - The dynamo captured graph now stashes arbitrary callable objects (the wrapper_fn) - this is equivalent to what SAC does today with policy_fn. - The `wrapper_fn` passed to `dynamo_bypassing_wrapper ` should have signature `Callable -> Callable` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153487 Approved by: https://github.com/ydwu4	2025-05-22 04:24:38 +00:00
Boyuan Feng	669b176d4c	[Graph Partition] support removed arguments, NoneLayout, and mutation (#153899 ) Graph partition relies on `read_writes` to collect partition inputs and outputs. There are three edge cases: 1. `NoneLayout` is not allocated so it cannot become a partition input or output. 2. Codegen may decide a buffer to be internal to a kernel (e.g., triton kernel). One example is some buffers internal to a FusedSchedulerNode. These buffers are never actually allocated as `buf_id`. 3. We should use mutation_real_name for graph partition inputs and outputs to match the behavior of other codegen. This PR supports these 3 cases. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153899 Approved by: https://github.com/eellison	2025-05-22 04:24:31 +00:00
Yidi Wu	d1fe198df6	[cond] support output the same unbacked symbol from two branches (#148206 ) Previously, we didn't track the unbacked symbols leaked out of true_branch and false_branch if they have the same shape expr. This cause the the fake output of cond operator itself doesn't set up its unbacked_bindings meta properly (because they're ignored). In this PR, we also check whether there're leaked out unbacked symbols and create new unbacked symbols for it and track it as output of cond. Pull Request resolved: https://github.com/pytorch/pytorch/pull/148206 Approved by: https://github.com/zou3519	2025-05-22 03:39:43 +00:00
Colin Peppler	fe285b9560	[aoti] fix corner case in unbacked replacements for atomically_apply_size_hint (#153768 ) ## PR There are a few cases that my previous PR (#153220) didn't cover. 1. The LHS/RHS matters. Today, if you do `torch._check(lhs == rhs)` then it will show up as a deferred runtime assert with `Eq(lhs, rhs)`. 2. There can be transitive replacements. For example, expr1 -> expr2 -> u0. `test_size_with_unbacked_add_expr_transitive` tests for this. 3. An unbacked symint expr may not have a replacement that's purely a symbol, for instance, it could be another expression. `test_size_with_unbacked_add_and_mul_expr` tests for this. ## Device assertion msg ``` /tmp/tmp07mu50tx/6y/c6ym2jzadwfigu3yexredb7qofviusz3p7ozcdjywvayhxgcqxkp.py:40: unknown: block: [8681,0,0], thread: [4,0,0] Assertion `index out of bounds: 0 <= tl.broadcast_to(tmp13, [XBLOCK]) < ks0` failed. ... /tmp/tmp07mu50tx/6y/c6ym2jzadwfigu3yexredb7qofviusz3p7ozcdjywvayhxgcqxkp.py:40: unknown: block: [8681,0,0], thread: [6,0,0] Assertion `index out of bounds: 0 <= tl.broadcast_to(tmp13, [XBLOCK]) < ks0` failed. ``` ## Autotuning code setup This is the autotuning code for a concat kernel which takes input tensors (`in_buf`) and writes them to the (`out_buf`). It's important to note the size of `in_buf0` is the same as `in_buf1` don't match along dim=0. This is bad because all concat inputs must share the same size for each dim except for the concat dim (here that's dim=1). ``` in_buf0 = generate_example_value(size=(u1 + s0, 256)) # concrete size is (17900, 256) in_buf1 = generate_example_value(size=(u0, 10)) # concrete size is (8192, 10) ... out_buf = generate_example_value(size=(u1 + s0, 266)) # concrete size is (17900, 256+10) triton_poi_fused_cat_1.run(in_buf0, in_buf1, ..., out_buf, xnumel=(u1 + s0) * 266 ...) ``` If we look into the kernel code, you'll see that `tmp9` loads `in_buf1` (our incorrectly shaped input tensor). There is also a mask to prevent OOB loads. - `tmp6` makes sure we're only loading with the `xindex` from 256 to 264. - `xmask` makes sure we're only loading with the `xindex` within `xnumel`. - `tmp6 & xmask` together is essentially checking `0 ≤ x0 < u1 + s0` and `256 ≤ x1 < 264`. The mask logic is correct, however, `in_buf1` has the shape `[8192, 10]` this means any load where `8192 ≤ x0 < u1 + s0` will be an OOB load. ``` def triton_poi_fused_cat_1(in_buf0, in_buf1, ... out_buf, xnumel, XBLOCK): xoffset = tl.program_id(0) * XBLOCK xindex = xoffset + tl.arange(0, XBLOCK) xmask = xindex < xnumel x0 = (xindex % 264) x1 = xindex // 264 ... tmp6 = x0 >= tl.full([1], value=256) tmp9 = tl.load(in_buf1 + (x1), tmp6 & xmask) # device assertion is thrown here tl.device_assert(((0 <= tl.broadcast_to(tmp13, [XBLOCK])) & (tl.broadcast_to(tmp13, [XBLOCK]) < ks0)) \| ~(xmask & tmp6), "index out of bounds: 0 <= tl.broadcast_to(tmp13, [XBLOCK]) < ks0") ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153768 Approved by: https://github.com/jingsh	2025-05-22 02:05:37 +00:00
Jiang, Yanbing	a264af8c71	Support fp8 output of _scaled_mm for CPU (#153600 ) This PR is to support fp8 output of torch._scaled_mm for CPU, and create related UTs with fp8 and bf16/fp16/fp32 output. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153600 Approved by: https://github.com/leslie-fang-intel, https://github.com/mingfeima, https://github.com/jansel	2025-05-22 01:15:39 +00:00
Gabriel Ferns	254293b777	Add `flag _metrics_log_runtime` to disable runtime metric logging by default (#153506 ) https://github.com/pytorch/pytorch/pull/152708 expanded support of `get_estimated_runtime` to many more types of `SchedulerNodes`. This caused an increase in compile time because we're always calling `get_estimated_runtime` to populate the metrics table. This PR adds a flag for this logging, which reduces the instruction count by 8%. Long term, we should probably merge metrics.py with TORCH_LOGS/tlparse (suggestion from @xmfan). Update: added support for TORCH_LOGS for the metrics logging. Test Plan: mm_loop.py and many existing tests cover. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153506 Approved by: https://github.com/eellison	2025-05-22 01:02:11 +00:00
PyTorch MergeBot	261897734a	Revert "cpp_wrapper: build non-performance-sensitive code at O1 (#148773 )" This reverts commit 3c89cfd46075e62c1725b43557612901a9cbb6fa. Reverted https://github.com/pytorch/pytorch/pull/148773 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it seems that pr_time_benchmark is regressed after this land ([comment](https://github.com/pytorch/pytorch/pull/148773#issuecomment-2899545140))	2025-05-22 00:11:14 +00:00
Max Podkorytov	7ef2c62fd3	[ROCm][Inductor][CK] Add ck-tile based universal gemm kernels to torch.mm autotune choices (#152341 ) This PR adds code generation for CK-tile based universal gemm kernels to the CK backend for Inductor, and adds these kernels to autotune choices. Unlike legacy-CK based kernels (which are generated by parsing the CK instances from CK library), we generate the set of instances by manually specifying the tuning parameters. This PR introduces a new template for code generation, and compilation/autotuning is handled by the existing infrastructure. Points of discussion: * For simplicity and reduced coupling with CK, the instance filter checks only data type and layout, and doesn't check the alignment requirement - meaning that more instances will be compiled than necessary - while keeping the code generation independent from internal CK logic which checks the alignment validity at runtime * CK-tile instances are enabled whenever legacy-CK instances are enabled. A config knob could be introduced to differentiate between the instance types if that's needed * Whether gemm problem size K is ever dynamic, since whenever it's not a compile-time constant, we need to perform a runtime dispatch between several kernels Testing Use the existing tests in `test/inductor/test_ck_backend.py` Pull Request resolved: https://github.com/pytorch/pytorch/pull/152341 Approved by: https://github.com/chenyang78	2025-05-21 23:59:16 +00:00
Ke Wen	87fc5af1f6	[c10d] Turn off default non-blocking API mode to work around hang in NCCL 2.26 (#154055 ) Work around issues like #153960, #152623 NCCL 2.26 seems to introduce random hang in non-blocking API mode. This PR opts out of non-blocking mode to work around it. Previously torch turned it on by default in eager init (i.e. `device_id` passed) to avoid init overhead. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154055 Approved by: https://github.com/atalman	2025-05-21 23:46:52 +00:00
Simon Fan	fae6f6c9ca	[aot] fix deepcopying of aot bwd containing real tensors (#153999 ) Previously when we lower backward AOT due to symints, the post grad passes would leave the bw_module in a non-runnable state. This caused issues when compiled autograd tried to trace at runtime. So we had inductor operate on a deepcopy of bw_module. But with https://github.com/pytorch/pytorch/issues/153993, we see that deepcopying real tensors will fail under fake mode due to the device type mismatch between the fake tensors ("meta" device) and the real tensor. So by disabling fake mode, we avoid these errors. This change is a strict improvement over current, but it does reveal that this deepcopy can theoretically cause OOMs. FIXES https://github.com/pytorch/pytorch/issues/153993 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153999 Approved by: https://github.com/jamesjwu, https://github.com/bdhirsh	2025-05-21 23:30:02 +00:00
William Wen	67f9feeee7	remove TestCustomOp.test_impl_device_cpu from dynamo expected failures (#154049 ) Fixes https://github.com/pytorch/pytorch/issues/153763 maybe? Pull Request resolved: https://github.com/pytorch/pytorch/pull/154049 Approved by: https://github.com/StrongerXi	2025-05-21 23:20:30 +00:00
Catherine Lee	5ee1242310	Follow up to #152209 , remove compat patch after docker image rename (#152958 ) Remove compat patch that lets PRs that haven't rebased base #152209 still have docker images. Merge ~2 weeks after the above PR was merged. ~80% of PRs have a merge base that is <2 weeks old Pull Request resolved: https://github.com/pytorch/pytorch/pull/152958 Approved by: https://github.com/huydhn	2025-05-21 23:11:29 +00:00
ishanjmukherjee	d82610c2af	docs: fix "should not to be" typo in `register_buffer` docstring (#153817 ) Corrects a small grammatical error in `register_buffer` docstring, from "... should not to be ..." to "... should not be ...". Docs-only change, so no runtime behavior, tests, or APIs are affected. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153817 Approved by: https://github.com/mikaylagawarecki	2025-05-21 22:46:50 +00:00
yang	b967b7b11e	Update rnn.py, fix `torch.nn.RNN` document error (#153620 ) I found the same issue as #147490 (@jibril-b-coulibaly). There's an equivalent in the [doc-string](https://docs.pytorch.org/docs/stable/generated/torch.nn.RNN.html#rnn) of `torch.nn.RNN`: ```python # Efficient implementation equivalent to the following with bidirectional=False def forward(x, hx=None): if batch_first: x = x.transpose(0, 1) seq_len, batch_size, _ = x.size() if hx is None: hx = torch.zeros(num_layers, batch_size, hidden_size) h_t_minus_1 = hx h_t = hx output = [] for t in range(seq_len): for layer in range(num_layers): h_t[layer] = torch.tanh( x[t] @ weight_ih[layer].T + bias_ih[layer] + h_t_minus_1[layer] @ weight_hh[layer].T + bias_hh[layer] ) output.append(h_t[-1]) h_t_minus_1 = h_t output = torch.stack(output) if batch_first: output = output.transpose(0, 1) return output, h_t ``` However there's something wrong. 1. Like mentioned in #147490, line 499 is wrong `fb55bac3de/torch/nn/modules/rnn.py (L499)` The input for RNNCell should be different for different layers. 2. The code contains several hidden reference-related issues that may result in unintended modifications to tensors. For example in line 504, this causes all elements in the final output list to point to the same tensor. `fb55bac3de/torch/nn/modules/rnn.py (L504)` 3. Some variable is not defined. Despite being a relatively minor issue in annotation, it can lead to significant confusion for those who are new to the concept. For example `weight_ih` in line 499 `fb55bac3de/torch/nn/modules/rnn.py (L499)` So, i write a runnable version to make it more clear: ```python # Efficient implementation equivalent to the following with bidirectional=False rnn = nn.RNN(input_size, hidden_size, num_layers) params = dict(rnn.named_parameters()) def forward(x, hx=None, batch_first=False): if batch_first: x = x.transpose(0, 1) seq_len, batch_size, _ = x.size() if hx is None: hx = torch.zeros(rnn.num_layers, batch_size, rnn.hidden_size) h_t_minus_1 = hx.clone() h_t = hx.clone() output = [] for t in range(seq_len): for layer in range(rnn.num_layers): input_t = x[t] if layer == 0 else h_t[layer - 1] h_t[layer] = torch.tanh( input_t @ params[f"weight_ih_l{layer}"].T + h_t_minus_1[layer] @ params[f"weight_hh_l{layer}"].T + params[f"bias_hh_l{layer}"] + params[f"bias_ih_l{layer}"] ) output.append(h_t[-1].clone()) h_t_minus_1 = h_t.clone() output = torch.stack(output) if batch_first: output = output.transpose(0, 1) return output, h_t ``` This code can reproduce the computation of torch.nn.RNN. For example: ```python import torch import torch.nn as nn torch.manual_seed(0) input_size, hidden_size, num_layers = 3, 5, 2 rnn = nn.RNN(input_size, hidden_size, num_layers) params = dict(rnn.named_parameters()) x = torch.randn(10, 4, 3) official_imp = rnn(x) my_imp = forward(x) assert torch.allclose(official_imp[0], my_imp[0]) assert torch.allclose(official_imp[1], my_imp[1]) ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153620 Approved by: https://github.com/mikaylagawarecki	2025-05-21 22:45:28 +00:00
Bin Bao	5b6e551c0f	[AOTI][refactor] Fix an anonymous namespace issue (#154033 ) Summary: Remove anonymous namespace in model_container.h to fix the following compiler warning, ``` warning: ‘torch::aot_inductor::AOTInductorModelContainer’ has a field ‘torch::aot_inductor::AOTInductorModelContainer::constant_folded_’ whose type uses the anonymous namespace [-Wsubobject-linkage] ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154033 Approved by: https://github.com/chenyang78	2025-05-21 22:29:09 +00:00
Yidi Wu	d356ca2466	[map] add inductor support by lowering to while_loop (#150971 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/150971 Approved by: https://github.com/zou3519 ghstack dependencies: #151034	2025-05-21 22:19:47 +00:00
Yidi Wu	cf1b38a017	[map] make proxy mode re-dispatch to fake key (#151034 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151034 Approved by: https://github.com/zou3519	2025-05-21 22:19:47 +00:00
Huy Do	c13eeaa718	Move inductor benchmark jobs to CUDA 12.8 (#154004 ) For benchmark jobs, we usually want to run with the latest support CUDA version to get the best performance. This is a request coming from NVIDIA where they are running inductor benchmarks on Blackwell with CUDA 12.8 (min support version) and looking for an apples-to-apples comparison. This also clean up references to CUDA 12.4 which have been sunset in PyTorch CI. ### Testing - H100 benchmark https://github.com/pytorch/pytorch/actions/runs/15151424588 - Micro benchmark https://github.com/pytorch/pytorch/actions/runs/15151445957 (I just realize that this is still running on A100, @yanboliang Do you want to run on H100 now that we have capacity there? It would also solve the problem of GPU memory) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154004 Approved by: https://github.com/atalman, https://github.com/ZainRizvi, https://github.com/seemethere	2025-05-21 22:17:10 +00:00
henrylhtsang	053ca7439a	[cutlass backend] Add serializer for cutlass ops (#153894 ) Differential Revision: [D74524786](https://our.internmc.facebook.com/intern/diff/D74524786/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153894 Approved by: https://github.com/ColinPeppler, https://github.com/mlazos	2025-05-21 22:01:40 +00:00
Natalia Gimelshein	401fa87ace	make only current thread allocate to pool in NcclPG (#153990 ) follow up to #153356 that fixes nccl allocation to pool Pull Request resolved: https://github.com/pytorch/pytorch/pull/153990 Approved by: https://github.com/kwen2501	2025-05-21 21:57:37 +00:00
angelayi	28bcd9eb30	[aoti] Initial Metal support (#153959 ) An example generated file: P1816629015 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153959 Approved by: https://github.com/malfet, https://github.com/desertfire ghstack dependencies: #153964	2025-05-21 21:55:59 +00:00
angelayi	918ae5d361	[aoti] Add MPS runner and shim (#153964 ) Added AOTIModelContainerRunnerMps and a shim for mps fallback ops. I also added a mps-specific shim which contains one operator, which will be used to set arguments being passed to the Metal kernel: ``` AOTI_TORCH_EXPORT AOTITorchError aoti_torch_mps_set_arg( AOTIMetalKernelFunctionHandle func, unsigned idx, AtenTensorHandle tensor); ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153964 Approved by: https://github.com/malfet, https://github.com/desertfire	2025-05-21 21:55:59 +00:00
Sidharth	0b79a8c1a9	[dynamo] renamed _fn for more clarity and put a comment of user compiler user (#154026 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154026 Approved by: https://github.com/williamwen42, https://github.com/StrongerXi	2025-05-21 21:12:51 +00:00
Scott Todd	0e5f2339d0	[ROCm][Windows] Run hipcc with compatibility flags. (#153986 ) See also https://github.com/ROCm/TheRock/issues/590. Including the `-Wno-ignored-attributes` flag here avoids 700MB of log warning spam while compiling and the `-fms-extensions` seems beneficial to include: https://clang.llvm.org/docs/MSVCCompatibility.html. Co-authored-by: Aaryaman Vasishta <jem456.vasishta@gmail.com> Co-authored-by: Scott Todd <scott.todd0@gmail.com> Pull Request resolved: https://github.com/pytorch/pytorch/pull/153986 Approved by: https://github.com/Skylion007, https://github.com/jeffdaily Co-authored-by: Aaryaman Vasishta <jem456.vasishta@gmail.com>	2025-05-21 20:26:52 +00:00
Benjamin Glass	3c89cfd460	cpp_wrapper: build non-performance-sensitive code at O1 (#148773 ) Builds on #148212, applying the same improvements to `cpp_wrapper` mode. Benchmark results: * [A100 Benchmarks](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Wed%2C%2014%20May%202025%2015%3A10%3A05%20GMT&stopTime=Wed%2C%2021%20May%202025%2015%3A10%3A05%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cuda%20(a100)&lBranch=gh/benjaminglass1/77/orig&lCommit=ca7d0a3f16e3c511534d2cd03d695be8524570d3&rBranch=main&rCommit=1075bb37d34e483763a09c7810790d5491441e13) * [x86 Benchmarks](https://hud.pytorch.org/benchmark/compilers?dashboard=torchinductor&startTime=Wed%2C%2014%20May%202025%2015%3A10%3A05%20GMT&stopTime=Wed%2C%2021%20May%202025%2015%3A10%3A05%20GMT&granularity=hour&mode=inference&dtype=bfloat16&deviceName=cpu%20(x86)&lBranch=gh/benjaminglass1/77/orig&lCommit=ca7d0a3f16e3c511534d2cd03d695be8524570d3&rBranch=main&rCommit=1075bb37d34e483763a09c7810790d5491441e13) Pull Request resolved: https://github.com/pytorch/pytorch/pull/148773 Approved by: https://github.com/desertfire	2025-05-21 20:23:04 +00:00
Ryan Guo	4c6f0fe22f	[dynamo] Properly handle `torch.script.jit` under `@staticmethod` (#153984 ) Fixes #153607. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153984 Approved by: https://github.com/williamwen42	2025-05-21 19:45:06 +00:00
James Wu	b184e3da9c	[easy] Fix internal only test (#154035 ) Internally static cuda launcher isn't enabled, so we need to always enable it Differential Revision: [D75146584](https://our.internmc.facebook.com/intern/diff/D75146584/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/154035 Approved by: https://github.com/Skylion007 ghstack dependencies: #153565	2025-05-21 19:00:55 +00:00
Yidi Wu	8e6e79fc1b	[hop_schema] support gen_schema for invoke_subgraph (#152984 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152984 Approved by: https://github.com/zou3519 ghstack dependencies: #151067, #152974	2025-05-21 18:55:46 +00:00
Yidi Wu	9c33899196	[hop_schema] add HopSchemaGenerator to make it easier to create hop schema (#152974 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152974 Approved by: https://github.com/zou3519 ghstack dependencies: #151067	2025-05-21 18:55:46 +00:00
Yidi Wu	1e0f19e173	auto functionalize base_hop (#151067 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/151067 Approved by: https://github.com/zou3519	2025-05-21 18:55:46 +00:00
Laith Sakka	11c0ffefcd	Cache code generation during triton template expansion and enable it for mm_template. (#151773 ) In a model, we see ~~ 40% of the time in mm/addmm tuning. The model have 2000 mm, many of which receives the same input shapes. with autotune enabled, this become expensive, while we already cache auto tuning results, we did not used to cache the generation of the python code and the loading for each config that we autotune on. This diff handles the code generation part (template expansions) a previous diff handled the loading part. This is expected to save 20% of the model I am working on. How do we do the caching? For a given configurations and input layout, the generated code is always the same. One caveat is that some other information collected during code generation are input dependent (namely depends on inputs names and symbol names in inputs). and not just layout. ! To handle those we use a record and replay approach, where we record the functions that are called during code generation that effect those outputs and replay them at a cache hit. Effect on the current benchmark on a local run on dev server. mm_loop. 24115830838 -> 18362098019 mm_loop_dynamic 30506097176-> 25697270062 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151773 Approved by: https://github.com/eellison	2025-05-21 18:55:41 +00:00
Deysha Rivera	aec7cc60d7	add graph_code_verbose_log artifact for fx passes (#153775 ) Fixes #153646 This PR refactors the logging behavior in the FX pass insert_deferred_runtime_asserts and runtime_assert.py to separate verbose/intermediate graph logs from the final output graph log. All verbose logs generated during the FX pass are now routed to a new artifact logger, graph_code_verbose, while only the final output graph remains logged to the original graph_code artifact. Changes - Added a new artifact logger: [graph_code_log = torch._logging.getArtifactLogger(__name__, "graph_code_verbose")] - Updated all verbose/intermediate FX pass logs in [insert_deferred_runtime_asserts] to use the new graph_code_verbose artifact. - Ensured that only the final output graph is logged to the original graph_code artifact. - No changes to the FX pass logic or output—only logging behavior is affected. Notes This change is backward-compatible and does not affect the functional behavior of FX passes. No changes to user-facing APIs. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153775 Approved by: https://github.com/williamwen42	2025-05-21 18:31:59 +00:00
Aleksei Nikiforov	d23f4ae7b5	s390x: use qemu issue workaround for runner registration too (#154030 ) s390x: use qemu issue workaround for runner registration too. Pull Request resolved: https://github.com/pytorch/pytorch/pull/154030 Approved by: https://github.com/seemethere	2025-05-21 18:30:25 +00:00
Tomasz Bohutyn	bb7e30c165	[MegaCache] Make MegaCache generic to allow external plugins registration (#152977 ) Implements #152976 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152977 Approved by: https://github.com/oulgen	2025-05-21 18:18:47 +00:00
James Wu	c31e239910	[precompile] Add BundledAOTAutogradCacheEntry (#152840 ) Finally, this PR adds BundledAOTAutogradCacheEntry. A BundledAOTAutogradCacheEntry is an AOTAutogradCacheEntry that saves the entire CompiledFxGraph directly in the entry. This has some advantages: - No more dependency on FxGraphCache at all - Clearing FxGraphCache does not result in AOTAutogradCache miss - Simpler logic, as BundledAOTAutogradCacheEntry has everything you need to load a full compiled python wrapper from a dynamo output We plan to use BundledAOTAutogradCacheEntry for precompile. There's also a question of whether we want to use it for regular caching — the main disadvantage of this is having to save the same CompiledFxGraph twice, once in Inductor cache and once for AOTAutogradCache. With MegaCaching, this could be a regression in total cache size (as well as a minor cold start regression, as you have to save the same graph twice). I will import this and measure the mega cache space complexity, and if it looks good I'll enable it by default for caching as well. On warm start, if AOTAutogradCache hits, you won't have to load inductor at all, so warm start overhead should be unaffected. Differential Revision: [D74593304](https://our.internmc.facebook.com/intern/diff/D74593304) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152840 Approved by: https://github.com/zhxchen17	2025-05-21 18:08:42 +00:00
PyTorch MergeBot	3eb8fa081a	Revert "[3/n][Optimus][Auto-AC] Support float8_e4m3fn quantization type and set scaling as the default (#153802 )" This reverts commit 32b1baa981fe53d13d77acbee509c51087abf107. Reverted https://github.com/pytorch/pytorch/pull/153802 on behalf of https://github.com/malfet due to It breaks ROCM testing, see `d23762974e/1` ([comment](https://github.com/pytorch/pytorch/pull/153802#issuecomment-2898695702))	2025-05-21 17:20:31 +00:00
henrylhtsang	d23762974e	[inductor][cutlass backend] Add 2 stage autotuning aka prescreening (#153335 ) Motivation: By default, we are tuning the cutlass backend kernels on 3 swizzles. There are runtime params, so they share the same underlying kernel, which saves a lot of compilation time. However, autotuning all combinations of {configs} x {swizzles} is still expensive. Observations: Winner of the {configs} x {swizzles} autotuning is the same as if we do a greedy search: first find the top X winners of {configs} with swizzle 2 (hardcoded), then autotune on the {top X winner configs} x {swizzles}. In other words, we can use a Greedy algorithm to reduce autotuning time. I attach the logs below. This somewhat depends on what X is, but a number like 5-10 works pretty well from empirical observations. Logs: Baseline: https://gist.github.com/henrylhtsang/9a604f150a270dc19524f72a5d4dfac2 ``` AUTOTUNE mm(2048x2048, 2048x2048) strides: [2048, 1], [1, 2048] dtypes: torch.bfloat16, torch.bfloat16 cuda_cutlass_gemm_1776 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1777 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1778 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1800 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1801 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1802 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_9012 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_9013 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_9014 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8940 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8941 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8942 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8934 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8935 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8936 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_2001 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_2002 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_2003 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1848 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1849 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1850 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8964 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8965 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8966 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8958 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8959 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8960 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1929 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1930 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1931 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1770 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1771 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1772 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1953 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1954 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1955 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1995 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1996 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1997 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1794 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1795 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1796 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1842 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1843 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1844 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_9006 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_9007 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_9008 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1923 0.0306 ms 95.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 ``` with prescreening: ``` AUTOTUNE mm(147456x6144, 6144x2048) strides: [6144, 1], [2048, 1] dtypes: torch.bfloat16, torch.bfloat16 cutlass_1a5e81af 4.5469 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.6328 ms 98.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.6836 ms 97.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_161b8b81 4.7224 ms 96.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_161b8b81 4.7234 ms 96.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_161b8b81 4.7274 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_853b6347 4.7369 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.7404 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_161b8b81 4.7711 ms 95.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_8bc6fbda 4.8148 ms 94.4% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_8bc6fbda 4.8159 ms 94.4% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_8bc6fbda 4.8214 ms 94.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_8bc6fbda 4.8302 ms 94.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_0a1c55af 4.8487 ms 93.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_0a1c55af 4.8527 ms 93.7% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_02780d72 4.8617 ms 93.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_0a1c55af 4.8737 ms 93.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_0a1c55af 4.8738 ms 93.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_02780d72 4.9348 ms 92.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_02780d72 4.9763 ms 91.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_853b6347 4.9805 ms 91.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_1a5e81af 5.0225 ms 90.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_853b6347 5.0271 ms 90.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_02780d72 5.0595 ms 89.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_853b6347 5.1434 ms 88.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.1574 ms 88.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_1a5e81af 5.1916 ms 87.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.2018 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.2019 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_c1ffa14b 5.2037 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_1a5e81af 5.5329 ms 82.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_aa6f899c 11.5046 ms 39.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 SingleProcess AUTOTUNE benchmarking takes 1.9526 seconds and 0.0352 seconds precompiling for 32 choices ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153335 Approved by: https://github.com/eellison	2025-05-21 17:12:05 +00:00
Bin Bao	72a3c8dfa8	[AOTI][reland] Add an option to specify custom op C shim (#153968 ) Summary: Reland https://github.com/pytorch/pytorch/pull/153851 after fixing a fuzzer test issue. Add an option to tell AOTInductor codegen to generate C shim functions for certain custom ops instead of relying on ProxyExecutor. The lib that defines custom ops need to implement corresponding C shim functions. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153968 Approved by: https://github.com/hl475	2025-05-21 15:57:57 +00:00
Aaron Gokaslan	b7d08defe9	[BE]: Type previously untyped decorators (#153726 ) This fixes decorator typing which unmasks a lot of typing issues in the codebase Pull Request resolved: https://github.com/pytorch/pytorch/pull/153726 Approved by: https://github.com/albanD	2025-05-21 15:56:19 +00:00
Nikita Shulga	6c2c527cd6	[BE] Remove extra semicolons from SymmetricMemory.hpp (#154034 ) Fixes ``` In file included from /Users/malfet/git/pytorch/pytorch/torch/csrc/distributed/c10d/SymmetricMemory.cpp:1: /Users/malfet/git/pytorch/pytorch/torch/csrc/distributed/c10d/SymmetricMemory.hpp:77:4: warning: extra ';' after member function definition [-Wextra-semi] 77 \| }; \| ^ /Users/malfet/git/pytorch/pytorch/torch/csrc/distributed/c10d/SymmetricMemory.hpp:81:4: warning: extra ';' after member function definition [-Wextra-semi] 81 \| }; \| ^ 2 warnings generated. ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/154034 Approved by: https://github.com/Skylion007	2025-05-21 14:33:30 +00:00
James Wu	33767eb391	Add option to statically launch user defined triton kernels (#153725 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153725 Approved by: https://github.com/oulgen, https://github.com/Mingming-Ding, https://github.com/jansel ghstack dependencies: #153565	2025-05-21 14:33:15 +00:00
Nikita Shulga	b73d77900e	[CI] Run limited h100 tests every 6 hours (#153900 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153900 Approved by: https://github.com/Skylion007	2025-05-21 13:40:03 +00:00
Yutao Xu	11f8511455	Update torch-xpu-ops commit pin (#153902 ) Update the torch-xpu-ops commit to `defce46ae7`, includes: - Resolve the aten::gamma accuracy gap compared to scipy - Optimize layernom_vectorized_impl by using adaptive wg selection for small shapes - [Intro async flag and use current stream avoid stream sync](https://github.com/intel/torch-xpu-ops/pull/1546) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153902 Approved by: https://github.com/Skylion007, https://github.com/EikanWang	2025-05-21 13:29:41 +00:00
ZhiweiYan-96	616e20b4f1	[Intel GPU] scalar tensor case handling in addmm, baddmm (#153051 ) # Motivation This PR adds scalar tensor (`t.elem()==1 and t.sizes().empty() == true and t.dim()=0` )handling in addmm, baddmm. The issue is found during the process of oneDNN upgradation. Found that the new version of oneDNN requires the post-binary (self in addmm) has same dimension as the one of output tensor. Now we need explicitly expand the shape of `self` tensor. Former version dnnl may help us do the broadcasting inside. This PR could fix issues in https://github.com/intel/torch-xpu-ops/issues/1612 and CI error in https://github.com/pytorch/pytorch/pull/151767. # Implementation We treat the scalar tensor as normal tensor by `unsqueeze` it as 1 dimension tensor. Accompanied with the existing shape handle logic, it would be further `unsqueeze` to 2D or 3D shape. UT testing ``` python test/inductor/test_torchinductor_opinfo.py TestInductorOpInfoXPU.test_comprehensive_addmm_xpu_float32 python test/inductor/test_torchinductor_opinfo.py TestInductorOpInfoXPU.test_comprehensive_addmv_xpu_float32 python test/inductor/test_torchinductor_opinfo.py TestInductorOpInfoXPU.test_comprehensive_baddbmm_xpu_float16 python test/inductor/test_torchinductor_opinfo.py TestInductorOpInfoXPU.test_comprehensive_baddbmm_xpu_float32 ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153051 Approved by: https://github.com/EikanWang, https://github.com/guangyey	2025-05-21 12:24:37 +00:00
Irem Yuksel	afd7a13bca	Migrate to new Windows Arm64 runners (#152099 ) This PR moves the Windows Arm64 nightly jobs to the new runner image, see [arm-windows-11-image](https://github.com/actions/partner-runner-images/blob/main/images/arm-windows-11-image.md ) Fixes #151671 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152099 Approved by: https://github.com/seemethere	2025-05-21 09:13:15 +00:00
Aaron Gokaslan	ffd49d538e	[BE][Ez]: Improve typing in torch/modules/container.py (#153728 ) Adds some missing type annotations Pull Request resolved: https://github.com/pytorch/pytorch/pull/153728 Approved by: https://github.com/albanD	2025-05-21 07:15:00 +00:00
Andy (An) Wang	a636a92ee9	[MTIA ATen Backend] Migrate "_unsafe_view" and "view" ops from out-of-tree to pytorch in-tree (#153670 ) Summary: # Context The MTIA New Aten Backend work is essentially to move MTIA operators from pytorch out-of-tree to in-tree, with following benefits: 1. Avoid duplicate code copied from pytorch, e.g. view ops implementation, util functions. 2. Utilize TensorIterator and structured kernel codegen, avoid manual implementation of broadcasting, dtype casting, asserting, etc. 3. Eliminate MTIA's own codegen flow, which is unnecessary complexity. 4. Overall make MTIA's aten backend more pytorch native. Differential Revision: D74672464 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153670 Approved by: https://github.com/albanD, https://github.com/nautsimon	2025-05-21 05:20:45 +00:00
xinan.lin	dcb3edd30d	[AOTI][XPU] Refactor AOTInductor runtime API for Intel GPU. (#153929 ) Simplify and improve code format for sycl_runtime_wrappers.h Pull Request resolved: https://github.com/pytorch/pytorch/pull/153929 Approved by: https://github.com/desertfire ghstack dependencies: #153924	2025-05-21 03:52:54 +00:00
PyTorch MergeBot	531d8f5fb6	Revert "[cuBLAS][cuBLASLt] Use cuBLAS default workspace size in Lt (#153556 )" This reverts commit 2b43d635d31f1338743885efd1a259f43bd2ee65. Reverted https://github.com/pytorch/pytorch/pull/153556 on behalf of https://github.com/eqy due to reverting, will add disable for reduced precision reduction ([comment](https://github.com/pytorch/pytorch/pull/153556#issuecomment-2896257521))	2025-05-21 02:09:11 +00:00
PyTorch MergeBot	1478d0185c	Revert "[CI][CUDA] Move cu118 distributed pull jobs to cu126, move cu124-sm75 to cu126-sm75 (#151594 )" This reverts commit 8cabd23b3d357ec38a400978bb5423efcb433f2a. Reverted https://github.com/pytorch/pytorch/pull/151594 on behalf of https://github.com/huydhn due to Sorry for reverting your change but it seems to fail a distributed test in trunk ([comment](https://github.com/pytorch/pytorch/pull/151594#issuecomment-2896230131))	2025-05-21 01:45:20 +00:00
Yu, Guangye	daa68e7a93	Update USE_XCCL option if USE_XPU is OFF (#153936 ) # Motivation Disable `USE_XCCL` when `USE_XPU` is turned `OFF` to ensure configuration consistency. This is required because XCCL depends on XPU functionality. Especially, ensure that `USE_XCCL` is correctly set to `OFF` when [caffe2_update_option(USE_XPU OFF)](`1075bb37d3/cmake/Dependencies.cmake (L97)`) is invoked. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153936 Approved by: https://github.com/Skylion007	2025-05-21 01:32:41 +00:00
PyTorch MergeBot	cf6e5d1881	Revert "[cuBLASLt] relax `addmm` cuBLASLt constraint (#153675 )" This reverts commit f9bb7cf72a43bec17b5fc2ccbe865aa130e760be. Reverted https://github.com/pytorch/pytorch/pull/153675 on behalf of https://github.com/eqy due to incorrect, cuBLASLt doesnt handle beta != 1.0 but this appears untested ([comment](https://github.com/pytorch/pytorch/pull/153675#issuecomment-2896188784))	2025-05-21 01:20:10 +00:00
Frost Mitchell	fe49b11e09	Add memory reporting for XPU to Memory Profiler (#152842 ) Adds support for XPU profile_memory in Pytorch Profiler. Currently, when `profile_memory=True` is passed to `torch.profiler.profile`, there is no XPU memory reported. For example, the profiling table printed by the code below is missing any `XPU Mem` columns: <details><summary>profiling.py</summary> <p> ```python import torch import torch.nn as nn import torch.optim as optim from torch.profiler import profile, ProfilerActivity class ToyModel(nn.Module): def __init__(self): super(ToyModel, self).__init__() self.conv1 = nn.Conv1d(20,20,15,padding="same") self.flatten = nn.Flatten() self.net1 = nn.Linear(2048, 4096) self.relu = nn.ReLU() self.net2 = nn.Linear(4096, 5) def forward(self, x): res = self.conv1(x) res = self.flatten(res) res = self.net1(res) return self.net2(self.relu(res)) def demo_basic(): model = ToyModel().to("xpu") loss_fn = nn.MSELoss().to("xpu") optimizer = optim.SGD(model.parameters(), lr=0.001) with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.XPU], profile_memory=True) as prof: for epoch in range(10): optimizer.zero_grad() outputs = model(torch.randn(20, 2048).to("xpu")) labels = torch.randn(20, 5).to("xpu") loss_fn(outputs, labels).backward() optimizer.step() print(prof.key_averages().table(max_name_column_width=100, sort_by="xpu_time_total", row_limit=100)) if __name__ == "__main__": demo_basic() ``` </p> </details> ``` ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg Self XPU Self XPU % XPU total XPU time avg CPU Mem Self CPU Mem # of Calls ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ gemm_kernel 0.00% 0.000us 0.00% 0.000us 0.000us 1.501ms 44.73% 1.501ms 25.024us 0 b 0 b 60 autograd::engine::evaluate_function: AddmmBackward0 0.12% 1.067ms 30.47% 260.929ms 13.046ms 0.000us 0.00% 1.009ms 50.448us 0 b 0 b 20 AddmmBackward0 0.09% 744.983us 15.99% 136.944ms 6.847ms 0.000us 0.00% 784.640us 39.232us 0 b 0 b 20 aten::mm 15.41% 131.956ms 15.79% 135.167ms 3.379ms 784.640us 23.37% 784.640us 19.616us 0 b 0 b 40 aten::linear 0.02% 156.361us 20.58% 176.187ms 8.809ms 0.000us 0.00% 741.760us 37.088us 0 b 0 b 20 aten::addmm 20.25% 173.371ms 20.52% 175.723ms 8.786ms 741.760us 22.10% 741.760us 37.088us 0 b 0 b 20 Optimizer.step#SGD.step 0.40% 3.429ms 5.55% 47.509ms 4.751ms 0.000us 0.00% 488.960us 48.896us 0 b 0 b 10 aten::_foreach_add_ 4.81% 41.162ms 5.15% 44.080ms 4.408ms 488.960us 14.57% 488.960us 48.896us 0 b 0 b 10 at::native::xpu::MultiTensorApplyKernelFunctor<at::n... 0.00% 0.000us 0.00% 0.000us 0.000us 422.880us 12.60% 422.880us 42.288us 0 b 0 b 10 autograd::engine::evaluate_function: ConvolutionBack... 0.03% 280.041us 4.36% 37.328ms 3.733ms 0.000us 0.00% 356.320us 35.632us 0 b 0 b 10 ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ Self CPU time total: 856.227ms Self XPU time total: 3.357ms ``` This PR updates the XPUCachingAllocator.cpp to report allocation events to the Profiler, and causes these to be printed in the table: ``` ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg Self XPU Self XPU % XPU total XPU time avg CPU Mem Self CPU Mem XPU Mem Self XPU Mem # of Calls ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ gemm_kernel 0.00% 0.000us 0.00% 0.000us 0.000us 1.436ms 43.64% 1.436ms 23.939us 0 b 0 b 0 b 0 b 60 autograd::engine::evaluate_function: AddmmBackward0 0.13% 1.186ms 29.92% 262.875ms 13.144ms 0.000us 0.00% 1.005ms 50.272us 0 b 0 b 320.94 Mb -4.69 Mb 20 AddmmBackward0 0.09% 815.288us 16.48% 144.802ms 7.240ms 0.000us 0.00% 790.720us 39.536us 0 b 0 b 325.47 Mb 0 b 20 aten::mm 15.86% 139.342ms 16.26% 142.875ms 3.572ms 790.720us 24.03% 790.720us 19.768us 0 b 0 b 325.47 Mb 325.47 Mb 40 aten::linear 0.02% 182.856us 20.46% 179.775ms 8.989ms 0.000us 0.00% 669.440us 33.472us 0 b 0 b 3.13 Mb 0 b 20 aten::addmm 20.10% 176.607ms 20.40% 179.210ms 8.961ms 669.440us 20.34% 669.440us 33.472us 0 b 0 b 3.13 Mb 3.13 Mb 20 Optimizer.step#SGD.step 0.42% 3.692ms 5.61% 49.267ms 4.927ms 0.000us 0.00% 486.640us 48.664us 0 b 0 b 0 b 0 b 10 aten::_foreach_add_ 4.83% 42.439ms 5.19% 45.574ms 4.557ms 486.640us 14.79% 486.640us 48.664us 0 b 0 b 0 b -20.00 Kb 10 at::native::xpu::MultiTensorApplyKernelFunctor<at::n... 0.00% 0.000us 0.00% 0.000us 0.000us 420.960us 12.79% 420.960us 42.096us 0 b 0 b 0 b 0 b 10 autograd::engine::evaluate_function: ConvolutionBack... 0.04% 310.719us 4.47% 39.279ms 3.928ms 0.000us 0.00% 339.520us 33.952us 0 b 0 b -2.89 Mb -3.12 Mb 10 ------------------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ ------------ Self CPU time total: 878.627ms Self XPU time total: 3.291ms ``` These XPU memory numbers match the same profiling results on CUDA. Pull Request resolved: https://github.com/pytorch/pytorch/pull/152842 Approved by: https://github.com/guangyey, https://github.com/sraikund16	2025-05-21 01:19:19 +00:00
Jane Xu	8817e5ac80	Render Example: and not Example:: in docs (#153978 ) Everything here is a grep except the changes in tools/autograd/load_derivatives.py which I manually corrected. The correct notation is: ``` Example:: >>> ... ``` It is common and wrong to have: ``` Example:: >>> ... ``` In the wrong example, we get these pesky double colons: ![image](https://github.com/user-attachments/assets/20ffd349-68bb-4552-966c-e23923350476) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153978 Approved by: https://github.com/soulitzer, https://github.com/malfet	2025-05-21 01:03:26 +00:00
atalman	0959869683	Bump triton pin for the release 3.3.1 of triton (#153951 ) Triton is pointing to latest triton pin : https://github.com/triton-lang/triton/tree/release/3.3.x XPU pointing to latest XPU pin: https://github.com/intel/intel-xpu-backend-for-triton/commits/release/3.3.x/ This version contains the fix for: Compilation Issue for RTX 5090 GPUs with Compute Capability = 120. https://github.com/triton-lang/triton/pull/6771 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153951 Approved by: https://github.com/davidberard98	2025-05-21 00:27:39 +00:00
Menglu Yu	32b1baa981	[3/n][Optimus][Auto-AC] Support float8_e4m3fn quantization type and set scaling as the default (#153802 ) Summary: 1. Customers now can test with float8_e4m3fn. 2. To play safe, we set the scaling version as the default. Test Plan: ### unit test ``` buck2 test 'fbcode//mode/dev-nosan' fbcode//caffe2/test/inductor:quantization ``` Buck UI: https://www.internalfb.com/buck2/f679f362-8bf4-454c-87df-a85cbc2ab2a8 Test UI: https://www.internalfb.com/intern/testinfra/testrun/5066549861047443 Network: Up: 16KiB Down: 3.9MiB (reSessionID-98badbfd-76f7-487f-ab1c-1ec4f850614d) Analyzing targets. Remaining 0/281 Executing actions. Remaining 0/5957 7.3s exec time total Command: test. Finished 3 local, 1 remote Time elapsed: 1:29.7s Tests finished: Pass 3. Fail 0. Fatal 0. Skip 0. Build failure 0 Differential Revision: D74910193 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153802 Approved by: https://github.com/nareshrajkumar866, https://github.com/Hahu803, https://github.com/Mingming-Ding	2025-05-21 00:21:54 +00:00
Nikita Shulga	58dc80dff6	[MPSInductor] Fix indexing calculation (#153997 ) By using `c10:🤘:floor_divie` primitive Which fixes `test_flip_cat_mps` test, and makes `doctr_reco_predictor` and `doctr_det_predictor` pass accuracy checks (at least locally, scheduled a workflow dispatch to validate it in CI) Before this change following script generated different compile and eager results ```python import torch def foo(unsqueeze, unsqueeze_1): cat_1 = torch.ops.aten.cat.default([unsqueeze, unsqueeze_1], 1) view = torch.ops.aten.view.default(cat_1, [4]) slice_5 = torch.ops.aten.slice.Tensor(view, 0, 0, 3) rev_1 = torch.ops.aten.flip.default(slice_5, [0]) return rev_1 if __name__ == "__main__": x = torch.arange(1.0, 3.0, device='mps').reshape(2, 1) y = torch.arange(5.0, 7.0, device='mps').reshape(2, 1) rc, (kernel,) = torch._inductor.utils.run_and_get_kernels(torch.compile(foo), x, y) print(kernel) print("Compile: ", rc) print("Eager: ", foo(x, y)) ``` After this change ``` ''' #include <c10/metal/utils.h> kernel void generated_kernel( device float* out_ptr0, constant float* in_ptr0, constant float* in_ptr1, uint xindex [[thread_position_in_grid]] ) { int x0 = xindex; auto tmp6 = in_ptr0[1 + (c10:🤘:floor_divide((-1)x0, 2))]; auto tmp11 = in_ptr1[1 + (c10:🤘:floor_divide((-1)x0, 2))]; auto tmp0 = (2 + ((-1)*x0)) % (2); auto tmp1 = static_cast<long>(tmp0); auto tmp2 = 0; auto tmp3 = tmp1 >= tmp2; auto tmp4 = 1; auto tmp5 = tmp1 < tmp4; auto tmp7 = tmp5 ? tmp6 : 0.0; auto tmp8 = tmp1 >= tmp4; auto tmp9 = 2; auto tmp10 = tmp1 < tmp9; auto tmp12 = tmp8 ? tmp11 : 0.0; auto tmp13 = tmp5 ? tmp7 : tmp12; out_ptr0[x0] = static_cast<float>(tmp13); } ''' Compile: tensor([2., 5., 1.], device='mps:0') Eager: tensor([2., 5., 1.], device='mps:0') ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153997 Approved by: https://github.com/dcci ghstack dependencies: #153970, #153971	2025-05-21 00:03:46 +00:00
Jane Xu	fc33da410f	Add torch/header_only_apis.txt and enforce they're tested (#153635 ) This PR adds enforcement of testing header only APIs. The benefit of torch/header_only_apis.txt is twofold: 1) this gives us a clear view of what we expect to be header only 2) this allows us to enforce testing The enforcement added in this PR is very basic--we literally string match that a symbol in `torch/header_only_apis.txt` is in a cpp test. This is meant to be a first step in verifying our APIs are properly tested and can get fancier over time. For now, I've added myself as a codeowner to learn what to look out for in terms of proper tests. Over time, I anticipate we can automate more steps, but right now let's just get something out the door. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153635 Approved by: https://github.com/albanD ghstack dependencies: #153965	2025-05-20 23:42:24 +00:00
Jane Xu	41a9aa6564	Remove janky (though at times useful) dlclose test (#153975 ) This test was never the shining star in class but it helped check that we can properly delete a stable library. But now that we are running it in CI this is not a good test to annoy people with as dlclose + parallelism is likely not the move. I will miss it locally though. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153975 Approved by: https://github.com/jbschlosser	2025-05-20 23:26:42 +00:00
PyTorch MergeBot	7b7604fdb4	Revert "[inductor][cutlass backend] Add 2 stage autotuning aka prescreening (#153335 )" This reverts commit 0c04492e3b142854fad8356a2a4d74f12e2c6c5d. Reverted https://github.com/pytorch/pytorch/pull/153335 on behalf of https://github.com/malfet due to Breaks lint, see `3742b7fb3a/1` ([comment](https://github.com/pytorch/pytorch/pull/153335#issuecomment-2896031661))	2025-05-20 23:12:11 +00:00
Slawomir Siwek	3742b7fb3a	Treat dim=[] same as dim=None (#153570 ) Fixes https://github.com/pytorch/pytorch/issues/153568 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153570 Approved by: https://github.com/ngimel	2025-05-20 22:44:29 +00:00
albanD	f7b8eadd9d	Add codeowner for merge rules (#152354 ) To ensure changes to merge rights are properly reviewed Also make the codeowner file valid by removing invalid users Pull Request resolved: https://github.com/pytorch/pytorch/pull/152354 Approved by: https://github.com/malfet	2025-05-20 22:24:23 +00:00
henrylhtsang	0c04492e3b	[inductor][cutlass backend] Add 2 stage autotuning aka prescreening (#153335 ) Motivation: By default, we are tuning the cutlass backend kernels on 3 swizzles. There are runtime params, so they share the same underlying kernel, which saves a lot of compilation time. However, autotuning all combinations of {configs} x {swizzles} is still expensive. Observations: Winner of the {configs} x {swizzles} autotuning is the same as if we do a greedy search: first find the top X winners of {configs} with swizzle 2 (hardcoded), then autotune on the {top X winner configs} x {swizzles}. In other words, we can use a Greedy algorithm to reduce autotuning time. I attach the logs below. This somewhat depends on what X is, but a number like 5-10 works pretty well from empirical observations. Logs: Baseline: https://gist.github.com/henrylhtsang/9a604f150a270dc19524f72a5d4dfac2 ``` AUTOTUNE mm(2048x2048, 2048x2048) strides: [2048, 1], [1, 2048] dtypes: torch.bfloat16, torch.bfloat16 cuda_cutlass_gemm_1776 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1777 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1778 0.0291 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1800 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1801 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1802 0.0293 ms 99.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_9012 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_9013 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_9014 0.0294 ms 98.9% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8940 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8941 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8942 0.0296 ms 98.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8934 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8935 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8936 0.0297 ms 98.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_2001 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_2002 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_2003 0.0297 ms 97.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1848 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1849 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1850 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8964 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8965 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8966 0.0298 ms 97.6% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_8958 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_8959 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_8960 0.0298 ms 97.5% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1929 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1930 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1931 0.0302 ms 96.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1770 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1771 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1772 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1953 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1954 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1955 0.0302 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_tnn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1995 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1996 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1997 0.0303 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1794 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1795 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1796 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1842 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_1843 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_1844 0.0303 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_9006 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cuda_cutlass_gemm_9007 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cuda_cutlass_gemm_9008 0.0304 ms 95.7% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cuda_cutlass_gemm_1923 0.0306 ms 95.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x1x1_0_tnn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 ``` with prescreening: ``` AUTOTUNE mm(147456x6144, 6144x2048) strides: [6144, 1], [2048, 1] dtypes: torch.bfloat16, torch.bfloat16 cutlass_1a5e81af 4.5469 ms 100.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.6328 ms 98.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.6836 ms 97.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_161b8b81 4.7224 ms 96.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_161b8b81 4.7234 ms 96.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_161b8b81 4.7274 ms 96.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_853b6347 4.7369 ms 96.0% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_aa6f899c 4.7404 ms 95.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_161b8b81 4.7711 ms 95.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_8bc6fbda 4.8148 ms 94.4% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_8bc6fbda 4.8159 ms 94.4% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_8bc6fbda 4.8214 ms 94.3% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_8bc6fbda 4.8302 ms 94.1% cutlass3x_sm90_tensorop_s64x256x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_0a1c55af 4.8487 ms 93.8% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_0a1c55af 4.8527 ms 93.7% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_02780d72 4.8617 ms 93.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_0a1c55af 4.8737 ms 93.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_0a1c55af 4.8738 ms 93.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_02780d72 4.9348 ms 92.1% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_02780d72 4.9763 ms 91.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_853b6347 4.9805 ms 91.3% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_1a5e81af 5.0225 ms 90.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_853b6347 5.0271 ms 90.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_02780d72 5.0595 ms 89.9% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_853b6347 5.1434 ms 88.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.1574 ms 88.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=8 cutlass_1a5e81af 5.1916 ms 87.6% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.2018 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=4 cutlass_c1ffa14b 5.2019 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=1 cutlass_c1ffa14b 5.2037 ms 87.4% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_256x128x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_1a5e81af 5.5329 ms 82.2% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_2x1x1_0_ttn_align8_stream_k_warpspecialized_cooperative_epi_tma swizzle=2 cutlass_aa6f899c 11.5046 ms 39.5% cutlass3x_sm90_tensorop_s64x128x16gemm_bf16_bf16_f32_void_bf16_128x256x64_1x2x1_0_ttn_align8_warpspecialized_cooperative_epi_tma swizzle=8 SingleProcess AUTOTUNE benchmarking takes 1.9526 seconds and 0.0352 seconds precompiling for 32 choices ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153335 Approved by: https://github.com/eellison	2025-05-20 22:19:02 +00:00
Bin Bao	2c2524f74b	[AOTI] Generate unique cubin file names when package_cpp_only (#153948 ) Summary: * When package_cpp_only is specified, generate kernel file names with unique kernel names to make the final packaged package files more readable. Assert on unique_kernel_names in case somehow it was explicitly set to False. * Fix a rocm test skip, see https://github.com/pytorch/pytorch/pull/153828 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153948 Approved by: https://github.com/angelayi, https://github.com/yushangdi	2025-05-20 22:07:53 +00:00
Wei Wang	8cabd23b3d	[CI][CUDA] Move cu118 distributed pull jobs to cu126, move cu124-sm75 to cu126-sm75 (#151594 ) This PR moves distributed cuda CI job from cuda 11.8 to cuda 12.6. In doing so, a few unit test failures were exposed, some if not all of which would take a while to root-cause and fix, so temporarily skip them after creating the issues. https://github.com/pytorch/pytorch/issues/153479 test_nan_assert tricky behavior (e.g. skip_but_pass_in_sandcastle, ubuntu 20.04 does not work, ubuntu 22.04 works, Amazon Linux 2023 skip - what is Sandcastle OS?) https://github.com/pytorch/pytorch/issues/153122 CUDA context related https://github.com/pytorch/pytorch/issues/153517 NCCL regression, future NCCL may fix it See: https://github.com/pytorch/pytorch/issues/147383 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151594 Approved by: https://github.com/eqy, https://github.com/atalman, https://github.com/cyyever	2025-05-20 21:56:47 +00:00
Eddie Yan	2b43d635d3	[cuBLAS][cuBLASLt] Use cuBLAS default workspace size in Lt (#153556 ) Also enables unified workspaces by default for non-FBCODE use cases. Default Lt workspace size is also updated to match cuBLAS logic for default, including for Blackwell (SM 10.0) and GeForce Blackwell (SM 12.0). Recommended defaults are documented here: https://docs.nvidia.com/cuda/cublas/#cublassetworkspace Pull Request resolved: https://github.com/pytorch/pytorch/pull/153556 Approved by: https://github.com/Skylion007, https://github.com/ngimel	2025-05-20 21:51:49 +00:00
Yiming Zhou	aeb734f519	[nativert] Move GraphSignature to pytorch core (#152969 ) Summary: Torch Native Runtime RFC: https://github.com/pytorch/rfcs/pull/72 Added an in-memory representation for input and output specs of a graph. The GraphSignature class models the input and output specs of an exported graph produced by torch.export, which holds the graph information deserialized from the pt2 archive package. Runtime relies on the GraphSignature for weight name lookup and weight loading. The serialization schema is defined in torch/_export/serde/schema.py See more at: https://docs.pytorch.org/docs/stable/export.html#torch.export.ExportGraphSignature Test Plan: Added tests under `test/cpp/nativert/test_graph_signature.cpp` Differential Revision: D73895378 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152969 Approved by: https://github.com/swolchok	2025-05-20 21:49:56 +00:00
Jane Xu	8f943046f8	[BE] light cleanups to linter logic (#153965 ) some BE cleanup on other lint things I saw while doing the top of the this stack Pull Request resolved: https://github.com/pytorch/pytorch/pull/153965 Approved by: https://github.com/soulitzer	2025-05-20 21:28:48 +00:00
Grace Cheng	deaf6c2f2f	Address the ignored warnings for `-Wmissing-field-initializers` in the file fbcode/caffe2/aten/src/ATen/native/cuda/RowwiseScaledMM.cu (#153958 ) Summary: the error message https://www.internalfb.com/sandcastle/workflow/698057942249983018/artifact/actionlog.698057942382778255.stderr.1?selectedLines=66-66-70-148 from D74892646 When switching the host compiler to Clang, maybe we should only silence these warnings in this file. Test Plan: sandcastle_green Differential Revision: D75029051 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153958 Approved by: https://github.com/Skylion007, https://github.com/eqy	2025-05-20 21:25:56 +00:00
Nikita Shulga	6cb7e4b5a5	[EZ] Update mps xfail reason (#153971 ) cummin is not implemented Pull Request resolved: https://github.com/pytorch/pytorch/pull/153971 Approved by: https://github.com/dcci, https://github.com/jansel ghstack dependencies: #153970	2025-05-20 21:15:14 +00:00
Nikita Shulga	03859242ce	[Testing] Fix `test_deterministic_`... on MPS (#153970 ) By decorated emitted kernels with `'''` rather than `"""` To match regex in `torch._inductor.utils.run_and_get_kernels` This fixes `test_deterministic_codegen_mps`, `test_deterministic_codegen_on_graph_break_mps` and `test_deterministic_codegen_with_suffix_mps` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153970 Approved by: https://github.com/dcci, https://github.com/jansel	2025-05-20 21:15:14 +00:00
soulitzer	3aa95b252a	Fix test_side_stream_backward_overlap flakiness (#153963 ) Fixes https://github.com/pytorch/pytorch/issues/153927 Although the autograd backward should always execute SideBw before MainBw, there is still a small chance the recorded events won't be in that order. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153963 Approved by: https://github.com/janeyx99, https://github.com/Skylion007 ghstack dependencies: #151079, #153412	2025-05-20 21:02:56 +00:00
PyTorch MergeBot	500a710422	Revert "Fixed an issue with XPU skip so the test_decompose_mem_bound_mm.py suite can be ran correctly (#153245 )" This reverts commit 2e56ce097a201ff3c69610cea953a9efce17d1b1. Reverted https://github.com/pytorch/pytorch/pull/153245 on behalf of https://github.com/yangw-dev due to tests failed internally [D75078034](https://www.internalfb.com/diff/D75078034) ([comment](https://github.com/pytorch/pytorch/pull/153245#issuecomment-2895785642))	2025-05-20 20:45:55 +00:00
Xu Han	179e7d8624	Fix vs2022 caused AVX512 illegal instruction issue. (#153480 ) Fixes #145702 Add `/d2implyavx512upperregs-` to disable compiler over-aggressive optimization, which caused involeved AVX512 register on AVX2 machine. Reference to: https://github.com/pytorch/pytorch/issues/145702#issuecomment-2874029459 Local test passed: <img width="1208" alt="image" src="https://github.com/user-attachments/assets/26f4cb91-6bb5-416f-aa35-c899eb1489b2" /> Pull Request resolved: https://github.com/pytorch/pytorch/pull/153480 Approved by: https://github.com/Blackhex, https://github.com/cyyever, https://github.com/atalman	2025-05-20 20:37:00 +00:00
Anita Katahoire	996c4d803d	Removing conda references from PyTorch Docs (#152702 ) Addresses #148339 Pull Request resolved: https://github.com/pytorch/pytorch/pull/152702 Approved by: https://github.com/svekars, https://github.com/albanD, https://github.com/atalman	2025-05-20 20:33:28 +00:00
Gantaphon Chalumporn	05bc78e64f	[submodule] Update fbgemm pinned version (#153950 ) Summary: Update fbgemm pinned version in PyTroch. Related update in fbgemm: D74434751 Included changes: Update fbgemm external dependencies directory in setup.py Add DISABLE_FBGEMM_AUTOVEC flag to disable fbgemm's autovec Test Plan: PyTorch OSS CI Differential Revision: D75073516 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153950 Approved by: https://github.com/Skylion007, https://github.com/ngimel	2025-05-20 20:24:27 +00:00
eqy	823a35807c	[CUDA][CUDNN] Dispatch to cuDNN for non-batch-splittable 64-bit NCHW convolutions (#153101 ) For #152816 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153101 Approved by: https://github.com/Skylion007	2025-05-20 20:19:03 +00:00
pbialecki	e8f8baf71f	set CUDA_MODULE_LOADING for older drivers only (#152695 ) `CUDA_MODULE_LOADING=LAZY` is the default for all drivers shipped with CUDA >=12.2 and we should check the driver version before setting the env variable. (the `LOG(WARNING)` has to be removed before merging) Pull Request resolved: https://github.com/pytorch/pytorch/pull/152695 Approved by: https://github.com/malfet, https://github.com/atalman, https://github.com/nWEIdia	2025-05-20 19:34:40 +00:00
Joel Schlosser	7587350458	Make python_agnostic cpp extension tests standalone (#153274 ) Related: #148920 This PR: * Introduces a new file `test/cpp_extensions/python_agnostic_extension/test/test_python_agnostic.py` with testing that follows the usual python testing patterns * This replaces the testing for python_agnostic in `test/test_cpp_extensions_aot.py` After this PR, it is now possible to run: ``` python test/cpp_extensions/python_agnostic_extension/test/test_python_agnostic.py ``` and the test will build the prerequisite wheel before running the tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153274 Approved by: https://github.com/janeyx99, https://github.com/cyyever ghstack dependencies: #153264	2025-05-20 19:18:09 +00:00
Joel Schlosser	3ecd444004	Support independent builds for cpp extension tests + apply to libtorch_agnostic tests (#153264 ) Related: #148920 This PR: * Provides a helper `install_cpp_extension(extension_root)` for building C++ extensions. This is intended to be used in `TestMyCppExtension.setUpClass()` * Updates libtorch_agnostic tests to use this * Deletes preexisting libtorch_agnostic tests from `test/test_cpp_extensions_aot.py` * Fixes `run_test.py` to actually run tests in `test/cpp_extensions/libtorch_agnostic_extension/test/test_libtorch_agnostic.py` to avoid losing coverage. This wasn't being run due to logic excluding tests that start with "cpp"; this is fixed now After this PR, it is now possible to run: ``` python test/cpp_extensions/libtorch_agnostic_extension/test/test_libtorch_agnostic.py ``` and the test will build the `libtorch_agnostic` extension before running the tests. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153264 Approved by: https://github.com/janeyx99	2025-05-20 19:18:09 +00:00
Tsung-Hsien Lee	f1f54c197d	[c10d] Simplify `new_subgroups()` by using `new_subgroups_by_enumeration()` (#153843 ) Summary: The code changes in each file of the diff include removing the `subgroups` and `cur_subgroup` variables, and replacing the while loop with a call to `new_subgroups_by_enumeration()`. Test Plan: contbuild & OSS CI Differential Revision: D75007368 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153843 Approved by: https://github.com/Skylion007, https://github.com/wz337	2025-05-20 19:15:20 +00:00
Bert Maher	2d20106922	[inductor] Support cutlass backend with remote execution (#153844 ) Meta-internal builds need to use RE to build with nvcc, since the trainers do not have nvcc (and its attendant build toolchain) installed. This diff enables building using an RE service (via the same code path used for Triton) Differential Revision: [D74907192](https://our.internmc.facebook.com/intern/diff/D74907192/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153844 Approved by: https://github.com/henrylhtsang	2025-05-20 19:05:23 +00:00
Dan Zimmerman	e0f8174001	[triton][fb] Move build_paths into triton_utils (#153652 ) Summary: TSA, this is just a small cleanup Test Plan: CI Differential Revision: D74835506 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153652 Approved by: https://github.com/Skylion007	2025-05-20 18:59:50 +00:00
Eddie Yan	f9bb7cf72a	[cuBLASLt] relax `addmm` cuBLASLt constraint (#153675 ) `beta == 1.0` doesn't seem to be required anymore https://github.com/pytorch/pytorch/issues/153590 `self.dim() == 1` restriction seems to still hold but not sure if that's due to a lack of handling on the PyTorch side or the cuBLASLt side, will investigate Pull Request resolved: https://github.com/pytorch/pytorch/pull/153675 Approved by: https://github.com/Skylion007	2025-05-20 18:43:38 +00:00
Svetlana Karslioglu	7c9d94e9bb	Redirect mobile_optimizer.rst to executorch (#153664 ) Redirect mobile_optimizer.rst to executorch Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/153664 Approved by: https://github.com/byjlw, https://github.com/malfet	2025-05-20 18:13:45 +00:00
xinan.lin	0087f5f0af	[AOTI][XPU] Embed SPRI-V files into .so (#153924 ) Following the design of #150739, this PR supports embed kernel SPIR-V files so AOTI is one step closer to generate a single binary. Fixes #153829 Fixes #153830 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153924 Approved by: https://github.com/desertfire	2025-05-20 17:38:53 +00:00
Yang Wang	335c89c6f1	[Monitoring] enable local logs and add mac test monitoring (#153454 ) Enable to run the upload utilzation logics using local pointer instead of reading from s3, this could be useful for rocm too, Pull Request resolved: https://github.com/pytorch/pytorch/pull/153454 Approved by: https://github.com/huydhn	2025-05-20 17:14:40 +00:00
henrylhtsang	b910d37ec6	[cutlass backend] Reduce log level for cutlass runtime error (#153457 ) Want to make sure we always call self.cleanup_run_fn() even if we crash. I think this is the reason why sometimes we get ``` in _dlclose TypeError: 'NoneType' object is not callable ``` Differential Revision: [D74629230](https://our.internmc.facebook.com/intern/diff/D74629230/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153457 Approved by: https://github.com/ColinPeppler	2025-05-20 17:03:17 +00:00
Shivam Raikundalia	6b5b69a468	[Memory Snapshot] Fix RecordFunction Callback Handling (#153839 ) Fixes #153571 Summary: 1. Set annotation callback to global to include all threads 2. Only init callbacks when enable == true and callbacks are empty under mutex 3. When enable == false, check if callbacks are present and if so remove them and set handle to 0 under mutex We don't expect memory snapshots to be called from several different threads (almost always called just from main) but we make sure to add thread safety in the off case that users do want to call it from different points of entry Test Plan: Ran basic snapshot and saw that the callbacks were registered properly Reviewed By: ngimel Differential Revision: D74771491 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153839 Approved by: https://github.com/ngimel, https://github.com/Skylion007	2025-05-20 17:01:00 +00:00
angelayi	ddfaab3b56	[aoti] Reset expr when generating cpp code (#153898 ) Maybe fixes https://github.com/pytorch/pytorch/issues/153896 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153898 Approved by: https://github.com/desertfire	2025-05-20 16:31:25 +00:00
Eddie Yan	5163bf0069	[CUDA][cuBLAS][cuBLASLt] avoid polluting prefer cuBLAS/Lt setting across tests (#153655 ) Some tests may not set the preferred backend, which leads to unexpected behavior when multiple tests are run vs. standalone Tests that should exercise both backends should explicitly parametrize this setting Pull Request resolved: https://github.com/pytorch/pytorch/pull/153655 Approved by: https://github.com/ngimel	2025-05-20 16:18:35 +00:00
PaulZhang12	a7c01d7f13	[Inductor] Subgraph check output strides (#153755 ) Make sure outputs strides of subgraph consistent with original gm. Without checking strides, it was possible for subgraph to produce nans with a reinterpret tensor on the output of the subgraph output, in which itself was not contiguous. Differential Revision: [D74691119](https://our.internmc.facebook.com/intern/diff/D74691119/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153755 Approved by: https://github.com/eellison ghstack dependencies: #153754	2025-05-20 16:07:18 +00:00
PaulZhang12	63e5d46478	[Inductor] Subgraph support dynamic input expressions (#153754 ) Support subgraph choice taking in inputs that have dynamic dimensions. Testing with decomposeK subgraph decomp Differential Revision: [D74484741](https://our.internmc.facebook.com/intern/diff/D74484741/) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153754 Approved by: https://github.com/eellison	2025-05-20 16:07:18 +00:00
iupaikov-amd	2e56ce097a	Fixed an issue with XPU skip so the test_decompose_mem_bound_mm.py suite can be ran correctly (#153245 ) Fixes #153239 Replaced custom decorator with the common one. Although the better way to skip the whole suite would be to add it to skip list in run_test.py Pull Request resolved: https://github.com/pytorch/pytorch/pull/153245 Approved by: https://github.com/guangyey, https://github.com/EikanWang, https://github.com/jeffdaily	2025-05-20 15:46:21 +00:00
Eddie Yan	ef958fa152	[cuDNN][cuDNN frontend] upgrade cuDNN frontend submodule to 1.12 (#153888 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153888 Approved by: https://github.com/Skylion007	2025-05-20 15:08:37 +00:00
PyTorch MergeBot	3102ae6798	Revert "[AOTI] Add an option to specify custom op C shim (#153851 )" This reverts commit 365ac49840105918c604a6b1c7e81c1ca59e37fb. Reverted https://github.com/pytorch/pytorch/pull/153851 on behalf of https://github.com/malfet due to Looks like it broke fuzzer test, but I could be wrong, see `c4d1ff02f8/1` ([comment](https://github.com/pytorch/pytorch/pull/153851#issuecomment-2894619773))	2025-05-20 14:23:50 +00:00
Nikita Shulga	c4d1ff02f8	[Lint] Update clang-format to 19.1.4 (#153889 ) All changes other than the one to `tools/linter/adapters/s3_init_config.json` are generated by newer clang-format Pull Request resolved: https://github.com/pytorch/pytorch/pull/153889 Approved by: https://github.com/cyyever, https://github.com/atalman	2025-05-20 14:12:46 +00:00
Aaron Gokaslan	d869ea11e0	[BE]: Update fmtlib submodule to 11.2.0 (#153853 ) Update fmtlib to 11.2.0 with a lot of miscellaneous fixes for various compilers. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153853 Approved by: https://github.com/malfet	2025-05-20 14:11:18 +00:00
James Wu	4b759d98f8	Recheck autotune cache on static cuda launcher load (#153565 ) When loading statically launchable triton kernels from FxGraphCache, since we don't instantiate a CachingAutotuner like we do normally, we need to recheck the autotune cache based on the existing compile results. If we get a hit, we take the compile result whose config matches the best config. Sometimes, the best config will have been from coordinate descent tuning. In this case, FxGraphCache today does not cache the resulting triton kernel, neither with static or without static cuda launcher. This is because coordinate descent tuning happens at runtime, and if the best config happens to not be one of the precompiled configs. Test Plan: New unit test that failed before Pull Request resolved: https://github.com/pytorch/pytorch/pull/153565 Approved by: https://github.com/aorenste	2025-05-20 14:00:43 +00:00
Michael Lazos	d68d4d31f4	[Cutlass] EVT tests update (#153926 ) Fixes internal EVT tests Pull Request resolved: https://github.com/pytorch/pytorch/pull/153926 Approved by: https://github.com/williamwen42	2025-05-20 10:03:10 +00:00
Michael Lazos	d44074f01a	[Dynamo] Fix einops regression (#153925 ) Fixes https://github.com/pytorch/pytorch/issues/153476 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153925 Approved by: https://github.com/williamwen42	2025-05-20 09:52:42 +00:00
Panagiotis Kourdis	44f19c7179	Record the XPU and XCCL build settings in the compiled binary (#147161 ) Fixes #ISSUE_NUMBER Currently the XPU and XCCL build settings are not recorded in the compiled binary and are not shown using the `torch.__config__.show()` which is a quick way to check if the binary has been built with such support. Below is the output adding them (see end of last line): ``` Python 3.12.8 \| packaged by conda-forge \| (main, Dec 5 2024, 14:24:40) [GCC 13.3.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> import torch >>> print(torch.__config__.show()) PyTorch built with: - GCC 13.3 - C++ Version: 201703 - Intel(R) oneAPI Math Kernel Library Version 2025.1-Product Build 20250203 for Intel(R) 64 architecture applications - Intel(R) MKL-DNN v3.5.3 (Git Hash 66f0cb9eb66affd2da3bf5f8d897376f04aae6af) - OpenMP 201511 (a.k.a. OpenMP 4.5) - LAPACK is enabled (usually provided by MKL) - CPU capability usage: AVX512 XPU backend - Build settings: BLAS_INFO=mkl, BUILD_TYPE=RelWithDebInfo, COMMIT_SHA=43eb39d7c832b5560f7bfa8d29cc7919ac21c0ca, CXX_COMPILER=/home/pkourdis/compilers/gcc-13.3.0/bin/c++, CXX_FLAGS= -D_GLIBCXX_USE_CXX11_ABI=1 -fvisibility-inlines-hidden -DUSE_PTHREADPOOL -DUSE_KINETO -DLIBKINETO_NOCUPTI -DLIBKINETO_NOROCTRACER -DLIBKINETO_NOXPUPTI=OFF -DUSE_PYTORCH_QNNPACK -DUSE_XNNPACK -DSYMBOLICATE_MOBILE_DEBUG_HANDLE -O2 -fPIC -Wall -Wextra -Werror=return-type -Werror=non-virtual-dtor -Werror=range-loop-construct -Werror=bool-operation -Wnarrowing -Wno-missing-field-initializers -Wno-unknown-pragmas -Wno-unused-parameter -Wno-strict-overflow -Wno-strict-aliasing -Wno-stringop-overflow -Wsuggest-override -Wno-psabi -Wno-error=old-style-cast -fdiagnostics-color=always -faligned-new -Wno-maybe-uninitialized -fno-math-errno -fno-trapping-math -Werror=format -Wno-dangling-reference -Wno-error=dangling-reference -Wno-error=redundant-move -DUSE_XPU -Wno-stringop-overflow, LAPACK_INFO=mkl, PERF_WITH_AVX=1, PERF_WITH_AVX2=1, TORCH_VERSION=2.7.0, USE_CUDA=0, USE_CUDNN=OFF, USE_CUSPARSELT=OFF, USE_EXCEPTION_PTR=1, USE_GFLAGS=OFF, USE_GLOG=OFF, USE_GLOO=ON, USE_MKL=ON, USE_MKLDNN=1, USE_MPI=0, USE_NCCL=OFF, USE_NNPACK=0, USE_OPENMP=ON, USE_ROCM=0, USE_ROCM_KERNEL_ASSERT=OFF, USE_XCCL=1, USE_XPU=1, ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/147161 Approved by: https://github.com/guangyey, https://github.com/EikanWang, https://github.com/albanD Co-authored-by: Yu, Guangye <106960996+guangyey@users.noreply.github.com>	2025-05-20 09:21:39 +00:00
PyTorch MergeBot	1075bb37d3	Revert "Fix fake tensor caching when output has unbacked (#153034 )" This reverts commit cb5f31a4a164a4fa1eaa627f9b15cdc18aa95ef1. Reverted https://github.com/pytorch/pytorch/pull/153034 on behalf of https://github.com/malfet due to Seems to have introduced flakiness in MacOS inductor tests, see https://github.com/pytorch/pytorch/issues/153891 ([comment](https://github.com/pytorch/pytorch/pull/153034#issuecomment-2893059329))	2025-05-20 06:02:38 +00:00
PyTorch MergeBot	9849c79fa2	Revert "FakeTensorMode dispatch shouldn't include bypass in exception context (#153780 )" This reverts commit aa84c037f0f473c91a79f48a5f278b7243f64b0e. Reverted https://github.com/pytorch/pytorch/pull/153780 on behalf of https://github.com/malfet due to Reverting to clearly revert https://github.com/pytorch/pytorch/pull/153034, that seems to have introduced flakiness in MacOS inductor tests, see https://github.com/pytorch/pytorch/issues/153891 ([comment](https://github.com/pytorch/pytorch/pull/153780#issuecomment-2893053304))	2025-05-20 05:59:42 +00:00
Bin Bao	365ac49840	[AOTI] Add an option to specify custom op C shim (#153851 ) Summary: Add an option to tell AOTInductor codegen to generate C shim functions for certain custom ops instead of relying on ProxyExecutor. The lib that defines custom ops need to implement corresponding C shim functions. Differential Revision: [D75014177](https://our.internmc.facebook.com/intern/diff/D75014177) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153851 Approved by: https://github.com/hl475	2025-05-20 05:12:09 +00:00
Sidharth	89ebd29fdc	[Dynamo] added warning message for tracing lru_cache wrapped functions (#153744 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153744 Approved by: https://github.com/williamwen42	2025-05-20 04:08:29 +00:00
angelayi	5ef90e14a3	[export] Remove unused constants (#153800 ) An internal test case ran into a weird issue when exporting, where the model imported a file which creates tensor constants upon importing [(code ptr)](https://fburl.com/code/xwmhxm7n). This causes the tracer to create some tensor constants even though it's not used in the model code. This PR updates the lift_constant_tensors pass to remove constant nodes that are not being used instead of lifting them as tensor constants. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153800 Approved by: https://github.com/dolpm, https://github.com/pianpwk	2025-05-20 03:15:27 +00:00
Yanli Zhao	a79e621c1c	[DDP] rebuilt bucket order when find_unused_parameters=true (#153404 ) Differential Revision: D72437251 Enable to rebuild bucket order when find_unused_parameters=true. It should be always better than not rebuilding bucket order when find_unused_parameters=True: 1. for cases where bucket order in the first iteration is the same as the parameter order, rebuilding bucket order will not change anything 2. for cases where bucket order in the first iteration is not the same as the parameter order, there could be two cases: a. bucket order will not change after 1st iteration even the graph is dynamic and there is unused parameter, in this case, rebuilding bucket order will have performance gain b. bucket order change after 1st iteration due to dynamic graph, in this case, both parameter order and bucket order in 1st iteration are not ideal, so rebuilding bucket order or not does not matter it can help case 2.a if enabling to rebuild bucket order when find_unused_parameters=true. meanwhile it will not hurt other cases in 1 and 2.b. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153404 Approved by: https://github.com/rohan-varma, https://github.com/fegin	2025-05-20 02:45:01 +00:00
Nikita Shulga	8b94d30b26	[Testing] Benchmark more tests for MPSInductor (#153897 ) And report HF tests as HF tests Pull Request resolved: https://github.com/pytorch/pytorch/pull/153897 Approved by: https://github.com/dcci	2025-05-20 02:41:38 +00:00
Nikita Shulga	1627951f24	[3.13] Remove all profiler related skips (#153857 ) As underlying issue were fixed by https://github.com/pytorch/pytorch/pull/153848 Fixes https://github.com/pytorch/pytorch/issues/142166 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153857 Approved by: https://github.com/williamwen42 ghstack dependencies: #153848	2025-05-20 01:19:52 +00:00
PyTorch MergeBot	b15720118a	Revert "Cache code generation during triton template expansion and enable it for mm_template. (#151773 )" This reverts commit 9180bb187c0e4c3ab3654e765fe33ad4c75a2b1a. Reverted https://github.com/pytorch/pytorch/pull/151773 on behalf of https://github.com/malfet due to It broke ROCm, see `f9aa3bae8c/1` ([comment](https://github.com/pytorch/pytorch/pull/151773#issuecomment-2892587039))	2025-05-20 00:42:53 +00:00
xinan.lin	f9aa3bae8c	[Inductor][XPU] Fallback bmm to mm when batch == 1, align with cuda. (#153770 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153770 Approved by: https://github.com/NikhilAPatel, https://github.com/EikanWang, https://github.com/jansel	2025-05-19 23:56:20 +00:00
PyTorch MergeBot	d81217be2e	Revert "Improve torch.ops typing (#153558 )" This reverts commit c5cba39d469151895cd0ecf7673b98e5072b69c2. Reverted https://github.com/pytorch/pytorch/pull/153558 on behalf of https://github.com/yangw-dev due to Your diff will not be landed to fbcode since we suspect it caused the following breakage in an internal test:[D75007157](https://www.internalfb.com/diff/D75007157) for instance: tests_gpu/lookup_gpu_index_test.py:232:8 Undefined attribute [16]: torch._ops._OpNamespace has no attribute simple_index_mm_batch ([comment](https://github.com/pytorch/pytorch/pull/153558#issuecomment-2892506789))	2025-05-19 23:32:36 +00:00
Menglu Yu	701e22112d	[PT2][Optimus][Observability] Refactor the logging to avoid excessive tlparse log (#153584 ) Summary: context: https://fb.workplace.com/groups/943185660584207/permalink/1215335930035844/ Test Plan: before: aps-aps-ig_v4_2t_2_make_baseline_30batch-735703723-f735706162 tlparse: https://manifold.edge.x2p.facebook.net/v0/read/tree/logs/aps-aps-ig_v4_2t_2_make_baseline_30batch-735703723-f735706162/attempt_0/version_0/rank_0/index.html?bucketName=tlparse_reports&apiKey=tlparse_reports-key&withPayload=1&timeoutMsec=10000&fbclid=IwZXh0bgNhZW0CMTEAAR575JfJZUtE7kQCqzIZVCYomv1q03JzuMFVok8qDA_FuGC8oZ6rhhb2EziSQA_aem_abITQJZQP45t51_r-J-cFw Differential Revision: D74776025 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153584 Approved by: https://github.com/jamesjwu	2025-05-19 22:57:29 +00:00
Natalia Gimelshein	c3e14ecdcd	[CachingHostAllocator] guard accesses to use_host_register by mutex (#153845 ) Per title Pull Request resolved: https://github.com/pytorch/pytorch/pull/153845 Approved by: https://github.com/mradmila, https://github.com/jeffdaily	2025-05-19 22:39:13 +00:00
Nikita Shulga	41564803c2	[Docs] Mention `version.txt` change for patch releases (#153860 ) Part of https://github.com/pytorch/pytorch/issues/151425 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153860 Approved by: https://github.com/Skylion007, https://github.com/seemethere	2025-05-19 22:35:33 +00:00
Nikita Shulga	08e716fc70	[BE] Fix `-Wextra-semi` warning (#153887 ) Introduced by https://github.com/pytorch/pytorch/pull/153645 Semicolon is not needed after closing curly bracket defining a class method. Not sure why CI did not catch it, but my local builds are now erroring out with ``` [19/97] Building CXX object caffe2/CMakeFiles/torch_cpu.dir/__/torch/csrc/jit/passes/dead_code_elimination.cpp.o In file included from /Users/nshulga/git/pytorch/pytorch/torch/csrc/jit/passes/dead_code_elimination.cpp:4: /Users/nshulga/git/pytorch/pytorch/torch/csrc/jit/ir/alias_analysis.h:356:64: warning: extra ';' after member function definition [-Wextra-semi] 356 \| ValueAndMemoryLocationSet(const AliasDb* db) : aliasDb_(db){}; \| ^ ``` Fixes #ISSUE_NUMBER Pull Request resolved: https://github.com/pytorch/pytorch/pull/153887 Approved by: https://github.com/wdvr, https://github.com/davidberard98	2025-05-19 22:25:03 +00:00
Dmitry Nikolaev	f419067e50	[ROCm] improve sparse addmm, enable complex (#153262 ) PR to: - enable complex data types for sparse matmul on ROCm - fix sparse addmm/baddbmm on ROCm - fix sparse hipification for ROCm - fix/enable sparse tests on ROCm (~40 tests total): ``` test_sparse_csr.py::TestSparseCSRCUDA::test_bmm_cuda_* test_sparse.py::TestSparseCUDA::test_sparse_matmul_cuda_* test_sparse_csr.py::TestSparseCSRCUDA::test_mm_cuda_float64 test_sparse_csr.py::TestSparseCSRCUDA::test_addmm_all_sparse_csr_SparseCS* test_sparse_csr.py::TestSparseCSRCUDA::test_addmm_sizes_all_sparse_csr_* ``` Pull Request resolved: https://github.com/pytorch/pytorch/pull/153262 Approved by: https://github.com/jeffdaily, https://github.com/pruthvistony	2025-05-19 22:23:18 +00:00
Catherine Lee	91cc93deae	[CI] Reuse old whl (#153838 ) ~50% of commits on main only touch python files unrelated to the object files in the whl, meaning that we could reuse old whls and put the current commit's python files into the whl. This PR does that in CI by identifying a previous job whose artifact and whls binaries can be reused. See https://docs.google.com/document/d/1nQ1FNJqnJuSFRiM2HvQ27zg6Vm-77n7LECp30zYfTDk/edit?tab=t.icom2lesr6es for more details? To reuse: * the changed files between the whl's commit and the current commit can only be python files in test/ or torch/ and not in torch/csrc * not on main branch or release branch * ci-force-rebuild not on PR * special abort issue is closed * artifact should exist Pros: * build time -> 6 min whenever this can be done Cons: * not sure if I have the right files * version + whl name still remains the same Testing: Unfortunately this PR's changed files are not on the list of acceptable changed files for reusing the whl, so I've been mangling it on other PRs to get things like https://github.com/pytorch/pytorch/actions/runs/15119214901/job/42497650394?pr=147470 (It is enabled on linux-focal-cuda12.6-py3.10-gcc11 / build and there are changes in common_utils.py to make sure the copying of python takes effect) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153838 Approved by: https://github.com/malfet	2025-05-19 21:47:33 +00:00
Nikita Shulga	c0343b1539	Fix profiler on cpython-3.13 (#153848 ) Per [PEP 667](https://peps.python.org/pep-0667/) `PyFrame_GetLocals` no longer returns dict, but rather instance of `PyFrameLocalsProxy_Type`, so calling `PyDict_GetItemString` is no longer valid(it will always return None) and must be replaced with `PyMapping_GetItemString` Tested by partially reverting https://github.com/pytorch/pytorch/pull/141674 full revert will be done in the followup PR Fixes https://github.com/pytorch/pytorch/issues/148273 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153848 Approved by: https://github.com/Skylion007	2025-05-19 21:20:53 +00:00
PyTorch MergeBot	8c40c9ffcb	Revert "[CI] Reuse old whl (#153838 )" This reverts commit cc48550e6f6fa8888b7d90d030b36c4e6d6581ab. Reverted https://github.com/pytorch/pytorch/pull/153838 on behalf of https://github.com/clee2000 due to testing on main is hard ([comment](https://github.com/pytorch/pytorch/pull/153838#issuecomment-2892272494))	2025-05-19 21:13:27 +00:00
David Berard	a237831bc2	[JIT] Optimize DCE by storing a MemoryLocations for an entire set<Value> (#153645 ) Summary: TL;DR: make DCE faster by replacing a Set<Value> with a MemoryLocations sparse bitset (representing all the memory locations stored by the collection of all values in the set). Details The goal of this PR is to optimize this function from AliasDb: ``` bool AliasDb::writesToAlias(Node* n, const ValueSet& vs) const { const auto writtenTo = getWrites(n); if (writtenTo.empty()) { return false; } MemoryLocations locs; for (const auto v : vs) { auto it = elementMap_.find(v); if (it != elementMap_.end()) { const auto& vlocs = memoryDAG_->getMemoryLocations(it->second); if (writtenTo.intersects(vlocs)) { return true; } } } return false; } ``` In the DCE use case, we have a ValueSet of live values, into which we insert `Value`s; and sometimes need to check whether a node mutates any of the live values using `writesToAlias`. Looping through all the values in the ValueSet and indexing into the elementMap_ is slow; so if we can pre-compute the MemoryLocations set, this speeds up the function. In some large model examples, I see ~15-25x speedups from this change. Implementation: To avoid exposing too many details of AliasDb, I introduce a friend class `ValueAndMemoryLocationSet`, which is an insert-only set of Values, which also maintains the corresponding MemoryLocations. Then in AliasDb, I use `ValueAndMemoryLocationSet` if we're using AliasDb for analysis, and otherwise use a `Set<Value>` if we don't have AliasDb. Test Plan: Rely on unit tests. Differential Revision: D74827086 Pull Request resolved: https://github.com/pytorch/pytorch/pull/153645 Approved by: https://github.com/eellison	2025-05-19 21:04:59 +00:00
Catherine Lee	cc48550e6f	[CI] Reuse old whl (#153838 ) ~50% of commits on main only touch python files unrelated to the object files in the whl, meaning that we could reuse old whls and put the current commit's python files into the whl. This PR does that in CI by identifying a previous job whose artifact and whls binaries can be reused. See https://docs.google.com/document/d/1nQ1FNJqnJuSFRiM2HvQ27zg6Vm-77n7LECp30zYfTDk/edit?tab=t.icom2lesr6es for more details? To reuse: * the changed files between the whl's commit and the current commit can only be python files in test/ or torch/ and not in torch/csrc * not on main branch or release branch * ci-force-rebuild not on PR * special abort issue is closed * artifact should exist Pros: * build time -> 6 min whenever this can be done Cons: * not sure if I have the right files * version + whl name still remains the same Testing: Unfortunately this PR's changed files are not on the list of acceptable changed files for reusing the whl, so I've been mangling it on other PRs to get things like https://github.com/pytorch/pytorch/actions/runs/15119214901/job/42497650394?pr=147470 (It is enabled on linux-focal-cuda12.6-py3.10-gcc11 / build and there are changes in common_utils.py to make sure the copying of python takes effect) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153838 Approved by: https://github.com/malfet	2025-05-19 20:56:44 +00:00
Laith Sakka	9180bb187c	Cache code generation during triton template expansion and enable it for mm_template. (#151773 ) In a model, we see ~~ 40% of the time in mm/addmm tuning. The model have 2000 mm, many of which receives the same input shapes. with autotune enabled, this become expensive, while we already cache auto tuning results, we did not used to cache the generation of the python code and the loading for each config that we autotune on. This diff handles the code generation part (template expansions) a previous diff handled the loading part. This is expected to save 20% of the model I am working on. How do we do the caching? For a given configurations and input layout, the generated code is always the same. One caveat is that some other information collected during code generation are input dependent (namely depends on inputs names and symbol names in inputs). and not just layout. ! To handle those we use a record and replay approach, where we record the functions that are called during code generation that effect those outputs and replay them at a cache hit. Effect on the current benchmark on a local run on dev server. mm_loop. 24115830838 -> 18362098019 mm_loop_dynamic 30506097176-> 25697270062 Pull Request resolved: https://github.com/pytorch/pytorch/pull/151773 Approved by: https://github.com/eellison	2025-05-19 20:38:04 +00:00
PyTorch MergeBot	1ccacc028d	Revert "[CI] Reuse old whl (#153838 )" This reverts commit 0716acff3a3692daaf31d97f91ae5aee70f10f24. Reverted https://github.com/pytorch/pytorch/pull/153838 on behalf of https://github.com/clee2000 due to forgot to comment some stuff out ([comment](https://github.com/pytorch/pytorch/pull/153838#issuecomment-2892195387))	2025-05-19 20:33:14 +00:00
Mikayla Gawarecki	6383ddcfa4	Update serialization docs (#153631 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153631 Approved by: https://github.com/albanD	2025-05-19 20:22:07 +00:00
Nikita Shulga	2fcbb903cb	[BE][EZ] Delete unsued conda-env-IOS.txt (#153849 ) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153849 Approved by: https://github.com/janeyx99, https://github.com/seemethere, https://github.com/Skylion007, https://github.com/ZainRizvi	2025-05-19 20:06:38 +00:00
eqy	6ae0c42278	[CUDA][cuBLASLt] Respect `allow[FP16/BF16]ReductionCuBLAS` in `cuBLASLt` (#153095 ) cuBLASLt matmuls have been silently allowing all reduction types, which meant that e.g., `allow_fp16_reduced_precision_reduction = False` had no effect. In practice split-K with reduced precision reductions were unlikely to happen as the default `CUBLASLT_WORKSPACE_SIZE` of 1MiB tends to prevent this. However this isn't guaranteed and we are on the path to increasing the default workspace size following #151163 This setting is effectively already tested in e.g., `test_cublas_addmm_size_100_cuda_float16` and `test_cublas_addmm_size_100_cuda_bfloat16` but the backend selection is not deterministic. Running the full `test_matmul_cuda.py` seems to exercise the Lt interface, but running a standalone test does not (apparently due to spurious alignment differences). Pull Request resolved: https://github.com/pytorch/pytorch/pull/153095 Approved by: https://github.com/cyyever, https://github.com/Skylion007	2025-05-19 20:05:37 +00:00
Aaron Gokaslan	e581e1c0f4	[BE][Ez]: Propogate some nodiscard in RNN (#153836 ) Follow up @cyyever #153805 to propagate [nodiscard] from the empty() method call. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153836 Approved by: https://github.com/eqy	2025-05-19 19:55:45 +00:00
Catherine Lee	0716acff3a	[CI] Reuse old whl (#153838 ) ~50% of commits on main only touch python files unrelated to the object files in the whl, meaning that we could reuse old whls and put the current commit's python files into the whl. This PR does that in CI by identifying a previous job whose artifact and whls binaries can be reused. See https://docs.google.com/document/d/1nQ1FNJqnJuSFRiM2HvQ27zg6Vm-77n7LECp30zYfTDk/edit?tab=t.icom2lesr6es for more details? To reuse: * the changed files between the whl's commit and the current commit can only be python files in test/ or torch/ and not in torch/csrc * not on main branch or release branch * ci-force-rebuild not on PR * special abort issue is closed * artifact should exist Pros: * build time -> 6 min whenever this can be done Cons: * not sure if I have the right files * version + whl name still remains the same Testing: Unfortunately this PR's changed files are not on the list of acceptable changed files for reusing the whl, so I've been mangling it on other PRs to get things like https://github.com/pytorch/pytorch/actions/runs/15119214901/job/42497650394?pr=147470 (It is enabled on linux-focal-cuda12.6-py3.10-gcc11 / build and there are changes in common_utils.py to make sure the copying of python takes effect) Pull Request resolved: https://github.com/pytorch/pytorch/pull/153838 Approved by: https://github.com/malfet	2025-05-19 19:26:08 +00:00
PyTorch MergeBot	674a85cf26	Revert "[Distributed][CI] Rework continuous TestCase (#153653 )" This reverts commit 0d5c628a6e96e0a960af39d1d0de4bf04df69c39. Reverted https://github.com/pytorch/pytorch/pull/153653 on behalf of https://github.com/kwen2501 due to More fixes needed ([comment](https://github.com/pytorch/pytorch/pull/153653#issuecomment-2891931028))	2025-05-19 18:29:27 +00:00
Ke Wen	0d5c628a6e	[Distributed][CI] Rework continuous TestCase (#153653 ) 1. Reworked `MultiProcContinousTest` to spawn processes during `setUpClass` instead of `main` (so that we can support multiple TestClass'es in one file). 2. The child processes are now an infinite loop, monitoring test IDs passed from main process via a task queue. Reciprocally, the child processes inform the main process completion of a test via a completion queue. 3. Added a test template. Pull Request resolved: https://github.com/pytorch/pytorch/pull/153653 Approved by: https://github.com/d4l3k, https://github.com/fegin, https://github.com/fduwjj	2025-05-19 18:20:42 +00:00

6345 changed files with 266959 additions and 106791 deletions

2

.bazelrc

View File

 @ -2,7 +2,7 @@ build --cxxopt=--std=c++17
 build --copt=-I.
 # Bazel does not support including its cc_library targets as system
 # headers. We work around this for generated code
 # (e.g. c10/macros/cmake_macros.h) by making the generated directory a
 # (e.g. torch/headeronly/macros/cmake_macros.h) by making the generated directory a
 # system include path.
 build --copt=-isystem --copt bazel-out/k8-fastbuild/bin
 build --copt=-isystem --copt bazel-out/darwin-fastbuild/bin

									
										7

.ci/aarch64_linux/aarch64_ci_build.sh
									
												View File
												
				@ -3,10 +3,8 @@ set -eux -o pipefail

				GPU_ARCH_VERSION=${GPU_ARCH_VERSION:-}

				if [[ "$GPU_ARCH_VERSION" == *"12.6"* ]]; then

				    export TORCH_CUDA_ARCH_LIST="9.0"

				elif [[ "$GPU_ARCH_VERSION" == *"12.8"* ]]; then

				    export TORCH_CUDA_ARCH_LIST="9.0;10.0;12.0"

				if [[ "$GPU_ARCH_VERSION" == *"12.9"* ]]; then

				    export TORCH_CUDA_ARCH_LIST="8.0;9.0;10.0;12.0"

				fi

				SCRIPTPATH="$( cd -- "$(dirname "$0")" >/dev/null 2>&1 ; pwd -P )"

				@ -27,6 +25,7 @@ if [ "$DESIRED_CUDA" = "cpu" ]; then

				    USE_PRIORITIZED_TEXT_FOR_LD=1 python /pytorch/.ci/aarch64_linux/aarch64_wheel_ci_build.py --enable-mkldnn

				else

				    echo "BASE_CUDA_VERSION is set to: $DESIRED_CUDA"

				    export USE_SYSTEM_NCCL=1

				    #USE_PRIORITIZED_TEXT_FOR_LD for enable linker script optimization https://github.com/pytorch/pytorch/pull/121975/files

				    USE_PRIORITIZED_TEXT_FOR_LD=1 python /pytorch/.ci/aarch64_linux/aarch64_wheel_ci_build.py --enable-mkldnn --enable-cuda

				fi

									
										101

.ci/aarch64_linux/aarch64_wheel_ci_build.py
									
												View File
												
				@ -31,33 +31,47 @@ def build_ArmComputeLibrary() -> None:

				        "build=native",

				    ]

				    acl_install_dir = "/acl"

				    acl_checkout_dir = "ComputeLibrary"

				    os.makedirs(acl_install_dir)

				    check_call(

				        [

				            "git",

				            "clone",

				            "https://github.com/ARM-software/ComputeLibrary.git",

				            "-b",

				            "v25.02",

				            "--depth",

				            "1",

				            "--shallow-submodules",

				        ]

				    )

				    acl_checkout_dir = os.getenv("ACL_SOURCE_DIR", "ComputeLibrary")

				    if os.path.isdir(acl_install_dir):

				        shutil.rmtree(acl_install_dir)

				    if not os.path.isdir(acl_checkout_dir) or not len(os.listdir(acl_checkout_dir)):

				        check_call(

				            [

				                "git",

				                "clone",

				                "https://github.com/ARM-software/ComputeLibrary.git",

				                "-b",

				                "v25.02",

				                "--depth",

				                "1",

				                "--shallow-submodules",

				            ]

				        )

				    check_call(

				        ["scons", "Werror=1", "-j8", f"build_dir=/{acl_install_dir}/build"]

				        + acl_build_flags,

				        ["scons", "Werror=1", f"-j{os.cpu_count()}"] + acl_build_flags,

				        cwd=acl_checkout_dir,

				    )

				    for d in ["arm_compute", "include", "utils", "support", "src"]:

				    for d in ["arm_compute", "include", "utils", "support", "src", "build"]:

				        shutil.copytree(f"{acl_checkout_dir}/{d}", f"{acl_install_dir}/{d}")

				def update_wheel(wheel_path, desired_cuda) -> None:

				def replace_tag(filename) -> None:

				    with open(filename) as f:

				        lines = f.readlines()

				    for i, line in enumerate(lines):

				        if line.startswith("Tag:"):

				            lines[i] = line.replace("-linux_", "-manylinux_2_28_")

				            print(f"Updated tag from {line} to {lines[i]}")

				            break

				    with open(filename, "w") as f:

				        f.writelines(lines)

				def package_cuda_wheel(wheel_path, desired_cuda) -> None:

				    """

				    Update the cuda wheel libraries

				    Package the cuda wheel libraries

				    """

				    folder = os.path.dirname(wheel_path)

				    wheelname = os.path.basename(wheel_path)

				@ -65,6 +79,7 @@ def update_wheel(wheel_path, desired_cuda) -> None:

				    os.system(f"unzip {wheel_path} -d {folder}/tmp")

				    libs_to_copy = [

				        "/usr/local/cuda/extras/CUPTI/lib64/libcupti.so.12",

				        "/usr/local/cuda/extras/CUPTI/lib64/libnvperf_host.so",

				        "/usr/local/cuda/lib64/libcudnn.so.9",

				        "/usr/local/cuda/lib64/libcublas.so.12",

				        "/usr/local/cuda/lib64/libcublasLt.so.12",

				@ -74,7 +89,7 @@ def update_wheel(wheel_path, desired_cuda) -> None:

				        "/usr/local/cuda/lib64/libcusparseLt.so.0",

				        "/usr/local/cuda/lib64/libcusolver.so.11",

				        "/usr/local/cuda/lib64/libcurand.so.10",

				        "/usr/local/cuda/lib64/libnvToolsExt.so.1",

				        "/usr/local/cuda/lib64/libnccl.so.2",

				        "/usr/local/cuda/lib64/libnvJitLink.so.12",

				        "/usr/local/cuda/lib64/libnvrtc.so.12",

				        "/usr/local/cuda/lib64/libcudnn_adv.so.9",

				@ -88,30 +103,19 @@ def update_wheel(wheel_path, desired_cuda) -> None:

				        "/usr/lib64/libgfortran.so.5",

				        "/acl/build/libarm_compute.so",

				        "/acl/build/libarm_compute_graph.so",

				        "/usr/local/lib/libnvpl_lapack_lp64_gomp.so.0",

				        "/usr/local/lib/libnvpl_blas_lp64_gomp.so.0",

				        "/usr/local/lib/libnvpl_lapack_core.so.0",

				        "/usr/local/lib/libnvpl_blas_core.so.0",

				    ]

				    if enable_cuda:

				    if "129" in desired_cuda:

				        libs_to_copy += [

				            "/usr/local/lib/libnvpl_lapack_lp64_gomp.so.0",

				            "/usr/local/lib/libnvpl_blas_lp64_gomp.so.0",

				            "/usr/local/lib/libnvpl_lapack_core.so.0",

				            "/usr/local/lib/libnvpl_blas_core.so.0",

				        ]

				        if "126" in desired_cuda:

				            libs_to_copy += [

				                "/usr/local/cuda/lib64/libnvrtc-builtins.so.12.6",

				                "/usr/local/cuda/lib64/libcufile.so.0",

				                "/usr/local/cuda/lib64/libcufile_rdma.so.1",

				            ]

				        elif "128" in desired_cuda:

				            libs_to_copy += [

				                "/usr/local/cuda/lib64/libnvrtc-builtins.so.12.8",

				                "/usr/local/cuda/lib64/libcufile.so.0",

				                "/usr/local/cuda/lib64/libcufile_rdma.so.1",

				            ]

				    else:

				        libs_to_copy += [

				            "/opt/OpenBLAS/lib/libopenblas.so.0",

				            "/usr/local/cuda/lib64/libnvrtc-builtins.so.12.9",

				            "/usr/local/cuda/lib64/libcufile.so.0",

				            "/usr/local/cuda/lib64/libcufile_rdma.so.1",

				        ]

				    # Copy libraries to unzipped_folder/a/lib

				    for lib_path in libs_to_copy:

				        lib_name = os.path.basename(lib_path)

				@ -120,6 +124,13 @@ def update_wheel(wheel_path, desired_cuda) -> None:

				            f"cd {folder}/tmp/torch/lib/; "

				            f"patchelf --set-rpath '$ORIGIN' --force-rpath {folder}/tmp/torch/lib/{lib_name}"

				        )

				    # Make sure the wheel is tagged with manylinux_2_28

				    for f in os.scandir(f"{folder}/tmp/"):

				        if f.is_dir() and f.name.endswith(".dist-info"):

				            replace_tag(f"{f.path}/WHEEL")

				            break

				    os.mkdir(f"{folder}/cuda_wheel")

				    os.system(f"cd {folder}/tmp/; zip -r {folder}/cuda_wheel/{wheelname} *")

				    shutil.move(

				@ -194,8 +205,10 @@ if __name__ == "__main__":

				    ).decode()

				    print("Building PyTorch wheel")

				    build_vars = "MAX_JOBS=5 CMAKE_SHARED_LINKER_FLAGS=-Wl,-z,max-page-size=0x10000 "

				    os.system("cd /pytorch; python setup.py clean")

				    build_vars = "CMAKE_SHARED_LINKER_FLAGS=-Wl,-z,max-page-size=0x10000 "

				    # MAX_JOB=5 is not required for CPU backend (see commit 465d98b)

				    if enable_cuda:

				        build_vars = "MAX_JOBS=5 " + build_vars

				    override_package_version = os.getenv("OVERRIDE_PACKAGE_VERSION")

				    desired_cuda = os.getenv("DESIRED_CUDA")

				@ -242,6 +255,6 @@ if __name__ == "__main__":

				        print("Updating Cuda Dependency")

				        filename = os.listdir("/pytorch/dist/")

				        wheel_path = f"/pytorch/dist/{filename[0]}"

				        update_wheel(wheel_path, desired_cuda)

				        package_cuda_wheel(wheel_path, desired_cuda)

				    pytorch_wheel_name = complete_wheel("/pytorch/")

				    print(f"Build Complete. Created {pytorch_wheel_name}..")

									
										6

.ci/caffe2/test.sh
									
												View File
												
				@ -5,7 +5,7 @@ source "$(dirname "${BASH_SOURCE[0]}")/common.sh"

				if [[ ${BUILD_ENVIRONMENT} == *onnx* ]]; then

				  pip install click mock tabulate networkx==2.0

				  pip -q install --user "file:///var/lib/jenkins/workspace/third_party/onnx#egg=onnx"

				  pip -q install "file:///var/lib/jenkins/workspace/third_party/onnx#egg=onnx"

				fi

				# Skip tests in environments where they are not built/applicable

				@ -147,8 +147,8 @@ export DNNL_MAX_CPU_ISA=AVX2

				if [[ "${SHARD_NUMBER:-1}" == "1" ]]; then

				  # TODO(sdym@meta.com) remove this when the linked issue resolved.

				  # py is temporary until https://github.com/Teemu/pytest-sugar/issues/241 is fixed

				  pip install --user py==1.11.0

				  pip install --user pytest-sugar

				  pip install py==1.11.0

				  pip install pytest-sugar

				  # NB: Warnings are disabled because they make it harder to see what

				  # the actual erroring test is

				  "$PYTHON" \

									
										101

.ci/docker/README.md
									
												View File
												
				@ -36,3 +36,104 @@ See `build.sh` for valid build environments (it's the giant switch).

				# Set flags (see build.sh) and build image

				sudo bash -c 'TRITON=1 ./build.sh pytorch-linux-bionic-py3.8-gcc9 -t myimage:latest

				```

				## [Guidance] Adding a New Base Docker Image

				### Background

				The base Docker images in directory `.ci/docker/` are built by the `docker-builds.yml` workflow. Those images are used throughout the PyTorch CI/CD pipeline. You should only create or modify a base Docker image if you need specific environment changes or dependencies before building PyTorch on CI.

				1. **Automatic Rebuilding**:

				   - The Docker image building process is triggered automatically when changes are made to files in the `.ci/docker/*` directory

				   - This ensures all images stay up-to-date with the latest dependencies and configurations

				2. **Image Reuse in PyTorch Build Workflows** (example: linux-build):

				   - The images generated by `docker-builds.yml` are reused in `_linux-build.yml` through the `calculate-docker-image` step

				   - The `_linux-build.yml` workflow:

				     - Pulls the Docker image determined by the `calculate-docker-image` step

				     - Runs a Docker container with that image

				     - Executes `.ci/pytorch/build.sh` inside the container to build PyTorch

				3. **Usage in Test Workflows** (example: linux-test):

				   - The same Docker images are also used in `_linux-test.yml` for running tests

				   - The `_linux-test.yml` workflow follows a similar pattern:

				     - It uses the `calculate-docker-image` step to determine which Docker image to use

				     - It pulls the Docker image and runs a container with that image

				     - It installs the wheels from the artifacts generated by PyTorch build jobs

				     - It executes test scripts (like `.ci/pytorch/test.sh` or `.ci/pytorch/multigpu-test.sh`) inside the container

				### Understanding File Purposes

				#### `.ci/docker/build.sh` vs `.ci/pytorch/build.sh`

				- **`.ci/docker/build.sh`**:

				  - Used for building base Docker images

				  - Executed by the `docker-builds.yml` workflow to pre-build Docker images for CI

				  - Contains configurations for different Docker build environments

				- **`.ci/pytorch/build.sh`**:

				  - Used for building PyTorch inside a Docker container

				  - Called by workflows like `_linux-build.yml` after the Docker container is started

				  - Builds PyTorch wheels and other artifacts

				#### `.ci/docker/ci_commit_pins/` vs `.github/ci_commit_pins`

				- **`.ci/docker/ci_commit_pins/`**:

				  - Used for pinning dependency versions during base Docker image building

				  - Ensures consistent environments for building PyTorch

				  - Changes here trigger base Docker image rebuilds

				- **`.github/ci_commit_pins`**:

				  - Used for pinning dependency versions during PyTorch building and tests

				  - Ensures consistent dependencies for PyTorch across different builds

				  - Used by build scripts running inside Docker containers

				### Step-by-Step Guide for Adding a New Base Docker Image

				#### 1. Add Pinned Commits (If Applicable)

				We use pinned commits for build stability. The `nightly.yml` workflow checks and updates pinned commits for certain repository dependencies daily.

				If your new Docker image needs a library installed from a specific pinned commit or built from source:

				1. Add the repository you want to track in `nightly.yml` and `merge-rules.yml`

				2. Add the initial pinned commit in `.ci/docker/ci_commit_pins/`. The text filename should match the one defined in step 1

				#### 2. Configure the Base Docker Image

				1. **Add new Base Docker image configuration** (if applicable):

				   Add the configuration in `.ci/docker/build.sh`. For example:

				   ```bash

				   pytorch-linux-jammy-cuda12.8-cudnn9-py3.12-gcc11-new1)

				     CUDA_VERSION=12.8.1

				     ANACONDA_PYTHON_VERSION=3.12

				     GCC_VERSION=11

				     VISION=yes

				     KATEX=yes

				     UCX_COMMIT=${_UCX_COMMIT}

				     UCC_COMMIT=${_UCC_COMMIT}

				     TRITON=yes

				     NEW_ARG_1=yes

				     ;;

				   ```

				2. **Add build arguments to Docker build command**:

				   If you're introducing a new argument to the Docker build, make sure to add it in the Docker build step in `.ci/docker/build.sh`:

				   ```bash

				   docker build \

				      ....

				      --build-arg "NEW_ARG_1=${NEW_ARG_1}"

				   ```

				3. **Update Dockerfile logic**:

				   Update the Dockerfile to use the new argument. For example, in `ubuntu/Dockerfile`:

				   ```dockerfile

				   ARG NEW_ARG_1

				   # Set up environment for NEW_ARG_1

				   RUN if [ -n "${NEW_ARG_1}" ]; then bash ./do_something.sh; fi

				   ```

				4. **Add the Docker configuration** in `.github/workflows/docker-builds.yml`:

				   The `docker-builds.yml` workflow pre-builds the Docker images whenever changes occur in the `.ci/docker/` directory. This includes the

				   pinned commit updates.

									
										17

.ci/docker/almalinux/Dockerfile
									
												View File
												
				@ -1,7 +1,7 @@

				ARG CUDA_VERSION=12.4

				ARG CUDA_VERSION=12.6

				ARG BASE_TARGET=cuda${CUDA_VERSION}

				ARG ROCM_IMAGE=rocm/dev-almalinux-8:6.3-complete

				FROM amd64/almalinux:8 as base

				FROM amd64/almalinux:8.10-20250519 as base

				ENV LC_ALL en_US.UTF-8

				ENV LANG en_US.UTF-8

				@ -11,6 +11,8 @@ ARG DEVTOOLSET_VERSION=11

				RUN yum -y update

				RUN yum -y install epel-release

				# install glibc-langpack-en make sure en_US.UTF-8 locale is available

				RUN yum -y install glibc-langpack-en

				RUN yum install -y sudo wget curl perl util-linux xz bzip2 git patch which perl zlib-devel openssl-devel yum-utils autoconf automake make gcc-toolset-${DEVTOOLSET_VERSION}-toolchain

				# Just add everything as a safe.directory for git since these will be used in multiple places with git

				RUN git config --global --add safe.directory '*'

				@ -50,10 +52,6 @@ ENV CUDA_VERSION=${CUDA_VERSION}

				# Make things in our path by default

				ENV PATH=/usr/local/cuda-${CUDA_VERSION}/bin:$PATH

				FROM cuda as cuda11.8

				RUN bash ./install_cuda.sh 11.8

				ENV DESIRED_CUDA=11.8

				FROM cuda as cuda12.6

				RUN bash ./install_cuda.sh 12.6

				ENV DESIRED_CUDA=12.6

				@ -62,6 +60,10 @@ FROM cuda as cuda12.8

				RUN bash ./install_cuda.sh 12.8

				ENV DESIRED_CUDA=12.8

				FROM cuda as cuda12.9

				RUN bash ./install_cuda.sh 12.9

				ENV DESIRED_CUDA=12.9

				FROM ${ROCM_IMAGE} as rocm

				ENV PYTORCH_ROCM_ARCH="gfx900;gfx906;gfx908;gfx90a;gfx942;gfx1030;gfx1100;gfx1101;gfx1102;gfx1200;gfx1201"

				ADD ./common/install_mkl.sh install_mkl.sh

				@ -76,7 +78,8 @@ RUN bash ./install_mnist.sh

				FROM base as all_cuda

				COPY --from=cuda11.8  /usr/local/cuda-11.8 /usr/local/cuda-11.8

				COPY --from=cuda12.6  /usr/local/cuda-12.6 /usr/local/cuda-12.6

				COPY --from=cuda12.4  /usr/local/cuda-12.8 /usr/local/cuda-12.8

				COPY --from=cuda12.8  /usr/local/cuda-12.8 /usr/local/cuda-12.8

				COPY --from=cuda12.9  /usr/local/cuda-12.9 /usr/local/cuda-12.9

				# Final step

				FROM ${BASE_TARGET} as final

									
										188

.ci/docker/build.sh
									
												View File
												
				@ -50,30 +50,23 @@ if [[ "$image" == *xla* ]]; then

				  exit 0

				fi

				if [[ "$image" == *-focal* ]]; then

				  UBUNTU_VERSION=20.04

				elif [[ "$image" == *-jammy* ]]; then

				if [[ "$image" == *-jammy* ]]; then

				  UBUNTU_VERSION=22.04

				elif [[ "$image" == *-noble* ]]; then

				  UBUNTU_VERSION=24.04

				elif [[ "$image" == *ubuntu* ]]; then

				  extract_version_from_image_name ubuntu UBUNTU_VERSION

				elif [[ "$image" == *centos* ]]; then

				  extract_version_from_image_name centos CENTOS_VERSION

				fi

				if [ -n "${UBUNTU_VERSION}" ]; then

				  OS="ubuntu"

				elif [ -n "${CENTOS_VERSION}" ]; then

				  OS="centos"

				else

				  echo "Unable to derive operating system base..."

				  exit 1

				fi

				DOCKERFILE="${OS}/Dockerfile"

				# When using ubuntu - 22.04, start from Ubuntu docker image, instead of nvidia/cuda docker image.

				if [[ "$image" == *cuda* && "$UBUNTU_VERSION" != "22.04" ]]; then

				  DOCKERFILE="${OS}-cuda/Dockerfile"

				elif [[ "$image" == *rocm* ]]; then

				if [[ "$image" == *rocm* ]]; then

				  DOCKERFILE="${OS}-rocm/Dockerfile"

				elif [[ "$image" == *xpu* ]]; then

				  DOCKERFILE="${OS}-xpu/Dockerfile"

				@ -98,9 +91,8 @@ tag=$(echo $image | awk -F':' '{print $2}')

				# configuration, so we hardcode everything here rather than do it

				# from scratch

				case "$tag" in

				  pytorch-linux-focal-cuda12.6-cudnn9-py3-gcc11)

				    CUDA_VERSION=12.6.3

				    CUDNN_VERSION=9

				  pytorch-linux-jammy-cuda12.4-cudnn9-py3-gcc11)

				    CUDA_VERSION=12.4

				    ANACONDA_PYTHON_VERSION=3.10

				    GCC_VERSION=11

				    VISION=yes

				@ -109,9 +101,18 @@ case "$tag" in

				    UCC_COMMIT=${_UCC_COMMIT}

				    TRITON=yes

				    ;;

				  pytorch-linux-focal-cuda12.4-cudnn9-py3-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.4.1

				    CUDNN_VERSION=9

				  pytorch-linux-jammy-cuda12.8-cudnn9-py3-gcc11)

				    CUDA_VERSION=12.8.1

				    ANACONDA_PYTHON_VERSION=3.10

				    GCC_VERSION=11

				    VISION=yes

				    KATEX=yes

				    UCX_COMMIT=${_UCX_COMMIT}

				    UCC_COMMIT=${_UCC_COMMIT}

				    TRITON=yes

				    ;;

				  pytorch-linux-jammy-cuda12.8-cudnn9-py3-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.8.1

				    ANACONDA_PYTHON_VERSION=3.10

				    GCC_VERSION=9

				    VISION=yes

				@ -121,9 +122,8 @@ case "$tag" in

				    TRITON=yes

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-focal-cuda12.4-cudnn9-py3.12-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.4.1

				    CUDNN_VERSION=9

				  pytorch-linux-jammy-cuda12.8-cudnn9-py3.12-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.8.1

				    ANACONDA_PYTHON_VERSION=3.12

				    GCC_VERSION=9

				    VISION=yes

				@ -133,9 +133,8 @@ case "$tag" in

				    TRITON=yes

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-focal-cuda12.4-cudnn9-py3.13-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.4.1

				    CUDNN_VERSION=9

				  pytorch-linux-jammy-cuda12.8-cudnn9-py3.13-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.8.1

				    ANACONDA_PYTHON_VERSION=3.13

				    GCC_VERSION=9

				    VISION=yes

				@ -145,56 +144,18 @@ case "$tag" in

				    TRITON=yes

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-focal-cuda12.6-cudnn9-py3-gcc9)

				    CUDA_VERSION=12.6.3

				    CUDNN_VERSION=9

				    ANACONDA_PYTHON_VERSION=3.10

				    GCC_VERSION=9

				    VISION=yes

				    KATEX=yes

				    UCX_COMMIT=${_UCX_COMMIT}

				    UCC_COMMIT=${_UCC_COMMIT}

				    TRITON=yes

				    ;;

				  pytorch-linux-focal-cuda12.6-cudnn9-py3-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.6.3

				    CUDNN_VERSION=9

				    ANACONDA_PYTHON_VERSION=3.10

				    GCC_VERSION=9

				    VISION=yes

				    KATEX=yes

				    UCX_COMMIT=${_UCX_COMMIT}

				    UCC_COMMIT=${_UCC_COMMIT}

				    TRITON=yes

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-focal-cuda12.6-cudnn9-py3.12-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.6.3

				    CUDNN_VERSION=9

				  pytorch-linux-jammy-cuda12.8-cudnn9-py3.12-gcc11-vllm)

				    CUDA_VERSION=12.8.1

				    ANACONDA_PYTHON_VERSION=3.12

				    GCC_VERSION=9

				    GCC_VERSION=11

				    VISION=yes

				    KATEX=yes

				    UCX_COMMIT=${_UCX_COMMIT}

				    UCC_COMMIT=${_UCC_COMMIT}

				    TRITON=yes

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-focal-cuda12.6-cudnn9-py3.13-gcc9-inductor-benchmarks)

				    CUDA_VERSION=12.6.3

				    CUDNN_VERSION=9

				    ANACONDA_PYTHON_VERSION=3.13

				    GCC_VERSION=9

				    VISION=yes

				    KATEX=yes

				    UCX_COMMIT=${_UCX_COMMIT}

				    UCC_COMMIT=${_UCC_COMMIT}

				    TRITON=yes

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-focal-cuda11.8-cudnn9-py3-gcc9)

				    CUDA_VERSION=11.8.0

				    CUDNN_VERSION=9

				  pytorch-linux-jammy-cuda12.8-cudnn9-py3-gcc9)

				    CUDA_VERSION=12.8.1

				    ANACONDA_PYTHON_VERSION=3.10

				    GCC_VERSION=9

				    VISION=yes

				@ -203,44 +164,24 @@ case "$tag" in

				    UCC_COMMIT=${_UCC_COMMIT}

				    TRITON=yes

				    ;;

				  pytorch-linux-focal-py3-clang10-onnx)

				  pytorch-linux-jammy-py3-clang12-onnx)

				    ANACONDA_PYTHON_VERSION=3.9

				    CLANG_VERSION=10

				    CLANG_VERSION=12

				    VISION=yes

				    ONNX=yes

				    ;;

				  pytorch-linux-focal-py3.9-clang10)

				  pytorch-linux-jammy-py3.9-clang12)

				    ANACONDA_PYTHON_VERSION=3.9

				    CLANG_VERSION=10

				    CLANG_VERSION=12

				    VISION=yes

				    TRITON=yes

				    ;;

				  pytorch-linux-focal-py3.11-clang10)

				    ANACONDA_PYTHON_VERSION=3.11

				    CLANG_VERSION=10

				    VISION=yes

				    TRITON=yes

				    ;;

				  pytorch-linux-focal-py3.9-gcc9)

				    ANACONDA_PYTHON_VERSION=3.9

				    GCC_VERSION=9

				    VISION=yes

				    TRITON=yes

				    ;;

				  pytorch-linux-jammy-rocm-n-1-py3)

				    ANACONDA_PYTHON_VERSION=3.10

				    GCC_VERSION=11

				    VISION=yes

				    ROCM_VERSION=6.3

				    NINJA_VERSION=1.9.0

				    TRITON=yes

				    KATEX=yes

				    UCX_COMMIT=${_UCX_COMMIT}

				    UCC_COMMIT=${_UCC_COMMIT}

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-jammy-rocm-n-py3)

				    ANACONDA_PYTHON_VERSION=3.10

				  pytorch-linux-jammy-rocm-n-py3 | pytorch-linux-noble-rocm-n-py3)

				    if [[ $tag =~ "jammy" ]]; then

				      ANACONDA_PYTHON_VERSION=3.10

				    else

				      ANACONDA_PYTHON_VERSION=3.12

				    fi

				    GCC_VERSION=11

				    VISION=yes

				    ROCM_VERSION=6.4

				@ -251,6 +192,19 @@ case "$tag" in

				    UCC_COMMIT=${_UCC_COMMIT}

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-noble-rocm-alpha-py3)

				    ANACONDA_PYTHON_VERSION=3.12

				    GCC_VERSION=11

				    VISION=yes

				    ROCM_VERSION=7.0

				    NINJA_VERSION=1.9.0

				    TRITON=yes

				    KATEX=yes

				    UCX_COMMIT=${_UCX_COMMIT}

				    UCC_COMMIT=${_UCC_COMMIT}

				    INDUCTOR_BENCHMARKS=yes

				    PYTORCH_ROCM_ARCH="gfx90a;gfx942;gfx950"

				    ;;

				  pytorch-linux-jammy-xpu-2025.0-py3)

				    ANACONDA_PYTHON_VERSION=3.9

				    GCC_VERSION=11

				@ -267,7 +221,7 @@ case "$tag" in

				    NINJA_VERSION=1.9.0

				    TRITON=yes

				    ;;

				    pytorch-linux-jammy-py3.9-gcc11-inductor-benchmarks)

				  pytorch-linux-jammy-py3.9-gcc11-inductor-benchmarks)

				    ANACONDA_PYTHON_VERSION=3.9

				    GCC_VERSION=11

				    VISION=yes

				@ -276,25 +230,13 @@ case "$tag" in

				    DOCS=yes

				    INDUCTOR_BENCHMARKS=yes

				    ;;

				  pytorch-linux-jammy-cuda11.8-cudnn9-py3.9-clang12)

				  pytorch-linux-jammy-cuda12.8-cudnn9-py3.9-clang12)

				    ANACONDA_PYTHON_VERSION=3.9

				    CUDA_VERSION=11.8

				    CUDNN_VERSION=9

				    CUDA_VERSION=12.8.1

				    CLANG_VERSION=12

				    VISION=yes

				    TRITON=yes

				    ;;

				  pytorch-linux-jammy-py3-clang12-asan)

				    ANACONDA_PYTHON_VERSION=3.9

				    CLANG_VERSION=12

				    VISION=yes

				    TRITON=yes

				    ;;

				  pytorch-linux-jammy-py3-clang15-asan)

				    ANACONDA_PYTHON_VERSION=3.10

				    CLANG_VERSION=15

				    VISION=yes

				    ;;

				  pytorch-linux-jammy-py3-clang18-asan)

				    ANACONDA_PYTHON_VERSION=3.10

				    CLANG_VERSION=18

				@ -327,21 +269,23 @@ case "$tag" in

				    GCC_VERSION=11

				    TRITON_CPU=yes

				    ;;

				  pytorch-linux-focal-linter)

				  pytorch-linux-jammy-linter)

				    # TODO: Use 3.9 here because of this issue https://github.com/python/mypy/issues/13627.

				    # We will need to update mypy version eventually, but that's for another day. The task

				    # would be to upgrade mypy to 1.0.0 with Python 3.11

				    PYTHON_VERSION=3.9

				    ;;

				  pytorch-linux-jammy-cuda11.8-cudnn9-py3.9-linter)

				  pytorch-linux-jammy-cuda12.8-cudnn9-py3.9-linter)

				    PYTHON_VERSION=3.9

				    CUDA_VERSION=11.8

				    CUDA_VERSION=12.8.1

				    ;;

				  pytorch-linux-jammy-aarch64-py3.10-gcc11)

				    ANACONDA_PYTHON_VERSION=3.10

				    GCC_VERSION=11

				    ACL=yes

				    VISION=yes

				    CONDA_CMAKE=yes

				    OPENBLAS=yes

				    # snadampal: skipping llvm src build install because the current version

				    # from pytorch/llvm:9.0.1 is x86 specific

				    SKIP_LLVM_SRC_BUILD_INSTALL=yes

				@ -351,6 +295,8 @@ case "$tag" in

				    GCC_VERSION=11

				    ACL=yes

				    VISION=yes

				    CONDA_CMAKE=yes

				    OPENBLAS=yes

				    # snadampal: skipping llvm src build install because the current version

				    # from pytorch/llvm:9.0.1 is x86 specific

				    SKIP_LLVM_SRC_BUILD_INSTALL=yes

				@ -365,7 +311,6 @@ case "$tag" in

				    fi

				    if [[ "$image" == *cuda* ]]; then

				      extract_version_from_image_name cuda CUDA_VERSION

				      extract_version_from_image_name cudnn CUDNN_VERSION

				    fi

				    if [[ "$image" == *rocm* ]]; then

				      extract_version_from_image_name rocm ROCM_VERSION

				@ -394,14 +339,6 @@ esac

				tmp_tag=$(basename "$(mktemp -u)" | tr '[:upper:]' '[:lower:]')

				#when using cudnn version 8 install it separately from cuda

				if [[ "$image" == *cuda*  && ${OS} == "ubuntu" ]]; then

				  IMAGE_NAME="nvidia/cuda:${CUDA_VERSION}-cudnn${CUDNN_VERSION}-devel-ubuntu${UBUNTU_VERSION}"

				  if [[ ${CUDNN_VERSION} == 9 ]]; then

				    IMAGE_NAME="nvidia/cuda:${CUDA_VERSION}-devel-ubuntu${UBUNTU_VERSION}"

				  fi

				fi

				no_cache_flag=""

				progress_flag=""

				# Do not use cache and progress=plain when in CI

				@ -418,7 +355,6 @@ docker build \

				       --build-arg "LLVMDEV=${LLVMDEV:-}" \

				       --build-arg "VISION=${VISION:-}" \

				       --build-arg "UBUNTU_VERSION=${UBUNTU_VERSION}" \

				       --build-arg "CENTOS_VERSION=${CENTOS_VERSION}" \

				       --build-arg "DEVTOOLSET_VERSION=${DEVTOOLSET_VERSION}" \

				       --build-arg "GLIBC_VERSION=${GLIBC_VERSION}" \

				       --build-arg "CLANG_VERSION=${CLANG_VERSION}" \

				@ -426,9 +362,6 @@ docker build \

				       --build-arg "PYTHON_VERSION=${PYTHON_VERSION}" \

				       --build-arg "GCC_VERSION=${GCC_VERSION}" \

				       --build-arg "CUDA_VERSION=${CUDA_VERSION}" \

				       --build-arg "CUDNN_VERSION=${CUDNN_VERSION}" \

				       --build-arg "TENSORRT_VERSION=${TENSORRT_VERSION}" \

				       --build-arg "GRADLE_VERSION=${GRADLE_VERSION}" \

				       --build-arg "NINJA_VERSION=${NINJA_VERSION:-}" \

				       --build-arg "KATEX=${KATEX:-}" \

				       --build-arg "ROCM_VERSION=${ROCM_VERSION:-}" \

				@ -446,6 +379,7 @@ docker build \

				       --build-arg "XPU_VERSION=${XPU_VERSION}" \

				       --build-arg "UNINSTALL_DILL=${UNINSTALL_DILL}" \

				       --build-arg "ACL=${ACL:-}" \

				       --build-arg "OPENBLAS=${OPENBLAS:-}" \

				       --build-arg "SKIP_SCCACHE_INSTALL=${SKIP_SCCACHE_INSTALL:-}" \

				       --build-arg "SKIP_LLVM_SRC_BUILD_INSTALL=${SKIP_LLVM_SRC_BUILD_INSTALL:-}" \

				       -f $(dirname ${DOCKERFILE})/Dockerfile \

									
										1

.ci/docker/centos-rocm/Dockerfile
									
												View File
												
				@ -39,6 +39,7 @@ RUN bash ./install_user.sh && rm install_user.sh

				# Install conda and other packages (e.g., numpy, pytest)

				ARG ANACONDA_PYTHON_VERSION

				ARG BUILD_ENVIRONMENT

				ENV ANACONDA_PYTHON_VERSION=$ANACONDA_PYTHON_VERSION

				ENV PATH /opt/conda/envs/py_$ANACONDA_PYTHON_VERSION/bin:/opt/conda/bin:$PATH

				COPY requirements-ci.txt /opt/conda/requirements-ci.txt

2

.ci/docker/ci_commit_pins/executorch.txt

View File

 @ -1 +1 @@
 b173722085b3f555d6ba4533d6bbaddfd7c71144
 aa978594cc155fa8af48cd949f5b5f1823a

2

.ci/docker/ci_commit_pins/nccl-cu12.txt

View File

 @ -1 +1 @@
 v2.26.5-1
 v2.27.5-1

1

.ci/docker/ci_commit_pins/torchbench.txt Normal file

View File

				`@ -0,0 +1 @@`
				`e03a63be43e33596f7f0a43b0f530353785e4a59`

2

.ci/docker/ci_commit_pins/triton-xpu.txt

View File

 @ -1 +1 @@
 bcc8265e677e5321606a3311bf71470f14456a8
 ae324eeac8e102a2b40370e341460f3791353398

2

.ci/docker/ci_commit_pins/triton.txt

View File

 @ -1 +1 @@
 ce50fade7e209553aba4898cd9b82aab83b
 f7888497a1eb9e98d4c07537f0d0bcfe180d1363

									
										4

.ci/docker/common/common_utils.sh
									
												View File
												
				@ -23,6 +23,10 @@ conda_install() {

				  as_jenkins conda install -q -n py_$ANACONDA_PYTHON_VERSION -y python="$ANACONDA_PYTHON_VERSION" $*

				}

				conda_install_through_forge() {

				  as_jenkins conda install -c conda-forge -q -n py_$ANACONDA_PYTHON_VERSION -y python="$ANACONDA_PYTHON_VERSION" $*

				}

				conda_run() {

				  as_jenkins conda run -n py_$ANACONDA_PYTHON_VERSION --no-capture-output $*

				}

									
										16

.ci/docker/common/install_base.sh
									
												View File
												
				@ -15,6 +15,9 @@ install_ubuntu() {

				  elif [[ "$UBUNTU_VERSION" == "22.04"* ]]; then

				    cmake3="cmake=3.22*"

				    maybe_libiomp_dev=""

				  elif [[ "$UBUNTU_VERSION" == "24.04"* ]]; then

				    cmake3="cmake=3.28*"

				    maybe_libiomp_dev=""

				  else

				    cmake3="cmake=3.5*"

				    maybe_libiomp_dev="libiomp-dev"

				@ -30,18 +33,6 @@ install_ubuntu() {

				    maybe_libomp_dev=""

				  fi

				  # HACK: UCC testing relies on libnccl library from NVIDIA repo, and version 2.16 crashes

				  # See https://github.com/pytorch/pytorch/pull/105260#issuecomment-1673399729

				  # TODO: Eliminate this hack, we should not relay on apt-get installation

				  # See https://github.com/pytorch/pytorch/issues/144768

				  if [[ "$UBUNTU_VERSION" == "20.04"* && "$CUDA_VERSION" == "11.8"* ]]; then

				    maybe_libnccl_dev="libnccl2=2.15.5-1+cuda11.8 libnccl-dev=2.15.5-1+cuda11.8 --allow-downgrades --allow-change-held-packages"

				  elif [[ "$UBUNTU_VERSION" == "20.04"* && "$CUDA_VERSION" == "12.4"* ]]; then

				    maybe_libnccl_dev="libnccl2=2.26.2-1+cuda12.4 libnccl-dev=2.26.2-1+cuda12.4 --allow-downgrades --allow-change-held-packages"

				  else

				    maybe_libnccl_dev=""

				  fi

				  # Install common dependencies

				  apt-get update

				  # TODO: Some of these may not be necessary

				@ -70,7 +61,6 @@ install_ubuntu() {

				    libasound2-dev \

				    libsndfile-dev \

				    ${maybe_libomp_dev} \

				    ${maybe_libnccl_dev} \

				    software-properties-common \

				    wget \

				    sudo \

									
										26

.ci/docker/common/install_conda.sh
									
												View File
												
				@ -4,12 +4,8 @@ set -ex

				# Optionally install conda

				if [ -n "$ANACONDA_PYTHON_VERSION" ]; then

				  BASE_URL="https://repo.anaconda.com/miniconda"

				  CONDA_FILE="Miniconda3-latest-Linux-x86_64.sh"

				  if [[ $(uname -m) == "aarch64" ]] || [[ "$BUILD_ENVIRONMENT" == *xpu* ]]; then

				    BASE_URL="https://github.com/conda-forge/miniforge/releases/latest/download"  # @lint-ignore

				    CONDA_FILE="Miniforge3-Linux-$(uname -m).sh"

				  fi

				  BASE_URL="https://github.com/conda-forge/miniforge/releases/latest/download"  # @lint-ignore

				  CONDA_FILE="Miniforge3-Linux-$(uname -m).sh"

				  MAJOR_PYTHON_VERSION=$(echo "$ANACONDA_PYTHON_VERSION" | cut -d . -f 1)

				  MINOR_PYTHON_VERSION=$(echo "$ANACONDA_PYTHON_VERSION" | cut -d . -f 2)

				@ -21,7 +17,6 @@ if [ -n "$ANACONDA_PYTHON_VERSION" ]; then

				      exit 1

				      ;;

				  esac

				  mkdir -p /opt/conda

				  chown jenkins:jenkins /opt/conda

				@ -64,11 +59,16 @@ if [ -n "$ANACONDA_PYTHON_VERSION" ]; then

				  # which is provided in libstdcxx 12 and up.

				  conda_install libstdcxx-ng=12.3.0 --update-deps -c conda-forge

				  # Miniforge installer doesn't install sqlite by default

				  if [[ "$BUILD_ENVIRONMENT" == *rocm* ]]; then

				    conda_install sqlite

				  fi

				  # Install PyTorch conda deps, as per https://github.com/pytorch/pytorch README

				  if [[ $(uname -m) == "aarch64" ]]; then

				    conda_install "openblas==0.3.29=*openmp*"

				  else

				    conda_install "mkl=2021.4.0 mkl-include=2021.4.0"

				  if [[ $(uname -m) != "aarch64" ]]; then

				    pip_install mkl==2024.2.0

				    pip_install mkl-static==2024.2.0

				    pip_install mkl-include==2024.2.0

				  fi

				  # Install llvm-8 as it is required to compile llvmlite-0.30.0 from source

				@ -82,6 +82,10 @@ if [ -n "$ANACONDA_PYTHON_VERSION" ]; then

				    conda_run ${SCRIPT_FOLDER}/install_magma_conda.sh $(cut -f1-2 -d'.' <<< ${CUDA_VERSION})

				  fi

				  if [[ "$UBUNTU_VERSION" == "24.04"* ]] ; then

				    conda_install_through_forge libstdcxx-ng=14

				  fi

				  # Install some other packages, including those needed for Python test reporting

				  pip_install -r /opt/conda/requirements-ci.txt

									
										36

.ci/docker/common/install_cpython.sh
									
												View File
												
				@ -3,11 +3,10 @@

				set -uex -o pipefail

				PYTHON_DOWNLOAD_URL=https://www.python.org/ftp/python

				PYTHON_DOWNLOAD_GITHUB_BRANCH=https://github.com/python/cpython/archive/refs/heads  # @lint-ignore

				GET_PIP_URL=https://bootstrap.pypa.io/get-pip.py

				# Python versions to be installed in /opt/$VERSION_NO

				CPYTHON_VERSIONS=${CPYTHON_VERSIONS:-"3.9.0 3.10.1 3.11.0 3.12.0 3.13.0 3.13.0t"}

				CPYTHON_VERSIONS=${CPYTHON_VERSIONS:-"3.9.0 3.10.1 3.11.0 3.12.0 3.13.0 3.13.0t 3.14.0 3.14.0t"}

				function check_var {

				    if [ -z "$1" ]; then

				@ -24,9 +23,8 @@ function do_cpython_build {

				    tar -xzf Python-$py_ver.tgz

				    local additional_flags=""

				    if [ "$py_ver" == "3.13.0t" ]; then

				    if [[ "$py_ver" == *"t" ]]; then

				        additional_flags=" --disable-gil"

				        mv cpython-3.13/ cpython-3.13t/

				    fi

				    pushd $py_folder

				@ -68,7 +66,7 @@ function do_cpython_build {

				        ln -s pip3 ${prefix}/bin/pip

				    fi

				    # install setuptools since python 3.12 is required to use distutils

				    ${prefix}/bin/pip install wheel==0.34.2 setuptools==68.2.2

				    ${prefix}/bin/pip install wheel==0.45.1 setuptools==80.9.0

				    local abi_tag=$(${prefix}/bin/python -c "from wheel.pep425tags import get_abbr_impl, get_impl_ver, get_abi_tag; print('{0}{1}-{2}'.format(get_abbr_impl(), get_impl_ver(), get_abi_tag()))")

				    ln -sf ${prefix} /opt/python/${abi_tag}

				}

				@ -76,24 +74,20 @@ function do_cpython_build {

				function build_cpython {

				    local py_ver=$1

				    check_var $py_ver

				    check_var $PYTHON_DOWNLOAD_URL

				    local py_ver_folder=$py_ver

				    local py_suffix=$py_ver

				    local py_folder=$py_ver

				    if [ "$py_ver" = "3.13.0t" ]; then

				        PY_VER_SHORT="3.13"

				        PYT_VER_SHORT="3.13t"

				        check_var $PYTHON_DOWNLOAD_GITHUB_BRANCH

				        wget $PYTHON_DOWNLOAD_GITHUB_BRANCH/$PY_VER_SHORT.tar.gz -O Python-$py_ver.tgz

				        do_cpython_build $py_ver cpython-$PYT_VER_SHORT

				    elif [ "$py_ver" = "3.13.0" ]; then

				        PY_VER_SHORT="3.13"

				        check_var $PYTHON_DOWNLOAD_GITHUB_BRANCH

				        wget $PYTHON_DOWNLOAD_GITHUB_BRANCH/$PY_VER_SHORT.tar.gz -O Python-$py_ver.tgz

				        do_cpython_build $py_ver cpython-$PY_VER_SHORT

				    else

				        wget -q $PYTHON_DOWNLOAD_URL/$py_ver_folder/Python-$py_ver.tgz

				        do_cpython_build $py_ver Python-$py_ver

				    # Special handling for nogil

				    if [[ "${py_ver}" == *"t" ]]; then

				        py_suffix=${py_ver::-1}

				        py_folder=$py_suffix

				    fi

				    # Only b3 is available now

				    if [ "$py_suffix" == "3.14.0" ]; then

				        py_suffix="3.14.0b3"

				    fi

				    wget -q $PYTHON_DOWNLOAD_URL/$py_folder/Python-$py_suffix.tgz -O Python-$py_ver.tgz

				    do_cpython_build $py_ver Python-$py_suffix

				    rm -f Python-$py_ver.tgz

				}

									
										97

.ci/docker/common/install_cuda.sh
									
												View File
												
				@ -10,6 +10,8 @@ else

				  arch_path='sbsa'

				fi

				NVSHMEM_VERSION=3.3.9

				function install_cuda {

				  version=$1

				  runfile=$2

				@ -40,18 +42,40 @@ function install_cudnn {

				  rm -rf tmp_cudnn

				}

				function install_118 {

				    CUDNN_VERSION=9.1.0.70

				    echo "Installing CUDA 11.8 and cuDNN ${CUDNN_VERSION} and NCCL and cuSparseLt-0.4.0"

				    install_cuda 11.8.0 cuda_11.8.0_520.61.05_linux

				function install_nvshmem {

				  cuda_major_version=$1      # e.g. "12"

				  nvshmem_version=$2         # e.g. "3.3.9"

				    install_cudnn 11 $CUDNN_VERSION

				  case "${arch_path}" in

				    sbsa)

				      dl_arch="aarch64"

				      ;;

				    x86_64)

				      dl_arch="x64"

				      ;;

				    *)

				      dl_arch="${arch}"

				      ;;

				  esac

				    CUDA_VERSION=11.8 bash install_nccl.sh

				  tmpdir="tmp_nvshmem"

				  mkdir -p "${tmpdir}" && cd "${tmpdir}"

				    CUDA_VERSION=11.8 bash install_cusparselt.sh

				  # nvSHMEM license: https://docs.nvidia.com/nvshmem/api/sla.html

				  filename="libnvshmem_cuda${cuda_major_version}-linux-${arch_path}-${nvshmem_version}"

				  url="https://developer.download.nvidia.com/compute/redist/nvshmem/${nvshmem_version}/builds/cuda${cuda_major_version}/txz/agnostic/${dl_arch}/${filename}.tar.gz"

				    ldconfig

				  # download, unpack, install

				  wget -q "${url}"

				  tar xf "${filename}.tar.gz"

				  cp -a "libnvshmem/include/"* /usr/local/cuda/include/

				  cp -a "libnvshmem/lib/"*     /usr/local/cuda/lib64/

				  # cleanup

				  cd ..

				  rm -rf "${tmpdir}"

				  echo "nvSHMEM ${nvshmem_version} for CUDA ${cuda_major_version} (${arch_path}) installed."

				}

				function install_124 {

				@ -69,12 +93,14 @@ function install_124 {

				}

				function install_126 {

				  CUDNN_VERSION=9.5.1.17

				  echo "Installing CUDA 12.6.3 and cuDNN ${CUDNN_VERSION} and NCCL and cuSparseLt-0.6.3"

				  CUDNN_VERSION=9.10.2.21

				  echo "Installing CUDA 12.6.3 and cuDNN ${CUDNN_VERSION} and NVSHMEM and NCCL and cuSparseLt-0.7.1"

				  install_cuda 12.6.3 cuda_12.6.3_560.35.05_linux

				  install_cudnn 12 $CUDNN_VERSION

				  install_nvshmem 12 $NVSHMEM_VERSION

				  CUDA_VERSION=12.6 bash install_nccl.sh

				  CUDA_VERSION=12.6 bash install_cusparselt.sh

				@ -82,35 +108,22 @@ function install_126 {

				  ldconfig

				}

				function prune_118 {

				    echo "Pruning CUDA 11.8 and cuDNN"

				    #####################################################################################

				    # CUDA 11.8 prune static libs

				    #####################################################################################

				    export NVPRUNE="/usr/local/cuda-11.8/bin/nvprune"

				    export CUDA_LIB_DIR="/usr/local/cuda-11.8/lib64"

				function install_129 {

				  CUDNN_VERSION=9.10.2.21

				  echo "Installing CUDA 12.9.1 and cuDNN ${CUDNN_VERSION} and NVSHMEM and NCCL and cuSparseLt-0.7.1"

				  # install CUDA 12.9.1 in the same container

				  install_cuda 12.9.1 cuda_12.9.1_575.57.08_linux

				    export GENCODE="-gencode arch=compute_35,code=sm_35 -gencode arch=compute_50,code=sm_50 -gencode arch=compute_60,code=sm_60 -gencode arch=compute_70,code=sm_70 -gencode arch=compute_75,code=sm_75 -gencode arch=compute_80,code=sm_80 -gencode arch=compute_86,code=sm_86 -gencode arch=compute_90,code=sm_90"

				    export GENCODE_CUDNN="-gencode arch=compute_35,code=sm_35 -gencode arch=compute_37,code=sm_37 -gencode arch=compute_50,code=sm_50 -gencode arch=compute_60,code=sm_60 -gencode arch=compute_61,code=sm_61 -gencode arch=compute_70,code=sm_70 -gencode arch=compute_75,code=sm_75 -gencode arch=compute_80,code=sm_80 -gencode arch=compute_86,code=sm_86 -gencode arch=compute_90,code=sm_90"

				  # cuDNN license: https://developer.nvidia.com/cudnn/license_agreement

				  install_cudnn 12 $CUDNN_VERSION

				    if [[ -n "$OVERRIDE_GENCODE" ]]; then

				        export GENCODE=$OVERRIDE_GENCODE

				    fi

				  install_nvshmem 12 $NVSHMEM_VERSION

				    # all CUDA libs except CuDNN and CuBLAS (cudnn and cublas need arch 3.7 included)

				    ls $CUDA_LIB_DIR/ | grep "\.a" | grep -v "culibos" | grep -v "cudart" | grep -v "cudnn" | grep -v "cublas" | grep -v "metis"  \

				      | xargs -I {} bash -c \

				                "echo {} && $NVPRUNE $GENCODE $CUDA_LIB_DIR/{} -o $CUDA_LIB_DIR/{}"

				  CUDA_VERSION=12.9 bash install_nccl.sh

				    # prune CuDNN and CuBLAS

				    $NVPRUNE $GENCODE_CUDNN $CUDA_LIB_DIR/libcublas_static.a -o $CUDA_LIB_DIR/libcublas_static.a

				    $NVPRUNE $GENCODE_CUDNN $CUDA_LIB_DIR/libcublasLt_static.a -o $CUDA_LIB_DIR/libcublasLt_static.a

				  CUDA_VERSION=12.9 bash install_cusparselt.sh

				    #####################################################################################

				    # CUDA 11.8 prune visual tools

				    #####################################################################################

				    export CUDA_BASE="/usr/local/cuda-11.8/"

				    rm -rf $CUDA_BASE/libnvvp $CUDA_BASE/nsightee_plugins $CUDA_BASE/nsight-compute-2022.3.0 $CUDA_BASE/nsight-systems-2022.4.2/

				  ldconfig

				}

				function prune_124 {

				@ -183,13 +196,15 @@ function prune_126 {

				function install_128 {

				  CUDNN_VERSION=9.8.0.87

				  echo "Installing CUDA 12.8.0 and cuDNN ${CUDNN_VERSION} and NCCL and cuSparseLt-0.6.3"

				  # install CUDA 12.8.0 in the same container

				  install_cuda 12.8.0 cuda_12.8.0_570.86.10_linux

				  echo "Installing CUDA 12.8.1 and cuDNN ${CUDNN_VERSION} and NVSHMEM and NCCL and cuSparseLt-0.7.1"

				  # install CUDA 12.8.1 in the same container

				  install_cuda 12.8.1 cuda_12.8.1_570.124.06_linux

				  # cuDNN license: https://developer.nvidia.com/cudnn/license_agreement

				  install_cudnn 12 $CUDNN_VERSION

				  install_nvshmem 12 $NVSHMEM_VERSION

				  CUDA_VERSION=12.8 bash install_nccl.sh

				  CUDA_VERSION=12.8 bash install_cusparselt.sh

				@ -201,13 +216,13 @@ function install_128 {

				while test $# -gt 0

				do

				    case "$1" in

				    11.8) install_118; prune_118

				        ;;

				    12.4) install_124; prune_124

				        ;;

				    12.6) install_126; prune_126

				    12.6|12.6.*) install_126; prune_126

				        ;;

				    12.8) install_128;

				    12.8|12.8.*) install_128;

				        ;;

				    12.9|12.9.*) install_129;

				        ;;

				    *) echo "bad argument $1"; exit 1

				        ;;

									
										26

.ci/docker/common/install_cudnn.sh
									
												View File
											
				@ -1,26 +0,0 @@

				#!/bin/bash

				if [[ -n "${CUDNN_VERSION}" ]]; then

				    # cuDNN license: https://developer.nvidia.com/cudnn/license_agreement

				    mkdir tmp_cudnn

				    pushd tmp_cudnn

				    if [[ ${CUDA_VERSION:0:4} == "12.8" ]]; then

				        CUDNN_NAME="cudnn-linux-x86_64-9.8.0.87_cuda12-archive"

				    elif [[ ${CUDA_VERSION:0:4} == "12.6" ]]; then

				        CUDNN_NAME="cudnn-linux-x86_64-9.5.1.17_cuda12-archive"

				    elif [[ ${CUDA_VERSION:0:2} == "12" ]]; then

				        CUDNN_NAME="cudnn-linux-x86_64-9.1.0.70_cuda12-archive"

				    elif [[ ${CUDA_VERSION:0:2} == "11" ]]; then

				        CUDNN_NAME="cudnn-linux-x86_64-9.1.0.70_cuda11-archive"

				    else

				        print "Unsupported CUDA version ${CUDA_VERSION}"

				        exit 1

				    fi

				    curl --retry 3 -OLs https://developer.download.nvidia.com/compute/cudnn/redist/cudnn/linux-x86_64/${CUDNN_NAME}.tar.xz

				    tar xf ${CUDNN_NAME}.tar.xz

				    cp -a ${CUDNN_NAME}/include/* /usr/local/cuda/include/

				    cp -a ${CUDNN_NAME}/lib/* /usr/local/cuda/lib64/

				    popd

				    rm -rf tmp_cudnn

				    ldconfig

				fi

									
										7

.ci/docker/common/install_cusparselt.sh
									
												View File
												
				@ -5,13 +5,13 @@ set -ex

				# cuSPARSELt license: https://docs.nvidia.com/cuda/cusparselt/license.html

				mkdir tmp_cusparselt && cd tmp_cusparselt

				if [[ ${CUDA_VERSION:0:4} =~ ^12\.[5-8]$ ]]; then

				if [[ ${CUDA_VERSION:0:4} =~ ^12\.[5-9]$ ]]; then

				    arch_path='sbsa'

				    export TARGETARCH=${TARGETARCH:-$(uname -m)}

				    if [ ${TARGETARCH} = 'amd64' ] || [ "${TARGETARCH}" = 'x86_64' ]; then

				        arch_path='x86_64'

				    fi

				    CUSPARSELT_NAME="libcusparse_lt-linux-${arch_path}-0.6.3.2-archive"

				    CUSPARSELT_NAME="libcusparse_lt-linux-${arch_path}-0.7.1.0-archive"

				    curl --retry 3 -OLs https://developer.download.nvidia.com/compute/cusparselt/redist/libcusparse_lt/linux-${arch_path}/${CUSPARSELT_NAME}.tar.xz

				elif [[ ${CUDA_VERSION:0:4} == "12.4" ]]; then

				    arch_path='sbsa'

				@ -21,9 +21,6 @@ elif [[ ${CUDA_VERSION:0:4} == "12.4" ]]; then

				    fi

				    CUSPARSELT_NAME="libcusparse_lt-linux-${arch_path}-0.6.2.3-archive"

				    curl --retry 3 -OLs https://developer.download.nvidia.com/compute/cusparselt/redist/libcusparse_lt/linux-${arch_path}/${CUSPARSELT_NAME}.tar.xz

				elif [[ ${CUDA_VERSION:0:4} == "11.8" ]]; then

				    CUSPARSELT_NAME="libcusparse_lt-linux-x86_64-0.4.0.7-archive"

				    curl --retry 3 -OLs https://developer.download.nvidia.com/compute/cusparselt/redist/libcusparse_lt/linux-x86_64/${CUSPARSELT_NAME}.tar.xz

				else

				    echo "Not sure which libcusparselt version to install for this ${CUDA_VERSION}"

				fi

									
										30

.ci/docker/common/install_inductor_benchmark_deps.sh
									
												View File
												
				@ -15,11 +15,37 @@ function install_timm() {

				  commit=$(get_pinned_commit timm)

				  pip_install "git+https://github.com/huggingface/pytorch-image-models@${commit}"

				  # Clean up

				  conda_run pip uninstall -y torch torchvision triton

				}

				function install_torchbench() {

				  local commit

				  commit=$(get_pinned_commit torchbench)

				  git clone https://github.com/pytorch/benchmark torchbench

				  pushd torchbench

				  git checkout "$commit"

				  python install.py --continue_on_fail

				  # TODO (huydhn): transformers-4.44.2 added by https://github.com/pytorch/benchmark/pull/2488

				  # is regressing speedup metric. This needs to be investigated further

				  pip install transformers==4.38.1

				  echo "Print all dependencies after TorchBench is installed"

				  python -mpip freeze

				  popd

				  chown -R jenkins torchbench

				}

				# Pango is needed for weasyprint which is needed for doctr

				conda_install pango

				# Stable packages are ok here, just to satisfy TorchBench check

				pip_install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128

				install_torchbench

				install_huggingface

				install_timm

				# Clean up

				conda_run pip uninstall -y torch torchvision torchaudio triton

									
										16

.ci/docker/common/install_onnx.sh
									
												View File
												
				@ -8,16 +8,6 @@ retry () {

				    "$@" || (sleep 10 && "$@") || (sleep 20 && "$@") || (sleep 40 && "$@")

				}

				# A bunch of custom pip dependencies for ONNX

				pip_install \

				  beartype==0.15.0 \

				  filelock==3.9.0 \

				  flatbuffers==2.0 \

				  mock==5.0.1 \

				  ninja==1.10.2 \

				  networkx==2.5 \

				  numpy==1.24.2

				# ONNXRuntime should be installed before installing

				# onnx-weekly. Otherwise, onnx-weekly could be

				# overwritten by onnx.

				@ -29,12 +19,8 @@ pip_install \

				  transformers==4.36.2

				pip_install coloredlogs packaging

				pip_install onnxruntime==1.18.1

				pip_install onnx==1.17.0

				pip_install onnxscript==0.2.2 --no-deps

				# required by onnxscript

				pip_install ml_dtypes

				pip_install onnxscript==0.3.1

				# Cache the transformers model to be used later by ONNX tests. We need to run the transformers

				# package to download the model. By default, the model is cached at ~/.cache/huggingface/hub/

									
										7

.ci/docker/common/install_openblas.sh
									
												View File
												
				@ -4,9 +4,9 @@

				set -ex

				cd /

				git clone https://github.com/OpenMathLib/OpenBLAS.git -b v0.3.29 --depth 1 --shallow-submodules

				git clone https://github.com/OpenMathLib/OpenBLAS.git -b "${OPENBLAS_VERSION:-v0.3.30}" --depth 1 --shallow-submodules

				OPENBLAS_CHECKOUT_DIR="OpenBLAS"

				OPENBLAS_BUILD_FLAGS="

				NUM_THREADS=128

				USE_OPENMP=1

				@ -14,9 +14,8 @@ NO_SHARED=0

				DYNAMIC_ARCH=1

				TARGET=ARMV8

				CFLAGS=-O3

				BUILD_BFLOAT16=1

				"

				OPENBLAS_CHECKOUT_DIR="OpenBLAS"

				make -j8 ${OPENBLAS_BUILD_FLAGS} -C ${OPENBLAS_CHECKOUT_DIR}

				make -j8 ${OPENBLAS_BUILD_FLAGS} install -C ${OPENBLAS_CHECKOUT_DIR}

									
										56

.ci/docker/common/install_rocm.sh
									
												View File
												
				@ -8,9 +8,11 @@ ver() {

				install_ubuntu() {

				    apt-get update

				    if [[ $UBUNTU_VERSION == 20.04 ]]; then

				      # gpg-agent is not available by default on 20.04

				      apt-get install -y --no-install-recommends gpg-agent

				    # gpg-agent is not available by default

				    apt-get install -y --no-install-recommends gpg-agent

				    if [[ $(ver $UBUNTU_VERSION) -ge $(ver 22.04) ]]; then

				        echo -e 'Package: *\nPin: release o=repo.radeon.com\nPin-Priority: 600' \

				            | sudo tee /etc/apt/preferences.d/rocm-pin-600

				    fi

				    apt-get install -y kmod

				    apt-get install -y wget

				@ -26,13 +28,27 @@ Pin: release o=repo.radeon.com

				Pin-Priority: 600

				EOF

				    # we want the patch version of 6.4 instead

				    if [[ $(ver $ROCM_VERSION) -eq $(ver 6.4) ]]; then

				        ROCM_VERSION="${ROCM_VERSION}.2"

				    fi

				    # Default url values

				    rocm_baseurl="http://repo.radeon.com/rocm/apt/${ROCM_VERSION}"

				    amdgpu_baseurl="https://repo.radeon.com/amdgpu/${ROCM_VERSION}/ubuntu"

				    # Special case for ROCM_VERSION == 7.0

				    if [[ $(ver "$ROCM_VERSION") -eq $(ver 7.0) ]]; then

				        rocm_baseurl="https://repo.radeon.com/rocm/apt/7.0_alpha2"

				        amdgpu_baseurl="https://repo.radeon.com/amdgpu/30.10_alpha2/ubuntu"

				    fi

				    # Add amdgpu repository

				    UBUNTU_VERSION_NAME=`cat /etc/os-release | grep UBUNTU_CODENAME | awk -F= '{print $2}'`

				    echo "deb [arch=amd64] https://repo.radeon.com/amdgpu/${ROCM_VERSION}/ubuntu ${UBUNTU_VERSION_NAME} main" > /etc/apt/sources.list.d/amdgpu.list

				    echo "deb [arch=amd64] ${amdgpu_baseurl} ${UBUNTU_VERSION_NAME} main" > /etc/apt/sources.list.d/amdgpu.list

				    # Add rocm repository

				    wget -qO - http://repo.radeon.com/rocm/rocm.gpg.key | apt-key add -

				    local rocm_baseurl="http://repo.radeon.com/rocm/apt/${ROCM_VERSION}"

				    echo "deb [arch=amd64] ${rocm_baseurl} ${UBUNTU_VERSION_NAME} main" > /etc/apt/sources.list.d/rocm.list

				    apt-get update --allow-insecure-repositories

				@ -66,25 +82,33 @@ EOF

				    done

				    # ROCm 6.3 had a regression where initializing static code objects had significant overhead

				    # CI no longer builds for ROCm 6.3, but

				    # ROCm 6.4 did not yet fix the regression, also HIP branch names are different

				    if [[ $(ver $ROCM_VERSION) -eq $(ver 6.3) ]] || [[ $(ver $ROCM_VERSION) -eq $(ver 6.4) ]]; then

				        if [[ $(ver $ROCM_VERSION) -eq $(ver 6.3) ]]; then

				            HIP_BRANCH=rocm-6.3.x

				            VER_STR=6.3

				    if [[ $(ver $ROCM_VERSION) -ge $(ver 6.4) ]] && [[ $(ver $ROCM_VERSION) -lt $(ver 7.0) ]]; then

				        if [[ $(ver $ROCM_VERSION) -eq $(ver 6.4.2) ]]; then

				            HIP_TAG=rocm-6.4.2

				            CLR_HASH=74d78ba3ac4bac235d02bcb48511c30b5cfdd457  # branch release/rocm-rel-6.4.2-statco-hotfix

				        elif [[ $(ver $ROCM_VERSION) -eq $(ver 6.4.1) ]]; then

				            HIP_TAG=rocm-6.4.1

				            CLR_HASH=efe6c35790b9206923bfeed1209902feff37f386  # branch release/rocm-rel-6.4.1-statco-hotfix

				        elif [[ $(ver $ROCM_VERSION) -eq $(ver 6.4) ]]; then

				            HIP_BRANCH=release/rocm-rel-6.4

				            VER_STR=6.4

				            HIP_TAG=rocm-6.4.0

				            CLR_HASH=600f5b0d2baed94d5121e2174a9de0851b040b0c  # branch release/rocm-rel-6.4-statco-hotfix

				        fi

				        # clr build needs CppHeaderParser but can only find it using conda's python

				        /opt/conda/bin/python -m pip install CppHeaderParser

				        git clone https://github.com/ROCm/HIP -b $HIP_BRANCH

				        python -m pip install CppHeaderParser

				        git clone https://github.com/ROCm/HIP -b $HIP_TAG

				        HIP_COMMON_DIR=$(readlink -f HIP)

				        git clone https://github.com/jeffdaily/clr -b release/rocm-rel-${VER_STR}-statco-hotfix

				        git clone https://github.com/jeffdaily/clr

				        pushd clr

				        git checkout $CLR_HASH

				        popd

				        mkdir -p clr/build

				        pushd clr/build

				        cmake .. -DCLR_BUILD_HIP=ON -DHIP_COMMON_DIR=$HIP_COMMON_DIR

				        # Need to point CMake to the correct python installation to find CppHeaderParser

				        cmake .. -DPython3_EXECUTABLE=/opt/conda/envs/py_${ANACONDA_PYTHON_VERSION}/bin/python3 -DCLR_BUILD_HIP=ON -DHIP_COMMON_DIR=$HIP_COMMON_DIR

				        make -j

				        cp hipamd/lib/libamdhip64.so.${VER_STR}.* /opt/rocm/lib/libamdhip64.so.${VER_STR}.*

				        cp hipamd/lib/libamdhip64.so.6.4.* /opt/rocm/lib/libamdhip64.so.6.4.*

				        popd

				        rm -rf HIP clr

				    fi

									
										7

.ci/docker/common/install_rocm_magma.sh
									
												View File
												
				@ -5,7 +5,12 @@ set -eou pipefail

				function do_install() {

				    rocm_version=$1

				    rocm_version_nodot=${1//./}

				    if [[ ${rocm_version} =~ ^[0-9]+\.[0-9]+\.[0-9]+$ ]]; then

				        # chop off any patch version

				        rocm_version="${rocm_version%.*}"

				    fi

				    rocm_version_nodot=${rocm_version//./}

				    # Version 2.7.2 + ROCm related updates

				    MAGMA_VERSION=a1625ff4d9bc362906bd01f805dbbe12612953f6

									
										14

.ci/docker/common/install_triton.sh
									
												View File
												
				@ -51,7 +51,12 @@ as_jenkins git clone --recursive ${TRITON_REPO} triton

				cd triton

				as_jenkins git checkout ${TRITON_PINNED_COMMIT}

				as_jenkins git submodule update --init --recursive

				cd python

				# Old versions of python have setup.py in ./python; newer versions have it in ./

				if [ ! -f setup.py ]; then

				  cd python

				fi

				pip_install pybind11==2.13.6

				# TODO: remove patch setup.py once we have a proper fix for https://github.com/triton-lang/triton/issues/4527

				@ -93,3 +98,10 @@ fi

				if [ -n "${NUMPY_VERSION}" ]; then

				  pip_install "numpy==${NUMPY_VERSION}"

				fi

				# IMPORTANT: helion needs to be installed without dependencies.

				# It depends on torch and triton. We don't want to install

				# triton and torch from production on Docker CI images

				if [[ "$ANACONDA_PYTHON_VERSION" != 3.9* ]]; then

				  pip_install helion --no-deps

				fi

									
										12

.ci/docker/common/install_xpu.sh
									
												View File
												
				@ -56,14 +56,10 @@ function install_ubuntu() {

				function install_rhel() {

				    . /etc/os-release

				    if [[ "${ID}" == "rhel" ]]; then

				        if [[ ! " 8.8 8.9 9.0 9.2 9.3 " =~ " ${VERSION_ID} " ]]; then

				            echo "RHEL version ${VERSION_ID} not supported"

				            exit

				        fi

				    elif [[ "${ID}" == "almalinux" ]]; then

				        # Workaround for almalinux8 which used by quay.io/pypa/manylinux_2_28_x86_64

				        VERSION_ID="8.8"

				    if [[ ! " 8.8 8.10 9.0 9.2 9.3 " =~ " ${VERSION_ID} " ]]; then

				        echo "RHEL version ${VERSION_ID} not supported"

				        exit

				    fi

				    dnf install -y 'dnf-command(config-manager)'

									
										15

.ci/docker/libtorch/Dockerfile
									
												View File
												
				@ -54,16 +54,6 @@ COPY ./ci_commit_pins/nccl-cu* /ci_commit_pins/

				COPY ./common/install_cusparselt.sh install_cusparselt.sh

				ENV CUDA_HOME /usr/local/cuda

				FROM cuda as cuda11.8

				RUN bash ./install_cuda.sh 11.8

				RUN bash ./install_magma.sh 11.8

				RUN ln -sf /usr/local/cuda-11.8 /usr/local/cuda

				FROM cuda as cuda12.4

				RUN bash ./install_cuda.sh 12.4

				RUN bash ./install_magma.sh 12.4

				RUN ln -sf /usr/local/cuda-12.4 /usr/local/cuda

				FROM cuda as cuda12.6

				RUN bash ./install_cuda.sh 12.6

				RUN bash ./install_magma.sh 12.6

				@ -74,6 +64,11 @@ RUN bash ./install_cuda.sh 12.8

				RUN bash ./install_magma.sh 12.8

				RUN ln -sf /usr/local/cuda-12.8 /usr/local/cuda

				FROM cuda as cuda12.9

				RUN bash ./install_cuda.sh 12.9

				RUN bash ./install_magma.sh 12.9

				RUN ln -sf /usr/local/cuda-12.9 /usr/local/cuda

				FROM cpu as rocm

				ARG ROCM_VERSION

				ARG PYTORCH_ROCM_ARCH

									
										4

.ci/docker/libtorch/build.sh
									
												View File
												
				@ -39,6 +39,10 @@ case ${DOCKER_TAG_PREFIX} in

				        DOCKER_GPU_BUILD_ARG=""

				        ;;

				    rocm*)

				        # we want the patch version of 6.4 instead

				        if [[ $(ver $GPU_ARCH_VERSION) -eq $(ver 6.4) ]]; then

				            GPU_ARCH_VERSION="${GPU_ARCH_VERSION}.2"

				        fi

				        BASE_TARGET=rocm

				        GPU_IMAGE=rocm/dev-ubuntu-22.04:${GPU_ARCH_VERSION}-complete

				        PYTORCH_ROCM_ARCH="gfx900;gfx906;gfx908;gfx90a;gfx942;gfx1030;gfx1100;gfx1101;gfx1102;gfx1200;gfx1201"

									
										2

.ci/docker/linter/Dockerfile
									
												View File
												
				@ -27,5 +27,7 @@ COPY ./common/install_linter.sh install_linter.sh

				RUN bash ./install_linter.sh

				RUN rm install_linter.sh

				RUN chown -R jenkins:jenkins /var/lib/jenkins/ci_env

				USER jenkins

				CMD ["bash"]

3

.ci/docker/manywheel/Dockerfile_2_28

View File

 @ -26,7 +26,7 @@ ADD ./common/install_openssl.sh install_openssl.sh
 RUN bash ./install_openssl.sh && rm install_openssl.sh
 # remove unncessary python versions
 # remove unnecessary python versions
 RUN rm -rf /opt/python/cp26-cp26m /opt/_internal/cpython-2.6.9-ucs2
 RUN rm -rf /opt/python/cp26-cp26mu /opt/_internal/cpython-2.6.9-ucs4
 RUN rm -rf /opt/python/cp33-cp33m /opt/_internal/cpython-3.3.6
 @ -103,6 +103,7 @@ ENV SSL_CERT_FILE=/opt/_internal/certs.pem
 # Install LLVM version
 COPY --from=openssl            /opt/openssl                          /opt/openssl
 COPY --from=base               /opt/python                           /opt/python
 COPY --from=base               /usr/local/lib/                       /usr/local/lib/
 COPY --from=base               /opt/_internal                        /opt/_internal
 COPY --from=base               /usr/local/bin/auditwheel             /usr/local/bin/auditwheel
 COPY --from=intel              /opt/intel                            /opt/intel

5

.ci/docker/manywheel/Dockerfile_2_28_aarch64

View File

 @ -2,7 +2,7 @@ FROM quay.io/pypa/manylinux_2_28_aarch64 as base
 ARG GCCTOOLSET_VERSION=13
 # Language variabes
 # Language variables
 ENV LC_ALL=en_US.UTF-8
 ENV LANG=en_US.UTF-8
 ENV LANGUAGE=en_US.UTF-8
 @ -58,12 +58,13 @@ RUN git config --global --add safe.directory "*"
 FROM base as openblas
 # Install openblas
 ARG OPENBLAS_VERSION
 ADD ./common/install_openblas.sh install_openblas.sh
 RUN bash ./install_openblas.sh && rm install_openblas.sh
 FROM base as final
 # remove unncessary python versions
 # remove unnecessary python versions
 RUN rm -rf /opt/python/cp26-cp26m /opt/_internal/cpython-2.6.9-ucs2
 RUN rm -rf /opt/python/cp26-cp26mu /opt/_internal/cpython-2.6.9-ucs4
 RUN rm -rf /opt/python/cp33-cp33m /opt/_internal/cpython-3.3.6

2

.ci/docker/manywheel/Dockerfile_cuda_aarch64

View File

 @ -60,7 +60,7 @@ RUN bash ./install_openssl.sh && rm install_openssl.sh
 ENV SSL_CERT_FILE=/opt/_internal/certs.pem
 FROM openssl as final
 # remove unncessary python versions
 # remove unnecessary python versions
 RUN rm -rf /opt/python/cp26-cp26m /opt/_internal/cpython-2.6.9-ucs2
 RUN rm -rf /opt/python/cp26-cp26mu /opt/_internal/cpython-2.6.9-ucs4
 RUN rm -rf /opt/python/cp33-cp33m /opt/_internal/cpython-3.3.6

21

.ci/docker/manywheel/Dockerfile_s390x

View File

 @ -5,7 +5,9 @@ ENV LC_ALL=C.UTF-8
 ENV LANG=C.UTF-8
 ENV LANGUAGE=C.UTF-8
 ARG DEVTOOLSET_VERSION=13
 # there is a bugfix in gcc >= 14 for precompiled headers and s390x vectorization interaction.
 # with earlier gcc versions test/inductor/test_cpu_cpp_wrapper.py will fail.
 ARG DEVTOOLSET_VERSION=14
 # Installed needed OS packages. This is to support all
 # the binary builds (torch, vision, audio, text, data)
 RUN yum -y install epel-release
 @ -58,7 +60,8 @@ RUN yum install -y \
   libxslt-devel \
   libxml2-devel \
   openssl-devel \
   valgrind
   valgrind \
   ninja-build
 ENV PATH=/opt/rh/gcc-toolset-${DEVTOOLSET_VERSION}/root/usr/bin:$PATH
 ENV LD_LIBRARY_PATH=/opt/rh/gcc-toolset-${DEVTOOLSET_VERSION}/root/usr/lib64:/opt/rh/gcc-toolset-${DEVTOOLSET_VERSION}/root/usr/lib:$LD_LIBRARY_PATH
 @ -103,9 +106,6 @@ CMD ["/bin/bash"]
 # install test dependencies:
 # - grpcio requires system openssl, bundled crypto fails to build
 RUN dnf install -y \
   protobuf-devel \
   protobuf-c-devel \
   protobuf-lite-devel \
   hdf5-devel \
   python3-h5py \
   git
 @ -120,15 +120,22 @@ RUN python3 -mpip install cmake==3.28.0
 # so just build it from upstream repository.
 # h5py is dependency of onnxruntime_training.
 # h5py==3.11.0 builds with hdf5-devel 1.10.5 from repository.
 # h5py 3.11.0 doesn't build with numpy >= 2.3.0.
 # install newest flatbuffers version first:
 # for some reason old version is getting pulled in otherwise.
 # packaging package is required for onnxruntime wheel build.
 RUN pip3 install flatbuffers && \
   pip3 install h5py==3.11.0 && \
   pip3 install cython 'pkgconfig>=1.5.5' 'setuptools>=77' 'numpy<2.3.0' && \
   pip3 install --no-build-isolation h5py==3.11.0 && \
   pip3 install packaging && \
   git clone https://github.com/microsoft/onnxruntime && \
   cd onnxruntime && git checkout v1.21.0 && \
   git submodule update --init --recursive && \
   ./build.sh --config Release --parallel 0 --enable_pybind --build_wheel --enable_training --enable_training_apis --enable_training_ops --skip_tests --allow_running_as_root && \
   wget https://github.com/microsoft/onnxruntime/commit/f57db79743c4d1a3553aa05cf95bcd10966030e6.patch && \
   patch -p1 < f57db79743c4d1a3553aa05cf95bcd10966030e6.patch && \
   ./build.sh --config Release --parallel 0 --enable_pybind \
   --build_wheel --enable_training --enable_training_apis \
   --enable_training_ops --skip_tests --allow_running_as_root \
   --compile_no_warning_as_error && \
   pip3 install ./build/Linux/Release/dist/onnxruntime_training-*.whl && \
   cd .. && /bin/rm -rf ./onnxruntime

									
										7

.ci/docker/manywheel/build.sh
									
												View File
												
				@ -27,6 +27,7 @@ fi

				MANY_LINUX_VERSION=${MANY_LINUX_VERSION:-}

				DOCKERFILE_SUFFIX=${DOCKERFILE_SUFFIX:-}

				OPENBLAS_VERSION=${OPENBLAS_VERSION:-}

				case ${image} in

				    manylinux2_28-builder:cpu)

				@ -40,6 +41,7 @@ case ${image} in

				        GPU_IMAGE=arm64v8/almalinux:8

				        DOCKER_GPU_BUILD_ARG=" --build-arg DEVTOOLSET_VERSION=13 --build-arg NINJA_VERSION=1.12.1"

				        MANY_LINUX_VERSION="2_28_aarch64"

				        OPENBLAS_VERSION="v0.3.30"

				        ;;

				    manylinuxcxx11-abi-builder:cpu-cxx11-abi)

				        TARGET=final

				@ -73,6 +75,10 @@ case ${image} in

				        DOCKERFILE_SUFFIX="_cuda_aarch64"

				        ;;

				    manylinux2_28-builder:rocm*)

				        # we want the patch version of 6.4 instead

				        if [[ $(ver $GPU_ARCH_VERSION) -eq $(ver 6.4) ]]; then

				            GPU_ARCH_VERSION="${GPU_ARCH_VERSION}.2"

				        fi

				        TARGET=rocm_final

				        MANY_LINUX_VERSION="2_28"

				        DEVTOOLSET_VERSION="11"

				@ -109,6 +115,7 @@ tmp_tag=$(basename "$(mktemp -u)" | tr '[:upper:]' '[:lower:]')

				DOCKER_BUILDKIT=1 docker build  \

				    ${DOCKER_GPU_BUILD_ARG} \

				    --build-arg "GPU_IMAGE=${GPU_IMAGE}" \

				    --build-arg "OPENBLAS_VERSION=${OPENBLAS_VERSION}" \

				    --target "${TARGET}" \

				    -t "${tmp_tag}" \

				    $@ \

58

.ci/docker/requirements-ci.txt

View File

 @ -16,6 +16,7 @@ click
 #test that import:
 coremltools==5.0b5 ; python_version < "3.12"
 coremltools==8.3 ; python_version == "3.12"
 #Description: Apple framework for ML integration
 #Pinned versions: 5.0b5
 #test that import:
 @ -41,18 +42,15 @@ fbscribelogger==0.1.7
 #Pinned versions: 0.1.6
 #test that import:
 flatbuffers==2.0 ; platform_machine != "s390x"
 flatbuffers==24.12.23
 #Description: cross platform serialization library
 #Pinned versions: 2.0
 #Pinned versions: 24.12.23
 #test that import:
 flatbuffers ; platform_machine == "s390x"
 #Description: cross platform serialization library; Newer version is required on s390x for new python version
 hypothesis==5.35.1
 # Pin hypothesis to avoid flakiness: https://github.com/pytorch/pytorch/issues/31136
 #Description: advanced library for generating parametrized tests
 #Pinned versions: 3.44.6, 4.53.2
 #Pinned versions: 5.35.1
 #test that import: test_xnnpack_integration.py, test_pruning_op.py, test_nn.py
 junitparser==2.1.1
 @ -66,6 +64,7 @@ lark==0.12.0
 #test that import:
 librosa>=0.6.2 ; python_version < "3.11"
 librosa==0.10.2 ; python_version == "3.12"
 #Description: A python package for music and audio analysis
 #Pinned versions: >=0.6.2
 #test that import: test_spectral_ops.py
 @ -93,10 +92,10 @@ librosa>=0.6.2 ; python_version < "3.11"
 #Pinned versions:
 #test that import:
 mypy==1.14.0
 mypy==1.16.0
 # Pin MyPy version because new errors are likely to appear with each release
 #Description: linter
 #Pinned versions: 1.14.0
 #Pinned versions: 1.16.0
 #test that import: test_typing.py, test_type_hints.py
 networkx==2.8.8
 @ -114,6 +113,7 @@ ninja==1.11.1.3
 numba==0.49.0 ; python_version < "3.9"
 numba==0.55.2 ; python_version == "3.9"
 numba==0.55.2 ; python_version == "3.10"
 numba==0.60.0 ; python_version == "3.12"
 #Description: Just-In-Time Compiler for Numerical Functions
 #Pinned versions: 0.54.1, 0.49.0, <=0.49.1
 #test that import: test_numba_integration.py
 @ -166,10 +166,10 @@ pillow==11.0.0
 #Pinned versions: 10.3.0
 #test that import:
 protobuf==3.20.2
 #Description:  Google’s data interchange format
 #Pinned versions: 3.20.1
 #test that import: test_tensorboard.py
 protobuf==5.29.4
 #Description:  Google's data interchange format
 #Pinned versions: 5.29.4
 #test that import: test_tensorboard.py, test/onnx/*
 psutil
 #Description: information on running processes and system utilization
 @ -221,9 +221,9 @@ pygments==2.15.0
 #Pinned versions: 2.12.0
 #test that import: the doctests
 #PyYAML
 #pyyaml
 #Description: data serialization format
 #Pinned versions:
 #Pinned versions: 6.0.2
 #test that import:
 #requests
 @ -233,7 +233,7 @@ pygments==2.15.0
 #rich
 #Description: rich text and beautiful formatting in the terminal
 #Pinned versions: 10.9.0
 #Pinned versions: 14.1.0
 #test that import:
 scikit-image==0.19.3 ; python_version < "3.10"
 @ -307,7 +307,7 @@ pytest-cpp==2.3.0
 #Pinned versions: 2.3.0
 #test that import:
 z3-solver==4.12.6.0
 z3-solver==4.15.1.0
 #Description: The Z3 Theorem Prover Project
 #Pinned versions:
 #test that import:
 @ -337,12 +337,12 @@ sympy==1.13.3
 #Pinned versions:
 #test that import:
 onnx==1.17.0
 #Description: Required by mypy and test_public_bindings.py when checking torch.onnx._internal
 onnx==1.18.0
 #Description: Required by onnx tests, and mypy and test_public_bindings.py when checking torch.onnx._internal
 #Pinned versions:
 #test that import:
 onnxscript==0.2.2
 onnxscript==0.3.1
 #Description: Required by mypy and test_public_bindings.py when checking torch.onnx._internal
 #Pinned versions:
 #test that import:
 @ -361,12 +361,11 @@ pwlf==2.2.1
 #Pinned versions: 2.2.1
 #test that import: test_sac_estimator.py
 # To build PyTorch itself
 astunparse
 PyYAML
 pyyaml
 pyzstd
 setuptools
 setuptools>=70.1.0
 six
 scons==4.5.2 ; platform_machine == "aarch64"
 @ -382,3 +381,16 @@ dataclasses_json==0.6.7
 cmake==4.0.0
 #Description: required for building
 tlparse==0.3.30
 #Description: required for log parsing
 cuda-bindings>=12.0,<13.0 ; platform_machine != "s390x"
 #Description: required for testing CUDAGraph::raw_cuda_graph(). See https://nvidia.github.io/cuda-python/cuda-bindings/latest/support.html for how this version was chosen. Note "Any fix in the latest bindings would be backported to the prior major version" means that only the newest version of cuda-bindings will get fixes. Depending on the latest version of 12.x is okay because all 12.y versions will be supported via "CUDA minor version compatibility". Pytorch builds against 13.z versions of cuda toolkit work with 12.x versions of cuda-bindings as well because newer drivers work with old toolkits.
 #test that import: test_cuda.py
 setuptools-git-versioning==2.1.0
 scikit-build==0.18.1
 pyre-extensions==0.0.32
 tabulate==0.9.0
 #Description: These package are needed to build FBGEMM and torchrec on PyTorch CI

19

.ci/docker/requirements-docs.txt

View File

 @ -1,11 +1,11 @@
 sphinx==5.3.0
 #Description: This is used to generate PyTorch docs
 #Pinned versions: 5.3.0
 -e git+https://github.com/pytorch/pytorch_sphinx_theme.git@pytorch_sphinx_theme2#egg=pytorch_sphinx_theme2
 -e git+https://github.com/pytorch/pytorch_sphinx_theme.git@722b7e6f9ca512fcc526ad07d62b3d28c50bb6cd#egg=pytorch_sphinx_theme2
 # TODO: sphinxcontrib.katex 0.9.0 adds a local KaTeX server to speed up pre-rendering
 # but it doesn't seem to work and hangs around idly. The initial thought is probably
 # something related to Docker setup. We can investigate this later
 # but it doesn't seem to work and hangs around idly. The initial thought that it is probably
 # something related to Docker setup. We can investigate this later.
 sphinxcontrib.katex==0.8.6
 #Description: This is used to generate PyTorch docs
 @ -15,9 +15,14 @@ sphinxext-opengraph==0.9.1
 #Description: This is used to generate PyTorch docs
 #Pinned versions: 0.9.1
 matplotlib==3.5.3
 sphinx_sitemap==2.6.0
 #Description: This is used to generate sitemap for PyTorch docs
 #Pinned versions: 2.6.0
 matplotlib==3.5.3 ; python_version < "3.13"
 matplotlib==3.6.3 ; python_version >= "3.13"
 #Description: This is used to generate PyTorch docs
 #Pinned versions: 3.5.3
 #Pinned versions: 3.6.3 if python > 3.12. Otherwise 3.5.3.
 tensorboard==2.13.0 ; python_version < "3.13"
 tensorboard==2.18.0 ; python_version >= "3.13"
 @ -45,8 +50,8 @@ IPython==8.12.0
 #Pinned versions: 8.12.0
 myst-nb==0.17.2
 #Description: This is used to generate PyTorch functorch docs
 #Pinned versions: 0.13.2
 #Description: This is used to generate PyTorch functorch and torch.compile docs.
 #Pinned versions: 0.17.2
 # The following are required to build torch.distributed.elastic.rendezvous.etcd* docs
 python-etcd==0.4.5

2

.ci/docker/triton_version.txt

View File

 @ -1 +1 @@
 .3.0
 .4.0

1

.ci/docker/triton_xpu_version.txt Normal file

View File

				`@ -0,0 +1 @@`
				`3.4.0`

									
										170

.ci/docker/ubuntu-cuda/Dockerfile
									
												View File
											
				@ -1,170 +0,0 @@

				ARG UBUNTU_VERSION

				ARG CUDA_VERSION

				ARG IMAGE_NAME

				FROM ${IMAGE_NAME} as base

				ARG UBUNTU_VERSION

				ARG CUDA_VERSION

				ENV DEBIAN_FRONTEND noninteractive

				# Install common dependencies (so that this step can be cached separately)

				COPY ./common/install_base.sh install_base.sh

				RUN bash ./install_base.sh && rm install_base.sh

				# Install user

				COPY ./common/install_user.sh install_user.sh

				RUN bash ./install_user.sh && rm install_user.sh

				# Install katex

				ARG KATEX

				COPY ./common/install_docs_reqs.sh install_docs_reqs.sh

				RUN bash ./install_docs_reqs.sh && rm install_docs_reqs.sh

				# Install conda and other packages (e.g., numpy, pytest)

				ARG ANACONDA_PYTHON_VERSION

				ENV ANACONDA_PYTHON_VERSION=$ANACONDA_PYTHON_VERSION

				ENV PATH /opt/conda/envs/py_$ANACONDA_PYTHON_VERSION/bin:/opt/conda/bin:$PATH

				COPY requirements-ci.txt /opt/conda/requirements-ci.txt

				COPY ./common/install_conda.sh install_conda.sh

				COPY ./common/common_utils.sh common_utils.sh

				COPY ./common/install_magma_conda.sh install_magma_conda.sh

				RUN bash ./install_conda.sh && rm install_conda.sh install_magma_conda.sh common_utils.sh /opt/conda/requirements-ci.txt

				# Install gcc

				ARG GCC_VERSION

				COPY ./common/install_gcc.sh install_gcc.sh

				RUN bash ./install_gcc.sh && rm install_gcc.sh

				# Install clang

				ARG CLANG_VERSION

				COPY ./common/install_clang.sh install_clang.sh

				RUN bash ./install_clang.sh && rm install_clang.sh

				# (optional) Install vision packages like OpenCV

				ARG VISION

				COPY ./common/install_vision.sh ./common/cache_vision_models.sh ./common/common_utils.sh ./

				RUN if [ -n "${VISION}" ]; then bash ./install_vision.sh; fi

				RUN rm install_vision.sh cache_vision_models.sh common_utils.sh

				ENV INSTALLED_VISION ${VISION}

				# (optional) Install UCC

				ARG UCX_COMMIT

				ARG UCC_COMMIT

				ENV UCX_COMMIT $UCX_COMMIT

				ENV UCC_COMMIT $UCC_COMMIT

				ENV UCX_HOME /usr

				ENV UCC_HOME /usr

				ADD ./common/install_ucc.sh install_ucc.sh

				RUN if [ -n "${UCX_COMMIT}" ] && [ -n "${UCC_COMMIT}" ]; then bash ./install_ucc.sh; fi

				RUN rm install_ucc.sh

				COPY ./common/install_openssl.sh install_openssl.sh

				ENV OPENSSL_ROOT_DIR /opt/openssl

				RUN bash ./install_openssl.sh

				ENV OPENSSL_DIR /opt/openssl

				ARG INDUCTOR_BENCHMARKS

				ARG ANACONDA_PYTHON_VERSION

				ENV ANACONDA_PYTHON_VERSION=$ANACONDA_PYTHON_VERSION

				COPY ./common/install_inductor_benchmark_deps.sh install_inductor_benchmark_deps.sh

				COPY ./common/common_utils.sh common_utils.sh

				COPY ci_commit_pins/huggingface.txt huggingface.txt

				COPY ci_commit_pins/timm.txt timm.txt

				RUN if [ -n "${INDUCTOR_BENCHMARKS}" ]; then bash ./install_inductor_benchmark_deps.sh; fi

				RUN rm install_inductor_benchmark_deps.sh common_utils.sh timm.txt huggingface.txt

				ARG TRITON

				FROM base as triton-builder

				# Install triton, this needs to be done before sccache because the latter will

				# try to reach out to S3, which docker build runners don't have access

				COPY ./common/install_triton.sh install_triton.sh

				COPY ./common/common_utils.sh common_utils.sh

				COPY ci_commit_pins/triton.txt triton.txt

				COPY triton_version.txt triton_version.txt

				RUN bash ./install_triton.sh

				FROM base as final

				COPY --from=triton-builder /opt/triton /opt/triton

				RUN if [ -n "${TRITON}" ]; then pip install /opt/triton/*.whl; chown -R jenkins:jenkins /opt/conda; fi

				RUN rm -rf /opt/triton

				ARG HALIDE

				# Build and install halide

				COPY ./common/install_halide.sh install_halide.sh

				COPY ./common/common_utils.sh common_utils.sh

				COPY ci_commit_pins/halide.txt halide.txt

				RUN if [ -n "${HALIDE}" ]; then bash ./install_halide.sh; fi

				RUN rm install_halide.sh common_utils.sh halide.txt

				# Install ccache/sccache (do this last, so we get priority in PATH)

				COPY ./common/install_cache.sh install_cache.sh

				ENV PATH /opt/cache/bin:$PATH

				# See https://github.com/pytorch/pytorch/issues/82174

				# TODO(sdym@fb.com):

				# check if this is needed after full off Xenial migration

				ENV CARGO_NET_GIT_FETCH_WITH_CLI true

				RUN bash ./install_cache.sh && rm install_cache.sh

				ENV CMAKE_CUDA_COMPILER_LAUNCHER=/opt/cache/bin/sccache

				# Add jni.h for java host build

				COPY ./common/install_jni.sh install_jni.sh

				COPY ./java/jni.h jni.h

				RUN bash ./install_jni.sh && rm install_jni.sh

				# Install Open MPI for CUDA

				COPY ./common/install_openmpi.sh install_openmpi.sh

				RUN if [ -n "${CUDA_VERSION}" ]; then bash install_openmpi.sh; fi

				RUN rm install_openmpi.sh

				# Include BUILD_ENVIRONMENT environment variable in image

				ARG BUILD_ENVIRONMENT

				ENV BUILD_ENVIRONMENT ${BUILD_ENVIRONMENT}

				# AWS specific CUDA build guidance

				ENV TORCH_CUDA_ARCH_LIST Maxwell

				ENV TORCH_NVCC_FLAGS "-Xfatbin -compress-all"

				ENV CUDA_PATH /usr/local/cuda

				# Install LLVM dev version (Defined in the pytorch/builder github repository)

				COPY --from=pytorch/llvm:9.0.1 /opt/llvm /opt/llvm

				# Install CUDNN

				ARG CUDNN_VERSION

				ARG CUDA_VERSION

				COPY ./common/install_cudnn.sh install_cudnn.sh

				RUN if [ -n "${CUDNN_VERSION}" ]; then bash install_cudnn.sh; fi

				RUN rm install_cudnn.sh

				# Install CUSPARSELT

				ARG CUDA_VERSION

				COPY ./common/install_cusparselt.sh install_cusparselt.sh

				RUN bash install_cusparselt.sh

				RUN rm install_cusparselt.sh

				# Install NCCL

				ARG CUDA_VERSION

				COPY ./common/install_nccl.sh install_nccl.sh

				COPY ./ci_commit_pins/nccl-cu* /ci_commit_pins/

				RUN bash install_nccl.sh

				RUN rm install_nccl.sh /ci_commit_pins/nccl-cu*

				ENV USE_SYSTEM_NCCL=1

				ENV NCCL_INCLUDE_DIR="/usr/local/cuda/include/"

				ENV NCCL_LIB_DIR="/usr/local/cuda/lib64/"

				# Install CUDSS

				ARG CUDA_VERSION

				COPY ./common/install_cudss.sh install_cudss.sh

				RUN bash install_cudss.sh

				RUN rm install_cudss.sh

				# Delete /usr/local/cuda-11.X/cuda-11.X symlinks

				RUN if [ -h /usr/local/cuda-11.6/cuda-11.6 ]; then rm /usr/local/cuda-11.6/cuda-11.6; fi

				RUN if [ -h /usr/local/cuda-11.7/cuda-11.7 ]; then rm /usr/local/cuda-11.7/cuda-11.7; fi

				RUN if [ -h /usr/local/cuda-12.1/cuda-12.1 ]; then rm /usr/local/cuda-12.1/cuda-12.1; fi

				RUN if [ -h /usr/local/cuda-12.4/cuda-12.4 ]; then rm /usr/local/cuda-12.4/cuda-12.4; fi

				USER jenkins

				CMD ["bash"]

									
										4

.ci/docker/ubuntu-rocm/Dockerfile
									
												View File
												
				@ -25,6 +25,7 @@ RUN bash ./install_docs_reqs.sh && rm install_docs_reqs.sh

				# Install conda and other packages (e.g., numpy, pytest)

				ARG ANACONDA_PYTHON_VERSION

				ARG BUILD_ENVIRONMENT

				ENV ANACONDA_PYTHON_VERSION=$ANACONDA_PYTHON_VERSION

				ENV PATH /opt/conda/envs/py_$ANACONDA_PYTHON_VERSION/bin:/opt/conda/bin:$PATH

				COPY requirements-ci.txt /opt/conda/requirements-ci.txt

				@ -97,8 +98,9 @@ COPY ./common/install_inductor_benchmark_deps.sh install_inductor_benchmark_deps

				COPY ./common/common_utils.sh common_utils.sh

				COPY ci_commit_pins/huggingface.txt huggingface.txt

				COPY ci_commit_pins/timm.txt timm.txt

				COPY ci_commit_pins/torchbench.txt torchbench.txt

				RUN if [ -n "${INDUCTOR_BENCHMARKS}" ]; then bash ./install_inductor_benchmark_deps.sh; fi

				RUN rm install_inductor_benchmark_deps.sh common_utils.sh timm.txt huggingface.txt

				RUN rm install_inductor_benchmark_deps.sh common_utils.sh timm.txt huggingface.txt torchbench.txt

				# (optional) Install non-default Ninja version

				ARG NINJA_VERSION

									
										2

.ci/docker/ubuntu-xpu/Dockerfile
									
												View File
												
				@ -72,7 +72,7 @@ ARG TRITON

				COPY ./common/install_triton.sh install_triton.sh

				COPY ./common/common_utils.sh common_utils.sh

				COPY ci_commit_pins/triton-xpu.txt triton-xpu.txt

				COPY triton_version.txt triton_version.txt

				COPY triton_xpu_version.txt triton_version.txt

				RUN if [ -n "${TRITON}" ]; then bash ./install_triton.sh; fi

				RUN rm install_triton.sh common_utils.sh triton-xpu.txt triton_version.txt

									
										9

.ci/docker/ubuntu/Dockerfile
									
												View File
												
				@ -98,8 +98,9 @@ COPY ./common/install_inductor_benchmark_deps.sh install_inductor_benchmark_deps

				COPY ./common/common_utils.sh common_utils.sh

				COPY ci_commit_pins/huggingface.txt huggingface.txt

				COPY ci_commit_pins/timm.txt timm.txt

				COPY ci_commit_pins/torchbench.txt torchbench.txt

				RUN if [ -n "${INDUCTOR_BENCHMARKS}" ]; then bash ./install_inductor_benchmark_deps.sh; fi

				RUN rm install_inductor_benchmark_deps.sh common_utils.sh timm.txt huggingface.txt

				RUN rm install_inductor_benchmark_deps.sh common_utils.sh timm.txt huggingface.txt torchbench.txt

				ARG TRITON

				ARG TRITON_CPU

				@ -147,6 +148,12 @@ RUN if [ -n "${ACL}" ]; then bash ./install_acl.sh; fi

				RUN rm install_acl.sh

				ENV INSTALLED_ACL ${ACL}

				ARG OPENBLAS

				COPY ./common/install_openblas.sh install_openblas.sh

				RUN if [ -n "${OPENBLAS}" ]; then bash ./install_openblas.sh; fi

				RUN rm install_openblas.sh

				ENV INSTALLED_OPENBLAS ${OPENBLAS}

				# Install ccache/sccache (do this last, so we get priority in PATH)

				ARG SKIP_SCCACHE_INSTALL

				COPY ./common/install_cache.sh install_cache.sh

									
										16

.ci/magma/Makefile
									
												View File
												
				@ -1,7 +1,7 @@

				SHELL=/usr/bin/env bash

				DOCKER_CMD ?= docker

				DESIRED_CUDA ?= 11.8

				DESIRED_CUDA ?= 12.8

				DESIRED_CUDA_SHORT = $(subst .,,$(DESIRED_CUDA))

				PACKAGE_NAME = magma-cuda

				CUDA_ARCH_LIST ?= -gencode arch=compute_50,code=sm_50 -gencode arch=compute_60,code=sm_60 -gencode arch=compute_70,code=sm_70 -gencode arch=compute_80,code=sm_80 -gencode arch=compute_86,code=sm_86 -gencode arch=compute_90,code=sm_90

				@ -16,15 +16,21 @@ DOCKER_RUN = set -eou pipefail; ${DOCKER_CMD} run --rm -i \

					magma/build_magma.sh

				.PHONY: all

				all: magma-cuda129

				all: magma-cuda128

				all: magma-cuda126

				all: magma-cuda118

				.PHONY:

				clean:

					$(RM) -r magma-*

					$(RM) -r output

				.PHONY: magma-cuda129

				magma-cuda129: DESIRED_CUDA := 12.9

				magma-cuda129: CUDA_ARCH_LIST += -gencode arch=compute_100,code=sm_100 -gencode arch=compute_120,code=sm_120

				magma-cuda129:

					$(DOCKER_RUN)

				.PHONY: magma-cuda128

				magma-cuda128: DESIRED_CUDA := 12.8

				magma-cuda128: CUDA_ARCH_LIST += -gencode arch=compute_100,code=sm_100 -gencode arch=compute_120,code=sm_120

				@ -35,9 +41,3 @@ magma-cuda128:

				magma-cuda126: DESIRED_CUDA := 12.6

				magma-cuda126:

					$(DOCKER_RUN)

				.PHONY: magma-cuda118

				magma-cuda118: DESIRED_CUDA := 11.8

				magma-cuda118: CUDA_ARCH_LIST += -gencode arch=compute_37,code=sm_37

				magma-cuda118:

					$(DOCKER_RUN)

									
										15

.ci/manywheel/build_common.sh
									
												View File
												
				@ -18,12 +18,10 @@ retry () {

				    $*  || (sleep 1 && $*) || (sleep 2 && $*) || (sleep 4 && $*) || (sleep 8 && $*)

				}

				PLATFORM="manylinux2014_x86_64"

				PLATFORM=""

				# TODO move this into the Docker images

				OS_NAME=$(awk -F= '/^NAME/{print $2}' /etc/os-release)

				if [[ "$OS_NAME" == *"CentOS Linux"* ]]; then

				    retry yum install -q -y zip openssl

				elif [[ "$OS_NAME" == *"AlmaLinux"* ]]; then

				if [[ "$OS_NAME" == *"AlmaLinux"* ]]; then

				    retry yum install -q -y zip openssl

				    PLATFORM="manylinux_2_28_x86_64"

				elif [[ "$OS_NAME" == *"Red Hat Enterprise Linux"* ]]; then

				@ -33,9 +31,11 @@ elif [[ "$OS_NAME" == *"Ubuntu"* ]]; then

				    # Comment out nvidia repositories to prevent them from getting apt-get updated, see https://github.com/pytorch/pytorch/issues/74968

				    # shellcheck disable=SC2046

				    sed -i 's/.*nvidia.*/# &/' $(find /etc/apt/ -type f -name "*.list")

				    retry apt-get update

				    retry apt-get -y install zip openssl

				else

				    echo "Unknown OS: '$OS_NAME'"

				    exit 1

				fi

				# We use the package name to test the package by passing this to 'pip install'

				@ -79,8 +79,6 @@ if [[ -e /opt/openssl ]]; then

				    export CMAKE_INCLUDE_PATH="/opt/openssl/include":$CMAKE_INCLUDE_PATH

				fi

				mkdir -p /tmp/$WHEELHOUSE_DIR

				export PATCHELF_BIN=/usr/local/bin/patchelf

				@ -99,6 +97,7 @@ if [[ -z "$PYTORCH_ROOT" ]]; then

				    exit 1

				fi

				pushd "$PYTORCH_ROOT"

				retry pip install -qUr requirements-build.txt

				python setup.py clean

				retry pip install -qr requirements.txt

				case ${DESIRED_PYTHON} in

				@ -152,7 +151,7 @@ if [[ "$USE_SPLIT_BUILD" == "true" ]]; then

				    BUILD_LIBTORCH_WHL=0 BUILD_PYTHON_ONLY=1 \

				    BUILD_LIBTORCH_CPU_WITH_DEBUG=$BUILD_DEBUG_INFO \

				    USE_NCCL=${USE_NCCL} USE_RCCL=${USE_RCCL} USE_KINETO=${USE_KINETO} \

				    python setup.py bdist_wheel -d /tmp/$WHEELHOUSE_DIR --cmake

				    CMAKE_FRESH=1 python setup.py bdist_wheel -d /tmp/$WHEELHOUSE_DIR

				    echo "Finished setup.py bdist_wheel for split build (BUILD_PYTHON_ONLY)"

				else

				    time CMAKE_ARGS=${CMAKE_ARGS[@]} \

									
										164

.ci/manywheel/build_cuda.sh
									
												View File
												
				@ -15,6 +15,9 @@ export INSTALL_TEST=0 # dont install test binaries into site-packages

				export USE_CUPTI_SO=0

				export USE_CUSPARSELT=${USE_CUSPARSELT:-1} # Enable if not disabled by libtorch build

				export USE_CUFILE=${USE_CUFILE:-1}

				export USE_SYSTEM_NCCL=1

				export NCCL_INCLUDE_DIR="/usr/local/cuda/include/"

				export NCCL_LIB_DIR="/usr/local/cuda/lib64/"

				# Keep an array of cmake variables to add to

				if [[ -z "$CMAKE_ARGS" ]]; then

				@ -36,10 +39,8 @@ if [[ -n "$DESIRED_CUDA" ]]; then

				    if [[ ${DESIRED_CUDA} =~ ^[0-9]+\.[0-9]+$ ]]; then

				        CUDA_VERSION=${DESIRED_CUDA}

				    else

				        # cu90, cu92, cu100, cu101

				        if [[ ${#DESIRED_CUDA} -eq 4 ]]; then

				            CUDA_VERSION="${DESIRED_CUDA:2:1}.${DESIRED_CUDA:3:1}"

				        elif [[ ${#DESIRED_CUDA} -eq 5 ]]; then

				        # cu126, cu128 etc...

				        if [[ ${#DESIRED_CUDA} -eq 5 ]]; then

				            CUDA_VERSION="${DESIRED_CUDA:2:2}.${DESIRED_CUDA:4:1}"

				        fi

				    fi

				@ -50,24 +51,23 @@ else

				fi

				cuda_version_nodot=$(echo $CUDA_VERSION | tr -d '.')

				EXTRA_CAFFE2_CMAKE_FLAGS+=("-DATEN_NO_TEST=ON")

				TORCH_CUDA_ARCH_LIST="5.0;6.0;7.0;7.5;8.0;8.6"

				case ${CUDA_VERSION} in

				    #removing sm_50-sm_60 as these architectures are deprecated in CUDA 12.8/9 and will be removed in future releases

				    #however we would like to keep sm_70 architecture see: https://github.com/pytorch/pytorch/issues/157517

				    12.8)

				        TORCH_CUDA_ARCH_LIST="7.5;8.0;8.6;9.0;10.0;12.0+PTX" #removing sm_50-sm_70 as these architectures are deprecated in CUDA 12.8 and will be removed in future releases

				        EXTRA_CAFFE2_CMAKE_FLAGS+=("-DATEN_NO_TEST=ON")

				        TORCH_CUDA_ARCH_LIST="7.0;7.5;8.0;8.6;9.0;10.0;12.0"

				        ;;

				    12.9)

				        TORCH_CUDA_ARCH_LIST="7.0;7.5;8.0;8.6;9.0;10.0;12.0+PTX"

				        # WAR to resolve the ld error in libtorch build with CUDA 12.9

				        if [[ "$PACKAGE_TYPE" == "libtorch" ]]; then

				            TORCH_CUDA_ARCH_LIST="7.5;8.0;9.0;10.0;12.0+PTX"

				        fi

				        ;;

				    12.6)

				        TORCH_CUDA_ARCH_LIST="${TORCH_CUDA_ARCH_LIST};9.0"

				        EXTRA_CAFFE2_CMAKE_FLAGS+=("-DATEN_NO_TEST=ON")

				        ;;

				    12.4)

				        TORCH_CUDA_ARCH_LIST="${TORCH_CUDA_ARCH_LIST};9.0"

				        EXTRA_CAFFE2_CMAKE_FLAGS+=("-DATEN_NO_TEST=ON")

				        ;;

				    11.8)

				        TORCH_CUDA_ARCH_LIST="${TORCH_CUDA_ARCH_LIST};3.7;9.0"

				        EXTRA_CAFFE2_CMAKE_FLAGS+=("-DATEN_NO_TEST=ON")

				        TORCH_CUDA_ARCH_LIST="5.0;6.0;7.0;7.5;8.0;8.6;9.0"

				        ;;

				    *)

				        echo "unknown cuda version $CUDA_VERSION"

				@ -91,14 +91,15 @@ fi

				mkdir -p "$PYTORCH_FINAL_PACKAGE_DIR" || true

				OS_NAME=$(awk -F= '/^NAME/{print $2}' /etc/os-release)

				if [[ "$OS_NAME" == *"CentOS Linux"* ]]; then

				    LIBGOMP_PATH="/usr/lib64/libgomp.so.1"

				elif [[ "$OS_NAME" == *"AlmaLinux"* ]]; then

				if [[ "$OS_NAME" == *"AlmaLinux"* ]]; then

				    LIBGOMP_PATH="/usr/lib64/libgomp.so.1"

				elif [[ "$OS_NAME" == *"Red Hat Enterprise Linux"* ]]; then

				    LIBGOMP_PATH="/usr/lib64/libgomp.so.1"

				elif [[ "$OS_NAME" == *"Ubuntu"* ]]; then

				    LIBGOMP_PATH="/usr/lib/x86_64-linux-gnu/libgomp.so.1"

				else

				    echo "Unknown OS: '$OS_NAME'"

				    exit 1

				fi

				DEPS_LIST=(

				@ -108,31 +109,12 @@ DEPS_SONAME=(

				    "libgomp.so.1"

				)

				# CUDA 11.8 have to ship the libcusparseLt.so.0 with the binary

				# since nvidia-cusparselt-cu11 is not available in PYPI

				if [[ $USE_CUSPARSELT == "1" && $CUDA_VERSION == "11.8" ]]; then

				        DEPS_SONAME+=(

				            "libcusparseLt.so.0"

				        )

				        DEPS_LIST+=(

				            "/usr/local/cuda/lib64/libcusparseLt.so.0"

				        )

				fi

				# Turn USE_CUFILE off for CUDA 11.8, 12.4 since nvidia-cufile-cu11 and 1.9.0.20 are

				# not available in PYPI

				if [[ $CUDA_VERSION == "11.8" || $CUDA_VERSION == "12.4" ]]; then

				    export USE_CUFILE=0

				fi

				# CUDA_VERSION 12.4, 12.6, 12.8

				# CUDA_VERSION 12.6, 12.8, 12.9

				if [[ $CUDA_VERSION == 12* ]]; then

				    export USE_STATIC_CUDNN=0

				    # Try parallelizing nvcc as well

				    export TORCH_NVCC_FLAGS="-Xfatbin -compress-all --threads 2"

				    if [[ -z "$PYTORCH_EXTRA_INSTALL_REQUIREMENTS" ]]; then

				        echo "Bundling with cudnn and cublas."

				        DEPS_LIST+=(

				@ -148,9 +130,12 @@ if [[ $CUDA_VERSION == 12* ]]; then

				            "/usr/local/cuda/lib64/libcublasLt.so.12"

				            "/usr/local/cuda/lib64/libcusparseLt.so.0"

				            "/usr/local/cuda/lib64/libcudart.so.12"

				            "/usr/local/cuda/lib64/libnvToolsExt.so.1"

				            "/usr/local/cuda/lib64/libnvrtc.so.12"

				            "/usr/local/cuda/lib64/libnvrtc-builtins.so"

				            "/usr/local/cuda/lib64/libcufile.so.0"

				            "/usr/local/cuda/lib64/libcufile_rdma.so.1"

				            "/usr/local/cuda/extras/CUPTI/lib64/libcupti.so.12"

				            "/usr/local/cuda/extras/CUPTI/lib64/libnvperf_host.so"

				        )

				        DEPS_SONAME+=(

				            "libcudnn_adv.so.9"

				@ -165,19 +150,17 @@ if [[ $CUDA_VERSION == 12* ]]; then

				            "libcublasLt.so.12"

				            "libcusparseLt.so.0"

				            "libcudart.so.12"

				            "libnvToolsExt.so.1"

				            "libnvrtc.so.12"

				            "libnvrtc-builtins.so"

				            "libcufile.so.0"

				            "libcufile_rdma.so.1"

				            "libcupti.so.12"

				            "libnvperf_host.so"

				        )

				        if [[ $USE_CUFILE == 1 ]]; then

				            DEPS_LIST+=(

				                "/usr/local/cuda/lib64/libcufile.so.0"

				                "/usr/local/cuda/lib64/libcufile_rdma.so.1"

				            )

				            DEPS_SONAME+=(

				                "libcufile.so.0"

				                "libcufile_rdma.so.1"

				            )

				        # Add libnvToolsExt only if CUDA version is not 12.9

				        if [[ $CUDA_VERSION != 12.9* ]]; then

				            DEPS_LIST+=("/usr/local/cuda/lib64/libnvToolsExt.so.1")

				            DEPS_SONAME+=("libnvToolsExt.so.1")

				        fi

				    else

				        echo "Using nvidia libs from pypi."

				@ -191,94 +174,21 @@ if [[ $CUDA_VERSION == 12* ]]; then

				            '$ORIGIN/../../nvidia/curand/lib'

				            '$ORIGIN/../../nvidia/cusolver/lib'

				            '$ORIGIN/../../nvidia/cusparse/lib'

				            '$ORIGIN/../../nvidia/cusparselt/lib'

				            '$ORIGIN/../../cusparselt/lib'

				            '$ORIGIN/../../nvidia/nccl/lib'

				            '$ORIGIN/../../nvidia/nvshmem/lib'

				            '$ORIGIN/../../nvidia/nvtx/lib'

				        )

				        if [[ $USE_CUFILE == 1 ]]; then

				            CUDA_RPATHS+=(

				                '$ORIGIN/../../nvidia/cufile/lib'

				            )

				        fi

				        CUDA_RPATHS=$(IFS=: ; echo "${CUDA_RPATHS[*]}")

				        export C_SO_RPATH=$CUDA_RPATHS':$ORIGIN:$ORIGIN/lib'

				        export LIB_SO_RPATH=$CUDA_RPATHS':$ORIGIN'

				        export FORCE_RPATH="--force-rpath"

				        export USE_STATIC_NCCL=0

				        export USE_SYSTEM_NCCL=1

				        export ATEN_STATIC_CUDA=0

				        export USE_CUDA_STATIC_LINK=0

				        export USE_CUPTI_SO=1

				        export NCCL_INCLUDE_DIR="/usr/local/cuda/include/"

				        export NCCL_LIB_DIR="/usr/local/cuda/lib64/"

				    fi

				elif [[ $CUDA_VERSION == "11.8" ]]; then

				    export USE_STATIC_CUDNN=0

				    # Try parallelizing nvcc as well

				    export TORCH_NVCC_FLAGS="-Xfatbin -compress-all --threads 2"

				    # Bundle ptxas into the wheel, see https://github.com/pytorch/pytorch/pull/119750

				    export BUILD_BUNDLE_PTXAS=1

				    if [[ -z "$PYTORCH_EXTRA_INSTALL_REQUIREMENTS" ]]; then

				        echo "Bundling with cudnn and cublas."

				        DEPS_LIST+=(

				            "/usr/local/cuda/lib64/libcudnn_adv.so.9"

				            "/usr/local/cuda/lib64/libcudnn_cnn.so.9"

				            "/usr/local/cuda/lib64/libcudnn_graph.so.9"

				            "/usr/local/cuda/lib64/libcudnn_ops.so.9"

				            "/usr/local/cuda/lib64/libcudnn_engines_runtime_compiled.so.9"

				            "/usr/local/cuda/lib64/libcudnn_engines_precompiled.so.9"

				            "/usr/local/cuda/lib64/libcudnn_heuristic.so.9"

				            "/usr/local/cuda/lib64/libcudnn.so.9"

				            "/usr/local/cuda/lib64/libcublas.so.11"

				            "/usr/local/cuda/lib64/libcublasLt.so.11"

				            "/usr/local/cuda/lib64/libcudart.so.11.0"

				            "/usr/local/cuda/lib64/libnvToolsExt.so.1"

				            "/usr/local/cuda/lib64/libnvrtc.so.11.2"    # this is not a mistake, it links to more specific cuda version

				            "/usr/local/cuda/lib64/libnvrtc-builtins.so.11.8"

				        )

				        DEPS_SONAME+=(

				            "libcudnn_adv.so.9"

				            "libcudnn_cnn.so.9"

				            "libcudnn_graph.so.9"

				            "libcudnn_ops.so.9"

				            "libcudnn_engines_runtime_compiled.so.9"

				            "libcudnn_engines_precompiled.so.9"

				            "libcudnn_heuristic.so.9"

				            "libcudnn.so.9"

				            "libcublas.so.11"

				            "libcublasLt.so.11"

				            "libcudart.so.11.0"

				            "libnvToolsExt.so.1"

				            "libnvrtc.so.11.2"

				            "libnvrtc-builtins.so.11.8"

				        )

				    else

				        echo "Using nvidia libs from pypi."

				        CUDA_RPATHS=(

				            '$ORIGIN/../../nvidia/cublas/lib'

				            '$ORIGIN/../../nvidia/cuda_cupti/lib'

				            '$ORIGIN/../../nvidia/cuda_nvrtc/lib'

				            '$ORIGIN/../../nvidia/cuda_runtime/lib'

				            '$ORIGIN/../../nvidia/cudnn/lib'

				            '$ORIGIN/../../nvidia/cufft/lib'

				            '$ORIGIN/../../nvidia/curand/lib'

				            '$ORIGIN/../../nvidia/cusolver/lib'

				            '$ORIGIN/../../nvidia/cusparse/lib'

				            '$ORIGIN/../../nvidia/nccl/lib'

				            '$ORIGIN/../../nvidia/nvtx/lib'

				            '$ORIGIN/../../nvidia/cufile/lib'

				        )

				        CUDA_RPATHS=$(IFS=: ; echo "${CUDA_RPATHS[*]}")

				        export C_SO_RPATH=$CUDA_RPATHS':$ORIGIN:$ORIGIN/lib'

				        export LIB_SO_RPATH=$CUDA_RPATHS':$ORIGIN'

				        export FORCE_RPATH="--force-rpath"

				        export USE_STATIC_NCCL=0

				        export USE_SYSTEM_NCCL=1

				        export ATEN_STATIC_CUDA=0

				        export USE_CUDA_STATIC_LINK=0

				        export USE_CUPTI_SO=1

				        export NCCL_INCLUDE_DIR="/usr/local/cuda/include/"

				        export NCCL_LIB_DIR="/usr/local/cuda/lib64/"

				    fi

				else

				    echo "Unknown cuda version $CUDA_VERSION"

									
										12

.ci/manywheel/build_libtorch.sh
									
												View File
												
				@ -22,9 +22,7 @@ retry () {

				# TODO move this into the Docker images

				OS_NAME=`awk -F= '/^NAME/{print $2}' /etc/os-release`

				if [[ "$OS_NAME" == *"CentOS Linux"* ]]; then

				    retry yum install -q -y zip openssl

				elif [[ "$OS_NAME" == *"AlmaLinux"* ]]; then

				if [[ "$OS_NAME" == *"AlmaLinux"* ]]; then

				    retry yum install -q -y zip openssl

				elif [[ "$OS_NAME" == *"Red Hat Enterprise Linux"* ]]; then

				    retry dnf install -q -y zip openssl

				@ -35,6 +33,9 @@ elif [[ "$OS_NAME" == *"Ubuntu"* ]]; then

				    sed -i 's/.*nvidia.*/# &/' $(find /etc/apt/ -type f -name "*.list")

				    retry apt-get update

				    retry apt-get -y install zip openssl

				else

				    echo "Unknown OS: '$OS_NAME'"

				    exit 1

				fi

				# Version: setup.py uses $PYTORCH_BUILD_VERSION.post$PYTORCH_BUILD_NUMBER if

				@ -91,6 +92,7 @@ if [[ -z "$PYTORCH_ROOT" ]]; then

				    exit 1

				fi

				pushd "$PYTORCH_ROOT"

				retry pip install -qUr requirements-build.txt

				python setup.py clean

				retry pip install -qr requirements.txt

				retry pip install -q numpy==2.0.1

				@ -102,7 +104,7 @@ if [[ "$DESIRED_CUDA" == *"rocm"* ]]; then

				    export ROCclr_DIR=/opt/rocm/rocclr/lib/cmake/rocclr

				fi

				echo "Calling setup.py install at $(date)"

				echo "Calling 'python -m pip install .' at $(date)"

				if [[ $LIBTORCH_VARIANT = *"static"* ]]; then

				    STATIC_CMAKE_FLAG="-DTORCH_STATIC=1"

				@ -118,7 +120,7 @@ fi

				        # TODO: Remove this flag once https://github.com/pytorch/pytorch/issues/55952 is closed

				        CFLAGS='-Wno-deprecated-declarations' \

				        BUILD_LIBTORCH_CPU_WITH_DEBUG=1 \

				        python setup.py install

				        python -m pip install --no-build-isolation -v .

				    mkdir -p libtorch/{lib,bin,include,share}

									
										25

.ci/manywheel/build_rocm.sh
									
												View File
												
				@ -95,6 +95,7 @@ ROCM_SO_FILES=(

				    "libroctracer64.so"

				    "libroctx64.so"

				    "libhipblaslt.so"

				    "libhipsparselt.so"

				    "libhiprtc.so"

				)

				@ -186,20 +187,28 @@ do

				    OS_SO_FILES[${#OS_SO_FILES[@]}]=$file_name # Append lib to array

				done

				ARCH=$(echo $PYTORCH_ROCM_ARCH | sed 's/;/|/g') # Replace ; separated arch list to bar for grep

				# rocBLAS library files

				ROCBLAS_LIB_SRC=$ROCM_HOME/lib/rocblas/library

				ROCBLAS_LIB_DST=lib/rocblas/library

				ARCH=$(echo $PYTORCH_ROCM_ARCH | sed 's/;/|/g') # Replace ; seperated arch list to bar for grep

				ARCH_SPECIFIC_FILES=$(ls $ROCBLAS_LIB_SRC | grep -E $ARCH)

				OTHER_FILES=$(ls $ROCBLAS_LIB_SRC | grep -v gfx)

				ROCBLAS_LIB_FILES=($ARCH_SPECIFIC_FILES $OTHER_FILES)

				ROCBLAS_ARCH_SPECIFIC_FILES=$(ls $ROCBLAS_LIB_SRC | grep -E $ARCH)

				ROCBLAS_OTHER_FILES=$(ls $ROCBLAS_LIB_SRC | grep -v gfx)

				ROCBLAS_LIB_FILES=($ROCBLAS_ARCH_SPECIFIC_FILES $ROCBLAS_OTHER_FILES)

				# hipblaslt library files

				HIPBLASLT_LIB_SRC=$ROCM_HOME/lib/hipblaslt/library

				HIPBLASLT_LIB_DST=lib/hipblaslt/library

				ARCH_SPECIFIC_FILES=$(ls $HIPBLASLT_LIB_SRC | grep -E $ARCH)

				OTHER_FILES=$(ls $HIPBLASLT_LIB_SRC | grep -v gfx)

				HIPBLASLT_LIB_FILES=($ARCH_SPECIFIC_FILES $OTHER_FILES)

				HIPBLASLT_ARCH_SPECIFIC_FILES=$(ls $HIPBLASLT_LIB_SRC | grep -E $ARCH)

				HIPBLASLT_OTHER_FILES=$(ls $HIPBLASLT_LIB_SRC | grep -v gfx)

				HIPBLASLT_LIB_FILES=($HIPBLASLT_ARCH_SPECIFIC_FILES $HIPBLASLT_OTHER_FILES)

				# hipsparselt library files

				HIPSPARSELT_LIB_SRC=$ROCM_HOME/lib/hipsparselt/library

				HIPSPARSELT_LIB_DST=lib/hipsparselt/library

				HIPSPARSELT_ARCH_SPECIFIC_FILES=$(ls $HIPSPARSELT_LIB_SRC | grep -E $ARCH)

				#HIPSPARSELT_OTHER_FILES=$(ls $HIPSPARSELT_LIB_SRC | grep -v gfx)

				HIPSPARSELT_LIB_FILES=($HIPSPARSELT_ARCH_SPECIFIC_FILES $HIPSPARSELT_OTHER_FILES)

				# ROCm library files

				ROCM_SO_PATHS=()

				@ -234,12 +243,14 @@ DEPS_SONAME=(

				DEPS_AUX_SRCLIST=(

				    "${ROCBLAS_LIB_FILES[@]/#/$ROCBLAS_LIB_SRC/}"

				    "${HIPBLASLT_LIB_FILES[@]/#/$HIPBLASLT_LIB_SRC/}"

				    "${HIPSPARSELT_LIB_FILES[@]/#/$HIPSPARSELT_LIB_SRC/}"

				    "/opt/amdgpu/share/libdrm/amdgpu.ids"

				)

				DEPS_AUX_DSTLIST=(

				    "${ROCBLAS_LIB_FILES[@]/#/$ROCBLAS_LIB_DST/}"

				    "${HIPBLASLT_LIB_FILES[@]/#/$HIPBLASLT_LIB_DST/}"

				    "${HIPSPARSELT_LIB_FILES[@]/#/$HIPSPARSELT_LIB_DST/}"

				    "share/libdrm/amdgpu.ids"

				)

									
										2

.ci/onnx/test.sh
									
												View File
												
				@ -19,7 +19,7 @@ git config --global --add safe.directory /var/lib/jenkins/workspace

				if [[ "$BUILD_ENVIRONMENT" == *onnx* ]]; then

				  # TODO: This can be removed later once vision is also part of the Docker image

				  pip install -q --user --no-use-pep517 "git+https://github.com/pytorch/vision.git@$(cat .github/ci_commit_pins/vision.txt)"

				  pip install -q --no-use-pep517 "git+https://github.com/pytorch/vision.git@$(cat .github/ci_commit_pins/vision.txt)"

				  # JIT C++ extensions require ninja, so put it into PATH.

				  export PATH="/var/lib/jenkins/.local/bin:$PATH"

				  # NB: ONNX test is fast (~15m) so it's ok to retry it few more times to avoid any flaky issue, we

									
										34

.ci/pytorch/build-mobile.sh
									
												View File
											
				@ -1,34 +0,0 @@

				#!/usr/bin/env bash

				# DO NOT ADD 'set -x' not to reveal CircleCI secret context environment variables

				set -eu -o pipefail

				# This script uses linux host toolchain + mobile build options in order to

				# build & test mobile libtorch without having to setup Android/iOS

				# toolchain/simulator.

				# shellcheck source=./common.sh

				source "$(dirname "${BASH_SOURCE[0]}")/common.sh"

				# shellcheck source=./common-build.sh

				source "$(dirname "${BASH_SOURCE[0]}")/common-build.sh"

				# Install torch & torchvision - used to download & trace test model.

				# Ideally we should use the libtorch built on the PR so that backward

				# incompatible changes won't break this script - but it will significantly slow

				# down mobile CI jobs.

				# Here we install nightly instead of stable so that we have an option to

				# temporarily skip mobile CI jobs on BC-breaking PRs until they are in nightly.

				retry pip install --pre torch torchvision \

				  -f https://download.pytorch.org/whl/nightly/cpu/torch_nightly.html \

				  --progress-bar off

				# Run end-to-end process of building mobile library, linking into the predictor

				# binary, and running forward pass with a real model.

				if [[ "$BUILD_ENVIRONMENT" == *-mobile-custom-build-static* ]]; then

				  TEST_CUSTOM_BUILD_STATIC=1 test/mobile/custom_build/build.sh

				elif [[ "$BUILD_ENVIRONMENT" == *-mobile-lightweight-dispatch* ]]; then

				  test/mobile/lightweight_dispatch/build.sh

				else

				  TEST_DEFAULT_BUILD=1 test/mobile/custom_build/build.sh

				fi

				print_sccache_stats

									
										67

.ci/pytorch/build.sh
									
												View File
												
				@ -11,10 +11,6 @@ source "$(dirname "${BASH_SOURCE[0]}")/common.sh"

				# shellcheck source=./common-build.sh

				source "$(dirname "${BASH_SOURCE[0]}")/common-build.sh"

				if [[ "$BUILD_ENVIRONMENT" == *-mobile-*build* ]]; then

				  exec "$(dirname "${BASH_SOURCE[0]}")/build-mobile.sh" "$@"

				fi

				echo "Python version:"

				python --version

				@ -27,6 +23,12 @@ cmake --version

				echo "Environment variables:"

				env

				# The sccache wrapped version of nvcc gets put in /opt/cache/lib in docker since

				# there are some issues if it is always wrapped, so we need to add it to PATH

				# during CI builds.

				# https://github.com/pytorch/pytorch/blob/0b6c0898e6c352c8ea93daec854e704b41485375/.ci/docker/common/install_cache.sh#L97

				export PATH="/opt/cache/lib:$PATH"

				if [[ "$BUILD_ENVIRONMENT" == *cuda* ]]; then

				  # Use jemalloc during compilation to mitigate https://github.com/pytorch/pytorch/issues/116289

				  export LD_PRELOAD=/usr/lib/x86_64-linux-gnu/libjemalloc.so.2

				@ -52,12 +54,6 @@ fi

				export USE_LLVM=/opt/llvm

				export LLVM_DIR=/opt/llvm/lib/cmake/llvm

				if [[ "$BUILD_ENVIRONMENT" == *executorch* ]]; then

				  # To build test_edge_op_registration

				  export BUILD_EXECUTORCH=ON

				  export USE_CUDA=0

				fi

				if ! which conda; then

				  # In ROCm CIs, we are doing cross compilation on build machines with

				  # intel cpu and later run tests on machines with amd cpu.

				@ -124,26 +120,8 @@ if [[ "$BUILD_ENVIRONMENT" == *libtorch* ]]; then

				fi

				# Use special scripts for Android builds

				if [[ "${BUILD_ENVIRONMENT}" == *-android* ]]; then

				  export ANDROID_NDK=/opt/ndk

				  build_args=()

				  if [[ "${BUILD_ENVIRONMENT}" == *-arm-v7a* ]]; then

				    build_args+=("-DANDROID_ABI=armeabi-v7a")

				  elif [[ "${BUILD_ENVIRONMENT}" == *-arm-v8a* ]]; then

				    build_args+=("-DANDROID_ABI=arm64-v8a")

				  elif [[ "${BUILD_ENVIRONMENT}" == *-x86_32* ]]; then

				    build_args+=("-DANDROID_ABI=x86")

				  elif [[ "${BUILD_ENVIRONMENT}" == *-x86_64* ]]; then

				    build_args+=("-DANDROID_ABI=x86_64")

				  fi

				  if [[ "${BUILD_ENVIRONMENT}" == *vulkan* ]]; then

				    build_args+=("-DUSE_VULKAN=ON")

				  fi

				  build_args+=("-DUSE_LITE_INTERPRETER_PROFILER=OFF")

				  exec ./scripts/build_android.sh "${build_args[@]}" "$@"

				fi

				if [[ "$BUILD_ENVIRONMENT" != *android* && "$BUILD_ENVIRONMENT" == *vulkan* ]]; then

				if [[ "$BUILD_ENVIRONMENT" == *vulkan* ]]; then

				  export USE_VULKAN=1

				  # shellcheck disable=SC1091

				  source /var/lib/jenkins/vulkansdk/setup-env.sh

				@ -198,10 +176,8 @@ fi

				# We only build FlashAttention files for CUDA 8.0+, and they require large amounts of

				# memory to build and will OOM

				if [[ "$BUILD_ENVIRONMENT" == *cuda* ]] && [[ 1 -eq $(echo "${TORCH_CUDA_ARCH_LIST} >= 8.0" | bc) ]] && [ -z "$MAX_JOBS_OVERRIDE" ]; then

				  echo "WARNING: FlashAttention files require large amounts of memory to build and will OOM"

				  echo "Setting MAX_JOBS=(nproc-2)/3 to reduce memory usage"

				  export MAX_JOBS="$(( $(nproc --ignore=2) / 3 ))"

				if [[ "$BUILD_ENVIRONMENT" == *cuda* ]] && [[ 1 -eq $(echo "${TORCH_CUDA_ARCH_LIST} >= 8.0" | bc) ]]; then

				  export BUILD_CUSTOM_STEP="ninja -C build flash_attention -j 2"

				fi

				if [[ "${BUILD_ENVIRONMENT}" == *clang* ]]; then

				@ -227,7 +203,7 @@ if [[ "${BUILD_ENVIRONMENT}" == *-pch* ]]; then

				    export USE_PRECOMPILED_HEADERS=1

				fi

				if [[ "${BUILD_ENVIRONMENT}" != *android* && "${BUILD_ENVIRONMENT}" != *cuda* ]]; then

				if [[ "${BUILD_ENVIRONMENT}" != *cuda* ]]; then

				  export BUILD_STATIC_RUNTIME_BENCHMARK=ON

				fi

				@ -257,6 +233,7 @@ if [[ "$BUILD_ENVIRONMENT" == *-bazel-* ]]; then

				  set -e -o pipefail

				  get_bazel

				  python3 tools/optional_submodules.py checkout_eigen

				  # Leave 1 CPU free and use only up to 80% of memory to reduce the change of crashing

				  # the runner

				@ -307,6 +284,22 @@ else

				    fi

				    pip_install_whl "$(echo dist/*.whl)"

				    if [[ "${BUILD_ADDITIONAL_PACKAGES:-}" == *vision* ]]; then

				      install_torchvision

				    fi

				    if [[ "${BUILD_ADDITIONAL_PACKAGES:-}" == *audio* ]]; then

				      install_torchaudio

				    fi

				    if [[ "${BUILD_ADDITIONAL_PACKAGES:-}" == *torchrec* || "${BUILD_ADDITIONAL_PACKAGES:-}" == *fbgemm* ]]; then

				      install_torchrec_and_fbgemm

				    fi

				    if [[ "${BUILD_ADDITIONAL_PACKAGES:-}" == *torchao* ]]; then

				      install_torchao

				    fi

				    if [[ "$BUILD_ENVIRONMENT" == *xpu* ]]; then

				      echo "Checking that xpu is compiled"

				      pushd dist/

				@ -394,10 +387,8 @@ else

				    # This is an attempt to mitigate flaky libtorch build OOM error. By default, the build parallelization

				    # is set to be the number of CPU minus 2. So, let's try a more conservative value here. A 4xlarge has

				    # 16 CPUs

				    if [ -z "$MAX_JOBS_OVERRIDE" ]; then

				      MAX_JOBS=$(nproc --ignore=4)

				      export MAX_JOBS

				    fi

				    MAX_JOBS=$(nproc --ignore=4)

				    export MAX_JOBS

				    # NB: Install outside of source directory (at the same level as the root

				    # pytorch folder) so that it doesn't get cleaned away prior to docker push.

									
										2

.ci/pytorch/check_binary.sh
									
												View File
												
				@ -313,7 +313,7 @@ if [[ "$(uname)" == 'Linux' &&  "$PACKAGE_TYPE" == 'manywheel' ]]; then

				  # Please see issue for reference: https://github.com/pytorch/pytorch/issues/152426

				  if [[ "$(uname -m)" == "s390x" ]]; then

				    cxx_abi="19"

				  elif [[ "$DESIRED_CUDA" != 'cu118' && "$DESIRED_CUDA" != 'xpu' && "$DESIRED_CUDA" != 'rocm'* ]]; then

				  elif [[ "$DESIRED_CUDA" != 'xpu' && "$DESIRED_CUDA" != 'rocm'* ]]; then

				    cxx_abi="18"

				  else

				    cxx_abi="16"

									
										7

.ci/pytorch/common-build.sh
									
												View File
												
				@ -13,6 +13,13 @@ if [[ "$BUILD_ENVIRONMENT" != *win-* ]]; then

				    fi

				    if which sccache > /dev/null; then

				        # Clear SCCACHE_BUCKET and SCCACHE_REGION if they are empty, otherwise

				        # sccache will complain about invalid bucket configuration

				        if [[ -z "${SCCACHE_BUCKET:-}" ]]; then

				          unset SCCACHE_BUCKET

				          unset SCCACHE_REGION

				        fi

				        # Save sccache logs to file

				        sccache --stop-server > /dev/null  2>&1 || true

				        rm -f ~/sccache_error.log || true

									
										2

.ci/pytorch/common.sh
									
												View File
												
				@ -15,6 +15,6 @@ if [[ "${BUILD_ENVIRONMENT}" == *rocm* ]]; then

				  export PYTORCH_TEST_WITH_ROCM=1

				fi

				# TODO: Renable libtorch testing for MacOS, see https://github.com/pytorch/pytorch/issues/62598

				# TODO: Reenable libtorch testing for MacOS, see https://github.com/pytorch/pytorch/issues/62598

				# shellcheck disable=SC2034

				BUILD_TEST_LIBTORCH=0

									
										153

.ci/pytorch/common_utils.sh
									
												View File
												
				@ -78,6 +78,34 @@ function pip_install_whl() {

				  fi

				}

				function pip_build_and_install() {

				  local build_target=$1

				  local wheel_dir=$2

				  local found_whl=0

				  for file in "${wheel_dir}"/*.whl

				  do

				    if [[ -f "${file}" ]]; then

				      found_whl=1

				      break

				    fi

				  done

				  # Build the wheel if it doesn't exist

				  if [ "${found_whl}" == "0" ]; then

				    python3 -m pip wheel \

				      --no-build-isolation \

				      --no-deps \

				      --no-use-pep517 \

				      -w "${wheel_dir}" \

				      "${build_target}"

				  fi

				  for file in "${wheel_dir}"/*.whl

				  do

				    pip_install_whl "${file}"

				  done

				}

				function pip_install() {

				  # retry 3 times

				@ -124,14 +152,7 @@ function get_pinned_commit() {

				function install_torchaudio() {

				  local commit

				  commit=$(get_pinned_commit audio)

				  if [[ "$1" == "cuda" ]]; then

				    # TODO: This is better to be passed as a parameter from _linux-test workflow

				    # so that it can be consistent with what is set in build

				    TORCH_CUDA_ARCH_LIST="8.0;8.6" pip_install --no-use-pep517 --user "git+https://github.com/pytorch/audio.git@${commit}"

				  else

				    pip_install --no-use-pep517 --user "git+https://github.com/pytorch/audio.git@${commit}"

				  fi

				  pip_build_and_install "git+https://github.com/pytorch/audio.git@${commit}" dist/audio

				}

				function install_torchtext() {

				@ -139,8 +160,8 @@ function install_torchtext() {

				  local text_commit

				  data_commit=$(get_pinned_commit data)

				  text_commit=$(get_pinned_commit text)

				  pip_install --no-use-pep517 --user "git+https://github.com/pytorch/data.git@${data_commit}"

				  pip_install --no-use-pep517 --user "git+https://github.com/pytorch/text.git@${text_commit}"

				  pip_build_and_install "git+https://github.com/pytorch/data.git@${data_commit}" dist/data

				  pip_build_and_install "git+https://github.com/pytorch/text.git@${text_commit}" dist/text

				}

				function install_torchvision() {

				@ -153,17 +174,19 @@ function install_torchvision() {

				    echo 'char* dlerror(void) { return "";}'|gcc -fpic -shared -o "${HOME}/dlerror.so" -x c -

				    LD_PRELOAD=${orig_preload}:${HOME}/dlerror.so

				  fi

				  pip_install --no-use-pep517 --user "git+https://github.com/pytorch/vision.git@${commit}"

				  if [[ "${BUILD_ENVIRONMENT}" == *cuda* ]]; then

				    # Not sure if both are needed, but why not

				    export FORCE_CUDA=1

				    export WITH_CUDA=1

				  fi

				  pip_build_and_install "git+https://github.com/pytorch/vision.git@${commit}" dist/vision

				  if [ -n "${LD_PRELOAD}" ]; then

				    LD_PRELOAD=${orig_preload}

				  fi

				}

				function install_tlparse() {

				  pip_install --user "tlparse==0.3.30"

				  PATH="$(python -m site --user-base)/bin:$PATH"

				}

				function install_torchrec_and_fbgemm() {

				  local torchrec_commit

				  torchrec_commit=$(get_pinned_commit torchrec)

				@ -178,25 +201,71 @@ function install_torchrec_and_fbgemm() {

				  if [[ "$BUILD_ENVIRONMENT" == *rocm* ]] ; then

				    # install torchrec first because it installs fbgemm nightly on top of rocm fbgemm

				    pip_install --no-use-pep517 --user "git+https://github.com/pytorch/torchrec.git@${torchrec_commit}"

				    pip_build_and_install "git+https://github.com/pytorch/torchrec.git@${torchrec_commit}" dist/torchrec

				    pip_uninstall fbgemm-gpu-nightly

				    # Set ROCM_HOME isn't available, use ROCM_PATH if set or /opt/rocm

				    ROCM_HOME="${ROCM_HOME:-${ROCM_PATH:-/opt/rocm}}"

				    # Find rocm_version.h header file for ROCm version extract

				    rocm_version_h="${ROCM_HOME}/include/rocm-core/rocm_version.h"

				    if [ ! -f "$rocm_version_h" ]; then

				        rocm_version_h="${ROCM_HOME}/include/rocm_version.h"

				    fi

				    # Error out if rocm_version.h not found

				    if [ ! -f "$rocm_version_h" ]; then

				        echo "Error: rocm_version.h not found in expected locations." >&2

				        exit 1

				    fi

				    # Extract major, minor and patch ROCm version numbers

				    MAJOR_VERSION=$(grep 'ROCM_VERSION_MAJOR' "$rocm_version_h" | awk '{print $3}')

				    MINOR_VERSION=$(grep 'ROCM_VERSION_MINOR' "$rocm_version_h" | awk '{print $3}')

				    PATCH_VERSION=$(grep 'ROCM_VERSION_PATCH' "$rocm_version_h" | awk '{print $3}')

				    ROCM_INT=$((MAJOR_VERSION * 10000 + MINOR_VERSION * 100 + PATCH_VERSION))

				    echo "ROCm version: $ROCM_INT"

				    export BUILD_ROCM_VERSION="$MAJOR_VERSION.$MINOR_VERSION"

				    pip_install tabulate  # needed for newer fbgemm

				    pip_install patchelf  # needed for rocm fbgemm

				    git clone --recursive https://github.com/pytorch/fbgemm

				    pushd fbgemm/fbgemm_gpu

				    git checkout "${fbgemm_commit}"

				    python setup.py install \

				      --package_variant=rocm \

				      -DHIP_ROOT_DIR="${ROCM_PATH}" \

				      -DCMAKE_C_FLAGS="-DTORCH_USE_HIP_DSA" \

				      -DCMAKE_CXX_FLAGS="-DTORCH_USE_HIP_DSA"

				    popd

				    local wheel_dir=dist/fbgemm_gpu

				    local found_whl=0

				    for file in "${wheel_dir}"/*.whl

				    do

				      if [[ -f "${file}" ]]; then

				        found_whl=1

				        break

				      fi

				    done

				    # Build the wheel if it doesn't exist

				    if [ "${found_whl}" == "0" ]; then

				      git clone --recursive https://github.com/pytorch/fbgemm

				      pushd fbgemm/fbgemm_gpu

				      git checkout "${fbgemm_commit}" --recurse-submodules

				      python setup.py bdist_wheel \

				        --build-variant=rocm \

				        -DHIP_ROOT_DIR="${ROCM_PATH}" \

				        -DCMAKE_C_FLAGS="-DTORCH_USE_HIP_DSA" \

				        -DCMAKE_CXX_FLAGS="-DTORCH_USE_HIP_DSA"

				      popd

				      # Save the wheel before cleaning up

				      mkdir -p dist/fbgemm_gpu

				      cp fbgemm/fbgemm_gpu/dist/*.whl dist/fbgemm_gpu

				    fi

				    for file in "${wheel_dir}"/*.whl

				    do

				      pip_install_whl "${file}"

				    done

				    rm -rf fbgemm

				  else

				    # See https://github.com/pytorch/pytorch/issues/106971

				    CUDA_PATH=/usr/local/cuda-12.1 pip_install --no-use-pep517 --user "git+https://github.com/pytorch/FBGEMM.git@${fbgemm_commit}#egg=fbgemm-gpu&subdirectory=fbgemm_gpu"

				    pip_install --no-use-pep517 --user "git+https://github.com/pytorch/torchrec.git@${torchrec_commit}"

				    pip_build_and_install "git+https://github.com/pytorch/torchrec.git@${torchrec_commit}" dist/torchrec

				    pip_build_and_install "git+https://github.com/pytorch/FBGEMM.git@${fbgemm_commit}#subdirectory=fbgemm_gpu" dist/fbgemm_gpu

				  fi

				}

				@ -212,34 +281,10 @@ function clone_pytorch_xla() {

				  fi

				}

				function checkout_install_torchbench() {

				  local commit

				  commit=$(get_pinned_commit torchbench)

				  git clone https://github.com/pytorch/benchmark torchbench

				  pushd torchbench

				  git checkout "$commit"

				  if [ "$1" ]; then

				    python install.py --continue_on_fail models "$@"

				  else

				    # Occasionally the installation may fail on one model but it is ok to continue

				    # to install and test other models

				    python install.py --continue_on_fail

				  fi

				  # TODO (huydhn): transformers-4.44.2 added by https://github.com/pytorch/benchmark/pull/2488

				  # is regressing speedup metric. This needs to be investigated further

				  pip install transformers==4.38.1

				  echo "Print all dependencies after TorchBench is installed"

				  python -mpip freeze

				  popd

				}

				function install_torchao() {

				  local commit

				  commit=$(get_pinned_commit torchao)

				  pip_install --no-use-pep517 --user "git+https://github.com/pytorch/ao.git@${commit}"

				  pip_build_and_install "git+https://github.com/pytorch/ao.git@${commit}" dist/ao

				}

				function print_sccache_stats() {

									
										123

.ci/pytorch/create_test_cert.py
									
												View File
											
				@ -1,123 +0,0 @@

				from datetime import datetime, timedelta, timezone

				from tempfile import mkdtemp

				from cryptography import x509

				from cryptography.hazmat.primitives import hashes, serialization

				from cryptography.hazmat.primitives.asymmetric import rsa

				from cryptography.x509.oid import NameOID

				temp_dir = mkdtemp()

				print(temp_dir)

				def genrsa(path):

				    key = rsa.generate_private_key(

				        public_exponent=65537,

				        key_size=2048,

				    )

				    with open(path, "wb") as f:

				        f.write(

				            key.private_bytes(

				                encoding=serialization.Encoding.PEM,

				                format=serialization.PrivateFormat.TraditionalOpenSSL,

				                encryption_algorithm=serialization.NoEncryption(),

				            )

				        )

				    return key

				def create_cert(path, C, ST, L, O, key):

				    subject = issuer = x509.Name(

				        [

				            x509.NameAttribute(NameOID.COUNTRY_NAME, C),

				            x509.NameAttribute(NameOID.STATE_OR_PROVINCE_NAME, ST),

				            x509.NameAttribute(NameOID.LOCALITY_NAME, L),

				            x509.NameAttribute(NameOID.ORGANIZATION_NAME, O),

				        ]

				    )

				    cert = (

				        x509.CertificateBuilder()

				        .subject_name(subject)

				        .issuer_name(issuer)

				        .public_key(key.public_key())

				        .serial_number(x509.random_serial_number())

				        .not_valid_before(datetime.now(timezone.utc))

				        .not_valid_after(

				            # Our certificate will be valid for 10 days

				            datetime.now(timezone.utc) + timedelta(days=10)

				        )

				        .add_extension(

				            x509.BasicConstraints(ca=True, path_length=None),

				            critical=True,

				        )

				        .sign(key, hashes.SHA256())

				    )

				    # Write our certificate out to disk.

				    with open(path, "wb") as f:

				        f.write(cert.public_bytes(serialization.Encoding.PEM))

				    return cert

				def create_req(path, C, ST, L, O, key):

				    csr = (

				        x509.CertificateSigningRequestBuilder()

				        .subject_name(

				            x509.Name(

				                [

				                    # Provide various details about who we are.

				                    x509.NameAttribute(NameOID.COUNTRY_NAME, C),

				                    x509.NameAttribute(NameOID.STATE_OR_PROVINCE_NAME, ST),

				                    x509.NameAttribute(NameOID.LOCALITY_NAME, L),

				                    x509.NameAttribute(NameOID.ORGANIZATION_NAME, O),

				                ]

				            )

				        )

				        .sign(key, hashes.SHA256())

				    )

				    with open(path, "wb") as f:

				        f.write(csr.public_bytes(serialization.Encoding.PEM))

				    return csr

				def sign_certificate_request(path, csr_cert, ca_cert, private_ca_key):

				    cert = (

				        x509.CertificateBuilder()

				        .subject_name(csr_cert.subject)

				        .issuer_name(ca_cert.subject)

				        .public_key(csr_cert.public_key())

				        .serial_number(x509.random_serial_number())

				        .not_valid_before(datetime.now(timezone.utc))

				        .not_valid_after(

				            # Our certificate will be valid for 10 days

				            datetime.now(timezone.utc) + timedelta(days=10)

				            # Sign our certificate with our private key

				        )

				        .sign(private_ca_key, hashes.SHA256())

				    )

				    with open(path, "wb") as f:

				        f.write(cert.public_bytes(serialization.Encoding.PEM))

				    return cert

				ca_key = genrsa(temp_dir + "/ca.key")

				ca_cert = create_cert(

				    temp_dir + "/ca.pem",

				    "US",

				    "New York",

				    "New York",

				    "Gloo Certificate Authority",

				    ca_key,

				)

				pkey = genrsa(temp_dir + "/pkey.key")

				csr = create_req(

				    temp_dir + "/csr.csr",

				    "US",

				    "California",

				    "San Francisco",

				    "Gloo Testing Company",

				    pkey,

				)

				cert = sign_certificate_request(temp_dir + "/cert.pem", csr, ca_cert, ca_key)

									
										2

.ci/pytorch/macos-build.sh
									
												View File
												
				@ -40,7 +40,7 @@ if [[ ${BUILD_ENVIRONMENT} == *"distributed"* ]]; then

				else

				  # Explicitly set USE_DISTRIBUTED=0 to align with the default build config on mac. This also serves as the sole CI config that tests

				  # that building with USE_DISTRIBUTED=0 works at all. See https://github.com/pytorch/pytorch/issues/86448

				  USE_DISTRIBUTED=0 USE_OPENMP=1 MACOSX_DEPLOYMENT_TARGET=11.0 WERROR=1 BUILD_TEST=OFF USE_PYTORCH_METAL=1 python setup.py bdist_wheel

				  USE_DISTRIBUTED=0 USE_OPENMP=1 MACOSX_DEPLOYMENT_TARGET=11.0 WERROR=1 BUILD_TEST=OFF USE_PYTORCH_METAL=1 python setup.py bdist_wheel --plat-name macosx_11_0_arm64

				fi

				if which sccache > /dev/null; then

				  print_sccache_stats

									
										10

.ci/pytorch/macos-common.sh
									
												View File
												
				@ -20,14 +20,4 @@ print_cmake_info() {

				  CONDA_INSTALLATION_DIR=$(dirname "$CMAKE_EXEC")

				  # Print all libraries under cmake rpath for debugging

				  ls -la "$CONDA_INSTALLATION_DIR/../lib"

				  export CMAKE_EXEC

				  # Explicitly add conda env lib folder to cmake rpath to address the flaky issue

				  # where cmake dependencies couldn't be found. This seems to point to how conda

				  # links $CMAKE_EXEC to its package cache when cloning a new environment

				  install_name_tool -add_rpath @executable_path/../lib "${CMAKE_EXEC}" || true

				  # Adding the rpath will invalidate cmake signature, so signing it again here

				  # to trust the executable. EXC_BAD_ACCESS (SIGKILL (Code Signature Invalid))

				  # with an exit code 137 otherwise

				  codesign -f -s - "${CMAKE_EXEC}" || true

				}

									
										128

.ci/pytorch/macos-test.sh
									
												View File
												
				@ -5,11 +5,6 @@ set -x

				# shellcheck source=./macos-common.sh

				source "$(dirname "${BASH_SOURCE[0]}")/macos-common.sh"

				if [[ -n "$CONDA_ENV" ]]; then

				  # Use binaries under conda environment

				  export PATH="$CONDA_ENV/bin":$PATH

				fi

				# Test that OpenMP is enabled

				pushd test

				if [[ ! $(python -c "import torch; print(int(torch.backends.openmp.is_available()))") == "1" ]]; then

				@ -162,9 +157,33 @@ test_jit_hooks() {

				  assert_git_not_dirty

				}

				# Shellcheck doesn't like it when you pass no arguments to a function

				# that can take args. See https://www.shellcheck.net/wiki/SC2120

				# shellcheck disable=SC2120

				checkout_install_torchbench() {

				  local commit

				  commit=$(cat .ci/docker/ci_commit_pins/torchbench.txt)

				  git clone https://github.com/pytorch/benchmark torchbench

				  pushd torchbench

				  git checkout "$commit"

				  if [ "$1" ]; then

				    python install.py --continue_on_fail models "$@"

				  else

				    # Occasionally the installation may fail on one model but it is ok to continue

				    # to install and test other models

				    python install.py --continue_on_fail

				  fi

				  echo "Print all dependencies after TorchBench is installed"

				  python -mpip freeze

				  popd

				}

				torchbench_setup_macos() {

				  git clone --recursive https://github.com/pytorch/vision torchvision

				  git clone --recursive https://github.com/pytorch/audio torchaudio

				  brew install jpeg-turbo libpng

				  pushd torchvision

				  git fetch

				@ -179,17 +198,15 @@ torchbench_setup_macos() {

				  git checkout "$(cat ../.github/ci_commit_pins/audio.txt)"

				  git submodule update --init --recursive

				  python setup.py clean

				  python setup.py develop

				  #TODO: Remove me, when figure out how to make TorchAudio find brew installed openmp

				  USE_OPENMP=0 python setup.py develop

				  popd

				  # Shellcheck doesn't like it when you pass no arguments to a function that can take args. See https://www.shellcheck.net/wiki/SC2120

				  # shellcheck disable=SC2119,SC2120

				  checkout_install_torchbench

				}

				conda_benchmark_deps() {

				  conda install -y astunparse numpy scipy ninja pyyaml setuptools cmake typing-extensions requests protobuf numba cython scikit-learn

				  conda install -y -c conda-forge librosa

				pip_benchmark_deps() {

				  python -mpip install --no-input requests cython scikit-learn six

				}

				@ -197,7 +214,7 @@ test_torchbench_perf() {

				  print_cmake_info

				  echo "Launching torchbench setup"

				  conda_benchmark_deps

				  pip_benchmark_deps

				  torchbench_setup_macos

				  TEST_REPORTS_DIR=$(pwd)/test/test-reports

				@ -224,7 +241,7 @@ test_torchbench_smoketest() {

				  print_cmake_info

				  echo "Launching torchbench setup"

				  conda_benchmark_deps

				  pip_benchmark_deps

				  # shellcheck disable=SC2119,SC2120

				  torchbench_setup_macos

				@ -232,53 +249,52 @@ test_torchbench_smoketest() {

				  mkdir -p "$TEST_REPORTS_DIR"

				  local device=mps

				  local models=(hf_T5 llama BERT_pytorch dcgan hf_GPT2 yolov3 resnet152 sam pytorch_unet stable_diffusion_text_encoder speech_transformer Super_SloMo)

				  local hf_models=(GoogleFnet YituTechConvBert)

				  local dtypes=(undefined float16 bfloat16 notset)

				  local dtype=${dtypes[$1]}

				  local models=(hf_T5 llama BERT_pytorch dcgan hf_GPT2 yolov3 resnet152 sam sam_fast pytorch_unet stable_diffusion_text_encoder speech_transformer Super_SloMo doctr_det_predictor doctr_reco_predictor timm_resnet timm_vovnet vgg16)

				  for backend in eager inductor; do

				    for dtype in notset float16 bfloat16; do

				      echo "Launching torchbench inference performance run for backend ${backend} and dtype ${dtype}"

				      local dtype_arg="--${dtype}"

				      if [ "$dtype" == notset ]; then

				          dtype_arg="--float32"

				      fi

				      touch "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_inference_${device}_performance.csv"

				      for model in "${models[@]}"; do

				    echo "Launching torchbench inference performance run for backend ${backend} and dtype ${dtype}"

				    local dtype_arg="--${dtype}"

				    if [ "$dtype" == notset ]; then

				        dtype_arg="--float32"

				    fi

				    touch "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_inference_${device}_performance.csv"

				    for model in "${models[@]}"; do

				      PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/torchbench.py \

				        --performance --only "$model" --backend "$backend" --inference --devices "$device" "$dtype_arg" \

				        --output "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_inference_${device}_performance.csv" || true

				      if [ "$backend" == "inductor" ]; then

				        PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/torchbench.py \

				          --performance --only "$model" --backend "$backend" --inference --devices "$device" "$dtype_arg" \

				          --output "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_inference_${device}_performance.csv" || true

				        if [ "$backend" == "inductor" ]; then

				          PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/torchbench.py \

				            --accuracy --only "$model" --backend "$backend" --inference --devices "$device" "$dtype_arg" \

				            --output "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_inference_${device}_accuracy.csv" || true

				        fi

				      done

				      for model in "${hf_models[@]}"; do

				        if [ "$backend" == "inductor" ]; then

				          PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/huggingface.py \

				            --performance --only "$model" --backend "$backend" --inference --devices "$device" "$dtype_arg" \

				            --output "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_inference_${device}_performance.csv" || true

				          PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/huggingface.py \

				            --accuracy --only "$model" --backend "$backend" --inference --devices "$device" "$dtype_arg" \

				            --output "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_inference_${device}_accuracy.csv" || true

				        fi

				      done

				          --accuracy --only "$model" --backend "$backend" --inference --devices "$device" "$dtype_arg" \

				          --output "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_inference_${device}_accuracy.csv" || true

				      fi

				    done

				    if [ "$backend" == "inductor" ]; then

				      PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/huggingface.py \

				        --performance --backend "$backend" --inference --devices "$device" "$dtype_arg" \

				        --output "$TEST_REPORTS_DIR/inductor_${backend}_huggingface_${dtype}_inference_${device}_performance.csv" || true

				      PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/huggingface.py \

				        --accuracy --backend "$backend" --inference --devices "$device" "$dtype_arg" \

				        --output "$TEST_REPORTS_DIR/inductor_${backend}_huggingface_${dtype}_inference_${device}_accuracy.csv" || true

				    fi

				    for dtype in notset amp; do

				      echo "Launching torchbench training performance run for backend ${backend} and dtype ${dtype}"

				      touch "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_training_${device}_performance.csv"

				      local dtype_arg="--${dtype}"

				      if [ "$dtype" == notset ]; then

				    if [ "$dtype" == notset ]; then

				      for dtype_ in notset amp; do

				        echo "Launching torchbench training performance run for backend ${backend} and dtype ${dtype_}"

				        touch "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype_}_training_${device}_performance.csv"

				        local dtype_arg="--${dtype_}"

				        if [ "$dtype_" == notset ]; then

				          dtype_arg="--float32"

				      fi

				      for model in "${models[@]}"; do

				        PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/torchbench.py \

				          --performance --only "$model" --backend "$backend" --training --devices "$device" "$dtype_arg" \

				          --output "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype}_training_${device}_performance.csv" || true

				        fi

				        for model in "${models[@]}"; do

				          PYTHONPATH="$(pwd)"/torchbench python benchmarks/dynamo/torchbench.py \

				            --performance --only "$model" --backend "$backend" --training --devices "$device" "$dtype_arg" \

				            --output "$TEST_REPORTS_DIR/inductor_${backend}_torchbench_${dtype_}_training_${device}_performance.csv" || true

				        done

				      done

				    done

				    fi

				  done

				@ -289,7 +305,7 @@ test_hf_perf() {

				  print_cmake_info

				  TEST_REPORTS_DIR=$(pwd)/test/test-reports

				  mkdir -p "$TEST_REPORTS_DIR"

				  conda_benchmark_deps

				  pip_benchmark_deps

				  torchbench_setup_macos

				  echo "Launching HuggingFace training perf run"

				@ -305,7 +321,7 @@ test_timm_perf() {

				  print_cmake_info

				  TEST_REPORTS_DIR=$(pwd)/test/test-reports

				  mkdir -p "$TEST_REPORTS_DIR"

				  conda_benchmark_deps

				  pip_benchmark_deps

				  torchbench_setup_macos

				  echo "Launching timm training perf run"

				@ -317,8 +333,6 @@ test_timm_perf() {

				  echo "timm benchmark on mps device completed"

				}

				install_tlparse

				if [[ $TEST_CONFIG == *"perf_all"* ]]; then

				  test_torchbench_perf

				  test_hf_perf

				@ -330,7 +344,7 @@ elif [[ $TEST_CONFIG == *"perf_hf"* ]]; then

				elif [[ $TEST_CONFIG == *"perf_timm"* ]]; then

				  test_timm_perf

				elif [[ $TEST_CONFIG == *"perf_smoketest"* ]]; then

				  test_torchbench_smoketest

				  test_torchbench_smoketest "${SHARD_NUMBER}"

				elif [[ $TEST_CONFIG == *"mps"* ]]; then

				  test_python_mps

				elif [[ $NUM_TEST_SHARDS -gt 1 ]]; then

									
										18

.ci/pytorch/run_glootls_test.sh
									
												View File
											
				@ -1,18 +0,0 @@

				#!/bin/bash

				CREATE_TEST_CERT="$(dirname "${BASH_SOURCE[0]}")/create_test_cert.py"

				TMP_CERT_DIR=$(python "$CREATE_TEST_CERT")

				openssl verify -CAfile "${TMP_CERT_DIR}/ca.pem" "${TMP_CERT_DIR}/cert.pem"

				export GLOO_DEVICE_TRANSPORT=TCP_TLS

				export GLOO_DEVICE_TRANSPORT_TCP_TLS_PKEY=${TMP_CERT_DIR}/pkey.key

				export GLOO_DEVICE_TRANSPORT_TCP_TLS_CERT=${TMP_CERT_DIR}/cert.pem

				export GLOO_DEVICE_TRANSPORT_TCP_TLS_CA_FILE=${TMP_CERT_DIR}/ca.pem

				time python test/run_test.py --include distributed/test_c10d_gloo --verbose -- ProcessGroupGlooTest

				unset GLOO_DEVICE_TRANSPORT

				unset GLOO_DEVICE_TRANSPORT_TCP_TLS_PKEY

				unset GLOO_DEVICE_TRANSPORT_TCP_TLS_CERT

				unset GLOO_DEVICE_TRANSPORT_TCP_TLS_CA_FILE

									
										5

.ci/pytorch/run_tests.sh
									
												View File
												
				@ -74,12 +74,13 @@ else

				fi

				# Environment initialization

				retry pip install -qUr requirements-build.txt

				if [[ "$(uname)" == Darwin ]]; then

				    # Install the testing dependencies

				    retry pip install -q future hypothesis ${NUMPY_PACKAGE} ${PROTOBUF_PACKAGE} pytest setuptools six typing_extensions pyyaml

				    retry pip install -q future hypothesis ${NUMPY_PACKAGE} ${PROTOBUF_PACKAGE} pytest

				else

				    retry pip install -qr requirements.txt || true

				    retry pip install -q hypothesis protobuf pytest setuptools || true

				    retry pip install -q hypothesis protobuf pytest || true

				    numpy_ver=1.15

				    case "$(python --version 2>&1)" in

				      *2* | *3.5* | *3.6*)

									
										2

.ci/pytorch/smoke_test/check_binary_symbols.py
									
												View File
												
				@ -93,7 +93,7 @@ def check_lib_symbols_for_abi_correctness(lib: str) -> None:

				            f"Found pre-cxx11 symbols, but there shouldn't be any, see: {pre_cxx11_symbols[:100]}"

				        )

				    if num_cxx11_symbols < 100:

				        raise RuntimeError("Didn't find enought cxx11 symbols")

				        raise RuntimeError("Didn't find enough cxx11 symbols")

				def main() -> None:

									
										3

.ci/pytorch/smoke_test/check_gomp.py
									
												View File
												
				@ -46,6 +46,9 @@ def get_gomp_thread():

				    # use the default gomp path of AlmaLinux OS

				    libgomp_path = "/usr/lib64/libgomp.so.1"

				    # if it does not exist, try Ubuntu path

				    if not os.path.exists(libgomp_path):

				        libgomp_path = f"/usr/lib/{os.uname().machine}-linux-gnu/libgomp.so.1"

				    os.environ["GOMP_CPU_AFFINITY"] = "0-3"

									
										27

.ci/pytorch/smoke_test/smoke_test.py
									
												View File
												
				@ -276,7 +276,7 @@ def smoke_test_cuda(

				            torch_nccl_version = ".".join(str(v) for v in torch.cuda.nccl.version())

				            print(f"Torch nccl; version: {torch_nccl_version}")

				        # Pypi dependencies are installed on linux ony and nccl is availbale only on Linux.

				        # Pypi dependencies are installed on linux only and nccl is available only on Linux.

				        if pypi_pkg_check == "enabled" and sys.platform in ["linux", "linux2"]:

				            compare_pypi_to_torch_versions(

				                "cudnn", find_pypi_package_version("nvidia-cudnn"), torch_cudnn_version

				@ -385,6 +385,29 @@ def smoke_test_compile(device: str = "cpu") -> None:

				    x_pt2 = torch.compile(model, mode="max-autotune")(x)

				def smoke_test_nvshmem() -> None:

				    if not torch.cuda.is_available():

				        print("CUDA is not available, skipping NVSHMEM test")

				        return

				    # Check if NVSHMEM is compiled in current build

				    try:

				        from torch._C._distributed_c10d import _is_nvshmem_available

				    except ImportError:

				        # Not built with NVSHMEM support.

				        # torch is not compiled with NVSHMEM prior to 2.9

				        if torch.__version__ < "2.9":

				            return

				        else:

				            # After 2.9: NVSHMEM is expected to be compiled in current build

				            raise RuntimeError("torch not compiled with NVSHMEM") from None

				    print("torch compiled with NVSHMEM")

				    # Check if NVSHMEM is available on current system.

				    print(f"NVSHMEM available at run time: {_is_nvshmem_available()}")

				def smoke_test_modules():

				    cwd = os.getcwd()

				    for module in MODULES:

				@ -479,6 +502,8 @@ def main() -> None:

				        options.pypi_pkg_check,

				    )

				    smoke_test_nvshmem()

				if __name__ == "__main__":

				    main()

									
										265

.ci/pytorch/test.sh
									
												View File
												
				@ -11,6 +11,8 @@ export TERM=vt100

				# shellcheck source=./common.sh

				source "$(dirname "${BASH_SOURCE[0]}")/common.sh"

				# shellcheck source=./common-build.sh

				source "$(dirname "${BASH_SOURCE[0]}")/common-build.sh"

				# Do not change workspace permissions for ROCm and s390x CI jobs

				# as it can leave workspace with bad permissions for cancelled jobs

				@ -163,8 +165,6 @@ elif [[ "$BUILD_ENVIRONMENT" == *xpu* ]]; then

				  export PYTORCH_TESTING_DEVICE_ONLY_FOR="xpu"

				  # setting PYTHON_TEST_EXTRA_OPTION

				  export PYTHON_TEST_EXTRA_OPTION="--xpu"

				  # Disable sccache for xpu test due to flaky issue https://github.com/pytorch/pytorch/issues/143585

				  sudo rm -rf /opt/cache

				fi

				if [[ "$TEST_CONFIG" == *crossref* ]]; then

				@ -196,12 +196,12 @@ if [[ "$BUILD_ENVIRONMENT" == *xpu* ]]; then

				  # shellcheck disable=SC1091

				  source /opt/intel/oneapi/mpi/latest/env/vars.sh

				  # Check XPU status before testing

				  xpu-smi discovery

				  timeout 30 xpu-smi discovery || true

				fi

				if [[ "$BUILD_ENVIRONMENT" != *-bazel-* ]] ; then

				  # JIT C++ extensions require ninja.

				  pip_install --user "ninja==1.10.2"

				  pip_install "ninja==1.10.2"

				  # ninja is installed in $HOME/.local/bin, e.g., /var/lib/jenkins/.local/bin for CI user jenkins

				  # but this script should be runnable by any user, including root

				  export PATH="$HOME/.local/bin:$PATH"

				@ -212,8 +212,6 @@ if [[ "$BUILD_ENVIRONMENT" == *aarch64* ]]; then

				  export VALGRIND=OFF

				fi

				install_tlparse

				# DANGER WILL ROBINSON.  The LD_PRELOAD here could cause you problems

				# if you're not careful.  Check this if you made some changes and the

				# ASAN test is not working

				@ -226,7 +224,7 @@ if [[ "$BUILD_ENVIRONMENT" == *asan* ]]; then

				    export PYTORCH_TEST_WITH_ASAN=1

				    export PYTORCH_TEST_WITH_UBSAN=1

				    # TODO: Figure out how to avoid hard-coding these paths

				    export ASAN_SYMBOLIZER_PATH=/usr/lib/llvm-15/bin/llvm-symbolizer

				    export ASAN_SYMBOLIZER_PATH=/usr/lib/llvm-18/bin/llvm-symbolizer

				    export TORCH_USE_RTLD_GLOBAL=1

				    # NB: We load libtorch.so with RTLD_GLOBAL for UBSAN, unlike our

				    # default behavior.

				@ -291,6 +289,12 @@ elif [[ $TEST_CONFIG == 'nogpu_AVX512' ]]; then

				  export ATEN_CPU_CAPABILITY=avx2

				fi

				if [[ "${TEST_CONFIG}" == "legacy_nvidia_driver" ]]; then

				  # Make sure that CUDA can be initialized

				  (cd test && python -c "import torch; torch.rand(2, 2, device='cuda')")

				  export USE_LEGACY_DRIVER=1

				fi

				test_python_legacy_jit() {

				  time python test/run_test.py --include test_jit_legacy test_jit_fuser_legacy --verbose

				  assert_git_not_dirty

				@ -324,6 +328,29 @@ test_python_smoke() {

				  assert_git_not_dirty

				}

				test_h100_distributed() {

				  # Distributed tests at H100

				  time python test/run_test.py --include distributed/_composable/test_composability/test_pp_composability.py  $PYTHON_TEST_EXTRA_OPTION --upload-artifacts-while-running

				  # This test requires multicast support

				  time python test/run_test.py --include distributed/_composable/fsdp/test_fully_shard_comm.py -k TestFullyShardAllocFromPG $PYTHON_TEST_EXTRA_OPTION --upload-artifacts-while-running

				  assert_git_not_dirty

				}

				test_h100_symm_mem() {

				  # symmetric memory test

				  time python test/run_test.py --include distributed/test_symmetric_memory.py  $PYTHON_TEST_EXTRA_OPTION --upload-artifacts-while-running

				  time python test/run_test.py --include distributed/test_nvshmem.py $PYTHON_TEST_EXTRA_OPTION --upload-artifacts-while-running

				  time python test/run_test.py --include distributed/test_nvshmem_triton.py $PYTHON_TEST_EXTRA_OPTION --upload-artifacts-while-running

				  time python test/run_test.py --include distributed/test_nccl.py $PYTHON_TEST_EXTRA_OPTION --upload-artifacts-while-running

				  assert_git_not_dirty

				}

				test_h100_cutlass_backend() {

				  # cutlass backend tests for H100

				  TORCHINDUCTOR_CUTLASS_DIR=$(realpath "./third_party/cutlass") python test/run_test.py --include inductor/test_cutlass_backend -k "not addmm" $PYTHON_TEST_EXTRA_OPTION --upload-artifacts-while-running

				  TORCHINDUCTOR_CUTLASS_DIR=$(realpath "./third_party/cutlass") python test/run_test.py --include inductor/test_cutlass_evt $PYTHON_TEST_EXTRA_OPTION --upload-artifacts-while-running

				}

				test_lazy_tensor_meta_reference_disabled() {

				  export TORCH_DISABLE_FUNCTIONALIZATION_META_REFERENCE=1

				  echo "Testing lazy tensor operations without meta reference"

				@ -352,12 +379,24 @@ test_dynamo_wrapped_shard() {

				  assert_git_not_dirty

				}

				test_einops() {

				  pip install einops==0.6.1

				  time python test/run_test.py --einops --verbose --upload-artifacts-while-running

				  pip install einops==0.7.0

				  time python test/run_test.py --einops --verbose --upload-artifacts-while-running

				  pip install einops==0.8.1

				  time python test/run_test.py --einops --verbose --upload-artifacts-while-running

				  assert_git_not_dirty

				}

				test_inductor_distributed() {

				  # Smuggle a few multi-gpu tests here so that we don't have to request another large node

				  echo "Testing multi_gpu tests in test_torchinductor"

				  python test/run_test.py -i inductor/test_torchinductor.py -k test_multi_gpu --verbose

				  python test/run_test.py -i inductor/test_aot_inductor.py -k test_non_default_cuda_device --verbose

				  python test/run_test.py -i inductor/test_aot_inductor.py -k test_replicate_on_devices --verbose

				  python test/run_test.py -i inductor/test_aot_inductor.py -k test_on_gpu_device1 --verbose

				  python test/run_test.py -i inductor/test_aot_inductor.py -k test_non_default_gpu_device --verbose

				  python test/run_test.py -i inductor/test_aot_inductor.py -k test_load_package_multiple_gpus --verbose

				  python test/run_test.py -i distributed/test_c10d_functional_native.py --verbose

				  python test/run_test.py -i distributed/tensor/test_dtensor_compile.py --verbose

				  python test/run_test.py -i distributed/tensor/parallel/test_micro_pipeline_tp.py --verbose

				@ -409,14 +448,21 @@ test_inductor_aoti() {

				    python3 tools/amd_build/build_amd.py

				  fi

				  if [[ "$BUILD_ENVIRONMENT" == *sm86* ]]; then

				    BUILD_AOT_INDUCTOR_TEST=1 TORCH_CUDA_ARCH_LIST=8.6 USE_FLASH_ATTENTION=OFF python setup.py develop

				    BUILD_COMMAND=(TORCH_CUDA_ARCH_LIST=8.6 USE_FLASH_ATTENTION=OFF python -m pip install --no-build-isolation -v -e .)

				    # TODO: Replace me completely, as one should not use conda libstdc++, nor need special path to TORCH_LIB

				    LD_LIBRARY_PATH=/opt/conda/envs/py_3.10/lib/:${TORCH_LIB_DIR}:$LD_LIBRARY_PATH

				    CPP_TESTS_DIR="${BUILD_BIN_DIR}" python test/run_test.py --cpp --verbose -i cpp/test_aoti_abi_check cpp/test_aoti_inference -dist=loadfile

				    TEST_ENVS=(CPP_TESTS_DIR="${BUILD_BIN_DIR}" LD_LIBRARY_PATH="/opt/conda/envs/py_3.10/lib:${TORCH_LIB_DIR}:${LD_LIBRARY_PATH}")

				  else

				    BUILD_AOT_INDUCTOR_TEST=1 python setup.py develop

				    CPP_TESTS_DIR="${BUILD_BIN_DIR}" LD_LIBRARY_PATH="${TORCH_LIB_DIR}" python test/run_test.py --cpp --verbose -i cpp/test_aoti_abi_check cpp/test_aoti_inference -dist=loadfile

				    BUILD_COMMAND=(python -m pip install --no-build-isolation -v -e .)

				    TEST_ENVS=(CPP_TESTS_DIR="${BUILD_BIN_DIR}" LD_LIBRARY_PATH="${TORCH_LIB_DIR}")

				  fi

				  # aoti cmake custom command requires `torch` to be installed

				  # initialize the cmake build cache and install torch

				  /usr/bin/env "${BUILD_COMMAND[@]}"

				  # rebuild with the build cache with `BUILD_AOT_INDUCTOR_TEST` enabled

				  /usr/bin/env CMAKE_FRESH=1 BUILD_AOT_INDUCTOR_TEST=1 "${BUILD_COMMAND[@]}"

				  /usr/bin/env "${TEST_ENVS[@]}" python test/run_test.py --cpp --verbose -i cpp/test_aoti_abi_check cpp/test_aoti_inference cpp/test_vec_half_AVX2 -dist=loadfile

				}

				test_inductor_cpp_wrapper_shard() {

				@ -429,47 +475,26 @@ test_inductor_cpp_wrapper_shard() {

				  TEST_REPORTS_DIR=$(pwd)/test/test-reports

				  mkdir -p "$TEST_REPORTS_DIR"

				  if [[ "$1" -eq "2" ]]; then

				    # For now, manually put the opinfo tests in shard 2, and all other tests in

				    # shard 1.  Run all CPU tests, as well as specific GPU tests triggering past

				    # bugs, for now.

				    python test/run_test.py \

				      --include inductor/test_torchinductor_opinfo \

				      -k 'linalg or to_sparse or TestInductorOpInfoCPU' \

				      --verbose

				    exit

				  fi

				  # Run certain inductor unit tests with cpp wrapper. In the end state, we

				  # should be able to run all the inductor unit tests with cpp_wrapper.

				  #

				  # TODO: I'm pretty sure that "TestInductorOpInfoCPU" is not a valid filter,

				  # but change that in another PR to more accurately monitor the increased CI

				  # usage.

				  python test/run_test.py \

				    --include inductor/test_torchinductor_opinfo \

				    -k 'linalg or to_sparse or TestInductorOpInfoCPU' \

				    --shard "$1" "$NUM_TEST_SHARDS" \

				    --verbose

				  python test/run_test.py \

				    --include inductor/test_torchinductor inductor/test_max_autotune inductor/test_cpu_repro \

				    --shard "$1" "$NUM_TEST_SHARDS" \

				    --verbose

				  python test/run_test.py --inductor \

				    --include test_torch \

				    -k 'take' \

				    --shard "$1" "$NUM_TEST_SHARDS" \

				    --verbose

				  python test/run_test.py --inductor --include test_torch -k 'take' --verbose

				  # Run inductor benchmark tests with cpp wrapper.

				  # Skip benchmark tests if it's in rerun-disabled-mode.

				  if [[ "${PYTORCH_TEST_RERUN_DISABLED_TESTS}" == "1" ]]; then

				    echo "skip dynamo benchmark tests for rerun-disabled-test"

				  else

				    echo "run dynamo benchmark tests with cpp wrapper"

				    python benchmarks/dynamo/timm_models.py --device cuda --accuracy --amp \

				    --training --inductor --disable-cudagraphs --only vit_base_patch16_224 \

				    --output "$TEST_REPORTS_DIR/inductor_cpp_wrapper_training.csv"

				    python benchmarks/dynamo/check_accuracy.py \

				      --actual "$TEST_REPORTS_DIR/inductor_cpp_wrapper_training.csv" \

				      --expected "benchmarks/dynamo/ci_expected_accuracy/${MAYBE_ROCM}inductor_timm_training.csv"

				    python benchmarks/dynamo/torchbench.py --device cuda --accuracy \

				      --bfloat16 --inference --inductor --only hf_T5 --output "$TEST_REPORTS_DIR/inductor_cpp_wrapper_inference.csv"

				    python benchmarks/dynamo/torchbench.py --device cuda --accuracy \

				      --bfloat16 --inference --inductor --only llama --output "$TEST_REPORTS_DIR/inductor_cpp_wrapper_inference.csv"

				    python benchmarks/dynamo/torchbench.py --device cuda --accuracy \

				      --bfloat16 --inference --inductor --only moco --output "$TEST_REPORTS_DIR/inductor_cpp_wrapper_inference.csv"

				    python benchmarks/dynamo/check_accuracy.py \

				      --actual "$TEST_REPORTS_DIR/inductor_cpp_wrapper_inference.csv" \

				      --expected "benchmarks/dynamo/ci_expected_accuracy/${MAYBE_ROCM}inductor_torchbench_inference.csv"

				  fi

				}

				# "Global" flags for inductor benchmarking controlled by TEST_CONFIG

				@ -482,7 +507,7 @@ DYNAMO_BENCHMARK_FLAGS=()

				pr_time_benchmarks() {

				  pip_install --user "fbscribelogger"

				  pip_install "fbscribelogger"

				  TEST_REPORTS_DIR=$(pwd)/test/test-reports

				  mkdir -p "$TEST_REPORTS_DIR"

				@ -590,7 +615,9 @@ test_perf_for_dashboard() {

				  local device=cuda

				  if [[ "${TEST_CONFIG}" == *cpu* ]]; then

				    if [[ "${TEST_CONFIG}" == *cpu_x86* ]]; then

				    if [[ "${TEST_CONFIG}" == *cpu_x86_zen* ]]; then

				      device=cpu_x86_zen

				    elif [[ "${TEST_CONFIG}" == *cpu_x86* ]]; then

				      device=cpu_x86

				    elif [[ "${TEST_CONFIG}" == *cpu_aarch64* ]]; then

				      device=cpu_aarch64

				@ -600,13 +627,19 @@ test_perf_for_dashboard() {

				    device=cuda_a10g

				  elif [[ "${TEST_CONFIG}" == *h100* ]]; then

				    device=cuda_h100

				  elif [[ "${TEST_CONFIG}" == *b200* ]]; then

				    device=cuda_b200

				  elif [[ "${TEST_CONFIG}" == *rocm* ]]; then

				    device=rocm

				  fi

				  for mode in "${modes[@]}"; do

				    if [[ "$mode" == "inference" ]]; then

				      dtype=bfloat16

				      if [[ "$device" == "cpu_x86" ]]; then

				        dtype=amp

				      else

				        dtype=bfloat16

				      fi

				    elif [[ "$mode" == "training" ]]; then

				      dtype=amp

				    fi

				@ -618,6 +651,10 @@ test_perf_for_dashboard() {

				        target_flag+=( --no-translation-validation)

				      fi

				      if [[ "$DASHBOARD_TAG" == *freezing-true* ]]; then

				        target_flag+=( --freezing)

				      fi

				      if [[ "$DASHBOARD_TAG" == *default-true* ]]; then

				        $TASKSET python "benchmarks/dynamo/$suite.py" \

				            "${target_flag[@]}" --"$mode" --"$dtype" --backend "$backend" --disable-cudagraphs "$@" \

				@ -766,6 +803,16 @@ test_dynamo_benchmark() {

				  if [[ "${TEST_CONFIG}" == *perf_compare* ]]; then

				    test_single_dynamo_benchmark "training" "$suite" "$shard_id" --training --amp "$@"

				  elif [[ "${TEST_CONFIG}" == *perf* ]]; then

				    # TODO (huydhn): Just smoke test some sample models

				    if [[ "${TEST_CONFIG}" == *b200* ]]; then

				      if [[ "${suite}" == "huggingface" ]]; then

				        export TORCHBENCH_ONLY_MODELS="DistillGPT2"

				      elif [[ "${suite}" == "timm_models" ]]; then

				        export TORCHBENCH_ONLY_MODELS="inception_v3"

				      elif [[ "${suite}" == "torchbench" ]]; then

				        export TORCHBENCH_ONLY_MODELS="hf_Bert"

				      fi

				    fi

				    test_single_dynamo_benchmark "dashboard" "$suite" "$shard_id" "$@"

				  else

				    if [[ "${TEST_CONFIG}" == *cpu* ]]; then

				@ -820,16 +867,7 @@ test_inductor_torchbench_smoketest_perf() {

				  done

				}

				test_inductor_get_core_number() {

				  if [[ "${TEST_CONFIG}" == *aarch64* ]]; then

				    echo "$(($(lscpu | grep 'Cluster(s):' | awk '{print $2}') * $(lscpu | grep 'Core(s) per cluster:' | awk '{print $4}')))"

				  else

				    echo "$(($(lscpu | grep 'Socket(s):' | awk '{print $2}') * $(lscpu | grep 'Core(s) per socket:' | awk '{print $4}')))"

				  fi

				}

				test_inductor_set_cpu_affinity(){

				  #set jemalloc

				  JEMALLOC_LIB="$(find /usr/lib -name libjemalloc.so.2)"

				  export LD_PRELOAD="$JEMALLOC_LIB":"$LD_PRELOAD"

				  export MALLOC_CONF="oversize_threshold:1,background_thread:true,metadata_thp:auto,dirty_decay_ms:-1,muzzy_decay_ms:-1"

				@ -841,14 +879,23 @@ test_inductor_set_cpu_affinity(){

				    export KMP_AFFINITY=granularity=fine,compact,1,0

				    export KMP_BLOCKTIME=1

				  fi

				  cores=$(test_inductor_get_core_number)

				  # Set number of cores to 16 on Aarch64 for performance runs.

				  # Use nproc here instead of lscpu because it takes into account cgroups slice

				  cpus=$(nproc)

				  thread_per_core=$(lscpu | grep 'Thread(s) per core:' | awk '{print $4}')

				  cores=$((cpus / thread_per_core))

				  # Set number of cores to 16 on aarch64 for performance runs

				  if [[ "${TEST_CONFIG}" == *aarch64* && $cores -gt 16 ]]; then

				    cores=16

				  fi

				  export OMP_NUM_THREADS=$cores

				  end_core=$((cores-1))

				  export TASKSET="taskset -c 0-$end_core"

				  # Handle cgroups slice start and end CPU

				  start_cpu=$(python -c 'import os; print(min(os.sched_getaffinity(0)))')

				  # Leaving one physical CPU for other tasks

				  end_cpu=$(($(python -c 'import os; print(max(os.sched_getaffinity(0)))') - thread_per_core))

				  export TASKSET="taskset -c $start_cpu-$end_cpu"

				}

				test_inductor_torchbench_cpu_smoketest_perf(){

				@ -893,12 +940,6 @@ test_torchbench_gcp_smoketest(){

				  popd

				}

				test_python_gloo_with_tls() {

				  source "$(dirname "${BASH_SOURCE[0]}")/run_glootls_test.sh"

				  assert_git_not_dirty

				}

				test_aten() {

				  # Test ATen

				  # The following test(s) of ATen have already been skipped by caffe2 in rocm environment:

				@ -945,6 +986,8 @@ test_without_numpy() {

				  if [[ "${TEST_CONFIG}" == *dynamo_wrapped* ]]; then

				    python -c "import sys;sys.path.insert(0, 'fake_numpy');import torch;torch.compile(lambda x:print(x))('Hello World')"

				  fi

				  # Regression test for https://github.com/pytorch/pytorch/pull/157734 (torch.onnx should be importable without numpy)

				  python -c "import sys;sys.path.insert(0, 'fake_numpy');import torch; import torch.onnx"

				  popd

				}

				@ -1131,6 +1174,12 @@ test_custom_backend() {

				test_custom_script_ops() {

				  echo "Testing custom script operators"

				  if [[ "$BUILD_ENVIRONMENT" == *s390x* ]]; then

				    echo "Skipping custom script operators until it's fixed"

				    return 0

				  fi

				  CUSTOM_OP_BUILD="${CUSTOM_TEST_ARTIFACT_BUILD_DIR}/custom-op-build"

				  pushd test/custom_operator

				  cp -a "$CUSTOM_OP_BUILD" build

				@ -1283,10 +1332,13 @@ EOF

				  # Step 2. Make sure that the public API test "test_correct_module_names" fails when an existing

				  # file is modified to introduce an invalid public API function.

				  EXISTING_FILEPATH="${TORCH_INSTALL_DIR}/nn/parameter.py"

				  # The filepath here must not have __all__ defined in it, otherwise the test will pass.

				  # If your PR introduces __all__ to torch/cuda/streams.py please point this to another file

				  # that does not have __all__ defined.

				  EXISTING_FILEPATH="${TORCH_INSTALL_DIR}/cuda/streams.py"

				  cp -v "${EXISTING_FILEPATH}" "${EXISTING_FILEPATH}.orig"

				  echo "${BAD_PUBLIC_FUNC}" >> "${EXISTING_FILEPATH}"

				  invalid_api="torch.nn.parameter.new_public_func"

				  invalid_api="torch.cuda.streams.new_public_func"

				  echo "Appended an invalid public API function to existing file ${EXISTING_FILEPATH}..."

				  check_public_api_test_fails \

				@ -1441,8 +1493,8 @@ test_bazel() {

				test_benchmarks() {

				  if [[ "$BUILD_ENVIRONMENT" == *cuda* && $TEST_CONFIG != *nogpu* ]]; then

				    pip_install --user "pytest-benchmark==3.2.3"

				    pip_install --user "requests"

				    pip_install "pytest-benchmark==3.2.3"

				    pip_install "requests"

				    BENCHMARK_DATA="benchmarks/.data"

				    mkdir -p ${BENCHMARK_DATA}

				    pytest benchmarks/fastrnns/test_bench.py --benchmark-sort=Name --benchmark-json=${BENCHMARK_DATA}/fastrnns_default.json --fuser=default --executor=default

				@ -1550,11 +1602,12 @@ test_operator_benchmark() {

				  test_inductor_set_cpu_affinity

				  cd benchmarks/operator_benchmark/pt_extension

				  python setup.py install

				  python -m pip install .

				  cd "${TEST_DIR}"/benchmarks/operator_benchmark

				  $TASKSET python -m benchmark_all_test --device "$1" --tag-filter "$2" \

				      --output-dir "${TEST_REPORTS_DIR}/operator_benchmark_eager_float32_cpu.csv"

				      --output-csv "${TEST_REPORTS_DIR}/operator_benchmark_eager_float32_cpu.csv" \

				      --output-json-for-dashboard "${TEST_REPORTS_DIR}/operator_benchmark_eager_float32_cpu.json" \

				  pip_install pandas

				  python check_perf_csv.py \

				@ -1569,7 +1622,13 @@ if ! [[ "${BUILD_ENVIRONMENT}" == *libtorch* || "${BUILD_ENVIRONMENT}" == *-baze

				fi

				if [[ "${TEST_CONFIG}" == *numpy_2* ]]; then

				  # Install numpy-2.0.2 and compatible scipy & numba versions

				  python -mpip install --pre numpy==2.0.2 scipy==1.13.1 numba==0.60.0

				  # Force re-install of pandas to avoid error where pandas checks numpy version from initial install and fails upon import

				  TMP_PANDAS_VERSION=$(python -c "import pandas; print(pandas.__version__)" 2>/dev/null)

				  if [ -n "$TMP_PANDAS_VERSION" ]; then

				    python -m pip install --pre numpy==2.0.2 scipy==1.13.1 numba==0.60.0 pandas=="$TMP_PANDAS_VERSION" --force-reinstall

				  else

				    python -m pip install --pre numpy==2.0.2 scipy==1.13.1 numba==0.60.0

				  fi

				  python test/run_test.py --include dynamo/test_functions.py dynamo/test_unspec.py test_binary_ufuncs.py test_fake_tensor.py test_linalg.py test_numpy_interop.py test_tensor_creation_ops.py test_torch.py torch_np/test_basic.py

				elif [[ "${BUILD_ENVIRONMENT}" == *aarch64* && "${TEST_CONFIG}" != *perf_cpu_aarch64* ]]; then

				  test_linux_aarch64

				@ -1623,52 +1682,40 @@ elif [[ "${TEST_CONFIG}" == *timm* ]]; then

				  id=$((SHARD_NUMBER-1))

				  test_dynamo_benchmark timm_models "$id"

				elif [[ "${TEST_CONFIG}" == cachebench ]]; then

				  install_torchaudio cuda

				  install_torchaudio

				  install_torchvision

				  checkout_install_torchbench nanogpt BERT_pytorch resnet50 hf_T5 llama moco

				  PYTHONPATH=$(pwd)/torchbench test_cachebench

				  PYTHONPATH=/torchbench test_cachebench

				elif [[ "${TEST_CONFIG}" == verify_cachebench ]]; then

				  install_torchaudio cpu

				  install_torchaudio

				  install_torchvision

				  checkout_install_torchbench nanogpt

				  PYTHONPATH=$(pwd)/torchbench test_verify_cachebench

				  PYTHONPATH=/torchbench test_verify_cachebench

				elif [[ "${TEST_CONFIG}" == *torchbench* ]]; then

				  if [[ "${TEST_CONFIG}" == *cpu* ]]; then

				    install_torchaudio cpu

				  else

				    install_torchaudio cuda

				  fi

				  install_torchaudio

				  install_torchvision

				  TORCH_CUDA_ARCH_LIST="8.0;8.6" pip_install git+https://github.com/pytorch/ao.git

				  install_torchao

				  id=$((SHARD_NUMBER-1))

				  # https://github.com/opencv/opencv-python/issues/885

				  pip_install opencv-python==4.8.0.74

				  if [[ "${TEST_CONFIG}" == *inductor_torchbench_smoketest_perf* ]]; then

				    checkout_install_torchbench hf_Bert hf_Albert timm_vision_transformer

				    PYTHONPATH=$(pwd)/torchbench test_inductor_torchbench_smoketest_perf

				    PYTHONPATH=/torchbench test_inductor_torchbench_smoketest_perf

				  elif [[ "${TEST_CONFIG}" == *inductor_torchbench_cpu_smoketest_perf* ]]; then

				    checkout_install_torchbench timm_vision_transformer phlippe_densenet basic_gnn_edgecnn \

				      llama_v2_7b_16h resnet50 timm_efficientnet mobilenet_v3_large timm_resnest \

				      functorch_maml_omniglot yolov3 mobilenet_v2 resnext50_32x4d densenet121 mnasnet1_0

				    PYTHONPATH=$(pwd)/torchbench test_inductor_torchbench_cpu_smoketest_perf

				    PYTHONPATH=/torchbench test_inductor_torchbench_cpu_smoketest_perf

				  elif [[ "${TEST_CONFIG}" == *torchbench_gcp_smoketest* ]]; then

				    checkout_install_torchbench

				    TORCHBENCHPATH=$(pwd)/torchbench test_torchbench_gcp_smoketest

				    TORCHBENCHPATH=/torchbench test_torchbench_gcp_smoketest

				  else

				    checkout_install_torchbench

				    # Do this after checkout_install_torchbench to ensure we clobber any

				    # nightlies that torchbench may pull in

				    if [[ "${TEST_CONFIG}" != *cpu* ]]; then

				      install_torchrec_and_fbgemm

				    fi

				    PYTHONPATH=$(pwd)/torchbench test_dynamo_benchmark torchbench "$id"

				    PYTHONPATH=/torchbench test_dynamo_benchmark torchbench "$id"

				  fi

				elif [[ "${TEST_CONFIG}" == *inductor_cpp_wrapper* ]]; then

				  install_torchaudio cuda

				  install_torchvision

				  checkout_install_torchbench hf_T5 llama moco

				  PYTHONPATH=$(pwd)/torchbench test_inductor_cpp_wrapper_shard "$SHARD_NUMBER"

				  test_inductor_aoti

				  PYTHONPATH=/torchbench test_inductor_cpp_wrapper_shard "$SHARD_NUMBER"

				  if [[ "$SHARD_NUMBER" -eq "1" ]]; then

				    test_inductor_aoti

				  fi

				elif [[ "${TEST_CONFIG}" == *inductor* ]]; then

				  install_torchvision

				  test_inductor_shard "${SHARD_NUMBER}"

				@ -1677,6 +1724,8 @@ elif [[ "${TEST_CONFIG}" == *inductor* ]]; then

				      test_inductor_distributed

				    fi

				  fi

				elif [[ "${TEST_CONFIG}" == *einops* ]]; then

				  test_einops

				elif [[ "${TEST_CONFIG}" == *dynamo_wrapped* ]]; then

				  install_torchvision

				  test_dynamo_wrapped_shard "${SHARD_NUMBER}"

				@ -1724,6 +1773,12 @@ elif [[ "${BUILD_ENVIRONMENT}" == *xpu* ]]; then

				  test_xpu_bin

				elif [[ "${TEST_CONFIG}" == smoke ]]; then

				  test_python_smoke

				elif [[ "${TEST_CONFIG}" == h100_distributed ]]; then

				  test_h100_distributed

				elif [[ "${TEST_CONFIG}" == "h100-symm-mem" ]]; then

				  test_h100_symm_mem

				elif [[ "${TEST_CONFIG}" == h100_cutlass_backend ]]; then

				  test_h100_cutlass_backend

				else

				  install_torchvision

				  install_monkeytype

									
										34

.ci/pytorch/win-arm64-build.ps1
									
										Normal file
									
												View File
												
				@ -0,0 +1,34 @@

				# If you want to rebuild, run this with $env:REBUILD=1

				# If you want to build with CUDA, run this with $env:USE_CUDA=1

				# If you want to build without CUDA, run this with $env:USE_CUDA=0

				# Check for setup.py in the current directory

				if (-not (Test-Path "setup.py")) {

				    Write-Host "ERROR: Please run this build script from PyTorch root directory."

				    exit 1

				}

				# Get the script's parent directory

				$ScriptParentDir = Split-Path -Parent $MyInvocation.MyCommand.Definition

				# Set TMP_DIR and convert to Windows path

				$env:TMP_DIR = Join-Path (Get-Location) "build\win_tmp"

				$env:TMP_DIR_WIN = $env:TMP_DIR  # Already in Windows format, no cygpath needed

				# Set final package directory with default fallback

				if (-not $env:PYTORCH_FINAL_PACKAGE_DIR) {

				    $env:PYTORCH_FINAL_PACKAGE_DIR = "C:\w\build-results"

				}

				# Create the final package directory if it doesn't exist

				if (-not (Test-Path $env:PYTORCH_FINAL_PACKAGE_DIR)) {

				    New-Item -Path $env:PYTORCH_FINAL_PACKAGE_DIR -ItemType Directory -Force | Out-Null

				}

				# Set script helpers directory

				$env:SCRIPT_HELPERS_DIR = Join-Path $ScriptParentDir "win-test-helpers\arm64"

				# Run the main build script

				& "$env:SCRIPT_HELPERS_DIR\build_pytorch.ps1"

				Write-Host "BUILD PASSED"

									
										24

.ci/pytorch/win-arm64-test.sh
									
										Normal file
									
												View File
												
				@ -0,0 +1,24 @@

				#!/bin/bash

				set -ex -o pipefail

				SCRIPT_PARENT_DIR=$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )

				# shellcheck source=./common.sh

				source "$SCRIPT_PARENT_DIR/common.sh"

				run_tests() {

				    echo Running smoke_test.py...

				    python ./.ci/pytorch/smoke_test/smoke_test.py --package torchonly

				    echo Running test_autograd.oy, test_nn.py, test_torch.py...

				    cd test

				    CORE_TEST_LIST=("test_autograd.py" "test_nn.py" "test_modules.py")

				    for t in "${CORE_TEST_LIST[@]}"; do

				        echo "Running test: $t"

				        python "$t" --verbose --save-xml --use-pytest -vvvv -rfEsxXP -p no:xdist

				    done

				}

				run_tests

				echo "TEST PASSED"

									
										2

.ci/pytorch/win-build.sh
									
												View File
												
				@ -31,7 +31,7 @@ PYLONG_API_CHECK=$?

				if [[ $PYLONG_API_CHECK == 0 ]]; then

				  echo "Usage of PyLong_{From,As}{Unsigned}Long API may lead to overflow errors on Windows"

				  echo "because \`sizeof(long) == 4\` and \`sizeof(unsigned long) == 4\`."

				  echo "Please include \"torch/csrc/utils/python_numbers.h\" and use the correspoding APIs instead."

				  echo "Please include \"torch/csrc/utils/python_numbers.h\" and use the corresponding APIs instead."

				  echo "PyLong_FromLong -> THPUtils_packInt32 / THPUtils_packInt64"

				  echo "PyLong_AsLong -> THPUtils_unpackInt (32-bit) / THPUtils_unpackLong (64-bit)"

				  echo "PyLong_FromUnsignedLong -> THPUtils_packUInt32 / THPUtils_packUInt64"

									
										98

.ci/pytorch/win-test-helpers/arm64/build_pytorch.ps1
									
										Normal file
									
												View File
												
				@ -0,0 +1,98 @@

				# TODO: we may can use existing build_pytorch.bat for arm64

				if ($env:DEBUG -eq "1") {

				    $env:BUILD_TYPE = "debug"

				} else {

				    $env:BUILD_TYPE = "release"

				}

				# This inflates our log size slightly, but it is REALLY useful to be

				# able to see what our cl.exe commands are. (since you can actually

				# just copy-paste them into a local Windows setup to just rebuild a

				# single file.)

				# log sizes are too long, but leaving this here in case someone wants to use it locally

				# $env:CMAKE_VERBOSE_MAKEFILE = "1"

				$env:INSTALLER_DIR = Join-Path $env:SCRIPT_HELPERS_DIR "installation-helpers"

				cd ..

				# Environment variables

				$env:SCCACHE_IDLE_TIMEOUT = "0"

				$env:SCCACHE_IGNORE_SERVER_IO_ERROR = "1"

				$env:CMAKE_BUILD_TYPE = $env:BUILD_TYPE

				$env:CMAKE_C_COMPILER_LAUNCHER = "sccache"

				$env:CMAKE_CXX_COMPILER_LAUNCHER = "sccache"

				$env:libuv_ROOT = Join-Path $env:DEPENDENCIES_DIR "libuv\install"

				$env:MSSdk = "1"

				if ($env:PYTORCH_BUILD_VERSION) {

				    $env:PYTORCH_BUILD_VERSION = $env:PYTORCH_BUILD_VERSION

				    $env:PYTORCH_BUILD_NUMBER = "1"

				}

				$env:CMAKE_POLICY_VERSION_MINIMUM = "3.5"

				# Set BLAS type

				if ($env:ENABLE_APL -eq "1") {

				    $env:BLAS = "APL"

				    $env:USE_LAPACK = "1"

				} elseif ($env:ENABLE_OPENBLAS -eq "1") {

				    $env:BLAS = "OpenBLAS"

				    $env:OpenBLAS_HOME = Join-Path $env:DEPENDENCIES_DIR "OpenBLAS\install"

				}

				# Change to source directory

				Set-Location $env:PYTORCH_ROOT

				# Copy libuv.dll

				Copy-Item -Path (Join-Path $env:libuv_ROOT "lib\Release\uv.dll") -Destination "torch\lib\uv.dll" -Force

				# Create virtual environment

				python -m venv .venv

				.\.venv\Scripts\Activate.ps1

				where.exe python

				# Python install dependencies

				python -m pip install --upgrade pip

				pip install setuptools pyyaml

				pip install -r requirements.txt

				# Set after installing psutil

				$env:DISTUTILS_USE_SDK = "1"

				# Print all environment variables

				Get-ChildItem Env:

				# Start and inspect sccache

				sccache --start-server

				sccache --zero-stats

				sccache --show-stats

				# Build the wheel

				python setup.py bdist_wheel

				if ($LASTEXITCODE -ne 0) { exit 1 }

				# Install the wheel locally

				$whl = Get-ChildItem -Path "dist\*.whl" | Select-Object -First 1

				if ($whl) {

				    python -mpip install --no-index --no-deps $whl.FullName

				}

				# Copy final wheel

				robocopy "dist" "$env:PYTORCH_FINAL_PACKAGE_DIR" *.whl

				# Export test times

				python tools/stats/export_test_times.py

				# Copy additional CI files

				robocopy ".additional_ci_files" "$env:PYTORCH_FINAL_PACKAGE_DIR\.additional_ci_files" /E

				# Save ninja log

				Copy-Item -Path "build\.ninja_log" -Destination $env:PYTORCH_FINAL_PACKAGE_DIR -Force

				# Final sccache stats and stop

				sccache --show-stats

				sccache --stop-server

				exit 0

									
										4

.ci/pytorch/win-test-helpers/build_pytorch.bat
									
												View File
												
				@ -10,7 +10,7 @@ set PATH=C:\Program Files\CMake\bin;C:\Program Files\7-Zip;C:\ProgramData\chocol

				:: able to see what our cl.exe commands are (since you can actually

				:: just copy-paste them into a local Windows setup to just rebuild a

				:: single file.)

				:: log sizes are too long, but leaving this here incase someone wants to use it locally

				:: log sizes are too long, but leaving this here in case someone wants to use it locally

				:: set CMAKE_VERBOSE_MAKEFILE=1

				@ -42,7 +42,7 @@ call choco upgrade -y cmake --no-progress --installargs 'ADD_CMAKE_TO_PATH=Syste

				if errorlevel 1 goto fail

				if not errorlevel 0 goto fail

				call pip install mkl-include==2021.4.0 mkl-devel==2021.4.0

				call pip install mkl==2024.2.0 mkl-static==2024.2.0 mkl-include==2024.2.0

				if errorlevel 1 goto fail

				if not errorlevel 0 goto fail

									
										2

.ci/pytorch/win-test-helpers/run_python_nn_smoketests.py
									
												View File
												
				@ -52,7 +52,7 @@ if __name__ == "__main__":

				            if os.path.exists(debugger):

				                command_args = [debugger, "-o", "-c", "~*g; q"] + command_args

				                command_string = " ".join(command_args)

				                print("Reruning with traceback enabled")

				                print("Rerunning with traceback enabled")

				                print("Command:", command_string)

				                subprocess.run(command_args, check=False)

				            sys.exit(e.returncode)

									
										7

.ci/pytorch/win-test.sh
									
												View File
												
				@ -38,10 +38,10 @@ if [[ "$BUILD_ENVIRONMENT" == *cuda* ]]; then

				fi

				# TODO: Move both of them to Windows AMI

				python -m pip install pytest-rerunfailures==10.3 pytest-cpp==2.3.0 tensorboard==2.13.0 pytest-subtests==0.13.1

				python -m pip install pytest-rerunfailures==10.3 pytest-cpp==2.3.0 tensorboard==2.13.0 protobuf==5.29.4 pytest-subtests==0.13.1

				# Install Z3 optional dependency for Windows builds.

				python -m pip install z3-solver==4.12.2.0

				python -m pip install z3-solver==4.15.1.0

				# Install tlparse for test\dynamo\test_structured_trace.py UTs.

				python -m pip install tlparse==0.3.30

				@ -52,6 +52,9 @@ python -m pip install parameterized==0.8.1

				# Install pulp for testing ilps under torch\distributed\_tools

				python -m pip install pulp==2.9.0

				# Install expecttest to merge https://github.com/pytorch/pytorch/pull/155308

				python -m pip install expecttest==0.3.0

				run_tests() {

				    # Run nvidia-smi if available

				    for path in '/c/Program Files/NVIDIA Corporation/NVSMI/nvidia-smi.exe' /c/Windows/System32/nvidia-smi.exe; do

									
										2

.ci/pytorch/windows/arm64/bootstrap_libuv.bat
									
												View File
												
				@ -7,7 +7,7 @@ if not exist "%DOWNLOADS_DIR%" mkdir %DOWNLOADS_DIR%

				if not exist "%DEPENDENCIES_DIR%" mkdir %DEPENDENCIES_DIR%

				:: activate visual studio

				call "%DEPENDENCIES_DIR%\VSBuildTools\VC\Auxiliary\Build\vcvarsall.bat" arm64

				call "C:\Program Files\Microsoft Visual Studio\2022\Enterprise\VC\Auxiliary\Build\vcvarsall.bat" arm64

				where cl.exe

				cd %DEPENDENCIES_DIR%

									
										2

.ci/pytorch/windows/arm64/bootstrap_openblas.bat
									
												View File
												
				@ -7,7 +7,7 @@ if not exist "%DOWNLOADS_DIR%" mkdir %DOWNLOADS_DIR%

				if not exist "%DEPENDENCIES_DIR%" mkdir %DEPENDENCIES_DIR%

				:: activate visual studio

				call "%DEPENDENCIES_DIR%\VSBuildTools\VC\Auxiliary\Build\vcvarsall.bat" arm64

				call "C:\Program Files\Microsoft Visual Studio\2022\Enterprise\VC\Auxiliary\Build\vcvarsall.bat" arm64

				where cl.exe

				:: Clone OpenBLAS

									
										2

.ci/pytorch/windows/arm64/bootstrap_tests.bat
									
												View File
												
				@ -2,7 +2,7 @@

				cd %PYTORCH_ROOT%

				:: activate visual studio

				call "%DEPENDENCIES_DIR%\VSBuildTools\VC\Auxiliary\Build\vcvarsall.bat" arm64

				call "C:\Program Files\Microsoft Visual Studio\2022\Enterprise\VC\Auxiliary\Build\vcvarsall.bat" arm64

				where cl.exe

				:: create virtual environment

									
										2

.ci/pytorch/windows/arm64/build_libtorch.bat
									
												View File
												
				@ -21,7 +21,7 @@ if %ENABLE_APL% == 1 (

				)

				:: activate visual studio

				call "%DEPENDENCIES_DIR%\VSBuildTools\VC\Auxiliary\Build\vcvarsall.bat" arm64

				call "C:\Program Files\Microsoft Visual Studio\2022\Enterprise\VC\Auxiliary\Build\vcvarsall.bat" arm64

				where cl.exe

				:: change to source directory

									
										2

.ci/pytorch/windows/arm64/build_pytorch.bat
									
												View File
												
				@ -21,7 +21,7 @@ if %ENABLE_APL% == 1 (

				)

				:: activate visual studio

				call "%DEPENDENCIES_DIR%\VSBuildTools\VC\Auxiliary\Build\vcvarsall.bat" arm64

				call "C:\Program Files\Microsoft Visual Studio\2022\Enterprise\VC\Auxiliary\Build\vcvarsall.bat" arm64

				where cl.exe

				:: change to source directory

									
										2

.ci/pytorch/windows/arm64/smoke_test.bat
									
												View File
												
				@ -33,7 +33,7 @@ pushd tmp

				set VC_VERSION_LOWER=14

				set VC_VERSION_UPPER=36

				call "%DEPENDENCIES_DIR%\VSBuildTools\VC\Auxiliary\Build\vcvarsall.bat" arm64

				call "C:\Program Files\Microsoft Visual Studio\2022\Enterprise\VC\Auxiliary\Build\vcvarsall.bat" arm64

				set install_root=%CD%

				set INCLUDE=%INCLUDE%;%install_root%\include;%install_root%\include\torch\csrc\api\include

									
										59

.ci/pytorch/windows/cuda118.bat
									
												View File
											
				@ -1,59 +0,0 @@

				@echo off

				set MODULE_NAME=pytorch

				IF NOT EXIST "setup.py" IF NOT EXIST "%MODULE_NAME%" (

				    call internal\clone.bat

				    cd %~dp0

				) ELSE (

				    call internal\clean.bat

				)

				IF ERRORLEVEL 1 goto :eof

				call internal\check_deps.bat

				IF ERRORLEVEL 1 goto :eof

				REM Check for optional components

				set USE_CUDA=

				set CMAKE_GENERATOR=Visual Studio 15 2017 Win64

				IF "%NVTOOLSEXT_PATH%"=="" (

				    IF EXIST "C:\Program Files\NVIDIA Corporation\NvToolsExt\lib\x64\nvToolsExt64_1.lib"  (

				        set NVTOOLSEXT_PATH=C:\Program Files\NVIDIA Corporation\NvToolsExt

				    ) ELSE (

				        echo NVTX ^(Visual Studio Extension ^for CUDA^) ^not installed, failing

				        exit /b 1

				    )

				)

				IF "%CUDA_PATH_V118%"=="" (

				    IF EXIST "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.8\bin\nvcc.exe" (

				        set "CUDA_PATH_V118=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.8"

				    ) ELSE (

				        echo CUDA 11.8 not found, failing

				        exit /b 1

				    )

				)

				IF "%BUILD_VISION%" == "" (

				    set TORCH_CUDA_ARCH_LIST=3.7+PTX;5.0;6.0;6.1;7.0;7.5;8.0;8.6;9.0

				    set TORCH_NVCC_FLAGS=-Xfatbin -compress-all

				) ELSE (

				    set NVCC_FLAGS=-D__CUDA_NO_HALF_OPERATORS__ --expt-relaxed-constexpr -gencode=arch=compute_35,code=sm_35 -gencode=arch=compute_50,code=sm_50 -gencode=arch=compute_60,code=sm_60 -gencode=arch=compute_70,code=sm_70 -gencode=arch=compute_75,code=sm_75 -gencode=arch=compute_80,code=compute_80 -gencode=arch=compute_86,code=compute_86 -gencode=arch=compute_90,code=compute_90

				)

				set "CUDA_PATH=%CUDA_PATH_V118%"

				set "PATH=%CUDA_PATH_V118%\bin;%PATH%"

				:optcheck

				call internal\check_opts.bat

				IF ERRORLEVEL 1 goto :eof

				if exist "%NIGHTLIES_PYTORCH_ROOT%" cd %NIGHTLIES_PYTORCH_ROOT%\..

				call  %~dp0\internal\copy.bat

				IF ERRORLEVEL 1 goto :eof

				call  %~dp0\internal\setup.bat

				IF ERRORLEVEL 1 goto :eof

									
										16

.ci/pytorch/windows/cuda124.bat → .ci/pytorch/windows/cuda129.bat
									
												View File
												
				@ -27,24 +27,24 @@ IF "%NVTOOLSEXT_PATH%"=="" (

				    )

				)

				IF "%CUDA_PATH_V124%"=="" (

				    IF EXIST "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.4\bin\nvcc.exe" (

				        set "CUDA_PATH_V124=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.4"

				IF "%CUDA_PATH_V129%"=="" (

				    IF EXIST "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9\bin\nvcc.exe" (

				        set "CUDA_PATH_V129=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.9"

				    ) ELSE (

				        echo CUDA 12.4 not found, failing

				        echo CUDA 12.9 not found, failing

				        exit /b 1

				    )

				)

				IF "%BUILD_VISION%" == "" (

				    set TORCH_CUDA_ARCH_LIST=6.1;7.0;7.5;8.0;8.6;9.0

				    set TORCH_CUDA_ARCH_LIST=7.0;7.5;8.0;8.6;9.0;10.0;12.0

				    set TORCH_NVCC_FLAGS=-Xfatbin -compress-all

				) ELSE (

				    set NVCC_FLAGS=-D__CUDA_NO_HALF_OPERATORS__ --expt-relaxed-constexpr -gencode=arch=compute_50,code=sm_50 -gencode=arch=compute_60,code=sm_60 -gencode=arch=compute_70,code=sm_70 -gencode=arch=compute_75,code=sm_75 -gencode=arch=compute_80,code=compute_80 -gencode=arch=compute_86,code=compute_86 -gencode=arch=compute_90,code=compute_90

				    set NVCC_FLAGS=-D__CUDA_NO_HALF_OPERATORS__ --expt-relaxed-constexpr -gencode=arch=compute_70,code=sm_70 -gencode=arch=compute_75,code=sm_75 -gencode=arch=compute_80,code=compute_80 -gencode=arch=compute_86,code=compute_86 -gencode=arch=compute_90,code=compute_90 -gencode=arch=compute_100,code=compute_100 -gencode=arch=compute_120,code=compute_120

				)

				set "CUDA_PATH=%CUDA_PATH_V124%"

				set "PATH=%CUDA_PATH_V124%\bin;%PATH%"

				set "CUDA_PATH=%CUDA_PATH_V129%"

				set "PATH=%CUDA_PATH_V129%\bin;%PATH%"

				:optcheck

									
										2

.ci/pytorch/windows/internal/check_deps.bat
									
												View File
												
				@ -65,7 +65,7 @@ for /F "usebackq delims=" %%i in (`python -c "import sys; print('{0[0]}{0[1]}'.f

				if  %PYVER% LSS 35 (

				    echo Warning: PyTorch for Python 2 under Windows is experimental.

				    echo Python x64 3.5 or up is recommended to compile PyTorch on Windows

				    echo Maybe you can create a virual environment if you have conda installed:

				    echo Maybe you can create a virtual environment if you have conda installed:

				    echo ^> conda create -n test python=3.6 pyyaml numpy

				    echo ^> activate test

				)

									
										1

.ci/pytorch/windows/internal/copy.bat
									
												View File
												
				@ -8,6 +8,7 @@ copy "%CUDA_PATH%\bin\cusolver*64_*.dll*" pytorch\torch\lib

				copy "%CUDA_PATH%\bin\cudnn*64_*.dll*" pytorch\torch\lib

				copy "%CUDA_PATH%\bin\nvrtc*64_*.dll*" pytorch\torch\lib

				copy "%CUDA_PATH%\extras\CUPTI\lib64\cupti64_*.dll*" pytorch\torch\lib

				copy "%CUDA_PATH%\extras\CUPTI\lib64\nvperf_host*.dll*" pytorch\torch\lib

				copy "C:\Program Files\NVIDIA Corporation\NvToolsExt\bin\x64\nvToolsExt64_1.dll*" pytorch\torch\lib

				copy "%PYTHON_LIB_PATH%\libiomp*5md.dll" pytorch\torch\lib

									
										82

.ci/pytorch/windows/internal/cuda_install.bat
									
												View File
												
				@ -23,66 +23,13 @@ set CUDNN_LIB_FOLDER="lib\x64"

				:: Skip all of this if we already have cuda installed

				if exist "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v%CUDA_VERSION_STR%\bin\nvcc.exe" goto set_cuda_env_vars

				if %CUDA_VER% EQU 118 goto cuda118

				if %CUDA_VER% EQU 124 goto cuda124

				if %CUDA_VER% EQU 126 goto cuda126

				if %CUDA_VER% EQU 128 goto cuda128

				if %CUDA_VER% EQU 129 goto cuda129

				echo CUDA %CUDA_VERSION_STR% is not supported

				exit /b 1

				:cuda118

				set CUDA_INSTALL_EXE=cuda_11.8.0_522.06_windows.exe

				if not exist "%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%" (

				    curl -k -L "https://ossci-windows.s3.amazonaws.com/%CUDA_INSTALL_EXE%" --output "%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%" & REM @lint-ignore

				    if errorlevel 1 exit /b 1

				    set "CUDA_SETUP_FILE=%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%"

				    set "ARGS=cuda_profiler_api_11.8 thrust_11.8 nvcc_11.8 cuobjdump_11.8 nvprune_11.8 nvprof_11.8 cupti_11.8 cublas_11.8 cublas_dev_11.8 cudart_11.8 cufft_11.8 cufft_dev_11.8 curand_11.8 curand_dev_11.8 cusolver_11.8 cusolver_dev_11.8 cusparse_11.8 cusparse_dev_11.8 npp_11.8 npp_dev_11.8 nvrtc_11.8 nvrtc_dev_11.8 nvml_dev_11.8 nvtx_11.8"

				)

				set CUDNN_FOLDER=cudnn-windows-x86_64-9.5.0.50_cuda11-archive

				set CUDNN_LIB_FOLDER="lib"

				set "CUDNN_INSTALL_ZIP=%CUDNN_FOLDER%.zip"

				if not exist "%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%" (

				    curl -k -L "http://s3.amazonaws.com/ossci-windows/%CUDNN_INSTALL_ZIP%" --output "%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%" & REM @lint-ignore

				    if errorlevel 1 exit /b 1

				    set "CUDNN_SETUP_FILE=%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%"

				)

				@REM cuDNN 8.3+ required zlib to be installed on the path

				echo Installing ZLIB dlls

				curl -k -L "http://s3.amazonaws.com/ossci-windows/zlib123dllx64.zip" --output "%SRC_DIR%\temp_build\zlib123dllx64.zip"

				7z x "%SRC_DIR%\temp_build\zlib123dllx64.zip" -o"%SRC_DIR%\temp_build\zlib"

				xcopy /Y "%SRC_DIR%\temp_build\zlib\dll_x64\*.dll" "C:\Windows\System32"

				goto cuda_common

				:cuda124

				set CUDA_INSTALL_EXE=cuda_12.4.0_551.61_windows.exe

				if not exist "%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%" (

				    curl -k -L "https://ossci-windows.s3.amazonaws.com/%CUDA_INSTALL_EXE%" --output "%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%" & REM @lint-ignore

				    if errorlevel 1 exit /b 1

				    set "CUDA_SETUP_FILE=%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%"

				    set "ARGS=cuda_profiler_api_12.4 thrust_12.4 nvcc_12.4 cuobjdump_12.4 nvprune_12.4 nvprof_12.4 cupti_12.4 cublas_12.4 cublas_dev_12.4 cudart_12.4 cufft_12.4 cufft_dev_12.4 curand_12.4 curand_dev_12.4 cusolver_12.4 cusolver_dev_12.4 cusparse_12.4 cusparse_dev_12.4 npp_12.4 npp_dev_12.4 nvrtc_12.4 nvrtc_dev_12.4 nvml_dev_12.4 nvjitlink_12.4 nvtx_12.4"

				)

				set CUDNN_FOLDER=cudnn-windows-x86_64-9.5.0.50_cuda12-archive

				set CUDNN_LIB_FOLDER="lib"

				set "CUDNN_INSTALL_ZIP=%CUDNN_FOLDER%.zip"

				if not exist "%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%" (

				    curl -k -L "http://s3.amazonaws.com/ossci-windows/%CUDNN_INSTALL_ZIP%" --output "%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%" & REM @lint-ignore

				    if errorlevel 1 exit /b 1

				    set "CUDNN_SETUP_FILE=%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%"

				)

				@REM cuDNN 8.3+ required zlib to be installed on the path

				echo Installing ZLIB dlls

				curl -k -L "http://s3.amazonaws.com/ossci-windows/zlib123dllx64.zip" --output "%SRC_DIR%\temp_build\zlib123dllx64.zip"

				7z x "%SRC_DIR%\temp_build\zlib123dllx64.zip" -o"%SRC_DIR%\temp_build\zlib"

				xcopy /Y "%SRC_DIR%\temp_build\zlib\dll_x64\*.dll" "C:\Windows\System32"

				goto cuda_common

				:cuda126

				@ -139,6 +86,33 @@ xcopy /Y "%SRC_DIR%\temp_build\zlib\dll_x64\*.dll" "C:\Windows\System32"

				goto cuda_common

				:cuda129

				set CUDA_INSTALL_EXE=cuda_12.9.1_576.57_windows.exe

				if not exist "%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%" (

				    curl -k -L "https://ossci-windows.s3.amazonaws.com/%CUDA_INSTALL_EXE%" --output "%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%" & REM @lint-ignore

				    if errorlevel 1 exit /b 1

				    set "CUDA_SETUP_FILE=%SRC_DIR%\temp_build\%CUDA_INSTALL_EXE%"

				    set "ARGS=cuda_profiler_api_12.9 thrust_12.9 nvcc_12.9 cuobjdump_12.9 nvprune_12.9 nvprof_12.9 cupti_12.9 cublas_12.9 cublas_dev_12.9 cudart_12.9 cufft_12.9 cufft_dev_12.9 curand_12.9 curand_dev_12.9 cusolver_12.9 cusolver_dev_12.9 cusparse_12.9 cusparse_dev_12.9 npp_12.9 npp_dev_12.9 nvrtc_12.9 nvrtc_dev_12.9 nvml_dev_12.9 nvjitlink_12.9 nvtx_12.9"

				)

				set CUDNN_FOLDER=cudnn-windows-x86_64-9.10.2.21_cuda12-archive

				set CUDNN_LIB_FOLDER="lib"

				set "CUDNN_INSTALL_ZIP=%CUDNN_FOLDER%.zip"

				if not exist "%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%" (

				    curl -k -L "http://s3.amazonaws.com/ossci-windows/%CUDNN_INSTALL_ZIP%" --output "%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%" & REM @lint-ignore

				    if errorlevel 1 exit /b 1

				    set "CUDNN_SETUP_FILE=%SRC_DIR%\temp_build\%CUDNN_INSTALL_ZIP%"

				)

				@REM cuDNN 8.3+ required zlib to be installed on the path

				echo Installing ZLIB dlls

				curl -k -L "http://s3.amazonaws.com/ossci-windows/zlib123dllx64.zip" --output "%SRC_DIR%\temp_build\zlib123dllx64.zip"

				7z x "%SRC_DIR%\temp_build\zlib123dllx64.zip" -o"%SRC_DIR%\temp_build\zlib"

				xcopy /Y "%SRC_DIR%\temp_build\zlib\dll_x64\*.dll" "C:\Windows\System32"

				goto cuda_common

				:cuda_common

				:: NOTE: We only install CUDA if we don't have it installed already.

				:: With GHA runners these should be pre-installed as part of our AMI process

									
										2

.ci/pytorch/windows/internal/install_python.bat
									
												View File
												
				@ -18,3 +18,5 @@ start /wait "" python-amd64.exe /quiet InstallAllUsers=1 PrependPath=0 Include_t

				if errorlevel 1 exit /b 1

				set "PATH=%CD%\Python\Scripts;%CD%\Python;%PATH%"

				%PYTHON_EXEC% -m pip install --upgrade pip setuptools packaging wheel

				if errorlevel 1 exit /b 1

									
										14

.ci/pytorch/windows/internal/smoke_test.bat
									
												View File
												
				@ -99,7 +99,6 @@ goto end

				:libtorch

				echo "install and test libtorch"

				if "%VC_YEAR%" == "2019" powershell internal\vs2019_install.ps1

				if "%VC_YEAR%" == "2022" powershell internal\vs2022_install.ps1

				if ERRORLEVEL 1 exit /b 1

				@ -111,10 +110,6 @@ pushd tmp\libtorch

				set VC_VERSION_LOWER=17

				set VC_VERSION_UPPER=18

				IF "%VC_YEAR%" == "2019" (

				    set VC_VERSION_LOWER=16

				    set VC_VERSION_UPPER=17

				)

				for /f "usebackq tokens=*" %%i in (`"%ProgramFiles(x86)%\Microsoft Visual Studio\Installer\vswhere.exe" -legacy -products * -version [%VC_VERSION_LOWER%^,%VC_VERSION_UPPER%^) -property installationPath`) do (

				    if exist "%%i" if exist "%%i\VC\Auxiliary\Build\vcvarsall.bat" (

				@ -153,14 +148,7 @@ if "%NVIDIA_GPU_EXISTS%" == "0" (

				    goto end

				)

				set BUILD_SPLIT_CUDA=

				if exist "%install_root%\lib\torch_cuda_cu.lib" if exist "%install_root%\lib\torch_cuda_cpp.lib" set BUILD_SPLIT_CUDA=ON

				if "%BUILD_SPLIT_CUDA%" == "ON" (

				    cl %PYTORCH_ROOT%\.ci\pytorch\test_example_code\check-torch-cuda.cpp torch_cpu.lib c10.lib torch_cuda_cu.lib torch_cuda_cpp.lib /EHsc /std:c++17 /link /INCLUDE:?warp_size@cuda@at@@YAHXZ /INCLUDE:?_torch_cuda_cu_linker_symbol_op_cuda@native@at@@YA?AVTensor@2@AEBV32@@Z

				) else (

				    cl %PYTORCH_ROOT%\.ci\pytorch\test_example_code\check-torch-cuda.cpp torch_cpu.lib c10.lib torch_cuda.lib /EHsc /std:c++17 /link /INCLUDE:?warp_size@cuda@at@@YAHXZ

				)

				cl %PYTORCH_ROOT%\.ci\pytorch\test_example_code\check-torch-cuda.cpp torch_cpu.lib c10.lib torch_cuda.lib /EHsc /std:c++17 /link /INCLUDE:?warp_size@cuda@at@@YAHXZ

				.\check-torch-cuda.exe

				if ERRORLEVEL 1 exit /b 1

									
										5

.ci/pytorch/windows/internal/vc_install_helper.bat
									
												View File
												
				@ -1,12 +1,7 @@

				if "%VC_YEAR%" == "2019" powershell windows/internal/vs2019_install.ps1

				if "%VC_YEAR%" == "2022" powershell windows/internal/vs2022_install.ps1

				set VC_VERSION_LOWER=17

				set VC_VERSION_UPPER=18

				if "%VC_YEAR%" == "2019" (

				    set VC_VERSION_LOWER=16

				    set VC_VERSION_UPPER=17

				)

				for /f "usebackq tokens=*" %%i in (`"%ProgramFiles(x86)%\Microsoft Visual Studio\Installer\vswhere.exe"  -products Microsoft.VisualStudio.Product.BuildTools -version [%VC_VERSION_LOWER%^,%VC_VERSION_UPPER%^) -property installationPath`) do (

				    if exist "%%i" if exist "%%i\VC\Auxiliary\Build\vcvarsall.bat" (

									
										48

.ci/pytorch/windows/internal/vs2019_install.ps1
									
												View File
											
				@ -1,48 +0,0 @@

				# https://developercommunity.visualstudio.com/t/install-specific-version-of-vs-component/1142479

				# https://docs.microsoft.com/en-us/visualstudio/releases/2019/history#release-dates-and-build-numbers

				# 16.8.6 BuildTools

				$VS_DOWNLOAD_LINK = "https://ossci-windows.s3.us-east-1.amazonaws.com/vs16.8.6_BuildTools.exe"

				$COLLECT_DOWNLOAD_LINK = "https://aka.ms/vscollect.exe"

				$VS_INSTALL_ARGS = @("--nocache","--quiet","--wait", "--add Microsoft.VisualStudio.Workload.VCTools",

				                                                     "--add Microsoft.Component.MSBuild",

				                                                     "--add Microsoft.VisualStudio.Component.Roslyn.Compiler",

				                                                     "--add Microsoft.VisualStudio.Component.TextTemplating",

				                                                     "--add Microsoft.VisualStudio.Component.VC.CoreIde",

				                                                     "--add Microsoft.VisualStudio.Component.VC.Redist.14.Latest",

				                                                     "--add Microsoft.VisualStudio.ComponentGroup.NativeDesktop.Core",

				                                                     "--add Microsoft.VisualStudio.Component.VC.Tools.x86.x64",

				                                                     "--add Microsoft.VisualStudio.ComponentGroup.NativeDesktop.Win81")

				curl.exe --retry 3 -kL $VS_DOWNLOAD_LINK --output vs_installer.exe

				if ($LASTEXITCODE -ne 0) {

				    echo "Download of the VS 2019 Version 16.8.5 installer failed"

				    exit 1

				}

				if (Test-Path "${env:ProgramFiles(x86)}\Microsoft Visual Studio\Installer\vswhere.exe") {

				    $existingPath = & "${env:ProgramFiles(x86)}\Microsoft Visual Studio\Installer\vswhere.exe" -products "Microsoft.VisualStudio.Product.BuildTools" -version "[16, 17)" -property installationPath

				    if ($existingPath -ne $null) {

				        if (!${env:CIRCLECI}) {

				            echo "Found correctly versioned existing BuildTools installation in $existingPath"

				            exit 0

				        }

				        echo "Found existing BuildTools installation in $existingPath, keeping it"

				    }

				}

				$process = Start-Process "${PWD}\vs_installer.exe" -ArgumentList $VS_INSTALL_ARGS -NoNewWindow -Wait -PassThru

				Remove-Item -Path vs_installer.exe -Force

				$exitCode = $process.ExitCode

				if (($exitCode -ne 0) -and ($exitCode -ne 3010)) {

				    echo "VS 2019 installer exited with code $exitCode, which should be one of [0, 3010]."

				    curl.exe --retry 3 -kL $COLLECT_DOWNLOAD_LINK --output Collect.exe

				    if ($LASTEXITCODE -ne 0) {

				        echo "Download of the VS Collect tool failed."

				        exit 1

				    }

				    Start-Process "${PWD}\Collect.exe" -NoNewWindow -Wait -PassThru

				    New-Item -Path "C:\w\build-results" -ItemType "directory" -Force

				    Copy-Item -Path "C:\Users\${env:USERNAME}\AppData\Local\Temp\vslogs.zip" -Destination "C:\w\build-results\"

				    exit 1

				}

									
										4

.ci/pytorch/windows/internal/xpu_install.bat
									
												View File
												
				@ -25,8 +25,8 @@ set XPU_EXTRA_INSTALLED=0

				set XPU_EXTRA_UNINSTALL=0

				if not [%XPU_VERSION%]==[] if [%XPU_VERSION%]==[2025.1] (

				    set XPU_BUNDLE_URL=https://registrationcenter-download.intel.com/akdlm/IRC_NAS/1a9fff3d-04c2-4d77-8861-3d86c774b66f/intel-deep-learning-essentials-2025.1.1.26_offline.exe

				    set XPU_BUNDLE_VERSION=2025.1.1+23

				    set XPU_BUNDLE_URL=https://registrationcenter-download.intel.com/akdlm/IRC_NAS/75d4eb97-914a-4a95-852c-7b9733d80f74/intel-deep-learning-essentials-2025.1.3.8_offline.exe

				    set XPU_BUNDLE_VERSION=2025.1.3+5

				)

				:: Check if XPU bundle is target version or already installed

									
										19

.ci/wheel/build_wheel.sh
									
												View File
												
				@ -127,7 +127,7 @@ export INSTALL_TEST=0 # dont install test binaries into site-packages

				export MACOSX_DEPLOYMENT_TARGET=10.15

				export CMAKE_PREFIX_PATH=${CONDA_PREFIX:-"$(dirname $(which conda))/../"}

				SETUPTOOLS_PINNED_VERSION="=46.0.0"

				SETUPTOOLS_PINNED_VERSION="==70.1.0"

				PYYAML_PINNED_VERSION="=5.3"

				EXTRA_CONDA_INSTALL_FLAGS=""

				CONDA_ENV_CREATE_FLAGS=""

				@ -135,7 +135,7 @@ RENAME_WHEEL=true

				case $desired_python in

				    3.13t)

				        echo "Using 3.13 deps"

				        SETUPTOOLS_PINNED_VERSION=">=68.0.0"

				        SETUPTOOLS_PINNED_VERSION=">=70.1.0"

				        PYYAML_PINNED_VERSION=">=6.0.1"

				        NUMPY_PINNED_VERSION="=2.1.0"

				        CONDA_ENV_CREATE_FLAGS="python-freethreading"

				@ -145,31 +145,31 @@ case $desired_python in

				        ;;

				    3.13)

				        echo "Using 3.13 deps"

				        SETUPTOOLS_PINNED_VERSION=">=68.0.0"

				        SETUPTOOLS_PINNED_VERSION=">=70.1.0"

				        PYYAML_PINNED_VERSION=">=6.0.1"

				        NUMPY_PINNED_VERSION="=2.1.0"

				        ;;

				    3.12)

				        echo "Using 3.12 deps"

				        SETUPTOOLS_PINNED_VERSION=">=68.0.0"

				        SETUPTOOLS_PINNED_VERSION=">=70.1.0"

				        PYYAML_PINNED_VERSION=">=6.0.1"

				        NUMPY_PINNED_VERSION="=2.0.2"

				        ;;

				    3.11)

				        echo "Using 3.11 deps"

				        SETUPTOOLS_PINNED_VERSION=">=46.0.0"

				        SETUPTOOLS_PINNED_VERSION=">=70.1.0"

				        PYYAML_PINNED_VERSION=">=5.3"

				        NUMPY_PINNED_VERSION="=2.0.2"

				        ;;

				    3.10)

				        echo "Using 3.10 deps"

				        SETUPTOOLS_PINNED_VERSION=">=46.0.0"

				        SETUPTOOLS_PINNED_VERSION=">=70.1.0"

				        PYYAML_PINNED_VERSION=">=5.3"

				        NUMPY_PINNED_VERSION="=2.0.2"

				        ;;

				    3.9)

				        echo "Using 3.9 deps"

				        SETUPTOOLS_PINNED_VERSION=">=46.0.0"

				        SETUPTOOLS_PINNED_VERSION=">=70.1.0"

				        PYYAML_PINNED_VERSION=">=5.3"

				        NUMPY_PINNED_VERSION="=2.0.2"

				        ;;

				@ -184,7 +184,8 @@ tmp_env_name="wheel_py$python_nodot"

				conda create ${EXTRA_CONDA_INSTALL_FLAGS} -yn "$tmp_env_name" python="$desired_python" ${CONDA_ENV_CREATE_FLAGS}

				source activate "$tmp_env_name"

				pip install "numpy=${NUMPY_PINNED_VERSION}"  "pyyaml${PYYAML_PINNED_VERSION}" requests ninja "setuptools${SETUPTOOLS_PINNED_VERSION}" typing_extensions

				retry pip install -r "${pytorch_rootdir}/requirements-build.txt"

				pip install "numpy=${NUMPY_PINNED_VERSION}"  "pyyaml${PYYAML_PINNED_VERSION}" requests ninja "setuptools${SETUPTOOLS_PINNED_VERSION}" typing-extensions

				retry pip install -r "${pytorch_rootdir}/requirements.txt" || true

				retry brew install libomp

				@ -206,7 +207,7 @@ if [[ "$USE_SPLIT_BUILD" == "true" ]]; then

				    BUILD_LIBTORCH_WHL=1 BUILD_PYTHON_ONLY=0 python setup.py bdist_wheel -d "$whl_tmp_dir"

				    echo "Finished setup.py bdist_wheel for split build (BUILD_LIBTORCH_WHL)"

				    echo "Calling setup.py bdist_wheel for split build (BUILD_PYTHON_ONLY)"

				    BUILD_PYTHON_ONLY=1 BUILD_LIBTORCH_WHL=0 python setup.py bdist_wheel -d "$whl_tmp_dir" --cmake

				    BUILD_LIBTORCH_WHL=0 BUILD_PYTHON_ONLY=1 CMAKE_FRESH=1 python setup.py bdist_wheel -d "$whl_tmp_dir"

				    echo "Finished setup.py bdist_wheel for split build (BUILD_PYTHON_ONLY)"

				else

				    python setup.py bdist_wheel -d "$whl_tmp_dir"

									
										5

.circleci/scripts/binary_populate_env.sh
									
												View File
												
				@ -75,8 +75,8 @@ TRITON_VERSION=$(cat $PYTORCH_ROOT/.ci/docker/triton_version.txt)

				# Here PYTORCH_EXTRA_INSTALL_REQUIREMENTS is already set for the all the wheel builds hence append TRITON_CONSTRAINT

				TRITON_CONSTRAINT="platform_system == 'Linux' and platform_machine == 'x86_64'"

				# CUDA 12.8 builds have triton for Linux and Linux aarch64 binaries.

				if [[ "$DESIRED_CUDA" == cu128 ]]; then

				# CUDA 12.9 builds have triton for Linux and Linux aarch64 binaries.

				if [[ "$DESIRED_CUDA" == "cu129" ]]; then

				  TRITON_CONSTRAINT="platform_system == 'Linux'"

				fi

				@ -105,6 +105,7 @@ fi

				# Set triton via PYTORCH_EXTRA_INSTALL_REQUIREMENTS for triton xpu package

				if [[ "$PACKAGE_TYPE" =~ .*wheel.* && -n "$PYTORCH_BUILD_VERSION" && "$PYTORCH_BUILD_VERSION" =~ .*xpu.* ]]; then

				    TRITON_VERSION=$(cat $PYTORCH_ROOT/.ci/docker/triton_xpu_version.txt)

				    TRITON_REQUIREMENT="pytorch-triton-xpu==${TRITON_VERSION}"

				    if [[ -n "$PYTORCH_BUILD_VERSION" && "$PYTORCH_BUILD_VERSION" =~ .*dev.* ]]; then

				        TRITON_SHORTHASH=$(cut -c1-8 $PYTORCH_ROOT/.ci/docker/ci_commit_pins/triton-xpu.txt)

									
										2

.circleci/scripts/binary_windows_build.sh
									
												View File
												
				@ -9,7 +9,7 @@ if [[ "$OS" != "windows-arm64" ]]; then

				    export USE_SCCACHE=1

				    export SCCACHE_BUCKET=ossci-compiler-cache

				    export SCCACHE_IGNORE_SERVER_IO_ERROR=1

				    export VC_YEAR=2019

				    export VC_YEAR=2022

				fi

				if [[ "$DESIRED_CUDA" == 'xpu' ]]; then

									
										2

.circleci/scripts/binary_windows_test.sh
									
												View File
												
				@ -4,7 +4,7 @@ set -eux -o pipefail

				source "${BINARY_ENV_FILE:-/c/w/env}"

				export CUDA_VERSION="${DESIRED_CUDA/cu/}"

				export VC_YEAR=2019

				export VC_YEAR=2022

				if [[ "$DESIRED_CUDA" == 'xpu' ]]; then

				    export VC_YEAR=2022

									
										157

.circleci/scripts/trigger_azure_pipeline.py
									
												View File
											
				@ -1,157 +0,0 @@

				# Documentation: https://docs.microsoft.com/en-us/rest/api/azure/devops/build/?view=azure-devops-rest-6.0

				import json

				import os

				import re

				import sys

				import time

				import requests

				AZURE_PIPELINE_BASE_URL = "https://aiinfra.visualstudio.com/PyTorch/"

				AZURE_DEVOPS_PAT_BASE64 = os.environ.get("AZURE_DEVOPS_PAT_BASE64_SECRET", "")

				PIPELINE_ID = "911"

				PROJECT_ID = "0628bce4-2d33-499e-bac5-530e12db160f"

				TARGET_BRANCH = os.environ.get("CIRCLE_BRANCH", "main")

				TARGET_COMMIT = os.environ.get("CIRCLE_SHA1", "")

				build_base_url = AZURE_PIPELINE_BASE_URL + "_apis/build/builds?api-version=6.0"

				s = requests.Session()

				s.headers.update({"Authorization": "Basic " + AZURE_DEVOPS_PAT_BASE64})

				def submit_build(pipeline_id, project_id, source_branch, source_version):

				    print("Submitting build for branch: " + source_branch)

				    print("Commit SHA1: ", source_version)

				    run_build_raw = s.post(

				        build_base_url,

				        json={

				            "definition": {"id": pipeline_id},

				            "project": {"id": project_id},

				            "sourceBranch": source_branch,

				            "sourceVersion": source_version,

				        },

				    )

				    try:

				        run_build_json = run_build_raw.json()

				    except json.decoder.JSONDecodeError as e:

				        print(e)

				        print(

				            "Failed to parse the response. Check if the Azure DevOps PAT is incorrect or expired."

				        )

				        sys.exit(-1)

				    build_id = run_build_json["id"]

				    print("Submitted bulid: " + str(build_id))

				    print("Bulid URL: " + run_build_json["url"])

				    return build_id

				def get_build(_id):

				    get_build_url = (

				        AZURE_PIPELINE_BASE_URL + f"/_apis/build/builds/{_id}?api-version=6.0"

				    )

				    get_build_raw = s.get(get_build_url)

				    return get_build_raw.json()

				def get_build_logs(_id):

				    get_build_logs_url = (

				        AZURE_PIPELINE_BASE_URL + f"/_apis/build/builds/{_id}/logs?api-version=6.0"

				    )

				    get_build_logs_raw = s.get(get_build_logs_url)

				    return get_build_logs_raw.json()

				def get_log_content(url):

				    resp = s.get(url)

				    return resp.text

				def wait_for_build(_id):

				    build_detail = get_build(_id)

				    build_status = build_detail["status"]

				    while build_status == "notStarted":

				        print("Waiting for run to start: " + str(_id))

				        sys.stdout.flush()

				        try:

				            build_detail = get_build(_id)

				            build_status = build_detail["status"]

				        except Exception as e:

				            print("Error getting build")

				            print(e)

				        time.sleep(30)

				    print("Bulid started: ", str(_id))

				    handled_logs = set()

				    while build_status == "inProgress":

				        try:

				            print("Waiting for log: " + str(_id))

				            logs = get_build_logs(_id)

				        except Exception as e:

				            print("Error fetching logs")

				            print(e)

				            time.sleep(30)

				            continue

				        for log in logs["value"]:

				            log_id = log["id"]

				            if log_id in handled_logs:

				                continue

				            handled_logs.add(log_id)

				            print("Fetching log: \n" + log["url"])

				            try:

				                log_content = get_log_content(log["url"])

				                print(log_content)

				            except Exception as e:

				                print("Error getting log content")

				                print(e)

				            sys.stdout.flush()

				        build_detail = get_build(_id)

				        build_status = build_detail["status"]

				        time.sleep(30)

				    build_result = build_detail["result"]

				    print("Bulid status: " + build_status)

				    print("Bulid result: " + build_result)

				    return build_status, build_result

				if __name__ == "__main__":

				    # Convert the branch name for Azure DevOps

				    match = re.search(r"pull/(\d+)", TARGET_BRANCH)

				    if match is not None:

				        pr_num = match.group(1)

				        SOURCE_BRANCH = f"refs/pull/{pr_num}/head"

				    else:

				        SOURCE_BRANCH = f"refs/heads/{TARGET_BRANCH}"

				    MAX_RETRY = 2

				    retry = MAX_RETRY

				    while retry > 0:

				        build_id = submit_build(PIPELINE_ID, PROJECT_ID, SOURCE_BRANCH, TARGET_COMMIT)

				        build_status, build_result = wait_for_build(build_id)

				        if build_result != "succeeded":

				            retry = retry - 1

				            if retry > 0:

				                print("Retrying... remaining attempt: " + str(retry))

				                # Wait a bit before retrying

				                time.sleep((MAX_RETRY - retry) * 120)

				                continue

				            else:

				                print("No more chance to retry. Giving up.")

				                sys.exit(-1)

				        else:

				            break

1

.clang-format

View File

 @ -120,6 +120,7 @@ UseTab:          Never
 Language: ObjC
 ColumnLimit: 120
 AlignAfterOpenBracket: Align
 IndentWidth: 2
 ObjCBlockIndentWidth: 2
 ObjCSpaceAfterProperty: false
 ObjCSpaceBeforeProtocolList: false

									
										4

.devcontainer/README.md
									
												View File
												
				@ -61,8 +61,8 @@ You are now all set to start developing with PyTorch in a DevContainer environme

				## Step 8: Build PyTorch

				To build pytorch from source, simply run:

				   ```

				   python setup.py develop

				   ```bash

				   python -m pip install --no-build-isolation -v -e .

				   ```

				The process involves compiling thousands of files, and would take a long time. Fortunately, the compiled objects can be useful for your next build. When you modify some files, you only need to compile the changed files the next time.

									
										24

.editorconfig
									
												View File
												
				@ -1,14 +1,36 @@

				root = true

				[*]

				charset = utf-8

				end_of_line = lf

				insert_final_newline = true

				# Python

				[*.py]

				[*.{py,pyi,py.in,pyi.in}]

				indent_style = space

				indent_size = 4

				# C/C++/CUDA

				[*.{cpp,hpp,cxx,cc,c,h,cu,cuh}]

				indent_style = space

				indent_size = 2

				# Objective-C

				[*.{mm,m,M}]

				indent_style = space

				indent_size = 2

				# Clang tools

				[.clang-{format,tidy}]

				indent_style = space

				indent_size = 2

				# Make

				[Makefile]

				indent_style = tab

				# Batch file

				[*.bat]

				indent_style = space

				indent_size = 2

				end_of_line = crlf

4

.flake8

View File

 @ -7,12 +7,12 @@ max-line-length = 120
 # C408 ignored because we like the dict keyword argument syntax
 # E501 is not flexible enough, we're using B950 instead
 ignore =
     E203,E305,E402,E501,E704,E721,E741,F405,F841,F999,W503,W504,C408,E302,W291,E303,
     E203,E305,E402,E501,E704,E721,E741,F405,F841,F999,W503,W504,C408,E302,W291,E303,F824,
     # shebang has extra meaning in fbcode lints, so I think it's not worth trying
     # to line this up with executable bit
     EXE001,
     # these ignores are from flake8-bugbear; please fix!
     B007,B008,B017,B019,B023,B028,B903,B904,B905,B906,B907
     B007,B008,B017,B019,B023,B028,B903,B904,B905,B906,B907,B908,B910
     # these ignores are from flake8-comprehensions; please fix!
     C407,
     # these ignores are from flake8-logging-format; please fix!

									
										5

.github/ISSUE_TEMPLATE/bug-report.yml
									
										vendored
									
												View File
												
				@ -12,7 +12,9 @@ body:

				    description: |

				      Please provide a clear and concise description of what the bug is.

				      If relevant, add a minimal example so that we can reproduce the error by running the code. It is very important for the snippet to be as succinct (minimal) as possible, so please take time to trim down any irrelevant code to help us debug efficiently. We are going to copy-paste your code and we expect to get the same result as you did: avoid any external data, and include the relevant imports, etc. For example:

				      If relevant, add a minimal example so that we can reproduce the error by running the code. It is very important for the snippet to be as succinct (minimal) as possible, so please take time to trim down any irrelevant code to help us debug efficiently.

				      Your example should be fully self-contained and not rely on any artifact that should be downloaded.

				      For example:

				      ```python

				      # All necessary imports at the beginning

				@ -26,6 +28,7 @@ body:

				      If the code is too long (hopefully, it isn't), feel free to put it in a public gist and link it in the issue: https://gist.github.com.

				      Please also paste or describe the results you observe instead of the expected results. If you observe an error, please paste the error message including the **full** traceback of the exception. It may be relevant to wrap error messages in ```` ```triple quotes blocks``` ````.

				      If your issue is related to numerical accuracy or reproducibility, please read the [numerical accuracy](https://docs.pytorch.org/docs/stable/notes/numerical_accuracy.html) and [reproducibility](https://docs.pytorch.org/docs/stable/notes/randomness.html) notes. If the difference is not expected as described in these documents, please provide appropriate justification on why one result is wrong and the other is correct.

				    placeholder: |

				      A clear and concise description of what the bug is.

									
										2

.github/ISSUE_TEMPLATE/disable-ci-jobs.md
									
										vendored
									
												View File
												
				@ -5,7 +5,7 @@ title: "DISABLED [WORKFLOW_NAME] / [PLATFORM_NAME] / [JOB_NAME]"

				labels: "module: ci"

				---

				> For example, DISABLED pull / win-vs2019-cpu-py3 / test (default). Once

				> For example, DISABLED pull / win-vs2022-cpu-py3 / test (default). Once

				> created, the job will be disabled within 15 minutes. You can check the

				> list of disabled jobs at https://ossci-metrics.s3.amazonaws.com/disabled-jobs.json

									
										111

.github/ISSUE_TEMPLATE/release-feature-request.yml
									
										vendored
									
										Normal file
									
												View File
												
				@ -0,0 +1,111 @@

				name: 🚀 Release highlight for proposed Feature

				description: Submit a Release highlight for proposed Feature

				labels: ["release-feature-request"]

				body:

				- type: textarea

				  attributes:

				    label: Release highlight for proposed Feature

				    description: >

				      Example: “A torch.special module, analogous to SciPy's special module.”

				- type: input

				  id: contact

				  attributes:

				    label: Point(s) of contact

				    description: How can we get in touch with you if we need more info?

				    placeholder: ex. github username

				  validations:

				    required: false

				- type: dropdown

				  attributes:

				    label: Release Mode (pytorch/pytorch features only)

				    description: |

				      If "out-of-tree", please include the GH repo name

				    options:

				      - In-tree

				      - Out-of-tree

				  validations:

				    required: true

				- type: textarea

				  attributes:

				    label: Out-Of-Tree Repo

				    description: >

				      please include the GH repo name

				  validations:

				    required: false

				- type: textarea

				  attributes:

				    label: Description and value to the user

				    description: >

				      Please provide a brief description of the feature and how it will benefit the user.

				  validations:

				    required: false

				- type: textarea

				  attributes:

				    label: Link to design doc, GitHub issues, past submissions, etc

				  validations:

				    required: false

				- type: textarea

				  attributes:

				    label: What feedback adopters have provided

				    description: >

				      Please list users/teams that have tried the feature and provided feedback. If that feedback motivated material changes (API, doc, etc..), a quick overview of the changes and the status (planned, in progress, implemented) would be helpful as well.

				  validations:

				    required: false

				- type: dropdown

				  attributes:

				    label: Plan for documentations / tutorials

				    description: |

				      Select One of the following options

				    options:

				      - Tutorial exists

				      - Will submit a PR to pytorch/tutorials

				      - Will submit a PR to a repo

				      - Tutorial is not needed

				  validations:

				    required: true

				- type: textarea

				  attributes:

				    label: Additional context for tutorials

				    description: >

				      Please provide a link for existing tutorial or link to a repo or context for why tutorial is not needed.

				  validations:

				    required: false

				- type: dropdown

				  attributes:

				    label: Marketing/Blog Coverage

				    description: |

				      Are you requesting feature Inclusion in the release blogs?

				    options:

				      - "Yes"

				      - "No"

				  validations:

				    required: true

				- type: textarea

				  attributes:

				    label: Are you requesting other marketing assistance with this feature?

				    description: >

				      E.g. supplementary blogs, social media amplification, etc.

				  validations:

				    required: false

				- type: textarea

				  attributes:

				    label: Release Version

				    description: >

				      Please include release version for marketing coverage.

				  validations:

				    required: false

				- type: textarea

				  attributes:

				    label: OS / Platform / Compute Coverage

				    description: >

				      Please list the platforms supported by the proposed feature. If the feature supports all the platforms, write "all". Goal of this section is to clearly share if this feature works in all PyTorch configurations or is it limited to only certain platforms/configurations (e.g. CPU only, GPU only, Linux only, etc...)

				  validations:

				    required: false

				- type: textarea

				  attributes:

				    label: Testing Support (CI, test cases, etc..)

				    description: >

				      Please provide an overview of test coverage. This includes unit testing and integration testing, but if E2E validation testing has been done to show that the feature works for a certain set of use cases or models please mention that as well.

				  validations:

				    required: false

Compare commits

3418 Commits starterTas ... codex/fix-

2 .bazelrc Unescape Escape View File

7 .ci/aarch64_linux/aarch64_ci_build.sh Unescape Escape View File

101 .ci/aarch64_linux/aarch64_wheel_ci_build.py Unescape Escape View File

6 .ci/caffe2/test.sh Unescape Escape View File

101 .ci/docker/README.md Unescape Escape View File

17 .ci/docker/almalinux/Dockerfile Unescape Escape View File

188 .ci/docker/build.sh Unescape Escape View File

1 .ci/docker/centos-rocm/Dockerfile Unescape Escape View File

2 .ci/docker/ci_commit_pins/executorch.txt Unescape Escape View File

2 .ci/docker/ci_commit_pins/nccl-cu12.txt Unescape Escape View File

1 .ci/docker/ci_commit_pins/torchbench.txt Normal file Unescape Escape View File

2 .ci/docker/ci_commit_pins/triton-xpu.txt Unescape Escape View File

2 .ci/docker/ci_commit_pins/triton.txt Unescape Escape View File

4 .ci/docker/common/common_utils.sh Unescape Escape View File

16 .ci/docker/common/install_base.sh Unescape Escape View File

26 .ci/docker/common/install_conda.sh Unescape Escape View File

36 .ci/docker/common/install_cpython.sh Unescape Escape View File

97 .ci/docker/common/install_cuda.sh Unescape Escape View File

26 .ci/docker/common/install_cudnn.sh Unescape Escape View File

7 .ci/docker/common/install_cusparselt.sh Unescape Escape View File

30 .ci/docker/common/install_inductor_benchmark_deps.sh Unescape Escape View File

16 .ci/docker/common/install_onnx.sh Unescape Escape View File

7 .ci/docker/common/install_openblas.sh Unescape Escape View File

56 .ci/docker/common/install_rocm.sh Unescape Escape View File

7 .ci/docker/common/install_rocm_magma.sh Unescape Escape View File

14 .ci/docker/common/install_triton.sh Unescape Escape View File

12 .ci/docker/common/install_xpu.sh Unescape Escape View File

15 .ci/docker/libtorch/Dockerfile Unescape Escape View File

4 .ci/docker/libtorch/build.sh Unescape Escape View File

2 .ci/docker/linter/Dockerfile Unescape Escape View File

3 .ci/docker/manywheel/Dockerfile_2_28 Unescape Escape View File

5 .ci/docker/manywheel/Dockerfile_2_28_aarch64 Unescape Escape View File

2 .ci/docker/manywheel/Dockerfile_cuda_aarch64 Unescape Escape View File

21 .ci/docker/manywheel/Dockerfile_s390x Unescape Escape View File

7 .ci/docker/manywheel/build.sh Unescape Escape View File

58 .ci/docker/requirements-ci.txt Unescape Escape View File

19 .ci/docker/requirements-docs.txt Unescape Escape View File

2 .ci/docker/triton_version.txt Unescape Escape View File

1 .ci/docker/triton_xpu_version.txt Normal file Unescape Escape View File

170 .ci/docker/ubuntu-cuda/Dockerfile Unescape Escape View File

4 .ci/docker/ubuntu-rocm/Dockerfile Unescape Escape View File

2 .ci/docker/ubuntu-xpu/Dockerfile Unescape Escape View File

9 .ci/docker/ubuntu/Dockerfile Unescape Escape View File

16 .ci/magma/Makefile Unescape Escape View File

15 .ci/manywheel/build_common.sh Unescape Escape View File

164 .ci/manywheel/build_cuda.sh Unescape Escape View File

12 .ci/manywheel/build_libtorch.sh Unescape Escape View File

25 .ci/manywheel/build_rocm.sh Unescape Escape View File

2 .ci/onnx/test.sh Unescape Escape View File

34 .ci/pytorch/build-mobile.sh Unescape Escape View File

67 .ci/pytorch/build.sh Unescape Escape View File

2 .ci/pytorch/check_binary.sh Unescape Escape View File

7 .ci/pytorch/common-build.sh Unescape Escape View File

2 .ci/pytorch/common.sh Unescape Escape View File

153 .ci/pytorch/common_utils.sh Unescape Escape View File

123 .ci/pytorch/create_test_cert.py Unescape Escape View File

2 .ci/pytorch/macos-build.sh Unescape Escape View File

10 .ci/pytorch/macos-common.sh Unescape Escape View File

128 .ci/pytorch/macos-test.sh Unescape Escape View File

18 .ci/pytorch/run_glootls_test.sh Unescape Escape View File

5 .ci/pytorch/run_tests.sh Unescape Escape View File

2 .ci/pytorch/smoke_test/check_binary_symbols.py Unescape Escape View File

3 .ci/pytorch/smoke_test/check_gomp.py Unescape Escape View File

27 .ci/pytorch/smoke_test/smoke_test.py Unescape Escape View File

265 .ci/pytorch/test.sh Unescape Escape View File

34 .ci/pytorch/win-arm64-build.ps1 Normal file Unescape Escape View File

24 .ci/pytorch/win-arm64-test.sh Normal file Unescape Escape View File

2 .ci/pytorch/win-build.sh Unescape Escape View File

98 .ci/pytorch/win-test-helpers/arm64/build_pytorch.ps1 Normal file Unescape Escape View File

4 .ci/pytorch/win-test-helpers/build_pytorch.bat Unescape Escape View File

2 .ci/pytorch/win-test-helpers/run_python_nn_smoketests.py Unescape Escape View File

7 .ci/pytorch/win-test.sh Unescape Escape View File

2 .ci/pytorch/windows/arm64/bootstrap_libuv.bat Unescape Escape View File

2 .ci/pytorch/windows/arm64/bootstrap_openblas.bat Unescape Escape View File

2 .ci/pytorch/windows/arm64/bootstrap_tests.bat Unescape Escape View File

2 .ci/pytorch/windows/arm64/build_libtorch.bat Unescape Escape View File

2 .ci/pytorch/windows/arm64/build_pytorch.bat Unescape Escape View File

2 .ci/pytorch/windows/arm64/smoke_test.bat Unescape Escape View File

3418 Commits

starterTas ... codex/fix-

2

.bazelrc

View File

7

.ci/aarch64_linux/aarch64_ci_build.sh

View File

101

.ci/aarch64_linux/aarch64_wheel_ci_build.py

View File

6

.ci/caffe2/test.sh

View File

101

.ci/docker/README.md

View File

17

.ci/docker/almalinux/Dockerfile

View File

188

.ci/docker/build.sh

View File

1

.ci/docker/centos-rocm/Dockerfile

View File

2

.ci/docker/ci_commit_pins/executorch.txt

View File

2

.ci/docker/ci_commit_pins/nccl-cu12.txt

View File

1

.ci/docker/ci_commit_pins/torchbench.txt Normal file

View File

2

.ci/docker/ci_commit_pins/triton-xpu.txt

View File

2

.ci/docker/ci_commit_pins/triton.txt

View File

4

.ci/docker/common/common_utils.sh

View File

16

.ci/docker/common/install_base.sh

View File

26

.ci/docker/common/install_conda.sh

View File

36

.ci/docker/common/install_cpython.sh

View File

97

.ci/docker/common/install_cuda.sh

View File

26

.ci/docker/common/install_cudnn.sh

View File

7

.ci/docker/common/install_cusparselt.sh

View File

30

.ci/docker/common/install_inductor_benchmark_deps.sh

View File

16

.ci/docker/common/install_onnx.sh

View File

7

.ci/docker/common/install_openblas.sh

View File

56

.ci/docker/common/install_rocm.sh

View File

7

.ci/docker/common/install_rocm_magma.sh

View File

14

.ci/docker/common/install_triton.sh

View File

12

.ci/docker/common/install_xpu.sh

View File

15

.ci/docker/libtorch/Dockerfile

View File

4

.ci/docker/libtorch/build.sh

View File

2

.ci/docker/linter/Dockerfile

View File

3

.ci/docker/manywheel/Dockerfile_2_28

View File

5

.ci/docker/manywheel/Dockerfile_2_28_aarch64

View File

2

.ci/docker/manywheel/Dockerfile_cuda_aarch64

View File

21

.ci/docker/manywheel/Dockerfile_s390x

View File

7

.ci/docker/manywheel/build.sh

View File

58

.ci/docker/requirements-ci.txt

View File

19

.ci/docker/requirements-docs.txt

View File

2

.ci/docker/triton_version.txt

View File

1

.ci/docker/triton_xpu_version.txt Normal file

View File

170

.ci/docker/ubuntu-cuda/Dockerfile

View File

4

.ci/docker/ubuntu-rocm/Dockerfile

View File

2

.ci/docker/ubuntu-xpu/Dockerfile

View File

9

.ci/docker/ubuntu/Dockerfile

View File

16

.ci/magma/Makefile

View File

15

.ci/manywheel/build_common.sh

View File

164

.ci/manywheel/build_cuda.sh

View File

12

.ci/manywheel/build_libtorch.sh

View File

25

.ci/manywheel/build_rocm.sh

View File

2

.ci/onnx/test.sh

View File

34

.ci/pytorch/build-mobile.sh

View File

67

.ci/pytorch/build.sh

View File

2

.ci/pytorch/check_binary.sh

View File

7

.ci/pytorch/common-build.sh

View File

2

.ci/pytorch/common.sh

View File

153

.ci/pytorch/common_utils.sh

View File

123

.ci/pytorch/create_test_cert.py

View File

2

.ci/pytorch/macos-build.sh

View File

10

.ci/pytorch/macos-common.sh

View File

128

.ci/pytorch/macos-test.sh

View File

18

.ci/pytorch/run_glootls_test.sh

View File

5

.ci/pytorch/run_tests.sh

View File

2

.ci/pytorch/smoke_test/check_binary_symbols.py

View File

3

.ci/pytorch/smoke_test/check_gomp.py

View File

27

.ci/pytorch/smoke_test/smoke_test.py

View File

265

.ci/pytorch/test.sh

View File

34

.ci/pytorch/win-arm64-build.ps1 Normal file

View File

24

.ci/pytorch/win-arm64-test.sh Normal file

View File

2

.ci/pytorch/win-build.sh

View File

98

.ci/pytorch/win-test-helpers/arm64/build_pytorch.ps1 Normal file

View File

4

.ci/pytorch/win-test-helpers/build_pytorch.bat

View File

2

.ci/pytorch/win-test-helpers/run_python_nn_smoketests.py

View File

7

.ci/pytorch/win-test.sh

View File

2

.ci/pytorch/windows/arm64/bootstrap_libuv.bat

View File

2

.ci/pytorch/windows/arm64/bootstrap_openblas.bat

View File

2

.ci/pytorch/windows/arm64/bootstrap_tests.bat

View File

2

.ci/pytorch/windows/arm64/build_libtorch.bat

View File

2

.ci/pytorch/windows/arm64/build_pytorch.bat

View File

2

.ci/pytorch/windows/arm64/smoke_test.bat

View File

59

.ci/pytorch/windows/cuda118.bat

View File