199467 Commits
Author SHA1 Message Date
Jian Cai b5d8801b1e [XLA][HLO Value Tracking] Maintain OriginalValue consistency in MemorySpaceAssignment
Root Cause:
`MemorySpaceAssignment::Process()` rewrites the HLO graph in two sequential loops over all `Allocation` objects:
1. Loop 1 calls the virtual method `Allocation::Process()` (`CopyAllocation::Process()`, `PinnedAllocation::Process()`, and `ParentAllocation::Process()`).
2. Loop 2 calls the virtual method `Allocation::PostProcess()` (`ParentAllocation::PostProcess()`).

Several HLO transformations across these two loops failed to maintain or propagate `OriginalValue`:
1. While-loop widening in `ParentAllocation::Process()` (Loop 1) and `ParentAllocation::PostProcess()` (Loop 2):
   When `ParentAllocation` passes a buffer from a parent computation into a `while` loop, it widens the `while` loop tuple from N elements to N+1 elements at `new_tuple_index`:
   - `ParentAllocation::Process()` widens the `while` loop inputs by creating a new input tuple (`new_while_operand = TupleUtil::ReplaceTupleWith(...)`) and mutating the shapes of `calling_instruction_` (the `kWhile` instruction), `while_condition->parameter_instruction(0)`, and `while_body->parameter_instruction(0)` in-place. Because the in-place mutated instructions kept their old N-element `OriginalValue` trees while their shapes became (N+1)-tuples, `HloVerifier` failed with shape mismatches. Meanwhile, `new_while_operand` was created with `original_value() == nullptr`.
   - To keep downstream users of the `while` loop valid, `ParentAllocation::Process()` creates `tuple_with_old_shape = TupleUtil::ExtractPrefix(calling_instruction_, new_tuple_index)` and replaces uses of `calling_instruction_` with `tuple_with_old_shape`. `tuple_with_old_shape` and its `GetTupleElement` operands were created with `original_value() == nullptr`, dropping value tracking for all users after the loop.
   - `ParentAllocation::PostProcess()` runs in Loop 2 (after all Loop 1 allocations inside `while_body` have finished) to widen the `while_body` root tuple (`new_while_body_root = TupleUtil::ReplaceTupleWith(added_element, old_body_root, ...)`), which also left `new_while_body_root` with `original_value() == nullptr` and dropped the appended element's `OriginalValue` at `{new_tuple_index}`.
2. Tuple replacement and async copy creation in `CopyAllocation::Process()` and `Allocation::UpdateUses()` (Loop 1):
   - `CopyAllocation::Process()` inserts `copy-start` and `copy-done` instructions to move buffers between default memory (HBM) and alternate memory (VMEM), but left `copy_done_` with `original_value() == nullptr` instead of copying `producing_instruction`'s `OriginalValue`.
   - `CopyAllocation::Process()`, `PinnedAllocation::Process()`, and `ParentAllocation::Process()` all call the helper `Allocation::UpdateUses()` during Loop 1 to rewire consumer instructions to use `producing_instruction`. When a consumer is a `kTuple` instruction (such as `while_body->root_instruction()` or `while_init`), `Allocation::UpdateUses()` calls `TupleUtil::ReplaceTupleWith`, which constructs a brand-new `kTuple` instruction with `original_value() == nullptr`. Consequently, any Loop 1 allocation updating an element of `while_body->root_instruction()` wiped out its `OriginalValue` to `nullptr` before Loop 2 (`ParentAllocation::PostProcess()`) even ran.

What This CL Fixes:
- Add `TupleTree::num_children(ShapeIndexView index = {})` in `tuple_tree.h` to query the number of children of a tuple node directly from the index table.
- Introduce `AppendToTupleOriginalValue(src_tuple, dest_tuple, appended_instr)` in `allocation.cc`, which uses `CopyOriginalValue` from `hlo_original_value_util.h` to map existing tuple elements from `src_tuple` to the widened `dest_tuple` and copies `appended_instr->original_value()->tree()` into the last element (`dest_tuple->shape().tuple_shapes().size() - 1`) via `CopyCompatibleSubtreeFrom`.
- In `ParentAllocation::Process()`, before widening `calling_instruction_` in-place, copy its N-element `OriginalValue` to `tuple_with_old_shape` via `CopyOriginalValue` and propagate subtrees to its `GetTupleElement` operands. Then widen `new_while_operand` with `AppendToTupleOriginalValue` and update `calling_instruction_`, `while_condition` parameter 0, and `while_body` parameter 0 via `AppendToWhileLoopOriginalValue`.
- In `AppendToWhileLoopOriginalValue` (`while_util.cc`), only update `while_body->root_instruction()` if its shape has already been widened to match `while_instr->shape()`, supporting two-phase while-loop widening where `while_body` root widening is deferred to `ParentAllocation::PostProcess()`.
- In `Allocation::AddGetTupleElements()`, propagate the `OriginalValue` subtree from `defining_position().instruction` at `defining_position().index` to the newly created `GetTupleElement` instruction.
- In `ParentAllocation::PostProcess()`, call `AppendToTupleOriginalValue(old_body_root, new_while_body_root, added_element)` so the widened `while_body` root tuple preserves both the body root's `OriginalValue` and the pass-through element's `OriginalValue`.
- In `CopyAllocation::Process()`, propagate `producing_instruction`'s `OriginalValue` to `copy_done_` (and any intermediate bitcast).
- In `Allocation::UpdateUses()`, when `TupleUtil::ReplaceTupleWith` replaces a `kTuple` instruction, clone `tuple_inst`'s `OriginalValue` onto `replacement_instruction` and update `{use.operand_index}` from `producing_instruction->original_value()`.

PiperOrigin-RevId: 987774950
2026-09-25 15:56:09 -07:00
Bixia Zheng 4fa7af7000 Clear sharding annotations on entry parameters at the end of GSPMD partitioning.
After GSPMD generates device local, it clears sharding except for two cases: (1) Send/Recv instructions and (2) Parameter instructions. Case (1) is due to the fact that GSPMD doesn't handle the single device sharding on Send/Recv instructions, and it leaves the sharding for later passses to handle it. Case (2) is due to the fact that there are many places in the compilation pipeline that find the module input sharding from the sharding on parameter instructions, even though the information is also available in the module spmd_parameters_shardings metadata.

This CL makes GSPMD also clear sharding on parameter instructions, and updates downstream users to read parameter shardings from the module's spmd_parameters_shardings metadata.

PiperOrigin-RevId: 987756729
2026-09-25 15:41:49 -07:00
Zaim Codi fadb2ca7b5 Add algebraic simplification logic for kShuffle::kRotate.
Extend the HLO optimization pipeline with algebraic simplifications for `HloOpcode::kShuffle` in `kRotate` mode, allowing compiler pipelines to
simplify rotations.

Specifically, this commit introduces:
* **AlgebraicSimplifier (`algebraic_simplifier.cc`)**: Add simplification rules for `kShuffle::kRotate` instructions
- Eliminating identity rotations (`norm_shift == 0` , empty shift dimensions, splat dimension)
- Folding/combining chained `kShuffle::kRotate` operations.
- Canonicalize shifts to be in `[0, dim_size)` and order dimensions in ascending order
* **Testing**: Add comprehensive unit tests in `AlgebraicSimplifierTest`.

PiperOrigin-RevId: 987747347
2026-09-25 15:37:09 -07:00
TensorFlower Gardener 159f4e5a03 Merge pull request #126911 from dynamo-pentester:patch-1
PiperOrigin-RevId: 987716993
2026-09-25 15:08:47 -07:00
Junwhan Ahn e3f6b57626 Replace IFRT mock usage with real client in //third_party/tensorflow/core/tfrt/ifrt:ifrt_device_utils_test
PiperOrigin-RevId: 987641321
2026-09-25 14:51:30 -07:00
Yulia Baturina 26e6f35a31 Update shardy commit.
PiperOrigin-RevId: 987639736
2026-09-25 14:30:06 -07:00
Zaim Codi 7a33d5986c Add first-class kShuffle instruction and builder API across HLO and platforms.
Introduce `HloOpcode::kShuffle` (`HloShuffleInstruction`), which moves the
elements of its operand around along a set of `dimensions`. A single opcode
covers the whole family of data-movement patterns (rotate, reverse, permute,
...), which keeps the semantic intent available to compiler analysis, layout
assignment, and target-specific lowerings instead of forcing an expansion into
slice and concatenate during graph construction.

The pattern to apply is selected by `xla.ShuffleMode`, a `oneof` pairing each
mode with exactly the attributes that parameterize it, so a mode cannot be
combined with another mode's attributes and adding a mode adds no fields to
`HloInstructionProto`. Switches over `mode()` are exhaustive and have no
`default`, so a new mode does not compile until every consumer has handled it.

Only `rotate` is implemented here. It takes one `shifts` entry per shuffled
dimension and rotates the elements to the left.

Specifically:
*   **HLO IR & Parser**: `HloShuffleInstruction` carries `dimensions` and a
    `ShuffleMode`, with text serialization and parsing. The text form is
    `shuffle(x), dimensions={0,1}, mode=rotate, shifts={2,5}`, where the name of
    a mode is the name of its field in `ShuffleMode`.
*   **Shape Inference & Verifier**: `ShapeInference::InferShuffleShape` and the
    verifier reject duplicate or out-of-range `dimensions`, a missing mode, and
    mode attributes inconsistent with the mode (for `rotate`, `dimensions` and
    `shifts` of unequal size).
*   **XlaBuilder API**: `xla::Shuffle` constructs shuffles from symbolic
    handles. The new `shuffle` utility provides `MakeRotateMode` to build a rotate
    mode and `NormalizeShift` to map an arbitrary shift onto `[0, dim_size)`.
*   **HLO-to-MHLO / StableHLO Translation**: `HloFunctionImporter` imports
    `kShuffle`, lowering each rotated dimension sequentially into
    `stablehlo.slice` and `stablehlo.concatenate`
    (`concat(slice(shift:), slice(0:shift))`).
*   **Pass Integration**: Register `kShuffle` across core compiler passes
    including instruction fusion, layout assignment, and sharding propagation.
PiperOrigin-RevId: 987630885
2026-09-25 14:12:36 -07:00
Junwhan Ahn 8bebdffc79 [PJRT:CPU] Use separate bounded thread pools for compilation, execution, and transfers.
- Introduce dedicated bounded thread pools in `PjRtCpuRawClient` for compilation (`compile_thread_pool_`) and execution (`execute_work_runner_`), separate from the general `async_work_runner_` used for buffer transfers and linearization/delinearization.
- Route `CompileInternal` to `compile_thread_pool_` and `CpuPjRtRawLoadedExecutable::Execute` to `execute_work_runner_`.

Separating execution, compilation, and transfer thread pools prevents multi-device collective launches from starving antecedent input buffer transfers on `async_work_runner_`, while keeping all three pools bounded so bursts of D2H/H2D buffer transfers remain throttled.

PiperOrigin-RevId: 987628676
2026-09-25 14:07:44 -07:00
Henning Becker 0668dfb00c [XLA:FFI] Make //xla/ffi:api an alias of //xla/ffi/api:api
PiperOrigin-RevId: 987625577
2026-09-25 13:47:11 -07:00
Majid Dadashi 5442ef5943 Add blockwise support in the litert converter.
Also includes a quant_spec metadata in tfl.fully_connected to allow encoding custom specs.

PiperOrigin-RevId: 987624321
2026-09-25 13:28:57 -07:00
Yue Sheng 659f756d1a [Mosaic TPU] Add some fold rules for the transpose op.
Reverts 2a3525476e

PiperOrigin-RevId: 987600180
2026-09-25 12:45:55 -07:00
TensorFlower Gardener 6eb71e6d0d Merge pull request #120176 from Ashutosh0x:fix/unsorted-segment-sum-overflow
PiperOrigin-RevId: 987596387
2026-09-25 12:39:13 -07:00
Sohaib Iftikhar a1ae4db3a8 [XLA:GPU]: Fix caching issue with kernels when unrolling while loops
When while loops are unrolled the same thunk is called with `Record` multiple times for different loop iterations. Currently `CustomCallRecordState` stored a single sequence of `XLA_FFI_Command*`, so during update only the last recorded sequence of commands was passed to the FFI record API, overwriting the last iteration and leaving earlier iterations of the while loop with stale buffer pointers.

Keying the state by the sink command pointer returned on `RecordCreate` (and passed back via `RecordUpdate::command`) so each unrolled iteration updates its own recorded command sequence.

PiperOrigin-RevId: 987596254
2026-09-25 12:09:05 -07:00
Jian Cai eaedaed010 [XLA][HLO Value Tracking] Maintain OriginalValue consistency across while loop transformations
- In WhileUtil::AppendToWhileLoopOriginalValue and WhileUtil::MakeInstructionsLiveIn, update while instruction, while init, body root, body parameter, and condition parameter to keep all loop instruction OriginalValues structurally consistent when widening while loops, and preserve call_hierarchy.
- In WhileLoopSimplifier (TryRemoveDeadTupleElements, TryRemoveConstantTupleElements, TryFlattenNestedTuples, and TrySimplifyInductionVariables), propagate and structurally update OriginalValues across all 5 loop boundary instructions and verify compatibility with the new shapes.

PiperOrigin-RevId: 987569790
2026-09-25 11:55:17 -07:00
Sohaib Iftikhar dfa4939546 [XLA:GPU]: Implement cufunction record for GPU backend.
PiperOrigin-RevId: 987558515
2026-09-25 11:51:14 -07:00
Aliia Khasanova 615d5ee70e Split collective_ops AOT goldens into per-arch directories
Goldens move to `executables/collective_ops_aot_test_2gpu_<arch>/`, one
directory per backend.

gb300 stays excluded through the existing `disabled_backends` entry
(b/491194726); no gb300 target is generated, so `DetectGpuArchToken()` has no
gb300 branch and a gb300 device would hit its `LOG(FATAL)`.

PiperOrigin-RevId: 987552652
2026-09-25 11:23:32 -07:00
Aliia Khasanova bf988457bb Add update_goldens.py to regenerate AOT compatibility goldens
Provides a script to automate regenerating and promoting AOT compatibility golden snapshots from Bazel undeclared test outputs into `<package>/executables/<target_name>/v<N+1>/`.

Also updates `AOTInterceptionPjrtClient::VerifyAgainstGolden` and `test_lib.cc` (`GetUndeclaredOutputsDir`) to reference `update_goldens.py` in golden comparison failure and precondition error messages. The golden
comparison failure message prints the failing test's full label (from
`TEST_TARGET`), which is what the script takes as its argument.

PiperOrigin-RevId: 987522513
2026-09-25 11:10:11 -07:00
Mohammed Anany 9c7723c4db Return a ready empty vector from typed JoinFutures on empty input
`Future<std::vector<T>>({})` cannot select the templated converting
constructors of `Future`, because a braced-init-list cannot deduce `U`.
Overload resolution therefore falls back to the move constructor from a
value-initialized `Future`, and both typed `JoinFutures` overloads returned an
invalid future for an empty span instead of a ready empty vector. Calling
`Map`, `OnReady` or `Await` on it dereferences a null `AsyncValue`.

PiperOrigin-RevId: 987497834
2026-09-25 11:06:13 -07:00
Eusebio Durán Montaña b08a6e3e6c Write ExecutableAndOptionsProto records uncompressed.
`serialized_executable` is already a snappy-compressed split
`GpuExecutableProto` (see `WriteSplitGpuExecutable`). Compressing the outer
container again saves no space (dropping it adds 0.06% on a 330 MB
executable), but it costs an extra compression pass on write and a full
decompression pass on load.

Riegeli stores the compression type per chunk, so readers don't change and
existing artifacts still load.

Alternative considered: leave the inner proto uncompressed and compress the
outer one instead. That would mean deciding per use case when to skip
compression, which adds logic for little gain, since the outer proto is
small apart from `serialized_executable`.

PiperOrigin-RevId: 987497408
2026-09-25 10:45:29 -07:00
Eusebio Durán Montaña 1d4cc1bac9 Parse split-proto merge records from a riegeli::Chain without flattening.
`HandleProtoMergeRecord` read each record into a flat `absl::string_view` and
then called `MergeFromString`. For records larger than a riegeli buffer (e.g.
an HLO module carrying hundreds of MB of constants) the decoded record spans
several blocks, and producing a flat view forces riegeli to copy it into a
contiguous scratch buffer first.

Read the record as a `riegeli::Chain` instead, which shares the decoded
blocks, and parse it with `riegeli::ParseMessage(..., set_merge(true))`,
which consumes the chain directly.

PiperOrigin-RevId: 987496597
2026-09-25 10:27:40 -07:00
Mikhail Goncharov 8a022e7d11 [XLA:GPU] simplify multiple expressions in tiling space
That will be used later to add constraints to one indexing map. Should also marginally imporove compilation time.

Also call simplify in tiling propagation only if we are not in "symbolic" case where we will not use them simplified anyway.

PiperOrigin-RevId: 987492063
2026-09-25 10:22:35 -07:00
Toli Yevtushenko 277107e9a3 [XLA] Update TODO comments on closed duplicate bugs to active canonical bugs
PiperOrigin-RevId: 987488947
2026-09-25 09:50:53 -07:00
Eusebio Durán Montaña 9b12b73b2e Document XLA:GPU compiler/runtime compatibility requirements.
Now that we have this compatibility window, we'll often need to ask OSS developers to change their PRs to avoid breaking it (e.g. https://github.com/openxla/xla/pull/46865). So we should have some documentation that we can point to.

PiperOrigin-RevId: 987466715
2026-09-25 09:38:55 -07:00
Adrian Kuegel e7a7327618 [XLA:GPU] Require bitcasts to keep bitwidth in CopyFusion
AlgebraicSimplifier can rewrite bitcast-convert instructions into bitcast even
when the element bitwidth changes (when the minor-most dimension is contiguous).
While PriorityFusion and GpuFusible guard against fusing width-changing bitcasts
using hlo_instruction_utils::KeepsBitwidth, CopyFusion previously skipped
through any single-user bitcast without checking its bitwidth.

Fusing a width-changing bitcast and its consuming copy into a producer fusion
breaks the MLIR fusion emitters because GetBitcastMap requires matching element
counts between roots and elemental lowering cannot emit a scalar
mhlo.bitcast_convert across different bitwidths.

Guard kBitcast checks in CopyFusion with hlo_instruction_utils::KeepsBitwidth.

PiperOrigin-RevId: 987463670
2026-09-25 09:35:40 -07:00
TensorFlower Gardener 712767ce01 Merge pull request #126091 from Abhirup0:fix-inter-op-threads-crash
PiperOrigin-RevId: 987452514
2026-09-25 09:15:29 -07:00
Ilya Tikhonovskiy 6d4756c6b2 [XLA:GPU] Add cuDNN scaled dot device test suite.
Adds cudnn_scaled_dot_test covering block-scaled dot support in cuDNN on Blackwell. Parameterized across MXFP8 and NVFP4 type combinations and transposition layouts, reporting lowering stages (composite, optimized HLO, cuDNN graph representation) and execution parity against reference decomposition.

PiperOrigin-RevId: 987447766
2026-09-25 08:57:37 -07:00
Christian Sigg de517cb79a Add defensive checks for P2P symmetric memory support and null windows.
Add pre-checks for P2P symmetric memory support via cudaDeviceCanAccessPeer,
post-checks for null windows returned from ncclCommWindowRegister, and
defensive null window checks in multimem_addr, peer_addr, and the destructor.

PiperOrigin-RevId: 987435182
2026-09-25 08:52:31 -07:00
Enver Kayaaslan d9ecfe9451 Reverts 84df1268af
PiperOrigin-RevId: 987427673
2026-09-25 08:33:07 -07:00
Alexander Belyaev 8c9cfc9f5a [XLA:CPU] Legalize narrow float vector memory operations in CPU xtile.
CPU targets without specific ISA extensions (such as AVX512-BF16) lack legal vector register classes for narrow floats (e.g., bf16, fp16). Using these types causes LLVM to scalarize them and generate libcalls.

PiperOrigin-RevId: 987420616
2026-09-25 08:16:45 -07:00
TensorFlower Gardener f9fab1c91d Merge pull request #127480 from gauravarun:fix-sparsesegment-grad-gpu-zero-size-oob
PiperOrigin-RevId: 987415547
2026-09-25 07:52:55 -07:00
Henning Becker 57dbb888ac [XLA] Use root native_custom_call_handler_allowlist in non-GPU XLA target visibility
Replace `"//xla/backends/gpu/libraries/native_custom_call_thunks:handler_allowlist"` in the `visibility` lists of core non-GPU XLA targets (`hlo`, `compiler`, `buffer_assignment`, `kernel_arguments`, `launch_dim`, `kernel_spec`) with `"//xla:native_custom_call_handler_allowlist"` and include `native_custom_call_thunks/...` in `native_custom_call_handler_allowlist` so non-GPU XLA packages do not reference a `package_group` inside `xla/backends/gpu/`.

PiperOrigin-RevId: 987404559
2026-09-25 07:34:29 -07:00
A. Unique TensorFlower 435c4e858e Automated Code Change
PiperOrigin-RevId: 987380727
2026-09-25 07:28:33 -07:00
linchen1-robot 6dc9dc2a5f PR #49033: [ROCm]Add ROCm HSACO compile-then-pack coverage and fix clang-offload-bundler target IDs.
Imported from GitHub PR https://github.com/openxla/xla/pull/49033

📝 Summary of Changes

- Add rocm-only //xla/stream_executor/gpu:gpu_hsaco_bundle_test: compile inlined AMDGPU LLVM IR with amdgpu::CompileToHsaco, then pack the HSACO with BundleGpuAsm(HsacoImage).
- Fix ROCm 7 clang-offload-bundler --targets= IDs in gpu/asm_compiler.cc (4-field host and HIP triples).
- Does not tag-drop //xla/stream_executor/cuda:subprocess_compilation_no_fakes_test (cuda-only). This is extra coverage, not a name-matched twin.

🎯 Justification
Add more unit test coverage on ROCm platform.

🚀 Kind of Contribution
🧪 Tests

Copybara import of the project:

--
981e93ac52fbeef919ccda3fdaf17cd7a2c6dd8f by Lin Chen1 <lin.chen1@amd.com>:

Add ROCm HSACO compile-then-pack coverage and fix clang-offload-bundler target IDs.

Signed-off-by: Lin Chen1 <lin.chen1@amd.com>

Merging this change closes #49033

PiperOrigin-RevId: 987378077
2026-09-25 07:12:43 -07:00
Levon Ter-Grigoryan d2a201ac0e [XLA:GPU] Include device failure details in PlatformUtil::GetStreamExecutors when no supported devices are found.
This simplifies debugging for example when VMM API is not available. See: https://github.com/openxla/xla/issues/49252

PiperOrigin-RevId: 987377773
2026-09-25 06:53:30 -07:00
TensorFlower Gardener c21d28ea07 Merge pull request #125856 from adi-IL:fix/lookup-table-export-signature-mismatch
PiperOrigin-RevId: 987377724
2026-09-25 06:48:04 -07:00
Christian Sigg f64587a591 Update TensorIR to v0.1.4.
PiperOrigin-RevId: 987373870
2026-09-25 06:17:48 -07:00
linchen1-robot 2dd5ba03ea PR #49372: [ROCm]Enable vectorization.hlo check on ROCm mi200 and mi350.
Imported from GitHub PR https://github.com/openxla/xla/pull/49372

📝 Summary of Changes
- Enable `//xla/backends/gpu/tests:vectorization.hlo.test` on mi200 and mi350. Same `gpu/` lit as CUDA. Removes those two specs from the disabled list.
- FileCheck LLVM `store <4 x i8>` at `--stage=llvm-before-optimizations`, the same stage and width as the H100/A100 lines.
- Gate `--stage=ptx` with `%if !IS_ROCM`. `AMDGPUCompiler::CompileTargetBinary` returns an empty asm text, and `hlo-opt` has no GCN stage. NVIDIA PTX checks stay.

🎯 Justification
Explain why this change is important and which workload benefits from this
change.

🚀 Kind of Contribution
🧪 Tests

Copybara import of the project:

--
9b122c87b8c8923be80148b641034d1a5330d3ad by linchen1 <lin.chen1@amd.com>:

Enable vectorization.hlo check on ROCm mi200 and mi350.

Signed-off-by: linchen1 <lin.chen1@amd.com>

Merging this change closes #49372

PiperOrigin-RevId: 987373576
2026-09-25 06:04:47 -07:00
dependabot[bot] 3445433c0b PR #49314: Bump astral-sh/setup-uv from 10.0.1 to 10.1.0
Imported from GitHub PR https://github.com/openxla/xla/pull/49314

Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 10.0.1 to 10.1.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a href="https://github.com/astral-sh/setup-uv/releases">astral-sh/setup-uv's releases</a>.</em></p>
<blockquote>
<h2>v10.1.0 🌈  New output <code>python-runtime-id</code>and respect NO_PROXY</h2>
<h2>Changes</h2>
<p>This release adds more bheind the scene security improvements and also 2 small improvements.</p>
<h3>NO_PROXY</h3>
<p>This action now respects <code>no_proxy/NO_PROXY</code> environment variables which were previously ignored.</p>
<h3>New output <code>python-runtime-id</code></h3>
<p>The new output <code>python-runtime-id</code> can be used to know which python version exactly was installed if you use <code>activate-environment</code>. See <a href="https://redirect.github.com/pyca/cryptography/pull/15572#discussion_r3913508686">pyca/cryptography#15572</a> for details on why this can be useful.</p>
<h2>🐛 Bug fixes</h2>
<ul>
<li>fix: respect no proxy directive <a href="https://github.com/mj0nez"><code>@​mj0nez</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1037">#1037</a>)</li>
<li>Use JSON + a typed wrapper instead of TS codegen <a href="https://github.com/woodruffw"><code>@​woodruffw</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1025">#1025</a>)</li>
</ul>
<h2>🚀 Enhancements</h2>
<ul>
<li>Expose a Python &quot;identity&quot; output <a href="https://github.com/woodruffw"><code>@​woodruffw</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1036">#1036</a>)</li>
<li>Verify downloads with astral-sh/versions checksums <a href="https://github.com/zaniebot"><code>@​zaniebot</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1033">#1033</a>)</li>
</ul>
<h2>🧰 Maintenance</h2>
<ul>
<li>chore: update known checksums for 0.12.12 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1041">#1041</a>)</li>
<li>chore: update known checksums for 0.12.10/0.12.11 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1038">#1038</a>)</li>
<li>chore: update known checksums for 0.12.9 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1035">#1035</a>)</li>
<li>chore: update known checksums for 0.12.7/0.12.8 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1031">#1031</a>)</li>
<li>chore: update known checksums for 0.12.6 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1030">#1030</a>)</li>
<li>chore: update known checksums for 0.12.5 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1020">#1020</a>)</li>
<li>Use self-repo syntax for all in-repo actions/reusable workflows <a href="https://github.com/woodruffw"><code>@​woodruffw</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1024">#1024</a>)</li>
<li>Pin one-shot tools <a href="https://github.com/woodruffw"><code>@​woodruffw</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1022">#1022</a>)</li>
<li>ci: remove obsolete direct push attempts <a href="https://github.com/eifinger"><code>@​eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1019">#1019</a>)</li>
</ul>
<h2>📚 Documentation</h2>
<ul>
<li>docs: update version references to v10.0.1 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1018">#1018</a>)</li>
</ul>
<h2>⬆️ Dependency updates</h2>
<ul>
<li>chore(deps-dev): roll up Dependabot updates <a href="https://github.com/eifinger"><code>@​eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1043">#1043</a>)</li>
<li>Harden npm install defaults <a href="https://github.com/zaniebot"><code>@​zaniebot</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1026">#1026</a>)</li>
<li>Add dependency cooldowns <a href="https://github.com/woodruffw"><code>@​woodruffw</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1021">#1021</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a href="https://github.com/astral-sh/setup-uv/commit/bec219d24cd3e171d82865faccec33120bb574f4"><code>bec219d</code></a> chore(deps-dev): roll up Dependabot updates (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1043">#1043</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/b90ec40d15bfa44c33c6700196eb6efcdddb4373"><code>b90ec40</code></a> fix: respect no proxy directive (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1037">#1037</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/421feb646df5262e7dd93bc54161edfa30372417"><code>421feb6</code></a> chore: update known checksums for 0.12.12 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1041">#1041</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/f634bf473ad85bf3e23a613f52c5fa9f363874fc"><code>f634bf4</code></a> Expose a Python &quot;identity&quot; output (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1036">#1036</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/a6772c8f0a09dc9e3582c70a994b0c55af921803"><code>a6772c8</code></a> chore: update known checksums for 0.12.10/0.12.11 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1038">#1038</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/e105c8fb1d7b13074b851babdaef4185243c6a07"><code>e105c8f</code></a> chore: update known checksums for 0.12.9 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1035">#1035</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/cd13f9217092d43a771cf9ba7b09bdd3da8d7c4d"><code>cd13f92</code></a> Verify downloads with astral-sh/versions checksums (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1033">#1033</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/3aef7b92c52cec135792ea1e95f4c77683d39e61"><code>3aef7b9</code></a> chore: update known checksums for 0.12.7/0.12.8 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1031">#1031</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/d08d816a1ea176d61a318eff45abd3dffef415b1"><code>d08d816</code></a> chore: update known checksums for 0.12.6 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1030">#1030</a>)</li>
<li><a href="https://github.com/astral-sh/setup-uv/commit/19b4d1e990bec64818914c40230bde93a0de300b"><code>19b4d1e</code></a> Harden npm install defaults (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1026">#1026</a>)</li>
<li>Additional commits viewable in <a href="https://github.com/astral-sh/setup-uv/compare/20cfd1bf945f4377ade1205e4dbc17946fc9a30d...bec219d24cd3e171d82865faccec33120bb574f4">compare view</a></li>
</ul>
</details>
<br />

[![Dependabot compatibility score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=astral-sh/setup-uv&package-manager=github_actions&previous-version=10.0.1&new-version=10.1.0)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

</details>
Copybara import of the project:

--
454adc369ab7b2a599e818b54fdb1ced4f464ea6 by dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>:

Bump astral-sh/setup-uv from 10.0.1 to 10.1.0

Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 10.0.1 to 10.1.0.
- [Release notes](https://github.com/astral-sh/setup-uv/releases)
- [Commits](https://github.com/astral-sh/setup-uv/compare/20cfd1bf945f4377ade1205e4dbc17946fc9a30d...bec219d24cd3e171d82865faccec33120bb574f4)

---
updated-dependencies:
- dependency-name: astral-sh/setup-uv
  dependency-version: 10.1.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

Merging this change closes #49314

PiperOrigin-RevId: 987366858
2026-09-25 06:01:11 -07:00
Levon Ter-Grigoryan f448a3eb39 [XLA:GPU] Add collective memory space to layout and use it as collective memory space color.
PiperOrigin-RevId: 987355696
2026-09-25 05:30:53 -07:00
Adrian Kuegel ee1d5b73a4 [XLA:GPU] Clamp per-dimension offsets in dynamic-slice analysis
HLO dynamic-slice and dynamic-update-slice clamp each dimension's start index
i to [0, dim_size[i] - slice_size[i]]. Previously, AnalyzeDynamicSlice and
VerifySliceOffset computed the 1D byte offset directly from unclamped start
indices, relying only on runtime 1D buffer clamping to [0, buffer_size -
slice_size]. When a slice is contiguous because all dimensions major to a
partially-sliced dimension have slice_size == 1, an out-of-bounds offset on a
minor dimension or size-1 dimension could stay within the total buffer bounds
and write or read at the wrong element offset.

Clamp constant offset operands per-dimension in AnalyzeDynamicSlice and
VerifySliceOffset, and for loop-dependent offsets verify that the runtime 1D
buffer clamping matches the per-dimension clamped HLO byte offset on every loop
iteration. Also preserve the sliced operand shape before walking through
bitcasts in DynamicSliceFusion::ResolveParameter so Parameter.shape matches the
slice rank and strides in ComputeSliceOffset.

Fixes #49120

PiperOrigin-RevId: 987353923
2026-09-25 05:19:13 -07:00
linchen1-robot 3312c76b6c PR #48977: [ROCm]Add new rocm_blas_lt_test for hipBLASLt device GEMM.
Imported from GitHub PR https://github.com/openxla/xla/pull/48977

📝 Summary of Changes

- Add //xla/stream_executor/rocm:rocm_blas_lt_test as the StreamExecutor twin of cuda_blas_lt_test.
- Keep cuda-only on the CUDA test. Use GpuPlatformName(), required F32 and S8S32 GEMM, skip S8S8F32 if hipBLASLt returns no algorithms.

🎯 Justification
Add more unit test coverage on ROCm platform.

🚀 Kind of Contribution
🧪 Tests

Copybara import of the project:

--
c9d526efeb59b64ace5a3d6b3d81dea2a38f2350 by Lin Chen1 <lin.chen1@amd.com>:

Add new rocm_blas_lt_test for hipBLASLt device GEMM.

Signed-off-by: Lin Chen1 <lin.chen1@amd.com>

--
7611d7aed1b43c18a91d4a45efbee51346b8215d by Lin Chen1 <lin.chen1@amd.com>:

fix clang-format error in rocm_blas_lt_test.

Signed-off-by: Lin Chen1 <lin.chen1@amd.com>

--
4e60098349f79f2403982075a07f624265573b28 by linchen1 <lin.chen1@amd.com>:

Add a trailing newline to rocm_blas_lt_test.cc.

Signed-off-by: linchen1 <lin.chen1@amd.com>

Merging this change closes #48977

PiperOrigin-RevId: 987322640
2026-09-25 05:15:51 -07:00
TensorFlower Gardener 9b379b2760 Merge pull request #126021 from vishwakt:fix-barrier-take-many-negative
PiperOrigin-RevId: 987297161
2026-09-25 04:54:53 -07:00
Ilya Tikhonovskiy 65c21117a1 [XLA:GPU] Move Triton scaled dot device tests to dedicated test file.
Extracts scaled dot tests from fusion_emitter_device_test into scaled_dot_device_test to isolate scaled dot test cases and reduce test compilation overhead.

PiperOrigin-RevId: 987286750
2026-09-25 04:37:37 -07:00
Henning Becker 9bb279da4c [XLA:Codegen] Remove unused GPU backend_configs_cc dependencies from codegen/tiling and codegen/xtile
Remove unused `#include "xla/service/gpu/backend_configs.proto.h"` includes and `//xla/service/gpu:backend_configs_cc` build dependencies from `//xla/codegen/tiling:symbolic_tile_analysis`, `//xla/codegen/tiling/experimental:tiled_hlo`, `//xla/codegen/xtile:block_level_parameters`, and `//xla/codegen/xtile:tiling_from_block_parameters` (left over after `BlockLevelFusionConfig` and `Tile` were migrated to `xtile_config.proto`).

PiperOrigin-RevId: 987284532
2026-09-25 04:32:36 -07:00
Henning Becker 420ff4c277 [XLA] Remove dead propagate_bitcast_layouts_to_backend_config option and GPU backend_configs_cc dep from mlir_hlo_to_hlo
`MlirToHloConversionOptions::propagate_bitcast_layouts_to_backend_config` defaults to `false` and is never set or referenced anywhere in the codebase. It was added as a temporary workaround for the legacy MHLO-based XLA:GPU elemental IR emitters (which have since been removed), and was the sole reason the hardware-independent `//xla/hlo/translate/mhlo_to_hlo:mlir_hlo_to_hlo` library included `xla/service/gpu/backend_configs.proto.h` (`xla::gpu::BitcastBackendConfig`) and depended on `//xla/service/gpu:backend_configs_cc`.

Remove the dead option, its unreachable `if` branch in `ExportXlaOperator(mhlo::BitcastOp)`, and the `//xla/service/gpu:backend_configs_cc` dependency.

PiperOrigin-RevId: 987284482
2026-09-25 04:03:24 -07:00
Henning Becker 435103ead5 [TF/XLA] Remove unused gpu_executable_run_options dependency from compiler/jit/kernels
Remove `"@xla//xla/service/gpu:gpu_executable_run_options"` from `XLA_OPS_DEPS` (`xla_ops` and `xla_ops_no_jit_rewrite_registration`) in `tensorflow/compiler/jit/kernels/BUILD`, as `GpuExecutableRunOptions` is not included or used by `xla_ops.{h,cc}`.

PiperOrigin-RevId: 987278014
2026-09-25 03:50:09 -07:00
A. Unique TensorFlower 019a0569e9 Automated Code Change
PiperOrigin-RevId: 987276596
2026-09-25 03:46:03 -07:00
A. Unique TensorFlower 3059318f88 Automated Code Change
PiperOrigin-RevId: 987275326
2026-09-25 03:26:39 -07:00
Dmitri Latushko 17c17ea4f0 Handle zero-sized inputs in tf.nn.lrn and its gradient op to prevent GPU crash.
PiperOrigin-RevId: 987255629
2026-09-25 03:08:22 -07:00
Junwhan Ahn ea22af286c Fix lifetime issues in ExecutionWatchdogScope async callbacks.
When `block_host_until_done` is false, `~ExecutionWatchdogScope()` releases the `HangWatchdog::Guard` asynchronously via `stream_->DoHostCallback(...)`, so the guard and its `on_timeout` and `pre_abort` callbacks outlive `ExecutionWatchdogScope` and `GpuExecutable::ExecuteThunks()`.

- Copy `ExecutionTimeoutHandler` by value in `ExecutionWatchdogScope::Arm()` instead of capturing a raw `const GpuExecutableRunOptions*` pointer that can dangle when callers allocate `GpuExecutableRunOptions` on the stack.
- Separate `ThunkExecutor::ProgressTracker` from `ThunkExecutor::ScopedProgressTracker` via `std::shared_ptr` so `pre_abort` in `GpuExecutable::ExecuteThunksImpl()` captures `tracker->tracker()` by value while `ScopedProgressTracker` remains a stack-local move-only RAII guard that uninstalls the thread-local pointer when leaving the dispatch scope.

PiperOrigin-RevId: 987251708
2026-09-25 03:03:00 -07:00