When while loops are unrolled the same thunk is called with `Record` multiple times for different loop iterations. Currently `CustomCallRecordState` stored a single sequence of `XLA_FFI_Command*`, so during update only the last recorded sequence of commands was passed to the FFI record API, overwriting the last iteration and leaving earlier iterations of the while loop with stale buffer pointers.
Keying the state by the sink command pointer returned on `RecordCreate` (and passed back via `RecordUpdate::command`) so each unrolled iteration updates its own recorded command sequence.
PiperOrigin-RevId: 987596254
- In WhileUtil::AppendToWhileLoopOriginalValue and WhileUtil::MakeInstructionsLiveIn, update while instruction, while init, body root, body parameter, and condition parameter to keep all loop instruction OriginalValues structurally consistent when widening while loops, and preserve call_hierarchy.
- In WhileLoopSimplifier (TryRemoveDeadTupleElements, TryRemoveConstantTupleElements, TryFlattenNestedTuples, and TrySimplifyInductionVariables), propagate and structurally update OriginalValues across all 5 loop boundary instructions and verify compatibility with the new shapes.
PiperOrigin-RevId: 987569790
Goldens move to `executables/collective_ops_aot_test_2gpu_<arch>/`, one
directory per backend.
gb300 stays excluded through the existing `disabled_backends` entry
(b/491194726); no gb300 target is generated, so `DetectGpuArchToken()` has no
gb300 branch and a gb300 device would hit its `LOG(FATAL)`.
PiperOrigin-RevId: 987552652
Provides a script to automate regenerating and promoting AOT compatibility golden snapshots from Bazel undeclared test outputs into `<package>/executables/<target_name>/v<N+1>/`.
Also updates `AOTInterceptionPjrtClient::VerifyAgainstGolden` and `test_lib.cc` (`GetUndeclaredOutputsDir`) to reference `update_goldens.py` in golden comparison failure and precondition error messages. The golden
comparison failure message prints the failing test's full label (from
`TEST_TARGET`), which is what the script takes as its argument.
PiperOrigin-RevId: 987522513
`Future<std::vector<T>>({})` cannot select the templated converting
constructors of `Future`, because a braced-init-list cannot deduce `U`.
Overload resolution therefore falls back to the move constructor from a
value-initialized `Future`, and both typed `JoinFutures` overloads returned an
invalid future for an empty span instead of a ready empty vector. Calling
`Map`, `OnReady` or `Await` on it dereferences a null `AsyncValue`.
PiperOrigin-RevId: 987497834
`serialized_executable` is already a snappy-compressed split
`GpuExecutableProto` (see `WriteSplitGpuExecutable`). Compressing the outer
container again saves no space (dropping it adds 0.06% on a 330 MB
executable), but it costs an extra compression pass on write and a full
decompression pass on load.
Riegeli stores the compression type per chunk, so readers don't change and
existing artifacts still load.
Alternative considered: leave the inner proto uncompressed and compress the
outer one instead. That would mean deciding per use case when to skip
compression, which adds logic for little gain, since the outer proto is
small apart from `serialized_executable`.
PiperOrigin-RevId: 987497408
`HandleProtoMergeRecord` read each record into a flat `absl::string_view` and
then called `MergeFromString`. For records larger than a riegeli buffer (e.g.
an HLO module carrying hundreds of MB of constants) the decoded record spans
several blocks, and producing a flat view forces riegeli to copy it into a
contiguous scratch buffer first.
Read the record as a `riegeli::Chain` instead, which shares the decoded
blocks, and parse it with `riegeli::ParseMessage(..., set_merge(true))`,
which consumes the chain directly.
PiperOrigin-RevId: 987496597
That will be used later to add constraints to one indexing map. Should also marginally imporove compilation time.
Also call simplify in tiling propagation only if we are not in "symbolic" case where we will not use them simplified anyway.
PiperOrigin-RevId: 987492063
Now that we have this compatibility window, we'll often need to ask OSS developers to change their PRs to avoid breaking it (e.g. https://github.com/openxla/xla/pull/46865). So we should have some documentation that we can point to.
PiperOrigin-RevId: 987466715
AlgebraicSimplifier can rewrite bitcast-convert instructions into bitcast even
when the element bitwidth changes (when the minor-most dimension is contiguous).
While PriorityFusion and GpuFusible guard against fusing width-changing bitcasts
using hlo_instruction_utils::KeepsBitwidth, CopyFusion previously skipped
through any single-user bitcast without checking its bitwidth.
Fusing a width-changing bitcast and its consuming copy into a producer fusion
breaks the MLIR fusion emitters because GetBitcastMap requires matching element
counts between roots and elemental lowering cannot emit a scalar
mhlo.bitcast_convert across different bitwidths.
Guard kBitcast checks in CopyFusion with hlo_instruction_utils::KeepsBitwidth.
PiperOrigin-RevId: 987463670
Adds cudnn_scaled_dot_test covering block-scaled dot support in cuDNN on Blackwell. Parameterized across MXFP8 and NVFP4 type combinations and transposition layouts, reporting lowering stages (composite, optimized HLO, cuDNN graph representation) and execution parity against reference decomposition.
PiperOrigin-RevId: 987447766
Add pre-checks for P2P symmetric memory support via cudaDeviceCanAccessPeer,
post-checks for null windows returned from ncclCommWindowRegister, and
defensive null window checks in multimem_addr, peer_addr, and the destructor.
PiperOrigin-RevId: 987435182
CPU targets without specific ISA extensions (such as AVX512-BF16) lack legal vector register classes for narrow floats (e.g., bf16, fp16). Using these types causes LLVM to scalarize them and generate libcalls.
PiperOrigin-RevId: 987420616
Replace `"//xla/backends/gpu/libraries/native_custom_call_thunks:handler_allowlist"` in the `visibility` lists of core non-GPU XLA targets (`hlo`, `compiler`, `buffer_assignment`, `kernel_arguments`, `launch_dim`, `kernel_spec`) with `"//xla:native_custom_call_handler_allowlist"` and include `native_custom_call_thunks/...` in `native_custom_call_handler_allowlist` so non-GPU XLA packages do not reference a `package_group` inside `xla/backends/gpu/`.
PiperOrigin-RevId: 987404559
Imported from GitHub PR https://github.com/openxla/xla/pull/49033📝 Summary of Changes
- Add rocm-only //xla/stream_executor/gpu:gpu_hsaco_bundle_test: compile inlined AMDGPU LLVM IR with amdgpu::CompileToHsaco, then pack the HSACO with BundleGpuAsm(HsacoImage).
- Fix ROCm 7 clang-offload-bundler --targets= IDs in gpu/asm_compiler.cc (4-field host and HIP triples).
- Does not tag-drop //xla/stream_executor/cuda:subprocess_compilation_no_fakes_test (cuda-only). This is extra coverage, not a name-matched twin.
🎯 Justification
Add more unit test coverage on ROCm platform.
🚀 Kind of Contribution
🧪 Tests
Copybara import of the project:
--
981e93ac52fbeef919ccda3fdaf17cd7a2c6dd8f by Lin Chen1 <lin.chen1@amd.com>:
Add ROCm HSACO compile-then-pack coverage and fix clang-offload-bundler target IDs.
Signed-off-by: Lin Chen1 <lin.chen1@amd.com>
Merging this change closes#49033
PiperOrigin-RevId: 987378077
Imported from GitHub PR https://github.com/openxla/xla/pull/49372📝 Summary of Changes
- Enable `//xla/backends/gpu/tests:vectorization.hlo.test` on mi200 and mi350. Same `gpu/` lit as CUDA. Removes those two specs from the disabled list.
- FileCheck LLVM `store <4 x i8>` at `--stage=llvm-before-optimizations`, the same stage and width as the H100/A100 lines.
- Gate `--stage=ptx` with `%if !IS_ROCM`. `AMDGPUCompiler::CompileTargetBinary` returns an empty asm text, and `hlo-opt` has no GCN stage. NVIDIA PTX checks stay.
🎯 Justification
Explain why this change is important and which workload benefits from this
change.
🚀 Kind of Contribution
🧪 Tests
Copybara import of the project:
--
9b122c87b8c8923be80148b641034d1a5330d3ad by linchen1 <lin.chen1@amd.com>:
Enable vectorization.hlo check on ROCm mi200 and mi350.
Signed-off-by: linchen1 <lin.chen1@amd.com>
Merging this change closes#49372
PiperOrigin-RevId: 987373576
HLO dynamic-slice and dynamic-update-slice clamp each dimension's start index
i to [0, dim_size[i] - slice_size[i]]. Previously, AnalyzeDynamicSlice and
VerifySliceOffset computed the 1D byte offset directly from unclamped start
indices, relying only on runtime 1D buffer clamping to [0, buffer_size -
slice_size]. When a slice is contiguous because all dimensions major to a
partially-sliced dimension have slice_size == 1, an out-of-bounds offset on a
minor dimension or size-1 dimension could stay within the total buffer bounds
and write or read at the wrong element offset.
Clamp constant offset operands per-dimension in AnalyzeDynamicSlice and
VerifySliceOffset, and for loop-dependent offsets verify that the runtime 1D
buffer clamping matches the per-dimension clamped HLO byte offset on every loop
iteration. Also preserve the sliced operand shape before walking through
bitcasts in DynamicSliceFusion::ResolveParameter so Parameter.shape matches the
slice rank and strides in ComputeSliceOffset.
Fixes#49120
PiperOrigin-RevId: 987353923
Imported from GitHub PR https://github.com/openxla/xla/pull/48977📝 Summary of Changes
- Add //xla/stream_executor/rocm:rocm_blas_lt_test as the StreamExecutor twin of cuda_blas_lt_test.
- Keep cuda-only on the CUDA test. Use GpuPlatformName(), required F32 and S8S32 GEMM, skip S8S8F32 if hipBLASLt returns no algorithms.
🎯 Justification
Add more unit test coverage on ROCm platform.
🚀 Kind of Contribution
🧪 Tests
Copybara import of the project:
--
c9d526efeb59b64ace5a3d6b3d81dea2a38f2350 by Lin Chen1 <lin.chen1@amd.com>:
Add new rocm_blas_lt_test for hipBLASLt device GEMM.
Signed-off-by: Lin Chen1 <lin.chen1@amd.com>
--
7611d7aed1b43c18a91d4a45efbee51346b8215d by Lin Chen1 <lin.chen1@amd.com>:
fix clang-format error in rocm_blas_lt_test.
Signed-off-by: Lin Chen1 <lin.chen1@amd.com>
--
4e60098349f79f2403982075a07f624265573b28 by linchen1 <lin.chen1@amd.com>:
Add a trailing newline to rocm_blas_lt_test.cc.
Signed-off-by: linchen1 <lin.chen1@amd.com>
Merging this change closes#48977
PiperOrigin-RevId: 987322640
Extracts scaled dot tests from fusion_emitter_device_test into scaled_dot_device_test to isolate scaled dot test cases and reduce test compilation overhead.
PiperOrigin-RevId: 987286750
Remove unused `#include "xla/service/gpu/backend_configs.proto.h"` includes and `//xla/service/gpu:backend_configs_cc` build dependencies from `//xla/codegen/tiling:symbolic_tile_analysis`, `//xla/codegen/tiling/experimental:tiled_hlo`, `//xla/codegen/xtile:block_level_parameters`, and `//xla/codegen/xtile:tiling_from_block_parameters` (left over after `BlockLevelFusionConfig` and `Tile` were migrated to `xtile_config.proto`).
PiperOrigin-RevId: 987284532
`MlirToHloConversionOptions::propagate_bitcast_layouts_to_backend_config` defaults to `false` and is never set or referenced anywhere in the codebase. It was added as a temporary workaround for the legacy MHLO-based XLA:GPU elemental IR emitters (which have since been removed), and was the sole reason the hardware-independent `//xla/hlo/translate/mhlo_to_hlo:mlir_hlo_to_hlo` library included `xla/service/gpu/backend_configs.proto.h` (`xla::gpu::BitcastBackendConfig`) and depended on `//xla/service/gpu:backend_configs_cc`.
Remove the dead option, its unreachable `if` branch in `ExportXlaOperator(mhlo::BitcastOp)`, and the `//xla/service/gpu:backend_configs_cc` dependency.
PiperOrigin-RevId: 987284482
Remove `"@xla//xla/service/gpu:gpu_executable_run_options"` from `XLA_OPS_DEPS` (`xla_ops` and `xla_ops_no_jit_rewrite_registration`) in `tensorflow/compiler/jit/kernels/BUILD`, as `GpuExecutableRunOptions` is not included or used by `xla_ops.{h,cc}`.
PiperOrigin-RevId: 987278014
When `block_host_until_done` is false, `~ExecutionWatchdogScope()` releases the `HangWatchdog::Guard` asynchronously via `stream_->DoHostCallback(...)`, so the guard and its `on_timeout` and `pre_abort` callbacks outlive `ExecutionWatchdogScope` and `GpuExecutable::ExecuteThunks()`.
- Copy `ExecutionTimeoutHandler` by value in `ExecutionWatchdogScope::Arm()` instead of capturing a raw `const GpuExecutableRunOptions*` pointer that can dangle when callers allocate `GpuExecutableRunOptions` on the stack.
- Separate `ThunkExecutor::ProgressTracker` from `ThunkExecutor::ScopedProgressTracker` via `std::shared_ptr` so `pre_abort` in `GpuExecutable::ExecuteThunksImpl()` captures `tracker->tracker()` by value while `ScopedProgressTracker` remains a stack-local move-only RAII guard that uninstalls the thread-local pointer when leaving the dispatch scope.
PiperOrigin-RevId: 987251708
`tf.linalg.band_part` instead of raising `InvalidArgumentError`, so that eager
execution matches the documented definition of the op and the XLA lowering.
Fixes#110316
PiperOrigin-RevId: 987133747
Compiling GPU convolutions invokes cuDNN v9 engine heuristics
(cudnnBackendGetAttribute -> traceback_api_add), which together with XLA GPU
compilation frames requires ~280+ KiB of stack space and overflows the previous
240 KiB stack size (breaking jax_api_pjrt_gpu_test after cl/986842563).
PiperOrigin-RevId: 987097843