Bazel processes labels in the `load` statements in the beginning of `cuda_configure.bzl` file in parallel, hence CUDA repositories are created in parallel as well.
Previously the CUDA repositories were created sequentially in the order of `repository_ctx.read()` operations inside the `cuda_configure` rule implementation. Bazel creates repositories only when there are direct usages of them, and the list of required CUDA repositories was not known until the repository rule `cuda_configure` was executed.
PiperOrigin-RevId: 745809642
This change is made to prevent hermetic CUDA repositories cache invalidation between the builds with `--config=cuda` and without it.
It should speed up Github presubmit jobs. Currently CPU and GPU jobs use the machines in the same pool, and they share the RBE cache.
Previously the cache was invalidated every time when `TF_NEED_CUDA` value changed between CPU and GPU builds, hence loading CUDA redistributions for GPU jobs took several minutes (see [this job](https://github.com/openxla/xla/actions/runs/14114621736/job/39541688832) for example: all the test results are cached, but CUDA redistributions were still downloaded).
With adding `--repo_env USE_CUDA_REDISTRIBUTIONS=1` to RBE CPU linux job configurations, we load some CUDA redistributions once in RBE cache, and then reuse it between the jobs.
PiperOrigin-RevId: 743700713
Mostly updating `WORKSPACE` and related files from `@tsl//third_party` -> `@xla//third_party`. Please tag me if this change breaks you.
PiperOrigin-RevId: 733489153
CUDA/CUDNN/NCCL repositories are created on a host machine, so if we need to download redistributions for other architectures in cross-compile scenario, we need to pass the target architecture name to repo rules, e.g. `--repo_env=CUDA_REDIST_TARGET_PLATFORM=aarch64`
PiperOrigin-RevId: 720254498
Usage example: provide NVIDIA wheel dependencies for ML wheels that have rpaths pointing to NVIDIA folders. When a user executes `pip install tensorflow[and_cuda]`, NVIDIA wheels are installed together with Tensorflow wheel. To reproduce this behavior in hermetic Python approach, we need to define `py_import` as follows (provided NVIDIA dependencies are defined in `requirements.in` and requirements lock files):
py_import(
name = "tf_py_import",
wheel = ":wheel",
deps = [
"@pypi_absl_py//:pkg",
"@pypi_astunparse//:pkg",
"@pypi_flatbuffers//:pkg",
"@pypi_gast//:pkg",
"@pypi_ml_dtypes//:pkg",
"@pypi_numpy//:pkg",
"@pypi_opt_einsum//:pkg",
"@pypi_packaging//:pkg",
"@pypi_protobuf//:pkg",
"@pypi_requests//:pkg",
"@pypi_termcolor//:pkg",
"@pypi_typing_extensions//:pkg",
"@pypi_wrapt//:pkg",
],
wheel_deps = [
"@pypi_nvidia_cublas_cu12//:whl",
"@pypi_nvidia_cuda_cupti_cu12//:whl",
"@pypi_nvidia_cuda_nvcc_cu12//:whl",
"@pypi_nvidia_cuda_nvrtc_cu12//:whl",
"@pypi_nvidia_cuda_runtime_cu12//:whl",
"@pypi_nvidia_cudnn_cu12//:whl",
"@pypi_nvidia_cufft_cu12//:whl",
"@pypi_nvidia_curand_cu12//:whl",
"@pypi_nvidia_cusolver_cu12//:whl",
"@pypi_nvidia_cusparse_cu12//:whl",
"@pypi_nvidia_nccl_cu12//:whl",
"@pypi_nvidia_nvjitlink_cu12//:whl",
],
)
PiperOrigin-RevId: 705666137
The CUDA and NCCL repositories are created on a host machine now and shared via Bazel cache between host and remote machines.
PiperOrigin-RevId: 671089856
Add `--@local_config_cuda//cuda:override_include_cuda_libs` to override settings for TF wheel.
Forbid building TF wheel with `--@local_config_cuda//cuda:include_cuda_libs=true`
PiperOrigin-RevId: 666848518
1) Hermetic CUDA rules allow building wheels with GPU support on a machine without GPUs, as well as running Bazel GPU tests on a machine with only GPUs and NVIDIA driver installed. When `--config=cuda` is provided in Bazel options, Bazel will download CUDA, CUDNN and NCCL redistributions in the cache, and use them during build and test phases.
[Default location of CUNN redistributions](https://developer.download.nvidia.com/compute/cudnn/redist/)
[Default location of CUDA redistributions](https://developer.download.nvidia.com/compute/cuda/redist/)
[Default location of NCCL redistributions](https://pypi.org/project/nvidia-nccl-cu12/#history)
2) To include hermetic CUDA rules in your project, add the following in the WORKSPACE of the downstream project dependent on XLA.
Note: use `@local_tsl` instead of `@tsl` in Tensorflow project.
```
load(
"@tsl//third_party/gpus/cuda/hermetic:cuda_json_init_repository.bzl",
"cuda_json_init_repository",
)
cuda_json_init_repository()
load(
"@cuda_redist_json//:distributions.bzl",
"CUDA_REDISTRIBUTIONS",
"CUDNN_REDISTRIBUTIONS",
)
load(
"@tsl//third_party/gpus/cuda/hermetic:cuda_redist_init_repositories.bzl",
"cuda_redist_init_repositories",
"cudnn_redist_init_repository",
)
cuda_redist_init_repositories(
cuda_redistributions = CUDA_REDISTRIBUTIONS,
)
cudnn_redist_init_repository(
cudnn_redistributions = CUDNN_REDISTRIBUTIONS,
)
load(
"@tsl//third_party/gpus/cuda/hermetic:cuda_configure.bzl",
"cuda_configure",
)
cuda_configure(name = "local_config_cuda")
load(
"@tsl//third_party/nccl/hermetic:nccl_redist_init_repository.bzl",
"nccl_redist_init_repository",
)
nccl_redist_init_repository()
load(
"@tsl//third_party/nccl/hermetic:nccl_configure.bzl",
"nccl_configure",
)
nccl_configure(name = "local_config_nccl")
```
PiperOrigin-RevId: 662981325
619575611 by A. Unique TensorFlower<gardener@tensorflow.org>:
Run buildifier on all files where it sorts loads differently
--
619498661 by A. Unique TensorFlower<gardener@tensorflow.org>:
[XLA:GPU][IndexAnalysis] Rename GetDefaultThreadIdToOutputIndexingMap to GetDefaultThreadIdIndexingMap.
The "output" part was a bit confusing. We use this function for threadId->input
mapping as well.
--
619490165 by A. Unique TensorFlower<gardener@tensorflow.org>:
Convert S8 to BF16 in one step without going though F32.
--
PiperOrigin-RevId: 619575611
The current macro substitution wasn't working and is resolving to %{nccl_version} rather than a version string. Instead, add code to create a NCCL config header (nccl_config.h), and use that to detect the .so version we should be opening.
This change should only affect NCCL when obtained via a stub, which is currently only used by an unreleased version of JAX.
PiperOrigin-RevId: 571081917
This replaces the "transfer script via cli argument" hack by the upload support for remote configurations that landed in Bazel 3.1.0.
PiperOrigin-RevId: 569079584
NCCL via a stub is enabled if the environment variable TF_NCCL_USE_STUB=1 is set.
The intent is to use this option in JAX to reduce the size of the CUDA wheels. NCCL takes up about 80MB in each compiled JAX wheel and takes a significant amount of time to build. It is possible other users of TSL may wish to do this also.
PiperOrigin-RevId: 568344701
This is fixing a UB issue which occurs with newer version of Clang (17+).
The fix is also upstreamed through https://github.com/NVIDIA/nccl/pull/916.
In addition I'm changing the handling of `enqueue.cc` which needs to be compiled
in cuda mode under clang. The previous solution with just passing in the `-x cuda` option fails with CUDA 12+.
I'm also correcting the version number that we set in the patch - not sure if this version is reported in some logs, but if it is, it should be correct.
PiperOrigin-RevId: 564811002
Imported from GitHub PR https://github.com/tensorflow/tensorflow/pull/59825
This PR is a POC demonstrating building TensorFlow so that it depends on NVIDIA's wheel-based library releases for the required CUDA libraries and runtime. Users will still need to install the CUDA driver separately, but all other NVIDIA library dependencies should be fetched automatically by pip.
For example, after building the wheel with the usual OSS build scripts (based on cuda 11.8), the resulting wheel could be installed in a fairly minimalist ubuntu 20.04 environment as follows:
```
docker run --rm -it --gpus all -v /path/to/tensorflow/pkg:/tf/pkg nvidia/cuda:12.0.1-base-ubuntu20.04
apt update && apt install -y --no-install-recommends curl python3 python3-pip
python3 -mpip install --upgrade pip
pip3 install /tf/pkg/tensorflow-*.whl
```
Notice that the CUDA driver's forward compatibility allows users to use the latest (here CUDA 12) base container/driver, while pip pulls in the necessary CUDA 11.8-based libraries as needed.
Currently, NVIDIA wheels are released only for linux platforms and the x86_64 architecture.
Copybara import of the project:
--
ff44df38cd6514f72a01de418003994d39f43e30 by Nathan Luehr <nluehr@nvidia.com>:
Depend on nvidia-pyindex packages
--
87bf0f71d4c806dc1fa65bb95b71d69d9299aa97 by Trevor Morris <tmorris@nvidia.com>:
Fix rpath to standalone nccl
While libtensorflow_cc.so has all of the rpaths to the nvidia standalone
libraries, _pywrap_tensorflow_internal.so only has those from the dependencies
listed here.
For whatever reason, nccl needs to be in pywrap_tensorflow_internal's rpaths,
so I've added a new dependency on a dummy target which just adds the rpaths.
--
885bd7d31ad0e35814da8f0590a9cf898afdf7f7 by Nathan Luehr <nluehr@nvidia.com>:
Install nv lib wheels based on 'with-cuda' extra dependency.
--
afdb0902f2fb521668e52c0b015a217055e72516 by Nathan Luehr <nluehr@nvidia.com>:
Use 'and-cuda' rather than 'with-cuda' as modifier
--
03262d48daf2fad4f36a013e29a6098aea304a4f by Nathan Luehr <nluehr@nvidia.com>:
Move rpath flag construction into helper function.
--
c915c60d727f84565d846017d22a5cccf4e54ccb by Nathan Luehr <nluehr@nvidia.com>:
Add explanatory comment for nvcc paths.
Merging this change closes#59825
COPYBARA_INTEGRATE_REVIEW=https://github.com/tensorflow/tensorflow/pull/59825 from nluehr:nv-wheel-deps c915c60d727f84565d846017d22a5cccf4e54ccb
PiperOrigin-RevId: 544733780
496953709 by A. Unique TensorFlower<gardener@tensorflow.org>:
496952678 by A. Unique TensorFlower<gardener@tensorflow.org>:
Fix iOS nightly release build. Add profiler.h in TensorFlowLiteC framework headers.
--
496952616 by A. Unique TensorFlower<gardener@tensorflow.org>:
Convert _placeholder_value into a public API.
--
496949861 by A. Unique TensorFlower<gardener@tensorflow.org>:
[XLA] Make HloModuleConfig own the string keys in `analysis_allowance_map_`
This fixes a use-after-free bug when deserializing an HloModuleConfig:
https://github.com/tensorflow/tensorflow/blob/f1251be0983f36616bcf6d3fa896d01dcbf82fdb/tensorflow/compiler/xla/service/hlo_module_config.cc#L304-L306
Without this change, the keys of `analysis_allowance_map_` have the
same lifetime as the input proto that's being deserialized, which is
probably very inconvenient semantics for the caller. In practice there
are not many keys and they're short, so I don't think the extra string
copy will have a noticeable performance impact.
--
496949804 by A. Unique TensorFlower<gardener@tensorflow.org>:
Update the JAX's docker image to base it on the new image that TF has.
--
496947227 by A. Unique TensorFlower<gardener@tensorflow.org>:
Remove unused imports.
--
496944423 by A. Unique TensorFlower<gardener@tensorflow.org>:
Bypass nvprune if compiling with CUDA Clang. Retry
--
496941078 by A. Unique TensorFlower<gardener@tensorflow.org>:
496940772 by A. Unique TensorFlower<gardener@tensorflow.org>:
Add `ParseFromString` method to `OpSharding` Python bindings.
--
496940632 by A. Unique TensorFlower<gardener@tensorflow.org>:
Remove PAT from Scorecards workflow
This should make the workflow work again. Will test after this lands.
Signed-off-by: Mihai Maruseac <mihaimaruseac@google.com>
--
496939937 by A. Unique TensorFlower<gardener@tensorflow.org>:
Automated visibility attribute cleanup.
--
496939490 by A. Unique TensorFlower<gardener@tensorflow.org>:
Remove all PTX except for the specified one.
--
496931612 by A. Unique TensorFlower<gardener@tensorflow.org>:
Integrate StableHLO at openxla/stablehlo@e8c1c04
--
496919498 by A. Unique TensorFlower<gardener@tensorflow.org>:
[JITRT] Add a regression_test for broadcasting.
This is a reduced version of broadcasting_2, which is much easier to
read/debug.
--
496918085 by A. Unique TensorFlower<gardener@tensorflow.org>:
Add RBE to linux builds to improve test times.
--
496914744 by A. Unique TensorFlower<gardener@tensorflow.org>:
Add serialization support to FeatureSpace.
--
496911817 by A. Unique TensorFlower<gardener@tensorflow.org>:
[IFRT] Remove BUILD file tests for IFRT, which is always enabled now.
--
496911684 by A. Unique TensorFlower<gardener@tensorflow.org>:
[GmlSt] Implement `reifyResultShapes` for `gml_st.materialize`.
--
496906503 by A. Unique TensorFlower<gardener@tensorflow.org>:
PR #58850: Update compat.py
Imported from GitHub PR https://github.com/tensorflow/tensorflow/pull/58850
Fixed the broken link of tf.compat.forward_compatible.
Copybara import of the project:
--
4589648939 by Tirumalesh <111861663+tiruk007@users.noreply.github.com>:
Update compat.py
Fixed the broken link of (tf.compat.forward_compatible)[https://www.tensorflow.org/api_docs/python/tf/compat/forward_compatible]
Merging this change closes#58850
--
496904759 by A. Unique TensorFlower<gardener@tensorflow.org>:
496895256 by A. Unique TensorFlower<gardener@tensorflow.org>:
[DelegatePerformance] Split the targets out into subpackages.
--
496893002 by A. Unique TensorFlower<gardener@tensorflow.org>:
Remove unused `tf-jitrt-symbolic-shape-optimization` pass
--
496887677 by A. Unique TensorFlower<gardener@tensorflow.org>:
[GmlSt] Tile map and fill ops again in a peeled loop for scalarization.
The perfectly-tiled part of the loop remains unchanged and will be vectorized. Ops in the peeled loop should be scalarized later.
--
496885897 by A. Unique TensorFlower<gardener@tensorflow.org>:
[MLIR:XLA] Move shape.bcast simplification into the MHLO's shape optimization pass
Also, remove the old shape optimization pass from the pipeline.
--
496880934 by A. Unique TensorFlower<gardener@tensorflow.org>:
[DelegatePerformance] Handle failures from parsing latency results.
--
496878266 by A. Unique TensorFlower<gardener@tensorflow.org>:
Internal fixes.
--
496866713 by A. Unique TensorFlower<gardener@tensorflow.org>:
Add failing tests for XlaCallModule.
The newly added tests are failing at the moment and are disabled.
Added here to help debugging.
--
496864860 by A. Unique TensorFlower<gardener@tensorflow.org>:
[XLA:GPU] [NFC] Remove unused function
--
496859249 by A. Unique TensorFlower<gardener@tensorflow.org>:
Cleaning up BUILD files to remove "loose" headers.
--
PiperOrigin-RevId: 496953709
To compile flatbuffers definition Bazel starts the flatbuffers compiler
flatc which is potentially built using a custom
toolchain and hence requires a set up LD_LIBRARY_PATH.
Ommitting the `use_default_shell_env` (defaulting to false) clears the
whole environment and the binary may try to use older system libs such
as /lib64/libstdc++.so causing it to fail in case it is (much) older
than the used libstdc++ from the custom toolchain which is very common
in HPC environments.
Hence I added `use_default_shell_env = True` as already done in e.g.
`_local_genrule_impl`.