Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
7fc73b3
cuda.core: require a per-major cuda-bindings floor at build and run time
Andy-Jost Sep 18, 2026
f4757aa
cuda.core: fence C++ on the CUDA major only, never on CUDA_VERSION
Andy-Jost Sep 18, 2026
b9ada0d
cuda.core: call the driver through cuda-bindings' resolved pointers (…
Andy-Jost Sep 18, 2026
4955393
cuda.core: gate features on the driver alone now that cuda-bindings h…
Andy-Jost Sep 18, 2026
65c63b0
cuda.core tests: gate the checkpoint helper tests on the CUDA 13 buil…
Andy-Jost Sep 18, 2026
cd1c42d
Merge remote-tracking branch 'origin/main' into ajost/core-bindings-r…
Andy-Jost Sep 18, 2026
af88337
cuda.core: regenerate stubs and drop exec from the floor tool and tes…
Andy-Jost Sep 18, 2026
b4fedec
cuda.core: import cuda.bindings for the build check through the #1824…
Andy-Jost Sep 18, 2026
b000b54
Merge remote-tracking branch 'origin/main' into ajost/core-bindings-r…
Andy-Jost Sep 22, 2026
b06cc7c
cuda.core: read linked LTOIR through cynvjitlink instead of probing t…
Andy-Jost Sep 22, 2026
ab980a9
cuda.core: declare the cuda-bindings floor once, in the pyproject ext…
Andy-Jost Sep 22, 2026
18c5a34
cuda.core: compare headers, not version strings, in the cuda-bindings…
Andy-Jost Sep 22, 2026
194f9ea
cuda.core: latch a failed driver-table fill, keep pending exceptions,…
Andy-Jost Sep 22, 2026
a37c10a
cuda.core: keep the NVML constant's stub annotation-only; finish the …
Andy-Jost Sep 22, 2026
6ba9004
cuda-bindings header check at build; final floor messages; bounds spe…
Andy-Jost Sep 24, 2026
b6cbcc0
Copyedit the docs, comments, docstrings and messages this PR adds (#2…
Andy-Jost Sep 24, 2026
1f05f42
Merge remote-tracking branch 'origin/main' into ajost/core-bindings-r…
Andy-Jost Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .github/workflows/test-wheel-linux.yml
Original file line number Diff line number Diff line change
Expand Up @@ -248,14 +248,14 @@ jobs:
fi

- name: Display structure of downloaded cuda-python artifacts
if: ${{ env.TEST_PYTHON == 'true' && env.BINDINGS_SOURCE != 'published' }}
if: ${{ env.TEST_PYTHON == 'true' && env.BINDINGS_SOURCE != 'floor' }}
run: |
pwd
ls -lah cuda_python*.whl cuda_pathfinder/

- name: Display structure of downloaded cuda.bindings artifacts
if: ${{ (env.TEST_BINDINGS == 'true' || env.TEST_CORE == 'true' || env.TEST_PYTHON == 'true') &&
env.BINDINGS_SOURCE != 'published' }}
env.BINDINGS_SOURCE != 'floor' }}
run: |
pwd
ls -lahR $CUDA_BINDINGS_ARTIFACTS_DIR
Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/test-wheel-windows.yml
Original file line number Diff line number Diff line change
Expand Up @@ -228,14 +228,14 @@ jobs:
fi

- name: Display structure of downloaded cuda-python artifacts
if: ${{ env.TEST_PYTHON == 'true' && env.BINDINGS_SOURCE != 'published' }}
if: ${{ env.TEST_PYTHON == 'true' && env.BINDINGS_SOURCE != 'floor' }}
run: |
Get-Location
Get-ChildItem cuda_python*.whl | Select-Object Mode, LastWriteTime, Length, FullName

- name: Display structure of downloaded cuda.bindings artifacts
if: ${{ (env.TEST_BINDINGS == 'true' || env.TEST_CORE == 'true' || env.TEST_PYTHON == 'true') &&
env.BINDINGS_SOURCE != 'published' }}
env.BINDINGS_SOURCE != 'floor' }}
run: |
Get-Location
Get-ChildItem -Recurse -Force $env:CUDA_BINDINGS_ARTIFACTS_DIR | Select-Object Mode, LastWriteTime, Length, FullName
Expand Down
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,8 @@ cuda_bindings/cuda/bindings/utils/_get_handle.pyx

# Version files from setuptools_scm
_version.py
# Generated by cuda_core/build_hooks.py at build time (see cuda/core/__init__.py).
cuda_core/cuda/core/_build_info.py

# Distribution / packaging
.Python
Expand Down
9 changes: 9 additions & 0 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,15 @@ repos:
pass_filenames: false
verbose: true

- id: check-cuda-core-bindings-floor
name: Check the cuda-bindings floor against ci/versions.yml and the docs
entry: python ./toolshed/check_cuda_core_bindings_floor.py
language: python
files: ^(cuda_core/pyproject\.toml|ci/versions\.yml|cuda_core/cuda/core/_bindings_floor\.py|cuda_core/docs/source/.*\.rst|toolshed/check_cuda_core_bindings_floor\.py)$
pass_filenames: false
additional_dependencies:
- "tomli>=1.1.0; python_version < '3.11'"

- id: check-spdx
name: Check SPDX Headers
entry: python ./toolshed/check_spdx.py
Expand Down
73 changes: 73 additions & 0 deletions ci/tools/cuda_core_bindings_floor.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
#!/usr/bin/env python3
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
#
# SPDX-License-Identifier: Apache-2.0

"""Print the cuda-bindings floor of a cuda-core wheel for one CUDA major.

cuda_core_bindings_floor.py --wheel dist/cuda_core-*.whl --major 13
-> 13.4.1

CI installs `cuda-bindings==<floor>` next to a freshly built cuda-core wheel to
test the oldest cuda-bindings that wheel supports (BINDINGS_SOURCE=floor in
ci/tools/env-vars). This script reads the floor from the wheel under test, not
from the checkout, so a nightly job that tests a wheel built from another
commit reads that wheel's floor.

Each build records its floor in the generated cuda/core/_build_info.py. A
single-major build places the file at top level. The merged wheel places it
under cuda/core/cu<major>/. This script parses the CUDA_BINDINGS_FLOOR literal
out of it and never runs it.
"""

from __future__ import annotations

import argparse
import ast
import sys
import zipfile
from pathlib import Path

MODULE = "_build_info.py"


def _literal(source: str, name: str):
"""The literal that `source` assigns to `name` at module level. Parses the source and never runs it."""
for node in ast.parse(source, MODULE).body:
if isinstance(node, ast.AnnAssign):
targets = [node.target]
elif isinstance(node, ast.Assign):
targets = node.targets
else:
continue
if node.value is not None and any(isinstance(t, ast.Name) and t.id == name for t in targets):
return ast.literal_eval(node.value)
raise SystemExit(f"{MODULE} does not assign {name}")


def floor_from_source(source: str, major: int) -> str:
if _literal(source, "CUDA_MAJOR") != major:
raise SystemExit(f"{MODULE} records a CUDA {_literal(source, 'CUDA_MAJOR')} build, not CUDA {major}")
return ".".join(str(part) for part in _literal(source, "CUDA_BINDINGS_FLOOR"))


def floor_from_wheel(wheel: Path, major: int) -> str:
with zipfile.ZipFile(wheel) as zf:
names = set(zf.namelist())
for candidate in (f"cuda/core/cu{major}/{MODULE}", f"cuda/core/{MODULE}"):
if candidate in names:
return floor_from_source(zf.read(candidate).decode("utf-8"), major)
raise SystemExit(f"{wheel.name} contains no build for CUDA {major}: it has no {MODULE}. Is it a cuda-core wheel?")


def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
parser.add_argument("--wheel", type=Path, required=True, help="the cuda-core wheel under test")
parser.add_argument("--major", type=int, required=True, help="CUDA major series (12 or 13)")
args = parser.parse_args(argv)
print(floor_from_wheel(args.wheel, args.major))
return 0


if __name__ == "__main__":
sys.exit(main())
30 changes: 24 additions & 6 deletions ci/tools/env-vars
Original file line number Diff line number Diff line change
Expand Up @@ -64,23 +64,41 @@ elif [[ "${1}" == "test" ]]; then
# BINDINGS_SOURCE controls which cuda-bindings to install at test time:
# main — use the just-built bindings wheel from this CI run
# backport — fetch bindings from the prior (N-1) branch
# published — install from PyPI (cuda-bindings==${TEST_CUDA_MAJOR}.${TEST_CUDA_MINOR}.*)
# floor — install from PyPI the oldest cuda-bindings that the cuda-core
# wheel under test supports, its per-major floor (see
# cuda_core/cuda/core/_bindings_floor.py and ci/tools/run-tests).
# Selected when the test CTK minor differs from the minor the
# wheel was built against. Those rows exercise a new cuda-core
# with the floor cuda-bindings and older CTK libraries, the skew
# that cuda-core supports. cuda-bindings older than the floor is
# unsupported and fails at import
# (https://github.com/NVIDIA/cuda-python/issues/2783).
#
# SKIP_CUDA_BINDINGS_TEST / SKIP_CYTHON_TEST control which *tests* to run
# (they do NOT affect installation — that's BINDINGS_SOURCE's job).

BUILD_CUDA_MINOR="$(cut -d '.' -f 2 <<< ${BUILD_CUDA_VER})"
TEST_CUDA_MINOR="$(cut -d '.' -f 2 <<< ${CUDA_VER})"
# CI builds the prior-major half of the cuda-core wheel against the prev_build
# toolkit in ci/versions.yml and the backport branch's bindings.
BUILD_PREV_CUDA_VER="$(sed -n '/prev_build:/,/version:/s/.*version: *"\([^"]*\)".*/\1/p' ci/versions.yml)"
BUILD_PREV_CUDA_MINOR="$(cut -d '.' -f 2 <<< ${BUILD_PREV_CUDA_VER})"

if [[ ${BUILD_CUDA_MAJOR} != ${TEST_CUDA_MAJOR} ]]; then
# Major mismatch (e.g. build=13.x, test=12.x): use the backport branch.
BINDINGS_SOURCE=backport
SKIP_CUDA_BINDINGS_TEST=1
SKIP_CYTHON_TEST=1
if [[ ${BUILD_PREV_CUDA_MINOR} != ${TEST_CUDA_MINOR} ]]; then
# Prior major, minor mismatch (e.g. built against 12.9, test=12.6): floor
# bindings from PyPI with the older CTK libraries.
BINDINGS_SOURCE=floor
else
# Prior major, same minor (e.g. build=13.x, test=12.9): the backport branch.
BINDINGS_SOURCE=backport
fi
elif [[ ${BUILD_CUDA_MINOR} != ${TEST_CUDA_MINOR} ]]; then
# Same major, minor mismatch (e.g. build=13.2, test=13.0): use published
# bindings from PyPI to test the real-world backward-compat scenario.
BINDINGS_SOURCE=published
# Same major, minor mismatch (e.g. build=13.4, test=13.0): floor bindings
# from PyPI with the older CTK libraries.
BINDINGS_SOURCE=floor
SKIP_CUDA_BINDINGS_TEST=1
SKIP_CYTHON_TEST=1
else
Expand Down
10 changes: 7 additions & 3 deletions ci/tools/run-tests
Original file line number Diff line number Diff line change
Expand Up @@ -74,10 +74,14 @@ elif [[ "${test_module}" == "core" || "${test_module}" == nightly-* ]]; then

# Resolve bindings based on BINDINGS_SOURCE (set by env-vars):
# main/backport → local wheel from artifacts dir
# published → install from PyPI by version
# floor → the oldest cuda-bindings that the cuda-core wheel under test
# supports, read from that wheel and installed from PyPI
BINDINGS_ARGS=()
if [[ "${BINDINGS_SOURCE}" == "published" ]]; then
BINDINGS_ARGS+=("cuda-bindings==${TEST_CUDA_MAJOR}.${TEST_CUDA_MINOR}.*")
if [[ "${BINDINGS_SOURCE}" == "floor" ]]; then
CORE_WHL_FOR_FLOOR=("${CUDA_CORE_ARTIFACTS_DIR}"/*.whl)
BINDINGS_FLOOR="$(python ci/tools/cuda_core_bindings_floor.py --wheel "${CORE_WHL_FOR_FLOOR[0]}" --major "${TEST_CUDA_MAJOR}")"
echo "cuda-bindings floor of ${CORE_WHL_FOR_FLOOR[0]##*/} for CUDA ${TEST_CUDA_MAJOR}: ${BINDINGS_FLOOR}"
BINDINGS_ARGS+=("cuda-bindings==${BINDINGS_FLOOR}")
else
BINDINGS_ARGS=("${CUDA_BINDINGS_ARTIFACTS_DIR}"/*.whl)
if [[ "${LOCAL_CTK}" != 1 ]]; then
Expand Down
82 changes: 82 additions & 0 deletions ci/tools/tests/test_cuda_core_bindings_floor.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
#
# SPDX-License-Identifier: Apache-2.0

import importlib.util
import zipfile
from pathlib import Path

import pytest

TOOLS = Path(__file__).resolve().parent.parent


def _load(name, path):
spec = importlib.util.spec_from_file_location(name, path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module


tool = _load("cuda_core_bindings_floor", TOOLS / "cuda_core_bindings_floor.py")

FLOORS = {12: (12, 9, 8), 13: (13, 4, 1)}


def _build_info(major):
floor = FLOORS[major]
return (
"# Generated by build_hooks.py at build time. Do not edit or commit.\n"
f"CUDA_MAJOR = {major}\n"
f"CUDA_VERSION = {major * 1000 + floor[1] * 10} # the cuda.h this build compiled against\n"
f"CUDA_BINDINGS_FLOOR = {floor!r}\n"
f"CUDA_BINDINGS_BUILD_VERSION = '{major}.{floor[1]}.{floor[2]}'\n"
)


def _wheel(tmp_path, entries):
path = tmp_path / "cuda_core-1.3.0-cp312-cp312-linux_x86_64.whl"
with zipfile.ZipFile(path, "w") as zf:
for name, major in entries.items():
zf.writestr(name, _build_info(major))
return path


@pytest.mark.agent_authored(model="claude-fable-5-1")
@pytest.mark.parametrize("major", [12, 13])
def test_reads_the_merged_wheel_layout(tmp_path, major):
wheel = _wheel(tmp_path, {"cuda/core/cu12/_build_info.py": 12, "cuda/core/cu13/_build_info.py": 13})
assert tool.floor_from_wheel(wheel, major) == ".".join(map(str, FLOORS[major]))


@pytest.mark.agent_authored(model="claude-fable-5-1")
def test_reads_a_single_major_wheel(tmp_path):
wheel = _wheel(tmp_path, {"cuda/core/_build_info.py": 13})
assert tool.floor_from_wheel(wheel, 13) == "13.4.1"


@pytest.mark.agent_authored(model="claude-fable-5-1")
def test_rejects_a_single_major_wheel_of_another_major(tmp_path):
wheel = _wheel(tmp_path, {"cuda/core/_build_info.py": 13})
with pytest.raises(SystemExit, match="records a CUDA 13 build, not CUDA 12"):
tool.floor_from_wheel(wheel, 12)


@pytest.mark.agent_authored(model="claude-fable-5-1")
def test_rejects_a_wheel_without_the_build_record(tmp_path):
wheel = _wheel(tmp_path, {})
with pytest.raises(SystemExit, match="contains no build for CUDA 13"):
tool.floor_from_wheel(wheel, 13)


@pytest.mark.agent_authored(model="claude-fable-5-1")
def test_rejects_a_build_record_without_the_floor():
with pytest.raises(SystemExit, match="does not assign CUDA_BINDINGS_FLOOR"):
tool.floor_from_source("CUDA_MAJOR = 13\n", 13)


@pytest.mark.agent_authored(model="claude-fable-5-1")
def test_cli_prints_the_floor(tmp_path, capsys):
wheel = _wheel(tmp_path, {"cuda/core/_build_info.py": 13})
assert tool.main(["--wheel", str(wheel), "--major", "13"]) == 0
assert capsys.readouterr().out.strip() == "13.4.1"
4 changes: 4 additions & 0 deletions cuda_bindings/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,10 @@ the `legacy_tests` subdirectory.

- `CUDA_HOME` or `CUDA_PATH` must point to a valid CUDA Toolkit for source
builds.
- The toolkit's `cuda.h` must have the same major.minor as the generated
sources, `CUDA_VERSION` in `cuda/bindings/cydriver.pxd`. `build_hooks.py`
checks this before cythonize and fails with a message that names both
versions.
- `CUDA_PYTHON_PARALLEL_LEVEL` controls build parallelism.
- Runtime behavior is affected by
`CUDA_PYTHON_CUDA_PER_THREAD_DEFAULT_STREAM` and
Expand Down
78 changes: 78 additions & 0 deletions cuda_bindings/build_hooks.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@
import functools
import glob
import os
import re
import shutil
import sys
import sysconfig
Expand All @@ -33,6 +34,13 @@
# Populated by _build_cuda_bindings(); consumed by setup.py.
_extensions = None

# The generated sources declare the types and functions of one CUDA header set,
# and cydriver.pxd records which. The install docs state the build rule that results.
_CYDRIVER_PXD = Path(__file__).resolve().parent / "cuda" / "bindings" / "cydriver.pxd"
_GENERATED_VERSION_RE = re.compile(r"^cdef enum:\s*CUDA_VERSION\s*=\s*(\d+)\s*$")
_CUDA_H_VERSION_RE = re.compile(r"^#\s*define\s+CUDA_VERSION\s+(\d+)\s*$")
_INSTALL_URL = "https://nvidia.github.io/cuda-python/cuda-bindings/latest/install.html#installing-from-source"


# Please keep in sync with the copy in cuda_core/build_hooks.py.
def _import_get_cuda_path_or_home():
Expand Down Expand Up @@ -80,6 +88,75 @@ def _get_cuda_path() -> str:
return cuda_path


# -----------------------------------------------------------------------
# CUDA header check


def _cuda_h_path(cuda_path: str) -> str:
"""The cuda.h under cuda_path, with symlinks such as /usr/local/cuda resolved for messages."""
return os.path.realpath(os.path.join(cuda_path, "include", "cuda.h"))


def _read_version_macro(path: str, pattern: re.Pattern) -> int | None:
"""The integer on the first line of ``path`` that matches ``pattern``, or None if no line does."""
with open(path, encoding="utf-8") as f:
for line in f:
m = pattern.match(line)
if m:
return int(m.group(1))
return None


def _read_cuda_h_version(cuda_path: str) -> int:
"""The CUDA_VERSION macro of the cuda.h under cuda_path, for example 13040 for 13.4."""
cuda_h = _cuda_h_path(cuda_path)
try:
version = _read_version_macro(cuda_h, _CUDA_H_VERSION_RE)
except OSError:
version = None
if version is None:
raise RuntimeError(
f"Cannot read CUDA_VERSION from {cuda_h}. "
"Ensure CUDA_PATH or CUDA_HOME points to a CUDA Toolkit with include/cuda.h."
)
return version


def _generated_cuda_version() -> int:
"""The CUDA_VERSION of the headers that this source tree was generated from.

Read from cuda/bindings/cydriver.pxd.
"""
version = _read_version_macro(str(_CYDRIVER_PXD), _GENERATED_VERSION_RE)
if version is None:
raise RuntimeError(f"Cannot read CUDA_VERSION from {_CYDRIVER_PXD}")
return version


def _major_minor(cuda_version: int) -> str:
"""13040 -> \"13.4\"."""
return f"{cuda_version // 1000}.{cuda_version // 10 % 100}"


def _check_cuda_headers(cuda_path: str) -> None:
"""Reject a toolkit whose cuda.h is not the major.minor that this source tree was generated from.

Against another minor, the C++ compile fails with a long list of
redefinition and undeclared-type errors that do not name the cause. See
https://github.com/NVIDIA/cuda-python/issues/2783. Runs before cythonize,
which is the first step that touches the source tree.
"""
generated = _generated_cuda_version()
needed, found = _major_minor(generated), _major_minor(_read_cuda_h_version(cuda_path))
if found != needed:
raise RuntimeError(
f"This cuda-bindings source tree needs CUDA {needed} headers, but {_cuda_h_path(cuda_path)} is "
f"CUDA {found}. This is a build-time requirement only: at run time cuda-bindings supports any "

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
f"CUDA {found}. This is a build-time requirement only: at run time cuda-bindings supports any "
f"CUDA {found}. This is a build-time requirement only: at run time this cuda-bindings build can be used with any "

Main concern: "supports" seems a tad too strong, e.g. I'd want to stay clear of leading someone to think that we're somehow supporting features added in toolkits with a minor version newer than the cuda-bindings minor version.

Minor concern: make "this cuda-bindings build" specific.

f"CUDA {generated // 1000}.x toolkit, see {_INSTALL_URL}. Point CUDA_PATH or CUDA_HOME at a "
f"CUDA {needed} toolkit, or build from cuda-bindings {found}.x sources."
)


# -----------------------------------------------------------------------
# Extension preparation helpers

Expand Down Expand Up @@ -142,6 +219,7 @@ def _build_cuda_bindings(debug=False):
global _extensions

cuda_path = _get_cuda_path()
_check_cuda_headers(cuda_path)

if os.environ.get("PARALLEL_LEVEL") is not None:
warn(
Expand Down
Loading
Loading