C++ Installation#
This page covers building and installing the cuNLS C++/CUDA library
(libcunls) for use from C++ applications.
Note
Using cuNLS from Python? Install the pycunls package instead; see
Installation. It does not require a separate C++
installation.
Prerequisites#
CUDA Toolkit (with
nvcc,cudart,cuBLAS,cuSPARSE,cuSOLVER)CMake >= 3.22
C++17 compiler
GNU Make (for the provided scripts)
NVIDIA GPU driver compatible with your CUDA Toolkit
Build locally#
From the repository root:
./scripts/build_cunls.sh <build_dir> <CMAKE_BUILD_TYPE = Release | Coverage> [install_dir]
Example (release build + install):
./scripts/build_cunls.sh build Release /tmp/cunls_install
By default, this builds a shared library (libcunls.so). Set the
EXTRA_CMAKE_ARGS environment variable to override CMake options, for example
to build a static library:
EXTRA_CMAKE_ARGS='-DBUILD_SHARED_LIBS=OFF' ./scripts/build_cunls.sh build Release /tmp/cunls_install
When an install_dir is provided, headers and the library binary are
installed there.
Build with Docker#
Install NVIDIA Container Toolkit:
https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.htmlRun:
./scripts/build_cunls_in_docker.sh <CMAKE_BUILD_TYPE = Release | Coverage> [local_install_dir]
The Docker build produces both shared and static library variants. Intermediate build directories live inside the container and are discarded automatically; only the final install directory is mounted to the host.
Install artifacts (default build_docker/, or the specified directory):
<install_dir>/
include/cunls/ # headers
lib/
libcunls.so # shared library
libcunls.a # static library (with bundled deps)
cmake/cunls/ # CMake package config
Direct CMake build (manual path)#
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/tmp/cunls_install
cmake --build build -j
cmake --install build
Pass -DBUILD_SHARED_LIBS=OFF to build a static library instead of a shared
one.
CUDA target and Arm platform selection#
cuNLS preserves caller-provided CUDA compiler and architecture settings.
Without one, /usr/local/cuda/bin/nvcc is used when it exists. A native
Jetson Orin build can compile an SM 87 cubin and avoid PTX JIT compatibility
requirements:
cmake -S . -B build -DCMAKE_CUDA_ARCHITECTURES=87-real
The cuDSS archive (CUDSS_PLATFORM=auto) follows the CUDA Toolkit target:
linux-x86_64, linux-aarch64 for Jetson toolkits, or linux-sbsa for
Arm server toolkits and CUDA 13 on Jetson. Toolkits installed without a
targets/ directory (e.g. distro packages under /usr) fall back to
linux-x86_64 on x86_64 and require an explicit platform on aarch64. Set
-DCUDSS_PLATFORM explicitly to override it.
Pass -DCUNLS_ENABLE_CUDSS=OFF to build without the cuDSS backend: the cuDSS
archive is not downloaded and SparseLinearSolverType::cuDSS throws
std::runtime_error when a solver is created. All other solvers work as
usual, and the C++ tests skip the cuDSS cases.
Notes#
BUILD_SHARED_LIBSdefaults toON(shared library). Set toOFFfor a static archive with bundled dependencies.BUILD_TESTINGis off by default.ENABLE_PROFILING=ONadds NVTX support.cuDSS is an optional runtime dependency. cunls never links against
libcudss.so; instead it loads it withdlopen()the first timeSparseLinearSolverType::cuDSSis used (seecmake/AddCUDSS.cmakeandcunls/common/cudss_dynamic.h). This means:Neither
libcunls.so/libcunls.anor thepycunlswheel bundle or depend on cuDSS at load time — every other solver (DenseLDLT,DenseCholesky,DenseQR,BlockSparsePCG,BlockTridiagonal) works with no cuDSS installed at all.To use the cuDSS solver, download a cuDSS release matching your CUDA major version from NVIDIA’s cuDSS redistributables (or install it via your platform’s package manager), then add the directory containing
libcudss.sotoLD_LIBRARY_PATHin the environment the process is launched with — e.g.export LD_LIBRARY_PATH=...in the shell before running your executable orpython. The dynamic loader readsLD_LIBRARY_PATHonce at process startup, so setting it from inside an already-running process (for example viaos.environ[...]in a live Python interpreter, before constructing acuDSSLinearSolver) has no effect ondlopen().Requesting the cuDSS solver without
libcudss.soreachable raises astd::runtime_error(Python:RuntimeError) describing how to fix it, rather than a linker error or crash.CI still downloads cuDSS at build time (for headers) and publishes its shared/static libraries as a separate
cudss-*artifact, distinct from thecunls-*andpycunls-*artifacts.
build_cunls.shsupports two environment variables for advanced use:CUNLS_SOURCE_DIR— override the CMake source directory (defaults to the parent of the build directory).EXTRA_CMAKE_ARGS— pass additional flags to the CMake configure step (e.g.-DBUILD_SHARED_LIBS=OFF).