cuNLS Documentation#
cuNLS provides GPU-accelerated nonlinear least-squares optimization for
batched geometric estimation problems. It is used primarily from Python
through the pycunls package (CuPy arrays, custom kernels with NVIDIA
Warp), and also offers a native C++/CUDA API.
Important
Capacity vs. active count. Factor and state batches are constructed with
their capacity (how many factors / states their buffers hold) and
start with zero active entries: call set_num_active_factors(n) /
set_num_active_states(n) (C++: SetNumActiveFactors /
SetNumActiveStates) before solving, and again whenever the problem
size changes. See Capacity and active count.