Factor API#

The factor module provides batched residual and Jacobian models for non-linear least squares. Each factor computes a residual vector from one or more states. States lie on manifolds; the factor API uses tangent-space dimensions for Jacobian layout and solver variables. This page documents the Python factor classes first, then the C++ API, which holds the full residual and Jacobian formulas. In the C++ API, links to the corresponding state batch types are in the Factor inputs section and in each factor’s Inputs subsection.

  • Python — pycunls

  • C++ — cunls/factor

Python API (pycunls)#

All Python factor batches inherit from the abstract FactorBatch base class. Every constructor argument documented as DevicePointer accepts either a cupy.ndarray (the device pointer is extracted automatically via .data.ptr) or a raw int GPU device address.

The residual formulas, Jacobian structure, and state layouts are identical to the C++ versions documented in C++ API below (each Python entry links to its C++ counterpart) — this section focuses on the Python constructor signatures, methods, and properties.

Important

Capacity vs. active count. Factor and state batches are constructed with their capacity (how many factors / states their buffers hold) and start with zero active entries: call set_num_active_factors(n) / set_num_active_states(n) (C++: SetNumActiveFactors / SetNumActiveStates) before solving, and again whenever the problem size changes. See Capacity and active count.

Common FactorBatch interface#

Every factor batch — built-in or user-defined — exposes the following read-only properties and methods.

Read-only properties

  • num_active_factors (int) — number of active factors (0 until set_num_active_factors; at most capacity).

  • capacity (int) — number of factors the measurement buffers hold (the constructor’s capacity); constant.

  • residuals_size (int) — residual dimension per factor (e.g. 2 for ReprojectionFactorBatch, 6 for SE3BetweenFactorBatch).

Methods

  • set_num_active_factors(num_active_factors) — sets the active factor count (at most capacity). Every batch starts with 0 active factors: call it before the first solve, and again whenever the count changes. Host-only; takes effect at the next minimize. Raises ValueError above the capacity.

  • state_sizes() -> list[int] — returns a list of tangent-space dimensions for each state consumed by one factor. For example, ReprojectionFactorBatch returns [6, 3] (SE(3) pose then \(\mathbb{R}^3\) point), and PnPFactorBatch returns [6] (pose only; 3-D points are fixed in the constructor).

Device buffers. Constructor and state-batch arguments typed DevicePointer take a CuPy array or a raw integer pointer. CuPy arrays are checked against the dtype the kernels read — float32 for data and measurements, int32 for ids and indices, uint64 for state-pointer tables — and a mismatch raises TypeError (cp.arange is int64, NumPy and CuPy default to float64). Raw integer pointers are not checked.

C++ reference: FactorBatch.

pycunls.PriorVectorFactorBatch1 / PriorVectorFactorBatch2 / PriorVectorFactorBatch3 / PriorVectorFactorBatch6#

Prior on a Euclidean vector. Residual = \(x - o\) with identity Jacobian. The suffix indicates the dimension.

Constructor

fb = pycunls.PriorVectorFactorBatch3(observations, capacity)
  • observations (DevicePointer) — contiguous GPU buffer of capacity × Dim floats holding the observed (target) vectors. The factor batch does not copy the data; the caller must keep the allocation alive.

  • capacity (int) — number of prior factors the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from the corresponding VectorStateBatch (see pycunls.VectorStateBatch1 / VectorStateBatch2 / VectorStateBatch3 / VectorStateBatch6).

C++ reference: PriorVectorFactorBatch<Dim>.

pycunls.SO2PriorFactorBatch#

Prior on a 2-D rotation. Residual = \(\mathrm{Log}(R_\mathrm{target}^\top R)\).

Constructor

fb = pycunls.SO2PriorFactorBatch(observations, capacity)
  • observations (DevicePointer) — capacity × 4 floats holding row-major 2×2 target rotation matrices.

  • capacity (int) — number of prior factors the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from SO2StateBatch (see pycunls.SE3StateBatch).

C++ reference: SO2PriorFactorBatch.

pycunls.SO3PriorFactorBatch#

Prior on a 3-D rotation. Residual = \(\mathrm{Log}(R_\mathrm{target}^\top R)\), Jacobian = \(J_r^{-1}(r)\).

Constructor

fb = pycunls.SO3PriorFactorBatch(observations, capacity)
  • observations (DevicePointer) — capacity × 9 floats holding row-major 3×3 target rotation matrices.

  • capacity (int) — number of prior factors the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from SO3StateBatch.

C++ reference: SO3PriorFactorBatch.

pycunls.SE3PriorFactorBatch#

Prior on a 3-D rigid transform. Residual = \(\mathrm{Log}(T_\mathrm{target}^{-1} T)\), Jacobian = \(J_r^{-1}(r)\).

Constructor

fb = pycunls.SE3PriorFactorBatch(observations, capacity)
  • observations (DevicePointer) — capacity × 16 floats holding row-major 4×4 target homogeneous matrices.

  • capacity (int) — number of prior factors the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from SE3StateBatch.

C++ reference: SE3PriorFactorBatch.

pycunls.SE2PriorFactorBatch#

Prior on a 2-D rigid transform. Residual = \(\mathrm{Log}(T_\mathrm{target}^{-1} T)\) (3-vector), Jacobian = \(J_r^{-1}(r)\).

Constructor

fb = pycunls.SE2PriorFactorBatch(observations, capacity)
  • observations (DevicePointer) — capacity × 9 floats holding row-major 3×3 target transforms.

  • capacity (int) — number of prior factors the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from SE2StateBatch (see pycunls.SE3StateBatch).

C++ reference: SE2PriorFactorBatch.

pycunls.Similarity2PriorFactorBatch#

Prior on a 2-D similarity. Residual = \(\mathrm{Log}(T_\mathrm{target}^{-1} T)\) (4-vector), Jacobian = \(J_r^{-1}(r)\).

Constructor

fb = pycunls.Similarity2PriorFactorBatch(observations, capacity)
  • observations (DevicePointer) — capacity × 9 floats holding row-major 3×3 target transforms.

  • capacity (int) — number of prior factors the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from Similarity2StateBatch (see pycunls.SE3StateBatch).

C++ reference: Similarity2PriorFactorBatch.

pycunls.Similarity3PriorFactorBatch#

Prior on a 3-D similarity. Residual = \(\mathrm{Log}(T_\mathrm{target}^{-1} T)\) (7-vector), Jacobian = \(J_r^{-1}(r)\).

Constructor

fb = pycunls.Similarity3PriorFactorBatch(observations, capacity)
  • observations (DevicePointer) — capacity × 16 floats holding row-major 4×4 target transforms.

  • capacity (int) — number of prior factors the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from Similarity3StateBatch (see pycunls.SE3StateBatch).

C++ reference: Similarity3PriorFactorBatch.

pycunls.SL4PriorFactorBatch#

Prior on an SL(4) transform. Residual = \(\mathrm{Log}(T_\mathrm{target}^{-1} T)\).

Constructor

fb = pycunls.SL4PriorFactorBatch(observations, capacity)
  • observations (DevicePointer) — capacity × 16 floats holding row-major 4×4 SL(4) target transforms.

  • capacity (int) — number of prior factors the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from SL4StateBatch.

C++ reference: SL4PriorFactorBatch.

pycunls.SE3BetweenFactorBatch#

Constrains the relative pose between two SE(3) frames. Residual = \(\mathrm{Log}(\Delta\, T_l^{-1} T_r)\), zero at \(\Delta = T_r^{-1} T_l\). Two states per factor.

Constructor

fb = pycunls.SE3BetweenFactorBatch(deltas, capacity)
  • deltas (DevicePointer) — capacity × 16 floats holding row-major 4×4 measured relative transforms \(\Delta\).

  • capacity (int) — number of between constraints the buffers hold; 0 are active until set_num_active_factors.

State layout: two states per factor — [T_left, T_right] — both from SE3StateBatch. The state-pointer list must therefore contain 2 × num_active_factors entries.

C++ reference: SE3BetweenFactorBatch.

pycunls.SE2BetweenFactorBatch#

Constrains the relative transform between two SE(2) frames. Residual = \(\mathrm{Log}(\Delta\, T_l^{-1} T_r)\), zero at \(\Delta = T_r^{-1} T_l\). Two states per factor.

Constructor

fb = pycunls.SE2BetweenFactorBatch(deltas, capacity)
  • deltas (DevicePointer) — capacity × 9 floats holding row-major 3×3 measured relative transforms \(\Delta\).

  • capacity (int) — number of between constraints the buffers hold; 0 are active until set_num_active_factors.

State layout: two states per factor — [T_left, T_right] — both from SE2StateBatch.

C++ reference: SE2BetweenFactorBatch.

pycunls.SO2BetweenFactorBatch#

Constrains the relative rotation between two SO(2) frames. Residual = \(\mathrm{Log}(R_l^\top R_r \Delta)\) (zero at \(R_r = R_l \Delta^\top\); the same convention as SO3BetweenFactorBatch). Two states per factor.

Constructor

fb = pycunls.SO2BetweenFactorBatch(deltas, capacity)
  • deltas (DevicePointer) — capacity × 4 floats holding row-major 2×2 measured relative rotations \(\Delta\).

  • capacity (int) — number of between constraints the buffers hold; 0 are active until set_num_active_factors.

State layout: two states per factor — [R_left, R_right] — both from SO2StateBatch.

C++ reference: SO2BetweenFactorBatch.

pycunls.SO3BetweenFactorBatch#

Constrains the relative rotation between two SO(3) frames. Residual = \(\mathrm{Log}(R_l^\top R_r \Delta)\) (zero at \(R_r = R_l \Delta^\top\); the same convention as SO2BetweenFactorBatch). Two states per factor.

Constructor

fb = pycunls.SO3BetweenFactorBatch(deltas, capacity)
  • deltas (DevicePointer) — capacity × 9 floats holding row-major 3×3 measured relative rotations \(\Delta\).

  • capacity (int) — number of between constraints the buffers hold; 0 are active until set_num_active_factors.

State layout: two states per factor — [R_left, R_right] — both from SO3StateBatch.

C++ reference: SO3BetweenFactorBatch.

pycunls.Similarity2BetweenFactorBatch#

Constrains the relative transform between two Sim(2) frames. Residual = \(\mathrm{Log}(\Delta\, T_l^{-1} T_r)\), zero at \(\Delta = T_r^{-1} T_l\). Two states per factor.

Constructor

fb = pycunls.Similarity2BetweenFactorBatch(deltas, capacity)
  • deltas (DevicePointer) — capacity × 9 floats holding row-major 3×3 measured relative transforms \(\Delta\).

  • capacity (int) — number of between constraints the buffers hold; 0 are active until set_num_active_factors.

State layout: two states per factor — [T_left, T_right] — both from Similarity2StateBatch.

C++ reference: Similarity2BetweenFactorBatch.

pycunls.Similarity3BetweenFactorBatch#

Constrains the relative transform between two Sim(3) frames. Residual = \(\mathrm{Log}(\Delta\, T_l^{-1} T_r)\), zero at \(\Delta = T_r^{-1} T_l\). Two states per factor.

Constructor

fb = pycunls.Similarity3BetweenFactorBatch(deltas, capacity)
  • deltas (DevicePointer) — capacity × 16 floats holding row-major 4×4 measured relative transforms \(\Delta\).

  • capacity (int) — number of between constraints the buffers hold; 0 are active until set_num_active_factors.

State layout: two states per factor — [T_left, T_right] — both from Similarity3StateBatch.

C++ reference: Similarity3BetweenFactorBatch.

pycunls.SL4BetweenFactorBatch#

Constrains the relative transform between two SL(4) frames. Residual = \(\mathrm{Log}(\Delta\, T_l^{-1} T_r)\), zero at \(\Delta = T_r^{-1} T_l\). Two states per factor.

Constructor

fb = pycunls.SL4BetweenFactorBatch(deltas, capacity)
  • deltas (DevicePointer) — capacity × 16 floats holding row-major 4×4 measured relative transforms \(\Delta\) (unit determinant).

  • capacity (int) — number of between constraints the buffers hold; 0 are active until set_num_active_factors.

State layout: two states per factor — [T_left, T_right] — both from SL4StateBatch.

C++ reference: SL4BetweenFactorBatch.

pycunls.VectorBetweenFactorBatch1 / VectorBetweenFactorBatch2 / VectorBetweenFactorBatch3 / VectorBetweenFactorBatch6#

Between factor on Euclidean vectors. Residual = \(x_l - x_r - \delta\). Two states per factor.

Constructor

fb = pycunls.VectorBetweenFactorBatch3(deltas, capacity)
  • deltas (DevicePointer) — capacity × Dim floats holding the measured difference vectors \(\delta\).

  • capacity (int) — number of between constraints the buffers hold; 0 are active until set_num_active_factors.

State layout: two states per factor from the corresponding VectorStateBatch (see pycunls.VectorStateBatch1 / VectorStateBatch2 / VectorStateBatch3 / VectorStateBatch6).

C++ reference: VectorBetweenFactorBatch<Dim>.

pycunls.ReprojectionFactorBatch#

Reprojection error for bundle adjustment. Observations must be in normalized image coordinates (intrinsic calibration already applied): \(r = \pi(T, P) - z\).

Constructor

fb = pycunls.ReprojectionFactorBatch(
    observations, capacity, z_threshold=1e-3)
  • observations (DevicePointer) — capacity × 2 floats holding normalized 2-D observations \((x_n, y_n)\).

  • capacity (int) — number of reprojection factors the buffers hold; 0 are active until set_num_active_factors.

  • z_threshold (float, default 1e-3) — minimum valid depth \(z\) in camera frame. Points with \(z < z_\text{threshold}\) produce zero residuals and Jacobians to avoid singularities.

State layout: two states per factor — [SE3 pose, R^3 point] — from SE3StateBatch and VectorStateBatch3 respectively. The pose is world_from_rig, the rig’s pose in the world (pose-convention). This binding has no camera_from_rig argument: the camera is the rig (identity extrinsic), \(P_{\mathrm{cam}} = T^{-1} P\).

C++ reference: ReprojectionFactorBatch.

pycunls.PnPFactorBatch#

PnP-style reprojection: fixed 3-D points in the constructor, one SE(3) state per correspondence (typically the same camera pose pointer repeated), world_from_rig (pose-convention).

Constructor (identity camera-from-rig)

fb = pycunls.PnPFactorBatch(
    observations, points_world, capacity, z_threshold=1e-3)

Constructor (with camera-from-rig extrinsics per factor)

fb = pycunls.PnPFactorBatch(
    observations, poses_camera_from_rig, points_world,
    capacity, z_threshold=1e-3)
  • observations — capacity × 2 normalized image coordinates.

  • points_world — capacity × 3 fixed world points (not optimized).

  • poses_camera_from_rig — capacity × 16 row-major SE(3) matrices (optional overload).

  • z_threshold — minimum valid depth in the camera frame (same role as ReprojectionFactorBatch).

State layout: one SE3StateBatch state per factor from SE3StateBatch.

C++ reference: PnPFactorBatch.

pycunls.ImuFactorBatch#

IMU factor between two keyframes. The raw samples between them are integrated and the intermediate states eliminated inside the factor at every evaluation (Schur complement). Theory, conventions and an example: IMU factor.

Constructor

p = pycunls.ImuParameters()
fb = pycunls.ImuFactorBatch(imu_samples, sample_offsets, num_samples, p, capacity)
  • imu_samples (DevicePointer, float32) — num_samples × 7: gyroscope [rad/s], specific force [m/s²] (IMU frame) and the step duration [s] of every sample, all factors back to back.

  • sample_offsets (DevicePointer, int32) — capacity + 1 CSR offsets: factor f uses samples [offsets[f], offsets[f + 1]), at least one (a factor without samples evaluates to zero rows).

  • num_samples (int) — samples in imu_samples. Only sizes the work split (num_samples / capacity is taken as the typical chain length).

  • parameters (ImuParameters) — gravity (3 floats, default [0, 0, -9.80665]), gyro_noise_density, accel_noise_density, integration_noise_density, gyro_bias_random_walk, accel_bias_random_walk (continuous-time, all positive; defaults: EuRoC ADIS16448) and body_from_imu (16 floats, row-major rig_from_imu, default identity).

  • capacity (int) — keyframe pairs the buffers hold; 0 are active until set_num_active_factors.

Arrays of another dtype raise TypeError; raw integer pointers are accepted unchecked.

State layout (per factor, in order): T_a (SE3StateBatch, world_from_rig, the pose state of the reprojection and PnP factors), v_a (VectorStateBatch3, world velocity of the IMU), b_a (VectorStateBatch6, [b_g; b_a]), T_b, v_b, b_b. state_sizes() returns [6, 3, 6, 6, 3, 6].

Residual (15): the whitened defect of keyframe b against the prediction (9), then the bias random walk (6).

C++ reference: ImuFactorBatch.

pycunls.PointToPointFactorBatch#

Point-to-point ICP factor. Residual = \(p - T q\).

Constructor

fb = pycunls.PointToPointFactorBatch(p_observations, q_observations, capacity)
  • p_observations (DevicePointer) — capacity × 3 floats holding target points \(p\).

  • q_observations (DevicePointer) — capacity × 3 floats holding source points \(q\).

  • capacity (int) — number of point correspondences the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from SE3StateBatch.

C++ reference: PointToPointFactorBatch.

pycunls.PointToPlaneFactorBatch#

Point-to-plane ICP factor. Residual = \(n_q^\top (p - T q)\): distance of the transformed source point from the plane through \(p\) with normal \(n_q\), given in the target frame (not rotated by \(T\)).

Constructor

fb = pycunls.PointToPlaneFactorBatch(
    p_observations, q_observations, nq_observations, capacity)
  • p_observations (DevicePointer) — target points (× 3 floats).

  • q_observations (DevicePointer) — source points (× 3 floats).

  • nq_observations (DevicePointer) — plane normals in the target frame (× 3 floats); a normal estimated at the source point must be rotated into the target frame first.

  • capacity (int) — number of correspondences the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from SE3StateBatch.

C++ reference: PointToPlaneFactorBatch.

pycunls.SymmetricPointToPlaneFactorBatch#

Symmetric point-to-plane ICP factor. Both frames contribute normals; \(N = n_p + n_q\).

Constructor

fb = pycunls.SymmetricPointToPlaneFactorBatch(
    p_observations, q_observations,
    np_observations, nq_observations, capacity)
  • p_observations (DevicePointer) — target points (× 3 floats).

  • q_observations (DevicePointer) — source points (× 3 floats).

  • np_observations (DevicePointer) — target normals (× 3 floats).

  • nq_observations (DevicePointer) — source normals (× 3 floats).

  • capacity (int) — number of correspondences the buffers hold; 0 are active until set_num_active_factors.

State layout: one state per factor from SE3StateBatch.

C++ reference: SymmetricPointToPlaneFactorBatch.

Motion priors#

Constant-velocity (CV) and constant-acceleration (CA) priors between consecutive poses, for SE(3), SO(3), SE(2) and SO(2) (ConstantVelocitySE3FactorBatch, ConstantAccelerationSO2FactorBatch, …). Residuals and state layout as in the C++ reference.

fb = pycunls.ConstantVelocitySE3FactorBatch(time_steps, capacity)
# Weighted by the closed-form process-noise information Q(dt)^-1:
fb = pycunls.ConstantVelocityInformationSE3FactorBatch(stream, time_steps, qc_diag, capacity)
fb.update(stream, n)  # after rewriting time_steps / qc_diag in place
  • time_steps (DevicePointer) — capacity float32 durations \(\Delta t_k = t_{k+1} - t_k\).

  • qc_diag (DevicePointer) — the continuous-time process-noise PSD diagonal, one float32 per tangent DOF (6 / 3 / 3 / 1), constant across the batch. Smaller values trust the motion model more.

  • stream (CudaStream) — the information matrices are computed on it at construction and by update.

State layout: CV [pose_k, pose_{k+1}, v_k, v_{k+1}], CA additionally a_k, a_{k+1}: poses from the matching Lie state batch (world_from_body, see pose-convention), velocities and accelerations (body frame) from VectorStateBatch6 / VectorStateBatch3 / VectorStateBatch1.

C++ reference: Motion prior factors.

pycunls.InformationFactorBatch#

Wraps any factor batch and left-multiplies residuals and Jacobians by per-factor square-root information matrices \(\Omega^{1/2}\). Unlike the C++ template, the Python class accepts any FactorBatch — no template specialization is needed.

C++ contrast: The C++ template InformationFactorBatch<T> also inherits T::sized_layout (a SizedFactorBatch with the same compile-time layout as T). The Python wrapper is a dynamic FactorBatch only.

Constructor

info_fb = pycunls.InformationFactorBatch(
    inner_factor, sqrt_information_matrices)
  • inner_factor (FactorBatch) — the factor batch to wrap. The wrapper delegates Evaluate to this factor first, then applies the information matrices. The inner factor must be kept alive for the lifetime of the wrapper.

  • sqrt_information_matrices (DevicePointer) — capacity × residual_size × residual_size contiguous floats holding one row-major square-root information matrix per factor.

Example

inner = pycunls.SE3BetweenFactorBatch(deltas, N)
info  = pycunls.InformationFactorBatch(inner, sqrt_info_gpu)
info.set_num_active_factors(N)  # forwarded to inner

problem.add_factor_batch(info, state_pointers)

C++ reference: InformationFactorBatch<T>.

pycunls.WeightedFactorBatch#

Wraps any factor batch and scales residuals and Jacobians by a scalar weight. Two construction modes are supported:

  1. Uniform weight (float) — the same scalar is applied to every factor.

  2. Per-factor weights (DevicePointer) — one weight per factor from a GPU array.

C++ contrast: WeightedFactorBatch<T> inherits T::sized_layout; the Python wrapper subclasses FactorBatch only.

Constructors

# Uniform weight
wfb = pycunls.WeightedFactorBatch(inner_factor, weight=2.0)

# Per-factor weights
wfb = pycunls.WeightedFactorBatch(inner_factor, weights=weights_gpu)
  • inner_factor (FactorBatch) — the factor batch to wrap.

  • weight (float) — uniform scalar weight applied to all factors.

  • weights (DevicePointer, keyword-only) — inner_factor.capacity contiguous floats, one weight per factor.

Exactly one of weight or weights must be provided.

Example

inner = pycunls.PriorVectorFactorBatch3(obs_gpu, N)
wfb   = pycunls.WeightedFactorBatch(inner, weight=5.0)
wfb.set_num_active_factors(N)  # forwarded to inner

problem.add_factor_batch(wfb, state_pointers)

C++ reference: WeightedFactorBatch<T>.

pycunls.CustomFactorBatch#

Base class for user-defined factors. Subclass this to implement a residual and Jacobian computation that is not available as a built-in factor.

Constructor

class MyFactor(pycunls.CustomFactorBatch):
    def __init__(self, capacity):
        super().__init__(
            residual_size=...,
            state_sizes=[...],
            capacity=capacity,
        )
  • residual_size (int) — dimension of the residual vector per factor.

  • state_sizes (Sequence[int]) — list of tangent-space dimensions for each state consumed by one factor (e.g. [1, 1] for a factor reading two scalar states).

  • capacity (int) — number of factor instances the buffers hold; 0 are active until set_num_active_factors.

Methods to override

  • evaluate(residuals_ptr, jacobians_ptr, state_pointers_ptr, stream_handle, factor_ids_ptr, num_factor_ids) -> bool — computes residuals and Jacobians on the GPU for n = num_factor_ids items (the same contract as C++ FactorBatch::Evaluate()): item t is factor f(t) evaluated at its own state pointers. All six arguments are raw int values:

    • residuals_ptr — device pointer to the output residual buffer. Layout: n × residual_size contiguous floats; item t writes row t.

    • jacobians_ptr — device pointer to the output Jacobian buffer. Layout: n × residual_size × sum(state_sizes) contiguous floats (row-major per item, blocks concatenated in state order). May be 0 (null) when the minimizer only needs residuals (e.g. for cost evaluation); in that case skip Jacobian writes.

    • state_pointers_ptr — device pointer to an array of float* pointers. The array has n × len(state_sizes) entries; item t’s state b is entry t * len(state_sizes) + b (the device address of that state’s ambient-space storage). Because Warp kernels cannot perform float** double-pointer indirection, custom factors typically gather state values into contiguous CuPy arrays before launching a kernel (see the Custom Warp Factor tutorial).

    • stream_handle — cudaStream_t cast to int. All GPU work must be launched on this stream.

    • factor_ids_ptr — device pointer to n int32 factor indices in [0, num_active_factors) with f(t) = factor_ids[t], or 0 (null) for f(t) = t % num_active_factors. Read measurements through f(t); everything else (state pointers, outputs) is indexed by t.

    • num_factor_ids — the item count n. Unlike C++, where 0 means num_active_factors, Python always receives the actual count (> 0).

    The regular minimizers call evaluate with factor_ids_ptr == 0 and num_factor_ids == num_active_factors; the RANSAC minimizers evaluate many items per factor, so a factor used with RANSAC must honor both arguments (see Custom Factors and States (Python and C++)).

    Return True on success. The default implementation raises NotImplementedError.

Skipping the Jacobian entirely. evaluate only has to write to jacobians_ptr when it is non-zero and you intend to supply an analytic Jacobian. A custom factor that never writes to it — even when jacobians_ptr is non-zero — still satisfies the contract, and can be registered with jacobian_mode_override=pycunls.JacobianMode.numeric in Problem.add_factor_batch() to have cuNLS differentiate it via finite differences instead. See Numeric (finite-difference) Jacobians for details and Warp factor code walkthrough for a worked Python example.

C++ reference: FactorBatch.

pycunls.warp.WarpFactorBatch#

Convenience base for custom factors implemented with NVIDIA Warp kernels. Inherits from CustomFactorBatch and provides helper methods for zero-copy pointer wrapping so you never need to manually construct wp.array objects from raw device addresses. Requires warp-lang.

Constructor

from pycunls.warp import WarpFactorBatch

class MyWarpFactor(WarpFactorBatch):
    def __init__(self, capacity):
        super().__init__(
            residual_size=...,
            state_sizes=[...],
            capacity=capacity,
            device="cuda:0",
        )
  • device (str, default "cuda:0") — Warp device string used when creating wp.array wrappers via wrap_array.

Helper methods (inherited — do not override)

  • wrap_array(ptr: int, dtype, shape) -> wp.array — zero-copy wrap of an existing GPU allocation as a Warp array. ptr is the device address, dtype a Warp data type (e.g. wp.float32), and shape an int or tuple giving the array dimensions. The returned wp.array shares the memory; no allocation or copy occurs.

  • factor_ids(factor_ids_ptr: int, num_items: int) -> wp.array — the factor index of every item as an int32 Warp array: wraps factor_ids_ptr when it is non-null, otherwise returns (and caches) arange(num_items) % num_active_factors. Kernels can then always read ids[t].

  • make_warp_stream(stream_handle: int) -> wp.Stream — wraps a raw cudaStream_t (passed as int) as a wp.Stream. Use the returned stream in wp.launch(..., stream=stream) to ensure the Warp kernel executes on the minimizer’s CUDA stream.

Methods to override

  • evaluate(residuals_ptr, jacobians_ptr, state_pointers_ptr, stream_handle, factor_ids_ptr, num_factor_ids) -> bool — same contract as CustomFactorBatch.evaluate. Typical implementations:

    1. Gather scattered state pointers into contiguous CuPy arrays (using a CuPy RawKernel or cp.ndarray indexing), one entry per item.

    2. Get per-item factor indices with self.factor_ids(factor_ids_ptr, num_factor_ids) and read measurements through them.

    3. Wrap the contiguous arrays and output buffers with self.wrap_array.

    4. Build a wp.Stream with self.make_warp_stream.

    5. Launch a @wp.kernel with dim=num_factor_ids on that stream.

See Custom Warp Factor for a complete example.

C++ API#

Factor inputs#

Each factor’s Inputs subsection below and the State API API list the required state batch types (e.g. VectorStateBatch<Dim>, SO3StateBatch).

FactorBatch#

Abstract base (cunls/factor/factor_batch.h).

\[r = f(x),\qquad J = \frac{\partial f}{\partial x}\]
bool Evaluate(
float *residuals,
float *jacobians,
float const *const *state_pointers,
cudaStream_t stream,
const int *factor_ids = nullptr,
size_t num_factor_ids = 0
) const#

Evaluates residuals and, optionally, Jacobians for a list of items.

Terms. \(N\) = NumActiveFactors() (factors/measurements in the batch), \(B\) = StateSizes().size() (states per factor), \(m\) = ResidualsSize(), \(J\) = sum of StateSizes() (Jacobian columns per factor). An item is one factor evaluated at one set of \(B\) states. The call evaluates \(n\) items \(t = 0 \ldots n-1\); item \(t\) reads the measurement of factor \(f(t)\) and its own state pointers, and writes its own output rows.

With the default arguments item \(t\) is simply factor \(t\) (\(n = N\), \(f(t) = t\)) — exactly the behavior of the classic 4-argument call, used by the regular minimizers. The last two arguments let one call evaluate the same factors at many state sets; the RANSAC minimizers evaluate every hypothesis this way.

Parameters:
  • residuals – [out] Device array of \(n \cdot m\) floats. Item \(t\) writes residuals[t * m + r] for \(r \in [0, m)\).

  • jacobians – [out] Device array of \(n \cdot m \cdot J\) floats, or nullptr when only residuals are needed. Item \(t\) writes a row-major \(m \times J\) block; element \((r, c)\) is jacobians[(t * m + r) * J + c]. Columns follow the states in order, each contributing its tangent size.

  • state_pointers – [in] Device array of \(n \cdot B\) device pointers. Item \(t\) reads state \(b\) from state_pointers[t * B + b]. Different items may point to the same state (e.g. every PnP factor points to the one camera pose).

  • stream – [in] CUDA stream on which all work is enqueued; the call may return before the work completes.

  • factor_ids – [in] Which factor each item evaluates. nullptr (default): \(f(t) = t \bmod N\) — with \(n = kN\) this evaluates the whole batch \(k\) times, copy \(c\) being items \([cN, (c+1)N)\). Otherwise a device array of \(n\) indices in \([0, N)\) with \(f(t) =\) factor_ids[t] (any order, repeats allowed).

  • num_factor_ids – [in] Number of items \(n\); 0 (default) means \(n = N\). When factor_ids is given it is that array’s length.

Returns:

[out] true on success.

Examples (\(N = 3\), \(B = 1\), \(m = 2\); ptrs[t] is the state pointer of item \(t\), Pk the state of set P for factor k, res rows the range of residuals item \(t\) writes):

1. Plain evaluation: Evaluate(res, jac, ptrs, stream)        n = 3
     item t          0     1     2
     factor f(t)     0     1     2
     ptrs[t]         x0    x1    x2
     res rows        [0,2) [2,4) [4,6)

2. Whole batch at two state sets P and Q:
   Evaluate(res, jac, ptrs, stream, nullptr, 6)              n = 6
     item t          0     1     2     3     4     5
     factor f(t)     0     1     2     0     1     2      (t % 3)
     ptrs[t]         P0    P1    P2    Q0    Q1    Q2
     res rows        [0,2) [2,4) [4,6) [6,8) [8,10) [10,12)

3. Chosen factors: ids = {2, 0, 2, 1} (device array)
   Evaluate(res, jac, ptrs, stream, ids, 4)                  n = 4
     item t          0     1     2     3
     factor f(t)     2     0     2     1
     ptrs[t]         P     P     Q     Q
     res rows        [0,2) [2,4) [4,6) [6,8)

Requirements for implementations. Launch one thread per item; index measurements by \(f(t)\) and everything else (state pointers, outputs) by \(t\). Item \(t\) must produce exactly what a plain evaluation produces for factor \(f(t)\) at item \(t\)’s states (all built-in batches are bitwise equal). Size any internal per-factor scratch for \(n\) items, not \(N\). Do not assume \(n = N\) or \(f(t) = t\). See Custom Factors and States (Python and C++) for a complete walkthrough (C++ and Python).

size_t ResidualsSize() const#
Returns:

[out] Residual dimension per factor.

std::vector<size_t> StateSizes() const#
Returns:

[out] State tangent dimensions consumed by each factor.

size_t NumActiveFactors() const#
Returns:

[out] Number of active factors: the first NumActiveFactors() measurements are used. 0 after construction, until SetNumActiveFactors.

size_t Capacity() const#
Returns:

[out] Number of factors the measurement buffers hold: the capacity passed to the constructor. Constant for the batch’s lifetime. A custom batch constructed without a capacity (one that overrides NumActiveFactors() instead) has a fixed size, and its capacity is NumActiveFactors().

void SetNumActiveFactors(size_t num_active_factors)#

Sets the active factor count. Every batch starts with 0 active factors: call this before the first solve, and again whenever the count changes (e.g. per frame, after rewriting the measurement buffers in place). Host-only (no allocation, no device work); takes effect at the next Minimize. Wrappers (InformationFactorBatch, WeightedFactorBatch) forward it to the wrapped batch.

Parameters:

num_active_factors – [in] Active count, at most Capacity().

Throws:
  • std::invalid_argument – if num_active_factors > Capacity().

  • std::logic_error – if a subclass overrides NumActiveFactors(), so the set would have no effect. Custom batches pass their capacity to SizedFactorBatch(capacity) instead.

Residual-only factors. Evaluate must support jacobians == nullptr (residual-only evaluation) — this is required for cost-only evaluation, and it is also all that’s needed to opt a factor into cuNLS’s numeric (finite-difference) Jacobians: a factor whose Evaluate never writes to jacobians at all still satisfies this interface, and can be solved by registering it with JacobianMode::kNumeric (see Numeric (finite-difference) Jacobians and JacobianMode / NumericDiffOptions) instead of implementing a Jacobian by hand.

SizedFactorBatch<kResidualSize, …kStateSizes>#

Compile-time convenience base (cunls/factor/sized_factor_batch.h) that fixes residual and state (tangent) dimensions at compile time.

Each specialization exposes sized_layout — an alias for the same SizedFactorBatch<kResidualSize, kStateSizes...> type. Wrapper templates such as InformationFactorBatch<T> and WeightedFactorBatch<T> inherit public T::sized_layout so they remain full SizedFactorBatch instances with the same layout as the inner batch T.

PriorVectorFactorBatch<Dim>#

Header: cunls/factor/prior/prior_vector_factor_batch.h

Prior on a Euclidean vector (e.g. bias, landmark). Pulls the state toward observed values.

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(r = x - o\)

\(\mathrm{Dim}\)

\(I\)

\(\mathrm{Dim} \times \mathrm{Dim}\)

\(\mathbb{R}^{\mathrm{Dim}}\)

Inputs: \(x\) = state vector, \(o\) = observation (constructor). State: one state from VectorStateBatch<Dim> (see State API).

Constructor:

PriorVectorFactorBatch(const Vector<Dim>* observations_ptr, size_t capacity)
  • observations_ptr — [in] Device pointer to observed vectors.

  • capacity — [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

SO2PriorFactorBatch#

Header: cunls/factor/prior/so2_prior_factor_batch.h

Prior on a 2D rotation (e.g. heading). Penalizes deviation from a target rotation.

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(r = \mathrm{Log}(R_{\mathrm{target}}^\top R)\)

1

\(1\)

\(1 \times 1\)

SO(2)

Inputs: \(R\) = current rotation (state). State: one state from SO2StateBatch (see State API).

SO2PriorFactorBatch(
const Matrix<2> *observations_ptr,
size_t capacity
)#
Parameters:
  • observations_ptr – [in] Device pointer to SO(2) observations (2×2 row-major).

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SO3PriorFactorBatch#

Header: cunls/factor/prior/so3_prior_factor_batch.h

Prior on a 3D rotation. Penalizes deviation from a target orientation.

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(r = \mathrm{Log}(R_{\mathrm{target}}^\top R)\)

3

\(J_r^{-1}(r)\)

\(3 \times 3\)

SO(3)

Inputs: \(R\) = current rotation (state). State: one state from SO3StateBatch (see State API).

SO3PriorFactorBatch(
const Matrix<3> *observations_ptr,
size_t capacity
)#
Parameters:
  • observations_ptr – [in] Device pointer to SO(3) observations (3×3 row-major).

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SE2PriorFactorBatch#

Header: cunls/factor/prior/se2_prior_factor_batch.h

Prior on 2D rigid transform. State: one state from SE2StateBatch (see State API).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(r = \mathrm{Log}(T_{\mathrm{target}}^{-1} T)\)

3

\(J_r^{-1}(r)\)

\(3 \times 3\)

SE(2)

SE3PriorFactorBatch#

Header: cunls/factor/prior/se3_prior_factor_batch.h

Prior on 3D rigid transform. State: one state from SE3StateBatch (see State API).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(r = \mathrm{Log}(T_{\mathrm{target}}^{-1} T)\)

6

\(J_r^{-1}(r)\)

\(6 \times 6\)

SE(3)

Similarity2PriorFactorBatch#

Header: cunls/factor/prior/similarity2_prior_factor_batch.h

Prior on 2D similarity transform. State: one state from Similarity2StateBatch (see State API).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(r = \mathrm{Log}(T_{\mathrm{target}}^{-1} T)\)

4

\(J_r^{-1}(r)\)

\(4 \times 4\)

Sim(2)

Similarity3PriorFactorBatch#

Header: cunls/factor/prior/similarity3_prior_factor_batch.h

Prior on 3D similarity transform. State: one state from Similarity3StateBatch (see State API).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(r = \mathrm{Log}(T_{\mathrm{target}}^{-1} T)\)

7

\(J_r^{-1}(r)\)

\(7 \times 7\)

Sim(3)

Constructors (all four prior classes above):

ClassName(const ObsType *observations_ptr, size_t capacity)#
Parameters:
  • observations_ptr – [in] Device pointer to observation transforms.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SL4PriorFactorBatch#

Header: cunls/factor/prior/sl4_prior_factor_batch.h

Prior on an SL(4) transform. State: one state from SL4StateBatch (see State API).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(r = \mathrm{Log}(T_{\mathrm{target}}^{-1} T)\)

15

\(I\)

\(15 \times 15\)

SL(4)

SL4PriorFactorBatch(
const SL4Transform *observations_ptr,
size_t capacity
)#
Parameters:
  • observations_ptr – [in] Device pointer to SL(4) target transforms (row-major 4×4).

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SE3BetweenFactorBatch#

Header: cunls/factor/between/se3_between_factor_batch.h

Constrains the relative pose between two SE(3) frames (e.g. odometry, loop closure).

\[r = \mathrm{Log}\bigl( \Delta \, T_{\mathrm{left}}^{-1} \, T_{\mathrm{right}} \bigr), \qquad r = 0 \iff \Delta = T_{\mathrm{right}}^{-1} T_{\mathrm{left}}\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

6

left/right SE(3) Jacobians

\(6 \times 12\)

SE(3) × SE(3)

Inputs: \(T_{\mathrm{left}}\), \(T_{\mathrm{right}}\) = two poses (states). State: two states from SE3StateBatch (see State API). \(\Delta\) = measured relative transform (constructor).

SE3BetweenFactorBatch(
const SE3Transform *pose_deltas_ptr,
size_t capacity
)#
Parameters:
  • pose_deltas_ptr – [in] Device pointer to measured relative transforms.

  • capacity – [in] Number of between constraints the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SE2BetweenFactorBatch#

Header: cunls/factor/between/se2_between_factor_batch.h

Constrains the relative transform between two SE(2) frames.

\[r = \mathrm{Log}\bigl( \Delta \, T_{\mathrm{left}}^{-1} \, T_{\mathrm{right}} \bigr), \qquad r = 0 \iff \Delta = T_{\mathrm{right}}^{-1} T_{\mathrm{left}}\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

3

left/right SE(2) Jacobians

\(3 \times 6\)

SE(2) × SE(2)

Inputs: \(T_{\mathrm{left}}\), \(T_{\mathrm{right}}\) = two poses (states). State: two states from SE2StateBatch (see State API). \(\Delta\) = measured relative transform (constructor).

SE2BetweenFactorBatch(
const Matrix<3> *pose_deltas_ptr,
size_t capacity
)#
Parameters:
  • pose_deltas_ptr – [in] Device pointer to measured relative transforms (row-major 3×3).

  • capacity – [in] Number of between constraints the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SO2BetweenFactorBatch#

Header: cunls/factor/between/so2_between_factor_batch.h

Constrains the relative rotation between two SO(2) frames (zero at \(R_{\mathrm{right}} = R_{\mathrm{left}} \Delta^{\top}\), as for SO(3)).

\[r = \mathrm{Log}\bigl( R_{\mathrm{left}}^{\top} \, R_{\mathrm{right}} \, \Delta \bigr)\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

1

left/right SO(2) Jacobians

\(1 \times 2\)

SO(2) × SO(2)

Inputs: \(R_{\mathrm{left}}\), \(R_{\mathrm{right}}\) = two rotations (states). State: two states from SO2StateBatch (see State API). \(\Delta\) = measured relative rotation (constructor).

SO2BetweenFactorBatch(
const Matrix<2> *rotation_deltas_ptr,
size_t capacity
)#
Parameters:
  • rotation_deltas_ptr – [in] Device pointer to measured relative rotations (row-major 2×2).

  • capacity – [in] Number of between constraints the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SO3BetweenFactorBatch#

Header: cunls/factor/between/so3_between_factor_batch.h

Constrains the relative rotation between two SO(3) frames.

\[r = \mathrm{Log}\bigl( R_{\mathrm{left}}^{\top} \, R_{\mathrm{right}} \, \Delta \bigr)\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

3

left/right SO(3) Jacobians

\(3 \times 6\)

SO(3) × SO(3)

Inputs: \(R_{\mathrm{left}}\), \(R_{\mathrm{right}}\) = two rotations (states). State: two states from SO3StateBatch (see State API). \(\Delta\) = measured relative rotation (constructor).

SO3BetweenFactorBatch(
const Matrix<3> *rotation_deltas_ptr,
size_t capacity
)#
Parameters:
  • rotation_deltas_ptr – [in] Device pointer to measured relative rotations (row-major 3×3).

  • capacity – [in] Number of between constraints the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

Similarity2BetweenFactorBatch#

Header: cunls/factor/between/similarity2_between_factor_batch.h

Constrains the relative transform between two Sim(2) frames.

\[r = \mathrm{Log}\bigl( \Delta \, T_{\mathrm{left}}^{-1} \, T_{\mathrm{right}} \bigr), \qquad r = 0 \iff \Delta = T_{\mathrm{right}}^{-1} T_{\mathrm{left}}\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

4

left/right Sim(2) Jacobians

\(4 \times 8\)

Sim(2) × Sim(2)

Inputs: \(T_{\mathrm{left}}\), \(T_{\mathrm{right}}\) = two transforms (states). State: two states from Similarity2StateBatch (see State API). \(\Delta\) = measured relative transform (constructor).

Similarity2BetweenFactorBatch(
const Matrix<3> *pose_deltas_ptr,
size_t capacity
)#
Parameters:
  • pose_deltas_ptr – [in] Device pointer to measured relative transforms (row-major 3×3).

  • capacity – [in] Number of between constraints the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

Similarity3BetweenFactorBatch#

Header: cunls/factor/between/similarity3_between_factor_batch.h

Constrains the relative transform between two Sim(3) frames.

\[r = \mathrm{Log}\bigl( \Delta \, T_{\mathrm{left}}^{-1} \, T_{\mathrm{right}} \bigr), \qquad r = 0 \iff \Delta = T_{\mathrm{right}}^{-1} T_{\mathrm{left}}\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

7

left/right Sim(3) Jacobians

\(7 \times 14\)

Sim(3) × Sim(3)

Inputs: \(T_{\mathrm{left}}\), \(T_{\mathrm{right}}\) = two transforms (states). State: two states from Similarity3StateBatch (see State API). \(\Delta\) = measured relative transform (constructor).

Similarity3BetweenFactorBatch(
const Matrix<4> *pose_deltas_ptr,
size_t capacity
)#
Parameters:
  • pose_deltas_ptr – [in] Device pointer to measured relative transforms (row-major 4×4).

  • capacity – [in] Number of between constraints the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SL4BetweenFactorBatch#

Header: cunls/factor/between/sl4_between_factor_batch.h

Constrains the relative transform between two SL(4) frames.

\[r = \mathrm{Log}\bigl( \Delta \, T_{\mathrm{left}}^{-1} \, T_{\mathrm{right}} \bigr), \qquad r = 0 \iff \Delta = T_{\mathrm{right}}^{-1} T_{\mathrm{left}}\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

15

left/right SL(4) Jacobians

\(15 \times 30\)

SL(4) × SL(4)

Inputs: \(T_{\mathrm{left}}\), \(T_{\mathrm{right}}\) = two transforms (states). State: two states from SL4StateBatch (see State API). \(\Delta\) = measured relative transform (constructor).

SL4BetweenFactorBatch(
const SL4Transform *pose_deltas_ptr,
size_t capacity
)#
Parameters:
  • pose_deltas_ptr – [in] Device pointer to measured relative transforms (row-major 4×4, unit determinant).

  • capacity – [in] Number of between constraints the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

VectorBetweenFactorBatch<Dim>#

Header: cunls/factor/between/vector_between_factor_batch.h

Constrains the difference between two Euclidean vector states.

\[r = x_{\mathrm{left}} - x_{\mathrm{right}} - \delta\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\(x_l - x_r - \delta\)

\(\mathrm{Dim}\)

\([I \;|\; {-}I]\)

\(\mathrm{Dim} \times 2\mathrm{Dim}\)

\(\mathbb{R}^{\mathrm{Dim}} \times \mathbb{R}^{\mathrm{Dim}}\)

Inputs: \(x_l\), \(x_r\) = two vector states. State: two states from VectorStateBatch<Dim> (see State API). \(\delta\) = measured difference (constructor).

Constructor:

VectorBetweenFactorBatch(const Vector<Dim>* deltas_ptr, size_t capacity)
  • deltas_ptr — [in] Device pointer to measured difference vectors.

  • capacity — [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Manifold facade#

The per-manifold classes above (SE3BetweenFactorBatch, SO2PriorFactorBatch, …) are the hand-optimized implementations. Two zero-cost, manifold-generic facades let callers write fewer distinct class names:

template<class Manifold>
class BetweenFactorBatch#

Header: cunls/factor/between/between_factor_batch.h. A compile-time alias for the matching XxxBetweenFactorBatch: each specialization adds no data members or virtual dispatch, so sizeof(BetweenFactorBatch<manifold::SE3>) == sizeof(SE3BetweenFactorBatch) and the two are interchangeable everywhere a FactorBatch* is expected. Manifold is one of the tags in cunls::manifold (SE3, SO3, SE2, SO2, Similarity2, Similarity3, SL4, or Vector<Dim>) and is usually deduced via CTAD from the deltas pointer’s own (now manifold-distinct) type, so <Manifold> need not be written explicitly:

cunls::BetweenFactorBatch between(deltas_ptr, capacity);  // manifold deduced

CTAD deduction is unavailable for Vector<Dim> (a C++ template-argument-deduction limitation: Vector’s int Dim cannot be deduced from the size_t extent of the underlying array type), so that one specialization requires <manifold::Vector<Dim>> explicitly.

template<class Manifold>
class PriorFactorBatch#

Header: cunls/factor/prior/prior_factor_batch.h. Same mechanism as BetweenFactorBatch<Manifold>, for XxxPriorFactorBatch instead of XxxBetweenFactorBatch; Manifold is deduced from the observations pointer’s type.

template<class Manifold>
class ConstantVelocityFactorBatch#

Header: cunls/factor/motion/constant_velocity_factor_batch.h. Same zero-cost specialization mechanism, but always requires the manifold as an explicit template argument: every ConstantVelocityXxxFactorBatch constructor is (const float* dt_ptr, size_t capacity), so there is no manifold-specific argument type to deduce from.

cunls::ConstantVelocityFactorBatch<cunls::manifold::SE3> factor(dt_ptr, capacity);
template<class Manifold>
class ConstantAccelerationFactorBatch#

Header: cunls/factor/motion/constant_acceleration_factor_batch.h. Identical mechanism and explicit-template-argument-only convention as ConstantVelocityFactorBatch<Manifold>.

Motion prior factors#

Constant-velocity (CV) and constant-acceleration (CA) motion-prior factors constrain how a pose evolves between two consecutive timestamps, given an explicit body-velocity (and, for CA, body-acceleration) state at each timestamp. Pose, velocity, and acceleration are kept as separate states (an SE3StateBatch/SO3StateBatch/SE2StateBatch/ SO2StateBatch for the pose, VectorStateBatch<Dim> for velocity/acceleration, Dim matching the pose’s tangent size) connected by one of the factors below.

For a pose group with Log/Exp and inverse-left-Jacobian \(J_l^{-1}\), and twist \(:= \mathrm{Log}(T_k^{-1} T_{k+1})\):

\[\begin{split}r_{\mathrm{pose}} &= \mathrm{twist} - \Delta t \, v_k \;\;(- \tfrac{1}{2}\Delta t^2 a_k \text{ for CA}) \\ r_{\mathrm{vel}} &= J_l^{-1}(\mathrm{twist}) \, v_{k+1} - v_k \;\;(- \Delta t \, a_k \text{ for CA}) \\ r_{\mathrm{accel}} &= J_l^{-1}(\mathrm{twist}) \, a_{k+1} - a_k \quad \text{(CA only)}\end{split}\]

i.e. the relative pose should match a first- (CV) or second-order (CA) Taylor prediction from the velocity/acceleration at \(k\), and the velocity/acceleration at \(k+1\), transported back into the local frame at \(k\) through the inverse left Jacobian, should match the value at \(k\). The Jacobians of \(r_{\mathrm{vel}}\) and \(r_{\mathrm{accel}}\) with respect to the pose states are treated as zero — a documented simplification (the residuals themselves are exact; only that specific curvature term is dropped). SO(2) is abelian (\(J_l^{-1} = 1\)), so its factors reduce to scalar arithmetic; SE(2) obtains \(J_l^{-1}\) from the identity \(J_l^{-1}(x) = J_r^{-1}(-x)\) rather than a dedicated left-Jacobian primitive.

ConstantVelocitySE3FactorBatch#

Header: cunls/factor/motion/constant_velocity_se3_factor_batch.h

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\([r_{\mathrm{pose}}; r_{\mathrm{vel}}]\)

12

closed-form (see above)

\(12 \times 24\)

SE(3) × SE(3) × \(\mathbb{R}^6\) × \(\mathbb{R}^6\)

Inputs: \(T_k, T_{k+1}\) = two states from SE3StateBatch; \(v_k, v_{k+1}\) = two states from VectorStateBatch<6> (body twist). \(\Delta t\) (constructor).

ConstantVelocitySE3FactorBatch(const float *dt_ptr, size_t capacity)#
Parameters:
  • dt_ptr – [in] Device pointer to per-factor time deltas.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

ConstantVelocitySO3FactorBatch#

Header: cunls/factor/motion/constant_velocity_so3_factor_batch.h. Same construction as ConstantVelocitySE3FactorBatch, specialized to SO(3) (vel = angular velocity in \(\mathbb{R}^3\)).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\([r_{\mathrm{pose}}; r_{\mathrm{vel}}]\)

6

closed-form (see above)

\(6 \times 12\)

SO(3) × SO(3) × \(\mathbb{R}^3\) × \(\mathbb{R}^3\)

Inputs: \(R_k, R_{k+1}\) = two states from SO3StateBatch; \(v_k, v_{k+1}\) = two states from VectorStateBatch<3>. \(\Delta t\) (constructor).

ConstantVelocitySO3FactorBatch(const float *dt_ptr, size_t capacity)#
Parameters:
  • dt_ptr – [in] Device pointer to per-factor time deltas.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

ConstantVelocitySE2FactorBatch#

Header: cunls/factor/motion/constant_velocity_se2_factor_batch.h. Same construction, specialized to SE(2).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\([r_{\mathrm{pose}}; r_{\mathrm{vel}}]\)

6

closed-form (see above)

\(6 \times 12\)

SE(2) × SE(2) × \(\mathbb{R}^3\) × \(\mathbb{R}^3\)

Inputs: \(T_k, T_{k+1}\) = two states from SE2StateBatch; \(v_k, v_{k+1}\) = two states from VectorStateBatch<3> (body twist). \(\Delta t\) (constructor).

ConstantVelocitySE2FactorBatch(const float *dt_ptr, size_t capacity)#
Parameters:
  • dt_ptr – [in] Device pointer to per-factor time deltas.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

ConstantVelocitySO2FactorBatch#

Header: cunls/factor/motion/constant_velocity_so2_factor_batch.h. SO(2) is abelian, so the residual/Jacobian reduce to scalar arithmetic.

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\([r_{\mathrm{pose}}; r_{\mathrm{vel}}]\)

2

closed-form scalars

\(2 \times 4\)

SO(2) × SO(2) × \(\mathbb{R}\) × \(\mathbb{R}\)

Inputs: \(\theta_k, \theta_{k+1}\) = two states from SO2StateBatch; \(v_k, v_{k+1}\) = two states from VectorStateBatch<1>. \(\Delta t\) (constructor).

ConstantVelocitySO2FactorBatch(const float *dt_ptr, size_t capacity)#
Parameters:
  • dt_ptr – [in] Device pointer to per-factor time deltas.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

ConstantAccelerationSE3FactorBatch#

Header: cunls/factor/motion/constant_acceleration_se3_factor_batch.h. Adds an acceleration state to ConstantVelocitySE3FactorBatch’s construction; see the general formula above.

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\([r_{\mathrm{pose}}; r_{\mathrm{vel}}; r_{\mathrm{accel}}]\)

18

closed-form (see above)

\(18 \times 36\)

SE(3) × SE(3) × \((\mathbb{R}^6)^4\)

Inputs: \(T_k, T_{k+1}\) = two states from SE3StateBatch; \(v_k, v_{k+1}, a_k, a_{k+1}\) = four states from VectorStateBatch<6>. \(\Delta t\) (constructor).

ConstantAccelerationSE3FactorBatch(
const float *dt_ptr,
size_t capacity
)#
Parameters:
  • dt_ptr – [in] Device pointer to per-factor time deltas.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

ConstantAccelerationSO3FactorBatch#

Header: cunls/factor/motion/constant_acceleration_so3_factor_batch.h. Same construction, specialized to SO(3).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\([r_{\mathrm{pose}}; r_{\mathrm{vel}}; r_{\mathrm{accel}}]\)

9

closed-form (see above)

\(9 \times 18\)

SO(3) × SO(3) × \((\mathbb{R}^3)^4\)

Inputs: \(R_k, R_{k+1}\) = two states from SO3StateBatch; \(v_k, v_{k+1}, a_k, a_{k+1}\) = four states from VectorStateBatch<3>. \(\Delta t\) (constructor).

ConstantAccelerationSO3FactorBatch(
const float *dt_ptr,
size_t capacity
)#
Parameters:
  • dt_ptr – [in] Device pointer to per-factor time deltas.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

ConstantAccelerationSE2FactorBatch#

Header: cunls/factor/motion/constant_acceleration_se2_factor_batch.h. Same construction, specialized to SE(2).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\([r_{\mathrm{pose}}; r_{\mathrm{vel}}; r_{\mathrm{accel}}]\)

9

closed-form (see above)

\(9 \times 18\)

SE(2) × SE(2) × \((\mathbb{R}^3)^4\)

Inputs: \(T_k, T_{k+1}\) = two states from SE2StateBatch; \(v_k, v_{k+1}, a_k, a_{k+1}\) = four states from VectorStateBatch<3>. \(\Delta t\) (constructor).

ConstantAccelerationSE2FactorBatch(
const float *dt_ptr,
size_t capacity
)#
Parameters:
  • dt_ptr – [in] Device pointer to per-factor time deltas.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

ConstantAccelerationSO2FactorBatch#

Header: cunls/factor/motion/constant_acceleration_so2_factor_batch.h. SO(2) is abelian, so the residual/Jacobian reduce to scalar arithmetic.

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

\([r_{\mathrm{pose}}; r_{\mathrm{vel}}; r_{\mathrm{accel}}]\)

3

closed-form scalars

\(3 \times 6\)

SO(2) × SO(2) × \(\mathbb{R}^4\)

Inputs: \(\theta_k, \theta_{k+1}\) = two states from SO2StateBatch; \(v_k, v_{k+1}, a_k, a_{k+1}\) = four states from VectorStateBatch<1>. \(\Delta t\) (constructor).

ConstantAccelerationSO2FactorBatch(
const float *dt_ptr,
size_t capacity
)#
Parameters:
  • dt_ptr – [in] Device pointer to per-factor time deltas.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

Motion prior covariance weighting#

Header: cunls/factor/information/motion_prior_information.h. The eight factors above are unweighted (unit information) on their own. To fuse the paper’s closed-form process-noise covariance \(Q(\Delta t)^{-1}\) into the residual/Jacobian, construct the corresponding named alias below instead of the plain factor — it is a drop-in replacement that takes two extra arguments (a stream for the one-off precomputation and the process-noise PSD) and internally computes the square-root information and composes it with the existing InformationFactorBatch<T> (see above), so callers never see the Kronecker-product math or manage a separate information buffer:

ConstantVelocityInformationSE3FactorBatch factor(
    stream, dt_ptr, qc_diag_ptr, capacity);
factor.SetNumActiveFactors(num_factors);  // active count, at most capacity

Available aliases (one per factor above, same residual/Jacobian shape as the wrapped factor): ConstantVelocityInformationSE3FactorBatch, ConstantVelocityInformationSO3FactorBatch, ConstantVelocityInformationSE2FactorBatch, ConstantVelocityInformationSO2FactorBatch, ConstantAccelerationInformationSE3FactorBatch, ConstantAccelerationInformationSO3FactorBatch, ConstantAccelerationInformationSE2FactorBatch, ConstantAccelerationInformationSO2FactorBatch.

template<class T, int Dim>
class MotionPriorInformationFactorBatch#

The generic template all eight aliases instantiate; T is the wrapped ConstantVelocityXxxFactorBatch/ ConstantAccelerationXxxFactorBatch and Dim its pose tangent size (6/3/3/1 for SE(3)/SO(3)/SE(2)/SO(2)). Prefer the named aliases; only spell this out directly for a factor/Dim combination that doesn’t have one yet.

MotionPriorInformationFactorBatch(
cudaStream_t stream,
const float *dt_ptr,
const float *qc_diag_ptr,
size_t capacity
)#
Parameters:
  • stream – [in] CUDA stream used to precompute the sqrt-information matrices at construction time.

  • dt_ptr – [in] Device pointer to per-factor time deltas; also forwarded to the wrapped factor’s own constructor.

  • qc_diag_ptr – [in] Device pointer to the continuous-time process-noise PSD diagonal (Dim floats), constant across the batch.

  • capacity – [in] Number of factors the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

The two function templates behind this wrapper, ComputeConstantVelocitySqrtInformation<Dim> and ComputeConstantAccelerationSqrtInformation<Dim>, are still available directly (same header) for callers who want the raw square-root information matrix without going through InformationFactorBatch — it’s a Kronecker product with an analytic Cholesky factor, no numerical linear algebra.

ReprojectionFactorBatch#

Header: cunls/factor/reprojection_factor_batch.h

Reprojection error for bundle adjustment. Observations in normalized image coordinates.

\[\begin{split}P_{\mathrm{cam}} = T_{\mathrm{cr}}\, T^{-1} P = T_{\mathrm{cr}} R^\top (P - t),\qquad r = \begin{bmatrix} P_{\mathrm{cam},x}/P_{\mathrm{cam},z} - x_n \\ P_{\mathrm{cam},y}/P_{\mathrm{cam},z} - y_n \end{bmatrix}\end{split}\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

2

chain rule on projection

\(2 \times 9\)

SE(3) × \(\mathbb{R}^3\)

Inputs: Pose \(T = (R, t)\) (state 1, world_from_rig, see pose-convention), 3D point \(P\) (state 2). State: SE3StateBatch then VectorStateBatch<3> (see State API). Observations \((x_n, y_n)\) and the optional camera-from-rig \(T_{\mathrm{cr}}\) (identity if omitted) from the constructor.

Jacobians (right perturbation \(T\,\mathrm{Exp}([\phi; \rho])\) in the rig frame), with \(P_{\mathrm{rig}} = R^\top (P - t)\) and \(A = \partial \pi / \partial P_{\mathrm{cam}}\; R_{\mathrm{cr}}\): \(\partial r / \partial \phi = A\,[P_{\mathrm{rig}}]_\times\), \(\partial r / \partial \rho = -A\), \(\partial r / \partial P = A R^\top\).

ReprojectionFactorBatch(
const Vector<2> *observations,
size_t capacity,
float z_threshold = 1e-3f
)#
ReprojectionFactorBatch(
const Vector<2> *observations,
const SE3Transform *poses_camera_from_rig,
size_t capacity,
float z_threshold = 1e-3f
)#
Parameters:
  • observations – [in] Device pointer to normalized observations.

  • poses_camera_from_rig – [in] Optional device pointer to camera extrinsics (second overload).

  • capacity – [in] Number of reprojection factors the buffers hold. 0 are active until SetNumActiveFactors.

  • z_threshold – [in] Minimum valid depth.

Returns:

Constructor has no return value.

PnPFactorBatch#

Header: cunls/factor/pnp_factor_batch.h

Fixed-structure Perspective-n-Point reprojection: the same normalized pinhole residual as ReprojectionFactorBatch, but each 3D landmark is held in device memory passed to the constructor (not a state variable). Only the SE(3) pose is optimized; the analytic Jacobian is therefore \(2 \times 6\).

\[\begin{split}P_{\mathrm{cam}} = T_{\mathrm{cam}\leftarrow\mathrm{world}}\, P_{\mathrm{world}},\qquad r = \begin{bmatrix} P_{\mathrm{cam},x}/P_{\mathrm{cam},z} - x_n \\ P_{\mathrm{cam},y}/P_{\mathrm{cam},z} - y_n \end{bmatrix}\end{split}\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(same pinhole model as reprojection)

2

pose chain rule only

\(2 \times 6\)

SE(3)

Inputs: Normalized observations \((x_n,y_n)\) and matching world points \(P_{\mathrm{world}}\) from the constructor (one pair per factor). State: a single SE3StateBatch state per factor (world_from_rig pose, the convention of every pose state; see pose-convention). Optional poses_camera_from_rig uses the same composition as ReprojectionFactorBatch.

PnPFactorBatch(
const Vector<2> *observations,
const Vector<3> *points_world,
size_t capacity,
float z_threshold = 1e-3f
)#
PnPFactorBatch(
const Vector<2> *observations,
const SE3Transform *poses_camera_from_rig,
const Vector<3> *points_world,
size_t capacity,
float z_threshold = 1e-3f
)#
Parameters:
  • observations – [in] Device pointer to normalized 2-D observations.

  • points_world – [in] Device pointer to fixed world points \(P\).

  • poses_camera_from_rig – [in] Optional per-factor rig extrinsics (second overload).

  • capacity – [in] Number of PnP correspondences the buffers hold. 0 are active until SetNumActiveFactors.

  • z_threshold – [in] Minimum valid camera-frame depth.

Returns:

Constructor has no return value.

ImuFactorBatch#

Header: cunls/factor/imu_factor_batch.h

IMU factor between two keyframes, with the raw samples between them marginalized inside the factor. Every evaluation integrates the Euler chain from keyframe \(a\) at the current bias, propagates the covariance \(\Sigma\) of the prediction, and returns the exact Schur complement of the chain onto the keyframe states. Theory, conventions and an example: IMU factor.

\[\begin{split}r_{0:9} = L^{-1} \begin{bmatrix} \mathrm{Log}(\hat R^\top R_b) \\ v_b - \hat v \\ p_b - \hat p \end{bmatrix}, \quad \Sigma = L L^\top, \qquad r_{9:15} = \frac{b_b - b_a}{\sigma_b \sqrt{T}}\end{split}\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

Whitened chain defect + bias random walk

15

Analytic

\(15 \times 30\)

SE(3), \(\mathbb{R}^3\), \(\mathbb{R}^6\) (twice)

States (per factor, in order): \(T_a\) (SE3StateBatch, world_from_rig), \(v_a\) (VectorStateBatch<3>, world velocity of the IMU), \(b_a = [b_g; b_a]\) (VectorStateBatch<6>), \(T_b\), \(v_b\), \(b_b\). The IMU’s pose in the world is \(T\,T_{bi}\) with \(T_{bi}\) = ImuParameters::body_from_imu.

struct ImuParameters#

Sensor model: Vector<3> gravity (default \((0, 0, -9.80665)\)), gyro_noise_density [rad/s/√Hz], accel_noise_density [m/s²/√Hz], integration_noise_density [m/√s], gyro_bias_random_walk [rad/s²/√Hz], accel_bias_random_walk [m/s³/√Hz] (continuous-time, all positive; defaults: EuRoC ADIS16448), SE3Transform body_from_imu (rig_from_imu, default identity).

ImuFactorBatch(
const float *imu_samples,
const int *sample_offsets,
size_t num_samples,
const ImuParameters &parameters,
size_t capacity
)#
Parameters:
  • imu_samples – [in] Device array of 7 floats per sample (gyroscope, specific force, step duration), all factors back to back.

  • sample_offsets – [in] Device array of capacity + 1 CSR offsets; factor f uses samples [offsets[f], offsets[f + 1]).

  • num_samples – [in] Samples the buffer holds; sizes the work split only.

  • parameters – [in] Gravity, noise densities and extrinsic (copied).

  • capacity – [in] Keyframe pairs the buffers hold. 0 are active until SetNumActiveFactors.

Throws:

std::invalid_argument – A buffer is null or a noise density is not positive.

PointToPointFactorBatch#

Header: cunls/factor/point_to_point_factor_batch.h

Point cloud registration (e.g. ICP). Residual = target point minus transformed source point.

\[r = p - T q = p - (R q + t)\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

3

\([\partial r/\partial\omega\;\partial r/\partial\rho] = [R[q]_\times\;{-}R]\)

\(3 \times 6\)

SE(3)

Inputs: \(T\) = pose (state). State: one state from SE3StateBatch (see State API). \(p\), \(q\) = target/source points (constructor).

PointToPointFactorBatch(
const Vector<3> *p_observations_ptr,
const Vector<3> *q_observations_ptr,
size_t capacity
)#
Parameters:
  • p_observations_ptr – [in] Device pointer to target points.

  • q_observations_ptr – [in] Device pointer to source points.

  • capacity – [in] Number of correspondences the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

PointToPlaneFactorBatch#

Header: cunls/factor/point_to_plane_factor_batch.h

Plane-based ICP: signed distance from the transformed source point to the plane through the target point with normal \(n_q\), given in the target frame (not rotated by \(T\)).

\[r = n_q^\top (p - T q) = n_q \cdot (p - (R q + t))\]

With \(n' = R^\top n_q\), the Jacobian row is \([n'^\top [q]_\times,\; -n'^\top]\).

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

1

\([n'^\top [q]_\times\;{-}n'^\top]\)

\(1 \times 6\)

SE(3)

Inputs: \(T\) = pose (state). State: one state from SE3StateBatch (see State API). \(p\), \(q\), \(n_q\) = target point, source point, plane normal in the target frame (constructor).

PointToPlaneFactorBatch(
const Vector<3> *p_observations_ptr,
const Vector<3> *q_observations_ptr,
const Vector<3> *nq_observations_ptr,
size_t capacity
)#
Parameters:
  • p_observations_ptr – [in] Device pointer to target points.

  • q_observations_ptr – [in] Device pointer to source points.

  • nq_observations_ptr – [in] Device pointer to plane normals in the target frame.

  • capacity – [in] Number of correspondences the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

SymmetricPointToPlaneFactorBatch#

Header: cunls/factor/symmetric_point_to_plane_factor_batch.h

Symmetric point-to-plane: both frames contribute normals; \(N = n_p + n_q\).

\[r = N^\top \bigl( T p - T^{-1} q \bigr) = \bigl( (R p + t) - R^\top(q - t) \bigr)^\top N\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

1

\(\mathrm{d}(T p,\, T^{-1} q)\) w.r.t. \(T\)

\(1 \times 6\)

SE(3)

Inputs: \(T\) = pose (state). State: one state from SE3StateBatch (see State API). \(p\), \(n_p\), \(q\), \(n_q\) = target/source points and normals (constructor).

SymmetricPointToPlaneFactorBatch(
const Vector<3> *p_observations_ptr,
const Vector<3> *q_observations_ptr,
const Vector<3> *np_observations_ptr,
const Vector<3> *nq_observations_ptr,
size_t capacity
)#
Parameters:
  • p_observations_ptr – [in] Device pointer to target points.

  • q_observations_ptr – [in] Device pointer to source points.

  • np_observations_ptr – [in] Device pointer to target normals.

  • nq_observations_ptr – [in] Device pointer to source normals.

  • capacity – [in] Number of correspondences the buffers hold. 0 are active until SetNumActiveFactors.

Returns:

Constructor has no return value.

InformationFactorBatch<T>#

Header: cunls/factor/information/information_factor_batch.h

Item parameters. Evaluate forwards factor_ids / num_factor_ids to the wrapped factor and weights item \(t\) with the matrix of its factor \(f(t)\), so the wrapper works under the RANSAC minimizers. The sqrt-information product uses deterministic CUDA kernels (fixed summation order per item). Residual sizes up to 96 stage each vector in shared memory; larger sizes read their inputs directly with a stream-ordered scratch buffer.

Inheritance: class InformationFactorBatch : public T::sized_layout — i.e. the same SizedFactorBatch<kResidualSize, ...> as the wrapped type T. Residual and state sizes come from that base; this class adds NumActiveFactors, storage for T, and an Evaluate that applies \(\Omega^{1/2}\) after the inner factor.

T must derive from some SizedFactorBatch (see type trait IsDerivedFromAnySizedFactorBatch).

Wraps a factor to apply a square-root information matrix \(\Omega^{1/2}\) (e.g. from measurement covariance).

\[r_{\mathrm{weighted}} = \Omega^{1/2} r,\qquad J_{\mathrm{weighted}} = \Omega^{1/2} J\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

(per \(T\))

(see above)

(per \(T\))

same as wrapped \(T\)

Inputs: Same state layout as the wrapped factor T. Constructor also takes sqrt_information_matrices_ptr (per-factor \(\Omega^{1/2}\)).

template<class ...Args>
InformationFactorBatch(
const Matrix<T::residual_size_> *sqrt_information_matrices_ptr,
size_t capacity,
Args&&... sized_factor_batch_args
)#
Parameters:
  • sqrt_information_matrices_ptr – [in] Device pointer to per-factor square-root information matrices.

  • capacity – [in] Number of square-root information matrices; must equal T::Capacity() after T is constructed. The active count is the wrapped batch’s (SetNumActiveFactors is forwarded).

  • sized_factor_batch_args – [in] Constructor arguments forwarded to wrapped factor T (same order as T’s constructor, or (weight, …) when T is WeightedFactorBatch<U>).

Returns:

Constructor has no return value.

WeightedFactorBatch<T>#

Header: cunls/factor/weighted_factor_batch.h

Inheritance: class WeightedFactorBatch : public T::sized_layout (same SizedFactorBatch specialization as T), with NumActiveFactors and Evaluate extended for scalar weighting.

T must derive from SizedFactorBatch.

Item parameters. Evaluate forwards factor_ids / num_factor_ids to the wrapped factor; with per-factor weights, item \(t\) is scaled by the weight of its factor \(f(t)\).

Wraps a factor to apply scalar weight(s) to residuals and Jacobians. Supports two modes: a single uniform weight applied to every factor, or per-factor weights from a device array.

\[r_{\mathrm{weighted}} = w \, r,\qquad J_{\mathrm{weighted}} = w \, J\]

Residual

Residual dim

Jacobian

Jacobian dims

Manifold

(see above)

(per \(T\))

(see above)

(per \(T\))

same as wrapped \(T\)

Inputs: Same state layout as the wrapped factor T. Constructor also takes either a single float weight or a const float* device pointer to per-factor weights.

template<class ...Args>
WeightedFactorBatch(
float weight,
Args&&... sized_factor_batch_args
)#

Uniform weight constructor. Multiplies every factor’s residual and Jacobian by the same scalar weight. The batch size is the inner factor’s NumActiveFactors().

Parameters:
  • weight – [in] Scalar weight applied to all factors.

  • sized_factor_batch_args – [in] Constructor arguments forwarded to wrapped factor T.

Returns:

Constructor has no return value.

template<class ...Args>
WeightedFactorBatch(
const float *per_factor_weights,
size_t capacity,
Args&&... sized_factor_batch_args
)#

Per-factor weight constructor. Factor i has its residual and Jacobian multiplied by per_factor_weights[i].

Parameters:
  • per_factor_weights – [in] Device pointer to per-factor weights (at least capacity floats).

  • capacity – [in] Number of weights; must equal T::Capacity() for the constructed inner batch.

  • sized_factor_batch_args – [in] Constructor arguments forwarded to wrapped factor T.

Returns:

Constructor has no return value.