Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
80 commits
Select commit Hold shift + click to select a range
1038138
direct KKT matrix support
yuwenchen95 Aug 18, 2026
e2572f7
cache reuse for linear objective updates
Iroy30 Aug 21, 2026
2c31258
cleanup
Iroy30 Aug 28, 2026
fa633cd
Fix device_scalar constructors after RMM change (#1795)
aliceb-nv Aug 25, 2026
62d595b
Fix CUDA 13 settings_ pointer access in sparse_cholesky.
Aug 31, 2026
13a5215
fix merge conflicts
Iroy30 Aug 31, 2026
4270b5c
resolve conflicts
Iroy30 Sep 1, 2026
050783d
move rmm to gpu path
Iroy30 Sep 1, 2026
6e20206
remove log files
Iroy30 Sep 1, 2026
7d61502
address reviews part 1
Iroy30 Sep 1, 2026
698afbe
remove logs
Iroy30 Sep 1, 2026
5cf8c5f
address reviews part 2
Iroy30 Sep 2, 2026
0ae8190
copyright update
Iroy30 Sep 2, 2026
d1b4b73
address reviews part 3
Iroy30 Sep 2, 2026
43b9b29
restructure and rename barrier.cu solve functions
Iroy30 Sep 3, 2026
0499d81
formatting checks fix
Iroy30 Sep 3, 2026
260d6a1
fix comments
Iroy30 Sep 3, 2026
faf0aa2
comment cleanup
Iroy30 Sep 3, 2026
bf7d115
Merge origin/main into cuopt_cache_reuse_FSI
Iroy30 Sep 8, 2026
854b857
Add GPU ruiz scaling
yuwenchen95 Sep 11, 2026
8711e6b
barrier: skip the A^T transpose on the ADAT path, the AD copy on the …
yuwenchen95 Sep 11, 2026
bf7f479
address reviews part 3
Iroy30 Sep 14, 2026
84b0bbb
Add update_rhs for barrier cache reuse
Iroy30 Sep 14, 2026
d9a579b
Restore the barrier sequence-update API design notes
Iroy30 Sep 14, 2026
a1a936f
Merge branch 'main' of https://github.com/NVIDIA/cuopt into add_updat…
Iroy30 Sep 14, 2026
b00c416
Gate update_rhs on range rows, not on slacks
Iroy30 Sep 14, 2026
8c66a93
Add sequence_solve tests for update_rhs
Iroy30 Sep 15, 2026
2782e35
Point the update API notes at the new sequence_solve tests
Iroy30 Sep 15, 2026
69191a6
barrier: replace the merge-sort device CSC→CSR with a scatter plus se…
yuwenchen95 Sep 15, 2026
9ea7635
Merge branch 'main' into gpu-scaling
yuwenchen95 Sep 15, 2026
f6da09d
barrier: build cusparse_view's CSR on device instead of on the host, …
yuwenchen95 Sep 15, 2026
64e84e8
barrier: skip the second CSC->CSR conversion in the ADAT path when th…
yuwenchen95 Sep 15, 2026
6a2accb
barrier: keep the Ruiz-scaled A and Q on device for SOCP instead of d…
yuwenchen95 Sep 15, 2026
efc0f55
barrier: drop device_A and d_original_A_values on the ADAT path when …
yuwenchen95 Sep 16, 2026
720ad1b
Merge remote-tracking branch 'origin/main' into gpu-scaling
yuwenchen95 Sep 16, 2026
fa03c34
Drop the no-op branch in the crush_user_rhs empty-row case
Sep 15, 2026
1e4d92b
Gate barrier cache reuse on the cache, not the current free-variable …
Sep 15, 2026
a40e97c
Update the sequence-update notes with what the free-variable fix esta…
Sep 15, 2026
7bba0b3
Merge branch 'main' of https://github.com/NVIDIA/cuopt into add_updat…
Sep 17, 2026
6d41719
Drop the Cython declarations left dead by the sequence_solve paramete…
Sep 17, 2026
8895540
Merge branch 'main' into gpu-scaling
yuwenchen95 Sep 18, 2026
5ce0235
Code cleanup
yuwenchen95 Sep 18, 2026
9b8e14a
clean up doc
Iroy30 Sep 18, 2026
a2ac329
clean up doc
Iroy30 Sep 18, 2026
d12fd7f
clean up doc
Iroy30 Sep 18, 2026
117cd99
Move device A,Q out of lp_data struct
yuwenchen95 Sep 18, 2026
648f805
Merge branch 'pr-1913' into test/rhs-pr1913
Sep 21, 2026
78884b0
enable socp for update rhs and lin obj
Iroy30 Sep 22, 2026
0e822d1
Merge upstream/main into add_update_apis.
Iroy30 Sep 30, 2026
d1880ab
Merge PR #1739 CSR augmented-matrix iterative refinement
Iroy30 Oct 1, 2026
e8fe40e
Merge PR #1941 barrier update_rhs API
Iroy30 Oct 2, 2026
80c103c
Merge PR #1979 SOCP support for barrier update APIs
Iroy30 Oct 2, 2026
31f6f97
Address barrier RHS update review comments.
Iroy30 Oct 2, 2026
3aeb182
address doc reviews, add return enum and template
Iroy30 Oct 5, 2026
9f6136b
formatting
Iroy30 Oct 5, 2026
9a2a6dd
Merge branch 'main' into add_update_apis
Iroy30 Oct 5, 2026
cdbd082
Merge remote-tracking branch 'origin/main'
Iroy30 Oct 5, 2026
d40feac
merge main
Iroy30 Oct 5, 2026
883b644
remove logs
Iroy30 Oct 5, 2026
4b77fac
address reviews
Oct 5, 2026
347276b
formatting
Iroy30 Oct 5, 2026
e5b11c6
Revert "Merge PR #1739 CSR augmented-matrix iterative refinement"
Iroy30 Oct 5, 2026
b92ede5
Merge branch 'pr-1941'
Iroy30 Oct 5, 2026
adf9cd6
Merge origin/main
Iroy30 Oct 6, 2026
bef9925
merge main
Iroy30 Oct 6, 2026
757e9eb
remove reverted
Iroy30 Oct 6, 2026
236ea7c
reword docs
Iroy30 Oct 6, 2026
8ba7357
fix merge issues
Iroy30 Oct 6, 2026
51b7b9d
clean templates i_t, f_t
Iroy30 Oct 6, 2026
5cb041a
address template and doc review
Iroy30 Oct 6, 2026
34f9782
address naming reviews
Iroy30 Oct 8, 2026
438fe0e
revert barrier_cache_t template
Iroy30 Oct 8, 2026
94dafd4
address reviews
Iroy30 Oct 8, 2026
3234877
address reviews
Iroy30 Oct 8, 2026
9a2372f
set eliminate_free_variables to False for sequence solve
Iroy30 Oct 8, 2026
cc9209b
address reviews Oct 8
Iroy30 Oct 9, 2026
2e1a10e
Merge origin/main into main.
Iroy30 Oct 9, 2026
5fec35c
address reviews
Iroy30 Oct 9, 2026
2f022ba
address format review and record cone head bounds in translate_soc
Iroy30 Oct 9, 2026
17173b9
address reviews and style formatting
Iroy30 Oct 9, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,7 @@ void apply_barrier_rhs(iteration_data_t<int, double>& data, double const* barrie
namespace cuopt {
namespace CUOPT_EXPORT mathematical_optimization {

template <typename i_t, typename f_t>
struct barrier_transform_t;

/**
Expand Down Expand Up @@ -63,9 +64,10 @@ class barrier_cache_t {
*/
barrier::iteration_data_t<int, double>* release_iteration_data();

void store_transform(std::unique_ptr<barrier_transform_t> transform);
[[nodiscard]] barrier_transform_t* transform();
[[nodiscard]] barrier_transform_t const* transform() const;
template <typename i_t, typename f_t>
void store_transform(std::unique_ptr<barrier_transform_t<i_t, f_t>> transform);
[[nodiscard]] barrier_transform_t<int, double>* transform();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We still have int, double here

[[nodiscard]] barrier_transform_t<int, double> const* transform() const;
/** True when an update API has staged new data that the next solve should reuse. */
[[nodiscard]] bool dirty() const;
void mark_clean();
Expand All @@ -77,13 +79,15 @@ class barrier_cache_t {
* Crush the input linear objective into cached iteration_data_t.c / d_c_ and mark dirty.
* Requires a stored transform and iteration_data from an Optimal solve.
*/
void update_linear_objective(double const* c, int n);
template <typename i_t, typename f_t>
void update_linear_objective(f_t const* c, i_t n);

/**
* Crush the input constraint RHS into cached iteration_data_t.b / d_b_ and mark dirty.
* Requires a stored transform and iteration_data from an Optimal solve.
*/
void update_rhs(double const* b, int m);
template <typename i_t, typename f_t>
void update_rhs(f_t const* b, i_t m);

private:
barrier_cache_t(std::unique_ptr<rmm::cuda_stream> stream, std::unique_ptr<raft::handle_t> handle);
Expand Down
55 changes: 47 additions & 8 deletions cpp/src/barrier/barrier.cu
Original file line number Diff line number Diff line change
Expand Up @@ -932,7 +932,7 @@ class iteration_data_t {
{
raft::common::nvtx::range form_scope("Barrier: LP Data: form augmented");
// Build the sparsity pattern of the augmented system
form_augmented(true);
form_augmented(augmented_form_t::build);
}
if (settings.concurrent_halt != nullptr && *settings.concurrent_halt == 1) { return; }
symbolic_status = chol->analyze(device_augmented);
Expand All @@ -948,9 +948,10 @@ class iteration_data_t {
}

// Attach this solve's settings and rewind iterate-dependent state so barrier can
// start with the new c / b. A and Q are unchanged; the previous solve
// left D and the KKT values at its last iterate. Reuse is QP-only (no cones),
// so form_*(false) updates values in the existing CSR; no symbolic rebuild.
// start with the new c / b. A and Q are unchanged. The previous solve left D
// and the KKT values at its last iterate. reset_for_reuse (or form_adat(false)) rewrites
// values in the existing CSR. The cone block is put back to the initial diagonal a cold
// start factorizes; Nesterov-Todd scaling is recomputed from the new point.
bool reset_iterate_state(const simplex_solver_settings_t<i_t, f_t>& settings)
{
if (chol == nullptr || symbolic_status != 0) { return false; }
Expand Down Expand Up @@ -987,7 +988,7 @@ class iteration_data_t {
}

if (use_augmented) {
form_augmented(false);
form_augmented(augmented_form_t::reset_for_reuse);
} else {
form_adat(false);
}
Expand Down Expand Up @@ -1063,7 +1064,12 @@ class iteration_data_t {
return degree;
}

void form_augmented(bool first_call = false)
// build: first call, device CSR and metadata.
// reset_for_reuse: values only; cone block back to the cold-start matrix.
// update: values only; cone block from the current Nesterov-Todd scaling.
enum class augmented_form_t { build, reset_for_reuse, update };

void form_augmented(augmented_form_t mode = augmented_form_t::update)
{
i_t n = A.n;
i_t m = A.m;
Expand All @@ -1074,7 +1080,7 @@ class iteration_data_t {
const i_t p = augmented_expansion_count();
i_t factorization_size = augmented_system_size(n, m);

if (first_call) {
if (mode == augmented_form_t::build) {
raft::common::nvtx::range scope("Barrier: augmented: device CSR build");

const size_t n_sparse_cone_entries =
Expand Down Expand Up @@ -1165,7 +1171,40 @@ class iteration_data_t {
});
RAFT_CHECK_CUDA(handle_ptr->get_stream().get());

if (has_soc) {
if (has_soc && mode == augmented_form_t::reset_for_reuse) {
// Cold initial_point factorizes this diagonal, then the first Newton step
// rebuilds the Nesterov-Todd Hessian. Zero w and eta so a dense block
// scatter and the matrix-free product both see that same initial matrix.
auto stream = handle_ptr->get_stream();
thrust::fill(rmm::exec_policy(stream), cones().w.begin(), cones().w.end(), f_t(0));
thrust::fill(rmm::exec_policy(stream), cones().eta.begin(), cones().eta.end(), f_t(0));
if (cones().has_sparse_cones()) {
restore_initial_sparse_cone_block(cones(),
device_augmented.x,
cone_kkt_data_.sparse_Hs_diag,
cone_kkt_data_.sparse_hessian_diag,
cone_kkt_data_.sparse_hessian_Q,
cone_kkt_data_.sparse_exp_v_col,
cone_kkt_data_.sparse_exp_u_col,
cone_kkt_data_.sparse_exp_v_row,
cone_kkt_data_.sparse_exp_u_row,
cone_kkt_data_.sparse_expansion_D,
stream,
dual_perturb);
RAFT_CHECK_CUDA(stream.get());
}
if (cones().n_dense_cones() > 0) {
scatter_dense_hessian_into_augmented(cones(),
device_augmented.x,
cone_kkt_data_.cone_csr_indices,
cone_kkt_data_.cone_Q_values,
cone_kkt_data_.dense_block_offsets,
cone_kkt_data_.dense_cone_ids,
stream,
dual_perturb);
RAFT_CHECK_CUDA(stream.get());
}
} else if (has_soc) {
if (cones().has_sparse_cones()) {
scatter_sparse_hessian_into_augmented(cones(),
device_augmented.x,
Expand Down
55 changes: 35 additions & 20 deletions cpp/src/barrier/barrier_cache.cu
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,9 @@ using barrier_iteration_data_t = barrier::iteration_data_t<int, double>;
using barrier_iteration_data_ptr =
std::unique_ptr<barrier_iteration_data_t, void (*)(barrier_iteration_data_t*)>;

static void require_cache(barrier_transform_t const* transform,
barrier_iteration_data_t const* data,
template <typename i_t, typename f_t>
static void require_cache(barrier_transform_t<i_t, f_t> const* transform,
barrier::iteration_data_t<i_t, f_t> const* data,
char const* api)
{
cuopt_expects(transform != nullptr,
Expand All @@ -38,7 +39,8 @@ static void require_cache(barrier_transform_t const* transform,
}

// Re-adds the first solve's barrier-minus-crush shift so the update lands in the presolved model.
static void add_shift(std::vector<double>& crushed, std::vector<double> const& shift)
template <typename f_t>
static void add_shift(std::vector<f_t>& crushed, std::vector<f_t> const& shift)
{
cuopt_expects(shift.size() == crushed.size(),
error_type_t::ValidationError,
Expand All @@ -59,7 +61,7 @@ struct barrier_cache_t::impl {
std::unique_ptr<rmm::cuda_stream> stream;
std::unique_ptr<raft::handle_t> handle;
// Destroy iteration_data before transform: it may const-ref A/Q stored on the transform.
std::unique_ptr<barrier_transform_t> transform;
std::unique_ptr<barrier_transform_t<int, double>> transform;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

int -> i_t, double -> f_t

barrier_iteration_data_ptr iteration_data;
bool linear_objective_dirty{false};
bool rhs_dirty{false};
Expand Down Expand Up @@ -107,14 +109,18 @@ barrier_iteration_data_t* barrier_cache_t::release_iteration_data()
return impl_->iteration_data.release();
}

void barrier_cache_t::store_transform(std::unique_ptr<barrier_transform_t> transform)
template <typename i_t, typename f_t>
void barrier_cache_t::store_transform(std::unique_ptr<barrier_transform_t<i_t, f_t>> transform)
{
impl_->transform = std::move(transform);
}

barrier_transform_t* barrier_cache_t::transform() { return impl_->transform.get(); }
barrier_transform_t<int, double>* barrier_cache_t::transform() { return impl_->transform.get(); }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

int -> i_t, double -> f_t


barrier_transform_t const* barrier_cache_t::transform() const { return impl_->transform.get(); }
barrier_transform_t<int, double> const* barrier_cache_t::transform() const

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

int -> i_t, double -> f_t

{
return impl_->transform.get();
}

bool barrier_cache_t::dirty() const
{
Expand All @@ -131,19 +137,20 @@ void barrier_cache_t::mark_clean()

bool barrier_cache_t::rhs_infeasible() const { return impl_->rhs_infeasible; }

void barrier_cache_t::update_linear_objective(double const* c, int n)
template <typename i_t, typename f_t>
void barrier_cache_t::update_linear_objective(f_t const* c, i_t n)
{
require_cache(impl_->transform.get(), impl_->iteration_data.get(), "update_linear_objective");
// Cached Q and c are in minimization space.
std::vector<double> user_objective;
std::vector<f_t> user_objective;
if (impl_->transform->maximize && c != nullptr && n > 0) {
user_objective.assign(c, c + n);
for (double& value : user_objective) {
for (f_t& value : user_objective) {
value = -value;
}
c = user_objective.data();
}
std::vector<double> crushed;
std::vector<f_t> crushed;
try {
crushed = crush_user_linear_objective(*impl_->transform, c, n);
} catch (std::invalid_argument const& e) {
Expand All @@ -154,35 +161,36 @@ void barrier_cache_t::update_linear_objective(double const* c, int n)
auto const& linear_obj_shift = impl_->transform->linear_obj_shift;
auto const& column_scales = impl_->transform->column_scales;
auto const& translated_lower = impl_->transform->presolve_info.removed_lower_bounds;
simplex::lp_problem_t<int, double>& barrier_lp = *impl_->transform->barrier_lp;
simplex::lp_problem_t<i_t, f_t>& barrier_lp = *impl_->transform->barrier_lp;
if (!translated_lower.empty() && linear_obj_shift.size() == crushed.size() &&
column_scales.size() == crushed.size() && barrier_lp.objective.size() == crushed.size()) {
double obj_constant_delta = 0.0;
f_t obj_constant_delta = 0.0;
std::size_t const n_lower = std::min(translated_lower.size(), crushed.size());
for (std::size_t j = 0; j < n_lower; ++j) {
double const crushed_before = barrier_lp.objective[j] - linear_obj_shift[j];
f_t const crushed_before = barrier_lp.objective[j] - linear_obj_shift[j];
obj_constant_delta += (crushed[j] - crushed_before) * column_scales[j] * translated_lower[j];
}
barrier_lp.obj_constant += obj_constant_delta;
}
add_shift(crushed, linear_obj_shift);
// The next solve builds its solver from barrier_lp, so keep its objective and the cached
// iteration workspace on the same c.
std::vector<double>& barrier_objective = barrier_lp.objective;
std::vector<f_t>& barrier_objective = barrier_lp.objective;
cuopt_expects(barrier_objective.size() == crushed.size(),
error_type_t::ValidationError,
"update_linear_objective: crushed objective size does not match the cached "
"barrier LP.");
barrier_objective = crushed;
barrier::apply_barrier_linear_objective(
*impl_->iteration_data, crushed.data(), static_cast<int>(crushed.size()));
*impl_->iteration_data, crushed.data(), static_cast<i_t>(crushed.size()));
impl_->linear_objective_dirty = true;
}

void barrier_cache_t::update_rhs(double const* b, int m)
template <typename i_t, typename f_t>
void barrier_cache_t::update_rhs(f_t const* b, i_t m)
{
require_cache(impl_->transform.get(), impl_->iteration_data.get(), "update_rhs");
std::vector<double> crushed;
std::vector<f_t> crushed;
std::string error;
crush_rhs_status_t const status = crush_user_rhs(*impl_->transform, b, m, crushed, error);
if (status == crush_rhs_status_t::infeasible) {
Expand All @@ -198,14 +206,21 @@ void barrier_cache_t::update_rhs(double const* b, int m)
add_shift(crushed, impl_->transform->rhs_shift);
// barrier_lp->rhs also seeds the next solve's Mehrotra start, so keep it and the cached
// workspace on the same b.
std::vector<double>& barrier_rhs = impl_->transform->barrier_lp->rhs;
std::vector<f_t>& barrier_rhs = impl_->transform->barrier_lp->rhs;
cuopt_expects(barrier_rhs.size() == crushed.size(),
error_type_t::ValidationError,
"update_rhs: crushed RHS size does not match the cached barrier LP.");
barrier_rhs = crushed;
barrier::apply_barrier_rhs(
*impl_->iteration_data, crushed.data(), static_cast<int>(crushed.size()));
*impl_->iteration_data, crushed.data(), static_cast<i_t>(crushed.size()));
impl_->rhs_dirty = true;
}

#ifdef DUAL_SIMPLEX_INSTANTIATE_DOUBLE

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we really need #ifdef here?

@Iroy30 Iroy30 Oct 9, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was included to addtress #1979 (comment)

template void barrier_cache_t::store_transform<int, double>(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We usually wrap template initialization in an ifdef. Please take a look at some of the other files in the barrier directory for an example

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

updated

std::unique_ptr<barrier_transform_t<int, double>>);
template void barrier_cache_t::update_linear_objective<int, double>(double const*, int);
template void barrier_cache_t::update_rhs<int, double>(double const*, int);
#endif

} // namespace cuopt::mathematical_optimization
Loading
Loading