Skip to content
Merged
4 changes: 3 additions & 1 deletion .grok/skills/unsga3-oracle/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,9 +19,11 @@ Repo root: Unsga3. Confirm `Unsga3.slnx` / `tools/OracleCompare` exist before ru
| zdt2 | 12 | 52 | **250** | `--pymoo-mode` (`PymooCompatible`) |
| dtlz2 | 12 | 92 | 150 | `--pymoo-mode` (`PymooCompatible`) |

DTLZ2: C# `Dtlz2Problem(k: 10)` ⇒ **n_var=12**. `run_pymoo_oracle.py` passes `n_var=12`. pymoo’s own default is n_var=10 (k=8). Do not treat the published 15-seed pymoo column as that matched run, and do not rewrite `docs/WILCOXON-RESULTS.md` until the seeds are re-run.

ZDT2 **gens=100** is an early-stress snapshot (collapse on Bend, C#, and pymoo), not the quality bar. Quality protocol matches unsga3-bend A/B (gens=250, PymooCompatible). `RankNicheDistance` is an optional unpublished Wilcoxon ZDT2 mating mode — do not silently switch all ZDT defaults to it. ZDT1 and DTLZ2 unchanged.

IGD = **mean** nearest Euclidean distance (pymoo-compatible). Docs: `docs/EQUIVALENCE.md`, `docs/RESEARCH-STANDARDS.md`.
IGD = **mean** nearest Euclidean distance (pymoo-compatible). C# scores the full non-dominated front; `run_pymoo_oracle.py` scores pymoo `res.F` (niche optimum). Compare them only on a shared front definition and a shared reference set. Docs: `docs/EQUIVALENCE.md`, `docs/RESEARCH-STANDARDS.md`.

Requires: .NET 10 SDK; Python 3 + `pip install pymoo` for pymoo side / multi-seed.

Expand Down
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,16 @@ and this project adheres to [Semantic Versioning](https://semver.org/).
### Changed

- Docs and XML comments describe `initialPopulation` and hybrid loops in generic terms (domain-adapter warm-start / grid-seed). No product-repo names.
- DTLZ2 pymoo oracle passes `n_var=12` (k=10) to match `Dtlz2Problem`. The published 15-seed table used pymoo’s default `n_var=10` and is not rewritten. Seed 1 was remeasured at n_var=12.
- Documented that ZDT IGD compares the C# full non-dominated front with pymoo `res.F`. Added `ReferenceDirectionThinning.OnePerDirection` as a cardinality aid, not a parity claim.
- Documented the collapsed-nadir fallback (nadir = ideal + 1 when the span stays ≤ 1e-6). pymoo 0.6.2 stops at the worst point in the population. Behavior is unchanged and covered by a unit test.
- Documented that default `RankNicheDistance` is not Seada and Deb Algorithm 2. `PymooCompatible` matches the paper's same-niche split; p_c stays 1.0 (paper experiments use 0.9). The default tournament is unchanged.
- Locked the infeasible-point hyperplane rule with a fixture: feasible (1, 1) beside infeasible (0, 0) sets ideal to (0, 0). The rule is unchanged.
- Documented that mating calls `PrepareForSelection`, which re-associates survivors. pymoo keeps the niche ids from survival. A fixture locks the current ids on a five-point pool.
- Documented that `WithDasDennis(1, 1)` throws because N must be at least 2. Single-objective runs pass an explicit population size. N = 1 is not accepted.
- Indicator edges: Euclidean distance rejects a shorter vector instead of ignoring the extra coordinates. IGD+ has a hand-case test. `ParetoFronts.Zdt1(1)` (and the other single-point samplers) throw. ZDT6's 0.280775 floor and ZDT3's 0.1822287280 endpoint are documented and locked.
- Duplicate elimination no longer spins when mutation cannot change the decision vector. The key is `G12` significant digits, not 12 decimal places. After the attempt cap, remaining offspring slots may be duplicates.
- Equivalence docs no longer say the published Wilcoxon table is within 1–2% of pymoo. The DTLZ2 seed-1 guard is 2× the published mismatched scalar 0.00350, so a regression to about 2.9× fails. Das–Dennis `Count` uses a checked 64-bit combination and throws when the value does not fit in `int`. Two-layer reference directions remain absent.

## [0.1.4] — 2026-09-19

Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@
**U-NSGA-III** (Unified NSGA-III) for .NET — single-, multi-, and many-objective evolutionary optimization with Das–Dennis reference directions, SBX crossover, polynomial mutation, and **niching-based tournament selection** ([Seada & Deb, 2016](https://ieeexplore.ieee.org/document/7271063)).

> **v0.1.4** — production-usable core with pymoo-aligned normalization. ZDT2 quality protocol is gens=250 (docs/defaults; no algorithm/API change).
> **15-seed IGD vs pymoo `UNSGA3`:** ZDT1 **median 0.053 vs 0.070** (we win; MWU *p*≈0.05); DTLZ2 **median 0.0045 vs 0.0028** (~1.6×, same order; pymoo still ahead).
> **15-seed IGD vs pymoo `UNSGA3`:** ZDT1 **median 0.053 vs 0.070** (MWU *p*≈0.05) compares the full C# non-dominated front with pymoo `res.F`. DTLZ2 **median 0.0045 vs 0.0028** (~1.6×) compares C# n_var=12 with pymoo's default n_var=10. Neither pair is a same-set, same-problem ranking. Notes: [`docs/ORACLE-RESULTS.md`](docs/ORACLE-RESULTS.md).
> Details: [`docs/WILCOXON-RESULTS.md`](docs/WILCOXON-RESULTS.md) · single-seed notes: [`docs/ORACLE-RESULTS.md`](docs/ORACLE-RESULTS.md)

```text
Expand Down
32 changes: 23 additions & 9 deletions docs/EQUIVALENCE.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Equivalence vs pymoo / MATLAB

Goal: prove this port is a faithful U-NSGA-III (Seada & Deb 2016), not a look-alike.
Goal: state where this port follows Seada & Deb 2016 and pymoo, and where it does not. Survival, Das–Dennis directions, and the SBX/PM shapes are the pymoo-shaped core. The default tournament is not Algorithm 2.

See also **[RESEARCH-STANDARDS.md](RESEARCH-STANDARDS.md)** for the literature + pymoo protocol,
**[ORACLE-RESULTS.md](ORACLE-RESULTS.md)** for single-seed numbers, and
Expand All @@ -17,10 +17,14 @@ See also **[RESEARCH-STANDARDS.md](RESEARCH-STANDARDS.md)** for the literature +

1. **Fixed operators:** SBX η=30, PM η=20, p_c=1.0, p_m=1/n, p_var(SBX)=0.5
2. **Same reference set:** Das–Dennis partitions identical to the oracle
3. **Same pop size / generations / seed** (or 15–31 seeds for statistics)
4. **Metrics:** IGD (primary), IGD+, HV (M=2, document ref point), front plots for M≤3
5. **Tolerance:** median IGD within ~1–2× of pymoo on ZDT/DTLZ is the practical bar.
15-seed: ZDT1 median **better** than pymoo (ratio 0.76, MWU n.s.); DTLZ2 median ~**1.6×** (pymoo still ahead).
3. **Same decision dimension:** DTLZ2 uses k=10 so `n_var = M + k − 1` (12 when M=3). pymoo’s default `n_var=10` (k=8) is a known mismatch; the oracle passes `n_var=12`.
4. **Same pop size / generations / seed** (or 15–31 seeds for statistics)
5. **Metrics:** IGD (primary), IGD+, HV (M=2, document ref point), front plots for M≤3.
Score the **same front definition** and the **same reference set**. C# reports the full non-dominated front. pymoo `res.F` is the survival niche set (about one point per filled direction). `ReferenceDirectionThinning.OnePerDirection` can match cardinality; it does not reproduce `res.F` and is not a parity claim.
6. **Tolerance:** do not read the published table as median IGD within 1–2% of pymoo.
ZDT1 median ratio 0.764107, Mann–Whitney U = 65, p = 0.0512394 (the pymoo column is `res.F`).
DTLZ2 median ratio 1.58638, U = 222, p = 6.15164×10⁻⁶ (the pymoo column is n_var=10). That comparison rejects equal distributions at α = 0.05.
The CI guard is 2× the published mismatched DTLZ2 scalar 0.00350, which fails a regression to about 2.9×. It is not a same-problem equivalence claim.

Published A/B budgets (ZDT1 / DTLZ2 unchanged; ZDT2 matches unsga3-bend protocol honesty):

Expand Down Expand Up @@ -59,17 +63,27 @@ ZDT2 **gens=100** is an early-stress snapshot (collapse on Bend, C#, and pymoo),

| Item | This library | pymoo |
|------|--------------|-------|
| Tournament (default) | rank → niche count → dist | — |
| Tournament (`PymooCompatible`) | same niche → rank/dist; else random | `comp_by_rank_and_ref_line_dist` |
| Duplicate elimination | default **on** | `eliminate_duplicates=True` |
| Tournament (default `RankNicheDistance`) | rank → niche count → dist, including across niches | not Algorithm 2 |
| Tournament (`PymooCompatible`) | same niche → rank then dist; else random; distance tie is a coin flip | `comp_by_rank_and_ref_line_dist` (paper keeps the second parent on a distance tie) |
| SBX p_c | **1.0** (pymoo `SBX(prob=1.0)`) | paper section 4 uses **0.9** |
| Mating pool | N independent tournaments with replacement | two shuffled consecutive-pair passes |
| Niche ids at mating | `PrepareForSelection` re-normalizes survivors and re-associates | ids written during survival are kept |
| `WithDasDennis(1, 1)` | throws. One objective has a single direction, and N must be ≥ 2 | pass `populationSize` ≥ 2 for the single-objective degeneration |
| Reference layers | single-layer Das–Dennis only. Two-layer directions are absent | many-objective NSGA-III adds an inside layer for larger M |
| Duplicate elimination | default **on**. Key is `G12` (12 significant digits), not 12 decimal places. Attempts are capped when mutation cannot produce a new key; remaining slots may be duplicates | `eliminate_duplicates=True` |
| Survival RNG | optional RNG niche pick | random among equal niches |
| IGD | **mean** nearest distance | same (verified pymoo 0.6.2) |
| Scored set | full non-dominated front | `res.F` niche optimum |
| Hyperplane norm | persistent ideal, ND extremes, correct ASF | `HyperplaneNormalization` |
| Collapsed nadir | if the span is still ≤ 1e-6, nadir = ideal + 1 | stop at worst-of-population |
| Infeasible points in the hyperplane | ideal and worst from the whole pool, including infeasible points. ASF extremes use the ND index set when supplied | pymoo niching can restrict the normalized set to feasible members |

## DTLZ2 gap history

| Stage | C# pymoo-mode IGD | vs pymoo 0.0035 |
|-------|-------------------|-----------------|
| Pre-fix (wrong ASF) | 0.017 | ~5× |
| ASF + persistent ideal | 0.0052 | ~1.5× |
| + duplicate elimination | **0.0040** | **~1.15×** |
| + duplicate elimination | **0.0040** | **~1.15× vs pymoo n_var=10** |

The ~1.15× denominator is pymoo at **n_var=10** (k=8), not the C# problem (n_var=12, k=10). A matched seed-1 pair is recorded in [ORACLE-RESULTS.md](ORACLE-RESULTS.md). The 15-seed table is still the mismatched-k run.
54 changes: 49 additions & 5 deletions docs/ORACLE-RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,8 @@ pip install pymoo
python tools/oracle/run_pymoo_oracle.py --problem zdt1 --partitions 12 --pop 52 --gens 100 --seed 1
python tools/oracle/run_pymoo_oracle.py --problem zdt2 --partitions 12 --pop 52 --seed 1
python tools/oracle/run_pymoo_oracle.py --problem dtlz2 --partitions 12 --pop 92 --gens 150 --seed 1
# DTLZ2 default is n_var=12 (k=10), matching Dtlz2Problem. pymoo's own default is n_var=10 (k=8).
# Pass --n-var 10 only to reproduce the historical mismatched column.

# C#
dotnet run --project tools/OracleCompare -c Release -- --problem zdt1 --partitions 12 --pop 52 --gens 100 --seed 1
Expand All @@ -39,8 +41,27 @@ C# never published a hard ZDT2 oracle / Wilcoxon table. The unpublished Wilcoxon

| Problem | Settings | pymoo IGD | C# default IGD | C# `PymooCompatible` IGD | Verdict |
|---------|----------|-----------|----------------|--------------------------|---------|
| **ZDT1** | p=12, pop=52, 100 gen | **0.0629** (n=13 ND) | **0.0514** (n=52) | — | **Default wins** |
| **DTLZ2** | p=12, pop=92, 150 gen | **0.00350** (n=91) | 0.0070 (n=92) | **0.00403** (n=92) | **~1.15× pymoo** (pymoo-mode) |
| **ZDT1** | p=12, pop=52, 100 gen | **0.0629** (`res.F`, n=13) | **0.0514** (full ND front, n=52) | — | Different sets. Not an algorithm ranking. |
| **DTLZ2** | p=12, pop=92, 150 gen, **mismatched k** | **0.00350** (n=91, pymoo **n_var=10**, k=8) | 0.0070 (n=92, n_var=12) | **0.00403** (n=92, n_var=12) | Historical pair only. Not a same-problem ratio. |

### ZDT fronts (same run, different sets)

C# `OracleCompare` scores the **full feasible non-dominated front** (here n=52) against `ParetoFronts.Zdt1(500)`. pymoo's oracle scores **`res.F`**, the survival niche set (here n=13, one per Das–Dennis direction), against pymoo's 100-point `pareto_front()`. The published 0.0514 vs 0.0629 pair is those two reporters. It is not evidence that the algorithm is better by ~0.011 IGD.

Seed 1 remeasured **2026-09-22**, pymoo 0.6.2. The C# console reprinted the published scalar.

| Set | Reference front | n | IGD |
|-----|-----------------|--:|----:|
| C# non-dominated front | library 500-point ZDT1 | 52 | 0.051430749249856716 (console 0.0514307) |
| C# non-dominated front | pymoo 100-point PF | 52 | 0.05119280568479224 |
| pymoo final population, non-dominated | pymoo 100-point PF | 52 | 0.05378307132263516 |
| pymoo `res.F` | pymoo 100-point PF | 13 | 0.0628633417931784 |
| pymoo `res.F` | library 500-point ZDT1 | 13 | 0.06276449352608372 |
| pymoo population ND | library 500-point ZDT1 | 52 | 0.05381261662420749 |

On the shared 100-point PF, the full-front pair is C# 0.05119280568479224 and pymoo 0.05378307132263516. One seed cannot carry a ranking. Switching the C# front from the 500-point sampler to that 100-point PF changes its IGD by 0.051430749249856716 − 0.05119280568479224 = 2.37943565064476×10⁻⁴, which is much smaller than the 13-versus-52 gap on pymoo's own PF (0.0628633417931784 − 0.05378307132263516 = 0.00908027047054324).

`ReferenceDirectionThinning.OnePerDirection` keeps the raw objective vector closest (perpendicular distance) to each Das–Dennis direction. On this C# front that helper kept 13 points and scored **0.06357076535717451** against the 100-point PF. That set is not `res.F`. The 15-seed pymoo column is still `res.F`, so its median ratio inherits the same asymmetry. Do not rewrite [WILCOXON-RESULTS.md](WILCOXON-RESULTS.md) until those seeds are re-run on a shared front definition.

### DTLZ2 multi-seed (C# `PymooCompatible`, same protocol)

Expand All @@ -53,7 +74,27 @@ C# never published a hard ZDT2 oracle / Wilcoxon table. The unpublished Wilcoxon
| 5 | 0.00466 |
| **mean** | **~0.00485** |

All seeds stay in the same band as pymoo’s single-seed 0.0035 (within ~1.4–1.6×).
Those five C# seeds are `Dtlz2Problem(k: 10)` (n_var=12). The 0.0035 figure they were compared with is pymoo at **n_var=10** (k=8). That is not a same-problem band. The 15-seed file is unchanged until a matched re-run (see below).

### DTLZ2 n_var (known mismatch, seed 1 remeasured)

`Dtlz2Problem(nObjectives: 3, k: 10)` builds **n = 12**. Deb et al. suggest k = 10. pymoo 0.6.2 `get_problem("dtlz2", n_obj=3)` defaults to **n_var=10** (k = 8). The harness used to omit `n_var`, so the published seed-1 pair and `docs/WILCOXON-RESULTS.md` compare those two dimensions. `tools/oracle/run_pymoo_oracle.py` now passes **n_var=12**.

Published mismatched seed 1 (already in the Wilcoxon table; not re-interpreted as parity):

| Solver | n_var | k | IGD |
|--------|------:|--:|----:|
| C# `PymooCompatible` | 12 | 10 | 0.00403168 |
| pymoo default | 10 | 8 | 0.00349879 |

Seed 1 remeasured **2026-09-22** with pymoo 0.6.2 after the oracle passes `n_var=12`. Console figures are the `G6` print; the second number is the meta-file value. Front sizes are what each reporter wrote (`res.F` vs full non-dominated front).

| Solver | n_var | k | Console IGD | Meta IGD | Front |
|--------|------:|--:|------------:|---------:|------:|
| C# `PymooCompatible` | 12 | 10 | 0.00403168 | 0.004031675764658275 | 92 |
| pymoo `n_var=12` | 12 | 10 | 0.00308392 | 0.003083921253245871 | 91 |

Ratio of the two meta IGDs: 0.004031675764658275 / 0.003083921253245871 = **1.30732**. That is one seed, and the fronts still differ by one point (92 vs 91). It is not a 15-seed ranking and it does not replace the Wilcoxon table.

## Root cause of the old ~5× DTLZ2 gap (fixed)

Expand All @@ -78,7 +119,7 @@ Deep-dive vs pymoo `HyperplaneNormalization` / `ReferenceDirectionSurvival` (pym
|------|-----|
| ZDT1 seed=1, 100 gen, default tournament | IGD ≤ 1.5 × 0.0629 |
| ZDT2 seed=2, 250 gen, default `RankNicheDistance` | IGD < 0.75 (loose CI smoke, not oracle parity) |
| DTLZ2 seed=1, 150 gen, pymoo-mode | IGD ≤ 3 × 0.00350 (currently ~1.15×) |
| DTLZ2 seed=1, 150 gen, pymoo-mode | IGD ≤ 2 × 0.00350. The 0.00350 scalar is the mismatched n_var=10 run. 3× still passed a regression to about 2.9×. This bar does not claim same-problem equivalence. |
| DTLZ2 short smoke (80 gen) | IGD < 0.15 |

ZDT2 quality A/B is gens=250 + `PymooCompatible` (not the loose smoke bar). ZDT1 / DTLZ2 shipping bars are unchanged.
Expand All @@ -88,7 +129,10 @@ ZDT2 quality A/B is gens=250 + `PymooCompatible` (not the loose smoke bar). ZDT1
| Item | Status |
|------|--------|
| IGD mean-distance | **aligned** |
| ASF / hyperplane normalization | **aligned** |
| ASF / axis intercepts | **aligned** |
| Collapsed nadir (span ≤ 1e-6) | **delta**: nadir = ideal + 1 after the worst-of-pop fallback. pymoo 0.6.2 stops at worst-of-pop. Locked by `Collapsed_span_sets_nadir_to_ideal_plus_one` (`{2, 2+1e-8}` → nadir 3). |
| Infeasible points | **current rule, locked**: ideal and worst include them. Fixture is feasible (1, 1) vs infeasible (0, 0) → ideal (0, 0). No constrained benchmark yet. |
| Mating re-association | **delta**: after survival, `PrepareForSelection` normalizes the survivors again and overwrites niche ids. pymoo keeps the survival ids. Fixture: `PrepareForSelection_overwrites_survival_niche_ids`. |
| Persistent ideal + ND extremes | **aligned** |
| `TournamentMode.PymooCompatible` | **implemented** |
| Duplicate elimination | **implemented** (default on) |
Expand Down
Loading
Loading