From b8f94477c24f2393d7c705285a58b87ccd68db6e Mon Sep 17 00:00:00 2001 From: Polichinl Date: Mon, 24 Aug 2026 13:29:18 +0200 Subject: [PATCH 1/6] =?UTF-8?q?docs(adr):=20ADR-029=20=E2=80=94=20why=20th?= =?UTF-8?q?e=20frames=20are=20numpy,=20and=20not=20xarray?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The container choice was never written down. Roughly 9,400 lines of governance prose in this repository and the word "xarray" appeared nowhere in it, so the first question any newcomer asks had no answer — which is how a settled decision gets relitigated by accident. Eight alternatives, each judged against one constraint: every dependency the DAG root takes, every consumer takes, and one consumer is an HTTP API server. Three claims were checked against live package metadata rather than asserted from memory, and two of them came back different from the first draft: - **xarray** declares `pandas>=2.2` as a hard requirement (2026.7.0). That is the whole argument against it — ADR-002's numpy floor exists precisely so a model repo does not transitively install the pandas world, and faoapi #242 is still undoing that on the consumer side. - **torch** 2.13.0 declares the CUDA stack as *unconditional* on Linux — cuda-toolkit, cudnn, cusparselt, nccl, nvshmem, triton — not as an optional extra. Stronger than the draft claimed. The second argument is better still: this platform already migrated away from torch in this exact domain, and the reconcile parity test's own docstring records that it "needs neither torch" because the oracle is a fixture. - **polars** was grouped with pandas and xarray on dependency weight. That was wrong. polars 1.44.0 declares one hard dependency, its own compiled runtime. The ADR now says so explicitly, because leaving a wrong reason in place would have made the right conclusion unciteable. Protocols-only is closed by measurement rather than argument: ADR-028 recorded that `runtime_checkable` checks member presence only, which is how a frame passed as an index for eleven releases. **A live decision, not a closed one.** Each alternative that could return names the evidence that would return it — a C-71 or C-73 OOM receipt makes zarr/dask worth costing *inside* `io/` without touching the frame type; a non-Python consumer makes format-as-contract worth costing. Four alternatives get no trigger, because they lose on grounds evidence does not change. The Consequences section names what declining xarray actually costs: no out-of-core path, no groupby/resample, hand-written alignment, numpy version skew as our problem (C-19/C-24), and ~3,700 lines someone else would have maintained. No code change. CONFORMANCE_FLOOR stays 2.0.0. Co-Authored-By: Claude Opus 5 (1M context) --- docs/ADRs/029_why_numpy_and_not_xarray.md | 299 ++++++++++++++++++++++ docs/ADRs/README.md | 1 + 2 files changed, 300 insertions(+) create mode 100644 docs/ADRs/029_why_numpy_and_not_xarray.md diff --git a/docs/ADRs/029_why_numpy_and_not_xarray.md b/docs/ADRs/029_why_numpy_and_not_xarray.md new file mode 100644 index 0000000..d76d31d --- /dev/null +++ b/docs/ADRs/029_why_numpy_and_not_xarray.md @@ -0,0 +1,299 @@ +# ADR-029: The frames are numpy arrays plus explicit identifier arrays — not xarray, not a DataFrame, not a tensor + +**Status:** Accepted (a live decision — see *What would change the answer*) +**Date:** 2026-08-24 +**Deciders:** VIEWS platform maintainers +**Consulted:** ADR-002's dependency rules; the 2026-06 leaf and summarize postmortems; ADR-028's measured evidence on structural typing +**Informed:** every repository that pins `views-frames` + +--- + +## In short + +A frame is a numpy array with a small object holding two integer identifier arrays beside it. +Nothing else. The only runtime dependency this package has is `numpy>=1.26,<3`. + +**The obvious question is why this is not xarray**, which does labeled dimensions and +coordinate alignment properly, is maintained by people who think about it full time, and +would have replaced most of `SpatioTemporalIndex` with `.sel()`. This ADR answers that, and +the seven other alternatives a reasonable person would raise. + +It was never written down before. Roughly 9,400 lines of governance prose in this repository +and the word "xarray" appeared nowhere in it until this document — so a newcomer asking the +first question anyone asks found no answer, which is how a settled decision gets relitigated +by accident. + +**Nothing changes as a result of this ADR.** It records a decision already made and names the +evidence that would reopen it. + +--- + +## What the leaf actually has to be + +Every alternative below is judged against one constraint, so it is worth stating first. + +`views-frames` sits at the **root of the platform dependency DAG**. Every other repository +depends toward it and nothing depends away from it. That is not an accident of layout, it is +ADR-002's decision, and it has a consequence most container choices ignore: + +> **Every dependency this package takes, every consumer takes.** + +`ADR-002` states the rule directly: + +> The numpy-only floor (CRP) means a model that wants a `PredictionFrame` does not +> transitively install the pandas/reporting world. + +The consumers are not all research code. **views-faoapi is an HTTP API server.** Whatever this +package imports, that server installs and ships. So the question for each alternative is not +"is it a better array container" — several are — but "is it a better array container *by +enough* to justify putting itself, its transitive dependencies, its release cadence and its +pin surface into every repository on the platform, including the ones that only want to +describe a forecast over HTTP". + +Two further constraints, both earned rather than assumed: + +- **The axis meaning must not be able to drift.** ADR-012 fixes the trailing axis as the sample + axis. The platform's worst incidents traced to code that inferred which axis was time or + space and then drifted silently when the data changed shape. A container where an axis is + identified by a string that anyone can rename is a container where that can happen again. +- **Failure must be loud and at construction.** ADR-008. A container that accepts and coerces + by default is working against this. + +--- + +## Decision + +The three frames hold a `float32` numpy array and a `SpatioTemporalIndex`, which holds two +`int64` identifier arrays and a `SpatialLevel`. Alignment is hand-written numpy +(`searchsorted` over a packed row view). The only runtime dependency is numpy. + +`pyarrow` is permitted, as an **optional extra**, **only** under `io/` — it is a codec, never +the frame type. That boundary is enforced by `tests/test_import_enforcement.py` and by the +`import-linter` contracts in `pyproject.toml`. + +--- + +## Alternatives considered + +### 1. xarray — the real contender + +`xr.DataArray` with dims `(row, sample)` and coordinates for `time` and `unit` is very close to +what this package hand-rolls. `.sel()` replaces `searchsorted`. Automatic coordinate alignment +replaces `reindex`. `xr.align` replaces the same-level join. The cross-level remap would still +be consumer-injected, because that is a domain rule rather than a container feature. Broadcast +semantics, `groupby`, `resample`, and unit-aware slicing all arrive free, and several of them +are things this package deliberately does not offer. + +**It loses on the dependency, and only on the dependency.** xarray declares `pandas>=2.2` as a +hard requirement (checked against xarray 2026.7.0 on 2026-08-24, not assumed). Adopting it +puts pandas back into every consumer of the contract — the precise outcome `ADR-002` names as +the reason for the numpy floor, and the thing views-faoapi #242 is still working to undo on the +consumer side. A leaf that re-imports what its consumers are removing is not a leaf. + +**The secondary loss is that xarray's flexibility is this package's liability.** In xarray a +dimension is a string. Nothing prevents `sample` becoming `draw`, nothing prevents it moving +position, nothing prevents a frame arriving with the dimension absent. ADR-012 exists because +that class of drift already cost this platform real incidents. Enforcing the sample-axis +invariant on top of xarray would mean validating at every boundary — which is the work this +package does anyway, minus the dependency. + +**What is genuinely lost by declining it** — stated plainly, because this is the alternative +that costs the most to refuse: + +- lazy and out-of-core evaluation via dask, which is the direct answer to register C-71 and + C-73; +- zarr and netCDF, which would make `io/` mostly disappear; +- `groupby`/`resample`/interpolation, all currently absent or hand-rolled; +- roughly 3,700 lines of source that someone else would maintain. + +This is the alternative most likely to come back, and *What would change the answer* below says +under exactly what evidence. + +### 2. No in-memory type at all — the file format is the contract + +Publish a schema (parquet, zarr or netCDF plus a written layout spec) and let every consumer +use whatever container it likes. ADR-013's wire contract is already half of this. + +It is the most intellectually honest alternative, because a *data contract* arguably should be +about bytes rather than Python objects, and it is the only option here that serves a non-Python +consumer. + +**It loses because the contract stops being executable.** ADR-016's mechanism is that consumers +run *the same* `assert_*` functions in their own CI, so "all consumers agree" is checked rather +than asserted. With format-as-contract every consumer writes its own loader and its own +validation, and the drift this repository spent two audits closing reappears in six places at +once instead of one. It also gives up fail-loud-at-construction entirely: a malformed frame +becomes a malformed *file*, discovered whenever someone happens to look. + +The strongest evidence against it is local. 2.0.0 exists because a frame could be built with a +non-index and the published checker certified it. That was caught because there *is* a shared +constructor and a shared checker to attack. Under format-as-contract there would have been six +independent versions of that bug and no single place to fix it. + +### 3. torch.Tensor + +Superficially attractive: the model repositories already live in torch, so a torch-native frame +would remove a conversion at the boundary that matters most. + +**It loses twice, and the second reason is the interesting one.** + +First, weight — and the measurement is worse than the intuition. torch 2.13.0 declares the CUDA +stack as **unconditional** hard dependencies on Linux, not as an optional extra: +`cuda-toolkit[cublas,cudart,cufft,cufile,cupti,curand,cusolver,cusparse,nvjitlink,nvrtc,nvtx]`, +`nvidia-cudnn-cu13`, `nvidia-cusparselt-cu13`, `nvidia-nccl-cu13`, `nvidia-nvshmem-cu13` and +`triton` (checked 2026-08-24). Putting that at the DAG root means views-faoapi — an API server +whose job is to serialize a forecast over HTTP — pulls the CUDA toolchain to do it, on every +deploy, with no way to opt out short of a fork. This is the xarray argument about two orders of +magnitude worse. + +Second, and more decisive: **this platform already migrated away from torch in exactly this +domain.** `views_frames_reconcile` exists to be a *numpy* reconciler with bit-parity against +views-reporting's *torch* oracle. The parity test's own docstring records that it "needs neither +torch" at runtime, because the oracle's output is captured as a generated fixture. Torch is not +a runtime dependency, not a test dependency, and not an optional extra — it is a historical +reference implementation that was deliberately replaced. Reintroducing it as the frame type +would reverse a completed migration. + +There is also no technical pull. The leaf does no gradient work: `views_frames_summarize` is +sample-axis reduction, not a differentiable path, by charter (ADR-017). A tensor buys autograd +and device placement, and this package needs neither. + +### 4. pyarrow `Table` as the frame type + +Nearly free, since `io/arrow` already exists. Zero-copy, language-neutral, excellent IO, and it +would collapse the in-memory type and the wire format into one thing. + +**It loses on memory model.** Arrow is columnar; a posterior is a dense `(N, S)` float32 block +and a feature frame is `(N, F, S)`. Every estimator in `views_frames_summarize` sorts, slices +and reduces along the trailing axis — operations that are natural on a strided buffer and +awkward on a column store. The tower estimators in particular work in row blocks to bound peak +memory (C-22); expressing that over Arrow columns would be a fight on every function. + +Note this was already decided implicitly: `pyarrow` is an optional extra restricted to `io/` by +`pyproject.toml` and enforced by two independent mechanisms. This ADR makes that explicit rather +than leaving it inferable from an import contract. + +### 5. Protocols only — no concrete frame type + +Ship `Frame`, `Sampled`, `SpatioTemporalIndexed` and `Persistable` as `Protocol`s and let each +consumer bring its own container. Maximum flexibility, zero opinion about storage, and the +protocols already exist in `protocols.py`. + +**This one is closed by measurement rather than argument.** ADR-028 records what happened when +the contract relied on structural typing: the frame constructors accepted anything exposing +`n_rows`, so a *frame* passed as an *index* for eleven releases, and the published conformance +suite certified the result. `runtime_checkable` protocols check member **presence**, never +member type or semantics — which is exactly the hole. + +A protocols-only design makes that failure mode the whole design rather than a bug in it. The +protocols remain valuable as the surface consumers *type against* (ADR-009); they are not +sufficient as the surface consumers *construct*. + +### 6. pandas DataFrame with a MultiIndex + +The historical baseline, and the thing the platform actively left. + +A `(time, unit)` MultiIndex over a wide frame of `S` sample columns is the natural encoding, and +it is what several consumers used before the leaf existed. It loses on three counts, all with +receipts in this repository: + +- **The list-in-cell encoding.** Storing a posterior as a Python list per cell is the obvious + pandas move and it is a memory catastrophe — register C-40/C-66 record the blow-up, and + README §7 bans object dtype outright as a result. +- **Dependency weight**, the same CRP argument as xarray, with pandas being the specific package + named in ADR-002. +- **Index semantics that are too permissive.** A pandas index silently accepts duplicates, + reorders on join, and coerces dtypes — every one of which is a behaviour this package has an + explicit fail-loud rule against (C-21 on row uniqueness, ADR-008 on coercion). + +### 7. polars + +Arrow-backed, fast, and notably *without* pandas' index concept — which removes one of the three +objections above. Identifier columns would carry `time` and `unit` explicitly, which is closer +to this package's philosophy than pandas is. + +It loses for the same reason pyarrow does — columnar memory against a dense sample axis. + +**The dependency argument does not apply here, and saying so matters.** An earlier draft of this +ADR grouped polars with pandas and xarray on weight; that was wrong. polars 1.44.0 declares one +hard dependency, its own compiled runtime — no pandas, no numpy requirement of its own. It is a +genuinely light install. The case against it is the memory model alone, and it is the closest +any DataFrame library comes to being viable here. + +### 8. awkward-array + +The ragged-data library. Dismissed in a sentence: it would make the variable-length per-cell +encoding *efficient*, and this package bans that encoding on purpose. The data is rectangular by +construction — a dense grid of cells by a fixed sample count — and a container optimised for +raggedness would legitimize the shape C-40 exists to prevent. + +--- + +## What would change the answer + +This is a live decision. Each trigger names evidence, not opinion, and none of them is met today. + +| If this happens | Then re-cost | Scope | +|---|---|---| +| **A real OOM at grid scale** — register **C-71** (a full-pgm `cartesian` target, ~259k cells × months, fed to `reindex_fill` on a sampled frame) or **C-73** (an OOM *despite* per-month sharding) | **zarr or dask, inside `io/`** | Additive. The frame type does not change; the storage layer gains a lazy path. `io/` is already the designated place for evolving formats (ADR-002's internal layering). | +| **A non-Python consumer** appears on the contract | **format-as-contract** (option 2) | Would supersede this ADR. The executable-conformance argument only holds while every consumer can run Python. | +| **The leaf is required on an autograd path** | **torch** | Would first require amending ADR-017, which puts estimation outside the leaf's charter. Very unlikely by construction. | +| **The wire format becomes the dominant access path** — consumers reading parquet directly more often than they construct frames | **pyarrow as the frame type** | Would supersede this ADR and probably ADR-013 with it. | + +**No trigger is recorded for pandas, polars, protocols-only or awkward-array.** Those lose on +grounds that do not change with evidence: a dependency the DAG root cannot take, a memory model +that fights the sample axis, a typing discipline ADR-028 measured as insufficient, and an +encoding this package bans. + +**The bar is deliberately high**, consistent with `GOVERNANCE.md`'s closing note that a package +needing frequent MAJOR bumps is not abstract enough. Changing the frame's container type would +be a MAJOR affecting every consumer simultaneously — a far larger event than 2.0.0, which only +narrowed what construction accepts. + +--- + +## Consequences + +### Positive + +- The runtime dependency surface is one line: `numpy>=1.26,<3`. A consumer that wants a + `PredictionFrame` installs numpy and nothing else. +- The sample axis cannot be renamed or moved, because it is a position rather than a label. +- Construction validates and fails loud, because there is one constructor to validate in. +- The conformance suite is executable and shared, so "all consumers agree" is checked. + +### Negative — accepted deliberately + +- **Alignment is hand-written.** `searchsorted` over a packed row view, plus the cross-level + remap, is code this package maintains and xarray would have provided. +- **No out-of-core path.** C-71 and C-73 are open for this reason, and both would be closed for + free by dask. +- **No `groupby`, `resample` or interpolation.** Consumers do this themselves or go without. +- **numpy version skew is this package's problem.** The `floor` CI job exists because numpy + 1.26.4 and 2.x differ in generic stubs *and* in ~1-ulp histogram binning (registers C-19 and + C-24). A higher-level container would have absorbed that. +- **Roughly 3,700 lines of source and 9,400 of governance are maintained here** that a + third-party container would have carried. + +These are real costs and they are worth naming rather than minimising. The judgement is that a +contract at the root of a dependency DAG is the one place where owning the code is cheaper than +owning the dependency. + +--- + +## References + +- **ADR-002** — the dependency rules and the CRP/SDP argument this ADR applies to a container + choice; its "numpy-only floor" line is the direct ancestor of this decision. +- **ADR-008** — fail loud at construction, which permissive containers work against. +- **ADR-009** — protocols are the surface consumers type against. +- **ADR-012** — the trailing sample axis, the invariant a label-based container cannot hold. +- **ADR-013** — the wire contract, which is format-as-contract for the serialized form. +- **ADR-016** — executable conformance, the mechanism format-as-contract would give up. +- **ADR-017** — estimation is a sibling package and is not differentiable, which removes torch's + technical pull. +- **ADR-028** — the measured evidence that structural typing alone is insufficient. +- Registers **C-19**/**C-24** (numpy version skew), **C-21** (row uniqueness), **C-22** (row + blocking), **C-40**/**C-66** (the list-in-cell blow-up), **C-71**/**C-73** (the memory-wall + triggers above). +- `pyproject.toml` — `dependencies = ["numpy>=1.26,<3"]`, and the `arrow` optional extra. diff --git a/docs/ADRs/README.md b/docs/ADRs/README.md index adf1633..516dc8d 100644 --- a/docs/ADRs/README.md +++ b/docs/ADRs/README.md @@ -76,6 +76,7 @@ it resolves. - **ADR-024** — [Principled joint probabilistic reconciliation — design direction & deferral](024_principled_joint_reconciliation_design.md). Design-only (no code): per-draw `proportional` is an information-losing approximation that pairs grid-draw `s` with country-draw `s` across **independently-trained** models with no shared draw identity. Names the **draw-identity contract** principled reconciliation requires (shared ensemble / copula coupling), fixes the upgrade as a **future sibling module** (probabilistic reconciliation / MinT — never a change to `proportional`), and **defers** implementation until a consumer needs calibrated joint tails *and* the country model can supply the identity/coupling. Records C-62; corrects the ambiguous `proportional.py` "C-37" reference (views-frames C-37 is unrelated/resolved; the lineage is views-postprocessing C-37). - **ADR-025** — [Value-buffer immutability is by convention; only the index is enforced](025_value_buffer_immutability_by_convention.md). Corrects the "immutable value objects" contract to match the code: only the index (`time`/`unit`) is `setflags(write=False)`-enforced; the frame **value buffer is immutable by convention** (mutating `.values` in place is unsupported and may corrupt buffer-sharing frames), left writeable to preserve zero-copy / `mmap` (C-07). The `setflags`-enforce on `.values` would be "tightening an invariant" on a frozen-surface member (a MAJOR, GOVERNANCE/ADR-018) for a hole nothing exercises, so it was recorded as a **deferred MAJOR-rider**. **Amended 2026-08-18 (ADR-028): the enforce shipped** in 2.0.0, riding that MAJOR exactly as planned — though as a read-only *view*, not the `setflags` one-liner this ADR recorded, which would have made the caller's own array read-only via the C-07 zero-copy return. Resolves C-63 and C-66. - **ADR-026** — [Dense-grid fill is a leaf primitive](026_dense_grid_fill_primitive.md). Adds `frame.reindex_fill(other, *, fill_value)` (align with **no** superset requirement; present rows bit-exact, absent rows filled; `fill_value` keyword-only + required) on all three siblings, `SpatioTemporalIndex.cartesian(times, units, level)` (dense product index, time-major, explicit arrays only, fails loud on duplicated inputs), and the published `assert_reindex_fill_law`. Closes #203 (faoapi's "views-frames has no fill primitive", the last pandas blocker on FAO ingestion, faoapi #242); rejects a derivation rule and a `reindex` mode-kwarg (D-09/C-52 precedents). Additive MINOR (v1.10.0); `CONFORMANCE_FLOOR` stays 1.0.0. +- **ADR-029** — [The frames are numpy arrays plus explicit identifier arrays — not xarray, not a DataFrame, not a tensor](029_why_numpy_and_not_xarray.md). Records the container choice that was never written down: the word "xarray" appeared nowhere in this repository's ~9,400 lines of governance prose until this ADR, so the first question every newcomer asks had no answer. Judges eight alternatives — **xarray**, format-as-contract, **torch.Tensor**, pyarrow `Table`, protocols-only, pandas, polars, awkward-array — against one constraint: every dependency the DAG root takes, every consumer takes, and one consumer is an HTTP API server. xarray loses only on its hard `pandas>=2.2` requirement; torch loses on an unconditional CUDA stack *and* on reversing a migration this platform already completed; protocols-only is closed by ADR-028's measurement that `runtime_checkable` checks member presence only. A **live** decision: names the evidence that would reopen it (a C-71/C-73 OOM receipt → cost zarr/dask *inside* `io/`; a non-Python consumer → cost format-as-contract). Nothing changes; `CONFORMANCE_FLOOR` stays 2.0.0. - **ADR-028** — [2.0.0 — a frame's index must actually be an index](028_major_2_0_0_index_type_enforcement.md). **The first MAJOR, and the first move of `CONFORMANCE_FLOOR`** (`1.0.0` → `2.0.0`). Until 2.0.0 the constructors read one attribute off `index` — `n_rows` — so a frame, or any object with an `n_rows`, was accepted silently, and the published `assert_summarizer_contract` then **certified** the result: a checker issuing a false pass, the failure mode `docs/CICs/Conformance.md` says it must never have. Found by a falsification audit of the claim that this package was finished. Carries its riders per C-13: **C-66**'s value-buffer write-protection (as a read-only view) and **C-57**'s `map_estimate` non-finite guard; **declines C-43** with its reason recorded. Also fixes the alignment ops' `AttributeError` leak of a private attribute. A correct consumer changes its version constraint and re-locks; nothing else. Resolves C-66, C-57 and the audit's findings; amends ADR-018 and ADR-025. - **ADR-027** — [Frame construction stays two-step; the convenience factory is declined](027_decline_construction_convenience_factory.md). **Declines #113**, which asked for a one-line shortcut for building a `PredictionFrame`. Nothing changes: callers keep building the `SpatioTemporalIndex` and then the frame. The request was never built and never blocked anyone — the engine repos migrated using the two-step form — and because the surface is frozen (ADR-018), adding a shortcut would be permanent. Stating which rows your numbers describe is what this package should make callers say out loud (ADR-003), not hide. The ADR records the design already agreed for any future need, and what would justify revisiting it. Resolves C-52/C-53/C-54; docs-only, `CONFORMANCE_FLOOR` stays 1.0.0. From 0c39c344a819e89e64dda50fe75307ae23ad5def Mon Sep 17 00:00:00 2001 From: Polichinl Date: Mon, 24 Aug 2026 13:33:20 +0200 Subject: [PATCH 2/6] docs(adr): ADR-029 gains the numpy-API-compatible backends, and drops a wrong claim MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review question: "have you also considered CuPy, JAX, Zarr, and/or Dask Array?" No — and the omission was a whole class, not three names. The eight alternatives in the first draft all proposed a DIFFERENT container with a DIFFERENT API. These propose the SAME API with a different backend: dask.array, jax.numpy and cupy are each close enough to numpy that most of this package would port with modest edits. That is a different question and it deserved asking. **One claim in the first draft was simply wrong, and it is struck rather than quietly edited.** The draft said declining xarray costs out-of-core evaluation "via dask". Out-of-core is not xarray's to give: dask 2026.7.1's hard dependencies are click, cloudpickle, fsspec, packaging, partd, pyyaml, toolz and importlib_metadata — all pure-Python, no pandas. So declining xarray costs less than stated, and the real out-of-core decision is a separate and much cheaper one the draft never posed. The revisit trigger now names `dask.array` inside io/ rather than xarray, for exactly that reason. **A filter emerged that no amount of reasoning would have produced.** Three of the four candidates force numpy 2: zarr 3.3.0 numpy>=2 jax 0.11.1 numpy>=2.1, plus scipy cupy-cuda12x numpy<2.6,>=2.0 dask 2026.7.1 none in core <- the only one compatible This package declares numpy>=1.26,<3 and runs a `floor` job at 1.26.4 because consumers pin conservatively (C-19, C-24). Adopting any of the first three moves the platform's floor as a side effect of a container choice. **JAX gets an honest paragraph beyond the floor objection.** Its arrays are immutable natively, so the entire C-66 saga — deferred 2026-06-28, carried unshipped through eleven releases, landed in 2.0.0, and only then after finding that ADR-025's recorded one-liner would have locked the caller's array — would not have existed. That is a real loss and glossing it would have made the ADR less useful, not more persuasive. Also added zarr as a storage-layer option distinct from format-as-contract, since the first draft folded the two together and they are different decisions. 289 -> 380 lines. No code change; CONFORMANCE_FLOOR stays 2.0.0. Co-Authored-By: Claude Opus 5 (1M context) --- docs/ADRs/029_why_numpy_and_not_xarray.md | 95 +++++++++++++++++++++-- docs/ADRs/README.md | 2 +- 2 files changed, 89 insertions(+), 8 deletions(-) diff --git a/docs/ADRs/029_why_numpy_and_not_xarray.md b/docs/ADRs/029_why_numpy_and_not_xarray.md index d76d31d..0bcdb28 100644 --- a/docs/ADRs/029_why_numpy_and_not_xarray.md +++ b/docs/ADRs/029_why_numpy_and_not_xarray.md @@ -3,7 +3,7 @@ **Status:** Accepted (a live decision — see *What would change the answer*) **Date:** 2026-08-24 **Deciders:** VIEWS platform maintainers -**Consulted:** ADR-002's dependency rules; the 2026-06 leaf and summarize postmortems; ADR-028's measured evidence on structural typing +**Consulted:** ADR-002's dependency rules; the 2026-06 leaf and summarize postmortems; ADR-028's measured evidence on structural typing; live PyPI metadata for every candidate named **Informed:** every repository that pins `views-frames` --- @@ -16,7 +16,7 @@ Nothing else. The only runtime dependency this package has is `numpy>=1.26,<3`. **The obvious question is why this is not xarray**, which does labeled dimensions and coordinate alignment properly, is maintained by people who think about it full time, and would have replaced most of `SpatioTemporalIndex` with `.sel()`. This ADR answers that, and -the seven other alternatives a reasonable person would raise. +the nine other alternatives a reasonable person would raise. It was never written down before. Roughly 9,400 lines of governance prose in this repository and the word "xarray" appeared nowhere in it until this document — so a newcomer asking the @@ -100,8 +100,10 @@ package does anyway, minus the dependency. **What is genuinely lost by declining it** — stated plainly, because this is the alternative that costs the most to refuse: -- lazy and out-of-core evaluation via dask, which is the direct answer to register C-71 and - C-73; +- ~~lazy and out-of-core evaluation via dask~~ — **struck. This was wrong in the first draft and + the correction matters.** Out-of-core is not xarray's to give: `dask.array` provides it with a + numpy-compatible API and no pandas (see §9). Declining xarray costs less than this ADR first + claimed; the out-of-core question is a separate, cheaper decision that the draft never posed; - zarr and netCDF, which would make `io/` mostly disappear; - `groupby`/`resample`/interpolation, all currently absent or hand-rolled; - roughly 3,700 lines of source that someone else would maintain. @@ -220,6 +222,84 @@ hard dependency, its own compiled runtime — no pandas, no numpy requirement of genuinely light install. The case against it is the memory model alone, and it is the closest any DataFrame library comes to being viable here. +### 9. The numpy-API-compatible backends — Dask Array, JAX, CuPy + +A class the first draft of this ADR missed entirely, and the omission was worth catching. The +eight alternatives above all propose a *different container with a different API*. These three +propose **the same API with a different backend**: `dask.array`, `jax.numpy` and `cupy` are each +close enough to numpy that most of this package would port with modest edits. That is a +materially different question, and it deserved asking. + +**One filter removes most of the class before any design argument.** This package declares +`numpy>=1.26,<3` and runs a dedicated `floor` CI job at numpy 1.26.4, because consumers pin +conservatively and the floor is the boundary they actually pin to (registers C-19, C-24). +Measured on 2026-08-24: + +| Candidate | numpy requirement | Against the `>=1.26` floor | +|---|---|---| +| `zarr` 3.3.0 | `numpy>=2` | forces every consumer to numpy 2 | +| `jax` 0.11.1 | `numpy>=2.1`, plus `scipy` | forces every consumer to numpy 2 | +| `cupy-cuda12x` 14.2.0 | `numpy<2.6,>=2.0` | forces every consumer to numpy 2 | +| `dask` 2026.7.1 | none in core | **compatible** | + +Adopting any of the first three would move the floor for the whole platform as a side effect of +a container choice. That is not automatically disqualifying — the floor can be raised +deliberately — but it converts "swap the array backend" into a coordinated cross-repo bump, +which is the cost this package exists to avoid. + +**Dask Array** is the one that survives the filter, and it is the strongest missed alternative. +`dask` 2026.7.1's hard dependencies are `click`, `cloudpickle`, `fsspec`, `packaging`, `partd`, +`pyyaml`, `toolz` and `importlib_metadata` — all pure-Python, and **no pandas**. It answers C-71 +(the dense-grid allocation) and C-73 (read-all-to-RAM) directly, which is exactly what this ADR +first mis-attributed to xarray. + +It loses **as the frame type** for a reason that is about this package's philosophy rather than +about dask: **laziness is incompatible with fail-loud-at-construction.** ADR-008 requires that a +frame is never returned half-valid — `validate_values` runs in `__init__`, before the object is +usable. On a lazy array that check either triggers computation, defeating the laziness that was +the whole point, or defers, defeating ADR-008. `frame.values` would stop being a numpy array, +breaking every consumer's expectation and the `assert_frame_envelope` contract with it. And +chunk boundaries would interact with the row-blocking discipline the tower estimators use to +bound peak memory (C-22) in ways that need re-derivation rather than porting. + +**Dask belongs in `io/`, not in the frame** — and that is now the sharpest form of this ADR's +main revisit trigger, sharper than the xarray version the first draft recorded, precisely +because it costs no pandas. + +**JAX** deserves one honest paragraph beyond the floor objection, because it would have given +this package something it wanted and had to build. **JAX arrays are immutable natively.** The +entire C-66 saga — write-protection recorded as a deferred MAJOR-rider on 2026-06-28, carried +unshipped through eleven releases, finally landed in 2.0.0, and then only after discovering that +the one-line `setflags` fix ADR-025 recorded would have made the *caller's* array read-only — +would not have existed. Immutability would have been a property of the container rather than an +invariant this package enforces by hand. + +That is a real loss and it is worth naming. It does not outweigh forcing numpy 2 and scipy on +every consumer for a package that needs neither autograd nor XLA (ADR-017 puts estimation +outside the leaf's charter, and none of it is differentiable), but the trade should be recorded +rather than glossed. + +**CuPy** loses on the same grounds as torch, with one addition. The DAG-root argument applies — +a GPU array library at the root means views-faoapi needs a CUDA runtime to serialize a forecast +— though cupy's *declared* footprint is far lighter than torch's. The addition is capacity: GPU +memory is scarcer than host RAM, and C-71's motivating case is a full-pgm dense grid (~259k +cells × months × samples) that does not fit comfortably in host memory, let alone device memory. +The leaf's work is alignment and sample-axis reduction, not the dense linear algebra a GPU +exists for. + +### 10. zarr as the storage layer (not as the contract) + +Distinct from option 2, where the *format is* the contract. Here the frames stay numpy and +`io/` gains a chunked, compressed, partially-readable format alongside `npz` and `arrow`. + +This is a complement rather than a competitor, and it is the natural companion to `dask.array` +above — chunked storage feeding a chunked reader. It is not adopted today only because C-73's +trigger has not fired: nothing has yet reported an OOM despite per-month sharding, and `io/` +already has two formats to maintain. + +The floor objection applies to zarr 3 (`numpy>=2`), so if this arrives it either waits for the +platform's numpy floor to move on its own schedule, or pins `zarr<3`. + ### 8. awkward-array The ragged-data library. Dismissed in a sentence: it would make the variable-length per-cell @@ -235,7 +315,7 @@ This is a live decision. Each trigger names evidence, not opinion, and none of t | If this happens | Then re-cost | Scope | |---|---|---| -| **A real OOM at grid scale** — register **C-71** (a full-pgm `cartesian` target, ~259k cells × months, fed to `reindex_fill` on a sampled frame) or **C-73** (an OOM *despite* per-month sharding) | **zarr or dask, inside `io/`** | Additive. The frame type does not change; the storage layer gains a lazy path. `io/` is already the designated place for evolving formats (ADR-002's internal layering). | +| **A real OOM at grid scale** — register **C-71** (a full-pgm `cartesian` target, ~259k cells × months, fed to `reindex_fill` on a sampled frame) or **C-73** (an OOM *despite* per-month sharding) | **`dask.array` inside `io/`**, with `zarr` as its storage format | Additive. The frame type does not change; the storage layer gains a lazy path. Prefer dask over xarray here: it answers the same need and requires no pandas (§9). `io/` is the designated place for evolving formats (ADR-002). | | **A non-Python consumer** appears on the contract | **format-as-contract** (option 2) | Would supersede this ADR. The executable-conformance argument only holds while every consumer can run Python. | | **The leaf is required on an autograd path** | **torch** | Would first require amending ADR-017, which puts estimation outside the leaf's charter. Very unlikely by construction. | | **The wire format becomes the dominant access path** — consumers reading parquet directly more often than they construct frames | **pyarrow as the frame type** | Would supersede this ADR and probably ADR-013 with it. | @@ -266,8 +346,9 @@ narrowed what construction accepts. - **Alignment is hand-written.** `searchsorted` over a packed row view, plus the cross-level remap, is code this package maintains and xarray would have provided. -- **No out-of-core path.** C-71 and C-73 are open for this reason, and both would be closed for - free by dask. +- **No out-of-core path.** C-71 and C-73 are open for this reason. Both would be closed by + `dask.array` — which, unlike xarray, costs no pandas. This is the cheapest capability this + ADR declines, and the first draft mislabelled it as xarray's. - **No `groupby`, `resample` or interpolation.** Consumers do this themselves or go without. - **numpy version skew is this package's problem.** The `floor` CI job exists because numpy 1.26.4 and 2.x differ in generic stubs *and* in ~1-ulp histogram binning (registers C-19 and diff --git a/docs/ADRs/README.md b/docs/ADRs/README.md index 516dc8d..e6207ed 100644 --- a/docs/ADRs/README.md +++ b/docs/ADRs/README.md @@ -76,7 +76,7 @@ it resolves. - **ADR-024** — [Principled joint probabilistic reconciliation — design direction & deferral](024_principled_joint_reconciliation_design.md). Design-only (no code): per-draw `proportional` is an information-losing approximation that pairs grid-draw `s` with country-draw `s` across **independently-trained** models with no shared draw identity. Names the **draw-identity contract** principled reconciliation requires (shared ensemble / copula coupling), fixes the upgrade as a **future sibling module** (probabilistic reconciliation / MinT — never a change to `proportional`), and **defers** implementation until a consumer needs calibrated joint tails *and* the country model can supply the identity/coupling. Records C-62; corrects the ambiguous `proportional.py` "C-37" reference (views-frames C-37 is unrelated/resolved; the lineage is views-postprocessing C-37). - **ADR-025** — [Value-buffer immutability is by convention; only the index is enforced](025_value_buffer_immutability_by_convention.md). Corrects the "immutable value objects" contract to match the code: only the index (`time`/`unit`) is `setflags(write=False)`-enforced; the frame **value buffer is immutable by convention** (mutating `.values` in place is unsupported and may corrupt buffer-sharing frames), left writeable to preserve zero-copy / `mmap` (C-07). The `setflags`-enforce on `.values` would be "tightening an invariant" on a frozen-surface member (a MAJOR, GOVERNANCE/ADR-018) for a hole nothing exercises, so it was recorded as a **deferred MAJOR-rider**. **Amended 2026-08-18 (ADR-028): the enforce shipped** in 2.0.0, riding that MAJOR exactly as planned — though as a read-only *view*, not the `setflags` one-liner this ADR recorded, which would have made the caller's own array read-only via the C-07 zero-copy return. Resolves C-63 and C-66. - **ADR-026** — [Dense-grid fill is a leaf primitive](026_dense_grid_fill_primitive.md). Adds `frame.reindex_fill(other, *, fill_value)` (align with **no** superset requirement; present rows bit-exact, absent rows filled; `fill_value` keyword-only + required) on all three siblings, `SpatioTemporalIndex.cartesian(times, units, level)` (dense product index, time-major, explicit arrays only, fails loud on duplicated inputs), and the published `assert_reindex_fill_law`. Closes #203 (faoapi's "views-frames has no fill primitive", the last pandas blocker on FAO ingestion, faoapi #242); rejects a derivation rule and a `reindex` mode-kwarg (D-09/C-52 precedents). Additive MINOR (v1.10.0); `CONFORMANCE_FLOOR` stays 1.0.0. -- **ADR-029** — [The frames are numpy arrays plus explicit identifier arrays — not xarray, not a DataFrame, not a tensor](029_why_numpy_and_not_xarray.md). Records the container choice that was never written down: the word "xarray" appeared nowhere in this repository's ~9,400 lines of governance prose until this ADR, so the first question every newcomer asks had no answer. Judges eight alternatives — **xarray**, format-as-contract, **torch.Tensor**, pyarrow `Table`, protocols-only, pandas, polars, awkward-array — against one constraint: every dependency the DAG root takes, every consumer takes, and one consumer is an HTTP API server. xarray loses only on its hard `pandas>=2.2` requirement; torch loses on an unconditional CUDA stack *and* on reversing a migration this platform already completed; protocols-only is closed by ADR-028's measurement that `runtime_checkable` checks member presence only. A **live** decision: names the evidence that would reopen it (a C-71/C-73 OOM receipt → cost zarr/dask *inside* `io/`; a non-Python consumer → cost format-as-contract). Nothing changes; `CONFORMANCE_FLOOR` stays 2.0.0. +- **ADR-029** — [The frames are numpy arrays plus explicit identifier arrays — not xarray, not a DataFrame, not a tensor](029_why_numpy_and_not_xarray.md). Records the container choice that was never written down: the word "xarray" appeared nowhere in this repository's ~9,400 lines of governance prose until this ADR, so the first question every newcomer asks had no answer. Judges ten alternatives — **xarray**, format-as-contract, **torch.Tensor**, pyarrow `Table`, protocols-only, pandas, polars, awkward-array, the numpy-API-compatible backends (**Dask Array**, **JAX**, **CuPy**) and zarr-as-storage — against one constraint: every dependency the DAG root takes, every consumer takes, and one consumer is an HTTP API server. xarray loses only on its hard `pandas>=2.2` requirement; torch loses on an unconditional CUDA stack *and* on reversing a migration this platform already completed; protocols-only is closed by ADR-028's measurement that `runtime_checkable` checks member presence only. Three candidates (zarr 3, JAX, CuPy) are filtered out by a single measured fact — each requires numpy 2, which would move the platform's declared floor as a side effect of a container choice. A **live** decision: names the evidence that would reopen it (a C-71/C-73 OOM receipt → cost **`dask.array`** *inside* `io/`, which unlike xarray costs no pandas; a non-Python consumer → cost format-as-contract). Nothing changes; `CONFORMANCE_FLOOR` stays 2.0.0. - **ADR-028** — [2.0.0 — a frame's index must actually be an index](028_major_2_0_0_index_type_enforcement.md). **The first MAJOR, and the first move of `CONFORMANCE_FLOOR`** (`1.0.0` → `2.0.0`). Until 2.0.0 the constructors read one attribute off `index` — `n_rows` — so a frame, or any object with an `n_rows`, was accepted silently, and the published `assert_summarizer_contract` then **certified** the result: a checker issuing a false pass, the failure mode `docs/CICs/Conformance.md` says it must never have. Found by a falsification audit of the claim that this package was finished. Carries its riders per C-13: **C-66**'s value-buffer write-protection (as a read-only view) and **C-57**'s `map_estimate` non-finite guard; **declines C-43** with its reason recorded. Also fixes the alignment ops' `AttributeError` leak of a private attribute. A correct consumer changes its version constraint and re-locks; nothing else. Resolves C-66, C-57 and the audit's findings; amends ADR-018 and ADR-025. - **ADR-027** — [Frame construction stays two-step; the convenience factory is declined](027_decline_construction_convenience_factory.md). **Declines #113**, which asked for a one-line shortcut for building a `PredictionFrame`. Nothing changes: callers keep building the `SpatioTemporalIndex` and then the frame. The request was never built and never blocked anyone — the engine repos migrated using the two-step form — and because the surface is frozen (ADR-018), adding a shortcut would be permanent. Stating which rows your numbers describe is what this package should make callers say out loud (ADR-003), not hide. The ADR records the design already agreed for any future need, and what would justify revisiting it. Resolves C-52/C-53/C-54; docs-only, `CONFORMANCE_FLOOR` stays 1.0.0. From 3f6483e4ec9ed8951a91a0035513deb75298ff17 Mon Sep 17 00:00:00 2001 From: Polichinl Date: Mon, 24 Aug 2026 13:51:55 +0200 Subject: [PATCH 3/6] =?UTF-8?q?docs(adr):=20ADR-029=20leads=20with=20the?= =?UTF-8?q?=20real=20reason=20=E2=80=94=20numpy=20was=20a=20free=20depende?= =?UTF-8?q?ncy?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Maintainer review supplied the rationale the draft only half-had, and asked for the alternatives' footprints to be measured rather than guessed. Both are now in. **The primary argument was buried.** numpy was chosen because every repo on this platform already depended on it, so adding the contract added no transitive dependency at all. That is what the repo separation exists to protect: the model repos are deliberately allowed wildly different libraries and need not share an environment, so anything imported by *most* repos must be nearly free — it is the one place their dependency sets are forced to agree. That now leads the document. **Weight and coupling are separated, because they are different costs.** Weight is what a repo ships. Coupling is whether the dependency has a version some repo already pins for its own reasons — small packages can be expensive this way. pandas is the clearest case; torch is the rare one where both apply. **Footprints measured**, resolving each candidate's full transitive closure with `pip install --dry-run --report` on 3.11 and summing manylinux wheel sizes: numpy 1 pkg 18.6 MB (baseline, already everywhere) dask[array] 12 pkgs 21.6 MB +3.0 MB zarr 5 pkgs 27.7 MB +9.1 xarray 8 pkgs 35.1 MB +16.5 <- drags pandas pyarrow 1 pkg 53.1 MB +34.5 polars 2 pkgs 59.4 MB +40.8 jax 6 pkgs 151.8 MB +133.2 <- drags scipy cupy-cuda12x 3 pkgs 166.9 MB +148.3 torch 29 pkgs 3022.8 MB +3004 (162x numpy) Two corrections follow from the measurement. **xarray is not heavy** — 35 MB — so the draft's weight argument was the weak one; the real cost is the pandas coupling. **dask[array] is essentially free** at three megabytes over numpy, cheaper than zarr and half of xarray, with nothing in its closure another repo is likely to pin. Any framing that treats adopting dask as expensive is wrong on the evidence, and the ADR now says its objections are about design, not cost — a distinction that matters, because a cheap dependency deferred on design can be revisited by changing the design. **New section: the scaling wall, with numbers.** 10 features x 128 samples x 259,200 cells x 500 months = 165.888e9 cells = 663.6 GB float32. The maintainer's water adjustment inverted — 80% water means 20% *remains*, so land-only is 33.2e9 cells (132.7 GB), not 132.7e9. Conclusion unchanged; both are far past RAM. The per-shard figure is the useful one: one month at 10 features x 128 samples is 1.33 GB and fits, at 50 x 512 it is 26.5 GB and does not. **The wall arrives through feature and sample count, not months** — months add shards, features multiply every shard. On memmap: sufficient longer than the totals suggest, because `(N, S)` row-major with contiguous samples is exactly what a per-row reduction wants and what C-22's row blocking already does. It fails on densification (C-71 writes a dense buffer, memmap helps readers), on compression, and on chunked parallelism. Expected order of failure: storage volume first, densification second, reduction last — so the first fix needed is a chunked compressed format in `io/`, not a new frame type. 380 -> 529 lines. No code change; CONFORMANCE_FLOOR stays 2.0.0. Co-Authored-By: Claude Opus 5 (1M context) --- docs/ADRs/029_why_numpy_and_not_xarray.md | 196 ++++++++++++++++++++-- 1 file changed, 178 insertions(+), 18 deletions(-) diff --git a/docs/ADRs/029_why_numpy_and_not_xarray.md b/docs/ADRs/029_why_numpy_and_not_xarray.md index 0bcdb28..6c84942 100644 --- a/docs/ADRs/029_why_numpy_and_not_xarray.md +++ b/docs/ADRs/029_why_numpy_and_not_xarray.md @@ -13,6 +13,10 @@ A frame is a numpy array with a small object holding two integer identifier arrays beside it. Nothing else. The only runtime dependency this package has is `numpy>=1.26,<3`. +**numpy was chosen because it was free.** Every repository on this platform already depended on +it, so adding the contract added no transitive dependency at all. That is the whole argument, and +everything below is the working that supports it. + **The obvious question is why this is not xarray**, which does labeled dimensions and coordinate alignment properly, is maintained by people who think about it full time, and would have replaced most of `SpatioTemporalIndex` with `.sel()`. This ADR answers that, and @@ -23,8 +27,15 @@ and the word "xarray" appeared nowhere in it until this document — so a newcom first question anyone asks found no answer, which is how a settled decision gets relitigated by accident. -**Nothing changes as a result of this ADR.** It records a decision already made and names the -evidence that would reopen it. +**Nothing changes as a result of this ADR.** It records a decision already made, measures what +each alternative would have cost, and names the evidence that would reopen it. + +**It has a visible expiry.** *The scaling wall, with numbers* below sizes the data this platform +expects to hold — hundreds of gigabytes for a single global feature set — and concludes that +per-month sharding plus `numpy.memmap` carries the current access pattern further than the raw +totals suggest, but not past densification or storage volume. This decision is live, not closed, +and `dask[array]` is deferred on design grounds rather than on cost: it measures at three +megabytes over numpy. --- @@ -38,6 +49,36 @@ ADR-002's decision, and it has a consequence most container choices ignore: > **Every dependency this package takes, every consumer takes.** +### numpy was chosen because it was free + +This is the primary reason, and it is worth stating before any comparison of features. + +**Every repository on this platform already depends on numpy.** Adding `views-frames` to a repo +therefore adds no new transitive dependency at all — the array library was already installed, +already resolved, already pinned. The contract arrived at zero dependency cost. + +That property is what the whole repository separation exists to protect. The model repos are +deliberately allowed to hold *wildly* different libraries — that is the point of separating +them, and it is why they need not share an environment; orchestration calls across repo +boundaries rather than importing across them. The price of that freedom is that anything +imported by *most* repos must be nearly free, because it is the one place where their dependency +sets are forced to agree. + +**So the question for every alternative is not "is it a better array container".** Several are. +The question is what it costs to make every repository on the platform carry it — in bytes, and +more importantly in *version coupling*. + +### The two costs are different, and the second is the one that bites + +- **Weight** is what a repo downloads and ships. It matters when it is large in absolute terms. +- **Coupling** is whether the dependency has a version that some repo *already pins for its own + reasons*. A package that many repos already use at their own chosen version is expensive to + put at the DAG root even when it is small, because from then on every repo's pin must satisfy + the leaf's floor too. + +pandas is the clearest case of the second: it is small enough, and it is exactly the kind of +library a model repo already pins. torch is the rare case where both costs apply at once. + `ADR-002` states the rule directly: > The numpy-only floor (CRP) means a model that wants a `PredictionFrame` does not @@ -73,6 +114,38 @@ the frame type. That boundary is enforced by `tests/test_import_enforcement.py` --- +## What each alternative actually costs + +Measured on 2026-08-24 by resolving each candidate's full transitive closure with +`pip install --dry-run --report` on Python 3.11 and summing the manylinux wheel sizes from PyPI. +Download size, not installed size; the ranking is what matters, not the absolute figures. + +| Candidate | Packages | Download | Over numpy | Drags in a commonly-pinned library? | +|---|---:|---:|---:|---| +| **numpy** — the baseline, already present in every repo | 1 | 18.6 MB | — | — | +| `dask[array]` | 12 | 21.6 MB | **+3.0 MB** | no | +| `awkward` | 8 | 21.8 MB | +3.2 MB | no | +| `zarr` | 5 | 27.7 MB | +9.1 MB | no | +| `xarray` | 8 | 35.1 MB | +16.5 MB | **yes — pandas** | +| `pyarrow` | 1 | 53.1 MB | +34.5 MB | no | +| `polars` | 2 | 59.4 MB | +40.8 MB | no | +| `jax` | 6 | 151.8 MB | +133.2 MB | **yes — scipy** | +| `cupy-cuda12x` | 3 | 166.9 MB | +148.3 MB | a CUDA runtime | +| **`torch`** | **29** | **3,022.8 MB** | **+3,004 MB (162×)** | **CUDA stack, unconditional on Linux** | + +Three things this table settles that argument alone did not: + +- **`dask[array]` is essentially free** — twelve small pure-Python packages, three megabytes over + numpy, cheaper than zarr and half the weight of xarray. Any framing that treats adopting dask + as an expensive step is wrong on the evidence. Its objections are about design, not cost. +- **xarray is not heavy.** At 35 MB the weight argument against it is weak. The real cost is the + pandas coupling in the last column. +- **torch is in a different category entirely** — 162× numpy, and the only entry whose transitive + closure is dominated by things unrelated to arrays (`nvidia-cublas` 543 MB, `torch` 527 MB, + `nvidia-cudnn` 445 MB). + +--- + ## Alternatives considered ### 1. xarray — the real contender @@ -84,11 +157,20 @@ be consumer-injected, because that is a domain rule rather than a container feat semantics, `groupby`, `resample`, and unit-aware slicing all arrive free, and several of them are things this package deliberately does not offer. -**It loses on the dependency, and only on the dependency.** xarray declares `pandas>=2.2` as a -hard requirement (checked against xarray 2026.7.0 on 2026-08-24, not assumed). Adopting it -puts pandas back into every consumer of the contract — the precise outcome `ADR-002` names as -the reason for the numpy floor, and the thing views-faoapi #242 is still working to undo on the -consumer side. A leaf that re-imports what its consumers are removing is not a leaf. +**It loses on the dependency, and only on the dependency — but not for the reason weight would +suggest.** At 35 MB total, xarray is not a heavy install; an earlier draft of this ADR leaned on +size and that was the weaker argument. + +The real cost is **coupling**. xarray declares `pandas>=2.2` as a hard requirement (checked +against xarray 2026.7.0 on 2026-08-24). pandas is precisely the kind of library a model repo +*already pins for its own reasons* — so putting a pandas floor at the DAG root means every +repository's pandas pin must from then on also satisfy the leaf's. That is the dependency +alignment work the repo separation exists to avoid, reintroduced at the one point in the graph +that touches everything. + +It is also the outcome `ADR-002` names as the reason for the numpy floor, and the thing +views-faoapi #242 is still working to undo on the consumer side. A leaf that re-imports what its +consumers are removing is not a leaf. **The secondary loss is that xarray's flexibility is this package's liability.** In xarray a dimension is a string. Nothing prevents `sample` becoming `draw`, nothing prevents it moving @@ -222,6 +304,15 @@ hard dependency, its own compiled runtime — no pandas, no numpy requirement of genuinely light install. The case against it is the memory model alone, and it is the closest any DataFrame library comes to being viable here. +### 8. awkward-array + +The ragged-data library. Dismissed in a sentence: it would make the variable-length per-cell +encoding *efficient*, and this package bans that encoding on purpose. The data is rectangular by +construction — a dense grid of cells by a fixed sample count — and a container optimised for +raggedness would legitimize the shape C-40 exists to prevent. + +--- + ### 9. The numpy-API-compatible backends — Dask Array, JAX, CuPy A class the first draft of this ADR missed entirely, and the omission was worth catching. The @@ -247,11 +338,18 @@ a container choice. That is not automatically disqualifying — the floor can be deliberately — but it converts "swap the array backend" into a coordinated cross-repo bump, which is the cost this package exists to avoid. -**Dask Array** is the one that survives the filter, and it is the strongest missed alternative. -`dask` 2026.7.1's hard dependencies are `click`, `cloudpickle`, `fsspec`, `packaging`, `partd`, -`pyyaml`, `toolz` and `importlib_metadata` — all pure-Python, and **no pandas**. It answers C-71 -(the dense-grid allocation) and C-73 (read-all-to-RAM) directly, which is exactly what this ADR -first mis-attributed to xarray. +**Dask Array** is the one that survives the filter, and it is the strongest missed alternative +by a distance. `dask` 2026.7.1's hard dependencies are `click`, `cloudpickle`, `fsspec`, +`packaging`, `partd`, `pyyaml`, `toolz` and `importlib_metadata` — all pure-Python, and **no +pandas**. Measured: `dask[array]` resolves to twelve packages totalling **21.6 MB, three +megabytes over numpy** — cheaper than zarr, half the weight of xarray, and nothing in the +closure is a library another repo is likely to be pinning. + +**On the platform's own cost test, dask is nearly as free as numpy was.** It answers C-71 (the +dense-grid allocation) and C-73 (read-all-to-RAM) directly, which is exactly what this ADR first +mis-attributed to xarray. Every objection to it below is about **design**, not cost — and that +distinction should be kept sharp, because a cheap dependency deferred on design grounds can be +revisited by changing the design, whereas an expensive one cannot be revisited at all. It loses **as the frame type** for a reason that is about this package's philosophy rather than about dask: **laziness is incompatible with fail-loud-at-construction.** ADR-008 requires that a @@ -300,14 +398,70 @@ already has two formats to maintain. The floor objection applies to zarr 3 (`numpy>=2`), so if this arrives it either waits for the platform's numpy floor to move on its own schedule, or pins `zarr<3`. -### 8. awkward-array +--- -The ragged-data library. Dismissed in a sentence: it would make the variable-length per-cell -encoding *efficient*, and this package bans that encoding on purpose. The data is rectangular by -construction — a dense grid of cells by a fixed sample count — and a container optimised for -raggedness would legitimize the shape C-40 exists to prevent. +## The scaling wall, with numbers ---- +The decision above is correct for the data this platform holds today. It is worth writing down +what the data is expected to become, because the honest reason this ADR is *live* rather than +closed is that the current answer has a visible expiry. + +Take a plausible near-future shape — ten features, 128 samples each, a global 0.5 degree grid +(720 x 360 = 259,200 cells), five hundred months: + +``` +10 features x 128 samples x 259,200 cells x 500 months = 165,888,000,000 cells + x 4 bytes (float32) = 663.6 GB +``` + +Excluding ocean cells helps but does not rescue it. If roughly 80% of cells are water and are +dropped, **20% remains**: + +| | cells | float32 | +|---|---:|---:| +| Full 0.5 degree grid | 165.9 billion | **663.6 GB** | +| Land only (20% of cells retained) | 33.2 billion | **132.7 GB** | + +Neither fits in memory, and ten features at 128 samples is nowhere near a ceiling. + +**But the per-shard figure is the one that explains why nothing has broken yet**, and it is the +number to watch: + +| One month, all features and samples | float32 | +|---|---:| +| 10 features x 128 samples, full grid | **1.33 GB** | +| 10 features x 128 samples, land only | 265 MB | +| 50 features x 512 samples, full grid | **26.5 GB** | + +Per-month sharding is load-bearing — it is what keeps a working set in the low gigabytes, and it +is exactly the precondition register **C-73** names ("an OOM *despite* per-month sharding"). +**The wall arrives through feature count and sample count, not through months.** Adding months +adds shards; adding features or samples multiplies every shard. + +### Is `numpy.memmap` enough? + +Partly, and for longer than the raw totals suggest — but not for the thing likely to break first. + +**Where it works.** A frame is `(N, S)` row-major with samples contiguous, so a per-row +sample-axis reduction reads sequentially. That is the dominant access pattern in +`views_frames_summarize`, and the tower estimators already work in row blocks to bound peak +memory (C-22) — precisely the pattern memmap rewards. Reduction over a frame far larger than RAM +streams acceptably today. + +**Where it does not.** + +1. **Densification writes, it does not read.** `reindex_fill` allocates the full dense buffer up + front (register **C-71**). memmap does not help build a dense grid; it helps read one. +2. **No compression.** 663 GB stored raw, against what a chunked compressed format would hold. + On present trends this is the binding constraint *before* RAM is — a storage-volume problem + rather than a memory problem. +3. **No chunked parallelism.** memmap gives lazy paging, not a scheduler, and not partial reads + of a compressed chunk. + +So the expected order of failure is **storage volume first, densification second, reduction +last** — which matters, because the first fix needed is then a chunked compressed format in +`io/` (zarr), not a new frame type. That is additive, behind the existing `io/` boundary, and it +does not touch the public surface. ## What would change the answer @@ -320,6 +474,12 @@ This is a live decision. Each trigger names evidence, not opinion, and none of t | **The leaf is required on an autograd path** | **torch** | Would first require amending ADR-017, which puts estimation outside the leaf's charter. Very unlikely by construction. | | **The wire format becomes the dominant access path** — consumers reading parquet directly more often than they construct frames | **pyarrow as the frame type** | Would supersede this ADR and probably ADR-013 with it. | +**The dask trigger deserves a caveat the others do not.** Every alternative above is deferred on +some mix of cost and design. `dask[array]` is deferred on **design alone** — it costs three +megabytes. That makes it the one entry here whose deferral could be undone by changing this +package's mind rather than by waiting for the world to change, and the scaling section above +suggests the world will not wait long. + **No trigger is recorded for pandas, polars, protocols-only or awkward-array.** Those lose on grounds that do not change with evidence: a dependency the DAG root cannot take, a memory model that fights the sample axis, a typing discipline ADR-028 measured as insufficient, and an From 762a06260e5cabe2c81fea648ff6e8077c0a603b Mon Sep 17 00:00:00 2001 From: Polichinl Date: Mon, 24 Aug 2026 14:06:58 +0200 Subject: [PATCH 4/6] =?UTF-8?q?docs(adr):=20ADR-029=20=E2=80=94=20the=20in?= =?UTF-8?q?dex=20fits,=20the=20values=20do=20not?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The strongest argument in the document was missing from it. **The totals were the wrong number to reason from.** A frame is an index plus a value array and they scale completely differently: full global, 500 months one month shard index (time+unit) 2.07 GB (land 0.41 GB) 4.1 MB values 10f x 128s 663.6 GB 1.33 GB values 50f x 512s 13.3 TB 26.5 GB Values are 320x to 6,400x the index, and the index's size does not depend on features or samples at all — 4.1 MB per month whether the frame carries ten features or a hundred. This decides the streaming question. Alignment is what this package structurally does — searchsorted, reindex, intersect, is_superset_of, cartesian — and every one of them touches only `(time, unit)`. The operations needing random access need it over data that **fits in memory at global scale**. What does not fit is read sequentially in row blocks, which is the access pattern the estimators already use (C-22). So **streaming needs a different storage layer, not a different array library**. A lazy array's `.compute()` returns numpy; chunked stores feed numpy rather than replace it. The frame is the unit of work, not the unit of storage — which is why this decision survives a dataset three orders of magnitude larger than the one it was made for. The consequence is stated rather than left implicit: a single frame can never exceed memory, because `assert_frame_envelope` asserts `isinstance(values, np.ndarray)` and ADR-008 requires validation to complete at construction. "One frame holds the whole global dataset" is off the table permanently, and consumers expecting otherwise should be told. **Also added a scope bound.** At 50 features x 512 samples the grid is 13.3 TB (2.65 TB land-only) — past local-disk streaming and into object storage, catalogues and partitioned network reads, where the binding constraints are shard layout, partition keys and transfer cost. Infrastructure questions, not array-library questions. The library choice stops being the decision that matters somewhere before that threshold, and this ADR says so rather than implying its own scope is unlimited. 540 -> 596 lines. No code change; CONFORMANCE_FLOOR stays 2.0.0. Co-Authored-By: Claude Opus 5 (1M context) --- docs/ADRs/029_why_numpy_and_not_xarray.md | 67 +++++++++++++++++++++-- 1 file changed, 63 insertions(+), 4 deletions(-) diff --git a/docs/ADRs/029_why_numpy_and_not_xarray.md b/docs/ADRs/029_why_numpy_and_not_xarray.md index 6c84942..fdadc7d 100644 --- a/docs/ADRs/029_why_numpy_and_not_xarray.md +++ b/docs/ADRs/029_why_numpy_and_not_xarray.md @@ -30,10 +30,13 @@ by accident. **Nothing changes as a result of this ADR.** It records a decision already made, measures what each alternative would have cost, and names the evidence that would reopen it. -**It has a visible expiry.** *The scaling wall, with numbers* below sizes the data this platform -expects to hold — hundreds of gigabytes for a single global feature set — and concludes that -per-month sharding plus `numpy.memmap` carries the current access pattern further than the raw -totals suggest, but not past densification or storage volume. This decision is live, not closed, +**It has a visible expiry, but not the one the totals suggest.** *The scaling wall, with numbers* +below sizes the data this platform expects to hold — hundreds of gigabytes for a single global +feature set, terabytes at the upper end. The decisive fact there is an **asymmetry**: the index +is 2 GB at global scale and fits, while the values are 320x to 6,400x larger and do not. Alignment +touches only the index. So **streaming needs a different storage layer, not a different array +library** — which is why this decision survives a dataset three orders of magnitude larger than +the one it was made for. This decision is live, not closed, and `dask[array]` is deferred on design grounds rather than on cost: it measures at three megabytes over numpy. @@ -424,6 +427,46 @@ dropped, **20% remains**: Neither fits in memory, and ten features at 128 samples is nowhere near a ceiling. +### Why numpy survives this: the index fits, the values do not + +The totals above are the wrong number to reason from. The right one is an asymmetry that decides +the whole question. + +A frame is an index plus a value array, and **they scale completely differently.** + +| | full global grid, 500 months | one month shard | +|---|---:|---:| +| **Index** — `time` + `unit`, `int64` | **2.07 GB** (land only: 0.41 GB) | **4.1 MB** | +| Values — 10 features x 128 samples | 663.6 GB | 1.33 GB | +| Values — 50 features x 512 samples | 13.3 TB | 26.5 GB | + +The values are **320x to 6,400x** the index — and the index's size does not depend on features or +samples at all. It is 4.1 MB per month whether the frame carries ten features or a hundred. + +**This matters because alignment is what this package structurally does, and alignment only ever +touches the index.** `searchsorted`, `reindex`, `intersect`, `is_superset_of` and `cartesian` all +operate on `(time, unit)` and never read a value. Those are the operations needing random access, +and the data they need it over **fits in memory at global scale**. What does not fit is read +sequentially, in row blocks — the access pattern the estimators already use (C-22). + +**So streaming does not require a different array library. It requires a different storage layer +under the same one.** A lazy array's `.compute()` returns a numpy array; chunked stores feed numpy +rather than replace it. The working shape is: stream a chunk from a chunked store, materialize +numpy, build a frame, summarize, discard. Per-month sharding already is that shape — streaming +only makes the shards lazier and smaller. + +**The frame is the unit of work, not the unit of storage.** That is the sentence this ADR exists +to protect, and it is why the numpy decision survives a dataset three orders of magnitude larger +than the one it was made for. + +**The consequence, stated plainly rather than left implicit:** a single frame can never exceed +memory. `assert_frame_envelope` asserts `isinstance(values, np.ndarray)`, and ADR-008 requires +validation to complete at construction — a lazy value array satisfies neither without either +computing (defeating the laziness) or deferring (defeating the guarantee). "One frame holds the +whole global dataset" is off the table by design, permanently. A consumer expecting otherwise +should be told rather than left to find out. + + **But the per-shard figure is the one that explains why nothing has broken yet**, and it is the number to watch: @@ -463,6 +506,22 @@ last** — which matters, because the first fix needed is then a chunked compres `io/` (zarr), not a new frame type. That is additive, behind the existing `io/` boundary, and it does not touch the public surface. +### Where the library stops being the binding decision + +One more scale marker, because it bounds how much this ADR can usefully decide. + +At 50 features x 512 samples the full grid is **13.3 TB**, or **2.65 TB** land-only. That is past +"stream it from local disk" and into object storage, a catalogue, and partitioned reads over a +network. At that point the binding constraints are shard layout, partition keys, cache locality +and transfer cost — **infrastructure questions, not array-library questions.** + +**The library choice stops being the decision that matters somewhere before that threshold.** This +ADR is scoped to the range where it still is one. If the platform reaches TB-scale working sets, +the right response is not a revision of this document but a new decision about storage +architecture — with this ADR reduced to a note about what the in-memory unit looks like once a +chunk has been fetched. + + ## What would change the answer This is a live decision. Each trigger names evidence, not opinion, and none of them is met today. From 518154f89756cc054f2ad99b7090e5706e8620ed Mon Sep 17 00:00:00 2001 From: Polichinl Date: Mon, 24 Aug 2026 14:38:06 +0200 Subject: [PATCH 5/6] docs(adr): ADR-029 reads as a decision, not as a record of arriving at one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two changes, both about who the document is for. **Deciders is now named**: Simon Polichinel von der Maase. **Every trace of the document's own composition is gone.** The text carried "an earlier draft grouped polars with xarray on weight; that was wrong", a struck bullet explaining what a previous version mis-attributed, "a class the first draft missed entirely", and five more of the same shape. Those are interesting to the two people who watched it being written and noise to every other reader — and every other reader is reading it for the first time. An ADR that narrates its own revisions invites the reader to weigh the revisions. Swept as a class rather than fixed instance by instance, across four patterns: iteration references (draft/earlier/struck/mis-attributed), dialogue traces (first and second person, "review", "asked", "deserved asking"), authorial process voice ("JAX deserves one honest paragraph"), and unconverged hedging. Eleven sites plus the frontmatter. Re-swept after editing; clean. Each claim now stands on its evidence instead of on its history. The polars paragraph says polars is light and couples to nothing, rather than confessing that an earlier version said otherwise. The xarray section says the weight argument is the weak one, rather than reporting that it was once leaned on. The measured facts and their dates stay — provenance is not process. One hedge tightened while sweeping: the pyarrow trigger said adopting it "would supersede this ADR and probably ADR-013". Not probably — if the frame type became an Arrow table, the in-memory type and the wire format would be one thing, so ADR-013 goes definitionally. Said that instead. Verified section by section: 26 sections, each leading with its claim, none containing leakage from any of the four patterns. Co-Authored-By: Claude Opus 5 (1M context) --- docs/ADRs/029_why_numpy_and_not_xarray.md | 75 +++++++++++------------ 1 file changed, 35 insertions(+), 40 deletions(-) diff --git a/docs/ADRs/029_why_numpy_and_not_xarray.md b/docs/ADRs/029_why_numpy_and_not_xarray.md index fdadc7d..5963f4a 100644 --- a/docs/ADRs/029_why_numpy_and_not_xarray.md +++ b/docs/ADRs/029_why_numpy_and_not_xarray.md @@ -2,7 +2,7 @@ **Status:** Accepted (a live decision — see *What would change the answer*) **Date:** 2026-08-24 -**Deciders:** VIEWS platform maintainers +**Deciders:** Simon Polichinel von der Maase **Consulted:** ADR-002's dependency rules; the 2026-06 leaf and summarize postmortems; ADR-028's measured evidence on structural typing; live PyPI metadata for every candidate named **Informed:** every repository that pins `views-frames` @@ -22,10 +22,9 @@ coordinate alignment properly, is maintained by people who think about it full t would have replaced most of `SpatioTemporalIndex` with `.sel()`. This ADR answers that, and the nine other alternatives a reasonable person would raise. -It was never written down before. Roughly 9,400 lines of governance prose in this repository -and the word "xarray" appeared nowhere in it until this document — so a newcomer asking the -first question anyone asks found no answer, which is how a settled decision gets relitigated -by accident. +The decision itself predates this record. It is written down because "why not xarray" is the +first question any consumer asks, and a decision that lives only in the shape of the code is one +that gets relitigated by accident. **Nothing changes as a result of this ADR.** It records a decision already made, measures what each alternative would have cost, and names the evidence that would reopen it. @@ -139,8 +138,8 @@ Download size, not installed size; the ranking is what matters, not the absolute Three things this table settles that argument alone did not: - **`dask[array]` is essentially free** — twelve small pure-Python packages, three megabytes over - numpy, cheaper than zarr and half the weight of xarray. Any framing that treats adopting dask - as an expensive step is wrong on the evidence. Its objections are about design, not cost. + numpy, cheaper than zarr and half the weight of xarray. Adopting dask is not an expensive step; + its objections are about design, not cost. - **xarray is not heavy.** At 35 MB the weight argument against it is weak. The real cost is the pandas coupling in the last column. - **torch is in a different category entirely** — 162× numpy, and the only entry whose transitive @@ -160,9 +159,8 @@ be consumer-injected, because that is a domain rule rather than a container feat semantics, `groupby`, `resample`, and unit-aware slicing all arrive free, and several of them are things this package deliberately does not offer. -**It loses on the dependency, and only on the dependency — but not for the reason weight would -suggest.** At 35 MB total, xarray is not a heavy install; an earlier draft of this ADR leaned on -size and that was the weaker argument. +**It loses on the dependency, and only on the dependency — but not on weight.** At 35 MB total, +xarray is not a heavy install, and an argument from size would be the weak one. The real cost is **coupling**. xarray declares `pandas>=2.2` as a hard requirement (checked against xarray 2026.7.0 on 2026-08-24). pandas is precisely the kind of library a model repo @@ -185,10 +183,9 @@ package does anyway, minus the dependency. **What is genuinely lost by declining it** — stated plainly, because this is the alternative that costs the most to refuse: -- ~~lazy and out-of-core evaluation via dask~~ — **struck. This was wrong in the first draft and - the correction matters.** Out-of-core is not xarray's to give: `dask.array` provides it with a - numpy-compatible API and no pandas (see §9). Declining xarray costs less than this ADR first - claimed; the out-of-core question is a separate, cheaper decision that the draft never posed; +- **not** lazy or out-of-core evaluation. That is not xarray's to give: `dask.array` provides it + with a numpy-compatible API and no pandas (§9), so out-of-core is a separate and much cheaper + decision than this one; - zarr and netCDF, which would make `io/` mostly disappear; - `groupby`/`resample`/interpolation, all currently absent or hand-rolled; - roughly 3,700 lines of source that someone else would maintain. @@ -301,11 +298,11 @@ to this package's philosophy than pandas is. It loses for the same reason pyarrow does — columnar memory against a dense sample axis. -**The dependency argument does not apply here, and saying so matters.** An earlier draft of this -ADR grouped polars with pandas and xarray on weight; that was wrong. polars 1.44.0 declares one -hard dependency, its own compiled runtime — no pandas, no numpy requirement of its own. It is a -genuinely light install. The case against it is the memory model alone, and it is the closest -any DataFrame library comes to being viable here. +**The dependency argument does not apply here.** polars 1.44.0 declares one hard dependency, its +own compiled runtime — no pandas, and no numpy requirement of its own. It is a genuinely light +install, and it does not couple to a library other repos already pin. The case against it is the +memory model alone, which makes it the closest any DataFrame library comes to being viable +here. ### 8. awkward-array @@ -318,11 +315,10 @@ raggedness would legitimize the shape C-40 exists to prevent. ### 9. The numpy-API-compatible backends — Dask Array, JAX, CuPy -A class the first draft of this ADR missed entirely, and the omission was worth catching. The -eight alternatives above all propose a *different container with a different API*. These three -propose **the same API with a different backend**: `dask.array`, `jax.numpy` and `cupy` are each -close enough to numpy that most of this package would port with modest edits. That is a -materially different question, and it deserved asking. +A distinct class. The eight alternatives above propose a *different container with a different +API*. These three propose **the same API with a different backend**: `dask.array`, `jax.numpy` +and `cupy` are each close enough to numpy that most of this package would port with modest +edits. That makes them a different question from the rest, judged on different grounds. **One filter removes most of the class before any design argument.** This package declares `numpy>=1.26,<3` and runs a dedicated `floor` CI job at numpy 1.26.4, because consumers pin @@ -349,10 +345,10 @@ megabytes over numpy** — cheaper than zarr, half the weight of xarray, and not closure is a library another repo is likely to be pinning. **On the platform's own cost test, dask is nearly as free as numpy was.** It answers C-71 (the -dense-grid allocation) and C-73 (read-all-to-RAM) directly, which is exactly what this ADR first -mis-attributed to xarray. Every objection to it below is about **design**, not cost — and that -distinction should be kept sharp, because a cheap dependency deferred on design grounds can be -revisited by changing the design, whereas an expensive one cannot be revisited at all. +dense-grid allocation) and C-73 (read-all-to-RAM) directly. Every objection to it below is about +**design**, not cost — and that distinction should be kept sharp, because a cheap dependency +deferred on design grounds can be revisited by changing the design, whereas an expensive one +cannot be revisited at all. It loses **as the frame type** for a reason that is about this package's philosophy rather than about dask: **laziness is incompatible with fail-loud-at-construction.** ADR-008 requires that a @@ -363,22 +359,21 @@ breaking every consumer's expectation and the `assert_frame_envelope` contract w chunk boundaries would interact with the row-blocking discipline the tower estimators use to bound peak memory (C-22) in ways that need re-derivation rather than porting. -**Dask belongs in `io/`, not in the frame** — and that is now the sharpest form of this ADR's -main revisit trigger, sharper than the xarray version the first draft recorded, precisely -because it costs no pandas. +**Dask belongs in `io/`, not in the frame** — and in that position it is the sharpest form of +this ADR's main revisit trigger, precisely because it costs no pandas. -**JAX** deserves one honest paragraph beyond the floor objection, because it would have given -this package something it wanted and had to build. **JAX arrays are immutable natively.** The +**JAX** has one advantage the floor objection alone would hide, and it is a real one, because it +would have supplied something this package wanted and had to build by hand. **JAX arrays are +immutable natively.** The entire C-66 saga — write-protection recorded as a deferred MAJOR-rider on 2026-06-28, carried unshipped through eleven releases, finally landed in 2.0.0, and then only after discovering that the one-line `setflags` fix ADR-025 recorded would have made the *caller's* array read-only — would not have existed. Immutability would have been a property of the container rather than an invariant this package enforces by hand. -That is a real loss and it is worth naming. It does not outweigh forcing numpy 2 and scipy on -every consumer for a package that needs neither autograd nor XLA (ADR-017 puts estimation -outside the leaf's charter, and none of it is differentiable), but the trade should be recorded -rather than glossed. +That is a genuine loss. It does not outweigh forcing numpy 2 and scipy on every consumer for a +package that needs neither autograd nor XLA — ADR-017 puts estimation outside the leaf's charter, +and none of it is differentiable — but it is the strongest thing any rejected candidate offered. **CuPy** loses on the same grounds as torch, with one addition. The DAG-root argument applies — a GPU array library at the root means views-faoapi needs a CUDA runtime to serialize a forecast @@ -531,7 +526,7 @@ This is a live decision. Each trigger names evidence, not opinion, and none of t | **A real OOM at grid scale** — register **C-71** (a full-pgm `cartesian` target, ~259k cells × months, fed to `reindex_fill` on a sampled frame) or **C-73** (an OOM *despite* per-month sharding) | **`dask.array` inside `io/`**, with `zarr` as its storage format | Additive. The frame type does not change; the storage layer gains a lazy path. Prefer dask over xarray here: it answers the same need and requires no pandas (§9). `io/` is the designated place for evolving formats (ADR-002). | | **A non-Python consumer** appears on the contract | **format-as-contract** (option 2) | Would supersede this ADR. The executable-conformance argument only holds while every consumer can run Python. | | **The leaf is required on an autograd path** | **torch** | Would first require amending ADR-017, which puts estimation outside the leaf's charter. Very unlikely by construction. | -| **The wire format becomes the dominant access path** — consumers reading parquet directly more often than they construct frames | **pyarrow as the frame type** | Would supersede this ADR and probably ADR-013 with it. | +| **The wire format becomes the dominant access path** — consumers reading parquet directly more often than they construct frames | **pyarrow as the frame type** | Would supersede this ADR, and ADR-013 with it — the in-memory type and the wire format would become one thing. | **The dask trigger deserves a caveat the others do not.** Every alternative above is deferred on some mix of cost and design. `dask[array]` is deferred on **design alone** — it costs three @@ -566,8 +561,8 @@ narrowed what construction accepts. - **Alignment is hand-written.** `searchsorted` over a packed row view, plus the cross-level remap, is code this package maintains and xarray would have provided. - **No out-of-core path.** C-71 and C-73 are open for this reason. Both would be closed by - `dask.array` — which, unlike xarray, costs no pandas. This is the cheapest capability this - ADR declines, and the first draft mislabelled it as xarray's. + `dask.array`, which unlike xarray costs no pandas. This is the cheapest capability this ADR + declines. - **No `groupby`, `resample` or interpolation.** Consumers do this themselves or go without. - **numpy version skew is this package's problem.** The `floor` CI job exists because numpy 1.26.4 and 2.x differ in generic stubs *and* in ~1-ulp histogram binning (registers C-19 and From 435eb91dcb685c6f4fbb36fc688b078e8fe63b6c Mon Sep 17 00:00:00 2001 From: Polichinl Date: Mon, 24 Aug 2026 15:12:38 +0200 Subject: [PATCH 6/6] =?UTF-8?q?docs(adr):=20ADR-029=20=E2=80=94=20six=20cl?= =?UTF-8?q?aims=20that=20did=20not=20survive=20being=20attacked?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A falsification pass over each section, against the claim that the document contains no hand waving, ambiguity or decision gap. It did. Every finding was in the argument for numpy or in the scaling section — the parts written from reasoning rather than from measurement. The alternatives analysis survived intact. **The primary claim was an unverified universal, and partly circular.** The text said "every repository on this platform already depends on numpy". Checked across all 33 org repos: seven declare no numpy, and two of those — views-postprocessing and views-hydranet — are views-frames consumers whose route to numpy is partly this package itself. "It was already there" is circular for them. The replacement is a stronger fact that happens to be verifiable: **numpy declares an empty requirement set.** Adding it adds exactly one package and nothing else — no transitive closure, no version to reconcile. That is a property of numpy rather than an assumption about any repo, and it holds even where numpy is genuinely absent. Most repos do carry it already; the ADR now says so without the universal. **"The estimators already work in row blocks" was false for the most-used one.** Measured: `collapse.py` and `aggregate.py` have zero row-blocking. `collapse` is a single `reducer(frame.values, axis=-1)` — it materialises the whole array however it was loaded, and it is the operation any streaming story breaks on first. The memmap section now names which estimators block and which do not, and states that the access-pattern argument is reasoned rather than benchmarked, since no such benchmark exists here. **One revisit trigger could never fire.** "The wire format becomes the dominant access path — consumers reading parquet more often than they construct frames" is not observable; nothing on this platform counts that. A trigger nothing can evaluate defers its option forever while appearing to keep it open. Replaced with a condition visible in a consumer's source: a consumer shipping its own reader against the wire contract. **"Alignment only ever touches the index" was ambiguous and, on one reading, false.** Computing an alignment is index-only; applying one gathers from the value buffer, and `reindex_fill` allocates a dense buffer to do it — which is C-71, the very problem the same section calls the streaming blocker. The two are now separated, which also explains why `reindex_fill` is the operation that breaks first. **The 80% water figure was unsourced** and a 4x swing in the headline number turned on it. The sizing is now given across a range of land fractions, with the observation that no row changes the conclusion, and a pointer to PRIO-GRID for anyone who needs a real number. 625 lines. No code change; CONFORMANCE_FLOOR stays 2.0.0. Co-Authored-By: Claude Opus 5 (1M context) --- docs/ADRs/029_why_numpy_and_not_xarray.md | 69 ++++++++++++++++------- 1 file changed, 50 insertions(+), 19 deletions(-) diff --git a/docs/ADRs/029_why_numpy_and_not_xarray.md b/docs/ADRs/029_why_numpy_and_not_xarray.md index 5963f4a..7e3d173 100644 --- a/docs/ADRs/029_why_numpy_and_not_xarray.md +++ b/docs/ADRs/029_why_numpy_and_not_xarray.md @@ -55,9 +55,18 @@ ADR-002's decision, and it has a consequence most container choices ignore: This is the primary reason, and it is worth stating before any comparison of features. -**Every repository on this platform already depends on numpy.** Adding `views-frames` to a repo -therefore adds no new transitive dependency at all — the array library was already installed, -already resolved, already pinned. The contract arrived at zero dependency cost. +**numpy has no dependencies of its own.** `numpy` declares an empty requirement set, so adding +it to a repository adds exactly one package — about 19 MB — and nothing else. There is no +transitive closure to resolve, no version to reconcile against anything, and no second library +that arrives with it. That is what "free" means here, and it is a property of numpy rather than +an assumption about any particular repo. + +**Most repositories on the platform already carry it in any case**, directly or through +scipy/pandas/scikit-learn, so for them the contract costs literally nothing. It is worth being +exact rather than sweeping: a handful of repos — including two `views-frames` consumers, +views-postprocessing and views-hydranet — do *not* declare numpy themselves and reach it partly +*through* this package. For those, the argument is not "they already had it" but the paragraph +above: what they gained is one dependency-free package. That property is what the whole repository separation exists to protect. The model repos are deliberately allowed to hold *wildly* different libraries — that is the point of separating @@ -412,15 +421,21 @@ Take a plausible near-future shape — ten features, 128 samples each, a global x 4 bytes (float32) = 663.6 GB ``` -Excluding ocean cells helps but does not rescue it. If roughly 80% of cells are water and are -dropped, **20% remains**: +Excluding ocean cells helps but does not rescue it. The land fraction of a 0.5 degree global +grid is an input to this arithmetic, not a fact this ADR establishes, so the sizing is given +across a range rather than resting on one figure: -| | cells | float32 | +| Cells retained | cells | float32 | |---|---:|---:| -| Full 0.5 degree grid | 165.9 billion | **663.6 GB** | -| Land only (20% of cells retained) | 33.2 billion | **132.7 GB** | +| All (no land mask) | 165.9 billion | **663.6 GB** | +| 30% | 49.8 billion | **199.1 GB** | +| 25% | 41.5 billion | **165.9 GB** | +| 20% | 33.2 billion | **132.7 GB** | -Neither fits in memory, and ten features at 128 samples is nowhere near a ceiling. +**The conclusion does not depend on which row is right.** Every one of them is far past the +memory of any single machine, and ten features at 128 samples is nowhere near a ceiling. If a +precise land-cell count is ever needed for capacity planning, it should be taken from the +PRIO-GRID definition rather than from this table. ### Why numpy survives this: the index fits, the values do not @@ -438,11 +453,17 @@ A frame is an index plus a value array, and **they scale completely differently. The values are **320x to 6,400x** the index — and the index's size does not depend on features or samples at all. It is 4.1 MB per month whether the frame carries ten features or a hundred. -**This matters because alignment is what this package structurally does, and alignment only ever -touches the index.** `searchsorted`, `reindex`, `intersect`, `is_superset_of` and `cartesian` all -operate on `(time, unit)` and never read a value. Those are the operations needing random access, -and the data they need it over **fits in memory at global scale**. What does not fit is read -sequentially, in row blocks — the access pattern the estimators already use (C-22). +**This matters because of which half needs random access.** *Computing* an alignment touches +only the index: `SpatioTemporalIndex.searchsorted`, `reindex`, `intersect`, `is_superset_of` and +`cartesian` take an index, return positions, and never read a value. *Applying* one then gathers +from the value buffer — `frame.reindex(other)` and `reindex_fill` both move values, and +`reindex_fill` allocates a dense buffer to do it, which is exactly C-71. + +So the split is: **the random-access half operates on data that fits in memory at global scale, +and the half that does not fit is touched only by gather and by sequential row-blocked +reduction.** That is the property that makes a streaming storage layer viable underneath an +unchanged frame type — and the reason `reindex_fill` is the one operation that breaks first, since +it is the only place where the large half is both written and materialised whole. **So streaming does not require a different array library. It requires a different storage layer under the same one.** A lazy array's `.compute()` returns a numpy array; chunked stores feed numpy @@ -481,10 +502,20 @@ adds shards; adding features or samples multiplies every shard. Partly, and for longer than the raw totals suggest — but not for the thing likely to break first. **Where it works.** A frame is `(N, S)` row-major with samples contiguous, so a per-row -sample-axis reduction reads sequentially. That is the dominant access pattern in -`views_frames_summarize`, and the tower estimators already work in row blocks to bound peak -memory (C-22) — precisely the pattern memmap rewards. Reduction over a frame far larger than RAM -streams acceptably today. +sample-axis reduction reads sequentially — precisely the access pattern memmap rewards. The +estimators that already bound peak memory by working in row blocks (C-22) have exactly that +shape: `interval`, `point`, `tower`, `tower_point`, `bimodality`, `exceedance` and +`expected_shortfall`. + +**`collapse` and `aggregate` do not row-block**, and `collapse` is the most-used estimator in +the package. Its implementation is a single call — +`reducer(frame.values, axis=-1)` — which materialises the whole array regardless of how it was +loaded. Any streaming story has to either block `collapse` or accept that it is the operation +that breaks first. + +**This is reasoned from the access pattern, not measured.** No benchmark of memmap-backed +reduction at grid scale exists in this repository, and the claim should be read as "the layout +is right for it" rather than "it has been shown to work". **Where it does not.** @@ -526,7 +557,7 @@ This is a live decision. Each trigger names evidence, not opinion, and none of t | **A real OOM at grid scale** — register **C-71** (a full-pgm `cartesian` target, ~259k cells × months, fed to `reindex_fill` on a sampled frame) or **C-73** (an OOM *despite* per-month sharding) | **`dask.array` inside `io/`**, with `zarr` as its storage format | Additive. The frame type does not change; the storage layer gains a lazy path. Prefer dask over xarray here: it answers the same need and requires no pandas (§9). `io/` is the designated place for evolving formats (ADR-002). | | **A non-Python consumer** appears on the contract | **format-as-contract** (option 2) | Would supersede this ADR. The executable-conformance argument only holds while every consumer can run Python. | | **The leaf is required on an autograd path** | **torch** | Would first require amending ADR-017, which puts estimation outside the leaf's charter. Very unlikely by construction. | -| **The wire format becomes the dominant access path** — consumers reading parquet directly more often than they construct frames | **pyarrow as the frame type** | Would supersede this ADR, and ADR-013 with it — the in-memory type and the wire format would become one thing. | +| **A consumer ships its own reader against the wire contract** rather than constructing frames — visible in that repo's source, unlike "which access path is more common", which nothing on this platform measures | **pyarrow as the frame type** | Would supersede this ADR, and ADR-013 with it — the in-memory type and the wire format would become one thing. | **The dask trigger deserves a caveat the others do not.** Every alternative above is deferred on some mix of cost and design. `dask[array]` is deferred on **design alone** — it costs three