Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Waterline

Measures the surface area of reservoirs from satellite imagery, through cloud, without downloading the satellites.

A Sentinel-2 band file is about 200 MB and a reservoir occupies a fraction of a percent of it. Waterline reads only the tiles it needs, straight out of the public archive on S3, over HTTP range requests — 2.33 MB of a 191.9 MB scene, in three requests, to measure a reservoir.

When cloud makes optical imagery useless — which over India means most of the monsoon, when water is changing fastest — it switches to Sentinel-1 radar, which sees through weather. The two agree with each other to a third of a percent.


The result

Krishna Raja Sagara, the reservoir that supplies Mysore and feeds Bengaluru, measured across a full year from two independent sensors:

2025-04-15   73.68 km²   optical    LOW confidence
2025-05-07   63.21 km²   optical               <- dry-season minimum
2025-06-09   79.02 km²   radar                 <- monsoon onset
2025-06-21   92.95 km²   radar
2025-07-03   97.82 km²   radar                 <- at capacity in three weeks
2025-07-15   96.15 km²   radar
2025-07-27  106.86 km²   radar      LOW confidence
2025-08-08   97.12 km²   radar
2025-08-20   97.45 km²   radar
2025-09-01   97.16 km²   radar
2025-09-13   96.41 km²   radar
2025-09-25   96.21 km²   radar
2025-10-07   95.67 km²   radar
2025-10-19   96.65 km²   radar                 <- handover
2025-11-01   96.34 km²   optical               <- 0.33% apart, 13 days later
2025-12-13   96.02 km²   optical
2026-01-15   97.24 km²   optical
2026-02-04   90.97 km²   optical
2026-03-11   84.05 km²   optical
2026-03-26   78.40 km²   optical    LOW confidence

Every optical date between June and October is missing because there was no usable one: 36 Sentinel-2 scenes passed over KRS in that window with a median cloud cover of 92.3%, and exactly one was under 20%. Radar had twelve passes and produced twelve measurements. That gap is the entire reason the radar path exists.

Checked against the official gauge

Surface area measured from orbit, against live storage published by India's Central Water Commission — 1,090 measurements across 2025, paired with the nearest bulletin within three days:

reservoir n storage CV rank correlation
Supa (Kalinadi) 29 32.5% 0.955
Bhadra 27 36.4% 0.954
Kabini 54 27.8% 0.893
Krishnaraja Sagar 54 34.8% 0.816
Tattihalla 115 95.7% 0.780
Mani Dam 40 41.7% 0.497
Salaulim 38 29.3% 0.456
Vani Vilas Sagar 60 6.4% excluded — no range

Rank correlation rather than a difference in km², because area and volume are different quantities: a steep gorge gains almost no surface as it fills, a shallow basin gains a lot. The claim being tested is that satellite area tracks real storage, and for four reservoirs it tracks it above 0.8.

Vani Vilas is excluded rather than credited. Its storage moved 6.4% across an entire year, so the comparison had nothing to work with, and its 0.605 is noise that landed on a flattering number. Every correlation is reported next to the dynamic range that makes it meaningful — for the reason in decision 015, which is the most instructive thing in this repository.

Two sensors, one reservoir

The handover is the other strong evidence here. Radar on 19 October reads 96.649 km²; optical on 1 November reads 96.335 km². Different physics — backscattered microwaves against reflected sunlight — different geolocation machinery, independently chosen thresholds, thirteen days apart, agreeing to 0.33%. Neither shares a failure mode with the other that could produce that by accident.


How it works

  STAC catalogue ──▶ scene list ──▶ SQLite queue ──▶ worker pool
                                                          │
                                       ┌──────────────────┴─────────────────┐
                                       ▼                                    ▼
                                  optical path                         radar path
                                       │                                    │
                          HTTP range reads, 2 bands          HTTP range reads, VV
                                       │                                    │
                              NDWI = (G−N)/(G+N)               Lee speckle filter
                                       │                                    │
                                       │                          20·log10 → dB
                                       │                                    │
                                       └──────────┬─────────────────────────┘
                                                  ▼
                                    Otsu threshold from the histogram
                                                  ▼
                                    4-connected component labelling
                                                  ▼
                                  body whose centroid is in the catalogue
                                                  ▼
                                  area summed per pixel from the geolocation

Reading pixels. A Cloud-Optimized GeoTIFF places its header, IFD chain and tile index at the front of the file so a client can learn the whole layout in one request, then fetch any rectangle as a handful of byte ranges. Waterline implements that directly — TIFF parsing, tile indexing, Deflate, horizontal predictor — rather than delegating to GDAL, because that mechanism is the project. Adjacent tiles are coalesced into single requests, since a round trip to Oregon costs more than a few hundred kilobytes of wasted transfer.

Geolocation. Sentinel-2 is projected to UTM, so an affine transform locates it; the transverse Mercator series is implemented here rather than pulled from PROJ, and checked against invariants that do not depend on the implementation. Sentinel-1 GRD is not map projected at all — it arrives in radar ground-range geometry with a 21 × 10 lattice of 210 control points, and everything between them is interpolated. Forward that is bilinear; the inverse has no closed form and is solved with Newton's method.

Classification. Neither sensor has a usable fixed threshold. Water in NDWI should sit between +0.3 and +0.6; at KRS it reads +0.05, because the reservoir is silty and turbidity lifts near-infrared reflectance. So the threshold is chosen from each scene's own histogram by Otsu's method — and guarded by a bimodality test that is not Otsu's own separability measure, because that cannot tell one population from two. See docs/DECISIONS.md, 006 and 009.


What it cannot do

Stated plainly, because the numbers above are only worth something if the limits are on the same page.

  • Validation is rank-based, not absolute. Waterline measures surface area; the gauge publishes stored volume. They are related by a reservoir's bathymetry, so the two can be shown to move together but not to agree numerically. There is still no independent check of absolute area in km².
  • Two of eight gauged reservoirs disagree (Mani Dam 0.497, Salaulim 0.456), undiagnosed and published anyway.
  • The reference data is second-hand. CWC publishes PDFs and India-WRIS was unreachable, so the storage readings come from an aggregator that republishes the bulletins. Probably faithful, but not the same claim as reading them.
  • Waterbodies larger than one Sentinel-2 tile are refused, not measured. A reservoir spanning several MGRS tiles gets a different partial answer from each, and waterline declines rather than guessing. Mosaicking across tiles — including across UTM zones — is real work and is not done. See decision 014, which is also the story of how this was found.
  • Radar is uncalibrated and not terrain corrected. Digital numbers, not gamma-nought. Adaptive thresholding absorbs the missing calibration; the missing terrain correction means absolute shoreline position is a weaker claim than relative change between repeat passes. See decision 011.
  • A branched reservoir can be measured as one arm. Two Srisailam radar dates read about half the surrounding series. Recorded, not yet diagnosed.

The map

./build/RelWithDebInfo/waterline_server.exe --store run/waterline.db --catalogue data/validation.tsv

Serves the map on http://localhost:8080, read-only, off the same SQLite file a backfill may be writing to. Click a reservoir for its series: optical in blue, radar in amber, hollow where confidence was low, and the CWC storage curve behind it on its own axis.

--export page.html writes the whole thing as one self-contained file instead, data inlined, built from the same snapshot the live API returns.

A published snapshot, without the basemap the artifact sandbox blocks: https://claude.ai/artifact/D6nM8Up6HHczv6SWk88AKE

Building

Windows, MSVC, C++20. Dependencies come from vcpkg:

vcpkg install curl:x64-windows zlib:x64-windows sqlite3:x64-windows jsoncpp:x64-windows gtest:x64-windows benchmark:x64-windows
cmake -S . -B build -DCMAKE_TOOLCHAIN_FILE=C:/vcpkg/scripts/buildsystems/vcpkg.cmake -DVCPKG_TARGET_TRIPLET=x64-windows
cmake --build build --config RelWithDebInfo

There is no database server and no container to start. The queue and the results live in one SQLite file, so a clone plus a build is enough to run anything here — the same reasoning that rejected a dataset needing credentials.

Running

Inspect a scene and read one pixel, without downloading it:

./build/RelWithDebInfo/waterline_probe.exe

Measure one reservoir on one date:

./build/RelWithDebInfo/waterline_measure.exe

Measure it from radar instead, on a date buried under monsoon cloud:

./build/RelWithDebInfo/waterline_measure_sar.exe

Back-fill a catalogue of waterbodies over a date range, with four workers:

./build/RelWithDebInfo/waterline_backfill.exe data/karnataka.tsv run/waterline.db --bodies 20 --from 2025-09-01 --to 2025-12-31

Kill it at any point and run the same command again. It recovers whatever was in flight and does not repeat what was finished.

Tests

./build/RelWithDebInfo/waterline_tests.exe

104 tests, no network, clean under /W4 /WX. They check against things that do not depend on the implementation wherever possible: the projection against the published WGS84 quarter-meridian, polygon area against the closed form for a latitude/longitude rectangle, normalized indices against the bound that says they cannot leave [−1, 1].

Several exist specifically to pin findings that cost time to discover, and are named so:

  • Otsu.SeparabilityAloneCannotDetectBimodality
  • LeeFilter.KeepsAnEdgeSharpInsteadOfBlurringIt
  • Store.WorkClaimedByAProcessThatDiedIsRecoveredOnRestart
  • ConnectedComponents.DoesNotJoinRegionsThatOnlyTouchDiagonally

Layout

include/waterline/   public headers
src/                 implementation and the four executables
tests/               GoogleTest suites
data/                waterbody catalogue, from OpenStreetMap
docs/PLAN.md         six stages, each with an acceptance gate and its evidence
docs/DECISIONS.md    append-only design log

docs/DECISIONS.md is the thing to read if you read one file. Every entry records what was decided, what else was considered, and what it cost — including the ones that were wrong and had to be superseded.

Data

All of it public and anonymously accessible, no account required:

  • Sentinel-2 L2A COGs — sentinel-cogs, us-west-2, AWS Open Data
  • Sentinel-1 GRD — sentinel-s1-l1c, eu-central-1
  • Scene discovery — Earth Search
  • Waterbody catalogue — OpenStreetMap via Overpass, ODbL licensed

Copernicus Sentinel data, processed by ESA.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages