Skip to content

fix: api compression and egress accounting - #43

Merged
Zacgoose merged 2 commits into
mainfrom
dev
Sep 14, 2026
Merged

Zacgoose merged 2 commits into
mainfrom
dev

Conversation

@Zacgoose

Copy link
Copy Markdown
Contributor

This pull request introduces several enhancements and refactorings to the API egress accounting, dynamic response compression, and configuration systems. The changes separate API response compression from static asset compression, add more granular and configurable egress accounting (including time-bucketed table storage), and split the egress limiting logic into distinct policy and accounting middleware for accuracy and maintainability.

API Compression and Configuration Separation:

  • Added ApiSettings to allow independent configuration of dynamic /api response compression, including toggles and compression level, decoupled from static frontend compression settings. [1] [2]
  • Extended CraftRoles to distinguish between static and API compression, with new environment variable overrides and configuration options. [1] [2] [3] [4]

API Response Compression Improvements:

  • Introduced logic to resolve the API compression level from configuration or environment, defaulting to a safe value if unset or invalid, and applied it to both Brotli and gzip providers. [1] [2]

Egress Accounting Enhancements:

  • Expanded EgressLimitSettings to support table-based, time-bucketed egress accounting with configurable bucket width and retention, all overridable via environment variables.
  • Added resolved property accessors for table name, bucket width, and retention days, honoring environment overrides and enforcing minimums.

Middleware Refactoring for Egress Accounting:

  • Split the egress limiter into a policy middleware (ApiEgressLimiterMiddleware) that decides whether to shed or allow requests and marks allowed requests for billing, and a new outer middleware (ApiEgressWireCounterMiddleware) that measures and records the actual on-the-wire (post-compression) response size for billing. [1] [2] [3]

These changes improve configurability, accuracy, and maintainability of API response compression and egress billing in the system.

Make dynamic /api response compression a first-class, independently-toggled
capability, and fix the egress cap to bill the compressed on-the-wire size
instead of the pre-compression body.

Compression
- App:Api:Compression / CRAFT_API_COMPRESSION (default on), resolved in
  CraftRoles.ApiCompressionEnabled, independent of the static Frontend.Compression
  toggle — an origin behind a CDN that compresses static assets still compresses
  its API JSON, and vice versa.
- App:Api:CompressionLevel / CRAFT_API_COMPRESSION_LEVEL (Fastest | Optimal |
  SmallestSize | NoCompression) applied to both providers, surfaced on the startup
  "[System] Compression" line. Default Optimal: measured on 2 vCPU, Brotli Optimal
  ~8.4x vs Fastest ~4.9x at the same ~14% CPU and p95, so it is a near-free win.
  SmallestSize stays off by default (Brotli q11 pegged both cores, p95 in the tens
  of seconds); drop to Fastest/NoCompression on a very small SKU.
- Program.cs splits the response pipeline by path (UseWhen): /api and static
  compression are governed by separate toggles, so an /api body is never
  double-compressed.

Egress accounting (policy + counting, no compression logic)
- ApiEgressLimiterMiddleware now only classifies and sheds (429), flagging greenlit
  API requests via ApiEgressWireCounterMiddleware.ChargeItemKey.
- New ApiEgressWireCounterMiddleware runs outside the response compressor (outermost
  /api link) and records the flagged response's wire bytes, so the cap charges what
  actually leaves the box. It must sit outside compression because the compressor
  only flushes its trailing block on unwind — a counter inside undercounts.

Tests
- ApiCompressionPipelineTests: real Kestrel (dynamic port), six pipeline shapes
  (bare, wire counter, UseWhen, routed endpoint, charset, full realistic chain) all
  confirm /api responses come back compressed and the counter does not suppress it.
  Adds Microsoft.AspNetCore.TestHost as a permanent test dependency.
- Egress limiter/counter unit tests (shed vs flag, never touches UI, counts the
  compressed size when compression runs inside it, records only flagged requests,
  restores the body on throw). CraftRoles tests for the two toggles' independence
  and the level resolver (env override, unrecognised falls back to Fastest).

perf-harness
- run-compression.ps1: correctness + accounting proof (identity uncompressed vs gzip
  compressed, ledger delta == bytes received in both) then CPU-vs-bandwidth per
  encoding under k6 + docker stats.
- run-compression-levels.ps1: sweeps the level (recreating per level) reporting
  ratio, CPU and p95 for each level x encoding.
- api_load.js gains an ENC knob (Accept-Encoding request header).
Extend the egress ledger from a single instance-wide counter to per-client
accounting, and mirror a queryable history to a table so the product can show
usage over time and whether/when/how often the cap was hit - without changing
what the cap enforces.

Accounting
- Record now takes the caller's AppId; the limiter carries it on the charge flag
  (instead of a bare bool) and calls RecordShed on a 429. The ledger keeps per-AppId
  daily totals (bytes/requests/shed/last-seen), a shed count, and the cap-reached
  time alongside the instance total.
- The local file becomes v2: daily totals + per-client map + cap state (still the
  source of truth for the cap; v1 files still load). No per-flush samples in the file.

Table mirror (EgressTableSchema)
- Two row kinds in one table (RowKey prefix): bkt_ 15-minute buckets (per-client +
  instance aggregate) for the time-series, and day_ daily audit rows carrying the cap
  in force, cap-reached time, shed count and enforcing state. PartitionKey is the AppId
  (or instance-total), RowKey a sortable UTC stamp, so "last 24h" and retention are
  range queries. Written through Craft's own store - the same account CIPP reads with
  Get-CIPPTable - so CIPP queries it directly (no bridge).
- Mirrored on the flush loop after the file write, only for clients active since the
  last sync; finalised buckets are written then dropped from memory. Backfill on start
  reseeds the current bucket from the table so a mid-bucket restart doesn't regress its
  row. A periodic purge drops rows past the retention window. Table failures never
  affect the cap; the file is the truth.

Config under App:RateLimit:Egress: TableName (blank = file-only), BucketMinutes (15),
RetentionDays (7), with CRAFT_API_EGRESS_TABLE / _BUCKET_MINUTES / _RETENTION_DAYS
overrides.

Tests: per-client totals + shed + cap-reached in the file; table bucket/daily row
shape, aggregate == sum, backfill no-regression, and retention purge via an in-memory
store.
@Zacgoose
Zacgoose merged commit a5d2d36 into main Sep 14, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants