Phase 22: real lakehouse archival writer for payment_archival_jobs - #68
Merged
Merged
Conversation
Polls pending payment_archival_jobs rows, claims them idempotently
(UPDATE ... WHERE status='pending' RETURNING), exports committed
payment_queue rows in the job's tier window as CSV, PUTs to the
S3-compatible sink (RustFS via @aws-sdk/client-s3, an existing
dependency) at s3://$PAYMENT_ARCHIVE_SINK_BUCKET/{tier}/{date}/{jobId}.csv,
and marks the job completed with the real storageUri/bytesWritten/
transfersArchived — or failed with errorMessage. No fabricated URIs.
…startup (part 3) The part-1 commit left server/_core/index.ts truncated at 762 lines (dropping startPaymentWorker, runPaymentArchivalCron, balance drift cron, and all later startup wiring). This restores the full main content and adds startPaymentArchivalWriter()/stopPaymentArchivalWriter() beside the payment worker lifecycle (SIGTERM/SIGINT), plus corrects the tier comments (CSV, not Parquet).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The payment-archival cron (
runPaymentArchivalCroninserver/_core/index.ts) honestly enqueuespayment_archival_jobsrows withstatus='pending'(NULL storageUri, 0 bytes) whenPAYMENT_ARCHIVE_SINK_BUCKETis set — but nothing ever fulfilled them. This PR adds the real writer.Files
server/workers/paymentArchivalWriter.ts(new) — lakehouse writer, styled onserver/paymentWorker.ts:UPDATE payment_archival_jobs SET status='running' WHERE id=? AND status='pending' RETURNING— a concurrent claim updates 0 rows and is skipped (exactly-once archival).payment_queuerows in the job's[periodStart, periodEnd)tier window (same predicate the cron counted with), bounded at 500k rows/job.s3://$PAYMENT_ARCHIVE_SINK_BUCKET/{tier}/{YYYY-MM-DD}/{jobId}.csv.completedwith the real storageUri, actualbytesWritten(measured from the uploaded payload) and truetransfersArchivedcount — only after the object exists. Any sink error →status='failed'+errorMessage, storageUri stays NULL. No fabricated URIs.startPaymentArchivalWriter()/stopPaymentArchivalWriter()/ status getter; disabled with a single log line whenPAYMENT_ARCHIVE_SINK_BUCKETis unset.server/_core/index.ts— starts/stops the writer besidestartPaymentWorker(), withSIGTERM/SIGINTgraceful stop. Also corrects the tier comments (CSV, not Parquet). Note: the earlier commit on this branch had leftindex.tstruncated at 762 lines (droppingstartPaymentWorker, the archival cron and balance-drift cron); this PR restores the full file from latest main plus the writer wiring.server/workers/paymentArchivalWriter.test.ts(new) — vitest, DB and S3 mocked: disabled-without-bucket, happy path (real URI/bytes/count asserted against the actual payload), sink-error → failed + errorMessage + no storageUri, and the 0-row-claim concurrency skip.CSV vs Parquet decision
CSV.
package.jsonhas no parquet writer (no parquetjs/duckdb/arrow), so Parquet would require a new runtime dependency; the object key extension (.csv) reflects the real format. Format upgrade to Parquet can be a follow-up behind the same job contract.Sink client
Uses the already-declared
@aws-sdk/client-s3dependency pointed at the RustFS endpoint (PAYMENT_ARCHIVE_SINK_*env overrides, falling back toRUSTFS_ENDPOINT/ACCESS_KEY/SECRET_KEY/REGION). The document-vaultrustfsSvcClient.tswas deliberately not reused: it is scoped to the single vault bucket (RUSTFS_BUCKET) and force-prefixes keys with a caller namespace, so it cannot write to the operator-configured sink bucket at the required key pattern. No new runtime dependencies.Validation
drizzle/schema.ts(payment_archival_job_statusenum includespending/running/completed/failed;errorMessage,completedAt,storageUri,bytesWrittenbigint,createdAtdefaultNow all present).tsc/vitest not run locally (no sandbox network for pnpm install); CI should runpnpm check+pnpm test.