Skip to content

[Epic] P3 — Backup, Restore & PITR (v0.7) #166

Description

@habibtalib

Epic — v0.7 production-hardening, phase P3: Backup, Restore & PITR

Goal: an operator can take scheduled backups and restore to an exact point in time, verified.

High

  • PITR is inert scaffolding — execute_restore, create_base_snapshot, WalArchiver, SnapshotCatalog have zero production callers; no base-snapshot task is scheduled; resolve_pitr always returns None. Reads as shipped but cannot be invoked. Replay also ignores PitrTarget.replay_lsn (no upper bound). storage/snapshot_executor.rs:62, storage/snapshot_writer.rs:167, storage/snapshot.rs:147, wal/archiver.rs:99. Every caller of each is a test module. resolve_pitr is None end-to-end because nothing populates SnapshotCatalog in production. Replay bounds only the LOW side — dry_run_restore (storage/snapshot_restore.rs:61) checks replay_lsn < base_snapshot.begin_lsn and never an upper bound. The filename parser is broken independently: parse_segment_filename (wal/archiver.rs:233) splits on three hyphen-separated parts, expecting wal-{first}-{last}.seg; no writer produces that name. Done in Make cluster consensus, crash recovery, and restore crash-safe #403. PITR runs from nodedb restore (ctl/restore/) with --target-time, --target-lsn, or --cluster --restore-point. Replay stops at the target (ctl/restore/segment.rs, cut_rule.rs). The WAL archiver and base-snapshot cycle run in production (control/pitr/cycle.rs). Segment names round-trip (nodedb-wal/src/segment/meta.rs).

Medium

  • WAL-archive upload failures only warn!ed and truncation proceeded, leaving a hole in the archived stream. A failed upload now keeps the segment on local disk and holds truncation back at the first unarchived segment (control/checkpoint_archival.rs, failed_first_lsns), and files the wal_archival_failed_truncation_held diagnostic. 1744dec.
  • Hand-rolled UTC/leap-year math for PITR targets. parse_utc_timestamp now parses with chrono (DateTime::parse_from_rfc3339, then NaiveDateTime for the space and T forms) in storage/snapshot_restore.rs. 98dd82e.
  • Restore staleness gate bypassed on lock poisoning → stale envelope overwrites newer writes. Both directions failed open: the reader took .lock().ok()...unwrap_or(0), and advance_tenant_write_hlc skipped the write entirely under if let Ok. A poisoned mutex therefore stopped recording writes AND reported a high-water mark of 0, which every watermark clears. Both now recover the lock (unwrap_or_else(|p| p.into_inner())) — the map is a plain HashMap a panic elsewhere cannot corrupt — and the read is behind a single SharedState::tenant_write_hlc accessor. 36296e0.
  • No restore verification (post-restore row-count/checksum reconciliation). RestoreStats counts only what the envelope itself carried (orchestrate/restore.rs:96-105), never a comparison against the source. Done in Make cluster consensus, crash recovery, and restore crash-safe #403. RESTORE DATABASE checks per-collection row counts and digests against the backup (control/backup/verify/). nodedb restore checks segment CRCs only.
  • Backup is full-snapshot only — no incremental/differential, no scheduling/retention. control/backup/ exposes one entry point, backup_tenant (orchestrator.rs:37). Done in Make cluster consensus, crash recovery, and restore crash-safe #403. Scheduled backups with keep retention (control/backup/schedule/), and PITR base-snapshot retention. PITR base snapshots are incremental: only changed chunks upload. BACKUP DATABASE stays a full logical copy.

Success: scheduled base snapshots + WAL archiver wired; restore replays to a target LSN/timestamp; end-to-end "base + WAL → restore to T, assert state == state@T" test passes.

Re-verified against main @ 1ff3551 (2026-09-28): the archive-upload and UTC-parser items are fixed and checked. PITR wiring, the replay upper bound, restore verification, and incremental backups remain open.

Activity

  1. changed the title [-][Epic] P3 — Backup, Restore & PITR (v1.0)[/-] [+][Epic] P3 — Backup, Restore & PITR (v0.8)[/+] on Jul 5, 2026
  2. changed the title [-][Epic] P3 — Backup, Restore & PITR (v0.8)[/-] [+][Epic] P3 — Backup, Restore & PITR (v0.7)[/+] on Jul 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:storage-walWAL, durability, flush/replaytype:epicTracking issue spanning multiple sub-issues

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions