Skip to content

[P2] Multi-site delete-marker purge: in-flight or cross-site creations and post-crash minority residue #217

Description

@Vonng

Summary

Three source repairs on main (eb4f5e5b3, 254b19ac0, 358ab38fb; all after the published Server 20260903) close the reproduced intra-site path by which an exact-version purge of a delete marker was undone in a multi-site replication mesh. Two parts of the original defect remain open and are tracked here. They are inherited behavior, not regressions of the repairs, and the 2026-09-16 release decision does not depend on them.

What is fixed on main

  • A pure exact-version purge no longer creates the marker on drives that lack it, a retried purge of a missing version needs a write-quorum majority of absent drives before it is acknowledged, purge results merge "removed" and "reliably absent", and healing a delete marker preserves its stored replication and purge metadata (eb4f5e5b3).
  • A queued delete-marker creation (from the DELETE handler, a GET/HEAD/LIST heal, the scanner or MRF) is re-read against the source under the per-object replication lock before it is sent. A missing version, a non-marker version or a version under purge makes the task stale and it is dropped; an unconfirmed read is retried through MRF instead of being treated as absence (254b19ac0). This is the observed sequence from the three-site runs: purge acknowledged by both targets, then about 85 ms later a replicated creation for the same VersionID and original mtime, leaving the source with 0/16 marker copies and each target with 16/16.
  • The purge response reports a delete marker only for a stored marker (358ab38fb).

What remains open

  1. Creations already in flight, or replayed from another site. The source-side re-check cannot intercept a request that has already left the source, nor a creation originating at another site. Closing this needs a receiver-side decision: a persistent per-VersionID purge record, or a monotonic sequence/epoch that lets a receiver reject an older creation. Time-based expiry alone cannot cover indefinitely delayed or offline peers, so the design has to define peer confirmation or an epoch-based reclaim rule, plus multi-pool, rolling-upgrade and PGSTY-stack coordination. Rough size: several person-weeks of design, implementation, migration and adversarial acceptance.
  2. Minority residue after a crash. With the creation replicated and COMPLETED, one target process (4 of 16 drives) stopped, and the purge acknowledged by the remaining 12 drives, killing every process after the acknowledgement left the 4 marker copies in place. During 450 s and 7 scanner cycles per site after restart, no heal touched them. All 12 frontends read the data version, so this is not a visibility failure, but no persistent owner of the cleanup is proven, and it was not shown that the scanner ever selects that version. The dangling-version heal removes such copies when it runs on that object; the open question is which component guarantees that it runs, and when.

Acceptance

  • A written decision on the receiver-side barrier (persistent purge record vs. sequence/epoch), including reclaim rules and rolling-upgrade behavior across the maintained stack.
  • Deterministic regressions for: purge acknowledged, then an in-flight creation for the same VersionID; a creation replayed by a second site; majority-acknowledged purge followed by loss of all in-memory queues.
  • A three-site run on the final build covering the batched-recovery rounds, the late-create sequence and the crash sequence, each observed for at least the previous 450 s window with a stable tail.

Evidence

Fixed binaries, per-drive xl.meta snapshots, HTTP traces and scanner cycle counts for the runs described above are kept outside the repository by the maintainer (2026-09-16 three-site ARM64 container runs, EC 12+4, 12 processes, 48 drives). The reproduction used only ordinary S3 operations: PUT, DELETE without versionId, DELETE with the marker's versionId, and GET/ListObjectVersions on every frontend.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    • Status
      Backlog

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions