SafeStore is a Terraform-built S3 backup and recovery system that makes accidental deletion recoverable in minutes and survives a full regional AWS outage.
- Problem Statement
- Architecture
- What This Does NOT Do
- Prerequisites
- Setup / Deployment
- Repository Structure
- Usage
- Key Design Decisions
- Testing
- Cost
- Teardown
A developer ran aws s3 rm --recursive on the wrong folder. No versioning or backup existed, and the data was permanently lost. SafeStore was built by David Danso (David Danso Cloud Labs) to make that specific failure mode recoverable.
|
SafeStore provisions infrastructure across two AWS regions: us-east-1 (primary) and eu-west-1 (backup).
- Primary Bucket (
safestore-primary-<account-id>-us-east-1): versioning enabled, SSE-S3 encryption (AES256) enforced, all Public Access Block settings enabled. - Backup Bucket (
safestore-backup-<account-id>-eu-west-1): versioning enabled, SSE-S3 encryption enforced, Public Access Block enabled. Bucket policy denies non-HTTPS requests and blocksPutObject,DeleteObject, andDeleteObjectVersionfor all principals except the replication role. - Primary Logs Bucket (
safestore-logs-primary-<account-id>-us-east-1): receives S3 server access logs for the primary bucket. - Backup Logs Bucket (
safestore-logs-backup-<account-id>-eu-west-1): receives S3 server access logs for the backup bucket. Two separate log buckets are required because AWS requires access-log delivery to remain in the same region as the source bucket. - Replication: one-way asynchronous Cross-Region Replication (CRR) from primary to backup, authorized by a dedicated IAM role (
safestore-replication-role) that onlys3.amazonaws.comcan assume. - Lifecycle Management: primary and backup automatically expire noncurrent versions after 30 days and clean up delete markers once nothing is left beneath them. Both logs buckets expire objects after 90 days. Nothing accumulates indefinitely.
- Monitoring: a CloudWatch alarm watches
BucketSizeByteson primary, and — beyond what was originally required — a second alarm watches backup too, since delete markers aren't replicated and backup can grow independently of primary over time. This metric is reported by AWS once per day, not continuously.
- Does not backfill existing objects — objects uploaded before replication was configured are not automatically replicated to backup.
- Does not replicate delete markers — deleting an object in primary creates a delete marker locally; that deletion does not propagate to backup.
- Does not consolidate logs into a single bucket — AWS requires access-log target buckets to be in the same region as the source bucket.
- Does not alert in real time — CloudWatch reports S3
BucketSizeBytesonce per day, not continuously.
- AWS account with access to
us-east-1andeu-west-1 - A dedicated IAM user for deployment — not root — with MFA enabled
- AWS CLI installed and configured
- Terraform (AWS provider
hashicorp/aws ~> 6.0) - Python 3.x with Boto3 (
pip install boto3) - IAM permissions to create S3 buckets, IAM roles/policies, and CloudWatch alarms
- Clone the repository:
git clone https://github.com/DavidDanso/safe-store.git
cd safestore- Create a
terraform.tfvarsfile interraform/:
account_id = "YOUR_ACCOUNT_ID_HERE"
primary_region = "us-east-1"
backup_region = "eu-west-1"
alarm_threshold_gb = 1- Initialize Terraform:
cd terraform
terraform init- Review the plan:
terraform plan- Apply:
terraform apply- Run the verification scripts:
python3 ../scripts/recover.py
python3 ../scripts/check_replication.py/terraform → all infrastructure code
/scripts → recover.py, check_replication.py
/docs → ADR.md (architecture decision records)
BUILD_LOG.md → raw log of what broke and how I fixed it
TEST_RESULTS.md → Test case results with evidence
Recovery Verification — scripts/recover.py
python3 scripts/recover.py- Uploads a test object to the primary bucket.
- Deletes the object, creating a delete marker.
- Confirms the object returns 404 via
head_object. - Finds the current delete marker via
IsLatest. - Deletes the delete marker by its
VersionId, restoring the object. - Verifies the restored object matches the original size and content.
- Cleans up test versions, leaving the bucket clean.
Replication Verification — scripts/check_replication.py
python3 scripts/check_replication.py- Uploads a uniquely-named test object to primary.
- Polls the backup bucket until the object appears.
- Checks
ReplicationStatusmetadata on the primary object. - Reports pass/fail.
| Decision | Selection | Rationale |
|---|---|---|
| Delete Marker Replication | Disabled | Prevents an accidental deletion on primary from also wiping the backup copy. |
| Replication Strategy | Cross-Region Replication (CRR) | Same-Region Replication can't guarantee recovery if the primary region has an outage. |
| Retroactive Replication | Excluded | S3 replication doesn't backfill pre-existing objects by default; that requires a separate batch operation. |
| Logging Architecture | Two regional log buckets | AWS requires access-log target buckets to be in the same region as the source bucket. |
| Encryption Provider | SSE-S3 (AES256) | Avoids KMS key management cost and overhead; no requirement here for key rotation control or usage auditing. |
| Storage Monitoring | Alarm on both primary and backup | Delete markers aren't replicated, so backup can diverge in size from primary over time — primary-only monitoring could miss abnormal growth specific to backup. |
10 of 10 PRD test cases pass. recover.py and check_replication.py both run successfully against real infrastructure, with output confirming recovery and replication end-to-end. Full evidence recorded in TEST_RESULTS.md. Day-by-day build history — real errors and how they were fixed — is in BUILD_LOG.md.
- Budget ceiling: $0.05 total AWS spend.
- Confirmed via Cost Explorer, filtered by tag
Project=SafeStore.
No KMS key exists in this project, so there's no ongoing charge either way — safe to leave running for demos or tear down immediately.
- Destroy all infrastructure:
cd terraform
terraform destroy- Confirm nothing was left behind:
aws s3 ls— confirm all 4 buckets (primary, backup, primary logs, backup logs) are gone.- IAM console — confirm
safestore-replication-roleis gone. - CloudWatch, in both
us-east-1andeu-west-1— confirm theBucketSizeBytesalarms are gone.
David Kofi Danso
