Skip to content
krav01Public

About

Go Kubernetes operator and persistent REST API: Pod self-healing, serial template rollouts, and verified Docker/kind E2E tests.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Sunday System

Verify Sunday System

My Go and Kubernetes home assignment: a custom controller that restores a managed Pod after deletion, paired with a groceries API that keeps its data across Pod replacements.

Author: Vladimir Krauchuk (@krav01)
Stack: Go · Kubernetes · controller-runtime · Docker · kind · GitHub Actions

I split the project into two components:

  • EtherealPod, a namespaced Kubernetes custom resource that maintains one active Pod and exposes its restart count through kubectl get eps.
  • SundayApp, a small groceries API whose data is atomically persisted on a PersistentVolumeClaim, so deleting or replacing its Pod does not lose data.

A short note on my approach

I started with the failure cases rather than the HTTP endpoints: what happens when a container exits, when the whole Pod disappears, and where the data lives while that happens. That led to two separate responsibilities. Kubernetes restarts a failed container in the same Pod, while the custom controller replaces a Pod that is deleted or reaches a terminal phase.

I kept the solution small enough for a home assignment. I used the Go standard library where it was enough and introduced abstractions only at boundaries that benefited from testing. The result is deliberately small, but the tradeoffs and the next production steps are explicit.

Architecture

The repository is one Go module with two small binaries:

api/v1alpha1/       EtherealPod API types
cmd/controller/     controller-runtime composition root
cmd/sundayapp/      HTTP service composition root
internal/controller reconciliation logic
internal/httpapi/   HTTP transport and validation
internal/store/     atomic JSON file persistence
config/             CRD and RBAC resources
examples/           sample EtherealPod
e2e/                live-cluster acceptance check
docs/               architecture rationale and demo walkthrough
submission.yaml     deployable cluster resources
requirements.sh     repeatable local kind setup
.github/workflows/  remote Docker/kind verification

Dependencies are wired manually. The API depends on a narrow storage interface, while the controller owns Kubernetes-specific reconciliation. I kept these boundaries visible because they make the important behavior easy to test without adding a large application framework.

Go and controller-runtime provide the native Kubernetes control loop; direct Pod ownership keeps restart reporting unambiguous; atomic JSON persistence on a PVC with a lifetime writer lock provides single-writer persistence without operating a database. The alternatives I considered, failure semantics, and production tradeoffs are documented in docs/ARCHITECTURE.md.

Prerequisites and quick start

Install Docker, kind, and kubectl, then run:

git clone https://github.com/krav01/homework.git
cd homework
./requirements.sh

The script creates (or reuses) a kind cluster named sunday-system, builds and loads both images, installs the CRD, and applies submission.yaml. That manifest contains the controller, persistent storage, Service, and the EtherealPod running SundayApp. Re-running the script rolls the controller and application Pod so they use the newly built images while preserving API data.

Inspect it with:

kubectl get eps -n sunday-system
kubectl get pods -n sunday-system
kubectl port-forward -n sunday-system service/sunday-app 8080:80

API

Method Endpoint Behavior
POST /write Add a quantity of a product for a user; supports durable idempotency
GET /get_product_amount Sum a product's quantity across all users
DELETE /delete_product Remove a product from every user

Names and product names must contain lowercase ASCII letters. Amounts must be positive integers. POST /write accepts JSON, URL query parameters, or form parameters because the assignment does not prescribe an encoding. JSON uses only fields in the body; form fields take precedence over query parameters. Bodies are limited to 1 MiB. Malformed or null JSON and invalid form/query encoding return 400, oversized bodies return 413, and unsupported content types return 415. A body requires an explicit supported content type.

POST /write also accepts an optional Idempotency-Key header. The key may be up to 128 visible ASCII bytes. The first request atomically persists both the increment and a fingerprint of the operation. Repeating the same request with the same key returns the original result without incrementing again, including after a process or Pod replacement. Replayed responses include Idempotency-Replayed: true. Reusing the same key with a different user, product, or amount returns 409 Conflict.

curl -sS -X POST http://127.0.0.1:8080/write \
  -H 'Content-Type: application/json' \
  -H 'Idempotency-Key: example-write-1' \
  -d '{"user_id":"loki","product_name":"apple","amount":1}'

curl -sS 'http://127.0.0.1:8080/get_product_amount?product_name=apple'

curl -i -X DELETE \
  'http://127.0.0.1:8080/delete_product?product_name=apple'

Product amounts are summed across all users. Deleting a product removes it from every user. Missing products return amount 0 on GET and 404 on DELETE.

Persistence and retry semantics

The store writes one versioned JSON state using temporary-file write, file fsync, atomic rename, and directory fsync. Existing files from the original groceries-only format are still accepted and are migrated to the versioned format on the next successful write.

Idempotency records are stored in the same atomic state transition as the data they protect. The raw Idempotency-Key is not persisted; the store keeps its SHA-256 digest and a request fingerprint. This means a successful operation and its replay protection cannot diverge across a normal restart.

If directory sync fails after rename, reads reflect the replaced file and the store enters a durability-unknown state. Readiness then returns 503 and new writes are blocked until the store is reopened. For retryable increment operations, using a stable Idempotency-Key avoids applying a committed write twice after recovery.

Self-healing demonstration

Delete the managed application Pod and watch the controller recreate it:

pod=$(kubectl get pod -n sunday-system \
  -l sunday.system/etherealpod=sunday-app \
  -o jsonpath='{.items[0].metadata.name}')
kubectl delete pod -n sunday-system "$pod"
kubectl get pods -n sunday-system -w

Write an item before deletion and read it after the replacement becomes Ready to verify that PVC-backed data survived. The automated E2E additionally replays an idempotent write against the replacement Pod and verifies that the amount is not incremented twice.

Container crashes are handled by the enforced restartPolicy: Always; the controller reports the sum of init- and application-container restart counts in .status.restarts. Terminal Pods and accidental duplicate Pods are removed, and a replacement is created only after all previous managed Pods disappear. Changes to spec.template trigger a serial replacement with a brief service interruption. A stuck terminating Pod blocks replacement to protect the writer. Long resource names are shortened for Pod prefixes and labels; the full name is retained in the sunday.system/etherealpod-name annotation.

Development checks

Local checks require Go 1.26.6 or a compatible newer toolchain, as declared in go.mod, and make. The file store supports Linux and macOS with a filesystem that implements advisory flock locks and atomic rename. Cluster deployment builds Linux Go binaries inside Docker.

make build
make test
make test-race
make test-envtest # real kube-apiserver + etcd API semantics
make lint         # requires golangci-lint v2
make e2e          # requires a deployment created by ./requirements.sh

The fast unit tests use the controller-runtime fake client where appropriate. make test-envtest fills the integration layer with a real Kubernetes API server and etcd, verifying CRD admission and status-subresource semantics. The kind E2E then exercises the complete deployed system.

The E2E check validates CRD admission, writes data with an idempotency key, deletes the application Pod, verifies that the replacement can read the persisted value, replays the same write without duplicating it, rolls out a template change, checks that the zero restart count is explicit, then starts a deliberately crashing EtherealPod and verifies that RESTARTS increases.

GitHub Actions runs the same acceptance check on an ubuntu-latest runner. It builds both Docker images, creates a disposable kind cluster, deploys the system, runs the live acceptance scenario, and prints cluster diagnostics if anything fails.

For a short reviewer-facing walkthrough, including discussion points, follow docs/DEMO.md.

Each GitHub Actions acceptance run also records a real terminal demonstration. Open a successful Verify Sunday System run, download the sunday-system-demo artifact, and open sunday-demo.html locally. It replays the actual API responses, Pod replacement, template rollout, and restart count with Play/Pause and speed controls; it needs no external scripts or network. The artifact also contains an asciicast v2 recording and a plain-text transcript.

Automated security

CodeQL analyzes Go changes and runs weekly with the extended security query suite. govulncheck remains part of the required Go quality gate. Dependabot checks Go modules and GitHub Actions weekly and opens update PRs; updates are not auto-merged. The protected main branch requires both Go quality and the Docker/Kubernetes acceptance check, including for administrators.

Requirements traceability

Assignment requirement Implementation
Exactly one active Pod per EtherealPod Owner-reference discovery, indexed lookup, termination waits, and deterministic duplicate cleanup
Recovery after a crash Enforced restartPolicy: Always; terminal Pods are replaced
Recovery after Pod deletion Owned-Pod watch triggers creation of a replacement
NAME AGE RESTARTS in kubectl get eps CRD printer columns and status restart aggregation
Sunday data never disappears with its Pod Atomic versioned persistence on a PersistentVolumeClaim
Safe retries for increments Durable Idempotency-Key record committed atomically with the write
GET, POST, DELETE API Typed routes, input validation, persistence, and HTTP tests

Assumptions

  • “Exactly one running Pod” is implemented as exactly one non-terminal managed Pod during startup, because Kubernetes cannot make a replacement Running instantaneously. A Running Pod is preferred when duplicates are observed.
  • The restart column reflects the current managed Pod, matching the wording of the assignment; deliberate Pod replacement starts a new Kubernetes restart counter.
  • One SundayApp replica is managed by each EtherealPod. RWO restricts volume attachment by node, so it does not itself prevent concurrent writers. The store holds an exclusive sidecar file lock for its lifetime; a second process cannot open the same store. The .lock file must not be deleted while the application runs.
  • Corrupt or empty existing data files fail startup. If directory sync fails after rename, reads reflect the new file, but further writes return 500 until the store is reopened. A write without an idempotency key may already have changed data, so its value must be inspected before blindly retrying.
  • Idempotency records currently have no TTL or automatic compaction. That is acceptable for this bounded home assignment; a production implementation would define a retention window and cleanup strategy.
  • Replaying a previously successful idempotent request returns its historical result even if the product was later deleted. The key identifies the original operation, not the current product state.
  • Kubernetes must establish a CRD before resources of its new kind can be decoded. For that reason requirements.sh installs the CRD first and then applies the self-contained workload resources in submission.yaml.

Scope and next steps

I built this as a focused home assignment. The controller converges toward one active Pod; recovery still takes time for reconciliation and scheduling. The JSON store is intended for a single application writer. Serial template rollouts favor storage safety and include downtime; forced Pod deletion and filesystems without reliable locking are outside this guarantee.

For a production version, I would add authentication and rate limiting, idempotency-record retention, backups, and a database when multiple writers or larger datasets are required.

About

Go Kubernetes operator and persistent REST API: Pod self-healing, serial template rollouts, and verified Docker/kind E2E tests.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages