A FluxCD monorepo that delivers a full data platform — Kafka with Debezium CDC, StackGres PostgreSQL, ClickHouse, MinIO, Airflow, Trino, Superset, and a Prometheus/Loki observability stack — to two environments from a single branch, with the operational surface (secret flow, reconcile ordering, drift detection, upgrade policy, network isolation) designed and enforced rather than assumed.
This is a portfolio piece and it is built to be read as one. The distinction I care about: plenty of GitOps repos demonstrate that a directory tree can be laid out correctly. This one tries to demonstrate that the delivery mechanism itself has been thought through — what happens when a reconcile fails, when someone edits prod by hand, when a chart resolves to a version nobody reviewed, and how you would know.
Three things this README leads with rather than buries:
- Every decision is enforced, not just documented. Each ADR has a corresponding check in CI. See What CI proves.
- The failures found along the way are written down, including the ones in earlier versions of this repo's own documentation. See What was wrong and how it was caught.
- Named limitations, not implied completeness. See Scope and known gaps.
If you have two minutes: read What CI proves and
Scope and known gaps.
If you're doing technical diligence: make validate runs the entire CI
pipeline locally — kustomize build, schema validation, policy tests, and full
chart rendering with assertions against the rendered objects.
| Area | Where |
|---|---|
| GitOps delivery design (base/overlay, layered reconcile ordering, dependency health gating) | ADR-0001, Delivery layers |
| Secret management without secrets in git (ESO + Workload Identity) | ADR-0003 |
| Supply-chain determinism and upgrade policy | ADR-0005 |
| Policy-as-code enforcing architectural decisions | ADR-0007, policy/gitops.rego |
| Workload isolation with real resource governance | ADR-0004 |
| CDC pipeline design (logical decoding, slot management, connector RBAC) | base/data-platform/debezium/ |
| Network segmentation derived from the architecture diagram | network-policies.yaml |
| Cost-aware environment strategy with the tradeoff stated | ADR-0002 |
| Honest scope framing — naming what is not production-ready | Scope and known gaps |
Contents: What CI proves · What was wrong · Repo layout · Architecture · Running it · Secret inventory · Scope and known gaps
kustomize build proves a manifest is valid YAML. That is a low bar, and
clearing it is not the same as the config being correct. The checks in
.github/workflows/validate.yaml make
specific claims:
| Check | What it catches |
|---|---|
kustomize build on all 10 overlays |
Broken references, missing resources |
kubeconform with the community CRD catalog |
Invalid Flux, Strimzi, cert-manager, ESO resources |
conftest against policy/gitops.rego |
Any change that contradicts an ADR |
helm template every HelmRelease |
Values the chart does not understand |
| Assertions on rendered objects | Overlay intent that never reaches the cluster |
The last two are the ones that matter, because Helm ignores unknown values
silently. A patch can target a key that does not exist, render perfectly,
pass every YAML check, commit cleanly — and do nothing. Only rendering through
the real chart and inspecting the resulting Deployment closes that gap.
That is not hypothetical; it is the bug this repo shipped with. Reverting the prod Trino patch to its original form reproduces the failure:
== prod: Trino scaling reaches the Deployment ==
[FAIL] trino worker replicas == 3 (rendered: 1)
Rendering also pins a target Kubernetes version (KUBE_VERSION = 1.31.0 in
scripts/render-helmreleases.py), which is
how a chart pinned outside its supported range was caught before deploy rather
than as a stuck reconcile.
A portfolio that only shows the finished state hides the part that is actually worth reading. These were all real defects in this repository, found by reviewing it against the charts it deploys:
| Defect | Consequence | Now caught by |
|---|---|---|
Prod Trino patched coordinator.replicas / worker.replicas — neither key exists in the chart |
Git recorded a scaled prod cluster; Helm ran the default | Rendered-object assertion |
Trino catalogs used ${ENV:...} with no envFrom and no Secret |
All three catalogs fail to initialise; coordinator does not start | Chart render |
base/ hardcoded postgres-dev.databases.svc |
Prod Airflow and prod Trino pointed at the dev database | ADR-0006 + cross-overlay leak assertion |
Airflow shipped its PostgreSQL subchart enabled and an external metadataConnection, with no password |
Two databases deployed, neither usable | Reviewed against chart defaults |
Superset secret written to configOverrides.secret |
configOverrides values are Python source — a syntax error at import, and SUPERSET_SECRET_KEY never set |
Reviewed against chart contract |
StackGres HelmRepository URL served no index |
404 on every reconcile | Manual probe |
StackGres 1.13.x declares kubeVersion: … - 1.31.x |
Cannot install on a current cluster | Chart render at KUBE_VERSION |
Fluent Bit hardcoded cluster=dev in base |
Prod logs labelled as dev | postBuild substitution |
dependsOn without wait: true |
Layers started against unready operators, contradicting the documented ordering | Rego policy |
13 charts pinned to ranges (6.x, 61.x) |
Cluster contents could change with no commit | Rego policy (ADR-0005) |
| No resource requests or limits anywhere | ADR-0004's noisy-neighbour rationale was unmet within each pool | Rendered-workload assertion |
Prometheus and Loki on emptyDir |
All metric and log history lost on pod restart | Persistence configured |
Architecture diagram drew a CDC path with no KafkaConnector |
Documented pipeline did not exist | Built (see below) |
| ADR-0003 claimed Workload Identity was wired | It was not annotated anywhere | Wired; ADR corrected |
The CDC path is now real: wal_level=logical via SGPostgresConfig, a
KafkaConnector with a bounded table.include.list, declared KafkaTopic
resources with explicit retention, and the RoleBinding that lets Connect
resolve ${secrets:...} against the API server.
clusters/ # Flux entrypoints, one directory per cluster
dev/ prod/
flux-system/ # bootstrap output (GitRepository + sync Kustomization)
cluster-vars.yaml # postBuild substitution values (non-secret)
cluster-infra.yaml # ─┐
monitoring.yaml # │ Flux Kustomizations — reconcile order
databases.yaml # │ enforced by dependsOn + wait
data-platform.yaml # ─┘
infrastructure/
base/ # shared definitions — the single source for each component
cluster-infra/ # cert-manager (+ClusterIssuers), ingress-nginx,
# external-secrets, Flux Alerts
databases/ # StackGres (+CDC postgres config, Airflow bootstrap),
# ClickHouse, MinIO, NetworkPolicies
data-platform/ # Strimzi, Debezium (Connect/Connector/Topics/RBAC),
# Airflow, Trino, Superset, NetworkPolicies
monitoring/ # kube-prometheus-stack, Loki, Fluent Bit
dev/ prod/ # overlays: only what genuinely differs
# (instance counts, volume sizes, replication
# factors, retention, PDBs, vault URL)
policy/gitops.rego # ADR enforcement, run by conftest in CI
scripts/
render-helmreleases.py # renders every HelmRelease through its real chart
assert-overlay.py # asserts overlay intent survives into rendered objects
docs/
CONTEXT.md # architecture, delivery layers, CDC dependencies, scope
adr/ # 7 decision records
The overlays are deliberately thin. Anything identical across environments —
including every ExternalSecret — lives in base/, because the only thing
that differs is the vault behind the azure-keyvault ClusterSecretStore.
- A Kubernetes cluster on 1.31 (the version CI renders against). Dev is four Hetzner CPX31 nodes running k3s, ≈€56/month — see ADR-0002 for why four and not one.
- An Azure Key Vault and a managed identity federated to the ESO service account.
flux,helm,kubectl,python3withpyyaml.
Every base HelmRelease carries a nodeSelector. Without these labels
nothing schedules:
kubectl label node <node-1> nodepool=infra
kubectl label node <node-2> nodepool=kafka
kubectl label node <node-3> nodepool=data-storage
kubectl label node <node-4> nodepool=data-compute
kubectl taint node <node-1> nodepool=infra:NoSchedule
kubectl taint node <node-2> nodepool=kafka:NoSchedule
kubectl taint node <node-3> nodepool=data-storage:NoSchedule
kubectl taint node <node-4> nodepool=data-compute:NoScheduleFill Key Vault with the secret inventory below, set the
client ID in clusters/dev/cluster-vars.yaml, then:
flux bootstrap github \
--owner=salbifaza \
--repository=data-platform-gitops \
--branch=main \
--path=clusters/dev \
--personalFlux reconciles cluster-infra → monitoring → databases → data-platform,
waiting for each layer to be genuinely healthy before starting the next.
flux get kustomizations --watchmake validate # everything CI runs
make build # kustomize build all overlays
make policy # conftest against policy/
make render # helm template + rendered-object assertionsNo secret material is in this repository, encrypted or otherwise — conftest
fails the build on any committed Secret. These keys must exist in the Key
Vault named by the environment's ClusterSecretStore:
| Key Vault key | Consumed by |
|---|---|
clickhouse-password |
ClickHouse auth.password |
minio-root-user, minio-root-password |
MinIO root credentials |
airflow-fernet-key, airflow-webserver-secret-key |
Airflow chart |
airflow-metadata-password |
Airflow metadataConnection, and the SGScript that creates the role |
superset-secret-key |
Superset SUPERSET_SECRET_KEY |
trino-pg-user, trino-pg-password |
Trino PostgreSQL catalog |
trino-ch-user, trino-ch-password |
Trino ClickHouse catalog |
debezium-db-user, debezium-db-password |
Debezium KafkaConnector |
slack-webhook-url |
Flux notification Provider |
airflow-metadata-password is deliberately used twice: ESO templates it into
the bootstrap SQL that creates the Postgres role, and injects it into the
Airflow chart that logs in with it. One credential, one place, two consumers.
Stated plainly, because a demo that hides its scope is a worse signal than one that names it.
-
Verification is static, not runtime — for both environments. Everything CI proves is established without a cluster: overlays render, charts render, rendered objects carry the intended replicas and limits, resources match their API schemas, no change contradicts an ADR. Nothing here demonstrates that pods schedule, that the CDC connector reaches
RUNNING, or that a query returns rows.clusters/*/flux-system/gotk-components.yamlis still the bootstrap placeholder, so noflux bootstrapresult is recorded in git for either environment.This is a real gap and the most valuable next step. The sibling
lakehouse-iceberg-batchrepo closes the equivalent gap with amake smokethat stands the stack up in CI and asserts against live data; the Kubernetes analogue is akindcluster job runningflux installplus a reduced-footprint overlay. ADR-0002 records the cost reasoning for prod specifically. -
Chart versions are pinned but stale. Most sit at their mid-2024 minor. This was a deliberate sequencing choice: pin exactly first so the repo is deterministic, then let Renovate bring them current as reviewed PRs. Upgrading ESO past 0.14 in particular means migrating every
ExternalSecretfromv1beta1tov1, which is a change that deserves its own PR. -
The Bitnami ClickHouse chart is a supply-chain risk. Bitnami restructured its public catalog in 2025 and moved older images to a
bitnamilegacyregistry. The chart version pinned here is verified present in the index, but image pulls may needglobal.imageRegistryset. A production system would mirror these images rather than depend on a vendor's free tier. -
Trino does not query MinIO. The Iceberg/object-storage layer is the sibling
lakehouse-iceberg-batchproject's subject. MinIO is deployed here as the storage layer those tables would land in; wiring them together means running Polaris in-cluster. Named future work — see Scope boundaries. -
No Airflow DAGs in this repo. The chart is configured for git-sync from a separate DAG repository, which is the right separation (platform delivery and pipeline code have different review cadences) but does mean the orchestration layer here is a running scheduler with nothing scheduled.
-
NetworkPolicies restrict ingress only. Egress is deliberately not denied. Doing so without a matching DNS allow rule breaks name resolution for every pod, which is the most common way a NetworkPolicy rollout causes an outage. A production rollout would add egress rules incrementally, with DNS allowed first.
-
PSA enforcement is
baseline, notrestricted.monitoringruns atprivilegedbecause Fluent Bit and node-exporter need hostPath and hostNetwork. Every namespace audits and warns atrestricted, so the gap to tightening is visible rather than guessed at. -
Single-region, single-cluster per environment. No multi-cluster fleet management, no cross-region failover, no backup/restore story for the stateful workloads. StackGres supports scheduled backups to object storage and it is not configured here — that is the most significant missing piece for anything holding real data.
-
Alerting covers delivery, not the platform. Flux reconcile failures and drift reach Slack. There are no Prometheus alerting rules for consumer lag, broker under-replication, or query failure rates — the metrics are collected and nothing pages on them.
MIT — see LICENSE.