Skip to content

Latest commit

 

History

13 Commits

Folders and files

Repository files navigation

GitOps delivery for a Kubernetes data platform

A FluxCD monorepo that delivers a full data platform — Kafka with Debezium CDC, StackGres PostgreSQL, ClickHouse, MinIO, Airflow, Trino, Superset, and a Prometheus/Loki observability stack — to two environments from a single branch, with the operational surface (secret flow, reconcile ordering, drift detection, upgrade policy, network isolation) designed and enforced rather than assumed.

About this project

This is a portfolio piece and it is built to be read as one. The distinction I care about: plenty of GitOps repos demonstrate that a directory tree can be laid out correctly. This one tries to demonstrate that the delivery mechanism itself has been thought through — what happens when a reconcile fails, when someone edits prod by hand, when a chart resolves to a version nobody reviewed, and how you would know.

Three things this README leads with rather than buries:

  1. Every decision is enforced, not just documented. Each ADR has a corresponding check in CI. See What CI proves.
  2. The failures found along the way are written down, including the ones in earlier versions of this repo's own documentation. See What was wrong and how it was caught.
  3. Named limitations, not implied completeness. See Scope and known gaps.

If you have two minutes: read What CI proves and Scope and known gaps. If you're doing technical diligence: make validate runs the entire CI pipeline locally — kustomize build, schema validation, policy tests, and full chart rendering with assertions against the rendered objects.

Competencies this demonstrates

Area Where
GitOps delivery design (base/overlay, layered reconcile ordering, dependency health gating) ADR-0001, Delivery layers
Secret management without secrets in git (ESO + Workload Identity) ADR-0003
Supply-chain determinism and upgrade policy ADR-0005
Policy-as-code enforcing architectural decisions ADR-0007, policy/gitops.rego
Workload isolation with real resource governance ADR-0004
CDC pipeline design (logical decoding, slot management, connector RBAC) base/data-platform/debezium/
Network segmentation derived from the architecture diagram network-policies.yaml
Cost-aware environment strategy with the tradeoff stated ADR-0002
Honest scope framing — naming what is not production-ready Scope and known gaps

Contents: What CI proves · What was wrong · Repo layout · Architecture · Running it · Secret inventory · Scope and known gaps


What CI proves

kustomize build proves a manifest is valid YAML. That is a low bar, and clearing it is not the same as the config being correct. The checks in .github/workflows/validate.yaml make specific claims:

Check What it catches
kustomize build on all 10 overlays Broken references, missing resources
kubeconform with the community CRD catalog Invalid Flux, Strimzi, cert-manager, ESO resources
conftest against policy/gitops.rego Any change that contradicts an ADR
helm template every HelmRelease Values the chart does not understand
Assertions on rendered objects Overlay intent that never reaches the cluster

The last two are the ones that matter, because Helm ignores unknown values silently. A patch can target a key that does not exist, render perfectly, pass every YAML check, commit cleanly — and do nothing. Only rendering through the real chart and inspecting the resulting Deployment closes that gap.

That is not hypothetical; it is the bug this repo shipped with. Reverting the prod Trino patch to its original form reproduces the failure:

== prod: Trino scaling reaches the Deployment ==
  [FAIL] trino worker replicas == 3 (rendered: 1)

Rendering also pins a target Kubernetes version (KUBE_VERSION = 1.31.0 in scripts/render-helmreleases.py), which is how a chart pinned outside its supported range was caught before deploy rather than as a stuck reconcile.


What was wrong and how it was caught

A portfolio that only shows the finished state hides the part that is actually worth reading. These were all real defects in this repository, found by reviewing it against the charts it deploys:

Defect Consequence Now caught by
Prod Trino patched coordinator.replicas / worker.replicas — neither key exists in the chart Git recorded a scaled prod cluster; Helm ran the default Rendered-object assertion
Trino catalogs used ${ENV:...} with no envFrom and no Secret All three catalogs fail to initialise; coordinator does not start Chart render
base/ hardcoded postgres-dev.databases.svc Prod Airflow and prod Trino pointed at the dev database ADR-0006 + cross-overlay leak assertion
Airflow shipped its PostgreSQL subchart enabled and an external metadataConnection, with no password Two databases deployed, neither usable Reviewed against chart defaults
Superset secret written to configOverrides.secret configOverrides values are Python source — a syntax error at import, and SUPERSET_SECRET_KEY never set Reviewed against chart contract
StackGres HelmRepository URL served no index 404 on every reconcile Manual probe
StackGres 1.13.x declares kubeVersion: … - 1.31.x Cannot install on a current cluster Chart render at KUBE_VERSION
Fluent Bit hardcoded cluster=dev in base Prod logs labelled as dev postBuild substitution
dependsOn without wait: true Layers started against unready operators, contradicting the documented ordering Rego policy
13 charts pinned to ranges (6.x, 61.x) Cluster contents could change with no commit Rego policy (ADR-0005)
No resource requests or limits anywhere ADR-0004's noisy-neighbour rationale was unmet within each pool Rendered-workload assertion
Prometheus and Loki on emptyDir All metric and log history lost on pod restart Persistence configured
Architecture diagram drew a CDC path with no KafkaConnector Documented pipeline did not exist Built (see below)
ADR-0003 claimed Workload Identity was wired It was not annotated anywhere Wired; ADR corrected

The CDC path is now real: wal_level=logical via SGPostgresConfig, a KafkaConnector with a bounded table.include.list, declared KafkaTopic resources with explicit retention, and the RoleBinding that lets Connect resolve ${secrets:...} against the API server.


Repo layout

clusters/                          # Flux entrypoints, one directory per cluster
  dev/  prod/
    flux-system/                   # bootstrap output (GitRepository + sync Kustomization)
    cluster-vars.yaml              # postBuild substitution values (non-secret)
    cluster-infra.yaml             # ─┐
    monitoring.yaml                #  │ Flux Kustomizations — reconcile order
    databases.yaml                 #  │ enforced by dependsOn + wait
    data-platform.yaml             # ─┘

infrastructure/
  base/                            # shared definitions — the single source for each component
    cluster-infra/                 #   cert-manager (+ClusterIssuers), ingress-nginx,
                                   #   external-secrets, Flux Alerts
    databases/                     #   StackGres (+CDC postgres config, Airflow bootstrap),
                                   #   ClickHouse, MinIO, NetworkPolicies
    data-platform/                 #   Strimzi, Debezium (Connect/Connector/Topics/RBAC),
                                   #   Airflow, Trino, Superset, NetworkPolicies
    monitoring/                    #   kube-prometheus-stack, Loki, Fluent Bit
  dev/   prod/                     # overlays: only what genuinely differs
                                   #   (instance counts, volume sizes, replication
                                   #    factors, retention, PDBs, vault URL)

policy/gitops.rego                 # ADR enforcement, run by conftest in CI
scripts/
  render-helmreleases.py           # renders every HelmRelease through its real chart
  assert-overlay.py                # asserts overlay intent survives into rendered objects
docs/
  CONTEXT.md                       # architecture, delivery layers, CDC dependencies, scope
  adr/                             # 7 decision records

The overlays are deliberately thin. Anything identical across environments — including every ExternalSecret — lives in base/, because the only thing that differs is the vault behind the azure-keyvault ClusterSecretStore.


Running it

Prerequisites

  • A Kubernetes cluster on 1.31 (the version CI renders against). Dev is four Hetzner CPX31 nodes running k3s, ≈€56/month — see ADR-0002 for why four and not one.
  • An Azure Key Vault and a managed identity federated to the ESO service account.
  • flux, helm, kubectl, python3 with pyyaml.

Label and taint the nodes

Every base HelmRelease carries a nodeSelector. Without these labels nothing schedules:

kubectl label node <node-1> nodepool=infra
kubectl label node <node-2> nodepool=kafka
kubectl label node <node-3> nodepool=data-storage
kubectl label node <node-4> nodepool=data-compute

kubectl taint node <node-1> nodepool=infra:NoSchedule
kubectl taint node <node-2> nodepool=kafka:NoSchedule
kubectl taint node <node-3> nodepool=data-storage:NoSchedule
kubectl taint node <node-4> nodepool=data-compute:NoSchedule

Populate the vault, then bootstrap

Fill Key Vault with the secret inventory below, set the client ID in clusters/dev/cluster-vars.yaml, then:

flux bootstrap github \
  --owner=salbifaza \
  --repository=data-platform-gitops \
  --branch=main \
  --path=clusters/dev \
  --personal

Flux reconciles cluster-infra → monitoring → databases → data-platform, waiting for each layer to be genuinely healthy before starting the next.

flux get kustomizations --watch

Validate locally before pushing

make validate      # everything CI runs
make build         # kustomize build all overlays
make policy        # conftest against policy/
make render        # helm template + rendered-object assertions

Secret inventory

No secret material is in this repository, encrypted or otherwise — conftest fails the build on any committed Secret. These keys must exist in the Key Vault named by the environment's ClusterSecretStore:

Key Vault key Consumed by
clickhouse-password ClickHouse auth.password
minio-root-user, minio-root-password MinIO root credentials
airflow-fernet-key, airflow-webserver-secret-key Airflow chart
airflow-metadata-password Airflow metadataConnection, and the SGScript that creates the role
superset-secret-key Superset SUPERSET_SECRET_KEY
trino-pg-user, trino-pg-password Trino PostgreSQL catalog
trino-ch-user, trino-ch-password Trino ClickHouse catalog
debezium-db-user, debezium-db-password Debezium KafkaConnector
slack-webhook-url Flux notification Provider

airflow-metadata-password is deliberately used twice: ESO templates it into the bootstrap SQL that creates the Postgres role, and injects it into the Airflow chart that logs in with it. One credential, one place, two consumers.


Scope and known gaps

Stated plainly, because a demo that hides its scope is a worse signal than one that names it.

  • Verification is static, not runtime — for both environments. Everything CI proves is established without a cluster: overlays render, charts render, rendered objects carry the intended replicas and limits, resources match their API schemas, no change contradicts an ADR. Nothing here demonstrates that pods schedule, that the CDC connector reaches RUNNING, or that a query returns rows. clusters/*/flux-system/gotk-components.yaml is still the bootstrap placeholder, so no flux bootstrap result is recorded in git for either environment.

    This is a real gap and the most valuable next step. The sibling lakehouse-iceberg-batch repo closes the equivalent gap with a make smoke that stands the stack up in CI and asserts against live data; the Kubernetes analogue is a kind cluster job running flux install plus a reduced-footprint overlay. ADR-0002 records the cost reasoning for prod specifically.

  • Chart versions are pinned but stale. Most sit at their mid-2024 minor. This was a deliberate sequencing choice: pin exactly first so the repo is deterministic, then let Renovate bring them current as reviewed PRs. Upgrading ESO past 0.14 in particular means migrating every ExternalSecret from v1beta1 to v1, which is a change that deserves its own PR.

  • The Bitnami ClickHouse chart is a supply-chain risk. Bitnami restructured its public catalog in 2025 and moved older images to a bitnamilegacy registry. The chart version pinned here is verified present in the index, but image pulls may need global.imageRegistry set. A production system would mirror these images rather than depend on a vendor's free tier.

  • Trino does not query MinIO. The Iceberg/object-storage layer is the sibling lakehouse-iceberg-batch project's subject. MinIO is deployed here as the storage layer those tables would land in; wiring them together means running Polaris in-cluster. Named future work — see Scope boundaries.

  • No Airflow DAGs in this repo. The chart is configured for git-sync from a separate DAG repository, which is the right separation (platform delivery and pipeline code have different review cadences) but does mean the orchestration layer here is a running scheduler with nothing scheduled.

  • NetworkPolicies restrict ingress only. Egress is deliberately not denied. Doing so without a matching DNS allow rule breaks name resolution for every pod, which is the most common way a NetworkPolicy rollout causes an outage. A production rollout would add egress rules incrementally, with DNS allowed first.

  • PSA enforcement is baseline, not restricted. monitoring runs at privileged because Fluent Bit and node-exporter need hostPath and hostNetwork. Every namespace audits and warns at restricted, so the gap to tightening is visible rather than guessed at.

  • Single-region, single-cluster per environment. No multi-cluster fleet management, no cross-region failover, no backup/restore story for the stateful workloads. StackGres supports scheduled backups to object storage and it is not configured here — that is the most significant missing piece for anything holding real data.

  • Alerting covers delivery, not the platform. Flux reconcile failures and drift reach Slack. There are no Prometheus alerting rules for consumer lag, broker under-replication, or query failure rates — the metrics are collected and nothing pages on them.


License

MIT — see LICENSE.

About

FluxCD data platform portfolio — GitOps delivery of Kafka, ClickHouse, Airflow, Trino, Superset on k3s

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages