Skip to content

Latest commit

 

History

History

README.md

Error library

300+ Kubernetes production errors, one page each, organized by subsystem. Use the offline search tool to find an error fast:

python tools/search.py CrashLoopBackOff
python tools/search.py --category storage --severity Critical

Subsystems

Directory Covers
pods/ Pod lifecycle: CrashLoopBackOff, ImagePullBackOff, OOMKilled, Pending, container config
deployments/ Rollouts, replicas, progress deadlines
daemonsets/ Per-node scheduling and rollout issues
statefulsets/ Ordered pods, stable identity, volume templates
jobs/ Batch jobs, backoff limits, completions
cronjobs/ Schedules, concurrency, missed runs
nodes/ NodeNotReady, pressure conditions, evictions
networking/ CNI, pod-to-pod, DNS, NetworkPolicy, kube-proxy
ingress/ Ingress controllers, 502/503/404, TLS
services/ ClusterIP/NodePort/LoadBalancer, endpoints
storage/ CSI, mounts, attach/detach
persistent-volumes/ PV lifecycle, reclaim, binding
persistent-volume-claims/ PVC provisioning and binding
rbac/ Forbidden, roles, bindings, service accounts
security/ Pod Security, secrets, TLS, certificates
helm/ Releases, upgrades, hooks, rollbacks
cert-manager/ Certificates, issuers, ACME challenges
monitoring/ metrics-server, Prometheus, scraping
autoscaling/ HPA, VPA, cluster-autoscaler
api-server/ Availability, throttling, timeouts, webhooks
etcd/ Quorum, space, latency, compaction
scheduler/ FailedScheduling, affinity, taints, topology
controller-manager/ Controllers, leader election
admission/ Validating/mutating webhooks, policy denials
kubelet/ Node agent, PLEG, image GC, cgroups
container-runtime/ containerd/CRI-O, image pulls, sandbox

Page format

See _TEMPLATE.md. Every page targets one error and one primary search phrase, with a full diagnostic flow, read-only kubectl commands, fixes, recovery, validation, and prevention.

Spotted a gap? Request a new error page or open a PR — see CONTRIBUTING.