| title | Conversion Webhook Failed | |||||
|---|---|---|---|---|---|---|
| error_message | conversion webhook for <group>/<version>, Kind=<kind> failed | |||||
| category | api-server | |||||
| severity | High | |||||
| recovery_time | 10–40 min | |||||
| k8s_versions | 1.20+ | |||||
| tags |
|
|||||
| related |
|
Severity: High · Typical recovery time: 10–40 min · Affected versions: 1.20+
Internal error occurred: conversion webhook for example.com/v1beta1, Kind=Widget
failed: Post "https://conv-svc.ns.svc:443/convert?timeout=30s":
dial tcp 10.96.0.7:443: connect: connection refused
CRDs with multiple stored/served versions can use a conversion webhook to
translate objects between versions on read/write. When that webhook is
unreachable or errors, the apiserver cannot serve or persist affected custom
resources — kubectl get/apply on the CRD fails, and controllers reconciling
those resources stall. Because conversion runs on every access (including LISTs),
a broken conversion webhook can wedge an entire operator.
Applies to 1.20+ where apiextensions.k8s.io/v1 CRDs use
spec.conversion.strategy: Webhook. The conversion review API
(apiextensions.k8s.io/v1) is stable across these versions.
- Conversion webhook backend (often the operator) is down or crash-looping
- Service has no ready endpoints (selector/label mismatch)
- Stale or wrong
conversion.webhook.clientConfig.caBundle(TLS rejection) - Webhook timeout exceeded under large LISTs
- Webhook returns malformed/incomplete
ConversionReviewresponses
flowchart TD
A[conversion webhook failed] --> B{Backend pod healthy?}
B -- No --> C[Restart/scale operator]
B -- Yes --> D{Service has endpoints?}
D -- No --> E[Fix selector/labels]
D -- Yes --> F{TLS/caBundle valid?}
F -- No --> G[Refresh caBundle]
F -- Yes --> H[Check webhook logs for conversion errors]
Confirm which CRD's conversion is failing, that its strategy is Webhook, and
whether the backing Service is reachable with a valid CA bundle.
kubectl get crd <crd> -o jsonpath='{.spec.conversion.strategy}'
kubectl get crd <crd> -o yaml | grep -A15 conversion:
kubectl get endpoints -n <ns> <conversion-svc>
kubectl get pods -n <ns> -l <operator-selector>
kubectl logs -n <ns> deploy/<operator> --tail=80
kubectl get apiservices | grep <group>
kubectl get events -A --sort-by=.lastTimestamp | grep -i conversion$ kubectl get widgets -A
Error from server: conversion webhook for example.com/v1beta1, Kind=Widget failed:
... connect: connection refused
$ kubectl get endpoints -n ns conv-svc
NAME ENDPOINTS AGE
conv-svc <none> 3d
- Restore the operator/webhook backend that serves
/convert. - Fix the Service selector so endpoints populate.
- Refresh the CRD's
conversion...caBundleto the current signing CA (cert-manager CA injection commonly automates this). - Increase the conversion
timeoutSecondsif large LISTs time out.
- Check operator logs for conversion errors and the Service endpoints.
- Roll the operator Deployment if it is the webhook backend.
- Disruptive (last resort): if the conversion webhook is permanently broken
and you must regain access, switching the CRD
conversion.strategytoNonestops conversion. Blast radius: objects are served only in their stored version and may be returned incorrectly — data-integrity risk; use only with a recovery plan and revert once the webhook is fixed.
kubectl get <cr> returns objects without error and the operator resumes
reconciliation cleanly.
Run the operator HA with a PDB, automate caBundle injection, set generous conversion timeouts, test version conversions in CI, and alert on the conversion webhook Service's endpoint readiness.