Describe the bug
When an LLM API gateway pod stops (rollout, scale-down, node drain), requests that are still in flight are cut off earlier than intended.
- On SIGTERM the gateway waits for in-flight requests for at most its
WriteTimeout (60s). There is no separate setting for the shutdown drain.
- The Helm chart does not set
terminationGracePeriodSeconds, so Kubernetes uses its 30s default and kills the pod before that 60s drain can finish.
Long LLM responses, especially streams, are the most affected.
Expected behavior
The shutdown drain time is configurable on its own, and the chart's pod termination grace period is longer than it, so the configured drain actually runs.
Describe the bug
When an LLM API gateway pod stops (rollout, scale-down, node drain), requests that are still in flight are cut off earlier than intended.
WriteTimeout(60s). There is no separate setting for the shutdown drain.terminationGracePeriodSeconds, so Kubernetes uses its 30s default and kills the pod before that 60s drain can finish.Long LLM responses, especially streams, are the most affected.
Expected behavior
The shutdown drain time is configurable on its own, and the chart's pod termination grace period is longer than it, so the configured drain actually runs.