Skip to content

llm-api-gateway: graceful shutdown drain is tied to WriteTimeout and cut by the default pod grace period #2162

Description

@FamousDirector

Describe the bug
When an LLM API gateway pod stops (rollout, scale-down, node drain), requests that are still in flight are cut off earlier than intended.

  • On SIGTERM the gateway waits for in-flight requests for at most its WriteTimeout (60s). There is no separate setting for the shutdown drain.
  • The Helm chart does not set terminationGracePeriodSeconds, so Kubernetes uses its 30s default and kills the pod before that 60s drain can finish.

Long LLM responses, especially streams, are the most affected.

Expected behavior
The shutdown drain time is configurable on its own, and the chart's pod termination grace period is longer than it, so the configured drain actually runs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions