Skip to content

Stalled SUSPENDING actors have no system-driven recovery #1502

Description

@HavenXia

Split out of #791 (the deferred "janitor / retry budget" item) so #791 can close.

Problem:

If SuspendActor fails after the actor is marked SUSPENDING, the actor stays SUSPENDING until a client calls SuspendActor again. In that state it cannot be resumed.

A running-origin suspend at least has a graceful termination exit: if the worker pod goes away, DeleteWorker marks the actor CRASHED.

A paused-origin is pinned to a node name, and once the node is gone every retry fails with just internal errors.

Xref: #660 #791 #817 #798

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions