Keep redialing a restarted codespace and bring back all its loops - #616
Merged
Merged
Conversation
A codespace that stayed unreachable for four minutes paused every dial until a human selected one of its loops, so a restart that took longer left every loop on it dead. The pause now lets one dial through every five minutes, in the daemon and in an open pane. The dial that reaches the codespace again asks for the reboot probe, so finished loops are restored even with no pane open, and the liveness sweep now also re-ensures the workers of a piloted or armed composite. Co-Authored-By: Claude Opus 5.5 <[email protected]> Signed-off-by: scgopi <[email protected]>
The paused codespace now redials once a minute instead of every five. Plain ssh hosts are swept every 30 seconds, and their finished loops are probed on every sweep rather than only after a pane redials: their dials ride the ControlMaster and spend no quota. Codespaces stay on a minute. Co-Authored-By: Claude Opus 5.5 <[email protected]> Signed-off-by: scgopi <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Loops on a codespace that restarted no longer stay dead until a human selects one. A down codespace is still dialed once a minute after the 4-minute pause. When it answers again, every loop on it comes back, including finished loops and the workers inside a running composite. Plain ssh hosts are swept every 30 seconds, and their finished loops come back with no pane open.
Why
After a codespace restart, loops did not recover on their own. Someone had to create or select a loop to bring them back. Reading the restore path found three gaps:
CodespaceDialBreakerreturned.pausedforever after 240s of failed dials, and only the reconnect marker cleared itRebootProbeGate)ensureUnattendedSessionsAliveswept onlygraph.nodesChanges
CodespaceDialSchedule.slowRetryInterval(60s). Paused, the breaker lets through one dial per interval for the whole codespace, whichever read or ensure asks first. A failed slow dial leaves the outage clock alone. Enter or selecting a loop still reconnects at once.ZmxSessionLauncher.markRedialed). The next sweep then runs the reboot probe and restores finished loops, as if a pane had redialed.pilotCompositestarts.RebootProbeGatealways lets a plain ssh host through, so its finished loops are probed on every sweep. Those dials are multiplexed over the ControlMaster and spend no quota.Cost while a codespace stays down past 4 minutes: about one gh dial per minute from the daemon, plus one per open pane of that codespace. A plain ssh host with finished loops costs one multiplexed probe dial every 30s. A dial to a stopped codespace starts it, so a codespace with running loops is brought back up rather than left stopped.
Tests
RED: xcodebuild test -only-testing CodespaceDialScheduleTests CodespaceSelectionTests RemoteSessionResumeTests RemoteRebootRestoreTests with SSHReconnectLoop.swift and GraphStore.swift from origin/main and the breaker's slow retry and recovery hook removed -> 4 failed of 52 (aPausedCodespaceIsStillDialedOncePerSlowInterval, aCodespaceThatAnswersAfterAnOutageAsksForTheRebootProbe, aPausedPaneRedialsOnItsOwnAfterTheSlowInterval, theLivenessSweepRestartsTheLoopsInsideARunningComposite), TEST FAILED exit 65
GREEN: same focused xcodebuild test on this branch -> 52 tests in 4 suites passed, exit 0
GREEN: after the interval change, the same focused suites -> 54 tests in 4 suites passed, exit 0; its new tests (aPlainSSHHostIsProbedOnEverySweepWithNoPaneOpen, plainSSHHostsAreSweptEveryThirtySecondsAndCodespacesEveryMinute) fail to compile on 91a930c -> ProjectRegistry has no member sweeps
REGRESSION: full xcodebuild test of the graphcode scheme -> 2226 tests in 255 suites, 1 failure (MessageDeliveryTests.aSessionWhoseTaskEndedIsNeverTypedInto), which fails identically 3 of 3 runs on base c1febe4 without this change. graphcoded and graphcode-cli schemes build, make check exit 0.
The
MessageDeliveryTestsfailure is pre-existing on main and unrelated: it is a live-zmx send test, and this PR does not touch the send path.