Skip to content

Add automatic Claude account hot switching - #1

Open
SauersML wants to merge 10 commits into
ryanirl:mainfrom
SauersML:agent/automatic-hot-switch
Open

SauersML wants to merge 10 commits into
ryanirl:mainfrom
SauersML:agent/automatic-hot-switch

Conversation

@SauersML

@SauersML SauersML commented Jul 25, 2026 •

Copy link
Copy Markdown

Summary

  • add an opt-in hot-switch mode that keeps all running Claude sessions in one main config directory while rotating saved logins
  • discover, identify, label, and deduplicate Claude logins automatically, including a managed snapshot when the active login has no durable saved profile
  • keep automatic switching alive independently of the TUI with a headless cctop autoswitch supervisor
  • automatically switch to the healthy account with the most headroom at 1% remaining, with lock-safe credential syncing and delegated token refresh
  • preserve the original separate-config/session behavior when hot-switch mode is disabled
  • document the credential lifecycle and hot-switch security model

Why this avoids re-logins

An access token lasts ~12-15h, but the refresh token behind it is what keeps an account alive, and it rotates as Claude Code uses it. A valid main credential is copied back to its saved profile before normal swaps and limits polls, so a profile is never reactivated with a superseded refresh token. Only a genuinely dead refresh token still needs /login — nothing local can substitute for that OAuth round trip.

Correctness notes

Bugs found and fixed during review of this branch:

  • Torn switch could destroy a saved login. The main identity was written after the credentials moved, so a failure left main holding the target's token while the identity still named the previous account — and the next sync copied that token over the previous account's saved profile. The identity is now staged first and published by a single atomic os.replace last, so credential and identity can never disagree.

  • Auto-switch could ping-pong or strand usable headroom. Candidates must now have strictly more headroom than the active profile. That rejects a 92% to 95% regression while allowing a fully exhausted 100% profile to hand work to a still-usable 93% profile.

  • Snapshotted profiles leaked into normal mode. Discovery listed ~/.config/cctop/profiles/* unconditionally, but only hot-switch dedup folds them back into the account they came from — so with hot_switch off, every snapshotted account appeared twice in accounts, doctor, settings, config init, and the limits panel.

  • A security(1) subprocess write per limits poll. The active-profile sync rewrote the credential unconditionally; no-op writes are now skipped.

  • Lock wait could spin, and release could raise. An unreapable stale lock dir looped hot for the full 9s timeout, and a non-empty dir raised OSError out of the context manager on release.

  • Autoswitch stopped with a suspended TUI. Rotation now has a headless supervisor suitable for launchd, and Claude binary discovery works under launchd’s minimal PATH.

  • Wrong live-session counts, since in hot-switch mode all sessions run in one dir; and --account / --config-dir opted out of hot-switch in the JSON path but not the TUI.

  • Fresh unused profiles were excluded from auto-switching. Anthropic reports a never-started usage window as 0% with no reset timestamp. The display layer correctly labels that as no usage yet, but candidate selection incorrectly treated the whole account as unavailable. A successful API response with no started windows now ranks at 0% used, so the healthiest login is selected.

  • HTTP 429 hid useful state. The daemon and TUI now share last-good percentages, reset times, and reading age through a metadata-only cache. Locally expired profiles are refreshed before usage polling, so a dead refresh token reports needs re-login instead of being masked as rate-limited. A recent partial cache never hides configured accounts that lack a successful reading.

Also: doctor now reports the login dir it actually read the token/expiry/tier from, cctop switch rejects a stray argument instead of reading it as a profile name, and the "~/.claude is the special profile" rule lives in one place (authctl) rather than being restated at three call sites.

Validation

  • uv run pytest -q (102 passed) on the current branch
  • uv run ruff check src tests, uv run ruff format --check src tests, uv run mypy
  • the three switch/sync regression tests were confirmed to fail against the pre-fix code, not pass vacuously
  • live end-to-end round trip on a real 6-account fleet: switched away and back, verified the main credential matched the reactivated profile's record and the other profile retained its own distinct valid token, with no leftover lock dirs or temp files
  • live autoswitch recovery on that fleet: refreshed a fresh account whose API returned successful empty windows, confirmed the patched selector ranked it at 0%, switched the main identity and credential to it, and observed the same six running Claude processes consume that account

Sauers and others added 2 commits July 25, 2026 10:12
A torn switch could destroy a saved login: the main identity was written
after the credentials moved, so a failure left main holding the target's
token while the identity still named the previous account, and the next
sync copied that token over the previous account's saved profile. Stage
the identity first and publish it with a single atomic replace last, so
credential and identity can never disagree.

Auto-switch could ping-pong forever: the trigger was configurable but the
candidate filter was hardcoded at 99%, so a threshold of 10% would rotate
at 92% into a profile at 95% and re-fire every poll. Candidates now have
to clear the same threshold that triggered the rotation.

Snapshotted login profiles leaked into normal mode: discovery listed them
unconditionally, but only hot-switch dedup folds them back into the
account they came from, so with hot_switch off every snapshotted account
appeared twice in accounts, doctor, settings, config init, and limits.

Also: skip no-op credential writes, which cost a security(1) subprocess
per limits poll; stop the lock wait from spinning on an unreapable stale
dir and from raising out of the context manager on release; count live
sessions from the main dir in hot-switch mode instead of reporting a
profile's leftover registry as its own; report the login dir that doctor
actually read the token, expiry, and tier from; honor --account and
--config-dir in the TUI as the snapshot paths already did; and reject a
stray argument to `cctop switch` rather than reading it as a profile name.

Fold the "~/.claude is the special profile" rule back to one definition in
authctl, and say why a dead refresh token needs /login rather than only
prescribing it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@SauersML
SauersML marked this pull request as ready for review July 25, 2026 15:42
@SauersML

Copy link
Copy Markdown
Author

Fixed the live “Login expired · Please run /login” failure in 0bd9379.

Root cause: the supervisor judged only saved-profile usage. It did not validate the mutable main credential used by running Claude sessions, could sync that rejected credential back over its saved profile, and could select a locally expired target from stale usage data. A 429 from the usage endpoint also left it with no percentage-based recovery path.

The supervisor now:

  • refreshes the live main credential through Claude Code before syncing it;
  • never copies a credential proven dead over the saved profile;
  • immediately restores a ready saved profile when live refresh fails, including while usage reads are 429-limited;
  • rejects locally expired switch targets even if cached usage says they have headroom.

Validation: ruff format --check, ruff check, mypy, and pytest -q (94 passed). I also installed the exact branch build on the affected six-profile fleet, restarted the launchd supervisor, observed a complete 30-second cycle, and confirmed claude auth status --json remained logged in after automatic rotation.

@SauersML

Copy link
Copy Markdown
Author

Two additional live-fleet fixes are now on the branch:

  • 55b381c changes auto-switch selection from “must clear the configured trigger” to strict improvement. This allows a fully exhausted 100%-used login to hand work to a still-usable 93%-used profile, while still rejecting worse targets and preventing ping-pong.
  • ebc3fcd and 221865e add a shared, metadata-only last-good usage cache. The autoswitch daemon and TUI now retain and share percentages, reset times, and reading age during HTTP 429 backoff; a running TUI adopts newer daemon readings before making its own request.

The affected machine is installed from 221865e, its polling interval is restored from 30s to 180s, the launchd supervisor is healthy, and Claude is authenticated on the usable profile. Validation: ruff format --check, ruff check, mypy, and pytest -q (98 passed).

@SauersML

Copy link
Copy Markdown
Author

481bc57 fixes the last false 429 state observed on the live fleet. The two profiles that kept saying “rate limited, retrying” were locally logged out (no usable credential; epoch expiry). Hot-switch polling now checks local expiry and delegates refresh before making a usage request. When the refresh token is dead, cctop reports “needs re-login” and does not waste a usage request. A live source cycle showed cached percentages for all four valid profiles and “needs re-login” for exactly the two dead profiles. The installed daemon is running this commit; validation is green at 99 tests.

@SauersML

Copy link
Copy Markdown
Author

e3a5f66 fixes the two apparently missing Claude accounts. Discovery still contained all six profiles; a recent partial shared cache contained only the four profiles with successful readings, and the TUI incorrectly treated that partial set as fresh. Cache freshness now requires every configured account to be represented, so uncached/logged-out profiles remain visible as “needs re-login.” Installed on the affected machine; formatting, lint, mypy, and pytest are green (100 passed).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant