Skip to content

Issue #196: Retire the shell when a command's terminate Signal does not go through - #197

Merged
bertysentry merged 8 commits into
mainfrom
196-retire-shell-after-failed-terminate-signal
Oct 9, 2026
Merged

bertysentry merged 8 commits into
mainfrom
196-retire-shell-after-failed-terminate-signal

Conversation

@NassimBtk

@NassimBtk NassimBtk commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #196

Problem

A long-lived WinRMClient polling periodically ends up with every command refused with HTTP 500 (WSManFault 2150859174): the maximum number of concurrent operations for this user has been exceeded. It never recovers until the client is recreated.

The client reuses one remote shell for all its commands. Live tests on Windows Server 2008 R2 and 2022 (see the live test results comment) showed that every command run in a shell holds one of the user's WSMan operations until the shell is deleted, even when its terminate Signal succeeded, and that time does not release them. MaxConcurrentOperationsPerUser is 15 on 2008 R2 and 1500 later, shared by all connections of the user. Only close() deleted the shell.

On top of that, a command whose terminate Signal failed or was skipped stayed in the reused shell.

Fix

The client replaces its shell (retires it, deletes it, creates a new one) in three cases:

When How
Every maxCommandsPerShell commands New WinRMClient.Builder.maxCommandsPerShell(int), default 10 (1 = a shell per command, like winrs)
Windows Server 2008 R2 only: the quota refuses a Command in a shell that already ran commands The shell is deleted and the Command retried once in a new shell (it ran nothing). In a fresh shell the fault is reported
A command could not be terminated cleanly Completion Signal faulted, failed in transit, or skipped (less than 750 ms left in a bounded poll); early-close Signal failed or held past its 1 s hold

The new shell gets the same pinned working directory, environment and profile. Deleting the old one ends any process a previous command left running in it, as close() always did (documented).

The retired shell is deleted right before the next shell is created, or by close(). Its Delete is best effort and never fails the next command: a failure other than a WSMan fault drops the connection (a response rejected for its HTTP status is never decrypted, which would desync NTLM); if it is cancelled before being sent, the shell stays retired; no shell is created after the caller was told the command timed out.

The quota fault is recognized by its WSManFault code 2150859174, which only Windows Server 2008 R2 sends. Measured on real hosts:

Host WSManFault code SOAP subcode
Windows Server 2008 R2 (English) 2150859174 InternalError
Windows Server 2016 (English) none InternalError
Windows Server 2022 (French) none InternalError

From 2012 on, nothing reliable identifies the quota fault (only its translated reason text), so the quota retry runs on 2008 R2 only, and replacing the shell every N commands is the main fix. While one client fills the quota, the user's other connections are refused too, WQL included, which is another reason not to wait for the fault.

Tests

15 new tests against FakeWsmanServer. Each fails when the line of the fix it guards is undone:

  • WinRMClientTest: replacement every N commands (count restarting in each new shell), every 10 by default, quota retry, fault reported on a fresh shell, single retry
  • WinRMClientBuilderTest: maxCommandsPerShell validation
  • WsmanProtocolTest: faulted Signal (the issue's scenario), Signal lost in transit, close() deleting a retired shell, Delete answered with HTTP 503
  • StreamingApiTest: Signal skipped in a tiny poll, early-close Signal failed, held past its hold, or lost in transit
  • WsmanRetryTest: Delete cancelled in a connect-retry pause

mvn verify passes: 347 tests (9 skipped: live tests needing a Windows host), with no Checkstyle, PMD or SpotBugs findings.

Docs

  • commands.md: new "Shell reuse" section; index.md: client options row; preparing-the-host.md: quota row; migrating-from-winrm4j.md: shell reuse and replay statements corrected; CHANGELOG.md: "Fixed" entry.
  • ShellFileCopy comments no longer claim the quota budget recovers over time.

Not in this PR

  • No logging: the library has none anywhere.
  • ShellFileCopy/RemoteFiles quota wait-and-retry still match the numeric code only; the client's own shell replacement now runs before them.
  • A Command request whose response is lost can still leave an unknown command in the shell until its next replacement.

🤖 Generated with Claude Code

A long-lived client polling periodically ended up with every command
refused with WSManFault 2150859174 (maximum number of concurrent
operations for this user exceeded), until the client was recreated.

When the terminate Signal ending a command faulted, failed in transit,
or was skipped for lack of budget in a bounded poll, the client kept
reusing the shell, and the never-terminated command kept holding one of
the user's WSMan operations until the shell was deleted, which only
close() did. The same happened when the Signal stopping a still-running
command failed.

Such a shell is now retired: the next command deletes it and runs in a
fresh shell created with the same working directory, environment and
profile, and close() deletes it when no command follows. The Delete is
deferred rather than sent right away because a completed command's
cleanup must neither outlive the caller's poll budget nor fail it, and
the failed or skipped Signal leaves no budget, and possibly no
connection, for it. The command that follows restores its own timeouts
after the Delete (a timed-out Delete must not leave the streaming
per-leg deadline expired for the Create's reconnect) and checks for
cancellation before creating the new shell.

Closes #196

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@NassimBtk NassimBtk changed the title Retire the shell when a command's terminate Signal does not go through (#196) Issue #196: Retire the shell when a command's terminate Signal does not go through Oct 7, 2026
@NassimBtk

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e7b9e83dd3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/main/java/org/metricshub/winrm/light/WsmanClient.java Outdated
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-09T14:26:05.200967Z 52223dc Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Hooray!

Reviewed commit: e7b9e83dd3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Address Codex review: a Delete answered with an unexpected HTTP status
(e.g. 503) is rejected by send() before its sealed body is decrypted,
while HttpTransport.post() keeps the socket open because the exchange
completed. Swallowing that failure let the Create that follows reuse an
NTLM session whose incoming cipher stream was out of sync, failing the
new command with a checksum mismatch. Any failure of the best-effort
Delete other than a WSMan fault now closes the transport, so the Create
starts on a fresh, re-authenticated connection, as terminateCompleted()
already does for the Signal.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@NassimBtk

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7664221edb

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/main/java/org/metricshub/winrm/light/WsmanClient.java
Address Codex review: with connection retries enabled, the deferred
Delete of a retired shell can be cancelled by the command's wall-clock
deadline during a connect-retry pause, before the request is sent. The
retired shell ID was already cleared, so a later command created a new
shell without retrying the cleanup, leaving the old command's operation
allocated until the server's IdleTimeout. The ID is now restored on that
cancellation path, so the next command or close() still deletes it; the
restored interrupt makes the cancelled command abort before its Create,
so the shell is never retired alongside a new one.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@NassimBtk

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Another round soon, please!

Reviewed commit: d985232685

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/main/java/org/metricshub/winrm/light/WsmanClient.java
Address review: the OperationTimeout fault answering the terminate
Signal of a command closed before completion only says the service did
not finish processing the Signal within its 1 s hold, not that the
command is gone. Keeping the shell bet on the command having been
killed; if it was not, its operation stayed allocated, which is #196 on
this path. The shell is now retired on that fault too: close() still
succeeds, and the next command deletes the shell and creates a new one.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
NassimBtk and others added 2 commits October 9, 2026 11:28
Since #198, HttpTransport.connect() re-arms the streaming deadline on
every explicit reconnect, so the Create that follows a timed-out
retired-shell Delete no longer inherits an expired deadline. Remove the
configureTimeouts() call that worked around it (its comment was no
longer true) and keep the cancellation check. Also restore the blank
line between the #196 and #198 CHANGELOG entries.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@NassimBtk

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: 52223dc150

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@bertysentry

Copy link
Copy Markdown
Contributor

Live test results

Tested on real hosts against origin/main (baseline) and this branch (52223dc): anaxagore (Windows Server 2008 R2, MaxConcurrentOperationsPerUser = 15), tc-win2022 (Windows Server 2022, French, limit 1500) and tc-win2016 (HTTPS, domain account).

The fix works, and nothing regressed

The probe uses the public API: RemoteProcess.waitFor(760 ms) in a loop. When the Done arrives, less than 750 ms of the deadline is left, so the completion Signal is skipped. That is one of the paths this PR covers.

Test Host main this PR
25 commands, completion Signal skipped, one client, 1 s apart 2008 R2 15 OK, then every command fails with WSManFault 2150859174 25/25 OK
Same, 5 commands 2022 5/5 5/5
Same, 5 commands (HTTPS, SENTRY\dev-admin) 2016 — 5/5
close() with a retired shell 2022 — retired shell deleted (its winrshost.exe is gone)
workingDirectory + environment, Signal skipped, 3 commands 2022 — working directory and env var kept in every recreated shell; one shell alive at a time
Early close of a running command (ping -n 30) 2022 shell reused shell reused (no needless retirement)
WinRMLiveTest 2022 — 7 passed, 1 skipped (Kerberos delegation)
CLI exec / wql 2008 R2, 2022 — OK, exit codes passed through

The cost of a retired shell: the next command takes about 100 ms longer for the Delete and the Create (55 → 155 ms on 2022, 40 → 130 ms on 2008 R2).

But the operations also pile up when every Signal succeeds

Same loop, but each command is drained with waitFor(30 s), so every terminate Signal is sent and answered normally:

Host main this PR
2008 R2 15 OK, then every command fails (2150859174) 15 OK, then every command fails
2022 1500 OK, then every command fails 1500 OK, then every command fails
  • The Signals did succeed. On this branch, a failed Signal retires the shell, and the Delete resets the count. Reaching the limit means no Signal failed in 1500 commands.
  • Time does not release them. After 13 commands on 2008 R2, I waited 90 s with the shell still open: only 2 more commands passed.
  • One operation per command, not per Receive. A ping -n 4 command, which needs several Receives, still allowed exactly 15 commands.
  • Deleting the shell releases them. close() freed all 1500 at once on 2022.

So on a real host, every command run in a reused shell holds one WSMan operation until that shell is deleted, terminate Signal or not. A long-lived polling client reaches the limit after 1500 commands on a modern host, and after 15 on 2008 R2. Recreating the client is the only way out. That matches the symptom described in #196, and it happens without any failed Signal. The reproduction in the issue only used FakeWsmanServer, which cannot show this.

Side finding: on tc-win2022 (French), the quota fault is an HTTP 500 with no WSManFault element. getFaultCode() returns null, and getFaultReason() only has the translated text (Le nombre maximal d'opérations simultanées pour cet utilisateur a été dépassé…). Matching on 2150859174 only works on 2008 R2.

Suggestion

This PR is still correct for the paths it covers. But #196 needs the client to replace its shell on its own as well:

  • After N commands: retire the shell every N commands. The quota is shared by all of a user's connections, so this also stops one long-lived client from holding most of the quota and starving the others.
  • On a quota fault: when the Command request on a reused shell gets a quota fault, retire the shell and retry once in a new one. The existing shell-not-found recovery already does this. The check would need the SOAP Subcode (probably wsman:QuotaLimit) rather than the fault code, since 2022 sends none.

Either way, this PR's mechanism (retire now, Delete before the next Create or in close()) could carry it.

NassimBtk added a commit to MetricsHub/metricshub-community that referenced this pull request Oct 9, 2026
Live tests on Windows Server 2008 R2 (MaxConcurrentOperationsPerUser = 15)
showed the 16th command of a cycle failing with WSManFault 2150859174: the
terminate Signal fails on that host after every command, and winrm-java#196
leaks one operation in the shell each time. A fresh client per command deletes
its shell, and with it the leaked operation, after every command; a pooled
command client piled them up over the cycle.

Commands are back to one client each, as on main. WQL queries, the bulk of a
cycle's requests, keep their pooled clients: they never open a remote shell, so
a pooled WQL client leaves nothing on the host. That also removes the need for
the WQL/command pool split and for the CLI shutdown hook.

Pool commands too once winrm-java ships MetricsHub/winrm-java#197.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Address live tests on Windows Server 2008 R2 and 2022: every command run
in a shell holds one of the user's WSMan operations until the shell is
deleted, even when its terminate Signal succeeded; time does not release
them. A long-lived client therefore reached MaxConcurrentOperationsPerUser
(15 on 2008 R2, 1500 later) without any failed Signal, which is #196 too.

- The client replaces its shell every maxCommandsPerShell commands: a new
  WinRMClient.Builder option, 10 by default, plumbed through a new
  LightWinRMService.createInstance overload (the previous one delegates
  with the default). WsmanClient counts the commands run in the current
  shell and retires it before the next command once the count is reached;
  the existing flow deletes it and creates a new one with the same
  working directory, environment and profile.
- A Command the quota refuses in a shell that already ran commands
  retires and deletes that shell, then is retried once in a new one (it
  ran nothing). In a fresh shell the fault is reported.
- The quota fault is recognized by its code 2150859174 or, on hosts that
  send no WSManFault element (a French Windows Server 2022), by its SOAP
  subcode, now exposed as WinRMFaultException.getFaultSubcode().

Docs: new "Shell reuse" section in commands.md (including that deleting
a shell ends processes left running in it), client options row, fault
detail table, host quota row, winrm4j migration guide, CHANGELOG. The
ShellFileCopy quota-retry comments no longer claim the budget recovers
over time.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@bertysentry

Copy link
Copy Markdown
Contributor

Live test results on 0665a09

Replacing the shell every N commands works

Test Host Result
Default (maxCommandsPerShell 10), 25 commands, winrshost.exe listed after each one 2008 R2 New shell at commands 11 and 21, the old one deleted each time, 25/25 OK
maxCommandsPerShell(1), 4 commands 2022 New shell for every command
Completion Signal skipped (waitFor(760 ms)), 4 commands 2022 New shell after each command
WinRMLiveTest 2022 7 passed, 1 skipped (Kerberos delegation)
CLI exec 2008 R2 OK, exit code 3 passed through

The quota retry only works on 2008 R2

Test: client A, built with maxCommandsPerShell(Integer.MAX_VALUE), fills the quota. Then a second client B runs a WQL query and a command, and A runs one more command.

The quota fault, as each host sends it:

Host WSManFault code SOAP subcode
Windows Server 2008 R2 (English) 2150859174 InternalError
Windows Server 2016 (English) none InternalError
Windows Server 2022 (French) none InternalError
  • 2008 R2: the retry works. A's command Issue #15: Update all references of sentrysoftware.github.io #16 deleted its shell and succeeded in a new one (111 ms).
  • 2016 and 2022: the retry never fires. A's command is refused, and the fault goes to the caller. The QuotaLimit subcode is never sent. InternalError is generic, so it cannot identify this fault. From 2012 on, only the translated Reason text does.
  • While A held the quota (2008 R2), client B was refused for both its WQL query and its command. B's own retry cannot help: its shell is fresh and holds nothing. B could be another client using the same account, or an admin's PowerShell session. B worked again once A replaced its shell.

Why the every-N replacement should stay the main fix

  1. From Windows 2012 on, nothing reliable identifies the quota fault.
  2. The quota is per user and shared by all connections. Waiting for the fault means one client first takes the whole quota, and that blocks the others (measured above).
  3. WQL queries are refused too, and the retry only covers the Command request.
  4. It costs one Delete + Create (about 100 ms) every 10 commands, roughly 10 ms per command.

The retry is still worth keeping for 2008 R2. With a quota of 15, two clients of the same user at up to 10 commands each can exceed it, and the fault code is present there to detect it.

Suggestions

  • Drop the QuotaLimit subcode match: it fires on none of the hosts tested.
  • Consider removing WinRMFaultException.getFaultSubcode() from the public API: it was added for that match. Keep it only if it is wanted for its own sake.
  • In the PR description, replace "To verify on a real host" with the table above.

Address live tests on Windows Server 2008 R2, 2016 and 2022: no host
sends a QuotaLimit subcode with the quota fault. 2008 R2 sends WSManFault
code 2150859174 with the generic InternalError subcode; 2016 and 2022
send no WSManFault code at all, only InternalError, which cannot identify
this fault. So:

- The quota fault is recognized by its code only, which means the
  retry-in-a-new-shell runs on 2008 R2 only. Replacing the shell every
  maxCommandsPerShell commands stays the main fix, the only one that
  works on later versions.
- WinRMFaultException.getFaultSubcode(), added for that match, is
  removed (back to main), with its fault detail table row.
- The builder Javadoc, commands.md, preparing-the-host.md and the
  CHANGELOG scope the quota retry to Windows Server 2008 R2.
- Tests: the subcode test and its made-up fixture are gone; the quota
  tests use the fault as 2008 R2 sends it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@bertysentry
bertysentry merged commit 07ba361 into main Oct 9, 2026
5 checks passed
@bertysentry
bertysentry deleted the 196-retire-shell-after-failed-terminate-signal branch October 9, 2026 18:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Long-lived  WinRMClient keeps failing with WSManFault 2150859174 after a terminate Signal fails

2 participants