Problem
In streaming mode (failOnQuietTimeout, i.e. HttpTransport.inactivityTimeout()), the socket deadline is per request leg: post() re-arms deadlineEpochMillis at the start of every leg (HttpTransport.java#L350).
But a reconnect does not go through post(). When the connection is gone, WsmanClient.send() calls transport.connect() explicitly before the authentication exchange (WsmanClient.java#L1166), and ensureConnected() caps the TCP connect and the TLS handshake with boundedByDeadline(connectTimeoutMillis) (HttpTransport.java#L318). That cap uses the deadline left over from the previous leg. If that leg ran long (a Receive that waited most of the inactivity timeout before the server dropped the connection, or a round trip that timed out), the leftover deadline is nearly or fully expired, and boundedByDeadline() floors it to 1 ms. The reconnect then fails with a connect timeout that has nothing to do with the server: a spurious TimeoutException for the streaming caller, or a wasted retry when connectRetries is set.
Where it bites today
PR #197 hit it on its Delete-then-Create sequence (a retired shell's Delete timing out, then the Create's reconnect) and works around it by calling configureTimeouts() again between the two requests. That resets the deadline for that one spot only; every other reconnect inside a streaming operation (a Receive, a Send, a Signal following a dropped connection) still runs under the stale deadline.
Proposed fix
Arm the per-leg deadline in connect() the same way post() does, so a reconnect gets the whole inactivity timeout for itself:
void connect() throws IOException {
if (deadlinePerLeg) {
deadlineEpochMillis = Utils.getCurrentTimeMillis() + readTimeoutMillis;
}
ensureConnected();
}
The configureTimeouts() re-call added by #197 can then go.
Poll mode (pollTimeout()) is unaffected: there the deadline is deliberately shared by every leg of one poll.
Test
StreamingApiTest: a command whose Receive the fake server answers late (close to the inactivity timeout) and then drops the connection, followed by a request that must reconnect; before the fix the reconnect fails with a connect timeout, after it the operation completes.
Problem
In streaming mode (
failOnQuietTimeout, i.e.HttpTransport.inactivityTimeout()), the socket deadline is per request leg:post()re-armsdeadlineEpochMillisat the start of every leg (HttpTransport.java#L350).But a reconnect does not go through
post(). When the connection is gone,WsmanClient.send()callstransport.connect()explicitly before the authentication exchange (WsmanClient.java#L1166), andensureConnected()caps the TCP connect and the TLS handshake withboundedByDeadline(connectTimeoutMillis)(HttpTransport.java#L318). That cap uses the deadline left over from the previous leg. If that leg ran long (a Receive that waited most of the inactivity timeout before the server dropped the connection, or a round trip that timed out), the leftover deadline is nearly or fully expired, andboundedByDeadline()floors it to 1 ms. The reconnect then fails with a connect timeout that has nothing to do with the server: a spuriousTimeoutExceptionfor the streaming caller, or a wasted retry whenconnectRetriesis set.Where it bites today
PR #197 hit it on its Delete-then-Create sequence (a retired shell's Delete timing out, then the Create's reconnect) and works around it by calling
configureTimeouts()again between the two requests. That resets the deadline for that one spot only; every other reconnect inside a streaming operation (a Receive, a Send, a Signal following a dropped connection) still runs under the stale deadline.Proposed fix
Arm the per-leg deadline in
connect()the same waypost()does, so a reconnect gets the whole inactivity timeout for itself:The
configureTimeouts()re-call added by #197 can then go.Poll mode (
pollTimeout()) is unaffected: there the deadline is deliberately shared by every leg of one poll.Test
StreamingApiTest: a command whose Receive the fake server answers late (close to the inactivity timeout) and then drops the connection, followed by a request that must reconnect; before the fix the reconnect fails with a connect timeout, after it the operation completes.