Skip to content

FR-011 scope-latency: single-run false failure on PR #1380 (read_cache tail-latency outlier, not a regression) #1387

Description

@Dumbris

Summary

PR #1380 (108-a-profile-policy-model) tripped the advisory FR-011 merge-base regression check (.github/workflows/scope-latency.yml) on TestScopeLatency_ReadCache_ScopedVsAdmin:

read_cache: scoped p95 must not exceed admin's by more than FR-011's 20ms budget
(got admin=1.98966ms scoped=48.901918ms gap=46.912258ms)

Investigation

  • The workflow is advisory (continue-on-error: true) and is not in main's required status checks, so it did not block the merge.
  • PR 108-a's actual diff (internal/server/mcp_visibility.go, internal/server/profile_tool.go, internal/profile/**) does not touch handleReadCache, cacheAuthorization, or anything else on the read_cache authorization hot path — those files (cache_authz.go, mcp.go) are untouched by this PR.
  • Reproduced TestScopeLatency_ReadCache_ScopedVsAdmin locally on the merged branch 3 times in a row: consistently passed with gap ~1.2ms (well under the 20ms budget). A later local run under heavy concurrent CPU load (multiple sibling go test -race jobs on the same machine) showed absolute latencies inflate 10-30x (e.g. admin p95 27.8ms instead of ~2ms) while the admin-vs-scoped GAP stayed small and negative -- i.e. load inflates both arms together, it does not selectively slow the scoped path.
  • Conclusion: the single CI run that failed hit a tail-latency outlier (e.g. a GC pause or scheduler stall during the 200-call scoped-path timing loop) on a shared/noisy GitHub-hosted runner -- consistent with the flakiness this same suite already caused on PRs F and G (documented in the workflow file's own comments). It is not a regression caused by 108-a or by main.

Suggested follow-up (not done here, out of scope for 108-a)

  • Consider a few repeated CI runs of scope-latency.yml to gather real multi-run variance data before ever flipping continue-on-error to false, per the workflow's own stated policy.
  • Optionally reduce P95 sensitivity to a single-call outlier (e.g. median of a few repeated measurement passes) if this recurs.

No code change was made in response to this -- see PR #1380 discussion.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions