Summary
PR #1380 (108-a-profile-policy-model) tripped the advisory FR-011 merge-base regression check (.github/workflows/scope-latency.yml) on TestScopeLatency_ReadCache_ScopedVsAdmin:
read_cache: scoped p95 must not exceed admin's by more than FR-011's 20ms budget
(got admin=1.98966ms scoped=48.901918ms gap=46.912258ms)
Investigation
- The workflow is advisory (
continue-on-error: true) and is not in main's required status checks, so it did not block the merge.
- PR 108-a's actual diff (
internal/server/mcp_visibility.go, internal/server/profile_tool.go, internal/profile/**) does not touch handleReadCache, cacheAuthorization, or anything else on the read_cache authorization hot path — those files (cache_authz.go, mcp.go) are untouched by this PR.
- Reproduced
TestScopeLatency_ReadCache_ScopedVsAdmin locally on the merged branch 3 times in a row: consistently passed with gap ~1.2ms (well under the 20ms budget). A later local run under heavy concurrent CPU load (multiple sibling go test -race jobs on the same machine) showed absolute latencies inflate 10-30x (e.g. admin p95 27.8ms instead of ~2ms) while the admin-vs-scoped GAP stayed small and negative -- i.e. load inflates both arms together, it does not selectively slow the scoped path.
- Conclusion: the single CI run that failed hit a tail-latency outlier (e.g. a GC pause or scheduler stall during the 200-call scoped-path timing loop) on a shared/noisy GitHub-hosted runner -- consistent with the flakiness this same suite already caused on PRs F and G (documented in the workflow file's own comments). It is not a regression caused by 108-a or by main.
Suggested follow-up (not done here, out of scope for 108-a)
- Consider a few repeated CI runs of
scope-latency.yml to gather real multi-run variance data before ever flipping continue-on-error to false, per the workflow's own stated policy.
- Optionally reduce P95 sensitivity to a single-call outlier (e.g. median of a few repeated measurement passes) if this recurs.
No code change was made in response to this -- see PR #1380 discussion.
Summary
PR #1380 (
108-a-profile-policy-model) tripped the advisory FR-011 merge-base regression check (.github/workflows/scope-latency.yml) onTestScopeLatency_ReadCache_ScopedVsAdmin:Investigation
continue-on-error: true) and is not inmain's required status checks, so it did not block the merge.internal/server/mcp_visibility.go,internal/server/profile_tool.go,internal/profile/**) does not touchhandleReadCache,cacheAuthorization, or anything else on the read_cache authorization hot path — those files (cache_authz.go,mcp.go) are untouched by this PR.TestScopeLatency_ReadCache_ScopedVsAdminlocally on the merged branch 3 times in a row: consistently passed with gap ~1.2ms (well under the 20ms budget). A later local run under heavy concurrent CPU load (multiple siblinggo test -racejobs on the same machine) showed absolute latencies inflate 10-30x (e.g. admin p95 27.8ms instead of ~2ms) while the admin-vs-scoped GAP stayed small and negative -- i.e. load inflates both arms together, it does not selectively slow the scoped path.Suggested follow-up (not done here, out of scope for 108-a)
scope-latency.ymlto gather real multi-run variance data before ever flippingcontinue-on-errortofalse, per the workflow's own stated policy.No code change was made in response to this -- see PR #1380 discussion.