Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Tests srt-slurm #31 at
4f95eee1b9f50fc0dadbc7163f93800680bca0a7as a patch on the current InferenceX runtime.One Qwen3.5 FP8 fixed-sequence point per cluster: TP8, concurrency 4, 8k1k. MI355X uses a test-only TP8 override; these results are not for performance publication.
Runs the official AMD exporter on port 19500 to avoid existing cluster services, requires native power telemetry, and writes the single-node measurement window using the existing InferenceX helper. Python setup uses runner scratch because MI325X's home-directory mount is unavailable.
Validation
The disaggregated smoke uses the existing Qwen3.5 FP8 SGLang/MoRI recipe, with test-only TP8 prefill and one concurrency, to validate all 16 GPUs across two nodes. Required native telemetry is enabled.
Do not merge this smoke-test branch.