Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ base:
SGLANG_HICACHE_DEBUG_LOG: '1'
SGLANG_HICACHE_DEBUG_SAMPLE_RATE: '16384'
SGLANG_MOE_NVFP4_DISPATCH: '1'
SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE: '1'
SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE: '0'
SGLANG_CLIP_MAX_NEW_TOKENS_ESTIMATION: '8'
PIP_BREAK_SYSTEM_PACKAGES: '1'
args:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -111,7 +111,7 @@ base:
SGLANG_ENABLE_TP_MEMORY_INBALANCE_CHECK: '0'
SGLANG_DEEPEP_NUM_MAX_DISPATCH_TOKENS_PER_RANK: '512'
SGLANG_MOE_NVFP4_DISPATCH: '1'
SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE: '1'
SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE: '0'
SGLANG_DEFAULT_THINKING: '1'
SGLANG_CLIP_MAX_NEW_TOKENS_ESTIMATION: '8'
SGLANG_REASONING_EFFORT: max
Expand Down
10 changes: 10 additions & 0 deletions inferencex-e2e/perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9235,3 +9235,13 @@
- "Add DSV4-Pro GB200 Dynamo+SGLang AgentX: TP8 aggregate C1/C4, 1P1D DEP8/DEP16 C64/C128, 1P1D DEP16/DEP32 C256, and 2P1D DEP16/DEP32 C768/C1024/C1280."
- "Use SGLang nightly-dev-20260916-c9a8fba9, DSpark block size 6, and HiCache; omit enable-w4a4-mxfp4-megamoe while retaining the MegaMoE all-to-all backend and FP4 indexer."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3629

- config-keys:
- glm5.2-fp4-b200-dynamo-sglang-agentic-agg
- glm5.2-fp4-b200-dynamo-sglang-agentic-disagg
scenario-type:
- agentic-coding
description:
- "Disable FP8 conversion of the GLM-5.2 NextN/MTP draft MoE on B200 AgentX so the draft runs at its shipped precision; images, topology, workload and acceptance are unchanged."
- "关闭 B200 GLM-5.2 AgentX 对 NextN/MTP draft MoE 的 FP8 转换,使 draft 保持原始发布精度;镜像、拓扑、工作负载和验收不变。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3400
Loading