Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .github/codeowner-signoff-verify-prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -324,7 +324,7 @@ speculative decoding. From the PR diff, identify configs that are BOTH:
downloads, or config names containing `-mtp` / `eagle`.
Agentic replay does not reproduce real-world token-by-token traffic, so measured
acceptance there is not representative. Per the AgentX fairness guidelines
in `golden_al_distribution/README.md` on the checked-out default branch, such configs
in `infx/golden_al_distribution/README.md` on the checked-out default branch, such configs
must instead SIMULATE acceptance at the committed golden acceptance length (AL).
Verify BOTH:
- (a) SIMULATED ACCEPTANCE ENABLED. The launch config must pin a simulated/synthetic
Expand All @@ -343,7 +343,7 @@ Verify BOTH:
FAIL if an agentic spec-decode config runs real (unsimulated) acceptance.
Name the config/script and line.
- (b) AL VALUE MATCHES THE GOLDEN CURVE. Read the committed golden AL YAML for the
model in `golden_al_distribution/` on the default-branch checkout. Examples include
model in `infx/golden_al_distribution/` on the default-branch checkout. Examples include
`qwen3.5_mtp.yaml` and `minimaxm3_eagle3.yaml`. Confirm the pinned AL equals the golden value for that
model, thinking mode, and the config's `num_speculative_tokens` / MTP level (e.g.
qwen3.5 thinking_on with 3 speculative tokens -> 3.39). For TRT-LLM configs, compare
Expand Down
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,7 @@ check_env_vars IS_MULTINODE MODEL_NAME PRECISION

## SRT Slurm synthetic acceptance

- **Do not hard-code synthetic acceptance lengths in SRT recipes, master configs, or launchers.** InferenceX automatically selects the measured value from [`golden_al_distribution/`](golden_al_distribution/) for speculative AgentX throughput runs. Do not add manual `SYNTHETIC_ACCEPTANCE_LENGTH`, vLLM `synthetic_acceptance_length`, SGLang `SGLANG_SIMULATE_ACC_LEN`, or TRT-LLM `TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS` settings.
- **Do not hard-code synthetic acceptance lengths in SRT recipes, master configs, or launchers.** InferenceX automatically selects the measured value from [`infx/golden_al_distribution/`](infx/golden_al_distribution/) for speculative AgentX throughput runs. Do not add manual `SYNTHETIC_ACCEPTANCE_LENGTH`, vLLM `synthetic_acceptance_length`, SGLang `SGLANG_SIMULATE_ACC_LEN`, or TRT-LLM `TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS` settings.
- Submit recipes through [`apply_srt_recipe`](runners/slurm_utils.sh). Its [`infx/srt_slurm` connector](infx/srt_slurm/synthetic_acceptance.py) applies native SRT `--set` / `--unset` overrides; calling upstream `srtctl` directly does not perform InferenceX's automatic selection.
- Keep the actual speculative method, draft model, draft-token count, and relevant sampling settings explicit in the recipe. The connector combines the generation role's settings (decode, otherwise aggregated), after caller overrides, with `MODEL_PREFIX` and `THINKING_MODE` to select the golden curve. For Kimi DSpark, explicitly set `draft_sample_method` to `greedy` or `probabilistic`.
- Eval-only and non-AgentX runs use real verification; the connector removes stale synthetic settings. Non-speculative roles do not receive simulation settings. `RUN_EVAL` does not disable simulation for the throughput portion.
Expand Down
2 changes: 1 addition & 1 deletion benchmarks/multi_node/amd_utils/models.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -87,5 +87,5 @@ DeepSeek-V4-Pro-AgentX: &DeepSeek-V4-Pro-AgentX
# measured clean across c4-c256; the EAGLE 3-1-4 arm remains available through
# spec-decoding: mtp. Synthetic acceptance is checkpoint-specific in
# server_sglang.sh and does not depend on which of the two runs: thinking-on,
# draft lengths 1-8 are wired from golden_al_distribution/dsv4-pro-0813-dspark.yaml.
# draft lengths 1-8 are wired from infx/golden_al_distribution/dsv4-pro-0813-dspark.yaml.
DeepSeek-V4-Pro-0813-AgentX: *DeepSeek-V4-Pro-AgentX
6 changes: 3 additions & 3 deletions benchmarks/multi_node/amd_utils/server_sglang.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1418,15 +1418,15 @@ else
# Agentic trace replay doesn't reproduce real token-by-token traffic, so
# measured MTP/EAGLE acceptance there isn't representative (PR #2309
# review: https://github.com/SemiAnalysisAI/InferenceX/pull/2309#pullrequestreview-4778348624).
# Per the AgentX fairness guidelines (golden_al_distribution/README.md),
# Per the AgentX fairness guidelines (infx/golden_al_distribution/README.md),
# agentic throughput benchmarks simulate acceptance at the model's
# committed golden AL instead of measuring real (non-representative)
# acceptance. Eval runs (RUN_EVAL / EVAL_ONLY) need real acceptance so
# GSM8K scores reflect actual MTP behavior. AgentX uses one golden curve
# per checkpoint, thinking mode, and draft length, including when the
# supported PD draft implementation differs from the calibration engine.
# Sources (thinking_on): dsv4_mtp.yaml for the original checkpoint and
# golden_al_distribution/dsv4-pro-0813-dspark.yaml for Pro-0813.
# infx/golden_al_distribution/dsv4-pro-0813-dspark.yaml for Pro-0813.
DECODE_SIM_ACC_ENV=""
if [[ "$DECODE_MTP_SIZE" -gt 0 ]] && { [[ "${IS_AGENTIC}" == "1" ]] || [[ "${IS_AGENTIC:-}" == "true" ]]; }; then
if [[ "${EVAL_ONLY}" == "true" ]] || [[ "${RUN_EVAL}" == "true" ]]; then
Expand All @@ -1453,7 +1453,7 @@ else
if [[ -n "$DSV4_GOLDEN_AL" ]]; then
DECODE_SIM_ACC_ENV="SGLANG_SIMULATE_ACC_LEN=${DSV4_GOLDEN_AL} SGLANG_SIMULATE_ACC_METHOD=match-expected SGLANG_SIMULATE_ACC_TOKEN_MODE=real-draft-token"
else
echo "WARNING: agentic spec-decoding run (model=${MODEL_NAME}, algorithm=${SPEC_DECODING:-mtp}, DECODE_MTP_SIZE=${DECODE_MTP_SIZE}) has no golden AL wired in server_sglang.sh -- falling back to real (unsimulated, non-representative) acceptance. Add a case in server_sglang.sh and golden_al_distribution/ before shipping this arm. See golden_al_distribution/README.md." >&2
echo "WARNING: agentic spec-decoding run (model=${MODEL_NAME}, algorithm=${SPEC_DECODING:-mtp}, DECODE_MTP_SIZE=${DECODE_MTP_SIZE}) has no golden AL wired in server_sglang.sh -- falling back to real (unsimulated, non-representative) acceptance. Add a case in server_sglang.sh and infx/golden_al_distribution/ before shipping this arm. See infx/golden_al_distribution/README.md." >&2
fi
fi
fi
Expand Down
2 changes: 1 addition & 1 deletion benchmarks/multi_node/amd_utils/server_tilert.sh
Original file line number Diff line number Diff line change
Expand Up @@ -241,7 +241,7 @@ start_decode() {
local extra=( ${TILERT_DECODE_EXTRA_FLAGS} )
if [[ "$SPEC_DECODING" == "mtp" && "$EVAL_ONLY" != "true" && "$RUN_EVAL" != "true" ]]; then
check_env_vars MODEL_PREFIX THINKING_MODE
local curve="${WS_PATH%/benchmarks/*}/golden_al_distribution/${MODEL_PREFIX}_mtp.yaml"
local curve="${WS_PATH%/benchmarks/*}/infx/golden_al_distribution/${MODEL_PREFIX}_mtp.yaml"
TILERT_SIMULATE_ACC_LEN="$("$PY" - "$curve" "$THINKING_MODE" "$DECODE_MTP_SIZE" <<'PYEOF'
import sys, yaml
path, thinking, tokens = sys.argv[1], sys.argv[2], int(sys.argv[3])
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -125,7 +125,7 @@ base:
# One TP8 worker serves prefill and decode on a single node. The sweep matrix
# supplies the concurrency list; srt_agentic.sh replays every point against
# this one server. Acceptance is pinned to the golden thinking-on AL for three
# speculative tokens (golden_al_distribution/glm5.2_mtp.yaml).
# speculative tokens (infx/golden_al_distribution/glm5.2_mtp.yaml).
override_c1:
name: agg-b200-tp8-c1-mtp
roles:
Expand All @@ -139,7 +139,7 @@ override_c1:
# One TP8 worker serves prefill and decode on a single node. The sweep matrix
# supplies the concurrency list; srt_agentic.sh replays every point against
# this one server. Acceptance is pinned to the golden thinking-on AL for three
# speculative tokens (golden_al_distribution/glm5.2_mtp.yaml).
# speculative tokens (infx/golden_al_distribution/glm5.2_mtp.yaml).
override_c4:
name: agg-b200-tp8-c4-mtp
roles:
Expand All @@ -153,7 +153,7 @@ override_c4:
# One TP8 worker serves prefill and decode on a single node. The sweep matrix
# supplies the concurrency list; srt_agentic.sh replays every point against
# this one server. Acceptance is pinned to the golden thinking-on AL for three
# speculative tokens (golden_al_distribution/glm5.2_mtp.yaml).
# speculative tokens (infx/golden_al_distribution/glm5.2_mtp.yaml).
override_c8:
name: agg-b200-tp8-c8-mtp
roles:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -39,10 +39,10 @@ MAX_NUM_SEQS="64"
# Reserve device memory for KV cache and speculative verification.
GPU_MEM_UTIL="0.90"
TEMPERATURE="1.0"
# MUST match the golden config: golden_al_distribution/dsv4_mtp.yaml was measured
# MUST match the golden config: infx/golden_al_distribution/dsv4_mtp.yaml was measured
# with reasoning_effort=high.
# The published recipe uses greedy; probabilistic won at every level on Kimi-K3
# (golden_al_distribution/kimik3_dspark*.yaml). vLLM accepts exactly these two values
# (infx/golden_al_distribution/kimik3_dspark*.yaml). vLLM accepts exactly these two values
# (vllm/config/speculative.py: DraftSampleMethod).
case "$DRAFT_SAMPLE_METHOD" in
greedy|probabilistic) ;;
Expand Down
4 changes: 2 additions & 2 deletions configs/amd-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -654,7 +654,7 @@ dsr1-fp8-mi355x-sglang-disagg-mtp:

# Kimi-K3 MXFP4 agentic-coding benchmark on MI355X via ATOM with DSpark
# speculative decoding. Acceptance is pinned to the committed golden curve in
# golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml
# infx/golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml
# at both draft lengths used here: 7 draft tokens -> AL 3.84 at concurrency 1 and 4,
# 3 draft tokens -> AL 3.00 at concurrency 14 and 16.
# Companion to kimik3-fp4-mi355x-vllm-agentic-mtp: same checkpoint, same runner.
Expand Down Expand Up @@ -1411,7 +1411,7 @@ dsv41flash-fp4-mi355x-vllm-agentic-dspark:
# and draft length (docs/PR_REVIEW_CHECKLIST.md), and a submission may not
# substitute its own target. Nothing about that target is written here:
# AGENTS.md forbids hard-coding an acceptance length in a master config, so
# server_tilert.sh reads golden_al_distribution/glm5.3_mtp.yaml at launch and
# server_tilert.sh reads infx/golden_al_distribution/glm5.3_mtp.yaml at launch and
# fails the run if the curve or the draft length is missing.
# The curve is consumed in the same units as every other framework here, and
# AgentX replays run with thinking on.
Expand Down
16 changes: 8 additions & 8 deletions configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1502,7 +1502,7 @@ dsr1-fp8-b300-sglang-mtp:
kimik3-fp4-b300-vllm-agentic-dspark:
# TP8 x DCP8 with Mooncake as the external KV tier. The recipe drafts with
# DSpark level 7 at conc <= 8 and stops above it; synthetic acceptance is
# pinned to the committed golden AL 3.84 (golden_al_distribution/
# pinned to the committed golden AL 3.84 (infx/golden_al_distribution/
# kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml),
# and EVAL_ONLY switches to real block verification.
#
Expand Down Expand Up @@ -5177,7 +5177,7 @@ qwen3.5-fp8-h100-sglang-agentic:
# qwen3.5-fp8-h100-sglang-agentic: same TP8/EP8 GPU-resident and HiCache arms,
# plus SGLang EAGLE MTP (num-steps 3, eagle-topk 1, 4 draft tokens = 3
# speculative tokens) with simulated acceptance pinned to the golden AL 3.39
# (golden_al_distribution/qwen3.5_mtp.yaml, thinking_on, K=3) -- the same value
# (infx/golden_al_distribution/qwen3.5_mtp.yaml, thinking_on, K=3) -- the same value
# the GB300 Qwen3.5 AgentX srt-slurm recipes use. Image moves to
# lmsysorg/sglang:v0.5.16-cu130 because SGLANG_SIMULATE_ACC_TOKEN_MODE only
# exists from v0.5.16; v0.5.14 is the version the fixed-seq-len H100 MTP entry
Expand All @@ -5187,7 +5187,7 @@ qwen3.5-fp8-h100-sglang-agentic:
# policy that new agentic arms enable speculative decoding rather than running a
# separate STP baseline (docs/MODELS.md). SGLang EAGLE MTP (num-steps 3, eagle-topk 1,
# 4 draft tokens = 3 speculative tokens) with simulated acceptance pinned to the
# golden AL 3.39 (golden_al_distribution/qwen3.5_mtp.yaml, thinking_on, K=3).
# golden AL 3.39 (infx/golden_al_distribution/qwen3.5_mtp.yaml, thinking_on, K=3).
# Image is lmsysorg/sglang:v0.5.16-cu130: SGLANG_SIMULATE_ACC_TOKEN_MODE only
# exists from v0.5.16, and v0.5.14 is what the fixed-seq-len H200 MTP entry runs.
# Search space is the H100 AgentX shape widened for H200's 141 GB: the
Expand Down Expand Up @@ -7020,7 +7020,7 @@ glm5.2-fp8-h200-dynamo-sglang-agentic-mtp-agg:
# AgentX speculative-decoding policy in docs/MODELS.md. SGLang EAGLE runs off
# GLM-5.2's built-in nextn head (num-steps 3,
# eagle-topk 1, 4 draft tokens = 3 speculative tokens), with acceptance pinned
# to the golden AL 2.99 (golden_al_distribution/glm5.2_mtp.yaml, thinking_on,
# to the golden AL 2.99 (infx/golden_al_distribution/glm5.2_mtp.yaml, thinking_on,
# K=3) through SGLANG_SIMULATE_ACC_*. The pinned nightly includes FlashInfer
# 0.6.18's BF16 TRTLLM MoE allocation fix for small-batch Blackwell execution
# in GLM-5.2's unquantized EAGLE draft head.
Expand Down Expand Up @@ -7048,7 +7048,7 @@ glm5.2-fp4-b300-sglang-agentic-mtp:
# glm5.2-fp8-b200-sglang-agentic-mtp (#2863). Same spec-decode-only shape per
# the AgentX policy (docs/MODELS.md): SGLang EAGLE off GLM-5.2's built-in nextn head
# (num-steps 3, eagle-topk 1, 4 draft tokens = 3 speculative tokens) with
# acceptance pinned to the golden AL 2.99 (golden_al_distribution/glm5.2_mtp.yaml,
# acceptance pinned to the golden AL 2.99 (infx/golden_al_distribution/glm5.2_mtp.yaml,
# thinking_on, K=3), which was measured on glm-5.2-fp8, i.e. on this checkpoint.
# Pinned to the 2026-09-08 cu13 dev nightly (the first that carries sgl-project/sglang#38318, the guard for the EAGLE DSA fp8 read-door crash seen on the 2026-09-07 build), the same tag the B200/B300 GLM-5.2
# and Qwen3.5 SGLang AgentX recipes moved to that day. Runs on cluster:b300-dsxe
Expand Down Expand Up @@ -7078,7 +7078,7 @@ glm5.2-fp8-b300-sglang-agentic-mtp:
# policy that agentic arms enable speculative decoding rather than running a
# separate STP baseline (docs/MODELS.md). SGLang EAGLE off GLM-5.2's built-in nextn
# head (num-steps 3, eagle-topk 1, 4 draft tokens = 3 speculative tokens) with
# acceptance pinned to the golden AL 2.99 (golden_al_distribution/glm5.2_mtp.yaml,
# acceptance pinned to the golden AL 2.99 (infx/golden_al_distribution/glm5.2_mtp.yaml,
# thinking_on, K=3) through SGLANG_SIMULATE_ACC_*. SGLang v0.5.16 is the first
# release that reads SGLANG_SIMULATE_ACC_TOKEN_MODE. The pinned 2026-09-01
# nightly includes FlashInfer 0.6.18's BF16 TRTLLM MoE allocation fix, which is
Expand Down Expand Up @@ -7107,7 +7107,7 @@ glm5.2-fp4-b200-sglang-agentic-mtp:
# glm5.2-fp4-b200-sglang-agentic-mtp. Same spec-decode-only shape per the AgentX
# policy (docs/MODELS.md): SGLang EAGLE off GLM-5.2's built-in nextn head (num-steps
# 3, eagle-topk 1, 4 draft tokens = 3 speculative tokens) with acceptance pinned
# to the golden AL 2.99 (golden_al_distribution/glm5.2_mtp.yaml, thinking_on,
# to the golden AL 2.99 (infx/golden_al_distribution/glm5.2_mtp.yaml, thinking_on,
# K=3) through SGLANG_SIMULATE_ACC_*. That curve was measured on glm-5.2-fp8,
# i.e. on this checkpoint. Pinned to the 2026-09-08 cu13 dev nightly (the first that carries sgl-project/sglang#38318, the guard for the EAGLE DSA fp8 read-door crash seen on the 2026-09-07 build), the same
# tag the B200 Qwen3.5 FP8/FP4 SGLang AgentX recipes moved to that day.
Expand Down Expand Up @@ -7135,7 +7135,7 @@ glm5.2-fp8-b200-sglang-agentic-mtp:

# GLM-5.2 NVFP4 B200 AgentX on Dynamo + SGLang. EAGLE uses the model's
# built-in nextn head, with acceptance pinned to the golden thinking-on AL in
# golden_al_distribution/glm5.2_mtp.yaml. Both arms use HiCache write-back
# infx/golden_al_distribution/glm5.2_mtp.yaml. Both arms use HiCache write-back
# DRAM offload for the agentic working set.
glm5.2-fp4-b200-dynamo-sglang-agentic-agg:
image: lmsysorg/sglang:nightly-dev-20260910-00840301
Expand Down
2 changes: 1 addition & 1 deletion docs/MODELS.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ The A/B retirements below concern standalone non-spec-decode baselines maintaine

**Status: baseline retirement is not yet enacted.** Retire a redundant baseline only once replacement coverage exists for the affected model, hardware, and engine. Keep non-spec-decode configurations that contribute to the Pareto frontier; the existence of a spec-decode counterpart alone is not a reason to remove them.

**Going forward we no longer maintain separate non-spec-decode and spec-decode tracks solely as an A/B comparison.** The non-spec-decode arm existed as a neutral baseline back when acceptance length wasn't standardized. That is now solved. [`golden_al_distribution/`](../golden_al_distribution/) commits one golden acceptance-length curve per model, thinking mode, and draft length, measured on the SPEED-Bench `coding` category. When speculative decoding is enabled, AgentX pins submissions to that curve through synthetic acceptance (vLLM `synthetic_acceptance_length`, SGLang `SGLANG_SIMULATE_ACC_LEN`, TensorRT-LLM `TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS`, etc). With a fair, engine-independent acceptance target in place, spec-decode results are directly comparable on their own and a dedicated non-spec-decode baseline is redundant. New models, including Kimi-K3, do not require that separate baseline from day 0.
**Going forward we no longer maintain separate non-spec-decode and spec-decode tracks solely as an A/B comparison.** The non-spec-decode arm existed as a neutral baseline back when acceptance length wasn't standardized. That is now solved. [`infx/golden_al_distribution/`](../infx/golden_al_distribution/) commits one golden acceptance-length curve per model, thinking mode, and draft length, measured on the SPEED-Bench `coding` category. When speculative decoding is enabled, AgentX pins submissions to that curve through synthetic acceptance (vLLM `synthetic_acceptance_length`, SGLang `SGLANG_SIMULATE_ACC_LEN`, TensorRT-LLM `TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS`, etc). With a fair, engine-independent acceptance target in place, spec-decode results are directly comparable on their own and a dedicated non-spec-decode baseline is redundant. New models, including Kimi-K3, do not require that separate baseline from day 0.

**Publish the best Pareto points, whether speculative decoding is enabled or disabled.** Recipes may disable MTP, EAGLE/EAGLE3, DSpark, or another draft method when doing so produces a better operating point, for example at high throughput. Valid non-spec-decode results remain eligible for publication under the same [North-star Pareto policy](#north-star-pareto-policy). A frontier may therefore contain both spec-decode and non-spec-decode points.

Expand Down
Loading
Loading