Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/AGENT_OPERATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,7 +94,7 @@ Multinode disaggregated results add `prefill_gpu_energy_j`, `decode_gpu_energy_j

Every power result — valid or invalid, single-node or multinode — carries `power_metric_schema_version`. Version 2 defines each unprefixed `joules_per_*` field as whole-deployment GPU-board energy over the named denominator; role-scoped energy uses the explicit `prefill_*` / `decode_*` keys. Rows without the field predate the whole-deployment switch and their unprefixed joules are not comparable across topologies.

For srt-slurm recipes, `telemetry.enabled: true` with `telemetry.dcgm_exporter` enables official energy collection. The Git submodule pointer at `inferencex-e2e/utils/srt-slurm` is the source of truth for the shared srt-slurm commit, used by both power and non-power NVIDIA lanes. TileRT is the single documented fork exception. CI derives `POWER_PRODUCER_SHA` from the launcher stamp. The aggregate-power and AgentX power tests validate telemetry and provenance. These local tests do not prove hardware power collection. Eligible recipe-gated `dynamo-sglang` dcgm-power lanes are validated.
For srt-slurm recipes, `telemetry.enabled: true` with `telemetry.dcgm_exporter` enables official energy collection. The Git submodule pointer at `inferencex-e2e/utils/srt-slurm` is the source of truth for every srt-slurm job, including TileRT. CI derives `POWER_PRODUCER_SHA` from the launcher stamp. The aggregate-power and AgentX power tests validate telemetry and provenance. These local tests do not prove hardware power collection. Eligible recipe-gated `dynamo-sglang` dcgm-power lanes are validated.

Power audit artifacts are named `power_audit_<result>` and contain `power_validation_<result>.json` for single-node runs or `power_validation_<result>_*.json` for multinode runs. They are uploaded even when validation fails.

Expand Down
23 changes: 19 additions & 4 deletions inferencex-e2e/benchmarks/benchmark_lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -2003,19 +2003,34 @@ get_native_max_context_length() {
if [ -n "${MODEL_PATH:-}" ] && [ -d "${MODEL_PATH}" ]; then
model_path="${MODEL_PATH}"
fi
python3 -c "
python3 - "$model_path" <<'PY'
import json
import sys
from pathlib import Path

fields = ['max_position_embeddings', 'max_sequence_length', 'seq_length', 'n_positions']
try:
config = json.loads((Path(sys.argv[1]) / 'config.json').read_text())
for field in fields:
value = config.get(field)
if type(value) is int and value > 0:
print(value)
sys.exit(0)
except (OSError, ValueError, AttributeError):
pass

try:
from transformers import AutoConfig
config = AutoConfig.from_pretrained('${model_path}', trust_remote_code=True)
for attr in ['max_position_embeddings', 'max_sequence_length', 'seq_length', 'n_positions']:
config = AutoConfig.from_pretrained(sys.argv[1], trust_remote_code=True)
for attr in fields:
if hasattr(config, attr):
print(getattr(config, attr))
break
else:
print(0)
except Exception:
print(0)
"
PY
}

# Requested benchmark context capped at the model's native max. Sets
Expand Down

This file was deleted.

27 changes: 6 additions & 21 deletions inferencex-e2e/benchmarks/multi_node/amd_utils/env.sh
Original file line number Diff line number Diff line change
Expand Up @@ -2,21 +2,12 @@

source "$(dirname "${BASH_SOURCE[0]}")/../../benchmark_lib.sh" --validation-only
check_env_vars ENGINE
# MoRI-IO queue-pair tuning, the UCX RoCE GID index, SGLang router logging and the
# SGLang decode cuda-graph NCCL workaround. Only the SGLang MoRI KV path
# below reads these. ENGINE=tilert moves KV over mooncake and starts no SGLang
# router, so it is neither given nor reads them: validating them there would force
# the recipe to invent MoRI tuning for a transport it never uses.
if [[ "$ENGINE" != "tilert" ]]; then
check_env_vars \
MORI_IO_SQ_BACKOFF_TIMEOUT_US MORI_IO_QP_MAX_SEND_WR MORI_IO_QP_MAX_CQE MORI_IO_QP_MAX_SGE MORI_IO_TC_DISABLE \
UCX_IB_GID_INDEX MORI_APP_LOG_LEVEL SGLANG_ROUTER_STDOUT_LOGS TORCH_NCCL_BLOCKING_WAIT NCCL_BLOCKING_WAIT \
SGLANG_OPT_USE_AITER_INDEXER
fi
# Dual-engine environment setup for multi-node disaggregated serving.
#
# ENGINE=sglang-disagg or tilert selects the engine-specific block.
#
# SGLang MoRI environment.
check_env_vars \
MORI_IO_SQ_BACKOFF_TIMEOUT_US MORI_IO_QP_MAX_SEND_WR MORI_IO_QP_MAX_CQE MORI_IO_QP_MAX_SGE MORI_IO_TC_DISABLE \
UCX_IB_GID_INDEX MORI_APP_LOG_LEVEL SGLANG_ROUTER_STDOUT_LOGS TORCH_NCCL_BLOCKING_WAIT NCCL_BLOCKING_WAIT \
SGLANG_OPT_USE_AITER_INDEXER

# REQUIRED ENVIRONMENT VARIABLES:
# IBDEVICES - RDMA/InfiniBand device names (e.g., ionic_0,ionic_1,... or mlx5_0,mlx5_1,...)
# Set by runner or auto-detected from hostname.
Expand Down Expand Up @@ -119,10 +110,6 @@ else
fi
fi

if [[ "$ENGINE" == "tilert" ]]; then
echo "[INFO] tilert: IBDEVICES=$IBDEVICES NCCL_SOCKET_IFNAME=$NCCL_SOCKET_IFNAME NCCL_IB_HCA=$NCCL_IB_HCA"

else

export SGLANG_USE_AITER=1
export AITER_LOG_LEVEL=ERROR
Expand Down Expand Up @@ -237,5 +224,3 @@ else
export GPU_MAX_HW_QUEUES=2
fi
fi

fi
67 changes: 5 additions & 62 deletions inferencex-e2e/benchmarks/multi_node/amd_utils/job.slurm
Original file line number Diff line number Diff line change
Expand Up @@ -34,11 +34,7 @@ echo ""

# Use $(pwd) not BASH_SOURCE — sbatch copies the script to /var/spool/slurmd/
# at runtime, but the CWD remains the submit-time directory (amd_utils/).
if [[ "$ENGINE" == "tilert" ]]; then
MODELS_YAML="$(pwd)/models_tilert.yaml"
else
MODELS_YAML="$(pwd)/models.yaml"
fi
MODELS_YAML="$(pwd)/models.yaml"

if [[ ! -f "$MODELS_YAML" ]]; then
echo "Error: models YAML not found at $MODELS_YAML"
Expand All @@ -50,16 +46,6 @@ if [[ -z "${DOCKER_IMAGE_NAME:-}" ]]; then
exit 1
fi

if [[ "$ENGINE" == "tilert" && -z "${PREFILL_IMAGE:-}" ]]; then
echo "Error: ENGINE=tilert requires PREFILL_IMAGE (e.g. PREFILL_IMAGE=vllm/vllm-openai-rocm:nightly-<sha> in prefill.additional-settings)."
exit 1
fi
if [[ "$ENGINE" == "tilert" ]]; then
# server_tilert.sh takes the container-creation barrier timeout from the
# recipe (no 300s default as on the SGLang path); fail here, before sbatch
# work is done, rather than inside the container.
check_env_vars CONTAINER_BARRIER_TIMEOUT
fi

# Resolve the models.yaml entry the same way server_sglang.sh does: agentic runs
# (IS_AGENTIC) use the '<model>-AgentX' recipe, non-agentic disaggregated runs use
Expand Down Expand Up @@ -541,46 +527,9 @@ DOCKER_ENV_COMMON=(
)

# Engine-specific env vars
if [[ "$ENGINE" == "tilert" ]]; then
DOCKER_ENV_ENGINE=(
-e MODEL_PATH=$DOCKER_MODEL_PATH
-e PREFILL_IMAGE=${PREFILL_IMAGE}
-e TILERT_VERSION=${TILERT_VERSION}
-e TILERT_PROFILE=${TILERT_PROFILE}
-e TILERT_MODEL_TYPE=${TILERT_MODEL_TYPE}
-e TILERT_MODEL_PKG=${TILERT_MODEL_PKG}
-e TILERT_MAX_MODEL_LEN=${TILERT_MAX_MODEL_LEN}
-e TILERT_TRANSPORT=${TILERT_TRANSPORT}
-e TILERT_PARSER=${TILERT_PARSER}
-e TILERT_QUEUE_TIMEOUT=${TILERT_QUEUE_TIMEOUT}
-e TILERT_WEIGHTS_DIR=${TILERT_WEIGHTS_DIR}
-e TILERT_RDMA_STRICT=${TILERT_RDMA_STRICT}
-e TILERT_CONVERT_LOCK_WAIT=${TILERT_CONVERT_LOCK_WAIT}
-e TILERT_SIMULATE_ACC_METHOD=${TILERT_SIMULATE_ACC_METHOD}
-e \"TILERT_EXTRA_ENV=${TILERT_EXTRA_ENV:-}\"
-e SERVED_MODEL_NAME=${SERVED_MODEL_NAME}
-e PREFILL_KV_DTYPE=${PREFILL_KV_DTYPE}
-e PREFILL_BLOCK_SIZE=${PREFILL_BLOCK_SIZE}
-e PREFILL_SPEC_TOKENS=${PREFILL_SPEC_TOKENS}
-e DECODE_KV_DTYPE=${DECODE_KV_DTYPE}
-e GPU_MEM_UTIL=${GPU_MEM_UTIL}
-e PREFILL_PORT=${PREFILL_PORT}
-e DECODE_CTRL_PORT=${DECODE_CTRL_PORT}
-e DECODE_HTTP_PORT=${DECODE_HTTP_PORT}
-e DECODE_WAIT=${DECODE_WAIT}
-e PREFILL_WAIT=${PREFILL_WAIT}
-e ROUTER_WAIT=${ROUTER_WAIT}
-e SKIP_CONTAINER_BARRIER=${SKIP_CONTAINER_BARRIER}
# Golden-acceptance selection on agentic MTP runs; unset elsewhere.
-e THINKING_MODE=${THINKING_MODE:-}
-e IBDEVICES=${IBDEVICES:-}
-e PYTHONPYCACHEPREFIX=/tmp/pycache
)
else
DOCKER_ENV_ENGINE=(
-e SGLANG_WS_PATH=${WS_PATH}
)
fi
DOCKER_ENV_ENGINE=(
-e SGLANG_WS_PATH=${WS_PATH}
)

# HiCache / Mooncake settings are delivered via a bind-mounted config file rather
# than a long list of docker -e flags. Write it once to the shared benchmark-logs
Expand Down Expand Up @@ -650,12 +599,6 @@ echo \"Rank \$SLURM_PROCID on \$(hostname)\"
eval \"\$DOCKER_CMD_DETECT\"
echo \"[docker-detect] rank \$SLURM_PROCID: DOCKER_CMD=\$DOCKER_CMD\"

RANK_IMAGE=
if [[ \"$ENGINE\" == \"tilert\" && \"\$SLURM_PROCID\" -lt \"$xP\" ]]; then
RANK_IMAGE=\"$PREFILL_IMAGE\"
echo \"[tilert] rank \$SLURM_PROCID is a prefill rank; using PREFILL_IMAGE=\$RANK_IMAGE\"
fi

exec \$DOCKER_CMD run \
--init \
--stop-timeout 10 \
Expand Down Expand Up @@ -696,7 +639,7 @@ exec \$DOCKER_CMD run \
${CLIENT_DOCKER_ENV} \
--name \"$DOCKER_CONT_NAME\" \
--entrypoint \"\" \
\"\${RANK_IMAGE:-$DOCKER_IMAGE_NAME}\" bash -lc '
\"$DOCKER_IMAGE_NAME\" bash -lc '
set -o pipefail
mkdir -p /run_logs/slurm_job-'\"\$SLURM_JOB_ID\"'
'"$RUN_FILE_FULL"' 2>&1 | tee /run_logs/slurm_job-'\"\$SLURM_JOB_ID\"'/server_\$(hostname).log
Expand Down

This file was deleted.

7 changes: 1 addition & 6 deletions inferencex-e2e/benchmarks/multi_node/amd_utils/server.sh
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@ source "$(dirname "${BASH_SOURCE[0]}")/../../benchmark_lib.sh" --validation-only
# Multi-Engine Disaggregated Server Dispatcher
# Dispatches to the engine-specific server launcher based on ENGINE env var.
# ENGINE=sglang-disagg (default) -> server_sglang.sh (SGLang + MoRI)
# ENGINE=tilert -> server_tilert.sh (vLLM prefill + TileRT decode)

check_env_vars ENGINE WS_PATH
if [[ -f /config/hicache_mc.env ]]; then
Expand All @@ -16,8 +15,4 @@ export WS_PATH ENGINE

echo "[DISPATCHER] ENGINE=$ENGINE WS_PATH=$WS_PATH"

if [[ "$ENGINE" == "tilert" ]]; then
source "$WS_PATH/server_tilert.sh"
else
source "$WS_PATH/server_sglang.sh"
fi
source "$WS_PATH/server_sglang.sh"
Loading