Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
72 commits
Select commit Hold shift + click to select a range
6ac1247
fix: bound BFCL requests and guard cross-hardware eval parity
adibarra Sep 15, 2026
bae1ca9
docs: link BFCL cleanup changelog to PR 3156
adibarra Sep 15, 2026
8cffddb
feat: test BFCL through the native Responses API
adibarra Sep 15, 2026
a3107c9
fix: enable native MiniMax tool parsing for TRT evals
adibarra Sep 15, 2026
ffc0aeb
chore: merge main and update eval workflow contracts
adibarra Sep 16, 2026
b026639
test: probe native TRT Responses usage on rc26
adibarra Sep 16, 2026
f949f87
fix: select native FlashInfer for the MiniMax draft
adibarra Sep 16, 2026
4707d7d
chore: merge latest main changelog entries
adibarra Sep 16, 2026
a50e613
fix: align native eval recipes with upstream engine fixes
adibarra Sep 16, 2026
57eee75
chore: incorporate main while preserving changelog history
adibarra Sep 16, 2026
80c1425
test: add resident Kimi vLLM full eval probe
adibarra Sep 16, 2026
46a1a6d
fix: use an available native ROCm image for eval
adibarra Sep 16, 2026
9c48999
chore: synchronize main changelog additions
adibarra Sep 16, 2026
e0ae2c1
fix(kimi): use image-bundled LMCache on ROCm
adibarra Sep 16, 2026
bd4f72e
fix(ci): use job tokens for benchmark checkout
adibarra Sep 16, 2026
1730851
chore: merge latest main changes
adibarra Sep 16, 2026
751acc1
docs(ci): explain run statistics token permissions
adibarra Sep 16, 2026
82b4bcd
fix(recipes): align native TRT and LMCache configuration
adibarra Sep 16, 2026
cf34772
fix(evals): allow long Kimi turns and probe native CPU offload
adibarra Sep 16, 2026
6374e5e
style(evals): format native BFCL handler construction
adibarra Sep 16, 2026
13cb02a
fix(recipes): use native MiniMax CPU offload on both NVIDIA SKUs
adibarra Sep 16, 2026
58705ff
test(evals): verify native CPU cache restores on Kimi
adibarra Sep 17, 2026
47ed9f2
test(evals): add native MiniMax MI300X image probe
adibarra Sep 17, 2026
e1bcb7a
fix(recipes): respect decimal CPU offload budgets
adibarra Sep 17, 2026
c725da3
test(evals): probe resident Kimi on B300
adibarra Sep 17, 2026
e0ff976
test(evals): expand native Kimi B300 suite capacity
adibarra Sep 17, 2026
57ce9ac
chore: merge main into native evaluation branch
adibarra Sep 17, 2026
be1397a
test(evals): bound Kimi generation diagnostics
adibarra Sep 17, 2026
e2213f5
chore: merge main into native evaluation branch
adibarra Sep 17, 2026
905d1d6
chore: sync native evaluation branch with main
adibarra Sep 17, 2026
dcb188b
test: pin native TRT BFCL probe to official nightly
adibarra Sep 18, 2026
027908d
fix: use Enroot-compatible digest syntax for TRT probe
adibarra Sep 18, 2026
63e4502
fix: disable BFCL for TRT and restore original MiniMax images
adibarra Sep 18, 2026
e8a3f2d
docs: explain how to re-enable native TRT BFCL
adibarra Sep 18, 2026
a1da200
chore: merge main into native eval branch
adibarra Sep 21, 2026
9062e6d
chore: merge main into native eval branch
adibarra Sep 22, 2026
0628783
chore: update MiniMax TRT to the latest official nightly
adibarra Sep 22, 2026
368d7ea
chore: merge main into native eval branch
adibarra Sep 22, 2026
0a1e5bd
chore: merge main into native eval branch
adibarra Sep 22, 2026
de42b52
chore: merge main into native eval branch
adibarra Sep 22, 2026
fbdc72e
chore: merge main into native eval branch
adibarra Sep 23, 2026
ec86b66
chore: merge main into native eval branch
adibarra Sep 23, 2026
57175e8
chore: merge main into native eval branch
adibarra Sep 23, 2026
9039a3a
chore: merge main into native eval branch
adibarra Sep 23, 2026
1cd17a8
chore: merge main into native eval branch
adibarra Sep 24, 2026
8354258
chore: merge main into vendor evaluation branch
adibarra Sep 24, 2026
38ec671
chore: merge main into vendor evaluation branch
adibarra Sep 24, 2026
7c2ed99
chore: merge main into vendor evaluation branch
adibarra Sep 24, 2026
a79e6df
chore: merge main into vendor evaluation branch
adibarra Sep 24, 2026
97a78bc
chore: merge main into vendor evaluation branch
adibarra Sep 25, 2026
29cc25b
chore: merge main into vendor evaluation branch
adibarra Sep 25, 2026
4241ffe
chore: merge latest main into vendor evaluation branch
adibarra Sep 25, 2026
8f7ee7d
chore: merge main into vendor evaluation branch
adibarra Sep 25, 2026
1ab0675
test: handle git global options in launcher fixture
adibarra Sep 25, 2026
b5cf65d
chore: merge main into vendor evaluation branch
adibarra Sep 25, 2026
018a6a4
chore: merge main into vendor evaluation branch
adibarra Sep 25, 2026
7ef0343
chore: merge main into vendor evaluation branch
adibarra Sep 25, 2026
5a1551a
chore: merge main into vendor evaluation branch
adibarra Sep 25, 2026
c346f1f
chore: merge main into the BFCL parity branch
adibarra Sep 25, 2026
dbd59b0
feat: merge main and port BFCL evaluation to native SRT-Slurm
adibarra Sep 26, 2026
cecd882
chore: merge latest main launcher cleanup
adibarra Sep 26, 2026
1e369b7
chore: merge main model and tooling cleanup
adibarra Sep 26, 2026
3b8803f
chore(sync): merge main and relocate BFCL tests [skip-sweep]
adibarra Sep 27, 2026
16e67a1
chore(sync): merge main H100 image update [skip-sweep]
adibarra Sep 27, 2026
5385b81
chore(sync): merge main CollectiveX updates [skip-sweep]
adibarra Sep 27, 2026
437455d
chore(sync): merge main CollectiveX MoRI updates [skip-sweep]
adibarra Sep 27, 2026
fe28a8b
chore: merge main and preserve tool evals in project layout [skip-sweep]
adibarra Sep 28, 2026
9ca0af7
chore: merge main documentation updates [skip-sweep]
adibarra Sep 28, 2026
682b2c7
chore: merge main collective tooling updates [skip-sweep]
adibarra Sep 28, 2026
c412719
chore: merge main into BFCL evaluation branch [skip-sweep]
adibarra Sep 28, 2026
56cc583
chore: merge main SRT-Slurm update [skip-sweep]
adibarra Sep 28, 2026
47bde66
chore: merge main recipe updates [skip-sweep]
adibarra Sep 29, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions .github/workflows/benchmark-multinode-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ on:
required: false
default: "lm-eval"
eval-suite:
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_responses_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
type: string
required: false
default: ""
Expand Down Expand Up @@ -152,7 +152,7 @@ env:
DISAGG: ${{ fromJSON(inputs.config).disagg }}
PREFILL_HARDWARE: ${{ fromJSON(inputs.config).prefill.hardware }}
DECODE_HARDWARE: ${{ fromJSON(inputs.config).decode.hardware }}
RUN_EVAL: ${{ inputs.run-eval }}
RUN_EVAL: ${{ inputs.run-eval && !(inputs.eval-framework == 'bfcl' && (fromJSON(inputs.config).framework == 'trt' || fromJSON(inputs.config).framework == 'dynamo-trt')) }}
EVAL_ONLY: ${{ inputs.eval-only }}
EVAL_FRAMEWORK: ${{ inputs.eval-framework }}
EVAL_SUITE: ${{ inputs.eval-suite }}
Expand Down Expand Up @@ -202,6 +202,8 @@ permissions:

jobs:
benchmark:
# BFCL is disabled for TRT; skip eval-only jobs before reserving a GPU runner.
if: ${{ !(inputs.eval-only && inputs.eval-framework == 'bfcl' && (fromJSON(inputs.config).framework == 'trt' || fromJSON(inputs.config).framework == 'dynamo-trt')) }}
runs-on: >-
${{ fromJSON(
vars.PRIORITY_SCHEDULER_ENABLED == 'true' &&
Expand Down Expand Up @@ -517,7 +519,7 @@ jobs:
"$INFERENCEX_RESULTS_PYTHON" -P -m infx.evals.validate_scores --expected-concs "${expected_concs}"

- name: Cleanup eval outputs (post-upload)
if: ${{ always() && (inputs.run-eval || inputs.eval-only) }}
if: ${{ always() && (env.RUN_EVAL == 'true' || inputs.eval-only) }}
run: |
if [[ -f inferencex-e2e/configs/runners.yaml ]]; then cd inferencex-e2e; fi
rm -f meta_env.json || true
Expand Down
6 changes: 4 additions & 2 deletions .github/workflows/benchmark-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -51,7 +51,7 @@ on:
required: false
default: "lm-eval"
eval-suite:
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_responses_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
type: string
required: false
default: ""
Expand Down Expand Up @@ -131,7 +131,7 @@ env:
CONC: ${{ fromJSON(inputs.config).conc }}
SPEC_DECODING: ${{ fromJSON(inputs.config).spec-decoding }}
DISAGG: ${{ inputs.scenario-type == 'agentic-coding' && 'false' || fromJSON(inputs.config).disagg }}
RUN_EVAL: ${{ inputs.run-eval }}
RUN_EVAL: ${{ inputs.run-eval && !(inputs.eval-framework == 'bfcl' && (fromJSON(inputs.config).framework == 'trt' || fromJSON(inputs.config).framework == 'dynamo-trt')) }}
EVAL_ONLY: ${{ inputs.eval-only }}
EVAL_FRAMEWORK: ${{ inputs.eval-framework }}
EVAL_SUITE: ${{ inputs.eval-suite }}
Expand Down Expand Up @@ -164,6 +164,8 @@ permissions:

jobs:
benchmark:
# BFCL is disabled for TRT; skip eval-only jobs before reserving a GPU runner.
if: ${{ !(inputs.eval-only && inputs.eval-framework == 'bfcl' && (fromJSON(inputs.config).framework == 'trt' || fromJSON(inputs.config).framework == 'dynamo-trt')) }}
runs-on: >-
${{ fromJSON(
vars.PRIORITY_SCHEDULER_ENABLED == 'true' &&
Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/e2e-tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ on: # zizmor: ignore[concurrency-limits]
type: string
default: "auto"
eval-suite:
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_responses_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
required: false
type: string
default: ""
Expand Down Expand Up @@ -159,7 +159,7 @@ on: # zizmor: ignore[concurrency-limits]
type: string
default: "auto"
eval-suite:
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
description: "Eval suite (kimi_tool_call_schema, kimi_tool_call_schema_full, minimax_m3_smoke, minimax_m3_full, bfcl_smoke, bfcl_responses_smoke, bfcl_vllm_minimax_m3, or bfcl_vllm_kimi); empty for lm-eval and swebench"
required: false
type: string
default: ""
Expand Down
37 changes: 22 additions & 15 deletions inferencex-e2e/benchmarks/benchmark_lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -363,21 +363,6 @@ if not speed_bench_ok or not sample_ok:
PYEOF
}

disable_trtllm_detailed_perf_metrics() {
local trtllm_root
local py_executor
local detailed_metrics_gate="enabled=getattr(self.llm_args, 'return_perf_metrics', False))"

trtllm_root=$(python3 -c 'from importlib.util import find_spec; from pathlib import Path; print(Path(find_spec("tensorrt_llm").origin).parent)')
py_executor="$trtllm_root/_torch/pyexecutor/py_executor.py"
if ! grep -Fq "$detailed_metrics_gate" "$py_executor"; then
echo "Error: unsupported TensorRT-LLM detailed metrics implementation in $py_executor" >&2
return 1
fi
sed -i "s/enabled=getattr(self.llm_args, 'return_perf_metrics', False))/enabled=False)/" "$py_executor"
grep -Fq "self.perf_manager = PerfMetricsManager(" "$py_executor"
grep -Fq "enabled=False)" "$py_executor"
}

# Agentic replays must use the model's native context limit. Ignore inherited
# workflow or shell overrides so neither the server nor AIPerf applies a cap.
Expand Down Expand Up @@ -1714,20 +1699,39 @@ _run_bfcl_smoke_eval() {
_run_bfcl_suite_eval bfcl_smoke 4 900 false "$@"
}

_skip_bfcl_for_trt() {
case "${FRAMEWORK:-}" in
trt|dynamo-trt)
echo "SKIP: BFCL is disabled for ${FRAMEWORK}; no evaluation score was produced."
return 0
;;
*) return 1 ;;
esac
}

run_bfcl_eval() {
if _skip_bfcl_for_trt; then
return 0
fi
local eval_suite="${EVAL_SUITE:-bfcl_smoke}"
export EVAL_SUITE="$eval_suite"

case "$eval_suite" in
bfcl_smoke)
_run_bfcl_smoke_eval "$@"
;;
bfcl_responses_smoke)
_run_bfcl_suite_eval "$eval_suite" 4 900 false "$@"
;;
bfcl_vllm_minimax_m3)
_run_bfcl_suite_eval "$eval_suite" 8 7200 true "$@"
;;
bfcl_vllm_kimi)
_run_bfcl_suite_eval "$eval_suite" 16 14400 true "$@"
;;
bfcl_kimi_diagnostic)
_run_bfcl_suite_eval "$eval_suite" 16 600 true "$@"
;;
*)
echo "ERROR: unsupported BFCL suite '${eval_suite}'" >&2
export EVAL_RESULT_DIR=""
Expand Down Expand Up @@ -2922,6 +2926,9 @@ run_eval() {
fi

local framework="${EVAL_FRAMEWORK:-${cli_framework:-$scenario_default}}"
if [ "$framework" = "bfcl" ] && _skip_bfcl_for_trt; then
return 0
fi
case "$framework" in
kimi-vendor)
[ -n "${EVAL_SUITE:-}" ] || EVAL_SUITE="kimi_tool_call_schema"
Expand Down

This file was deleted.

Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ schema: 2
base:
model:
path: minimax-m3-nvfp4
container: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc23.post1
container: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc28.dev202609220000
precision: fp4
dynamo:
install: true
Expand Down Expand Up @@ -72,11 +72,11 @@ base:
implementation: msa
indexer_kv_dtype: fp8
sparse_disable_index_value: true
fuse_qkv_index_projection: true
kv_cache_config:
free_gpu_memory_fraction: 0.94
enable_block_reuse: true
block_reuse_policy: per_conversation
block_reuse_config:
policy: per_conversation
tokens_per_block: 128
use_kv_cache_manager_v2: true
dtype: fp8
Expand Down
Original file line number Diff line number Diff line change
@@ -1,15 +1,14 @@
# Kimi-K3 MXFP4 AgentX on MI355X with vLLM DSpark
# (https://recipes.vllm.ai/moonshotai/Kimi-K3). TP8 only: the 1.56 TB checkpoint
# is ~195 GB per GPU. The KV cache is GPU-resident through c4 and backed by
# vLLM's SimpleCPUOffloadConnector from c8. The DCP8 arm (c44-c70) ran without a
# draft model on the legacy script and was removed with it (#3461). No master
# config uses this recipe yet.
# vLLM's SimpleCPUOffloadConnector from c8. The DCP8 arm (c44-c70) runs without
# a draft model, preserving the configured MTP label.
base:
schema: 2
name: kimik3-fp4-mi355x-vllm-agentic
model:
path: hf:moonshotai/Kimi-K3
container: vllm/vllm-openai-rocm:nightly-rocm100-af1c01499b289be555c475669ba50a88e96d846e
container: vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398e9610d96c473931a26c0e12bbe0d0423
precision: fp4
resources:
gpu_type: mi355x
Expand Down Expand Up @@ -184,3 +183,66 @@ override_tp8_c14_simple:
CONC: '14'
KV_OFFLOADING: 'dram'
TOTAL_CPU_DRAM_GB: '1799'

override_tp8_dcp8_c44_simple:
roles:
agg:
args:
max-num-seqs: 88
max-num-batched-tokens: 8192
compilation-config: '{"mode":3,"cudagraph_mode":"FULL","max_cudagraph_capture_size":88,"custom_ops":["+fused_rms_norm_gated"],"cudagraph_capture_sizes":[2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88]}'
kv-transfer-config: '{"kv_connector":"SimpleCPUOffloadConnector","kv_role":"kv_both","kv_connector_extra_config":{"cpu_bytes_to_use_per_rank":224875000000,"lazy_offload":false}}'
decode-context-parallel-size: 8
dcp-comm-backend: a2a
attention-backend: ROCM_AITER_MLA
env:
VLLM_USE_BREAKABLE_CUDAGRAPH: '0'
PYTHONHASHSEED: '42'
benchmark:
env:
CONC: '44'
KV_OFFLOADING: dram
TOTAL_CPU_DRAM_GB: '1799'
SPEC_DECODING: mtp

override_tp8_dcp8_c48_simple:
roles:
agg:
args:
max-num-seqs: 96
max-num-batched-tokens: 8192
compilation-config: '{"mode":3,"cudagraph_mode":"FULL","max_cudagraph_capture_size":96,"custom_ops":["+fused_rms_norm_gated"],"cudagraph_capture_sizes":[2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96]}'
kv-transfer-config: '{"kv_connector":"SimpleCPUOffloadConnector","kv_role":"kv_both","kv_connector_extra_config":{"cpu_bytes_to_use_per_rank":224875000000,"lazy_offload":false}}'
decode-context-parallel-size: 8
dcp-comm-backend: a2a
attention-backend: ROCM_AITER_MLA
env:
VLLM_USE_BREAKABLE_CUDAGRAPH: '0'
PYTHONHASHSEED: '42'
benchmark:
env:
CONC: '48'
KV_OFFLOADING: dram
TOTAL_CPU_DRAM_GB: '1799'
SPEC_DECODING: mtp

override_tp8_dcp8_c70_simple:
roles:
agg:
args:
max-num-seqs: 140
max-num-batched-tokens: 8192
compilation-config: '{"mode":3,"cudagraph_mode":"FULL","max_cudagraph_capture_size":140,"custom_ops":["+fused_rms_norm_gated"],"cudagraph_capture_sizes":[2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140]}'
kv-transfer-config: '{"kv_connector":"SimpleCPUOffloadConnector","kv_role":"kv_both","kv_connector_extra_config":{"cpu_bytes_to_use_per_rank":224875000000,"lazy_offload":false}}'
decode-context-parallel-size: 8
dcp-comm-backend: a2a
attention-backend: ROCM_AITER_MLA
env:
VLLM_USE_BREAKABLE_CUDAGRAPH: '0'
PYTHONHASHSEED: '42'
benchmark:
env:
CONC: '70'
KV_OFFLOADING: dram
TOTAL_CPU_DRAM_GB: '1799'
SPEC_DECODING: mtp
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ base:
name: minimaxm3-fp4-b200-trtllm-agentic
model:
path: hf:nvidia/MiniMax-M3-NVFP4
container: nvcr.io#nvidia/tensorrt-llm/release:1.3.0rc23.post1
container: nvcr.io#nvidia/tensorrt-llm/release:1.3.0rc28.dev202609220000
precision: fp4
resources:
gpu_type: b200
Expand All @@ -21,9 +21,6 @@ base:
engine:
type: trtllm
served_model_name: nvidia/MiniMax-M3-NVFP4
# Keep Prometheus without rc23's per-step timing collector and accept BFCL's
# store=false field.
setup_script: minimaxm3-trtllm-rc23.sh
# TP8 loads, autotunes and captures graphs past the 30-minute default.
health_check:
max_attempts: 1440
Expand All @@ -33,7 +30,7 @@ base:
nodes: 1
workers: 1
# Server-layer flag; the local checkpoint mounts at /model.
extra_args: [--chat_template, /model/chat_template.jinja]
extra_args: [--chat_template, /model/chat_template.jinja, --tool_parser, minimax_m3]
args:
moe_expert_parallel_size: 1
max_seq_len: 1048576
Expand All @@ -55,11 +52,11 @@ base:
implementation: msa
indexer_kv_dtype: fp8
sparse_disable_index_value: true
fuse_qkv_index_projection: true
kv_cache_config:
free_gpu_memory_fraction: 0.94
enable_block_reuse: true
block_reuse_policy: per_conversation
block_reuse_config:
policy: per_conversation
tokens_per_block: 128
use_kv_cache_manager_v2: true
dtype: fp8
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ base:
name: minimaxm3-fp4-b300-trtllm-agentic
model:
path: hf:nvidia/MiniMax-M3-NVFP4
container: nvcr.io#nvidia/tensorrt-llm/release:1.3.0rc23.post1
container: nvcr.io#nvidia/tensorrt-llm/release:1.3.0rc28.dev202609220000
precision: fp4
resources:
gpu_type: b300
Expand All @@ -25,7 +25,7 @@ base:
agg:
nodes: 1
workers: 1
extra_args: [--chat_template, /model/chat_template.jinja]
extra_args: [--chat_template, /model/chat_template.jinja, --tool_parser, minimax_m3]
args:
moe_expert_parallel_size: 1
max_seq_len: 1048576
Expand All @@ -47,11 +47,11 @@ base:
implementation: msa
indexer_kv_dtype: fp8
sparse_disable_index_value: true
fuse_qkv_index_projection: true
kv_cache_config:
free_gpu_memory_fraction: 0.94
enable_block_reuse: true
block_reuse_policy: per_conversation
block_reuse_config:
policy: per_conversation
tokens_per_block: 128
use_kv_cache_manager_v2: true
dtype: fp8
Expand Down
Loading
Loading