Skip to content

[None][perf] Enable tiled FP8 context FMHA for head_dim 256 on SM120 - #19825

Open
amukkara wants to merge 1 commit into
NVIDIA:mainfrom
amukkara:fmha-sm120-tile
Open

amukkara wants to merge 1 commit into
NVIDIA:mainfrom
amukkara:fmha-sm120-tile

Conversation

@amukkara

@amukkara amukkara commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Use tiled fmha_v2 kernel for FP8 (E4M3) context attention with head_dim 256 on SM120/SM121. Targets nvidia/Qwen3.6-35B-A3B-NVFP4 which uses FP8 KV cache.

  • cpp/kernels/fmha_v2/setup.py: add a tiled 64x128 QMMA kernel for head_dim 256, alongside the existing non-tiled 64x32 kernel.
  • fmhaRunner.cpp: restrict the forced non-tiled FP8 selection to SM89, so SM120/SM121 use the existing head-size heuristic, which selects tiled kernels for head_dim >= 256.

Performance

fmha_v2 harness on RTX PRO 6K BSE, FP8 Q/K/V, head_dim 256, tiled vs non-tiled kernel.
The tiled kernel is faster on all 35 shapes tested: 2.03x - 3.52x, geomean 3.01x.

sweep range speedup
KV length (b=1, 16 Q / 4 KV heads) 65 -> 32k 2.15x -> 3.32x
batch size 4 -> 64 at KV 256 / 1k / 4k 2.03x -> 3.24x
head config 8/2, 32/8, 32/32 2.93x -> 3.23x
chunked prefill query chunk 128 / 512, KV 2k / 8k 3.38x -> 3.43x
packed QKV, padding mask, FP8 output KV 1k / 4k 2.99x -> 3.52x

Test Coverage

The tiled kernel passes the fmha_v2 harness reference check on SM120 (-epsilon 0.2) for paged KV and packed QKV; causal, padding, sliding-window and custom masks; ALiBi; GQA; chunked prefill; and FP8 and BF16 output.

e2e accuracy verified by tests/integration/defs/accuracy/test_llm_api_pytorch.py::TestQwen3_6_35B_A3B::test_nvfp4_w4a16

Dev Engineer Review

cpp/kernels/fmha_v2/setup.py removes unnecessary f-string prefixes from two constant strings. This does not change generated output or kernel selection. The supplied diff does not support the stated tiled-kernel or fmhaRunner.cpp changes.

QA Engineer Review

No test changes.

Per-File QA Perspective

  • cpp/kernels/fmha_v2/setup.py: The generated #endif text remains unchanged. No observable behavior change requires QA verification.

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

Signed-off-by: Anurag Mukkara <[email protected]>
@amukkara
amukkara marked this pull request as ready for review October 2, 2026 23:58
@amukkara
amukkara requested a review from a team as a code owner October 2, 2026 23:58
@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/TensorRT-LLM/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: e58ae387-4aac-4040-9c86-897acf072b3e
📥 Commits

Reviewing files that changed from the base of the PR and between f388b7c and 54f2dbc.

📒 Files selected for processing (2)
  • cpp/kernels/fmha_v2/setup.py
  • cpp/tensorrt_llm/kernels/contextFusedMultiHeadAttention/fmhaRunner.cpp

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


Walkthrough

The FMHA kernel enumerator adds an SM120 tiled configuration for head size 256. The runner excludes SM120/SM121 from the E4M3 non-tiled-kernel branch. Both files update their copyright year ranges through 2026.

Changes

SM120 FMHA kernel selection

Layer / File(s) Summary
Tiled kernel configuration and selection
cpp/kernels/fmha_v2/setup.py, cpp/tensorrt_llm/kernels/contextFusedMultiHeadAttention/fmhaRunner.cpp
The enumerator adds an SM120 tiled configuration for head size 256 with Q/KV loop steps 64/128. The runner limits the E4M3 non-tiled-kernel branch to SM89.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Feature

Suggested reviewers: brnguyen2

Merge Risk: ⚪ Minimal · up to 54f2d

No identified kernel-selection issue remains; this change is mergeable after normal checks.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 66.67% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main change: enabling tiled FP8 context FMHA for head_dim 256 on SM120. It is concise and uses the required ticket and type format.
Description check ✅ Passed The description explains the change and its purpose, reports performance results, lists relevant test coverage, and includes the repository checklist with the review box checked.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant