b11513: CUDA: improve top-k algorithm selection (#28713)
llama.cpp 的 CUDA 後端為 top-k 換了實現:用按行分塊的 radix select 取代 CUB 的逐行 DeviceTopKKernel,並按形狀自動選擇演算法。在 qwen4exp、34,816 token 的基準中,top-k 的 kernel 啟動次數從 1,671,253 次降到 2,329 次,耗時從 5,761.8 ms 降到 941.8 ms。
llama.cpp CUDA 後端的 top-k 改用 radix select 後,官方基準顯示 kernel 啟動次數減少約三個數量級、耗時降到約六分之一,並給出按行列形狀選演算法的閾值,跑本地長上下文推理的人值得關注。
原標題:b11513: CUDA: improve top-k algorithm selection (#28713)
閱讀原文
| 評分 | 68 / 68(平均 68,門檻 60) |
| 狀態 | 精選 |
|---|
全文翻譯
CUDA:面向大行數的基數 top-k
用按行的網格基數選擇替代 CUB 的逐行 DeviceTopKKernel,
由 GGML_CUDA_TOPK_RADIX_MIN_ROWS 控制。在 qwen4exp 的 34,816 個 token 上,這將
top-k 從 1,671,253 次啟動 / 5,761.8 ms 降至 2,329 次 / 941.8 ms。
CUDA:按形狀選擇 TOP_K 實現
用 #28547 中的決策邊界(如 #29278 中所實現)替換 nrows/ncols 特例:
短行用雙調排序,多個長行用基數選擇,單個長行用 DeviceTopK 或 CUB argsort。
這些閾值在建置時仍可覆蓋。
在該邊界之上還有兩項改進:
當行數適合一波塊(nrows <= SM 數量)時,對於填充後最多 1024 的行仍使用雙調排序;
基數選擇有大約十幾次啟動的固定成本,只有在更多行時才能攤薄
在 DeviceTopK 可用時,它最多處理兩行
基數選擇現在按塊處理行,因此其暫存記憶體保持有界,
雙調排序路徑也保留其分塊。HIP 和 MUSA 保留其先前的
閾值。
在 test-backend-ops 中圍繞雙調/基數交叉點新增效能用例。
CUDA:讓 top-k 註釋不那麼冗長
CUDA:從 supports_op 中移除 TOP_K 寬度限制
CUDA:如果可用,對單行 TOP_K 使用 DeviceTopK
CUDA:避免 TOP_K 雙調檢查中的 ncols 溢位
CUDA:在 argsort 和 top-k 之間共享行分塊輔助函式
CUDA:用 int64_t 進行 TOP_K 基數 blocks_per_row 計算
CUDA:將 GGML_CUDA_TOP_K_NROWS_THRESHOLD_DEVICETOPK 重新命名為 GGML_CUDA_TOP_K_NROWS_THRESHOLD
CUDA:在雙調和 CUB TOP_K 路徑之間共享一個排序輔助函式
CUDA:更新 TOP_K 的 TODO、閾值和分塊註釋
tests:新增跨越多個行塊的 TOP_K 用例
CUDA:在 TOP_K 基數迴圈中使用 int64_t col,修正閾值註釋
CUDA:將 TOP_K 和 ARGSORT 支援限制為 ne[0] <= INT_MAX
共同作者:praneshgo 227579474+praneshgo@users.noreply.github.com
共同作者:Pranesh Gonegandla pgonegandla@nvidia.com
由 AI 翻譯,以原文為準。
原文
CUDA: radix top-k for large row counts
Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select,
gated on GGML_CUDA_TOPK_RADIX_MIN_ROWS. On qwen4exp at 34,816 tokens this cuts
top-k from 1,671,253 launches / 5,761.8 ms to 2,329 / 941.8 ms.
CUDA: select the TOP_K implementation by shape
Replace the nrows/ncols special case with the decision boundary from #28547
(as implemented in #29278 ): bitonic for short rows, radix select for several
long rows, and DeviceTopK or CUB argsort for a single long row. The
thresholds stay overridable at build time.
Two refinements on top of that boundary:
bitonic stays in use for rows up to a padded 1024 while the rows fit in one
wave of blocks (nrows <= number of SMs); radix select pays a fixed cost of
about a dozen launches that only amortizes over more rows
with DeviceTopK available, it handles up to two rows
Radix select now processes rows in chunks so its scratch memory stays bounded,
and the bitonic path keeps its chunking. HIP and MUSA keep their previous
thresholds.
Add perf cases around the bitonic/radix crossover to test-backend-ops.
CUDA: make top-k comments less verbose
CUDA: remove the TOP_K width limit from supports_op
CUDA: use DeviceTopK for single-row TOP_K if available
CUDA: avoid ncols overflow in the TOP_K bitonic check
CUDA: share the row chunking helper between argsort and top-k
CUDA: do the TOP_K radix blocks_per_row math in int64_t
CUDA: rename GGML_CUDA_TOP_K_NROWS_THRESHOLD_DEVICETOPK to GGML_CUDA_TOP_K_NROWS_THRESHOLD
CUDA: share one sort helper between the bitonic and CUB TOP_K paths
CUDA: update the TOP_K TODO, threshold and chunking comments
tests: add TOP_K cases that span several row chunks
CUDA: use int64_t col in the TOP_K radix loops, fix threshold comment
CUDA: limit TOP_K and ARGSORT support to ne[0] <= INT_MAX
Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com
Co-authored-by: Pranesh Gonegandla pgonegandla@nvidia.com
相關報導
llama.cpp releases10/8 10:40AI 評分42
llama.cpp 發布 b11505,修復 Vulkan 後端 TOP_K 運算元在 +inf/NaN 輸入和 k=1 負值時的錯誤。原 bucket search 從 [0, 0xFF800000) 起算,+inf 與 NaN 永不被計數,可致 NVIDIA 上掛起或裝置丟失、AMD 上索引錯誤,並漏掉真實 top 值。
llama.cpp releases10/5 00:35AI 評分37
llama.cpp 發布 b11402 版本,CUDA 後端改為優先採用整塊(whole-tile)FlashAttention 排程,以提升兩階段核心效率。
llama.cpp releases10/3 07:28AI 評分42
llama.cpp 合併多項改動,將 indexer 分數視訊記憶體佔用減半:原先同時存在兩個 [n_pool, n_idx_h, n_tokens] f32 張量,現改為逐頭計算並原地累加為單個 [n_pool, n_tokens] 分數。同時 CUDA 支援 4 頭、Metal 以函式常量讀取頭數、Vulkan 按 64 鍵×8 token 分塊並改用 fp16 點積,indexer 改為呼叫 ggml_lightning_indexer 且鍵只讀一次。
r/LocalLLaMA top of the day10/2 14:47AI 評分65
llama.cpp 的 PR #29184 提出在 CUDA 後端把共享專家融合進 MMVQ,為部分 MoE 架構帶來提速,例如 Qwen 35B A3B。有評論稱該改動使 TG 速度提升約 5%,並指出僅對部分 MoE 架構有效。
llama.cpp releases10/6 14:34AI 評分45
llama.cpp 發布建置 b11448,ggml-cuda 後端把大塊 BF16/FP16 到 F32 的轉換改為分塊執行(#29442),並修正分塊 cuBLAS 矩陣乘法未遵循目標 stride 的問題,提交由 Johannes Gäßler 參與。改動僅涉及 CUDA 後端,影響用 NVIDIA GPU 做本地推理的場景;該建置同時覆蓋 macOS Apple Silicon 與 Linux、Windows 的 CUDA 12/13 等平台。