AISpot

b11513: CUDA: improve top-k algorithm selection (#28713)

llama.cpp releases10/8 12:34版本更新地端推論開源

llama.cpp 的 CUDA 後端為 top-k 換了實現:用按行分塊的 radix select 取代 CUB 的逐行 DeviceTopKKernel,並按形狀自動選擇演算法。在 qwen4exp、34,816 token 的基準中,top-k 的 kernel 啟動次數從 1,671,253 次降到 2,329 次,耗時從 5,761.8 ms 降到 941.8 ms。

llama.cpp CUDA 後端的 top-k 改用 radix select 後,官方基準顯示 kernel 啟動次數減少約三個數量級、耗時降到約六分之一,並給出按行列形狀選演算法的閾值,跑本地長上下文推理的人值得關注。

原標題:b11513: CUDA: improve top-k algorithm selection (#28713)
閱讀原文

評分68 / 68(平均 68,門檻 60)
狀態精選

全文翻譯

CUDA:面向大行數的基數 top-k 用按行的網格基數選擇替代 CUB 的逐行 DeviceTopKKernel, 由 GGML_CUDA_TOPK_RADIX_MIN_ROWS 控制。在 qwen4exp 的 34,816 個 token 上,這將 top-k 從 1,671,253 次啟動 / 5,761.8 ms 降至 2,329 次 / 941.8 ms。 CUDA:按形狀選擇 TOP_K 實現 用 #28547 中的決策邊界(如 #29278 中所實現)替換 nrows/ncols 特例: 短行用雙調排序,多個長行用基數選擇,單個長行用 DeviceTopK 或 CUB argsort。 這些閾值在建置時仍可覆蓋。 在該邊界之上還有兩項改進: 當行數適合一波塊(nrows <= SM 數量)時,對於填充後最多 1024 的行仍使用雙調排序; 基數選擇有大約十幾次啟動的固定成本,只有在更多行時才能攤薄 在 DeviceTopK 可用時,它最多處理兩行 基數選擇現在按塊處理行,因此其暫存記憶體保持有界, 雙調排序路徑也保留其分塊。HIP 和 MUSA 保留其先前的 閾值。 在 test-backend-ops 中圍繞雙調/基數交叉點新增效能用例。 CUDA:讓 top-k 註釋不那麼冗長 CUDA:從 supports_op 中移除 TOP_K 寬度限制 CUDA:如果可用,對單行 TOP_K 使用 DeviceTopK CUDA:避免 TOP_K 雙調檢查中的 ncols 溢位 CUDA:在 argsort 和 top-k 之間共享行分塊輔助函式 CUDA:用 int64_t 進行 TOP_K 基數 blocks_per_row 計算 CUDA:將 GGML_CUDA_TOP_K_NROWS_THRESHOLD_DEVICETOPK 重新命名為 GGML_CUDA_TOP_K_NROWS_THRESHOLD CUDA:在雙調和 CUB TOP_K 路徑之間共享一個排序輔助函式 CUDA:更新 TOP_K 的 TODO、閾值和分塊註釋 tests:新增跨越多個行塊的 TOP_K 用例 CUDA:在 TOP_K 基數迴圈中使用 int64_t col,修正閾值註釋 CUDA:將 TOP_K 和 ARGSORT 支援限制為 ne[0] <= INT_MAX 共同作者:praneshgo 227579474+praneshgo@users.noreply.github.com 共同作者:Pranesh Gonegandla pgonegandla@nvidia.com

由 AI 翻譯,以原文為準。

原文
CUDA: radix top-k for large row counts Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select, gated on GGML_CUDA_TOPK_RADIX_MIN_ROWS. On qwen4exp at 34,816 tokens this cuts top-k from 1,671,253 launches / 5,761.8 ms to 2,329 / 941.8 ms. CUDA: select the TOP_K implementation by shape Replace the nrows/ncols special case with the decision boundary from #28547 (as implemented in #29278 ): bitonic for short rows, radix select for several long rows, and DeviceTopK or CUB argsort for a single long row. The thresholds stay overridable at build time. Two refinements on top of that boundary: bitonic stays in use for rows up to a padded 1024 while the rows fit in one wave of blocks (nrows <= number of SMs); radix select pays a fixed cost of about a dozen launches that only amortizes over more rows with DeviceTopK available, it handles up to two rows Radix select now processes rows in chunks so its scratch memory stays bounded, and the bitonic path keeps its chunking. HIP and MUSA keep their previous thresholds. Add perf cases around the bitonic/radix crossover to test-backend-ops. CUDA: make top-k comments less verbose CUDA: remove the TOP_K width limit from supports_op CUDA: use DeviceTopK for single-row TOP_K if available CUDA: avoid ncols overflow in the TOP_K bitonic check CUDA: share the row chunking helper between argsort and top-k CUDA: do the TOP_K radix blocks_per_row math in int64_t CUDA: rename GGML_CUDA_TOP_K_NROWS_THRESHOLD_DEVICETOPK to GGML_CUDA_TOP_K_NROWS_THRESHOLD CUDA: share one sort helper between the bitonic and CUB TOP_K paths CUDA: update the TOP_K TODO, threshold and chunking comments tests: add TOP_K cases that span several row chunks CUDA: use int64_t col in the TOP_K radix loops, fix threshold comment CUDA: limit TOP_K and ARGSORT support to ne[0] <= INT_MAX Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com Co-authored-by: Pranesh Gonegandla pgonegandla@nvidia.com

相關報導

llama.cpp releases10/3 07:28AI 評分42

llama.cpp 最佳化 lightning indexer 視訊記憶體與多後端支援

llama.cpp 合併多項改動,將 indexer 分數視訊記憶體佔用減半:原先同時存在兩個 [n_pool, n_idx_h, n_tokens] f32 張量,現改為逐頭計算並原地累加為單個 [n_pool, n_tokens] 分數。同時 CUDA 支援 4 頭、Metal 以函式常量讀取頭數、Vulkan 按 64 鍵×8 token 分塊並改用 fp16 點積,indexer 改為呼叫 ggml_lightning_indexer 且鍵只讀一次。

llama.cpp releases10/6 14:34AI 評分45

llama.cpp 建置 b11448:ggml-cuda 分塊處理大塊 BF16/FP16 轉 F32

llama.cpp 發布建置 b11448,ggml-cuda 後端把大塊 BF16/FP16 到 F32 的轉換改為分塊執行(#29442),並修正分塊 cuBLAS 矩陣乘法未遵循目標 stride 的問題,提交由 Johannes Gäßler 參與。改動僅涉及 CUDA 後端,影響用 NVIDIA GPU 做本地推理的場景;該建置同時覆蓋 macOS Apple Silicon 與 Linux、Windows 的 CUDA 12/13 等平台。