AISpot

llama.cpp SYCL 為 GLM MLA 預填充接入 MKL,pp8192 提速 2.99 倍

llama.cpp releases10/7 11:43版本更新地端推論開源

llama.cpp 的 SYCL 後端放行了 GLM-4.7 Flash 的 MLA 形狀走 MKL flash attention,此前該形狀因 K/V 寬度不一致、head dim 上限 512 被排程器拒絕,預填充只能回退到更慢的 TILE kernel。在 Intel Arc Pro B70 上,pp8192 從 432.80 提升到 1292.29 tok/s(2.99 倍,+198.6%),tg256 在噪聲範圍內基本不變(45.67 對 45.65)。

給出了 Intel GPU 上 GLM MLA 預填充 2.99 倍加速的實測數字與具體放行條件,對在 Arc 顯示卡上用 llama.cpp 跑 GLM-4.7 Flash 的人最有用。

原標題:b11463
閱讀原文

評分68 / 68(平均 68,門檻 60)
狀態精選

全文翻譯

sycl:使用 MKL flash attention 加速 GLM MLA prefill(#29171) sycl:使用 MKL flash attention 加速 GLM MLA prefill GLM-4.7 Flash 使用一種 MLA 形狀:Q/K 頭寬 576、V 頭寬 512、 GQA 為 20、KV 為 F16。SYCL 排程器拒絕該形狀, 因為常規的 MKL flash-attention 門控要求 K/V 寬度匹配, 並將頭維度上限設為 512,因此提示詞處理會 回退到明顯更慢的 TILE 核心。 僅將已驗證的 576/576/512、GQA-20 F16 形狀接入現有的 MKL 流水線。讓所有其他不匹配的 K/V 形狀繼續留在其當前的 回退路徑上。 將 GLM 的 V 快取作為 K 行的更窄的跨步檢視處理。當行跨步 被填充時,選擇跨步的 F16 描述符,並且僅當 K/V 反量化緩衝區 的邏輯寬度匹配時才對其取別名。將跨步例外限制為 真正共享 K 行跨步的 V 檢視。 新增精確的 576/512、GQA-20 提示詞路徑後端測試。 在 master e613ef2 上的 Intel Arc Pro B70 上,pp8192 從 432.80 提升到 1292.29 tok/s(2.99 倍,+198.6%)。tg256 在噪聲範圍內 保持不變,為 45.67 對 45.65 tok/s。精確的 MLA 測試通過, 除錯輸出確認了 MKL 排程。 sycl:將 MKL flash attention 分數以 F16 儲存 將 QK GEMM 輸出保持為 F16 而不是 F32。線上 softmax 仍會 把每個分數轉換為 F32 以計算其最大值、指數和求和,因此 除分數舍入外,逐元素的數學運算沒有變化,而且 F32 矩陣此前只是被寫入,隨後又作為 F16 機率被消費。 F32 分數矩陣是該路徑上最大的 flash-attention 中間量; 將其儲存為 F16 可使其大小和流量減半。這建立在 1aa2954 的合併 softmax 載入之上,後者以協作方式讀取每個分數行, 因此更小的資料型別帶來了收益。 在 Intel Arc Pro B70 上使用前一個提交的排程進行測量, 引數為 -ngl 999 -b 4096 -ub 1024 -ctk f16 -ctv f16 -fa on: pp8192 1583.9 -> 1657.1 tok/s(+4.6%) pp64000 610.0 -> 684.5 tok/s(+12.2%) pp131072 ~354 -> 402.1 tok/s(+13.5%) 8k 上下文下的 tg256 保持不變(32.64),並且 FLASH_ATTN_EXT 測試套件 未顯示新的失敗。精確的 GLM MLA 後端用例在 CPU 上通過。 如果你更傾向於只引用成對的實測資料,請調整 ~354 這一基線數字 (131k 僅排程點來自等效的維護建置)。可選:根據貢獻指南 新增 Assisted-by:,因為 AI 對該變更有所貢獻。 回滾“sycl:將 MKL flash attention 分數以 F16 儲存” 本次回滾提交 265f974。 網站: https://llama.app 證明: https://github.com/ggml-org/llama.cpp/attestations/53529269 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, 啟用 KleidiAI) 已停用 macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 庫 Ubuntu x64 (CUDA 13) - CUDA 13.4 庫 Ubuntu arm64 (CUDA 13) - CUDA 13.4 庫 Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安裝指南 Android: Android arm64 (CPU) Android arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安裝指南 Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLL Windows x64 (CUDA 13) - CUDA 13.4 DLL Windows arm64 (CUDA 13) - CUDA 13.4 DLL Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: 已停用 openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

由 AI 翻譯,以原文為準。

原文
sycl: accelerate GLM MLA prefill with MKL flash attention ( #29171 ) sycl: accelerate GLM MLA prefill with MKL flash attention GLM-4.7 Flash uses an MLA shape with 576-wide Q/K heads, a 512-wide V head, GQA 20, and F16 KV. The SYCL dispatcher rejects this shape because the normal MKL flash-attention gate requires matching K/V widths and caps the head dimension at 512, so prompt processing falls back to the substantially slower TILE kernel. Admit only the validated 576/576/512, GQA-20 F16 shape to the existing MKL pipeline. Keep all other mismatched K/V shapes on their current fallback paths. Handle GLM's V cache as a narrower strided view of K rows. Select the strided F16 descriptor when row stride is padded, and alias K/V dequantization buffers only when their logical widths match. Restrict the stride exception to a real V view sharing K's row stride. Add the exact 576/512, GQA-20 prompt-path backend test. On an Intel Arc Pro B70 at master e613ef2 , pp8192 improves from 432.80 to 1292.29 tok/s (2.99x, +198.6%). tg256 remains unchanged within noise at 45.67 versus 45.65 tok/s. The exact MLA test passes and debug output confirms MKL dispatch. sycl: store MKL flash attention scores in F16 Keep the QK GEMM output in F16 instead of F32. The online softmax still converts each score to F32 for its max, exponent, and sum, so the per-element math is unchanged apart from score rounding, and the F32 matrix was being written only to be consumed as F16 probabilities. The F32 score matrix is the largest flash-attention intermediate on this path; storing it as F16 halves its size and traffic. This builds on the coalesced softmax loads from 1aa2954 , which read each score row cooperatively, so the smaller dtype pays off. Measured on an Intel Arc Pro B70 with the dispatch from the previous commit, -ngl 999 -b 4096 -ub 1024 -ctk f16 -ctv f16 -fa on: pp8192 1583.9 -> 1657.1 tok/s (+4.6%) pp64000 610.0 -> 684.5 tok/s (+12.2%) pp131072 ~354 -> 402.1 tok/s (+13.5%) tg256 at 8k context is unchanged (32.64), and the FLASH_ATTN_EXT suite shows no new failures. The exact GLM MLA backend cases pass against CPU. Adjust the ~354 baseline figure if you prefer citing only measured pairs (the 131k dispatch-only point came from the equivalent maintained build). Optionally add Assisted-by: per the contribution guidelines since AI contributed to the change. Revert "sycl: store MKL flash attention scores in F16" This reverts commit 265f974 . Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/53529269 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/7 12:11AI 評分42

llama.cpp 修復 SYCL 混用不同型號 GPU 的 FA 問題

llama.cpp 合併 PR #29071,修復 SYCL 後端在 Flash Attention(FA)中混用不同型號 GPU 時出現的問題。改動位於 ggml/src/ggml-sycl/ggml-sycl.cpp,提交由 Georgi Gerganov 共同署名。該修復包含在版本 b11464 中,此版本同時提供 Ubuntu x64(SYCL FP32、FP16)與 Windows x64(SYCL)等建置。

llama.cpp releases10/4 21:07AI 評分45

llama.cpp 為 x86 tinyBLAS 加入 BF16/FP16/FP32 K 尾部向量化

llama.cpp 的 ggml-cpu 後端在 x86 上為 tinyBLAS 增加了 BF16、FP16、FP32 的 K 尾部支援並做了向量化,同時讓 CPU 測試在啟用 use_ref 時跳過 tinyBLAS,以便與 vec_dot 路徑對比。改動以 PR #29806 合入,隨 b11398 建置發布,覆蓋 macOS、Linux、Windows、Android 等多平台 CPU 與 GPU 後端。