llama.cpp SYCL 為 GLM MLA 預填充接入 MKL,pp8192 提速 2.99 倍
llama.cpp 的 SYCL 後端放行了 GLM-4.7 Flash 的 MLA 形狀走 MKL flash attention,此前該形狀因 K/V 寬度不一致、head dim 上限 512 被排程器拒絕,預填充只能回退到更慢的 TILE kernel。在 Intel Arc Pro B70 上,pp8192 從 432.80 提升到 1292.29 tok/s(2.99 倍,+198.6%),tg256 在噪聲範圍內基本不變(45.67 對 45.65)。
給出了 Intel GPU 上 GLM MLA 預填充 2.99 倍加速的實測數字與具體放行條件,對在 Arc 顯示卡上用 llama.cpp 跑 GLM-4.7 Flash 的人最有用。
原標題:b11463
閱讀原文
| 評分 | 68 / 68(平均 68,門檻 60) |
| 狀態 | 精選 |
|---|
全文翻譯
sycl:使用 MKL flash attention 加速 GLM MLA prefill(#29171)
sycl:使用 MKL flash attention 加速 GLM MLA prefill
GLM-4.7 Flash 使用一種 MLA 形狀:Q/K 頭寬 576、V 頭寬 512、
GQA 為 20、KV 為 F16。SYCL 排程器拒絕該形狀,
因為常規的 MKL flash-attention 門控要求 K/V 寬度匹配,
並將頭維度上限設為 512,因此提示詞處理會
回退到明顯更慢的 TILE 核心。
僅將已驗證的 576/576/512、GQA-20 F16 形狀接入現有的
MKL 流水線。讓所有其他不匹配的 K/V 形狀繼續留在其當前的
回退路徑上。
將 GLM 的 V 快取作為 K 行的更窄的跨步檢視處理。當行跨步
被填充時,選擇跨步的 F16 描述符,並且僅當 K/V 反量化緩衝區
的邏輯寬度匹配時才對其取別名。將跨步例外限制為
真正共享 K 行跨步的 V 檢視。
新增精確的 576/512、GQA-20 提示詞路徑後端測試。
在 master e613ef2 上的 Intel Arc Pro B70 上,pp8192 從
432.80 提升到 1292.29 tok/s(2.99 倍,+198.6%)。tg256 在噪聲範圍內
保持不變,為 45.67 對 45.65 tok/s。精確的 MLA 測試通過,
除錯輸出確認了 MKL 排程。
sycl:將 MKL flash attention 分數以 F16 儲存
將 QK GEMM 輸出保持為 F16 而不是 F32。線上 softmax 仍會
把每個分數轉換為 F32 以計算其最大值、指數和求和,因此
除分數舍入外,逐元素的數學運算沒有變化,而且 F32
矩陣此前只是被寫入,隨後又作為 F16 機率被消費。
F32 分數矩陣是該路徑上最大的 flash-attention 中間量;
將其儲存為 F16 可使其大小和流量減半。這建立在
1aa2954 的合併 softmax 載入之上,後者以協作方式讀取每個分數行,
因此更小的資料型別帶來了收益。
在 Intel Arc Pro B70 上使用前一個提交的排程進行測量,
引數為 -ngl 999 -b 4096 -ub 1024 -ctk f16 -ctv f16 -fa on:
pp8192 1583.9 -> 1657.1 tok/s(+4.6%)
pp64000 610.0 -> 684.5 tok/s(+12.2%)
pp131072 ~354 -> 402.1 tok/s(+13.5%)
8k 上下文下的 tg256 保持不變(32.64),並且 FLASH_ATTN_EXT 測試套件
未顯示新的失敗。精確的 GLM MLA 後端用例在 CPU 上通過。
如果你更傾向於只引用成對的實測資料,請調整 ~354 這一基線數字
(131k 僅排程點來自等效的維護建置)。可選:根據貢獻指南
新增 Assisted-by:,因為 AI 對該變更有所貢獻。
回滾“sycl:將 MKL flash attention 分數以 F16 儲存”
本次回滾提交 265f974。
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/53529269
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, 啟用 KleidiAI) 已停用
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 庫
Ubuntu x64 (CUDA 13) - CUDA 13.4 庫
Ubuntu arm64 (CUDA 13) - CUDA 13.4 庫
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安裝指南
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 安裝指南
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLL
Windows x64 (CUDA 13) - CUDA 13.4 DLL
Windows arm64 (CUDA 13) - CUDA 13.4 DLL
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
已停用
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
由 AI 翻譯,以原文為準。
原文
sycl: accelerate GLM MLA prefill with MKL flash attention ( #29171 )
sycl: accelerate GLM MLA prefill with MKL flash attention
GLM-4.7 Flash uses an MLA shape with 576-wide Q/K heads, a 512-wide
V head, GQA 20, and F16 KV. The SYCL dispatcher rejects this shape
because the normal MKL flash-attention gate requires matching K/V
widths and caps the head dimension at 512, so prompt processing falls
back to the substantially slower TILE kernel.
Admit only the validated 576/576/512, GQA-20 F16 shape to the existing
MKL pipeline. Keep all other mismatched K/V shapes on their current
fallback paths.
Handle GLM's V cache as a narrower strided view of K rows. Select the
strided F16 descriptor when row stride is padded, and alias K/V
dequantization buffers only when their logical widths match. Restrict
the stride exception to a real V view sharing K's row stride.
Add the exact 576/512, GQA-20 prompt-path backend test.
On an Intel Arc Pro B70 at master e613ef2 , pp8192 improves from
432.80 to 1292.29 tok/s (2.99x, +198.6%). tg256 remains unchanged
within noise at 45.67 versus 45.65 tok/s. The exact MLA test passes
and debug output confirms MKL dispatch.
sycl: store MKL flash attention scores in F16
Keep the QK GEMM output in F16 instead of F32. The online softmax still
converts each score to F32 for its max, exponent, and sum, so the
per-element math is unchanged apart from score rounding, and the F32
matrix was being written only to be consumed as F16 probabilities.
The F32 score matrix is the largest flash-attention intermediate on this
path; storing it as F16 halves its size and traffic. This builds on the
coalesced softmax loads from 1aa2954 , which read each score row
cooperatively, so the smaller dtype pays off.
Measured on an Intel Arc Pro B70 with the dispatch from the previous
commit, -ngl 999 -b 4096 -ub 1024 -ctk f16 -ctv f16 -fa on:
pp8192 1583.9 -> 1657.1 tok/s (+4.6%)
pp64000 610.0 -> 684.5 tok/s (+12.2%)
pp131072 ~354 -> 402.1 tok/s (+13.5%)
tg256 at 8k context is unchanged (32.64), and the FLASH_ATTN_EXT suite
shows no new failures. The exact GLM MLA backend cases pass against CPU.
Adjust the ~354 baseline figure if you prefer citing only measured pairs (the 131k dispatch-only point came from the equivalent maintained build). Optionally add Assisted-by: per the contribution guidelines since AI contributed to the change.
Revert "sycl: store MKL flash attention scores in F16"
This reverts commit 265f974 .
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53529269
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/7 12:11AI 評分42
llama.cpp 合併 PR #29071,修復 SYCL 後端在 Flash Attention(FA)中混用不同型號 GPU 時出現的問題。改動位於 ggml/src/ggml-sycl/ggml-sycl.cpp,提交由 Georgi Gerganov 共同署名。該修復包含在版本 b11464 中,此版本同時提供 Ubuntu x64(SYCL FP32、FP16)與 Windows x64(SYCL)等建置。
llama.cpp releases10/5 05:35AI 評分37
llama.cpp 發布 b11402 版本,CUDA 後端改為優先採用整塊(whole-tile)FlashAttention 排程,以提升兩階段核心效率。
llama.cpp releases10/7 07:02AI 評分37
llama.cpp 發布建置 b11460,SYCL 後端新增 IQ3_S 的多列 MMVQ 支援(PR #29500)。該版本繼續提供覆蓋 macOS/iOS、Linux、Android、Windows 與 openEuler 的預編譯包,其中 Ubuntu x64 與 Windows x64 包含 SYCL FP32、FP16 建置。
llama.cpp releases10/4 21:07AI 評分45
llama.cpp 的 ggml-cpu 後端在 x86 上為 tinyBLAS 增加了 BF16、FP16、FP32 的 K 尾部支援並做了向量化,同時讓 CPU 測試在啟用 use_ref 時跳過 tinyBLAS,以便與 vec_dot 路徑對比。改動以 PR #29806 合入,隨 b11398 建置發布,覆蓋 macOS、Linux、Windows、Android 等多平台 CPU 與 GPU 後端。
llama.cpp releases10/7 11:13AI 評分35
llama.cpp 發布 b11462 建置,本次合併的主要改動是 SYCL 後端的 fattn_kv_buffers 清理(#27689)。該版本繼續提供 macOS/iOS、Linux、Android、Windows、openEuler 的預編譯包,覆蓋 CUDA 12/13、ROCm 10.0、Vulkan、OpenVINO、SYCL,以及驍龍平台的 CPU、Adreno GPU 與 Hexagon NPU。