llama.cpp Vulkan 後端加入量化 K/V 稀疏 Flash Attention
llama.cpp 合併 PR #29639,為 Vulkan 後端加入針對量化 K/V 快取的稀疏 Flash Attention,並把稀疏 FA 的索引壓縮改為單次掃描。原實現每行 mask 用一個 workgroup、按 BLOCK_SIZE 分塊各做一次掃描,解碼時單個 workgroup 要跑 KV/1024 次受 barrier 限制的迭代,128k 單元格下開銷超過它服務的稀疏注意力本身。
Vulkan 後端新增量化 K/V 的稀疏 Flash Attention,並給出索引壓縮從逐塊多次掃描改為單次掃描的具體做法,對在非 NVIDIA GPU 上跑長上下文字地推理的人最有用。
原標題:b11413
閱讀原文
| 評分 | 68 / 50(平均 59,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
vulkan: sparse flash attention for quantized K/V ( #29639 )
vulkan: sparse flash attention for quantized K/V
Assisted-by: Claude
vulkan: single-scan sparse FA index compaction
The compaction ran one workgroup per mask row and walked the row in
BLOCK_SIZE chunks, with a workgroup scan per chunk. For decode that is
one workgroup doing KV/1024 barrier-bound iterations, so at 128k cells
it cost more than the sparse attention it feeds.
Split the row into contiguous segments instead: one per subgroup with
ballot counting over coalesced loads, or one per thread without
subgroups. A single scan over the segment counts then gives each
segment its output offset. The index list stays ascending.
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52797762
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases● 精選10/3 01:28AI 評分68
llama.cpp 合併 PR #29570,在 Metal 後端為 F16 KV 快取加入基於張量 API 的 flash attention 核心。該核心覆蓋 DK=DV=512、DK=576/DV=512、DK=192/DV=128 等配置,並支援 attention sinks、ALiBi 與 logit softcap。改動隨 b11362 建置發布,覆蓋 macOS Apple Silicon、iOS 及 Linux、Windows、Android 等多平台建置。
llama.cpp releases10/5 12:46AI 評分36
llama.cpp 發布建置 b11414,本版列出的程式碼改動是 Vulkan 後端修復 prealloc_y 在 flash attention 與 soft_max 之間複用時殘留舊資料的問題(PR #29591),提交標註 Assisted-by: Claude。
llama.cpp releases● 精選10/5 16:55AI 評分68
llama.cpp 的 Hexagon 後端在 row-split 多核模式下把 flash_attn 從按 Q token 切分改為按 KV head 切分,讓每個核心只讀自己負責的那部分 KV cache,不再重複讀取全部 KV cache。該行為由 GGML_HEXAGON_FA_HEAD_SPLIT 控制,預設開啟;當 n_kv_heads 不能被核心數整除時回退到原 token-block 切分(如 4 核上只有 2 個 KV head 的 Gemma-4)。
llama.cpp releases10/5 15:50AI 評分42
llama.cpp 發布 b11424 建置版本,修復 Vulkan 後端 Flash Attention 的共享記憶體越界寫問題(#29988)。該版本照例提供 macOS/iOS、Linux、Windows、Android 的預編譯包,涵蓋 Vulkan、CUDA 12/13、ROCm 10.0、OpenVINO、SYCL、OpenCL 等後端,並附驍龍 CPU/Adreno GPU/Hexagon NPU 的安裝指引。
llama.cpp releases10/3 00:18AI 評分31
llama.cpp 合併 PR #28531,在共享記憶體為 32KB 的三星 GPU 上停用 Vulkan 後端的大矩陣乘分塊(large matmul tile),以規避該硬體配置下的問題。該改動由 Claude Opus 輔助完成,並附有建置證明連結。