所屬事件:llama.cpp 合併 PR:Qwen 索引器視訊記憶體減半(2 個來源 · 2 則報導)
2026 年 10 月 3 日,llama.cpp 合併了 PR #29825(作者 ServeurpersoCom),把 Qwen Flash Next 的 lightning indexer 打分視訊記憶體佔用減半,隨版本 b11372 發布。據 PR 說明,此前索引器一次性算出所有 head 的分數並保留一份副本… 看完整事件
llama.cpp 最佳化 lightning indexer 視訊記憶體與多後端支援
llama.cpp 合併了 qwen4exp 的 lightning indexer 最佳化,把索引器打分視訊記憶體減半:原先同時存在兩個 [n_pool, n_idx_h, n_tokens] f32 張量,現在每個 head 單獨計算並原地累加進一個 [n_pool, n_tokens] 分數。同時 CUDA 支援 4 heads、Metal 用函式常量傳 head 數、Vulkan 按 keys×tokens 分塊並向量化 fp16 點積。
你在 Mac Studio 上跑本地模型時,長上下文視訊記憶體佔用和 Metal/Vulkan 後端效能會直接受益,升級 llama.cpp 即可獲得這些最佳化。
原標題:b11372
閱讀原文
| 評分 | 62 / 82(平均 72,門檻 60) |
|---|---|
| 狀態 | 精選 |
全文翻譯
qwen4exp:將索引器分數記憶體減半(#29825)
qwen4exp:將索引器分數記憶體減半
索引器在一個乘積中對所有頭進行評分,並對其副本進行修正,
因此兩個 [n_pool, n_idx_h, n_tokens] f32 張量同時存在,
這是長上下文下圖中最大的緩衝區。現在每個頭獲得
自己的乘積,修正並就地求和為一個 [n_pool, n_tokens]
分數。
qwen4exp:讓分配器複用索引器分數緩衝區
處理來自 CISC 的審查:在索引器頭迴圈中使用普通的 ggml_add 和 ggml_relu。
當它們的源沒有其他消費者時,圖分配器已經就地執行它們,
因此不需要 _inplace 變體。計算緩衝區和速度不變。
cuda:在閃電索引器中支援 4 個頭
將 4 個頭也分派到向量核心,對於 wmma 瓦片來說太少,
並在 supports_op 中接受它們。test-backend-ops 覆蓋 4 個頭。
metal:將閃電索引器頭數作為函式常量
核心從函式常量讀取頭數,並將最後一個頭瓦片零填充,
因此任何頭數都能執行,64 個頭保持不變。
qwen4exp:使用閃電索引器計算索引器分數
處理來自 am17an 的審查:修正後的頭分數的未加權和
按 1/sqrt(head_dim) 縮放,就是每個頭權重都設為該縮放的閃電索引器,
因此索引器在池化鍵上呼叫 ggml_lightning_indexer,並使用 f16 池掩碼。
鍵對所有頭只讀取一次,不物化每個頭的分數。
vulkan:將閃電索引器在鍵和令牌上分塊
一個工作組對 64 個鍵與 8 個令牌進行評分:鍵在共享記憶體中暫存一次,
查詢一次一個頭,每個呼叫擁有一個鍵用於兩個令牌,
因此任何點積都不需要跨呼叫歸約。子組變體和扁平分派已移除,
網格為鍵 x 令牌 x 流。
向量化 vulkan 載入並使用 fp16 點積
共同作者:Ruben Ortlam rortlam@redhat.com
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/52404802
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
由 AI 翻譯,以原文為準。
原文
qwen4exp : halve the indexer score memory ( #29825 )
qwen4exp : halve the indexer score memory
The indexer scored all heads in one product and rectified a copy of it,
so two [n_pool, n_idx_h, n_tokens] f32 tensors were live at once, the
largest buffers of the graph at long context. Each head now gets its
own product, rectified and summed in place into one [n_pool, n_tokens]
score.
qwen4exp: let the allocator reuse the indexer score buffers
Address review from CISC: use plain ggml_add and ggml_relu in the
indexer head loop. The graph allocator already runs them in place when
their source has no other consumer, so the _inplace variants are not
needed. The compute buffer and the speed are unchanged.
cuda: support 4 heads in the lightning indexer
Dispatch 4 heads to the vector kernel, too few for a wmma tile, and
accept them in supports_op. test-backend-ops covers 4 heads.
metal: take the lightning indexer head count as a function constant
The kernel reads the head count from a function constant and zero fills
the last head tile, so any head count runs and 64 heads is unchanged.
qwen4exp: compute the indexer score with the lightning indexer
Address review from am17an: the unweighted sum of the rectified head
scores scaled by 1/sqrt(head_dim) is the lightning indexer with every
head weight set to that scale, so the indexer calls
ggml_lightning_indexer on the pooled keys with an f16 pool mask. The
keys are read once for all heads and no per head score is
materialized.
vulkan: tile the lightning indexer over keys and tokens
A workgroup scores 64 keys against 8 tokens: the keys are staged once
in shared memory, the queries one head at a time, and each invocation
owns one key for two tokens, so no dot product needs a cross invocation
reduction. The subgroup variant and the flat dispatch are gone, the grid
is keys x tokens x streams.
vectorize vulkan loads and use fp16 dot product
Co-authored-by: Ruben Ortlam rortlam@redhat.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52404802
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI