AISpot
所屬事件:llama.cpp 合併 PR:Qwen 索引器視訊記憶體減半(2 個來源 · 2 則報導)

2026 年 10 月 3 日,llama.cpp 合併了 PR #29825(作者 ServeurpersoCom),把 Qwen Flash Next 的 lightning indexer 打分視訊記憶體佔用減半,隨版本 b11372 發布。據 PR 說明,此前索引器一次性算出所有 head 的分數並保留一份副本… 看完整事件

llama.cpp 最佳化 lightning indexer 視訊記憶體與多後端支援

llama.cpp releases10/3 12:28版本更新開發工具地端推論開源

llama.cpp 合併了 qwen4exp 的 lightning indexer 最佳化,把索引器打分視訊記憶體減半:原先同時存在兩個 [n_pool, n_idx_h, n_tokens] f32 張量,現在每個 head 單獨計算並原地累加進一個 [n_pool, n_tokens] 分數。同時 CUDA 支援 4 heads、Metal 用函式常量傳 head 數、Vulkan 按 keys×tokens 分塊並向量化 fp16 點積。

你在 Mac Studio 上跑本地模型時,長上下文視訊記憶體佔用和 Metal/Vulkan 後端效能會直接受益,升級 llama.cpp 即可獲得這些最佳化。

原標題:b11372
閱讀原文

同一事件共有 2 則報導(2 個來源),看事件全貌

評分62 / 82(平均 72,門檻 60)
狀態精選

全文翻譯

qwen4exp:將索引器分數記憶體減半(#29825) qwen4exp:將索引器分數記憶體減半 索引器在一個乘積中對所有頭進行評分,並對其副本進行修正, 因此兩個 [n_pool, n_idx_h, n_tokens] f32 張量同時存在, 這是長上下文下圖中最大的緩衝區。現在每個頭獲得 自己的乘積,修正並就地求和為一個 [n_pool, n_tokens] 分數。 qwen4exp:讓分配器複用索引器分數緩衝區 處理來自 CISC 的審查:在索引器頭迴圈中使用普通的 ggml_add 和 ggml_relu。 當它們的源沒有其他消費者時,圖分配器已經就地執行它們, 因此不需要 _inplace 變體。計算緩衝區和速度不變。 cuda:在閃電索引器中支援 4 個頭 將 4 個頭也分派到向量核心,對於 wmma 瓦片來說太少, 並在 supports_op 中接受它們。test-backend-ops 覆蓋 4 個頭。 metal:將閃電索引器頭數作為函式常量 核心從函式常量讀取頭數,並將最後一個頭瓦片零填充, 因此任何頭數都能執行,64 個頭保持不變。 qwen4exp:使用閃電索引器計算索引器分數 處理來自 am17an 的審查:修正後的頭分數的未加權和 按 1/sqrt(head_dim) 縮放,就是每個頭權重都設為該縮放的閃電索引器, 因此索引器在池化鍵上呼叫 ggml_lightning_indexer,並使用 f16 池掩碼。 鍵對所有頭只讀取一次,不物化每個頭的分數。 vulkan:將閃電索引器在鍵和令牌上分塊 一個工作組對 64 個鍵與 8 個令牌進行評分:鍵在共享記憶體中暫存一次, 查詢一次一個頭,每個呼叫擁有一個鍵用於兩個令牌, 因此任何點積都不需要跨呼叫歸約。子組變體和扁平分派已移除, 網格為鍵 x 令牌 x 流。 向量化 vulkan 載入並使用 fp16 點積 共同作者:Ruben Ortlam rortlam@redhat.com 網站: https://llama.app 證明: https://github.com/ggml-org/llama.cpp/attestations/52404802 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

由 AI 翻譯,以原文為準。

原文
qwen4exp : halve the indexer score memory ( #29825 ) qwen4exp : halve the indexer score memory The indexer scored all heads in one product and rectified a copy of it, so two [n_pool, n_idx_h, n_tokens] f32 tensors were live at once, the largest buffers of the graph at long context. Each head now gets its own product, rectified and summed in place into one [n_pool, n_tokens] score. qwen4exp: let the allocator reuse the indexer score buffers Address review from CISC: use plain ggml_add and ggml_relu in the indexer head loop. The graph allocator already runs them in place when their source has no other consumer, so the _inplace variants are not needed. The compute buffer and the speed are unchanged. cuda: support 4 heads in the lightning indexer Dispatch 4 heads to the vector kernel, too few for a wmma tile, and accept them in supports_op. test-backend-ops covers 4 heads. metal: take the lightning indexer head count as a function constant The kernel reads the head count from a function constant and zero fills the last head tile, so any head count runs and 64 heads is unchanged. qwen4exp: compute the indexer score with the lightning indexer Address review from am17an: the unweighted sum of the rectified head scores scaled by 1/sqrt(head_dim) is the lightning indexer with every head weight set to that scale, so the indexer calls ggml_lightning_indexer on the pooled keys with an f16 pool mask. The keys are read once for all heads and no per head score is materialized. vulkan: tile the lightning indexer over keys and tokens A workgroup scores 64 keys against 8 tokens: the keys are staged once in shared memory, the queries one head at a time, and each invocation owns one key for two tokens, so no dot product needs a cross invocation reduction. The subgroup variant and the flat dispatch are gone, the grid is keys x tokens x streams. vectorize vulkan loads and use fp16 dot product Co-authored-by: Ruben Ortlam rortlam@redhat.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52404802 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI