AISpot

llama.cpp 修復共享序列 k-pool 資料競爭

llama.cpp releases10/6 07:18版本更新地端推論開源開發工具

llama.cpp 合併 PR #29994,修復 CPU 後端在共享序列(seq_cp)場景下 k-pool 的散射寫入資料競爭:多個共享同一 cell 的序列會從不同 scatter 條目寫同一 rep 行。

原標題:b11435
閱讀原文

評分45 / 50(平均 47,門檻 60)
狀態未入選

原文

llama: fix k-pool scatter data race on shared sequences ( #29994 ) llama: re-pool each shared k-pool rep once With shared cells every pool is re-pooled, and since the pooled keys are always scattered, the pools a seq_cp shares between sequences wrote the same rep row from several scatter entries, a data race on the CPU backend. Mark each rep once: the sharing sequences read the same row through pool_cells. llama: assert whole-sequence seq_cp in the hybrid idx memory The recurrent state is always copied whole whatever the range, and a k-pool cell shared by a partial copy could carry two pool groupings with a single pooled row. Every caller copies whole sequences, so reject partial ranges instead of supporting them. llama: drop the k-pool cache_safe mode With whole-sequence seq_cp, sequences sharing cells share their pools too, so the pooled row of a shared rep is valid for all of them. Mark each rep once in every ubatch instead of re-pooling everything while cells are shared, which removes the sharing scan and the stale-all workarounds in seq_rm, state_read and state_drop. seq_cp now only stales the destination. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/53070668 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/5 09:36AI 評分48

llama.cpp 修復 k-pool 模型圖重分配導致的解碼中止

llama.cpp 發布 b11412,修復 k-pool 模型(qwen4exp、glm5-next)解碼時意外重新預留計算圖並中止的問題。原因是兩模型按 cache_safe、n_tokens 等 reserve 無法預知的狀態分支,解碼圖與預留圖節點數不一致(如 7564 對 7762),解碼時被迫按當前狀態重預留、丟掉最壞情況尺寸,在 GGML_SCHED_DEBUG_REALLOC=1 下直接 abort。

llama.cpp releases10/7 12:11AI 評分42

llama.cpp 修復 SYCL 混用不同型號 GPU 的 FA 問題

llama.cpp 合併 PR #29071,修復 SYCL 後端在 Flash Attention(FA)中混用不同型號 GPU 時出現的問題。改動位於 ggml/src/ggml-sycl/ggml-sycl.cpp,提交由 Georgi Gerganov 共同署名。該修復包含在版本 b11464 中,此版本同時提供 Ubuntu x64(SYCL FP32、FP16)與 Windows x64(SYCL)等建置。

llama.cpp releases10/5 15:50AI 評分42

llama.cpp 發布 b11424,修復 Vulkan Flash Attention 共享記憶體越界寫

llama.cpp 發布 b11424 建置版本,修復 Vulkan 後端 Flash Attention 的共享記憶體越界寫問題(#29988)。該版本照例提供 macOS/iOS、Linux、Windows、Android 的預編譯包,涵蓋 Vulkan、CUDA 12/13、ROCm 10.0、OpenVINO、SYCL、OpenCL 等後端,並附驍龍 CPU/Adreno GPU/Hexagon NPU 的安裝指引。

llama.cpp releases10/4 13:34AI 評分36

llama.cpp 修復 Vulkan 後端 RDNA4 矩陣向量調優

llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。