AISpot

llama.cpp Vulkan 後端修復 TOP_K 的 inf/NaN 與負值錯誤

llama.cpp releases10/8 10:40版本更新地端推論開源

llama.cpp 發布 b11505,修復 Vulkan 後端 TOP_K 運算元在 +inf/NaN 輸入和 k=1 負值時的錯誤。原 bucket search 從 [0, 0xFF800000) 起算,+inf 與 NaN 永不被計數,可致 NVIDIA 上掛起或裝置丟失、AMD 上索引錯誤,並漏掉真實 top 值。

原標題:b11505
閱讀原文

評分45 / 40(平均 42,門檻 60)
狀態未入選

原文

vulkan : fix TOP_K for +inf/NaN inputs and k = 1 on negative values ( #30107 ) The bucket search in topk_nary_search.comp started from the range [0, 0xFF800000), which ends just below the ordered-uint mapping of +inf, so +inf and NaN were never counted. A workgroup block with fewer than k countable values left the ballot empty and the shader read uninitialized shared state (hang/device lost on NVIDIA, wrong indices on AMD), and a few +inf in a block were selected without being counted, dropping real top values. Map NaN to -inf on input, start from [0, 0xFFFFFFFF) so every value is counted, and clamp the top bucket's end (2^32) instead of wrapping to 0. The k = 1 path compared float bits as signed integers, which orders negative values backwards; compare floats instead. Add test_top_k_inf to test-backend-ops: negative values, fewer than k +inf and many -inf, for k = 1, 10, 40. Assisted-by: Claude Opus 5.5 Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/53971594 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/4 08:34AI 評分36

llama.cpp 修復 Vulkan 後端 RDNA4 矩陣向量調優

llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。

llama.cpp releases10/5 04:19AI 評分12

llama.cpp 發布 b11407,修復 Vulkan 測試編譯錯誤

llama.cpp 發布 b11407 版本,修復了在開啟 -DGGML_VULKAN_RUN_TESTS=ON 時 Vulkan 後端出現的未宣告識別符號問題。該版本同步提供 macOS、Linux、Windows、Android 等多平台預編譯包,涵蓋 CPU、CUDA 12/13、Vulkan、ROCm 10.0、OpenVINO、SYCL 及 Snapdragon 的 Adreno GPU 與 Hexagon NPU 等後端。

llama.cpp releases10/5 10:50AI 評分42

llama.cpp 發布 b11424,修復 Vulkan Flash Attention 共享記憶體越界寫

llama.cpp 發布 b11424 建置版本,修復 Vulkan 後端 Flash Attention 的共享記憶體越界寫問題(#29988)。該版本照例提供 macOS/iOS、Linux、Windows、Android 的預編譯包,涵蓋 Vulkan、CUDA 12/13、ROCm 10.0、OpenVINO、SYCL、OpenCL 等後端,並附驍龍 CPU/Adreno GPU/Hexagon NPU 的安裝指引。

llama.cpp releases10/6 12:07AI 評分41

llama.cpp 修復 Vulkan 後端在 1.0 loader 上的空指標呼叫

llama.cpp 合併 PR #29872,在 Vulkan 後端初始化時檢查 vkEnumerateInstanceVersion 是否為空指標。Vulkan 1.0 的 loader(例如 Android 8.1)不提供該函式,此前會讓後端初始化直接呼叫空指標,現在這類 loader 被與所有低於 1.2 的 loader 同樣處理。該補丁修復 issue #29871,並隨本次建置的 macOS、Linux、Windows、Android 等多平台產物一同發布。

llama.cpp releases10/7 02:53AI 評分42

llama.cpp 發布 b11461,修復 AMD 核顯 Vulkan 檢查點讀取慢

llama.cpp 發布 b11461,主要改動是修復 AMD 整合顯示卡在 Vulkan 後端下檢查點(checkpoint)讀取過慢的問題,對應 PR #30049。該版本同時更新了 macOS/iOS、Linux、Android、Windows、openEuler 的預編譯包,覆蓋 Vulkan、CUDA 12/13、ROCm 10.0、OpenVINO、SYCL,以及 Snapdragon 的 CPU、Adreno GPU 與 Hexagon NPU 後端。