AISpot

llama.cpp 修復 ggml CLAMP 非連續檢視定址錯誤

llama.cpp releases10/6 19:51版本更新地端推論開源開發工具

ggml 修復了 CLAMP 運算元在非連續檢視上的定址錯誤,涉及 CPU 與 CUDA 後端。CUDA 此前按 ggml_nelements 做平坦處理、忽略檢視 stride,CPU 則把第 j 行按 j*nb01 定址、忽略 nb02/nb03;現在兩者都改為遵循 dims 1..3 的 stride,且 CUDA 的 supports_op 要求行連續,與 Metal 一致。

原標題:b11449
閱讀原文

評分47 / 45(平均 46,門檻 60)
狀態未入選

原文

ggml: fix CLAMP on non-contiguous views (CPU, CUDA) ( #29517 ) ggml: fix CLAMP on non-contiguous views (CPU, CUDA) CUDA clamped ggml_nelements values flat and ignored the view strides. CPU addressed row j as j*nb01 and ignored nb02/nb03. Both now follow the strides of dims 1..3; CUDA supports_op requires contiguous rows, like Metal. test_clamp gains a non-contiguous view case. cuda: clamp kernel uses fastdiv for the view strides Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/53297891 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/6 07:57AI 評分45

llama.cpp b11436 修復 OpenVINO 後端多項 GPU 迴歸

llama.cpp b11436 修復 ggml-openvino 後端的 CI 測試失敗與多項 GPU 迴歸。上游 #29622 給輸入 embedding 圖加入混合 token/embd 分支,後端改為只從計算節點建置 OV 模型,並把同型別連續 DUP 按 CONT 翻譯以避免圖被拆分,修復了 CPU 與 GPU 上首個單 token decode 失敗;

llama.cpp releases10/6 19:34AI 評分45

llama.cpp 建置 b11448:ggml-cuda 分塊處理大塊 BF16/FP16 轉 F32

llama.cpp 發布建置 b11448,ggml-cuda 後端把大塊 BF16/FP16 到 F32 的轉換改為分塊執行(#29442),並修正分塊 cuBLAS 矩陣乘法未遵循目標 stride 的問題,提交由 Johannes Gäßler 參與。改動僅涉及 CUDA 後端,影響用 NVIDIA GPU 做本地推理的場景;該建置同時覆蓋 macOS Apple Silicon 與 Linux、Windows 的 CUDA 12/13 等平台。

llama.cpp releases10/3 02:32AI 評分37

llama.cpp 修復 ggml-cpu 中 soft_max_back 別名輸出錯誤

llama.cpp 合併 PR #27096,修復 ggml-cpu 後端 GGML_OP_SOFT_MAX_BACK 在 dst 與 src1(y)別名時輸出靜默錯誤的問題:原實現先覆蓋 y 再讀取,導致結果錯誤。補丁改為單次融合迴圈,先讀兩個輸入再寫出,並新增迴歸測試強制觸發該別名。CUDA 核心因先完成歸約再寫出,不受影響。

llama.cpp releases10/5 15:50AI 評分42

llama.cpp 發布 b11424,修復 Vulkan Flash Attention 共享記憶體越界寫

llama.cpp 發布 b11424 建置版本,修復 Vulkan 後端 Flash Attention 的共享記憶體越界寫問題(#29988)。該版本照例提供 macOS/iOS、Linux、Windows、Android 的預編譯包,涵蓋 Vulkan、CUDA 12/13、ROCm 10.0、OpenVINO、SYCL、OpenCL 等後端,並附驍龍 CPU/Adreno GPU/Hexagon NPU 的安裝指引。