AISpot

llama.cpp 修復 ggml-cpu soft_max_back 別名輸出錯誤

llama.cpp releases10/3 02:32版本更新地端推論開源開發工具

llama.cpp 合併 PR #27096,修復 ggml-cpu 中 GGML_OP_SOFT_MAX_BACK 在 dst 與 src1(y)別名時輸出靜默錯誤的問題。原因是原實現先覆蓋 y 再讀取,改為單次融合迴圈先讀兩個源再寫,並新增迴歸測試強制觸發該別名。CUDA 核心因先完成歸約再寫入,不受影響。

你在 Mac Studio 上用 MLX、llama.cpp 跑本地模型時,這類 CPU 後端數值 bug 會直接影響訓練或反向傳播結果,升級到含該修復的版本可避免靜默錯誤。

原標題:b11365
閱讀原文

評分62 / 55(平均 58,門檻 60)
狀態未入選

原文

ggml-cpu : fix soft_max_back wrong output when dst aliases src1 ( #27096 ) ggml-cpu : fix soft_max_back wrong output when dst aliases src1 GGML_OP_SOFT_MAX_BACK is listed in ggml_op_can_inplace, so the graph allocator may assign dst to alias either src0 (dy) or src1 (y). The result was built in several steps: ggml_vec_cpy_f32 (nc, dx, dy); ggml_vec_acc1_f32 (nc, dx, -dot_y_dy); ggml_vec_mul_f32 (nc, dx, dx, y); ggml_vec_scale_f32(nc, dx, scale); When dst aliases src1, the first step overwrites y and the third step then reads the overwritten values, so the output is silently wrong. Aliasing dst with src0 is unaffected. The CUDA kernel completes its reduction before writing and is already safe. Replace the sequence with a single fused loop that reads both sources before writing, which is correct under either aliasing. Add a regression test that marks dy as a graph output so the allocator is forced to alias dst with y, asserts that the alias actually happened, and compares against values computed on the host. cont : remove comment Co-authored-by: Georgi Gerganov ggerganov@gmail.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52336742 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/3 22:35AI 評分32

llama.cpp 修復 Windows 下 mtmd 棄用警告

llama.cpp 發布 b11381 版本,修復了 mtmd 在 Windows 上觸發的 strdup 棄用警告(#29863),由 Hugging Face 的 Adrien Gallouët 提交。該版本繼續提供 macOS Apple Silicon(arm64)、Linux、Windows、Android 等多平台預編譯包,其中 macOS Apple Silicon 的 KleidiAI 加速版被標記為 DISABLED。