AISpot

llama.cpp 修復 Metal 後端 MUL_MAT+ADD 融合錯誤

llama.cpp releases10/7 15:28版本更新地端推論開源開發工具

llama.cpp 合併 PR #30100,修復 Metal 後端 MUL_MAT+ADD 融合中殘差運算元的選擇錯誤:當 ADD 的兩個運算元都是矩陣乘法輸出(x = W1@u + W2@v)時,舊邏輯會選中自身從未寫入的輸出緩衝區,導致結果錯誤。Clef 決策模型因此在 Metal 上機率趨近均勻,billing 輸出 0.28,而 CPU 後端為 0.977,且同一檔案在 CPU 上結果正確,說明並非量化問題。

官方確認 Metal 後端 MUL_MAT+ADD 融合的殘差判定會讀入錯誤緩衝區,並用新增測試量化影響(28 例失敗 27 例),對在 Apple Silicon 上用 Metal 跑本地模型的人最值得留意。

原標題:b11475
閱讀原文

評分73 / 75(平均 74,門檻 60)
狀態精選

全文翻譯

metal:修復當殘差本身是 MUL_MAT 時的 MUL_MAT+ADD 融合(#30100) metal:修復當殘差本身是 MUL_MAT 時的 MUL_MAT+ADD 融合 ggml_metal_op_mul_mat_mma 將融合後的 MUL_MAT+ADD 的殘差選取為 “其 op 不是 MUL_MAT 的 ADD 運算元”。當 ADD 的兩個運算元 都是矩陣乘法輸出(x = W1 @ u + W2 @ v)時,該判斷對兩者都成立,因此 殘差解析為融合矩陣乘法自身的、從未寫入的輸出,並且 核心會加上該緩衝區中當時儲存的任何內容。 融合檢查(ggml_metal_mul_mat_add_operand)已經按身份選擇 運算元;讓編碼器也這樣做。 Clef 決策模型在其 head 中遇到此問題(proj_option_context @ ctx + proj_option_lexical @ lex,9 個選項行):在 Metal 上,/v1/systemone 的機率坍縮至接近均勻分佈(billing 0.28,而 CPU 後端給出 0.977,Cloudflare_clef-flash Q8_0),對每種記憶體佈局都是確定性的, 使用 GGML_METAL_FUSION_DISABLE=1 時正確。不是量化問題:同一檔案在 CPU 上是正確的。 向 test-backend-ops 新增一個 MUL_MAT_ADD 模式,其中殘差是第二個 矩陣乘法;在 Metal 上,此更改前它在 28 個用例中失敗 27 個(唯一通過 的是 f16 n=2,低於 MMA 行閾值,因此不會發生任何融合)。 更新 tests/test-backend-ops.cpp 共同作者:Georgi Gerganov ggerganov@gmail.com 共同作者:Georgi Gerganov ggerganov@gmail.com 網站: https://llama.app 證明: https://github.com/ggml-org/llama.cpp/attestations/53594480 macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI enabled) 已停用 - macOS Intel (x64) - iOS XCFramework Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 庫 - Ubuntu x64 (CUDA 13) - CUDA 13.4 庫 - Ubuntu arm64 (CUDA 13) - CUDA 13.4 庫 - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 安裝指南 Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 安裝指南 Windows: - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - CUDA 12.4 DLL - Windows x64 (CUDA 13) - CUDA 13.4 DLL - Windows arm64 (CUDA 13) - CUDA 13.4 DLL - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0) openEuler: - 已停用 - openEuler x86 (310p) - openEuler x86(910b,ACL Graph) - openEuler aarch64(310p) - openEuler aarch64(910b,ACL Graph) - UI: - UI

由 AI 翻譯,以原文為準。

原文
metal : fix MUL_MAT+ADD fusion when the residual is itself a MUL_MAT ( #30100 ) metal : fix MUL_MAT+ADD fusion when the residual is itself a MUL_MAT ggml_metal_op_mul_mat_mma picks the residual of a fused MUL_MAT+ADD as "the ADD operand whose op is not MUL_MAT". When both operands of the ADD are mat-mul outputs (x = W1 @ u + W2 @ v), that test is true for both, so the residual resolves to the fused mat-mul's own, never-written output and the kernel adds whatever that buffer holds. The fusion check (ggml_metal_mul_mat_add_operand) already selects the operand by identity; make the encoder do the same. Clef decision models hit this in their head (proj_option_context @ ctx + proj_option_lexical @ lex, 9 option rows): on Metal, /v1/systemone probabilities collapse toward uniform (billing 0.28 where the CPU backend gives 0.977, Cloudflare_clef-flash Q8_0), deterministic per memory layout, correct with GGML_METAL_FUSION_DISABLE=1. Not a quantization issue: the same file is right on CPU. Add a MUL_MAT_ADD mode to test-backend-ops where the residual is a second mat-mul; on Metal it fails 27 of 28 cases before this change (the one pass is f16 n=2, under the MMA row threshold, so nothing fuses). Update tests/test-backend-ops.cpp Co-authored-by: Georgi Gerganov ggerganov@gmail.com Co-authored-by: Georgi Gerganov ggerganov@gmail.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/53594480 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/3 14:27AI 評分27

llama.cpp 修復 f16 ADD_ADD 不穩定測試併發布多平台建置

llama.cpp 合併 PR #29904,通過改用融合 ADD 容差修復 f16 下 ADD_ADD 測試的 flaky 問題。隨附的建置覆蓋 macOS Apple Silicon、iOS、Linux、Android、Windows 及 openEuler,包含 CUDA 12/13、ROCm 10.0、Vulkan、OpenVINO、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 版本和 openEuler 被標記為 DISABL…

llama.cpp releases● 精選10/7 15:59AI 評分62

llama.cpp Metal 把 few-row MMA 核心擴充套件到 BF16、MXFP4 等型別

llama.cpp 的 Metal 後端將通用 few-row MMA 矩陣乘法核心擴充套件到更多權重量化型別,BF16、Q1_0、Q2_0、MXFP4、Q2_K、Q3_K、TQ2_0 及 IQ 系列現已納入。該核心適用於任何帶 16-weight 反量化器的型別,各型別從在 M3 Ultra 上能勝過現有核心的行數起啟用:TQ2_0 為 5 行,BF16 為 4 行,MXFP4、Q2_0、Q2_K、IQ4_NL 為 3 行,其餘為 2 行。

llama.cpp releases10/4 13:34AI 評分36

llama.cpp 修復 Vulkan 後端 RDNA4 矩陣向量調優

llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。