llama.cpp 修復 Metal 後端 MUL_MAT+ADD 融合錯誤
llama.cpp 合併 PR #30100,修復 Metal 後端 MUL_MAT+ADD 融合中殘差運算元的選擇錯誤:當 ADD 的兩個運算元都是矩陣乘法輸出(x = W1@u + W2@v)時,舊邏輯會選中自身從未寫入的輸出緩衝區,導致結果錯誤。Clef 決策模型因此在 Metal 上機率趨近均勻,billing 輸出 0.28,而 CPU 後端為 0.977,且同一檔案在 CPU 上結果正確,說明並非量化問題。
官方確認 Metal 後端 MUL_MAT+ADD 融合的殘差判定會讀入錯誤緩衝區,並用新增測試量化影響(28 例失敗 27 例),對在 Apple Silicon 上用 Metal 跑本地模型的人最值得留意。
原標題:b11475
閱讀原文
| 評分 | 73 / 75(平均 74,門檻 60) |
| 狀態 | 精選 |
|---|
全文翻譯
metal:修復當殘差本身是 MUL_MAT 時的 MUL_MAT+ADD 融合(#30100)
metal:修復當殘差本身是 MUL_MAT 時的 MUL_MAT+ADD 融合
ggml_metal_op_mul_mat_mma 將融合後的 MUL_MAT+ADD 的殘差選取為
“其 op 不是 MUL_MAT 的 ADD 運算元”。當 ADD 的兩個運算元
都是矩陣乘法輸出(x = W1 @ u + W2 @ v)時,該判斷對兩者都成立,因此
殘差解析為融合矩陣乘法自身的、從未寫入的輸出,並且
核心會加上該緩衝區中當時儲存的任何內容。
融合檢查(ggml_metal_mul_mat_add_operand)已經按身份選擇
運算元;讓編碼器也這樣做。
Clef 決策模型在其 head 中遇到此問題(proj_option_context @ ctx +
proj_option_lexical @ lex,9 個選項行):在 Metal 上,/v1/systemone
的機率坍縮至接近均勻分佈(billing 0.28,而 CPU 後端給出 0.977,Cloudflare_clef-flash Q8_0),對每種記憶體佈局都是確定性的,
使用 GGML_METAL_FUSION_DISABLE=1 時正確。不是量化問題:同一檔案在 CPU 上是正確的。
向 test-backend-ops 新增一個 MUL_MAT_ADD 模式,其中殘差是第二個
矩陣乘法;在 Metal 上,此更改前它在 28 個用例中失敗 27 個(唯一通過
的是 f16 n=2,低於 MMA 行閾值,因此不會發生任何融合)。
更新 tests/test-backend-ops.cpp
共同作者:Georgi Gerganov ggerganov@gmail.com
共同作者:Georgi Gerganov ggerganov@gmail.com
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/53594480
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) 已停用
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (CUDA 12) - CUDA 12.8 庫
- Ubuntu x64 (CUDA 13) - CUDA 13.4 庫
- Ubuntu arm64 (CUDA 13) - CUDA 13.4 庫
- Ubuntu x64 (ROCm 10.0)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
- Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 安裝指南
Android:
- Android arm64 (CPU)
- Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 安裝指南
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLL
- Windows x64 (CUDA 13) - CUDA 13.4 DLL
- Windows arm64 (CUDA 13) - CUDA 13.4 DLL
- Windows x64 (Vulkan)
- Windows arm64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (ROCm 10.0)
openEuler:
- 已停用
- openEuler x86 (310p)
- openEuler x86(910b,ACL Graph)
- openEuler aarch64(310p)
- openEuler aarch64(910b,ACL Graph)
- UI:
- UI
由 AI 翻譯,以原文為準。
原文
metal : fix MUL_MAT+ADD fusion when the residual is itself a MUL_MAT ( #30100 )
metal : fix MUL_MAT+ADD fusion when the residual is itself a MUL_MAT
ggml_metal_op_mul_mat_mma picks the residual of a fused MUL_MAT+ADD as
"the ADD operand whose op is not MUL_MAT". When both operands of the ADD
are mat-mul outputs (x = W1 @ u + W2 @ v), that test is true for both, so
the residual resolves to the fused mat-mul's own, never-written output and
the kernel adds whatever that buffer holds.
The fusion check (ggml_metal_mul_mat_add_operand) already selects the
operand by identity; make the encoder do the same.
Clef decision models hit this in their head (proj_option_context @ ctx +
proj_option_lexical @ lex, 9 option rows): on Metal, /v1/systemone
probabilities collapse toward uniform (billing 0.28 where the CPU backend
gives 0.977, Cloudflare_clef-flash Q8_0), deterministic per memory layout,
correct with GGML_METAL_FUSION_DISABLE=1. Not a quantization issue: the
same file is right on CPU.
Add a MUL_MAT_ADD mode to test-backend-ops where the residual is a second
mat-mul; on Metal it fails 27 of 28 cases before this change (the one pass
is f16 n=2, under the MMA row threshold, so nothing fuses).
Update tests/test-backend-ops.cpp
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53594480
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/6 18:33AI 評分43
llama.cpp 發布 b11446,Metal 後端修復量化 flash attention 中 threadgroup memory 超額佔用的問題,對應 PR #29340。
llama.cpp releases10/3 14:27AI 評分27
llama.cpp 合併 PR #29904,通過改用融合 ADD 容差修復 f16 下 ADD_ADD 測試的 flaky 問題。隨附的建置覆蓋 macOS Apple Silicon、iOS、Linux、Android、Windows 及 openEuler,包含 CUDA 12/13、ROCm 10.0、Vulkan、OpenVINO、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 版本和 openEuler 被標記為 DISABL…
llama.cpp releases● 精選10/7 15:59AI 評分62
llama.cpp 的 Metal 後端將通用 few-row MMA 矩陣乘法核心擴充套件到更多權重量化型別,BF16、Q1_0、Q2_0、MXFP4、Q2_K、Q3_K、TQ2_0 及 IQ 系列現已納入。該核心適用於任何帶 16-weight 反量化器的型別,各型別從在 M3 Ultra 上能勝過現有核心的行數起啟用:TQ2_0 為 5 行,BF16 為 4 行,MXFP4、Q2_0、Q2_K、IQ4_NL 為 3 行,其餘為 2 行。
llama.cpp releases10/4 13:34AI 評分36
llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。
llama.cpp releases10/5 06:24AI 評分42
llama.cpp 合併 PR #29633,在 CUDA 後端針對小批次場景的 thin f16/bf16 mul_mat 改用 MMVF 指令,並調整了核心選擇邏輯。該改動由 NVIDIA 工程師提交,影響 Ubuntu 與 Windows 的 CUDA 12/13 建置,面向在 NVIDIA GPU 上執行本地推理的使用者。