所屬事件:
llama.cpp Metal 新增少行 MMA 核心,加速投機解碼(1 個來源 · 2 則報導)
2026 年 10 月 5 日至 7 日,llama.cpp 連續合併兩個 PR,為 Metal 後端加入少行 MMA(矩陣乘)核心,用於加速投機解碼中的小批次矩陣乘法。起因是:在沒有 tensor API 的 Metal 裝置上,投機解碼每步要驗證少量 draft token,此前只能用 mat-vec 核心處理,耗… 看完整事件
llama.cpp Metal 把 few-row MMA 核心擴充套件到 BF16、MXFP4 等型別
llama.cpp 的 Metal 後端將通用 few-row MMA 矩陣乘法核心擴充套件到更多權重量化型別,BF16、Q1_0、Q2_0、MXFP4、Q2_K、Q3_K、TQ2_0 及 IQ 系列現已納入。該核心適用於任何帶 16-weight 反量化器的型別,各型別從在 M3 Ultra 上能勝過現有核心的行數起啟用:TQ2_0 為 5 行,BF16 為 4 行,MXFP4、Q2_0、Q2_K、IQ4_NL 為 3 行,其餘為 2 行。
公開了各量化型別在 M3 Ultra 上的具體行數閾值與相對 master 的實測耗時比,讓在 Apple Silicon 上用 Metal 跑本地推理、關注 MUL_MAT 效能的人能判斷哪些格式會提速。
原標題:b11476
閱讀原文
同一事件共有 2 則報導(1 個來源),看事件全貌
| 評分 | 52 / 72(平均 62,門檻 60) |
| 狀態 | 精選 |
|---|
全文翻譯
metal:用於其餘 src0 型別的 few-row MMA 矩陣乘法(#30065)
通用 few-row MMA 核心適用於任何具有 16 權重的反量化器的型別,因此它現在也接受 BF16、Q1_0、Q2_0、MXFP4、Q2_K、Q3_K、TQ2_0 和 IQ 型別。每種型別從它在 M3 Ultra 上優於當前核心的行數開始:TQ2_0 為 5 行,BF16 為 4 行,MXFP4、Q2_0、Q2_K 和 IQ4_NL 為 3 行,其他為 2 行。
test-backend-ops perf -o MUL_MAT, m=4096, k=14336, M3 Ultra,此更改相較 master 的耗時(每次為兩次交錯執行的平均值):從閾值到 8 行為 0.23 到 0.98,9 到 16 行為 0.24 到 0.33,1 和 512 行為 0.99 到 1.01。
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/53605242
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) 已停用
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (CUDA 12) - CUDA 12.8 庫
- Ubuntu x64 (CUDA 13) - CUDA 13.4 庫
- Ubuntu arm64 (CUDA 13) - CUDA 13.4 庫
- Ubuntu x64 (ROCm 10.0)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
- Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 設定指南
Android:
- Android arm64 (CPU)
- Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 設定指南
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.4 DLLs
- Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
- Windows x64 (Vulkan)
- Windows arm64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (ROCm 10.0)
openEuler:
- 已停用
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI:
- UI
由 AI 翻譯,以原文為準。
原文
metal : few-row MMA mat-mul for the remaining src0 types ( #30065 )
The generic few-row MMA kernel works for any type with a 16-weight
dequantizer, so it now also takes BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K,
TQ2_0 and the IQ types. Each type starts at the row count where it beats
the current kernels on an M3 Ultra: 5 rows for TQ2_0, 4 for BF16, 3
for MXFP4, Q2_0, Q2_K and IQ4_NL, and 2 for the others.
test-backend-ops perf -o MUL_MAT, m=4096, k=14336, M3 Ultra, time of this
change over master (mean of two interleaved runs each): 0.23 to 0.98 from
the threshold to 8 rows, 0.24 to 0.33 at 9 to 16 rows, and 0.99 to 1.01 at
1 and 512 rows.
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53605242
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/5 06:24AI 評分42
llama.cpp 合併 PR #29633,在 CUDA 後端針對小批次場景的 thin f16/bf16 mul_mat 改用 MMVF 指令,並調整了核心選擇邏輯。該改動由 NVIDIA 工程師提交,影響 Ubuntu 與 Windows 的 CUDA 12/13 建置,面向在 NVIDIA GPU 上執行本地推理的使用者。
llama.cpp releases10/4 21:07AI 評分45
llama.cpp 的 ggml-cpu 後端在 x86 上為 tinyBLAS 增加了 BF16、FP16、FP32 的 K 尾部支援並做了向量化,同時讓 CPU 測試在啟用 use_ref 時跳過 tinyBLAS,以便與 vec_dot 路徑對比。改動以 PR #29806 合入,隨 b11398 建置發布,覆蓋 macOS、Linux、Windows、Android 等多平台 CPU 與 GPU 後端。
llama.cpp releases● 精選10/3 01:28AI 評分68
llama.cpp 合併 PR #29570,在 Metal 後端為 F16 KV 快取加入基於張量 API 的 flash attention 核心。該核心覆蓋 DK=DV=512、DK=576/DV=512、DK=192/DV=128 等配置,並支援 attention sinks、ALiBi 與 logit softcap。改動隨 b11362 建置發布,覆蓋 macOS Apple Silicon、iOS 及 Linux、Windows、Android 等多平台建置。
llama.cpp releases● 精選10/7 15:28AI 評分74
llama.cpp 合併 PR #30100,修復 Metal 後端 MUL_MAT+ADD 融合中殘差運算元的選擇錯誤:當 ADD 的兩個運算元都是矩陣乘法輸出(x = W1@u + W2@v)時,舊邏輯會選中自身從未寫入的輸出緩衝區,導致結果錯誤。Clef 決策模型因此在 Metal 上機率趨近均勻,billing 輸出 0.28,而 CPU 後端為 0.977,且同一檔案在 CPU 上結果正確,說明並非量化問題。
llama.cpp releases10/5 14:23AI 評分49
llama.cpp 發布 b11417 建置,CUDA 後端針對 NVFP4 型別的 mmq 累加做了最佳化,改進 mmq_vec_dot_fp4_fp4_mma 以提升效能,並修正 mma_block_scaled_fp4 迴圈的縮排。該版本建置覆蓋 macOS/iOS、Linux、Windows、Android 與 openEuler,包含 CUDA 12.8、CUDA 13.4、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;