AISpot
所屬事件:llama.cpp Metal 新增少行 MMA 核心,加速投機解碼(1 個來源 · 2 則報導)

2026 年 10 月 5 日至 7 日,llama.cpp 連續合併兩個 PR,為 Metal 後端加入少行 MMA(矩陣乘)核心,用於加速投機解碼中的小批次矩陣乘法。起因是:在沒有 tensor API 的 Metal 裝置上,投機解碼每步要驗證少量 draft token,此前只能用 mat-vec 核心處理,耗… 看完整事件

llama.cpp Metal 把 few-row MMA 核心擴充套件到 BF16、MXFP4 等型別

llama.cpp releases10/7 15:59版本更新地端推論開源

llama.cpp 的 Metal 後端將通用 few-row MMA 矩陣乘法核心擴充套件到更多權重量化型別,BF16、Q1_0、Q2_0、MXFP4、Q2_K、Q3_K、TQ2_0 及 IQ 系列現已納入。該核心適用於任何帶 16-weight 反量化器的型別,各型別從在 M3 Ultra 上能勝過現有核心的行數起啟用:TQ2_0 為 5 行,BF16 為 4 行,MXFP4、Q2_0、Q2_K、IQ4_NL 為 3 行,其餘為 2 行。

公開了各量化型別在 M3 Ultra 上的具體行數閾值與相對 master 的實測耗時比,讓在 Apple Silicon 上用 Metal 跑本地推理、關注 MUL_MAT 效能的人能判斷哪些格式會提速。

原標題:b11476
閱讀原文

同一事件共有 2 則報導(1 個來源),看事件全貌

評分52 / 72(平均 62,門檻 60)
狀態精選

全文翻譯

metal:用於其餘 src0 型別的 few-row MMA 矩陣乘法(#30065) 通用 few-row MMA 核心適用於任何具有 16 權重的反量化器的型別,因此它現在也接受 BF16、Q1_0、Q2_0、MXFP4、Q2_K、Q3_K、TQ2_0 和 IQ 型別。每種型別從它在 M3 Ultra 上優於當前核心的行數開始:TQ2_0 為 5 行,BF16 為 4 行,MXFP4、Q2_0、Q2_K 和 IQ4_NL 為 3 行,其他為 2 行。 test-backend-ops perf -o MUL_MAT, m=4096, k=14336, M3 Ultra,此更改相較 master 的耗時(每次為兩次交錯執行的平均值):從閾值到 8 行為 0.23 到 0.98,9 到 16 行為 0.24 到 0.33,1 和 512 行為 0.99 到 1.01。 網站: https://llama.app 證明: https://github.com/ggml-org/llama.cpp/attestations/53605242 macOS/iOS: - macOS Apple Silicon (arm64) - macOS Apple Silicon (arm64, KleidiAI enabled) 已停用 - macOS Intel (x64) - iOS XCFramework Linux: - Ubuntu x64 (CPU) - Ubuntu arm64 (CPU) - Ubuntu s390x (CPU) - Ubuntu x64 (Vulkan) - Ubuntu arm64 (Vulkan) - Ubuntu x64 (CUDA 12) - CUDA 12.8 庫 - Ubuntu x64 (CUDA 13) - CUDA 13.4 庫 - Ubuntu arm64 (CUDA 13) - CUDA 13.4 庫 - Ubuntu x64 (ROCm 10.0) - Ubuntu x64 (OpenVINO) - Ubuntu x64 (SYCL FP32) - Ubuntu x64 (SYCL FP16) - Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 設定指南 Android: - Android arm64 (CPU) - Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - 設定指南 Windows: - Windows x64 (CPU) - Windows arm64 (CPU) - Windows arm64 (OpenCL Adreno) - Windows x64 (CUDA 12) - CUDA 12.4 DLLs - Windows x64 (CUDA 13) - CUDA 13.4 DLLs - Windows arm64 (CUDA 13) - CUDA 13.4 DLLs - Windows x64 (Vulkan) - Windows arm64 (Vulkan) - Windows x64 (OpenVINO) - Windows x64 (SYCL) - Windows x64 (ROCm 10.0) openEuler: - 已停用 - openEuler x86 (310p) - openEuler x86 (910b, ACL Graph) - openEuler aarch64 (310p) - openEuler aarch64 (910b, ACL Graph) UI: - UI

由 AI 翻譯,以原文為準。

原文
metal : few-row MMA mat-mul for the remaining src0 types ( #30065 ) The generic few-row MMA kernel works for any type with a 16-weight dequantizer, so it now also takes BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K, TQ2_0 and the IQ types. Each type starts at the row count where it beats the current kernels on an M3 Ultra: 5 rows for TQ2_0, 4 for BF16, 3 for MXFP4, Q2_0, Q2_K and IQ4_NL, and 2 for the others. test-backend-ops perf -o MUL_MAT, m=4096, k=14336, M3 Ultra, time of this change over master (mean of two interleaved runs each): 0.23 to 0.98 from the threshold to 8 rows, 0.24 to 0.33 at 9 to 16 rows, and 0.99 to 1.01 at 1 and 512 rows. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/53605242 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/4 21:07AI 評分45

llama.cpp 為 x86 tinyBLAS 加入 BF16/FP16/FP32 K 尾部向量化

llama.cpp 的 ggml-cpu 後端在 x86 上為 tinyBLAS 增加了 BF16、FP16、FP32 的 K 尾部支援並做了向量化,同時讓 CPU 測試在啟用 use_ref 時跳過 tinyBLAS,以便與 vec_dot 路徑對比。改動以 PR #29806 合入,隨 b11398 建置發布,覆蓋 macOS、Linux、Windows、Android 等多平台 CPU 與 GPU 後端。

llama.cpp releases● 精選10/3 01:28AI 評分68

llama.cpp 為 Metal 新增 F16 KV 張量 API flash attention 核心

llama.cpp 合併 PR #29570,在 Metal 後端為 F16 KV 快取加入基於張量 API 的 flash attention 核心。該核心覆蓋 DK=DV=512、DK=576/DV=512、DK=192/DV=128 等配置,並支援 attention sinks、ALiBi 與 logit softcap。改動隨 b11362 建置發布,覆蓋 macOS Apple Silicon、iOS 及 Linux、Windows、Android 等多平台建置。

llama.cpp releases● 精選10/7 15:28AI 評分74

llama.cpp 修復 Metal 後端 MUL_MAT+ADD 融合錯誤

llama.cpp 合併 PR #30100,修復 Metal 後端 MUL_MAT+ADD 融合中殘差運算元的選擇錯誤:當 ADD 的兩個運算元都是矩陣乘法輸出(x = W1@u + W2@v)時,舊邏輯會選中自身從未寫入的輸出緩衝區,導致結果錯誤。Clef 決策模型因此在 Metal 上機率趨近均勻,billing 輸出 0.28,而 CPU 後端為 0.977,且同一檔案在 CPU 上結果正確,說明並非量化問題。

llama.cpp releases10/5 14:23AI 評分49

llama.cpp 最佳化 CUDA 後端 NVFP4 的 mmq 累加

llama.cpp 發布 b11417 建置,CUDA 後端針對 NVFP4 型別的 mmq 累加做了最佳化,改進 mmq_vec_dot_fp4_fp4_mma 以提升效能,並修正 mma_block_scaled_fp4 迴圈的縮排。該版本建置覆蓋 macOS/iOS、Linux、Windows、Android 與 openEuler,包含 CUDA 12.8、CUDA 13.4、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;