llama.cpp 為 Metal 加入少行 MMA 矩陣乘核心
llama.cpp 合併 PR #29869,為 Metal 後端新增支援 2 到 16 行 src1 的 mat-mul 核心,用於加速投機解碼中的少行矩陣乘。此前 Metal 在無 tensor API 時用 mat-vec 核心處理這些運算,耗時隨 src1 行數增長,導致 M3 Ultra 上 DFlash2 解碼慢於序列解碼。
這是 llama.cpp Metal 後端針對投機解碼少行矩陣乘的具體效能修復,含各量化型別啟用行數閾值與 M3 Ultra 實測背景,對在 Apple Silicon 上跑本地推理和投機解碼的人有直接參考價值。
原標題:b11404
閱讀原文
| 評分 | 66 / 55(平均 60,門檻 60) |
| 狀態 | 精選 |
|---|
全文翻譯
metal : few-row MMA mat-mul ( #29869 )
metal : few-row MMA mat-mul and batched copies for speculative decoding
推測解碼每步驗證少量草稿 token。沒有 tensor API 時,Metal 用 mat-vec 核心執行這些矩陣乘法,其耗時隨每個 src1 行增長,因此 M3 Ultra 上的 DFlash2 解碼比序列解碼還慢。
為 8x8 simdgroup 矩陣上的 2..16 個 src1 行新增 mat-mul 核心:每個權重對所有行只反量化一次,執行緒組的 simdgroup 拆分 K。Q4_0、Q8_0 和 Q5_K 有各自的核心,F32、F16、Q4_1、Q5_0、Q5_1、Q4_K 和 Q6_K 使用基於 16 權重反量化器的通用路徑,而 2 行的 Q4_0 使用 mat-vec 核心的 2 行變體
僅在 MTLGPUFamilyApple7+ 且沒有 tensor API 時使用它們,從在 M3 Ultra 上它們勝過 mat-vec 核心的行數開始(F32:6,F16、Q4_K、Q5_0、Q5_1:3,其他型別:2)
融合表:MUL_MAT + ADD 在 MMA 儲存中新增同形狀殘差,同一對張量之間最多 16 個相鄰同佈局 f32 複製作為一次排程執行
融合檢查和 ggml_graph_optimize 接收裝置屬性,因此重排序僅在能夠融合它的裝置上打包 MUL_MAT + ADD,且在每個 src1 行數下都如此
當重排序打包一個組時,檢視不計入 GGML_METAL_FUSION_MAX,因此 16 個帶中間檢視的迴圈狀態快照複製保持為一個組
編碼器檢查融合組的內部節點是否存在併發,按範圍跟蹤已寫入的檢視,並且不將 CPY 的目標計為讀取
CONCAT 在行數很少時將長行拆分到多個執行緒組
測試:test-backend-ops 中的 few-row MUL_MAT、MUL_MAT_ADD、CPY_BATCH 和 CONCAT 用例(帶用於複製順序的 prepare_graph 鉤子)、test-metal-graph-optimize、test-metal-cpy-batch-alias
metal : remove the CPY_BATCH fusion and the memory range changes
按評審建議,移除批次複製融合及其核心和測試,並還原記憶體範圍更改。記憶體範圍、圖重排序和 CPY 編碼器再次與 master 上相同。
cont : clean-up
cont : drop has_tensor gate
cont : clean-up operand/residual logic
cont : drop Q4_0 ne11=2 special-case
cont : add kernels/mul_mv_mma.metal
cont : consolidate mma pipeline selection logic
cont : decouple fusion logic from device props
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52722020
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
由 AI 翻譯,以原文為準。
原文
metal : few-row MMA mat-mul ( #29869 )
metal : few-row MMA mat-mul and batched copies for speculative decoding
Speculative decoding verifies a few draft tokens per step. Without the tensor API, Metal ran these mat-muls with the mat-vec kernels, whose time grows with every src1 row, so DFlash2 decoding on an M3 Ultra was slower than serial decoding.
add mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices: each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4_0, Q8_0 and Q5_K have their own kernels, F32, F16, Q4_1, Q5_0, Q5_1, Q4_K and Q6_K use a generic path over the 16-weight dequantizers, and Q4_0 at 2 rows uses a 2-row variant of the mat-vec kernel
use them only on MTLGPUFamilyApple7+ without the tensor API, from the row count at which they beat the mat-vec kernels on an M3 Ultra (F32: 6, F16, Q4_K, Q5_0, Q5_1: 3, other types: 2)
fusion table: MUL_MAT + ADD adds a same-shape residual in the MMA store, and up to 16 adjacent same-layout f32 copies between the same two tensors run as one dispatch
the fusion checks and ggml_graph_optimize take the device props, so the reorder packs MUL_MAT + ADD only on devices that can fuse it, at every src1 row count
views do not count toward GGML_METAL_FUSION_MAX when the reorder packs a group, so 16 recurrent state snapshot copies with views between them stay one group
the encoder checks the inner nodes of a fused group for concurrency, tracks written views by their extent, and does not count the destination of a CPY as a read
CONCAT splits long rows across threadgroups when there are few rows
tests: few-row MUL_MAT, MUL_MAT_ADD, CPY_BATCH and CONCAT cases in test-backend-ops (with a prepare_graph hook for the copy order), test-metal-graph-optimize, test-metal-cpy-batch-alias
metal : remove the CPY_BATCH fusion and the memory range changes
Remove the batched copy fusion with its kernel and tests, and revert the
memory range changes, as suggested in review. The memory ranges, the
graph reorder and the CPY encoder are again the same as on master.
cont : clean-up
cont : drop has_tensor gate
cont : clean-up operand/residual logic
cont : drop Q4_0 ne11=2 special-case
cont : add kernels/mul_mv_mma.metal
cont : consolidate mma pipeline selection logic
cont : decouple fusion logic from device props
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52722020
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases● 精選10/3 01:28AI 評分68
llama.cpp 合併 PR #29570,在 Metal 後端為 F16 KV 快取加入基於張量 API 的 flash attention 核心。該核心覆蓋 DK=DV=512、DK=576/DV=512、DK=192/DV=128 等配置,並支援 attention sinks、ALiBi 與 logit softcap。改動隨 b11362 建置發布,覆蓋 macOS Apple Silicon、iOS 及 Linux、Windows、Android 等多平台建置。
llama.cpp releases10/5 06:24AI 評分42
llama.cpp 合併 PR #29633,在 CUDA 後端針對小批次場景的 thin f16/bf16 mul_mat 改用 MMVF 指令,並調整了核心選擇邏輯。該改動由 NVIDIA 工程師提交,影響 Ubuntu 與 Windows 的 CUDA 12/13 建置,面向在 NVIDIA GPU 上執行本地推理的使用者。
llama.cpp releases10/3 00:18AI 評分31
llama.cpp 合併 PR #28531,在共享記憶體為 32KB 的三星 GPU 上停用 Vulkan 後端的大矩陣乘分塊(large matmul tile),以規避該硬體配置下的問題。該改動由 Claude Opus 輔助完成,並附有建置證明連結。
llama.cpp releases10/4 13:34AI 評分36
llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。
llama.cpp releases10/4 13:56AI 評分37
llama.cpp 發布建置版本 b11390,主要修復了 CUDA 後端在 n_expert 遠大於 n_ubatch 時的 MMQ 記憶體故障(#29941)。