AISpot

llama.cpp 為 Metal 加入少行 MMA 矩陣乘核心

llama.cpp releases10/5 06:29版本更新地端推論開源開發工具

llama.cpp 合併 PR #29869,為 Metal 後端新增支援 2 到 16 行 src1 的 mat-mul 核心,用於加速投機解碼中的少行矩陣乘。此前 Metal 在無 tensor API 時用 mat-vec 核心處理這些運算,耗時隨 src1 行數增長,導致 M3 Ultra 上 DFlash2 解碼慢於序列解碼。

這是 llama.cpp Metal 後端針對投機解碼少行矩陣乘的具體效能修復,含各量化型別啟用行數閾值與 M3 Ultra 實測背景,對在 Apple Silicon 上跑本地推理和投機解碼的人有直接參考價值。

原標題:b11404
閱讀原文

評分66 / 55(平均 60,門檻 60)
狀態精選

全文翻譯

metal : few-row MMA mat-mul ( #29869 ) metal : few-row MMA mat-mul and batched copies for speculative decoding 推測解碼每步驗證少量草稿 token。沒有 tensor API 時,Metal 用 mat-vec 核心執行這些矩陣乘法,其耗時隨每個 src1 行增長,因此 M3 Ultra 上的 DFlash2 解碼比序列解碼還慢。 為 8x8 simdgroup 矩陣上的 2..16 個 src1 行新增 mat-mul 核心:每個權重對所有行只反量化一次,執行緒組的 simdgroup 拆分 K。Q4_0、Q8_0 和 Q5_K 有各自的核心,F32、F16、Q4_1、Q5_0、Q5_1、Q4_K 和 Q6_K 使用基於 16 權重反量化器的通用路徑,而 2 行的 Q4_0 使用 mat-vec 核心的 2 行變體 僅在 MTLGPUFamilyApple7+ 且沒有 tensor API 時使用它們,從在 M3 Ultra 上它們勝過 mat-vec 核心的行數開始(F32:6,F16、Q4_K、Q5_0、Q5_1:3,其他型別:2) 融合表:MUL_MAT + ADD 在 MMA 儲存中新增同形狀殘差,同一對張量之間最多 16 個相鄰同佈局 f32 複製作為一次排程執行 融合檢查和 ggml_graph_optimize 接收裝置屬性,因此重排序僅在能夠融合它的裝置上打包 MUL_MAT + ADD,且在每個 src1 行數下都如此 當重排序打包一個組時,檢視不計入 GGML_METAL_FUSION_MAX,因此 16 個帶中間檢視的迴圈狀態快照複製保持為一個組 編碼器檢查融合組的內部節點是否存在併發,按範圍跟蹤已寫入的檢視,並且不將 CPY 的目標計為讀取 CONCAT 在行數很少時將長行拆分到多個執行緒組 測試:test-backend-ops 中的 few-row MUL_MAT、MUL_MAT_ADD、CPY_BATCH 和 CONCAT 用例(帶用於複製順序的 prepare_graph 鉤子)、test-metal-graph-optimize、test-metal-cpy-batch-alias metal : remove the CPY_BATCH fusion and the memory range changes 按評審建議,移除批次複製融合及其核心和測試,並還原記憶體範圍更改。記憶體範圍、圖重排序和 CPY 編碼器再次與 master 上相同。 cont : clean-up cont : drop has_tensor gate cont : clean-up operand/residual logic cont : drop Q4_0 ne11=2 special-case cont : add kernels/mul_mv_mma.metal cont : consolidate mma pipeline selection logic cont : decouple fusion logic from device props Co-authored-by: Georgi Gerganov ggerganov@gmail.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52722020 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

由 AI 翻譯,以原文為準。

原文
metal : few-row MMA mat-mul ( #29869 ) metal : few-row MMA mat-mul and batched copies for speculative decoding Speculative decoding verifies a few draft tokens per step. Without the tensor API, Metal ran these mat-muls with the mat-vec kernels, whose time grows with every src1 row, so DFlash2 decoding on an M3 Ultra was slower than serial decoding. add mat-mul kernels for 2..16 src1 rows on 8x8 simdgroup matrices: each weight is dequantized once for all rows, and the simdgroups of a threadgroup split K. Q4_0, Q8_0 and Q5_K have their own kernels, F32, F16, Q4_1, Q5_0, Q5_1, Q4_K and Q6_K use a generic path over the 16-weight dequantizers, and Q4_0 at 2 rows uses a 2-row variant of the mat-vec kernel use them only on MTLGPUFamilyApple7+ without the tensor API, from the row count at which they beat the mat-vec kernels on an M3 Ultra (F32: 6, F16, Q4_K, Q5_0, Q5_1: 3, other types: 2) fusion table: MUL_MAT + ADD adds a same-shape residual in the MMA store, and up to 16 adjacent same-layout f32 copies between the same two tensors run as one dispatch the fusion checks and ggml_graph_optimize take the device props, so the reorder packs MUL_MAT + ADD only on devices that can fuse it, at every src1 row count views do not count toward GGML_METAL_FUSION_MAX when the reorder packs a group, so 16 recurrent state snapshot copies with views between them stay one group the encoder checks the inner nodes of a fused group for concurrency, tracks written views by their extent, and does not count the destination of a CPY as a read CONCAT splits long rows across threadgroups when there are few rows tests: few-row MUL_MAT, MUL_MAT_ADD, CPY_BATCH and CONCAT cases in test-backend-ops (with a prepare_graph hook for the copy order), test-metal-graph-optimize, test-metal-cpy-batch-alias metal : remove the CPY_BATCH fusion and the memory range changes Remove the batched copy fusion with its kernel and tests, and revert the memory range changes, as suggested in review. The memory ranges, the graph reorder and the CPY encoder are again the same as on master. cont : clean-up cont : drop has_tensor gate cont : clean-up operand/residual logic cont : drop Q4_0 ne11=2 special-case cont : add kernels/mul_mv_mma.metal cont : consolidate mma pipeline selection logic cont : decouple fusion logic from device props Co-authored-by: Georgi Gerganov ggerganov@gmail.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52722020 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases● 精選10/3 01:28AI 評分68

llama.cpp 為 Metal 新增 F16 KV 張量 API flash attention 核心

llama.cpp 合併 PR #29570,在 Metal 後端為 F16 KV 快取加入基於張量 API 的 flash attention 核心。該核心覆蓋 DK=DV=512、DK=576/DV=512、DK=192/DV=128 等配置,並支援 attention sinks、ALiBi 與 logit softcap。改動隨 b11362 建置發布,覆蓋 macOS Apple Silicon、iOS 及 Linux、Windows、Android 等多平台建置。

llama.cpp releases10/4 13:34AI 評分36

llama.cpp 修復 Vulkan 後端 RDNA4 矩陣向量調優

llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。