llama.cpp 最佳化 CUDA 後端 NVFP4 的 mmq 累加
llama.cpp 發布 b11417 建置,CUDA 後端針對 NVFP4 型別的 mmq 累加做了最佳化,改進 mmq_vec_dot_fp4_fp4_mma 以提升效能,並修正 mma_block_scaled_fp4 迴圈的縮排。該版本建置覆蓋 macOS/iOS、Linux、Windows、Android 與 openEuler,包含 CUDA 12.8、CUDA 13.4、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;
原標題:b11417
閱讀原文
| 評分 | 50 / 48(平均 49,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
CUDA: Optimize accumulation in mmq for NVFP4 type ( #29857 )
ggml_cuda: optimize accumulation in mmq_vec_dot_fp4_fp4_mma for better performance
remove whitespace
fix: correct indentation in mma_block_scaled_fp4 loop
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52836411
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/4 13:56AI 評分37
llama.cpp 發布建置版本 b11390,主要修復了 CUDA 後端在 n_expert 遠大於 n_ubatch 時的 MMQ 記憶體故障(#29941)。
llama.cpp releases10/4 13:34AI 評分36
llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。
llama.cpp releases10/5 06:24AI 評分42
llama.cpp 合併 PR #29633,在 CUDA 後端針對小批次場景的 thin f16/bf16 mul_mat 改用 MMVF 指令,並調整了核心選擇邏輯。該改動由 NVIDIA 工程師提交,影響 Ubuntu 與 Windows 的 CUDA 12/13 建置,面向在 NVIDIA GPU 上執行本地推理的使用者。
r/LocalLLaMA top of the day10/2 19:47AI 評分65
llama.cpp 的 PR #29184 提出在 CUDA 後端把共享專家融合進 MMVQ,為部分 MoE 架構帶來提速,例如 Qwen 35B A3B。有評論稱該改動使 TG 速度提升約 5%,並指出僅對部分 MoE 架構有效。
llama.cpp releases10/4 15:12AI 評分16
llama.cpp 發布 b11391 建置版本,唯一程式碼改動是將 CUDA 後端的 blocks_per_col 變數移到實際使用位置(PR #29939),由 Hugging Face 的 Adrien Gallouët 提交。