llama.cpp 建置 b11448:ggml-cuda 分塊處理大塊 BF16/FP16 轉 F32
llama.cpp 發布建置 b11448,ggml-cuda 後端把大塊 BF16/FP16 到 F32 的轉換改為分塊執行(#29442),並修正分塊 cuBLAS 矩陣乘法未遵循目標 stride 的問題,提交由 Johannes Gäßler 參與。改動僅涉及 CUDA 後端,影響用 NVIDIA GPU 做本地推理的場景;該建置同時覆蓋 macOS Apple Silicon 與 Linux、Windows 的 CUDA 12/13 等平台。
原標題:b11448
閱讀原文
| 評分 | 45 / 45(平均 45,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
cuda: BF16/FP16 conversion to f32 chunking ( #29442 )
ggml-cuda: chunk large BF16/FP16 to F32 conversions
Update ggml/src/ggml-cuda/ggml-cuda.cu
Co-authored-by: Johannes Gäßler johannesg@5d6.de
Update ggml/src/ggml-cuda/ggml-cuda.cu
Co-authored-by: Johannes Gäßler johannesg@5d6.de
ggml-cuda: respect dst stride in chunked cuBLAS matmul
Co-authored-by: Johannes Gäßler johannesg@5d6.de
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53292370
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/4 15:12AI 評分16
llama.cpp 發布 b11391 建置版本,唯一程式碼改動是將 CUDA 後端的 blocks_per_col 變數移到實際使用位置(PR #29939),由 Hugging Face 的 Adrien Gallouët 提交。
llama.cpp releases10/4 21:07AI 評分45
llama.cpp 的 ggml-cpu 後端在 x86 上為 tinyBLAS 增加了 BF16、FP16、FP32 的 K 尾部支援並做了向量化,同時讓 CPU 測試在啟用 use_ref 時跳過 tinyBLAS,以便與 vec_dot 路徑對比。改動以 PR #29806 合入,隨 b11398 建置發布,覆蓋 macOS、Linux、Windows、Android 等多平台 CPU 與 GPU 後端。
llama.cpp releases10/5 05:35AI 評分37
llama.cpp 發布 b11402 版本,CUDA 後端改為優先採用整塊(whole-tile)FlashAttention 排程,以提升兩階段核心效率。
llama.cpp releases10/4 20:50AI 評分23
llama.cpp 發布建置版本 b11397,主要變更是將 CUDA 後端的 neu_padded 移動到實際使用位置(PR #29940),作者為 Hugging Face 的 Adrien Gallouët。該版本繼續提供 macOS、Linux、Windows、Android 等多平台預編譯包,覆蓋 Apple Silicon、CUDA 12/13、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;
llama.cpp releases10/5 06:24AI 評分42
llama.cpp 合併 PR #29633,在 CUDA 後端針對小批次場景的 thin f16/bf16 mul_mat 改用 MMVF 指令,並調整了核心選擇邏輯。該改動由 NVIDIA 工程師提交,影響 Ubuntu 與 Windows 的 CUDA 12/13 建置,面向在 NVIDIA GPU 上執行本地推理的使用者。