AISpot

llama.cpp b11456:CUDA 緩衝區初始化改用每執行緒流

llama.cpp releases10/6 22:09版本更新地端推論開源

llama.cpp 發布建置 b11456,其中 ggml-cuda 在緩衝區初始化的 padding memset 中改用每執行緒 stream(#28782),CI 也重新為 ROCm 啟用 test-backend-ops 的 -j 並行測試。該建置覆蓋 macOS/iOS、Linux、Android、Windows 與 openEuler,含 CPU、Vulkan、CUDA 12/13、ROCm 10.0、OpenVINO、SYCL 等後端;

原標題:b11456
閱讀原文

評分45 / 38(平均 41,門檻 60)
狀態未入選

原文

ggml-cuda: use per-thread stream for buffer-init padding memset ( #28782 ) ggml-cuda: use per-thread stream for buffer-init padding memset ci : re-enable test-backend-ops -j for ROCm Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/53340384 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/5 16:21AI 評分43

llama.cpp 發布 b11425:CUDA alloc_deps 檢查改為與 batch 無關

llama.cpp 發布 b11425 建置,主要變更是把 CUDA 後端的 alloc_deps 檢查改為與 batch 無關(PR #29986,修復 #29980)。該版本同時更新各平台預編譯產物,覆蓋 macOS Apple Silicon 與 Intel、iOS XCFramework、Android arm64,以及 Windows 和 Ubuntu 上的 CUDA 12、CUDA 13、Vulkan、ROCm 10.0、SYCL、OpenVINO 等後端。

llama.cpp releases10/4 20:50AI 評分23

llama.cpp 發布 b11397,CUDA 後端調整 neu_padded 位置

llama.cpp 發布建置版本 b11397,主要變更是將 CUDA 後端的 neu_padded 移動到實際使用位置(PR #29940),作者為 Hugging Face 的 Adrien Gallouët。該版本繼續提供 macOS、Linux、Windows、Android 等多平台預編譯包,覆蓋 Apple Silicon、CUDA 12/13、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;

llama.cpp releases10/6 19:34AI 評分45

llama.cpp 建置 b11448:ggml-cuda 分塊處理大塊 BF16/FP16 轉 F32

llama.cpp 發布建置 b11448,ggml-cuda 後端把大塊 BF16/FP16 到 F32 的轉換改為分塊執行(#29442),並修正分塊 cuBLAS 矩陣乘法未遵循目標 stride 的問題,提交由 Johannes Gäßler 參與。改動僅涉及 CUDA 後端,影響用 NVIDIA GPU 做本地推理的場景;該建置同時覆蓋 macOS Apple Silicon 與 Linux、Windows 的 CUDA 12/13 等平台。