AISpot

llama.cpp 為 WebGPU 後端加入 f16 支援

llama.cpp releases10/4 01:37版本更新地端推論開源開發工具

llama.cpp 在 WebGPU 後端為 fill/set_rows 操作加入 f16 支援(PR #29897),併發布 b11382 版本。該版本覆蓋 macOS Apple Silicon、iOS、Linux、Android、Windows 等平台,其中 macOS Apple Silicon 的 KleidiAI 建置被停用。對使用 WebGPU 推理的使用者,f16 支援可減少視訊記憶體佔用並提升相容性。

你準備用 Mac Studio 跑本地模型,llama.cpp 的 WebGPU 後端改進會直接影響瀏覽器或跨平台推理的 f16 效能與記憶體佔用。

原標題:b11382
閱讀原文

評分72 / 55(平均 63,門檻 60)
狀態精選

全文翻譯

webgpu:為 fill/set_rows 新增 f16 支援(#29897) 網站: https://llama.app 證明: https://github.com/ggml-org/llama.cpp/attestations/52497832 macOS/iOS: macOS Apple Silicon(arm64) macOS Apple Silicon(arm64,啟用 KleidiAI)已停用 macOS Intel(x64) iOS XCFramework Linux: Ubuntu x64(CPU) Ubuntu arm64(CPU) Ubuntu s390x(CPU) Ubuntu x64(Vulkan) Ubuntu arm64(Vulkan) Ubuntu x64(CUDA 12)- CUDA 12.8 庫 Ubuntu x64(CUDA 13)- CUDA 13.4 庫 Ubuntu arm64(CUDA 13)- CUDA 13.4 庫 Ubuntu x64(ROCm 10.0) Ubuntu x64(OpenVINO) Ubuntu x64(SYCL FP32) Ubuntu x64(SYCL FP16) Linux arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU)- 安裝指南 Android: Android arm64(CPU) Android arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU)- 安裝指南 Windows: Windows x64(CPU) Windows arm64(CPU) Windows arm64(OpenCL Adreno) Windows x64(CUDA 12)- CUDA 12.4 DLL Windows x64(CUDA 13)- CUDA 13.4 DLL Windows arm64(CUDA 13)- CUDA 13.4 DLL Windows x64(Vulkan) Windows x64(OpenVINO) Windows x64(SYCL) Windows x64(ROCm 10.0) openEuler: 已停用 openEuler x86(310p) openEuler x86(910b,ACL Graph) openEuler aarch64(310p) openEuler aarch64(910b,ACL Graph) UI: UI

由 AI 翻譯,以原文為準。

原文
webgpu: add f16 support to fill/set_rows ( #29897 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52497832 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/3 00:18AI 評分34

llama.cpp b11355 在 32KB 共享記憶體的三星 GPU 上停用 Vulkan 大 matmul tile

llama.cpp 發布 b11355 版:對共享記憶體為 32KB 的三星 GPU,Vulkan 後端停用大型 matmul tile(#28531),該提交由 Claude Op 協助完成。此版本中 macOS Apple Silicon 的 KleidiAI 啟用建置(arm64, KleidiAI enabled)與 openEuler 建置標記為 DISABLED;