AISpot

llama.cpp 為 Metal 加入 F16 KV 的 tensor API flash attention 核心

llama.cpp releases10/3 01:28版本更新地端推論開源開發工具

llama.cpp 合併 PR #29570,在 Metal 後端新增基於 tensor API 的 flash attention 核心,支援 F16 KV 快取。同一批提交還補充了 DK=DV=512、DK=576/DV=512、DK=192/DV=128 等形狀的 tensor FA 核心,並支援 attention sinks、ALiBi 與 logit softcap。

你準備用 Mac Studio 跑本地開源模型,這條直接關係到 Apple Silicon 上長上下文推理的視訊記憶體佔用與速度,可等新版 llama.cpp 發布後實測對比 MLX 與 Ollama 的表現。

原標題:b11362
閱讀原文

評分78 / 82(平均 80,門檻 60)
狀態精選

全文翻譯

metal:為 F16 KV 新增 tensor API flash attention 核心(#29570) metal:為 F16 KV 新增 tensor API flash attention 核心 cont:為 DK=DV=512 和 DK=576、DV=512 新增 tensor FA 核心 cont:在 tensor FA 核心中支援 attention sinks、ALiBi 和 logit softcap cont:為 DK=192、DV=128 新增 tensor FA 核心 網站: https://llama.app 證明: https://github.com/ggml-org/llama.cpp/attestations/52326292 macOS/iOS: macOS Apple Silicon(arm64) macOS Apple Silicon(arm64,啟用 KleidiAI)已停用 macOS Intel(x64) iOS XCFramework Linux: Ubuntu x64(CPU) Ubuntu arm64(CPU) Ubuntu s390x(CPU) Ubuntu x64(Vulkan) Ubuntu arm64(Vulkan) Ubuntu x64(CUDA 12)- CUDA 12.8 庫 Ubuntu x64(CUDA 13)- CUDA 13.4 庫 Ubuntu arm64(CUDA 13)- CUDA 13.4 庫 Ubuntu x64(ROCm 10.0) Ubuntu x64(OpenVINO) Ubuntu x64(SYCL FP32) Ubuntu x64(SYCL FP16) Linux arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU)- 設定指南 Android: Android arm64(CPU) Android arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU)- 設定指南 Windows: Windows x64(CPU) Windows arm64(CPU) Windows arm64(OpenCL Adreno) Windows x64(CUDA 12)- CUDA 12.4 DLL Windows x64(CUDA 13)- CUDA 13.4 DLL Windows arm64(CUDA 13)- CUDA 13.4 DLL Windows x64(Vulkan) Windows x64(OpenVINO) Windows x64(SYCL) Windows x64(ROCm 10.0) openEuler: 已停用 openEuler x86(310p) openEuler x86(910b,ACL Graph) openEuler aarch64(310p) openEuler aarch64(910b,ACL Graph) UI: UI

由 AI 翻譯,以原文為準。

原文
metal : add tensor API flash attention kernel for F16 KV ( #29570 ) metal : add tensor API flash attention kernel for F16 KV cont : add tensor FA kernels for DK=DV=512 and DK=576, DV=512 cont : support attention sinks, ALiBi and logit softcap in the tensor FA kernel cont : add tensor FA kernel for DK=192, DV=128 Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52326292 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI