llama.cpp 為 Metal 加入 F16 KV 的 tensor API flash attention 核心
llama.cpp 合併 PR #29570,在 Metal 後端新增基於 tensor API 的 flash attention 核心,支援 F16 KV 快取。同一批提交還補充了 DK=DV=512、DK=576/DV=512、DK=192/DV=128 等形狀的 tensor FA 核心,並支援 attention sinks、ALiBi 與 logit softcap。
你準備用 Mac Studio 跑本地開源模型,這條直接關係到 Apple Silicon 上長上下文推理的視訊記憶體佔用與速度,可等新版 llama.cpp 發布後實測對比 MLX 與 Ollama 的表現。
原標題:b11362
閱讀原文
| 評分 | 78 / 82(平均 80,門檻 60) |
|---|---|
| 狀態 | 精選 |
全文翻譯
metal:為 F16 KV 新增 tensor API flash attention 核心(#29570)
metal:為 F16 KV 新增 tensor API flash attention 核心
cont:為 DK=DV=512 和 DK=576、DV=512 新增 tensor FA 核心
cont:在 tensor FA 核心中支援 attention sinks、ALiBi 和 logit softcap
cont:為 DK=192、DV=128 新增 tensor FA 核心
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/52326292
macOS/iOS:
macOS Apple Silicon(arm64)
macOS Apple Silicon(arm64,啟用 KleidiAI)已停用
macOS Intel(x64)
iOS XCFramework
Linux:
Ubuntu x64(CPU)
Ubuntu arm64(CPU)
Ubuntu s390x(CPU)
Ubuntu x64(Vulkan)
Ubuntu arm64(Vulkan)
Ubuntu x64(CUDA 12)- CUDA 12.8 庫
Ubuntu x64(CUDA 13)- CUDA 13.4 庫
Ubuntu arm64(CUDA 13)- CUDA 13.4 庫
Ubuntu x64(ROCm 10.0)
Ubuntu x64(OpenVINO)
Ubuntu x64(SYCL FP32)
Ubuntu x64(SYCL FP16)
Linux arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU)- 設定指南
Android:
Android arm64(CPU)
Android arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU)- 設定指南
Windows:
Windows x64(CPU)
Windows arm64(CPU)
Windows arm64(OpenCL Adreno)
Windows x64(CUDA 12)- CUDA 12.4 DLL
Windows x64(CUDA 13)- CUDA 13.4 DLL
Windows arm64(CUDA 13)- CUDA 13.4 DLL
Windows x64(Vulkan)
Windows x64(OpenVINO)
Windows x64(SYCL)
Windows x64(ROCm 10.0)
openEuler:
已停用
openEuler x86(310p)
openEuler x86(910b,ACL Graph)
openEuler aarch64(310p)
openEuler aarch64(910b,ACL Graph)
UI:
UI
由 AI 翻譯,以原文為準。
原文
metal : add tensor API flash attention kernel for F16 KV ( #29570 )
metal : add tensor API flash attention kernel for F16 KV
cont : add tensor FA kernels for DK=DV=512 and DK=576, DV=512
cont : support attention sinks, ALiBi and logit softcap in the tensor FA kernel
cont : add tensor FA kernel for DK=192, DV=128
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52326292
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI