AISpot

llama.cpp 新增啟用統計與 GGUF imatrix 支援

llama.cpp releases10/4 11:42版本更新地端推論開源開發工具

llama.cpp 合併 PR #14891,為 GGUF 格式 imatrix 引入基於啟用值的統計計算,新增 --activation-statistics 引數。該功能可計算啟用熵、餘弦相似度、L2 範數及歐氏-餘弦分數(ECS),並支援外部 NextN 草稿檔案(-md / --model-draft)。預設不啟用啟用統計以避免 imatrix 體積翻倍,統計報告佈局與排序也按 imatrix 型別更新。

原標題:b11388
閱讀原文

評分58 / 42(平均 50,門檻 60)
狀態未入選

原文

imatrix: calculate activation-based statistics for new format (GGUF) imatrices ( #14891 ) Use activations to calculate the stats Determine calculation mode Compute entropy for activations Compute cosine similarity based on activations Compute l2 norm Add compute_layer_statistics() function Update aggregated statistic report layout Fix printing l2 norm when calc_mode = 1 Refactor variable name Compute aggregated (per layer) l2 norm Update aggregated sum of squared activations per layer Make ZD Score two-tailed Update report layout Reverse conditional logic to match convention Rename report heading Add --activation-statistics parameter Add Euclidean–Cosine Score (ECS) Add --activation-statistics logic to avoid doubling the imatrix size by default Update stats output sort based on imatrix type Process external NextN draft files (-md / --model-draft) Refactor to use new llama_batch_ext Co-authored-by: compilade git@compilade.net Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52565482 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

r/LocalLLaMA top of the day10/3 07:18AI 評分70

llama.cpp 新 PR 將 Qwen4 索引器視訊記憶體減半

llama.cpp 的 PR #29825 通過最佳化索引器分數記憶體,把 Qwen Flash Next 等 Qwen4 模型的視訊記憶體佔用減半。實測在上下文 131072、ub 2048 時從 3.0 GiB 降至 1.5 GiB,262144、ub 4096 時從 11.7 GiB 降至 5.6 GiB。作者稱 logits 逐位一致、pp/tg 速度不變,僅分數因 fp32 舍入有差異;有使用者反饋解碼速度約降 10%。

llama.cpp releases10/4 13:34AI 評分36

llama.cpp 修復 Vulkan 後端 RDNA4 矩陣向量調優

llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。