llama.cpp 新增啟用統計與 GGUF imatrix 支援
llama.cpp 合併 PR #14891,為 GGUF 格式 imatrix 引入基於啟用值的統計計算,新增 --activation-statistics 引數。該功能可計算啟用熵、餘弦相似度、L2 範數及歐氏-餘弦分數(ECS),並支援外部 NextN 草稿檔案(-md / --model-draft)。預設不啟用啟用統計以避免 imatrix 體積翻倍,統計報告佈局與排序也按 imatrix 型別更新。
原標題:b11388
閱讀原文
| 評分 | 58 / 42(平均 50,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
imatrix: calculate activation-based statistics for new format (GGUF) imatrices ( #14891 )
Use activations to calculate the stats
Determine calculation mode
Compute entropy for activations
Compute cosine similarity based on activations
Compute l2 norm
Add compute_layer_statistics() function
Update aggregated statistic report layout
Fix printing l2 norm when calc_mode = 1
Refactor variable name
Compute aggregated (per layer) l2 norm
Update aggregated sum of squared activations per layer
Make ZD Score two-tailed
Update report layout
Reverse conditional logic to match convention
Rename report heading
Add --activation-statistics parameter
Add Euclidean–Cosine Score (ECS)
Add --activation-statistics logic to avoid doubling the imatrix size by default
Update stats output sort based on imatrix type
Process external NextN draft files (-md / --model-draft)
Refactor to use new llama_batch_ext
Co-authored-by: compilade git@compilade.net
Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52565482
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/3 09:36AI 評分43
llama.cpp 合併 PR #29831,為 clef 決策模型加入初始支援,目前僅限純文字輸入。該改動同時更新了 gguf-py 常量並清理了靜態圖相關程式碼,隨附的建置產物覆蓋 macOS、Linux、Windows、Android 與 openEuler 等多個平台。
r/LocalLLaMA top of the day10/3 07:18AI 評分70
llama.cpp 的 PR #29825 通過最佳化索引器分數記憶體,把 Qwen Flash Next 等 Qwen4 模型的視訊記憶體佔用減半。實測在上下文 131072、ub 2048 時從 3.0 GiB 降至 1.5 GiB,262144、ub 4096 時從 11.7 GiB 降至 5.6 GiB。作者稱 logits 逐位一致、pp/tg 速度不變,僅分數因 fp32 舍入有差異;有使用者反饋解碼速度約降 10%。
llama.cpp releases10/4 15:12AI 評分16
llama.cpp 發布 b11391 建置版本,唯一程式碼改動是將 CUDA 後端的 blocks_per_col 變數移到實際使用位置(PR #29939),由 Hugging Face 的 Adrien Gallouët 提交。
r/LocalLLaMA top of the day10/2 19:47AI 評分65
llama.cpp 的 PR #29184 提出在 CUDA 後端把共享專家融合進 MMVQ,為部分 MoE 架構帶來提速,例如 Qwen 35B A3B。有評論稱該改動使 TG 速度提升約 5%,並指出僅對部分 MoE 架構有效。
llama.cpp releases10/4 13:34AI 評分36
llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。