llama.cpp 修復 qkx3 量化縮放搜尋的非法舍入
llama.cpp 合併 PR #29817,在 qkx3 量化級別送入 nearest_int 前先鉗制到 [0, nmax],避免 imatrix 縮放搜尋產生無窮、NaN 或越界值時觸發 Debug 建置斷言。該問題由 #29804 報告,修復同時為 q2_K、q4_K、q5_K、q4_1、q5_1 的退化 imatrix 分組補充迴歸測試。合法範圍內的取值行為不變,macOS Apple Silicon(arm64)等預編譯包已隨本次發布更新。
本地用 llama.cpp 在 Apple Silicon 上跑量化模型時,這個修復能避免退化 imatrix 分組導致的量化崩潰或斷言失敗,建議更新到含該提交的建置。
原標題:b11366
閱讀原文
| 評分 | 62 / 55(平均 58,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
ggml-quants : avoid invalid rounding in qkx3 scale search ( #29817 )
ggml-quants : avoid invalid rounding in qkx3 scale search
The imatrix scale search can produce an infinite, NaN, or otherwise out-of-range value when the fitted minimum collapses to the maximum or makes the range extremely small. That value is then passed to nearest_int and can trip its assertion in Debug builds.
Clamp the quantization level to [0, nmax] before rounding so valid in-range values behave the same as before while invalid scale-search results no longer reach nearest_int.
Add regression coverage for degenerate imatrix groups across q2_K, q4_K, q5_K, q4_1, and q5_1.
Fixes #29804 .
Assisted-by: Claude Opus 5.5
tests: print degenerate imatrix quant types
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52340657
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases● 精選10/2 14:11AI 評分60
llama.cpp b11345 為 Hexagon 後端新增 Q2_K 和 Q3_K 量化型別支援,對應 PR #29717,並統一了 src1_row_size 的分配;改動由 Max Krasnyansky(maxk@qti.qualcomm.com)參與提交。
r/LocalLLaMA top of the day● 精選10/3 07:18AI 評分79
llama.cpp 的 PR #29825 把 Qwen Flash Next 的 indexer score 視訊記憶體佔用減半,例如 131072 上下文、ub 2048 時從 3.0 GiB 降到 1.5 GiB,262144 上下文、ub 4096 時從 11.7 GiB 降到 5.6 GiB。作者稱 logits 完全一致,僅分數因 fp32 舍入有差異,pp/tg 速度不變,test-llama-archs 通過。
llama.cpp releases● 精選10/2 23:23AI 評分65
llama.cpp 發布 b11352 版本,主要改動是最佳化 qwen4exp 的掩碼(mask)建置(#29824),並將同一改動應用到 GLM5-next。該版本繼續提供 macOS Apple Silicon(arm64)、iOS、Linux、Windows、Android 等平台建置,涵蓋 CUDA 12/13、Vulkan、ROCm 10.0、OpenVINO、SYCL 與 Snapdragon 後端。
llama.cpp releases10/2 14:52AI 評分48
llama.cpp 發布建置 b11346,本次改動是修復 qwen4exp 的測試(PR #29819)。該版本繼續提供 macOS Apple Silicon(arm64)與 Intel(x64)、iOS XCFramework 預編譯包,Linux 與 Windows 覆蓋 CPU、Vulkan、CUDA 12.8/13.4、ROCm 10.0、OpenVINO、SYCL 等後端,另有 Android arm64 與 openEuler 版本。
llama.cpp releases● 精選10/3 01:28AI 評分80
llama.cpp 合併 PR #29570,在 Metal 後端新增基於 tensor API 的 flash attention 核心,支援 F16 KV 快取。同一批提交還補充了 DK=DV=512、DK=576/DV=512、DK=192/DV=128 等形狀的 tensor FA 核心,並支援 attention sinks、ALiBi 與 logit softcap。