AISpot

llama.cpp 為 SpacemiT X60 補上 Q8_0 IME1 核心

llama.cpp releases10/5 08:34版本更新地端推論開源開發工具

llama.cpp 合併 PR #28479,為 SpacemiT X60 新增 Q8_0 的 IME1 矩陣核心與 repack 路徑。此前該平台 IME 加速只覆蓋 Q4_0/Q4_1/Q4_K,且建置時 GGML_CPU_REPACK=OFF,Q8_0 無加速路徑,prefill 比 Q4_0 慢約十倍。

首次讓 SpacemiT X60 上的 Q8_0 獲得 IME1 加速,prefill 吞吐提升近 9 倍,對在 RISC-V 開發板本地跑 llama.cpp 的使用者有直接參考價值。

原標題:b11408
閱讀原文

評分68 / 45(平均 56,門檻 60)
狀態未入選

原文

ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 ( #28479 ) ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 On the SpacemiT X60, IME matrix acceleration only covered Q4_0/Q4_1/Q4_K. Q8_0 had no IME1 kernel, and since the SpacemiT build sets GGML_CPU_REPACK=OFF there was no repack path compiled in either, so Q8_0 had no accelerated path at all and ran roughly ten times slower than Q4_0 for prefill on the same board. add make_block_q8_0x16 and the Q8_0 repack entry: interleave the weights into the 16-column layout the IME1 vmadot sequence expects add ime1::gemm_kernel_i8i8, an int8 x int8 IME1 kernel with a single-row and a 4-row A path; the 4-row path loads each B panel once and reuses it across 4 rows of A add quantize_a_4row_i8 for the 4-row activation quantization wire both into forward_mul_mat and the repack factory for Q8_0 docs: mark Q8_0 as supported on X60 Correctness was checked against a quant-exact integer reference for K = 32 up to 4096, with a max relative error of about 1e-6, and by checking that generation stays coherent across several prompts. Tested on Milk-V Jupiter (SpacemiT X60), Bianbu 2.1.1, gcc 14.2, with Qwen2.5-0.5B-Instruct Q8_0. llama-bench -t 4 under taskset -c 0-3, 5 repetitions on an idle board: pp128 goes from 10.70 to 93.87 t/s. Q4_0 is unchanged at 106.40 -> 107.51 t/s, as expected since this does not touch that path. ggml-cpu : move q8_0_16x32 decl to IME1 section ggml-cpu : align q8_0 IME1 kernel assignments Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52757256 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/3 02:58AI 評分40

llama.cpp 修復 qkx3 量化縮放搜尋的非法舍入

llama.cpp 合併 PR #29817,修復 ggml-quants 中 qkx3 量化縮放搜尋的非法舍入問題:當 imatrix 擬合的最小值塌縮為最大值或範圍極小時,會產生無窮、NaN 或越界值,傳入 nearest_int 後可能觸發 Debug 建置斷言。修復方式是在舍入前把量化級別鉗制到 [0, nmax],並新增 q2_K、q4_K、q5_K、q4_1、q5_1 的退化 imatrix 組迴歸測試,對應 issue #29804。

llama.cpp releases10/2 14:11AI 評分43

llama.cpp 為 Hexagon NPU 增加 q2_k 與 q3_k 量化支援

llama.cpp 合併 PR #29717,為 Hexagon 後端新增 q2_k 和 q3_k 兩種量化型別支援,並統一了 src1_row_size 的分配方式。該改動由高通工程師 Max Krasnyansky 參與提交,面向在驍龍 Hexagon NPU 上執行本地推理的使用者,相關建置覆蓋 Linux arm64 與 Android arm64 的 Snapdragon CPU、Adreno GPU、Hexagon NPU 組合。

r/LocalLLaMA top of the day10/3 07:18AI 評分70

llama.cpp 新 PR 將 Qwen4 索引器視訊記憶體減半

llama.cpp 的 PR #29825 通過最佳化索引器分數記憶體,把 Qwen Flash Next 等 Qwen4 模型的視訊記憶體佔用減半。實測在上下文 131072、ub 2048 時從 3.0 GiB 降至 1.5 GiB,262144、ub 4096 時從 11.7 GiB 降至 5.6 GiB。作者稱 logits 逐位一致、pp/tg 速度不變,僅分數因 fp32 舍入有差異;有使用者反饋解碼速度約降 10%。