AISpot
所屬事件:llama.cpp 批次內混合 embd 與原始 token,並修復 OpenVINO 迴歸(1 個來源 · 2 則報導)

llama.cpp 於 2026 年 10 月 5 日合併 PR #29622(build b11400),支援在單一批次中同時傳入嵌入向量(embd)與原始 token;次日發布的 build b11436 修復了該改動給 ggml-openvino 後端帶來的迴歸。據 llama.cpp releases 說明,#… 看完整事件

b11436: ggml-openvino: fix CI tests; fix GPU regressions. (#30037)

llama.cpp releases10/6 07:57版本更新地端推論開源

llama.cpp 的 ggml-openvino 後端在 build b11436 修復 CI 測試與 GPU 迴歸。上游 #29622 給輸入 embedding 圖加了混合 token/embd 分支,其中的 DUP 不受支援,導致排程器拆圖、首次單 token decode 在 CPU 與 GPU 失敗,現改為只從 compute 節點建置 OV 模型。

原標題:b11436: ggml-openvino: fix CI tests; fix GPU regressions. (#30037)
閱讀原文

同一事件共有 2 則報導(1 個來源),看事件全貌

評分48 / 48(平均 48,門檻 60)
狀態未入選

原文

ggml-openvino: skip unselected graph branches and support DUP Upstream #29622 adds a mixed token/embd branch to every input embedding graph through ggml_build_forward_select(). Its nodes are not flagged for compute, but the backend translated them anyway, and the DUP in that branch was unsupported, so the scheduler split the graph and passed the embeddings across the split with a fixed token count. The first single-token decode then failed (test-thread-safety on CPU and GPU). Build the OV model from the compute nodes only, and translate a same-type contiguous DUP like CONT so the graph stays on one backend. ggml-openvino: make inp_scale_rows token dim dynamic #29622 also moves the per-token embedding scale (gemma3, gemma3n, gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim and pad it per chunk on the static (NPU) path. ggml-openvino: skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU plugin fails to compile that form for some row counts with "clFinish, error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on the MUL_MAT cases added in #29869 (e.g. m=1000, n=2, k=1024). Model weights use a u4 zero point and are not affected. Report these cases as unsupported on GPU until the plugin is fixed. Op tests check support before allocating, so the check matches unbound weights only; model loading probes with a dummy buffer and keeps its weights on the GPU. ggml-openvino: create FILL in the output type translate_fill always built an f32 constant, so an f16 FILL produced f32 data and the copy back overran the f16 output buffer. Use the output type for the constant. ggml-openvino: reject CONCAT with a quantized type Quantized inputs are dequantized when translated, so the backend cannot write a quantized CONCAT output. Report it as unsupported, as for CPY to a quantized type. ggml-openvino: handle the single recurrent state gather of build_rs #29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU ("is_axis_valid(axis, r)" in a Concat). For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the active-state gather, keep the rank-4 layout of reshapes that read a view of it, and map the copy of the empty extra-state view to the single-slot remainder writeback. Do not warn about the dynamic dim of empty views. openvino: align eltwise operand ranks to work around a GPU-plugin defect openvino: match the MoE fusion on the rank-3 stateful graph ggml-openvino: do not unsqueeze an RMS norm output in AlignEltwiseOperandRanks The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract whose operand ranks differ. In gemma-3 the lower-rank operand of the post-attention residual add is the norm output, and unsqueezing it makes the GPU plugin compute the layer wrongly: gemma-3 returns empty answers on GPU with stateful execution. Skip the rewrite when the lower-rank operand is an RMS norm output. docs : update OpenVINO validated models Co-authored-by: Mustafa Cavus mustafa.cavus@intel.com

相關報導

llama.cpp releases10/5 09:19AI 評分12

llama.cpp 發布 b11407,修復 Vulkan 測試編譯錯誤

llama.cpp 發布 b11407 版本,修復了在開啟 -DGGML_VULKAN_RUN_TESTS=ON 時 Vulkan 後端出現的未宣告識別符號問題。該版本同步提供 macOS、Linux、Windows、Android 等多平台預編譯包,涵蓋 CPU、CUDA 12/13、Vulkan、ROCm 10.0、OpenVINO、SYCL 及 Snapdragon 的 Adreno GPU 與 Hexagon NPU 等後端。

llama.cpp releases● 精選10/3 12:54AI 評分70

llama.cpp 的 OpenVINO 後端更新至 2026.4.1 並大幅最佳化 MoE 效能

ggml-openvino 後端更新到 2026.4.1,重點最佳化 Qwen3.5 MoE 推理效能,並新增推理效能分析、預設使用遠端輸出張量、磁碟快取模型直載(cache_only)等能力。在 Arc B390 上以 q4_asym64_all 重量化執行 gemma-4-26B-A4B,pp512 從 66.16 t/s 提升到 1608.73 t/s,tg128 從 25.94 提升到 26.46 t/s,困惑度基本不變。

llama.cpp releases10/4 20:50AI 評分23

llama.cpp 發布 b11397,CUDA 後端調整 neu_padded 位置

llama.cpp 發布建置版本 b11397,主要變更是將 CUDA 後端的 neu_padded 移動到實際使用位置(PR #29940),作者為 Hugging Face 的 Adrien Gallouët。該版本繼續提供 macOS、Linux、Windows、Android 等多平台預編譯包,覆蓋 Apple Silicon、CUDA 12/13、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;

llama.cpp releases10/4 13:34AI 評分36

llama.cpp 修復 Vulkan 後端 RDNA4 矩陣向量調優

llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。

llama.cpp releases10/5 09:36AI 評分48

llama.cpp 修復 k-pool 模型圖重分配導致的解碼中止

llama.cpp 發布 b11412,修復 k-pool 模型(qwen4exp、glm5-next)解碼時意外重新預留計算圖並中止的問題。原因是兩模型按 cache_safe、n_tokens 等 reserve 無法預知的狀態分支,解碼圖與預留圖節點數不一致(如 7564 對 7762),解碼時被迫按當前狀態重預留、丟掉最壞情況尺寸,在 GGML_SCHED_DEBUG_REALLOC=1 下直接 abort。