b11436: ggml-openvino: fix CI tests; fix GPU regressions. (#30037)
llama.cpp 的 ggml-openvino 後端在 build b11436 修復 CI 測試與 GPU 迴歸。上游 #29622 給輸入 embedding 圖加了混合 token/embd 分支,其中的 DUP 不受支援,導致排程器拆圖、首次單 token decode 在 CPU 與 GPU 失敗,現改為只從 compute 節點建置 OV 模型。
原標題:b11436: ggml-openvino: fix CI tests; fix GPU regressions. (#30037)
閱讀原文
同一事件共有 2 則報導(1 個來源),看事件全貌
| 評分 | 48 / 48(平均 48,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
ggml-openvino: skip unselected graph branches and support DUP
Upstream #29622 adds a mixed token/embd branch to every input
embedding graph through ggml_build_forward_select(). Its nodes are
not flagged for compute, but the backend translated them anyway,
and the DUP in that branch was unsupported, so the scheduler split
the graph and passed the embeddings across the split with a fixed
token count. The first single-token decode then failed
(test-thread-safety on CPU and GPU).
Build the OV model from the compute nodes only, and translate a
same-type contiguous DUP like CONT so the graph stays on one backend.
ggml-openvino: make inp_scale_rows token dim dynamic
#29622 also moves the per-token embedding scale (gemma3, gemma3n,
gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim
and pad it per chunk on the static (NPU) path.
ggml-openvino: skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights
Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU
plugin fails to compile that form for some row counts with "clFinish,
error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on
the MUL_MAT cases added in #29869 (e.g. m=1000, n=2, k=1024). Model
weights use a u4 zero point and are not affected.
Report these cases as unsupported on GPU until the plugin is fixed.
Op tests check support before allocating, so the check matches unbound
weights only; model loading probes with a dummy buffer and keeps its
weights on the GPU.
ggml-openvino: create FILL in the output type
translate_fill always built an f32 constant, so an f16 FILL produced
f32 data and the copy back overran the f16 output buffer. Use the
output type for the constant.
ggml-openvino: reject CONCAT with a quantized type
Quantized inputs are dequantized when translated, so the backend cannot
write a quantized CONCAT output. Report it as unsupported, as for CPY
to a quantized type.
ggml-openvino: handle the single recurrent state gather of build_rs
#29856 changed build_rs to gather all recurrent states with one GET_ROWS
on the s_copy leaf and take the ubatch and extra states as views of it.
The stateful path matched only the previous form, a GET_ROWS per view of
s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU
("is_axis_valid(axis, r)" in a Concat).
For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the
active-state gather, keep the rank-4 layout of reshapes that read a view
of it, and map the copy of the empty extra-state view to the single-slot
remainder writeback. Do not warn about the dynamic dim of empty views.
openvino: align eltwise operand ranks to work around a GPU-plugin defect
openvino: match the MoE fusion on the rank-3 stateful graph
ggml-openvino: do not unsqueeze an RMS norm output in AlignEltwiseOperandRanks
The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract
whose operand ranks differ. In gemma-3 the lower-rank operand of the
post-attention residual add is the norm output, and unsqueezing it makes
the GPU plugin compute the layer wrongly: gemma-3 returns empty answers
on GPU with stateful execution.
Skip the rewrite when the lower-rank operand is an RMS norm output.
docs : update OpenVINO validated models
Co-authored-by: Mustafa Cavus mustafa.cavus@intel.com
相關報導
llama.cpp releases10/5 09:19AI 評分12
llama.cpp 發布 b11407 版本,修復了在開啟 -DGGML_VULKAN_RUN_TESTS=ON 時 Vulkan 後端出現的未宣告識別符號問題。該版本同步提供 macOS、Linux、Windows、Android 等多平台預編譯包,涵蓋 CPU、CUDA 12/13、Vulkan、ROCm 10.0、OpenVINO、SYCL 及 Snapdragon 的 Adreno GPU 與 Hexagon NPU 等後端。
llama.cpp releases● 精選10/3 12:54AI 評分70
ggml-openvino 後端更新到 2026.4.1,重點最佳化 Qwen3.5 MoE 推理效能,並新增推理效能分析、預設使用遠端輸出張量、磁碟快取模型直載(cache_only)等能力。在 Arc B390 上以 q4_asym64_all 重量化執行 gemma-4-26B-A4B,pp512 從 66.16 t/s 提升到 1608.73 t/s,tg128 從 25.94 提升到 26.46 t/s,困惑度基本不變。
llama.cpp releases10/4 20:50AI 評分23
llama.cpp 發布建置版本 b11397,主要變更是將 CUDA 後端的 neu_padded 移動到實際使用位置(PR #29940),作者為 Hugging Face 的 Adrien Gallouët。該版本繼續提供 macOS、Linux、Windows、Android 等多平台預編譯包,覆蓋 Apple Silicon、CUDA 12/13、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;
llama.cpp releases10/4 13:34AI 評分36
llama.cpp 發布 b11389 版本,修復了 Vulkan 後端在 RDNA4 架構上的矩陣向量運算調優問題(PR #29934)。該版本繼續提供覆蓋 macOS、Linux、Windows、Android 等平台的預編譯二進位制,包括 Vulkan、CUDA、ROCm、SYCL 等後端,其中 macOS Apple Silicon 的 KleidiAI 啟用版和 openEuler 建置被停用。
llama.cpp releases10/5 09:36AI 評分48
llama.cpp 發布 b11412,修復 k-pool 模型(qwen4exp、glm5-next)解碼時意外重新預留計算圖並中止的問題。原因是兩模型按 cache_safe、n_tokens 等 reserve 無法預知的狀態分支,解碼圖與預留圖節點數不一致(如 7564 對 7762),解碼時被迫按當前狀態重預留、丟掉最壞情況尺寸,在 GGML_SCHED_DEBUG_REALLOC=1 下直接 abort。