llama.cpp 的 OpenVINO 後端更新,MoE 推理大幅提速
ggml-openvino 後端更新至 2026.4.1,新增 MoE 路由融合與 GDN 歸一化融合,並預設在 GPU 上啟用 MoE 融合。在 Arc B390 上跑 gemma-4-26B-A4B,配合 q4_asym64 重新量化,pp512 從 66.16 t/s 提升到 1608.73 t/s,困惑度基本不變。同時修復了 stateful 執行下 rank-3 軸處理導致的 MoE 崩潰,並改進裝置列舉與記憶體上報。
原標題:b11374
閱讀原文
| 評分 | 55 / 30(平均 42,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. ( #29852 )
ggml-openvino : Qwen3.5 MoE perf ( #312 )
Squash of ravi9#312 :
ggml-openvino: add detailed inference profiling (Yu, Zijun)
ggml-openvino: use remote output tensors by default (Yu, Zijun)
ggml-openvino: optimize single-sequence recurrent state (Yu, Zijun)
opt1: remove recurrent reset for single sequence, opt2: direct gdn outputs (break parallel sequence) (Yu, Zijun)
fix parallel sequences (Yu, Zijun)
ggml-openvino: simplify graph cache key (ynimmaga)
enable stateful for qwen35 single sequence (Yu, Zijun)
Fix after rebasing (Yu, Zijun)
Add k-requant option q4_asym64 (Yu, Zijun)
Fix qwen35 llama-bench -p 0 (Yu, Zijun)
Simplify RESHAPE translation (Yu, Zijun)
openvino: fuse MoE routing (Yu, Zijun)
openvino: fuse GDN qk normalization (Yu, Zijun)
openvino: enable GPU MoE fusion by default (Yu, Zijun)
ggml-openvino: add cache_only mode to import cached compiled model on disk directly (Yu, Zijun)
openvino : report the device allocation limit to ggml (Łukasz Ślusarczyk)
Fix windows build (Yu, Zijun)
Co-authored-by: ynimmaga ynimmaga@users.noreply.github.com
Co-authored-by: Łukasz Ślusarczyk lukasz.slusarczyk@intel.com
ggml-openvino: Update doc of compiled model cache
openvino: implement PRD-compliant device enumeration and memory reporting
openvino: fix multi-device listing issues from review
Only the device selected by GGML_OPENVINO_DEVICE reports as GPU; the
other OpenVINO devices report as IGPU so llama.cpp does not offload to
them. Initializing a non-selected device logs a warning.
Name devices OPENVINO again and show the OpenVINO id in the
description. Raw "CPU" names shadowed the ggml CPU backend.
Support GPU.N: create the OpenCL queue on OpenVINO's own context for
the selected device, and replace "GPU"/"NPU" string comparisons with
ggml_openvino_is_gpu()/ggml_openvino_is_npu().
An unavailable GGML_OPENVINO_DEVICE is now an error that lists the
available devices, instead of silently falling back to CPU.
Memory: cap iGPU/NPU free memory at system available memory, fall back
to system memory instead of 0/0 when the plugin lacks memory
properties, and ignore host USM allocations in GPU usage.
Initialize the device config once under a lock, even if OpenCL setup
fails.
Fix supports_op return type for non-selected devices (build error).
openvino : take USM entry points from the selected device platform
clGetExtensionFunctionAddressForPlatform was called on the first platform
returned by clGetPlatformIDs. The address it returns is only valid for the
platform it was queried on, and the first platform is not always the one that
holds the device OpenVINO selected.
On a host whose first platform comes from another vendor the lookup returns
null, and then every read, write and memset on a GPU buffer fails with
"clEnqueueMemcpyINTEL not available".
Look both entry points up in init(), on the platform of the device OpenVINO
picked, and keep them in the device config next to the command queue.
Assisted-by: Claude Opus 5
openvino: fuse MoE experts for models with a fused gate_up weight
FuseMoeCompressed only matches models whose gate and up projections are
separate GatherMatmul ops. gemma-4 packs both into one expert weight and
splits the result after the GEMM, so its MoE block stayed unfused and ran
the expert GEMMs as per-token GEMVs.
Add FuseMoeCompressedFusedGateUp, which matches that shape
(one GatherMatmul -> Slice/Slice -> Gelu(ERF) -> Multiply) and folds it into
the same MOECompressed op, using GEMM3_SWIGLU with GEGLU_ERF. The fused
weight, scale and zero point are split into gate/up halves by copying raw
bytes, since a graph Slice would be rewritten to StridedSlice and constant
folded, whose reference evaluator crashes on sub-byte types.
gemma-4 also applies a per-expert output scale to the down projection before
the router weights. MOECompressed takes only one per-expert weight, so that
scale is folded into the routing weights, which is exact.
The op reads the zero point straight off a weight port and needs an integer
Constant there, so the matcher requires one and leaves natively quantized
experts (exact f16 zp) to the unfused path.
gemma-4-26B-A4B on Arc B390, GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all,
llama-bench -p 512 -n 128 -r 2, against a GGML_OPENVINO_MOE_OP=0 baseline:
pp512 66.16 -> 1608.73 t/s, tg128 25.94 -> 26.46 t/s. Perplexity over 12
chunks is unchanged (1451.3 +/- 177.9 unfused vs 1427.6 +/- 175.1 fused).
No effect without that requant option, on models with separate gate/up
weights, or on CPU. test-backend-ops -b OPENVINO0 is unchanged by this
commit: two MUL_MAT_ID m_v cases fail, the same two on the unmodified base.
openvino: fix rank-3 axis handling so MoE works under stateful execution
Stateful execution drops the leading size-1 batch dim, so OV tensors are rank
3 while GgmlOvDecoder::get_shape/get_stride still report GGML_MAX_DIMS=4
reversed entries. Several MoE ops derive OV axis indices straight from that
metadata, so they picked the wrong axis. A MoE model with
GGML_OPENVINO_STATEFUL_EXECUTION=1 aborts while building the graph:
Check 'is_axis_valid(axis, r)' failed at src/core/src/validation_util.cpp:336
While validating node 'opset11::TopK ... _ffn_moe_probs ...'
Axis 3 out of the tensor rank range [-3, 2].
Fix idiom throughout: take the axis from the real OV rank, or shift a
metadata-derived axis down by metadata_rank - actual_rank.
argsort.cpp the router top-k axis is 2 on rank 3, not 3. This is the
abort quoted above.
add.cpp the MoE expert-sum bypass collapses the 8-ADD chain into one
ReduceSum on hardcoded axis 2, which on rank 3 reduces n_embd
instead of the expert axis. Now rank-2, with the following
Unsqueeze at rank-3.
get_rows.cpp squeezing a hardcoded {0,1} also strips the batch dim
whenever it is 1, which is every decode step. Squeeze down to
the trailing two dims instead.
mul_mat_id.cpp pick the reshape dims by actual rank, and skip the trailing
Unsqueeze that re-adds the batch dim.
view.cpp the expert-plane slice had the Slice axis, dst_ov_axis, the
ShapeOf+Gather index and the Reshape target all rank-4.
utils.cpp process_view_input_new's "translate_view already resolved
this VIEW, skip re-slicing" shortcut required equal ranks. 4
vs 3 never matched, so every resolved expert plane got
re-sliced. Now compares the common trailing dims. Same axis
shift for the Slice in the view-chain walker.
Stateless is unchanged by construction: every edit is gated on the actual
rank, so axis_shift == 0 reproduces the previous code exactly. Checked on
OV-CPU by diffing greedy output against the unmodified base for dense
gemma-4-E2B, granite-1b-a400m and gemma-4-26B-A4B; all identical.
granite-1b-a400m on OV-CPU aborts with the error above before this change;
after it, it generates and is byte-identical to stateless. Dense gemma-4-E2B
is identical stateless vs stateful both before and after. test-backend-ops
-b OPENVINO0 is unchanged: two pre-existing MUL_MAT_ID m_v cases fail, the
same two on the unmodified base.
gemma-4-26B-A4B is a poor correctness vehicle here. On OV it already drifts
into degenerate repetition a few tokens in, in stateless as much as stateful,
and the two modes diverge somewhere inside that degenerate region instead of
matching token for token. Each mode is self-reproducible across runs.
Known limitation: FuseMoeCompressedFusedGateUp does not match the rank-3
graph, so a MoE model run with GGML_OPENVINO_STATEFUL_EXECUTION=1 loses the
prefill fusion while gaining decode. gemma-4-26B-A4B on Arc B390,
GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all, llama-bench -p 512 -n 128 -r 2:
unfused (GGML_OPENVINO_MOE_OP=0) pp512 66.16 tg128 25.94
fused, stateless (default) pp512 1608.73 tg128 26.46
fused, stateful pp512 66.18 tg128 29.91
Stateful is opt-in and off by default, and MoE did not run there at all
before this, so nothing that previously worked regresses. Making the pass
match rank 3 is the follow-up.
OpenVINO Backend: Upgrade graph cache to use node_idx, src_idx, node type
ggml-openvino : enable more comprehensive conv fusion
enable conv ops
Reject kernel size 0 and support IM2COL_3D
openvino : abort when the GPU remote context cannot be created
init() logged the error and returned, which left the device name a GPU but
remote_context empty. The remote buffer and tensor paths assert only on the
device being a GPU and then dereference that empty optional.
Those paths have no host fallback, and a device that OpenVINO listed should
have a working OpenCL context, so stop instead of continuing. An OpenCL stack
that is broken as a whole is still caught earlier by the device availability
check, which falls back to CPU.
Assisted-by: Claude Opus 5
openvino : fix build warnings
The single-argument form of the OpenVINO RTTI macros is the intended one, but
their selector macro leaves VA_ARGS empty, which -Wpedantic reports on
every pass and op header. Turn that warning off for this backend only, the
way ggml-cuda and ggml-sycl already do for their own third-party warnings.
Also drop a break and a dead assignment around a GGML_ABORT, which is noreturn.
Assisted-by: Claude Opus 5
OpenVINO Backend: Support common MTMD ops
ggml-openvino: give a reshaping view its own ov::Tensor
ggml-openvino : compute HARDSIGMOID and EXPM1 in f32
HARDSIGMOID used a 1/6 constant in the input type, which is not exact
in bf16, and EXPM1 lost precision for small inputs in f16. Both now
compute in f32 and convert back, except on NPU where the f32 path
gives wrong results.
Fixes the HARDSIGMOID/EXPM1 test-backend-ops failures on GPU.
ggml-openvino : update device selection and --list-devices
Show the selecting GGML_OPENVINO_DEVICE value and active device in
--list-devices, startup logs, and backend tests.
Clarify OpenVINO selection uses GGML_OPENVINO_DEVICE, not -dev.
openvino : remove unreachable OpenCL queue checks
A remote buffer exists only on a GPU device, and init() aborts there if the
queue cannot be created, so the queue is never null at these call sites.
Assisted-by: Claude Opus 5
openvino : update OpenVINO to 2026.4.1 and GPU drivers to 26.35.39758.10
docs : update OpenVINO validated models and GPU driver version
ggml-openvino : skip empty views when giving a reshaping view its own tensor
A zero-size view can sit at the end of a GPU USM buffer (Qwen3.5 recurrent cache). Wrapping it as a remote tensor throws "shared USM buffer has smaller size (0)".
Assisted-by: Claude
ggml-openvino : rebind the cached decoder when llama passes a different graph
llama keeps separate graphs for batches with and without outputs. llama-server splits the prompt into chunks for context checkpoints, so a cached decoder could be reused with a graph built in other memory and bind the previous chunk's input tensors. SWA and recurrent models then lost most of the prompt in llama-cli and llama-server.
Assisted-by: Claude
docs : update OpenVINO validated models
Smoke test on Lunar Lake (32 GB) with the two fixes above. Re-add the Qwen3.5 and gemma models.
Assisted-by: Claude
Co-authored-by: Yu, Zijun zijun.yu@intel.com
Co-authored-by: ynimmaga ynimmaga@users.noreply.github.com
Co-authored-by: Łukasz Ślusarczyk lukasz.slusarczyk@intel.com
Co-authored-by: haarika-madaka haarika.madaka@intel.com
Co-authored-by: Mustafa Cavus mustafa.cavus@intel.com
Co-authored-by: Mostafa Faheem mostafaaafaheem@gmail.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52408343
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases● 精選10/2 23:23AI 評分65
llama.cpp 發布 b11352 版本,主要改動是最佳化 qwen4exp 的掩碼(mask)建置(#29824),並將同一改動應用到 GLM5-next。該版本繼續提供 macOS Apple Silicon(arm64)、iOS、Linux、Windows、Android 等平台建置,涵蓋 CUDA 12/13、Vulkan、ROCm 10.0、OpenVINO、SYCL 與 Snapdragon 後端。
llama.cpp releases● 精選10/3 07:27AI 評分62
llama.cpp 發布建置版本 b11370,CUDA 後端新增將共享專家(shared experts)融合進 MMVQ 的改動(#29184),並修復緩衝區空值檢查、把 stride_col_dst 移入融合引數。
r/LocalLLaMA top of the day10/3 15:16AI 評分72
基於 ExLlamaV3 的 Kyojin 引擎在 AMD Strix Halo(gfx1151、ROCm)上讓兩個 300B 級 MoE 模型各裝進一臺 128GB 迷你主機:GLM-5.3-Flash 佔 99.7GB,3.5K 上下文預填充 580 tok/s,解碼 26 到 30 tok/s;MiMo-V2.6-Flash 佔 105GB,解碼最高 44 tok/s(程式碼場景)。
llama.cpp releases● 精選10/3 09:36AI 評分72
llama.cpp 合併 PR #29831,為 clef 決策模型加入純文字推理支援,並更新 gguf-py 常量。該版本覆蓋 macOS Apple Silicon、Linux、Windows、Android 等平台,其中 Apple Silicon 的 KleidiAI 加速建置被停用。
r/LocalLLaMA top of the day10/3 21:06AI 評分75
開源推理引擎 TensorSharp 在 RTX 3080 Laptop(16GB 視訊記憶體)、32GB 記憶體加 SSD 的普通筆記本上跑起了 Qwen3.8 Flash Next 176B,無需 128GB 以上記憶體工作站或多 GPU。它靠量化加 MoE 感知的統一排程,協調視訊記憶體、記憶體、SSD 與快取分層。