llama.cpp 合併 GLM5-Next MTP 支援,可作 draft 模型
llama.cpp 合併 PR #29928,為 GLM5-Next 加入多 token 預測(MTP)的 NextN 計算圖,使其能作為投機解碼的 draft 模型使用。改動包括支援只含 NextN 層或只含主幹張量的拆分 GGUF 檔案載入,並把 MTP 圖按輸出行裁剪,無輸出行的 4-token catch-up 核心耗時從 6.9 ms 降到 2.9 ms,貪心輸出雜湊與 draft 接受率不變。
給出了 GLM5-Next MTP 在 llama.cpp 的具體實現與核心耗時數字(4-token catch-up 從 6.9 ms 降至 2.9 ms),並支援拆分 GGUF 當 draft 模型,對在本地跑 GLM5-Next 的人最有用。
原標題:b11474
閱讀原文
| 評分 | 78 / 78(平均 78,門檻 60) |
| 狀態 | 精選 |
|---|
原文
feat: add GLM5Next MTP, optimize ( #29928 )
llama : add GLM5-Next NextN (MTP) graph
Build the GLM5-Next multi-token-prediction head as graph_mtp: the NextN block
embeds enorm(tok)+hnorm(h) through eh_proj, runs one plain DSA layer and the
shared lm_head, reusing the trunk's builders through the no_build tag ctor.
llama_memory_recurrent also tolerates a partial seq_rm when the context holds
no recurrent layers, which is what the MTP draft context needs.
Assisted-by: Claude
llama : glm5-next: skip dead compute in headless NextN forwards
A NextN forward with no output rows (the MTP catch-up and the draft-context
prefill) persists only through its cache writes, so the headless graph keeps
the MLA latent, indexer key|gate and pooled-key writes and drops the query
path, the indexer selection, the attention body, the FFN and the LM head. The
4-token catch-up falls from 6.9 ms to 0.33 ms of kernels; the greedy output
hashes and the draft acceptance are unchanged.
Assisted-by: Claude
llama : glm5-next: fix NextN extraction contracts and shared-tail rollback
Three fixes from the architectural review. The headless graph prune now also
requires that no unmasked nextn extraction is live, because that mode reads
n_tokens hidden rows regardless of the logits flags. Masked extraction
publishes the hidden rows gathered by the output ids, so a batch whose output
flags are not a prefix exports the right rows. A partial recurrent rollback
whose tail cell is shared with another sequence is now rejected instead of
silently moving that sequence's tail.
Assisted-by: Claude
llama : glm5-next: tidy comments in the MTP changes
Assisted-by: Claude
llama : glm5-next: crop the MTP graph to the output rows instead of pruning it
Replace the headless NextN prune with the crop pattern the other MTP
graphs use: gather the attention output and the block input at the
output ids before the position-wise FFN and the shared head. A NextN
forward with no output rows (the MTP catch-up and the draft-context
prefill) then runs the FFN and the head over zero rows. The 4-token
catch-up falls from 6.9 ms to 2.9 ms of kernels; greedy output hashes
are unchanged.
Assisted-by: Claude
glm5-next: use the nextn crop helpers in the MTP graph
Replace the local crop condition and the masked select of t_h_nextn
with crop_before_nextn and crop_after_nextn, so the MTP graph narrows
its rows the same way as the main graph and the other models.
Describe the shared cell and empty filter branches of the recurrent
partial rollback.
glm5-next: load MTP-only and trunk-only GGUF files
Make the trunk tensors optional when the file only holds the NextN
layer, and the NextN tensors optional when the file only holds the
trunk, so the split MTP GGUF loads as a draft model.
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
Co-authored-by: Pascal admin@serveurperso.com
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53581440
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/6 14:10AI 評分45
llama.cpp 發布 b11440,修復投機 MTP 解碼中 NextN 提取標誌變化導致的排程器預留失效報錯。啟用 MTP 後,NextN 提取在目標與草稿上下文的排程器已預留之後才開啟;未掩碼提取時主幹圖保留最後一層的全部 token 而非裁剪到輸出行,首次解碼按該批次形狀重分配,下一批更寬的批次就會觸發 GGML_SCHED_DEBUG_REALLOC。修復方式是在標誌變化時作廢預留,讓下一次計算按新圖形狀重新預留,並隨各平台預編譯包發布。
llama.cpp releases10/7 15:30AI 評分50
llama.cpp 把此前各模型各自複製的 nextn 張量檢測邏輯,統一提取為 llama_model_base 的 nextn_flags 輔助函式:探測第一個 trunk 層和第一個 NextN 層,並在未載入 MTP 時加上 TENSOR_SKIP。
llama.cpp releases● 精選10/5 17:58AI 評分82
llama.cpp 發布 v0.6.0,新增 llama_batch_ext 擴充套件 batch API,支援混合 token/embedding 輸入與 MTP、deepstack 狀態嵌入。本次加入 320B 的 GLM-5.3-Flash(GLM5-Next)文字+視覺混合模型、Clef 決策模型,以及 Qwen4Exp 的 MTP 投機解碼,官方稱在 DGX Spark 上解碼約快 1.5 倍。
llama.cpp releases10/6 14:27AI 評分42
llama.cpp 發布 b11443 建置,其中提交 #30017 把各模型重複的 nextn 行裁剪邏輯合併為 llm_graph_context 上的 crop_before_nextn() 與 crop_after_nextn() 兩個共享 helper。
llama.cpp releases10/6 20:08AI 評分42
llama.cpp 發布 b11450,RPC 後端新增 -sm tensor 選項(#26610),支援在 RPC 場景下按 tensor 拆分模型。該版本還修復 Apple RDMA 的 flush 問題,把 graph_uids 移入 rpc_dispatcher,停止 dispatcher 執行緒空轉,並提升 RPC 主版本號。