AISpot

b11477: llama: share the nextn tensor flags between models (#30097)

llama.cpp releases10/7 15:30版本更新開源地端推論

llama.cpp 把此前各模型各自複製的 nextn 張量檢測邏輯,統一提取為 llama_model_base 的 nextn_flags 輔助函式:探測第一個 trunk 層和第一個 NextN 層,並在未載入 MTP 時加上 TENSOR_SKIP。

原標題:b11477: llama: share the nextn tensor flags between models (#30097)
閱讀原文

評分50 / 50(平均 50,門檻 60)
狀態未入選

原文

llama: share the nextn tensor flags between models Follow-up of the TODO in glm5-next: move the trunk-only and MTP-only detection that each model copied into a nextn_flags helper of llama_model_base. It probes the first trunk layer and the first NextN layer, and adds TENSOR_SKIP when MTP is not loaded. qwen4exp probes hc_attn_norm since it has no attn_norm. deepseek4, nemotron-h, qwen35, qwen35moe, qwen3next and qwen4exp now also accept a trunk-only file, like the other models. llama: avoid capturing structured bindings in the nextn flags Lambdas that capture structured bindings need C++20, and GCC 15 rejects them under -Werror, so the models read the trunk and MTP flags into plain variables.

相關報導

llama.cpp releases● 精選10/7 14:45AI 評分78

llama.cpp 合併 GLM5-Next MTP 支援,可作 draft 模型

llama.cpp 合併 PR #29928,為 GLM5-Next 加入多 token 預測(MTP)的 NextN 計算圖,使其能作為投機解碼的 draft 模型使用。改動包括支援只含 NextN 層或只含主幹張量的拆分 GGUF 檔案載入,並把 MTP 圖按輸出行裁剪,無輸出行的 4-token catch-up 核心耗時從 6.9 ms 降到 2.9 ms,貪心輸出雜湊與 draft 接受率不變。

llama.cpp releases10/6 14:10AI 評分45

llama.cpp 修復 NextN 提取引發的排程器重分配報錯

llama.cpp 發布 b11440,修復投機 MTP 解碼中 NextN 提取標誌變化導致的排程器預留失效報錯。啟用 MTP 後,NextN 提取在目標與草稿上下文的排程器已預留之後才開啟;未掩碼提取時主幹圖保留最後一層的全部 token 而非裁剪到輸出行,首次解碼按該批次形狀重分配,下一批更寬的批次就會觸發 GGML_SCHED_DEBUG_REALLOC。修復方式是在標誌變化時作廢預留,讓下一次計算按新圖形狀重新預留,並隨各平台預編譯包發布。

llama.cpp releases10/4 20:50AI 評分23

llama.cpp 發布 b11397,CUDA 後端調整 neu_padded 位置

llama.cpp 發布建置版本 b11397,主要變更是將 CUDA 後端的 neu_padded 移動到實際使用位置(PR #29940),作者為 Hugging Face 的 Adrien Gallouët。該版本繼續提供 macOS、Linux、Windows、Android 等多平台預編譯包,覆蓋 Apple Silicon、CUDA 12/13、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;