llama.cpp 修復 NextN 提取引發的排程器重分配報錯
llama.cpp 發布 b11440,修復投機 MTP 解碼中 NextN 提取標誌變化導致的排程器預留失效報錯。啟用 MTP 後,NextN 提取在目標與草稿上下文的排程器已預留之後才開啟;未掩碼提取時主幹圖保留最後一層的全部 token 而非裁剪到輸出行,首次解碼按該批次形狀重分配,下一批更寬的批次就會觸發 GGML_SCHED_DEBUG_REALLOC。修復方式是在標誌變化時作廢預留,讓下一次計算按新圖形狀重新預留,並隨各平台預編譯包發布。
原標題:b11440
閱讀原文
| 評分 | 45 / 45(平均 45,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
llama : re-reserve the sched when the nextn extraction flags change ( #30020 )
The speculative MTP init enables NextN extraction on the target and draft
contexts after both were created and their schedulers reserved. With
unmasked extraction the trunk graph keeps every token through the last
layer instead of cropping to the output rows, so the first decode
reallocates to that batch's shape and the next, wider batch trips
GGML_SCHED_DEBUG_REALLOC. Invalidate the reserve when the flags change so
the next compute re-reserves with the new graph shape.
Assisted-by: Claude
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53181663
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/5 09:36AI 評分48
llama.cpp 發布 b11412,修復 k-pool 模型(qwen4exp、glm5-next)解碼時意外重新預留計算圖並中止的問題。原因是兩模型按 cache_safe、n_tokens 等 reserve 無法預知的狀態分支,解碼圖與預留圖節點數不一致(如 7564 對 7762),解碼時被迫按當前狀態重預留、丟掉最壞情況尺寸,在 GGML_SCHED_DEBUG_REALLOC=1 下直接 abort。
llama.cpp releases10/4 13:56AI 評分37
llama.cpp 發布建置版本 b11390,主要修復了 CUDA 後端在 n_expert 遠大於 n_ubatch 時的 MMQ 記憶體故障(#29941)。
llama.cpp releases10/6 14:27AI 評分42
llama.cpp 發布 b11443 建置,其中提交 #30017 把各模型重複的 nextn 行裁剪邏輯合併為 llm_graph_context 上的 crop_before_nextn() 與 crop_after_nextn() 兩個共享 helper。
llama.cpp releases10/4 11:22AI 評分42
llama.cpp 發布 b11387 版本,修復了在截斷後溫度大於 0 時 n-gram 草稿被拒絕的問題(PR #29924),由 NVIDIA 的 Pranesh Gonegandla 參與提交。
llama.cpp releases10/4 17:25AI 評分40
llama.cpp 合併 PR #29942,修復 chat-peg-parser 中 pending_tool_call 重置時未清空 current_tool 指標導致的 use-after-free 與 id 緩衝區二次釋放。該問題在 TOOL_CLOSE 之後收到 TOOL_ID 節點時觸發,修復方式是在重置時清除指標。