llama.cpp 修復圖重分配導致的排程中止
llama.cpp 合併 PR #29856,把 build_rs 中分散的 get_rows 合併為一次收集全部 n_rs 個 recurrent states,使預留空間按 n_rs 最大值計算,避免 ubatch 單元不連續時在節點數不變的情況下觸發圖重分配、進而在 GGML_SCHED_NO_REALLOC 下中止。
你在 Mac Studio 上用 llama.cpp 或 MLX 跑本地模型時,這類排程中止會直接讓長上下文或多序列推理崩掉,升級到 b11375 可避開該問題。
原標題:b11375
閱讀原文
| 評分 | 82 / 72(平均 77,門檻 60) |
| 狀態 | 精選 |
|---|
全文翻譯
graph:一次性收集迴圈狀態,使預留覆蓋每個拆分(#29856)
build_rs 用自己的 get_rows 收集了額外狀態(n_rs - n_seqs 行)。最壞情況下的預留有 n_rs == n_seqs,因此該節點被設定為零行,任何單元不連續的 ubatch 都會在節點數不變的情況下強制進行圖重新分配,而這在 GGML_SCHED_NO_REALLOC 下會中止。
現在單個 get_rows 收集 n_rs 個狀態:ubatch 狀態和額外狀態都是它的檢視,其大小僅取決於 n_rs,而預留已將其設定為最大值。自訂 getter(mamba ssm_scan)從第二個狀態收集,因此單序列 ubatch 不復制任何狀態。檢視在輸入中每個圖建置一次,以保持圖的主機開銷不變。
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/52413277
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
由 AI 翻譯,以原文為準。
原文
graph: gather the recurrent states once so the reserve covers every split ( #29856 )
build_rs gathered the extra states (n_rs - n_seqs rows) with their own
get_rows. The worst-case reserve has n_rs == n_seqs, so that node was
sized at zero rows, and any ubatch whose cells are not contiguous forced
a graph reallocation at an unchanged node count, which aborts under
GGML_SCHED_NO_REALLOC.
A single get_rows now gathers the n_rs states: the ubatch states and the
extra states are views of it, and its size only depends on n_rs, which
the reserve already sets to the maximum. A custom getter (mamba ssm_scan)
gathers from the second state, so a single sequence ubatch copies no
state. The views are built once per graph in the input to keep the host
overhead of the graph unchanged.
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52413277
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/3 16:54AI 評分53
llama.cpp 合併 PR #29903,通過把 n_batch 限制在 n_ubatch 以內,修復了 server 的 laya abort 崩潰(#29902)。該修復由 Claude 輔助完成,並移除了相關測試與 embeddings 條件。
llama.cpp releases10/3 02:32AI 評分58
llama.cpp 合併 PR #27096,修復 ggml-cpu 中 GGML_OP_SOFT_MAX_BACK 在 dst 與 src1(y)別名時輸出靜默錯誤的問題。原因是原實現先覆蓋 y 再讀取,改為單次融合迴圈先讀兩個源再寫,並新增迴歸測試強制觸發該別名。CUDA 核心因先完成歸約再寫入,不受影響。
llama.cpp releases10/3 14:27AI 評分35
llama.cpp 合併 PR #29904,通過改用融合 ADD 容差修復了 f16 精度下 ADD_ADD 測試偶發失敗的問題。該改動只涉及 CI 測試容差,不改變推理邏輯,但會隨下一次建置進入各平台發布包,包括 macOS Apple Silicon arm64、Ubuntu CUDA 12/13、Windows CUDA 12/13 等。
llama.cpp releases10/3 02:58AI 評分58
llama.cpp 合併 PR #29817,在 qkx3 量化級別送入 nearest_int 前先鉗制到 [0, nmax],避免 imatrix 縮放搜尋產生無窮、NaN 或越界值時觸發 Debug 建置斷言。該問題由 #29804 報告,修復同時為 q2_K、q4_K、q5_K、q4_1、q5_1 的退化 imatrix 分組補充迴歸測試。合法範圍內的取值行為不變,macOS Apple Silicon(arm64)等預編譯包已隨本次發布更新。
r/LocalLLaMA top of the day● 精選10/3 07:18AI 評分79
llama.cpp 的 PR #29825 把 Qwen Flash Next 的 indexer score 視訊記憶體佔用減半,例如 131072 上下文、ub 2048 時從 3.0 GiB 降到 1.5 GiB,262144 上下文、ub 4096 時從 11.7 GiB 降到 5.6 GiB。作者稱 logits 完全一致,僅分數因 fp32 舍入有差異,pp/tg 速度不變,test-llama-archs 通過。