AISpot

llama.cpp 修復圖重分配導致的排程中止

llama.cpp releases10/3 13:40版本更新地端推論開源開發工具

llama.cpp 合併 PR #29856,把 build_rs 中分散的 get_rows 合併為一次收集全部 n_rs 個 recurrent states,使預留空間按 n_rs 最大值計算,避免 ubatch 單元不連續時在節點數不變的情況下觸發圖重分配、進而在 GGML_SCHED_NO_REALLOC 下中止。

你在 Mac Studio 上用 llama.cpp 或 MLX 跑本地模型時,這類排程中止會直接讓長上下文或多序列推理崩掉,升級到 b11375 可避開該問題。

原標題:b11375
閱讀原文

評分82 / 72(平均 77,門檻 60)
狀態精選

全文翻譯

graph:一次性收集迴圈狀態,使預留覆蓋每個拆分(#29856) build_rs 用自己的 get_rows 收集了額外狀態(n_rs - n_seqs 行)。最壞情況下的預留有 n_rs == n_seqs,因此該節點被設定為零行,任何單元不連續的 ubatch 都會在節點數不變的情況下強制進行圖重新分配,而這在 GGML_SCHED_NO_REALLOC 下會中止。 現在單個 get_rows 收集 n_rs 個狀態:ubatch 狀態和額外狀態都是它的檢視,其大小僅取決於 n_rs,而預留已將其設定為最大值。自訂 getter(mamba ssm_scan)從第二個狀態收集,因此單序列 ubatch 不復制任何狀態。檢視在輸入中每個圖建置一次,以保持圖的主機開銷不變。 網站: https://llama.app 證明: https://github.com/ggml-org/llama.cpp/attestations/52413277 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

由 AI 翻譯,以原文為準。

原文
graph: gather the recurrent states once so the reserve covers every split ( #29856 ) build_rs gathered the extra states (n_rs - n_seqs rows) with their own get_rows. The worst-case reserve has n_rs == n_seqs, so that node was sized at zero rows, and any ubatch whose cells are not contiguous forced a graph reallocation at an unchanged node count, which aborts under GGML_SCHED_NO_REALLOC. A single get_rows now gathers the n_rs states: the ubatch states and the extra states are views of it, and its size only depends on n_rs, which the reserve already sets to the maximum. A custom getter (mamba ssm_scan) gathers from the second state, so a single sequence ubatch copies no state. The views are built once per graph in the input to keep the host overhead of the graph unchanged. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52413277 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/3 02:58AI 評分58

llama.cpp 修復 qkx3 量化縮放搜尋的非法舍入

llama.cpp 合併 PR #29817,在 qkx3 量化級別送入 nearest_int 前先鉗制到 [0, nmax],避免 imatrix 縮放搜尋產生無窮、NaN 或越界值時觸發 Debug 建置斷言。該問題由 #29804 報告,修復同時為 q2_K、q4_K、q5_K、q4_1、q5_1 的退化 imatrix 分組補充迴歸測試。合法範圍內的取值行為不變,macOS Apple Silicon(arm64)等預編譯包已隨本次發布更新。

r/LocalLLaMA top of the day● 精選10/3 07:18AI 評分79

llama.cpp 新 PR 將 Qwen 索引器視訊記憶體減半

llama.cpp 的 PR #29825 把 Qwen Flash Next 的 indexer score 視訊記憶體佔用減半,例如 131072 上下文、ub 2048 時從 3.0 GiB 降到 1.5 GiB,262144 上下文、ub 4096 時從 11.7 GiB 降到 5.6 GiB。作者稱 logits 完全一致,僅分數因 fp32 舍入有差異,pp/tg 速度不變,test-llama-archs 通過。