AISpot

llama.cpp 修復 server 崩潰:限制 n_batch 不超過 n_ubatch

llama.cpp releases10/3 16:54版本更新地端推論開源開發工具

llama.cpp 合併 PR #29903,通過把 n_batch 限制在 n_ubatch 以內,修復了 server 的 laya abort 崩潰(#29902)。該修復由 Claude 輔助完成,並移除了相關測試與 embeddings 條件。

你準備用 Mac Studio 跑本地模型,這個修復直接關係到 llama.cpp server 在 Apple Silicon 上的穩定性,升級後能避免 n_batch 配置不當導致的崩潰。

原標題:b11379
閱讀原文

評分62 / 45(平均 53,門檻 60)
狀態未入選

原文

server : fix laya abort by limiting n_batch to n_ubatch ( #29903 ) server : fix laya abort by limiting n_batch to n_ubatch Fixes #29902 Assisted-by: Claude fix(review) : rm tests, embeddings cond Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52435853 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/3 02:58AI 評分58

llama.cpp 修復 qkx3 量化縮放搜尋的非法舍入

llama.cpp 合併 PR #29817,在 qkx3 量化級別送入 nearest_int 前先鉗制到 [0, nmax],避免 imatrix 縮放搜尋產生無窮、NaN 或越界值時觸發 Debug 建置斷言。該問題由 #29804 報告,修復同時為 q2_K、q4_K、q5_K、q4_1、q5_1 的退化 imatrix 分組補充迴歸測試。合法範圍內的取值行為不變,macOS Apple Silicon(arm64)等預編譯包已隨本次發布更新。

r/LocalLLaMA top of the day● 精選10/3 07:18AI 評分79

llama.cpp 新 PR 將 Qwen 索引器視訊記憶體減半

llama.cpp 的 PR #29825 把 Qwen Flash Next 的 indexer score 視訊記憶體佔用減半,例如 131072 上下文、ub 2048 時從 3.0 GiB 降到 1.5 GiB,262144 上下文、ub 4096 時從 11.7 GiB 降到 5.6 GiB。作者稱 logits 完全一致,僅分數因 fp32 舍入有差異,pp/tg 速度不變,test-llama-archs 通過。

llama.cpp releases● 精選10/3 00:57AI 評分68

llama.cpp 新增 /v1/systemone API,支援 laya 等 5 個模型

llama.cpp 發布 b11361,在 server 中新增 /v1/systemone API,支援 laya、julia-1、lev、openjev、kev 五個模型,並附帶模型轉換指令碼。該版本加入了共享 prompt 字首與視覺輸入支援,新增用於測試的小模型 openjev tiny,文件中明確說明不支援 date_facts。macOS Apple Silicon(arm64)建置正常,但啟用 KleidiAI 的 arm64 版本被標記為 DISABLED;