llama.cpp server 拒絕部分媒體截斷
llama.cpp 發布 b11415,其 server 端現在會拒絕部分媒體截斷(partial media truncation)的請求。此次提交只保留 keep_first 修復,去掉了 mtmd 測試輔助改動——clip_image_f32_batch 改為按值儲存條目後該改動已無法編譯;同時刪除視覺測試,因為沒有測試夾具能在複用快取的情況下切到兩個相鄰媒體塊之間,該測試與修復本身無關。
原標題:b11415
閱讀原文
| 評分 | 45 / 45(平均 45,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
server: reject partial media truncation ( #24076 )
server: reject partial media truncation
server: keep only the keep_first fix
Drop the mtmd test helper change, which no longer builds since
clip_image_f32_batch stores its entries by value, and drop the
vision test: no test fixture reaches a cut between two adjacent
media chunks with a reused cache (tinygemma3 uses SWA and wraps
images in text tokens, tinyopenjev and small-test are recurrent),
so the test passed or failed independently of the fix.
Co-authored-by: Pascal admin@serveurperso.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52813031
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/4 11:22AI 評分42
llama.cpp 發布 b11387 版本,修復了在截斷後溫度大於 0 時 n-gram 草稿被拒絕的問題(PR #29924),由 NVIDIA 的 Pranesh Gonegandla 參與提交。
llama.cpp releases10/3 16:54AI 評分20
llama.cpp 合併 PR #29903,通過將 n_batch 限制為不超過 n_ubatch 來修復 server 崩潰問題,對應 issue #29902。該修復由 Claude 輔助完成,並移除了相關測試與 embeddings 條件。新版本已覆蓋 macOS、Linux、Windows、Android 等平台,包括 Apple Silicon、CUDA、ROCm、Vulkan 等後端。
llama.cpp releases● 精選10/5 15:05AI 評分66
llama.cpp 發布建置 b11418,其 server 新增對 Clef 的視覺輸入支援(PR #29969),並修復影像 token 上限、abort 與 yield_to_queue 資料變更等問題。
llama.cpp releases10/3 22:35AI 評分18
llama.cpp 合併提交 #29863,修復 mtmd 在 Windows 上使用已棄用 strdup 函式產生的編譯警告,由 Hugging Face 的 Adrien Gallouët 提交。該提交隨新版本發布,覆蓋 macOS、Linux、Windows、Android 及 openEuler 等多個平台建置,其中 macOS Apple Silicon 的 KleidiAI 啟用版與 openEuler 部分建置被標記為 DISABLED。
llama.cpp releases10/4 13:56AI 評分37
llama.cpp 發布建置版本 b11390,主要修復了 CUDA 後端在 n_expert 遠大於 n_ubatch 時的 MMQ 記憶體故障(#29941)。