所屬事件:llama.cpp 合併 PR:Qwen 索引器視訊記憶體減半(2 個來源 · 2 則報導)
2026 年 10 月 3 日,llama.cpp 合併了 PR #29825(作者 ServeurpersoCom),把 Qwen Flash Next 的 lightning indexer 打分視訊記憶體佔用減半,隨版本 b11372 發布。據 PR 說明,此前索引器一次性算出所有 head 的分數並保留一份副本… 看完整事件
llama.cpp 新 PR 將 Qwen 索引器視訊記憶體減半
llama.cpp 的 PR #29825 把 Qwen Flash Next 的 indexer score 視訊記憶體佔用減半,例如 131072 上下文、ub 2048 時從 3.0 GiB 降到 1.5 GiB,262144 上下文、ub 4096 時從 11.7 GiB 降到 5.6 GiB。作者稱 logits 完全一致,僅分數因 fp32 舍入有差異,pp/tg 速度不變,test-llama-archs 通過。
你準備用 Mac Studio 跑本地模型,這項改動能直接降低長上下文推理的視訊記憶體佔用,讓更多層留在 GPU 上。
原標題:qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp
閱讀原文
| 評分 | 80 / 78(平均 79,門檻 76) |
|---|---|
| 狀態 | 精選 |