AISpot

b11536: ui: Models Manager Follow-up Improvements (#30228)

llama.cpp releases版本更新地端推論開源開發工具

llama.cpp 新版本 b11536 讓 server 的 GET /models 直接返回每個模型的訓練上下文長度(context_length),資料取自 GGUF 後設資料,無需載入模型或請求 Hugging Face Hub。新增的 common_get_gguf_n_ctx_train 只讀後設資料、相容 u32 與 u64,檔案缺失或無效時返回 0,router 列表因此可離線攜帶該欄位。

原標題:b11536: ui: Models Manager Follow-up Improvements (#30228)
閱讀原文

評分45 / 45(平均 45,門檻 60)
狀態未入選

原文

common : read a GGUF's trained context from its metadata common_get_gguf_n_ctx_train opens only the file's metadata (no_alloc, like common_get_decision_type) and reads .context_length, so a caller can learn the trained context without loading the model. It accepts both u32 and u64 values and returns 0 when the file is missing, unreadable, invalid, or reports no context length. Assisted-by: pi:zai-org/GLM-5.3-Flash server : report the trained context in the models listing update_caps already resolves the model file offline to read its modalities, so it now reads the trained context from the same GGUF metadata, and GET /models reports it as context_length when it is known. A router listing then carries the context without any Hub request, which lets the UI sort and filter by it offline. Assisted-by: pi:zai-org/GLM-5.3-Flash ui : take the trained context from the models listing The router now reports context_length per model, so the option mapping fills contextLength from it and the manager reads it before the Hub record. The Context column, the context sort and the context filter then work with the Hugging Face Hub API turned off. A browser suite guards the sort and the search, the Hub-cache driven context filter and the re-sort when details arrive after the sort was clicked. Assisted-by: pi:zai-org/GLM-5.3-Flash ui : mark favorite models with a heart A favorited model shows a rose heart in the selector even before its row is hovered, and the crossed heart takes its place on hover, so unfavoriting stays one hover away. The manager table marks its favorited rows with the same heart after the badges and capabilities. Assisted-by: pi:zai-org/GLM-5.3-Flash fix: UI text nit fix: UI nits fix: Favorite models grouping in models table feat: Remove sorting from Status column in Models Table server: read the GGUF metadata once per model Read the decision type and the trained context in a single GGUF open, accept only a UINT32 context length like the model loader, and reset n_ctx_train with the other caps so a failed refresh drops it. fix: Post-review fixes Co-authored-by: Pascal [email protected]

相關報導

llama.cpp releasesAI 評分45

llama.cpp b11436 修復 OpenVINO 後端多項 GPU 迴歸

llama.cpp b11436 修復 ggml-openvino 後端的 CI 測試失敗與多項 GPU 迴歸。上游 #29622 給輸入 embedding 圖加入混合 token/embd 分支,後端改為只從計算節點建置 OV 模型,並把同型別連續 DUP 按 CONT 翻譯以避免圖被拆分,修復了 CPU 與 GPU 上首個單 token decode 失敗;

X @ollama● 精選AI 評分65

@ollama: .@GoogleDeepMind EmbeddingGemma 2 is now available on Ollama.

Google DeepMind 的 EmbeddingGemma 2 已上線 Ollama,執行 ollama pull embeddinggemma-2 即可拉取。訊息由 Ollama 官方帳號在 10 月 6 日發布,官方稱該模型面向消費級裝置,並且是多模態模型。原文未披露模型體積、上下文長度與具體支援的模態等細節。

X @ggerganov● 精選AI 評分77

@ggerganov: The new v0.6.0 release packs a lot of good stuff:

llama.cpp 發布 v0.6.0,作者 ggerganov 列出了這一版的主要更新:新增 Clef 支援,涵蓋文字與視覺;對 Qwen3.8-Flash-Next 提供高品質支援;Metal 後端效能大幅提升;並引入新的 llama_batch_ext API。配套官網也完成改版,地址為 llama.app。原文未給出各項改動的具體效能數字。