llama.cpp 發布 v0.6.0,新增 llama_batch_ext 擴充套件 batch API,支援混合 token/embedding 輸入與 MTP、deepstack 狀態嵌入。本次加入 320B 的 GLM-5.3-Flash(GLM5-Next)文字+視覺混合模型、Clef 決策模型,以及 Qwen4Exp 的 MTP 投機解碼,官方稱在 DGX Spark 上解碼約快 1.5 倍。
官方發布說明給出擴充套件 batch API、320B 新模型、Metal 新核心與 ggml v0.26.0 等具體改動,並附 DGX Spark 約 1.5 倍、Apple GPU 約 3 倍的實測加速數字,適合本地或服務端跑模型的 llama.cpp 使用者評估升級。
概述
llama.cpp v0.6.0 引入了新的 llama_batch_ext 擴充套件批處理 API(帶有 llama_process),用於混合 token/嵌入輸入和 MTP/deepstack 狀態嵌入;增加了對 GLM-5.3-Flash (GLM5-Next) 320B 混合模型、Clef 決策模型(文字和視覺)以及面向 Qwen4Exp 的 MTP 投機解碼的支援;發布了新的 /v1/systemone 伺服器 API,用於決策模型(laya、julia-1、lev、openjev、kev、nimble);通過 Hugging Face Hub 資料層和模型下載流水線全面改造了 Web UI;為 F16 KV 新增了 Metal tensor-API flash attention 核心,在 Vulkan 上為量化 K/V 新增了稀疏 flash attention,並將 ggml 更新到 v0.26.0。
亮點
- 新的 llama_batch_ext 擴充套件批處理 API 與 llama_process(),支援混合 token/嵌入批處理,以及用於 MTP 和 deepstack 模型的逐 token“state”嵌入 #24669
- 新模型:GLM-5.3-Flash (GLM5-Next),一個 320B 文字+視覺混合模型 #27773,以及 Clef 決策模型,已完全支援文字和視覺 #29831 #29969
- Qwen4Exp:現已提供高品質支援,包括 MTP 投機解碼(在 DGX Spark 上約 1.5 倍解碼加速)以及各種正確性修復 #29761 #29751
- llama 和 server:新的 /v1/systemone API 支援五個決策模型——laya、julia-1、lev、openjev(+視覺)、kev #29818
- Metal:用於 F16 KV 的新 tensor API flash attention 核心 #29570
- Metal:用於投機解碼和批處理解碼的新 few-row MMA mat-mul 核心,在 Apple GPU 上 mat-mul 最高約快 3 倍 #29869
- 新的 llama_prefetch_rows(),在 Qwen4Exp 和 Gemma4 中使用基於 MADVISE 的 PLE 張量預取 #29599
API 變更
include/llama.h:新的 llama_batch_ext 批處理 API,包含 llama_embd、llama_process() 和 llama_process_type #24669;新增 llama_get_causal_attn() #28876;會話格式升級到 LLAMA_SESSION_VERSION 11 和 LLAMA_STATE_SEQ_VERSION 4
include/llama-cpp.h:為新的擴充套件批處理 API 新增了 llama_batch_ext_ptr 和 deleter #24669
tools/mtmd/mtmd.h:mtmd_get_memory_usage() 現在返回帶有 image_max_tokens 和 use_non_causal 的 mtmd_memory_usage 結構體 #29773
tools/server:用於決策模型的新 /v1/systemone 端點 #29818,且 /v1/embeddings 現在接受帶型別的視覺/音訊/影片內容 #29556
新模型
GLM-5.3-Flash (GLM5-Next):320B KDA/DSA 混合文字+視覺模型,具有 mHC 和 MoE #27773
Clef 決策模型,已完全支援文字和視覺 #29831 #29969
Ling 3.0 VL,已併入 BailingMoeV3 架構 #29151
Nimble 決策模型 #29844
為 LFM2.5-Encoder-230M/350M 註冊了 Lfm2BidirectionalForMaskedLM #29862
為 reranker 新增了 classifier_pooling 支援 #29627
核心變更
將 examples、投機解碼、mtmd 和 server 遷移到新的 llama_batch_ext API #29385 #29601;批處理現在同時接受 embd 和原始 token #29622
新增了 llama_prec_policy 以及模型驅動的 W4A4 (NVFP4/MXFP4) mul_mat 路徑 #24364
Qwen4Exp:將 indexer 分數記憶體減半 #29825,最佳化了 mask 構造 #29824,重新啟用了 -sm 張量 #28569;GLM5-Next:為失效 indexer 槽位使用唯一的 scatter 行 #29745
KV cache:修復了恢復不匹配 KV cache 旋轉的問題 #28498,修復了恢復失敗後的 K/V 和迴圈狀態清理 #27530,以及迴圈記憶體中的無效斷言 #29799
投機解碼:為 simple draft 和 MTP 提供機率取樣 #27694,修復了截斷後在 temp > 0 時 n-gram 草稿被拒絕的問題 #29924,保留層輸入的原始批處理順序 #29019,在 EOG 處停止接受草稿 token #29638
k-pool 模型:修復了意外的計算圖重新分配 #29958,並將 kpool re-pool 邊界限制到現有池 #29805
修復了 K/V 頭大小不均勻時 fused qkv 的張量拆分 #29294,一次性收集迴圈狀態,使預留覆蓋每個拆分 #29856,並正確處理訓練中的 KV #28520
Context:切換 causal_attn 時不要重新預留排程器 #28751
DFlash 草稿:在轉換期間寫入 Gemma 嵌入縮放 #29802,併為 MiMo 新增 dflash 支援 #29650
多模態變更
將非因果模型的 max_image 上限設為 n_ubatch #29773
修復了 LFM2 audio 中的 mel 前處理器 #29403
將輸入處理遷移至 llama_batch_ext #29385
伺服器變更
新增 /v1/systemone 決策模型 API,帶有專用決策流水線 #29818,並擴充套件至 nimble 決策模型 #29844
為 Clef 支援視覺輸入 #29969
GET /v1/models 和 GET /models 現在會在新的 architecture 物件中報告模型輸入/輸出模態 #29987
/v1/embeddings:接受型別化內容(視覺/音訊/影片)輸入 #29556,並對無效的嵌入請求返回 HTTP 400 #29060
允許因果 LLM 重排序器(Qwen3、Qwen3-VL)進行 RANK 池化批處理拆分 #28876
拒絕部分媒體截斷 #24076
日誌記錄:自包含顏色,並在 router 模式下將子命令日誌拆分 #29895,允許 preset 設定日誌檔案 #29334,修復了 preset 允許列表中被停用的 LLAMA_ARG_HF_REPO_FILE 鍵 #29938
通過將 n_batch 限制為 n_ubatch 修復了 laya 中止問題 #29903
當未提供 UI 時,移除內建 UI 的 service worker #29565
UI 變更
新的模型下載流水線 #27959,Hugging Face Hub 資料層 #27947,模型記憶體適配估算 #27957,以及用於 sidecar、量化與能力解析的模型 id 語法 #27946
型別安全的 API 型別、fetch 輔助工具和可下載模型儲存管道 #29582
共享模型顯示基礎元件 #29644
修復了預覽和下載中缺失的 svg use 和 animation 元素 #28962
在聊天訊息統計中一致使用 toLocaleString() 格式化 #27990
ggml 變更
ggml 更新至 v0.26.0:新的 alloc_buffer_n / get_alloc_size_n 緩衝區分配 API,SYCL、Vulkan 和 Metal 上的稀疏 flash attention 核心,以及重大的 lightning indexer 改進(分數記憶體減半、分塊、MUSA 支援)。CPU 後端新增 BF16 運算和分塊 k-quant mul_mat,CUDA 新增模型驅動的 W4A4(NVFP4/MXFP4)mul_mat 路徑以及 MMVQ 共享專家融合,Hexagon 後端新增取樣器和更多量化型別。通過更嚴格的 GGUF 大小驗證,模型載入更快,Windows ARM64 MSVC 建置已啟用,WebGPU/OpenVINO/OpenCL/SYCL 獲得了大量新核心、運算和修復。
資源
每夜建置:b11429
更多資訊
ggml-org 專案的發布與版本管理
自 v0.5.0 以來的變更日誌
d812350 llama.cpp:將版本提升至 0.6.0(#29997)
4d60b4d common、server:在 GET /models 中報告模型輸入/輸出模態(#29987)
c06f841 sync:ggml
f05c8b2 ggml:將版本提升至 0.26.0(ggml/1652)
e117148 CUDA:使 alloc_deps 檢查與批次無關(#29986)
6c59c40 vulkan:修復 Flash Attention 共享記憶體寫出越界(#29988)
3c9e747 vulkan:回退 mul_mat_id tile selection PR #29182(#29936)
b809b88 cuda:在 MUSA 上使用向量 lightning indexer 核心(#29990)
994e8f2 ci:為 make-release 工作流新增“Require Docker”標誌(#29989)
9d853bb webui:在聊天訊息統計中一致使用 toLocaleString() 格式(#27990)
8f9ae20 ci:在虛擬 Metal 裝置上停用失敗的測試(#29993)
9871df5 server:支援 Clef 的視覺輸入(#29969)
8b2fbaf CUDA:最佳化 NVFP4 型別在 mmq 中的累加(#29857)
2ed93db ci:在 docker 建置中停用未使用的 qemu(#29984)
8e16421 server:拒絕部分媒體截斷(#24076)
806eee9 vulkan:修復 flash attention 與 soft_max 之間過期的 prealloc_y 重用(#29591)
b3daa07 vulkan:用於量化 K/V 的稀疏 flash attention(#29639)
c173a53 llama:修復 k-pool 模型中意外的圖重新分配(#29958)
2107910 kv-cache:通過儲存精確旋轉後設資料修復恢復不匹配的 KV cache 旋轉(#28498)
e5983d6 ci:winget URL 必須是單獨的字串(#29978)
4ca6b76 ci:修復 docker 工作流權限(#29979)
9f12cd4 ggml-cpu:為 SpacemiT X60 新增 Q8_0 IME1 矩陣核心(#28479)
ebe18be vulkan:修復 -DGGML_VULKAN_RUN_TESTS=ON 時未宣告的識別符號(#29912)
8216c84 webgpu:為 Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 新增 MMVQ 支援(#29483)
1b43d31 cuda: 將 lightning indexer 在 keys 和 tokens 上按 4 個頭進行分塊 ( #29901 )
a3a1c47 metal : few-row MMA mat-mul ( #29869 )
9d3aba6 CUDA: 在小批次大小下對 thin f16/bf16 mul_mat 使用 MMVF ( #29633 )
d89651a CUDA: 為高效的兩階段 kernel 優先採用 whole-tile FlashAttention 排程 ( #29435 )
a7fb71f log, server: 自包含顏色,在 router 模式下將子命令從日誌中分離 ( #29895 )
0bb496d llama: 在 batch 中同時支援 embd + raw tokens ( #29622 )
2ca15f5 CUDA: 重構 swizzling 程式碼 ( #29612 )
a7b94df ggml-cpu: 在 x86 上的 tinyBLAS 中支援 BF16/FP16/FP32 K tails ( #29806 )
0eb6d9a cuda : 將 neu_padded 移動到其被使用的位置 ( #29940 )
2e7c58c ci : windows llvm 建置需要 ninja multi-config ( #29959 )
7f2dd88 ci : 新增 windows arm64 vulkan release ( #29954 )
bf79dbb AGENTS.md : 改版 ( #29656 )
dbe4c3e chat-peg-parser : 當 pending_tool_call 被重置時清除 current_tool ( #29942 )
46847e6 ci : 設定預設權限 ( #29945 )
2bc5635 cuda : 將 blocks_per_col 移動到其被使用的位置 ( #29939 )
dd26678 CUDA: 修復當 n_expert >> n_ubatch 時的 MMQ 記憶體錯誤 ( #29941 )
16c163d vulkan: 修復 rdna4 mat_vec 調優 ( #29934 )
0504396 imatrix: 為新格式 (GGUF) imatrices 計算基於啟用的統計資訊 ( #14891 )
8330e96 spec : 修復截斷後在 temp > 0 時 n-gram drafts 被拒絕的問題 ( #29924 )
6716df6 common : 為路徑轉換準備 load_from_models_dir() ( #29674 )
bf9a0cc server : 修復預設允許列表中失效的 LLAMA_ARG_HF_REPO_FILE key ( #29938 )
0faee50 ci : 推送 tag 需要 deploy key ( #29937 )
f98b31c ci : 改進發布流程 ( #29913 )
11fe021 webgpu: 為 fill/set_rows 新增 f16 支援 ( #29897 )
836d571 mtmd : 修復 Windows 上已棄用的 strdup 警告 ( #29863 )
eec18f5 vendor : 將 cpp-httplib 更新到 0.59.0 ( #29886 )
1537a0a server : 通過將 n_batch 限制為 n_ubatch 修復 laya 中止 ( #29903 )
edd6e2b common : 新增 common_is_tty() 輔助函式並修復 Windows 上已棄用的警告 ( #29860 )
9bf55f4 chat : 在 Ling 3.0 parser 中遵循 json_schema ( #29813 )
a55e952 ci: 通過使用 fused ADD 容差修復不穩定的 ADD_ADD f16 ( #29904 )
436f6f8 graph: 一次性收集迴圈狀態,使預留覆蓋每個 split ( #29856 )
b92761a ggml-openvino: 更新到 2026.4.1,最佳化效能,擴充套件運算元,改進裝置列表。 ( #29852 )
cb7934c model : 新增 LFM2.5-Encoder-350M 和 LFM2.5-Encoder-230M ( #29862 )
889edf4 qwen4exp : 將 indexer score 記憶體減半 ( #29825 )
99b9548 model: 新增對 clef decision model 的支援(純文字) ( #29831 )
bed0a85 CUDA: 將 shared experts 融合進 MMVQ ( #29184 )
4ebdf2c ci : 為 cuda 任務使用 t4-medium ( #29842 )
1fb7ef3 spec : 為 simple draft 和 MTP 新增機率取樣 ( #27694 )
134b2bb ggml-cuda : 修復 cpy 轉置路徑破壞非連續 dst 的問題 ( #27663 )
2923cf2 ggml-quants : 避免 qkx3 scale 搜尋中的無效舍入 ( #29817 )
dd4c286 ggml-cpu : 修復當 dst 與 src1 別名時 soft_max_back 輸出錯誤 ( #27096 )
46ca246 model: 支援 nimble decision model ( #29844 )
d8fbd25 readme : 新增 cmd 安裝命令 ( #29850 )
926862e metal : 為 F16 KV 新增 tensor API flash attention kernel ( #29570 )
a4cb4c6 llama, server: 新增 /v1/systemone API (models: laya, julia-1, lev, openjev, kev) ( #29818 )
70849ee common : 通過使用 u8path() 移除 fs_open_ifstream() ( #29841 )
8d81559 llama : 消除 unused-result 警告 ( #29839 )
6805ae3 llama : 使用 GGML_ABORT 代替 throw ( #29840 )
a8c9a4e opencl: 對 bf16 使用 sigmoid f16 ( #29787 )
392ded6 SYCL: Q8_0 DMMV ESIMD 和 MMVQ wide load ( #29186 )
9e258a6 vulkan: 在具有 32KB 共享記憶體的 Samsung GPU 上停用 large matmul tile ( #28531 )
b933289 sycl: 為 D=512 FA vec kernels 使用大型暫存器檔案 ( #29062 )
c328acc sycl : 不要使用緩慢的 oneDNN reference matmul 和 fattn ( #28985 )
4e2713c qwen4exp : 最佳化 mask 構造 ( #29824 )
631109b ggml : 為 buffer type 介面新增 alloc_buffer_n ( #23671 )
- 254b177 ci:修復缺失的 zdnn 後端檢查 ( #29837 )
- fb4b273 vulkan:為管線編譯問題新增日誌記錄 ( #29794 )
- 207bdab pyproject:向 uv torch 源新增 linux 平台標記 ( #29177 )
- 5fc4f3c hexagon:安裝重建後的 HTP skels ( #29828 )
- 159c651 qwen4exp:修復測試 ( #29819 )
- a868c3e hexagon:新增 q2_k 和 q3_k 量化型別支援 ( #29717 )
- ec7630a CUDA:修復 2 個損壞的 Volta FA 用例 ( #29803 )
- 78e2964 llama:引用 segment 文件 [no ci] ( #29074 )
- f1cee99 common,rpc:修復在有缺陷的 libstdc++ 上通過符號連結建立快取目錄 ( #29816 )
- 68e79bd skill:關於模型特定 CLI 引數 + 測試的說明 ( #29808 )
- e358d59 ci:通過更新 qwen4exp 基線修復 Fusion / metal ( #29812 )
- 81e39ad llama:將 kpool 重新池化邊界限制到現有池 ( #29805 )
- dcd387a hexagon:為 CPY 和 CONCAT 共享跨步 DMA 複製,通過 DMA 實現任意維度 CONCAT ( #29685 )
- d775ebf server:對無效嵌入請求返回 HTTP 400 ( #29060 )
- 2b36825 convert:為 DFlash 草稿寫入 Gemma 嵌入縮放 ( #29802 )
- 42d9581 cuda:將 sm70 路由到 Turing MMVQ nwarps 表 ( #29753 )
- 13b4d71 metal:釋放臨時私有傳輸緩衝區 ( #29777 )
- 4b1622a webgpu:為 MUL_MAT/MUL_MAT_ID/GET_ROWS 新增 bfloat16 支援- #29358 ( #29358 )
- 869034b llama:修復迴圈記憶體中的無效斷言 ( #29799 )
- b56f34a CUDA:在 cublass 路徑上處理 NVFP4 的計算型別 ( #29173 )
- c061df1 Qwen4Exp:新增 MTP ( #29761 )
- 66e0c17 llama:修復 qwen4exp ( #29751 )
- 7677678 CUDA:使 CCCL 可配置 + 為 CI 任務將其固定到 3.4.3 ( #29792 )
- 552f18f mtmd:對於 non_causal 模型將 max_image 限制為 ubatch ( #29773 )
- 5503b04 meta:用 FILL 而非 SCALE 清除不活躍的 AllReduce 分片 ( #29793 )
- def4d40 jinja:跳過複製迴圈作用域,除非迴圈過濾器需要它 ( #29776 )
- 32dd62e llama-mmap:使用 direct-io 時避免每個張量的第二次全尺寸複製 ( #29749 )
- f11d642 HIP:在 fattn_mma dqk 576 中避免將 CDNA 視為 gqa_ratio 20 的 dgx spark ( #29572 )
- 3aa0ce9 hex-workqueue:修復 seqn 與 idx_read/write 不同步的競態條件 ( #29785 )
- b0aca3c BLAS:記錄 AOCL-BLAS 建置並將裝置標記為 AOCL-BLAS ( #29640 )
- b8f96c3 common:新增 LLM-jp-4.1 Harmony 方言處理器 ( #29681 )
- 3ec4df4 opencl:為 Adreno E17 編譯器將 vec 子組廣播標記為支援 ( #29698 )
- db33d3c vocab:為 PLaMo-2 和 PLaMo-3 遵循 BOS/EOS 設定 ( #29734 )
- 7dad6db llama-bench:修復詳細度過濾器以顯示 GGML_LOG_ERROR ( #28229 )
- 2232bc8 metal:對 mxfp4 mul-mat 使用 bf16 數學 ( #29770 )
- 79625e0 llama-bench:修復文件 ( #29464 )
- 66bcc27 docs:重新整理 CPU 運算元支援矩陣 ( #29666 )
- 10f340d model:為 qwen4exp 重新啟用 -sm tensor ( #28569 )
- 0c1e570 webgpu:修復 SSM_SCAN 繫結別名 ( #29750 )
- f7b384c ggml-opencl:用 std::vector 替換 alloca() ( #29765 )
- f872b59 cuda:保護 iq4_nl 反量化行核心免受短行影響 ( #29683 )
- a4d880f Hexagon:通過支援安全散射模式最佳化 ALLREDUCE ( #29757 )
- feb9a3d args:修復 cli 下載 mmproj 引數 ( #28977 )
- 4453b53 llama:為推測解碼層輸入保留原始批次順序 ( #29019 )
- 4f31296 test-llama-archs:切換 causal_attn 以捕獲圖形狀變化 ( #29724 )
- b016f46 convert:修復 Qwen3.5 V-head 重排序導致的 LoRA 轉換崩潰 ( #28324 )
- 81ff93e llama:正確處理訓練時的 KV ( #28520 )
- 60e9cf7 batch:將其餘範例遷移到 llama_batch_ext ( #29601 )
- 05af0d2 glm5-next:為失效索引器槽提供唯一的散射行 ( #29745 )
- 2149c00 ggml/gguf:修復整數溢位 ( #29384 )
- 876c75b codeowners:移除前 ZenDNN 負責人 ( #29747 )
- b046420 cli:在 stdin EOF 時退出並移除控制台範圍的 Ctrl+C 廣播 ( #29722 )
- 22bdcc4 mimo:支援 dflash(convert + 特徵提取) ( #29650 )
- ca2e203 jinja:支援強制轉換的陣列屬性 ( #29574 )
- bdeb855 ggml-et:移除無用的 alloca() ( #29663 )
- 3b3d022 ci : 通過縮短 hrm_text fixture 修復 Models Backend Check(#29744)
- 185103d llama: llama_prefetch_rows(#29599)
- 2090f60 ggml : 新增 BF16 一元、GLU、二元和縮放運算(CPU、CUDA)(#29675)
- 90c908d cpu: 在 mul_mat 的 src1 中接受 BF16(#28937)
- 8df332d model-conversion : 為執行 org model 指令碼新增 --add-bos(#29558)
- 4a096b8 ui : 共享模型顯示原語
原文
Overview
llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process ) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.
Highlights
New llama_batch_ext extended batch API with llama_process() , supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models #24669
New models: GLM-5.3-Flash (GLM5-Next), a 320B text+vision hybrid model #27773 , and the Clef decision model, fully supported with both text and vision #29831 #29969
Qwen4Exp: high-quality support is now available, with MTP speculative decoding (~1.5x decode speedup on DGX Spark) and various correctness fixes #29761 #29751
llama and server: new /v1/systemone API supporting five decision models - laya, julia-1, lev, openjev (+vision), kev #29818
Metal: new tensor API flash attention kernel for F16 KV #29570
Metal: new few-row MMA mat-mul kernels for speculative and batched decoding, up to ~3x faster mat-mul on Apple GPUs #29869
New llama_prefetch_rows() using MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 #29599
API changes
include/llama.h : new llama_batch_ext batch API with llama_embd , llama_process() and llama_process_type #24669 , new llama_get_causal_attn() #28876 , session formats bumped to LLAMA_SESSION_VERSION 11 and LLAMA_STATE_SEQ_VERSION 4
include/llama-cpp.h : added llama_batch_ext_ptr and deleter for the new extended batch API #24669
tools/mtmd/mtmd.h : mtmd_get_memory_usage() now returns an mtmd_memory_usage struct with image_max_tokens and use_non_causal #29773
tools/server : new /v1/systemone endpoint for decision models #29818 and /v1/embeddings now accepts typed vision/audio/video content #29556
New models
GLM-5.3-Flash (GLM5-Next): 320B KDA/DSA hybrid text+vision model with mHC and MoE #27773
Clef decision model, fully supported with both text and vision #29831 #29969
Ling 3.0 VL, folded into the BailingMoeV3 architecture #29151
Nimble decision model #29844
Registered Lfm2BidirectionalForMaskedLM for LFM2.5-Encoder-230M/350M #29862
Added classifier_pooling support for rerankers #29627
Core changes
Migrated examples, speculative decoding, mtmd and server to the new llama_batch_ext API #29385 #29601 ; batches now accept both embd and raw tokens #29622
Added llama_prec_policy and a model-driven W4A4 (NVFP4/MXFP4) mul_mat path #24364
Qwen4Exp: halved indexer score memory #29825 , optimized mask constructions #29824 , re-enabled the -sm tensor #28569 ; GLM5-Next: unique scatter rows for dead indexer slots #29745
KV cache: fixed restoring mismatched KV cache rotation #28498 , fixed K/V and recurrent state cleanup after failed restores #27530 and an invalid assert in recurrent memory #29799
Speculative decoding: probabilistic sampling for simple draft and MTP #27694 , fixed n-gram drafts rejected at temp > 0 after truncation #29924 , preserved original batch order for layer inputs #29019 , stop accepting draft tokens at EOG #29638
k-pool models: fixed unexpected graph reallocation #29958 and clamped kpool re-pool bound to existing pools #29805
Fixed tensor split for fused qkv with uneven K/V head sizes #29294 , gather recurrent states once so the reserve covers every split #29856 , and properly handle KV on training #28520
Context: do not re-reserve the scheduler when toggling causal_attn #28751
DFlash drafts: write Gemma embedding scale during conversion #29802 and add dflash support for MiMo #29650
Multi-modality changes
Cap max_image to n_ubatch for non-causal models #29773
Fixed the mel preprocessor in LFM2 audio #29403
Migrated input processing to llama_batch_ext #29385
Server changes
New /v1/systemone decision-model API with dedicated decision pipeline #29818 , extended to the nimble decision model #29844
Support vision input for Clef #29969
GET /v1/models and GET /models now report model input/output modalities in a new architecture object #29987
/v1/embeddings : accept typed content (vision/audio/video) input #29556 and return HTTP 400 for invalid embedding requests #29060
Allow RANK pooling batch splitting for causal LLM rerankers (Qwen3, Qwen3-VL) #28876
Reject partial media truncation #24076
Logging: self-contained colors and split child commands from logs in router mode #29895 , allow preset to set log file #29334 , fixed dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list #29938
Fixed laya abort by limiting n_batch to n_ubatch #29903
Remove the built-in UI's service worker when the UI is not served #29565
UI changes
New model download pipeline #27959 , Hugging Face Hub data layer #27947 , model memory-fit estimation #27957 and model id grammar for sidecars, quants and capability parsing #27946
Type-safe API types, fetch helpers and download-ready models store plumbing #29582
Shared model display primitives #29644
Fixed missing svg use and animation elements in preview and download #28962
Use toLocaleString() formatting consistently across chat message statistics #27990
ggml changes
ggml updated to v0.26.0 : a new alloc_buffer_n / get_alloc_size_n buffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes.
Assets
Nightly build: b11429
More info
Releases and versioning of ggml-org projects
Changelog since v0.5.0
d812350 llama.cpp : bump version to 0.6.0 ( #29997 )
4d60b4d common, server : report model input/output modalities in GET /models ( #29987 )
c06f841 sync : ggml
f05c8b2 ggml : bump version to 0.26.0 (ggml/1652)
e117148 CUDA: make the alloc_deps check batch independent ( #29986 )
6c59c40 vulkan: fix Flash Attention shmem write out of bounds ( #29988 )
3c9e747 vulkan: revert mul_mat_id tile selection PR #29182 ( #29936 )
b809b88 cuda: use the vector lightning indexer kernel on MUSA ( #29990 )
994e8f2 ci : add "Require Docker" flag to make-release workflow ( #29989 )
9d853bb webui: Use toLocaleString() format consistently across chat message statistics ( #27990 )
8f9ae20 ci : disable failing test on virtual Metal device ( #29993 )
9871df5 server: support vision input for Clef ( #29969 )
8b2fbaf CUDA: Optimize accumulation in mmq for NVFP4 type ( #29857 )
2ed93db ci : disable unused qemu in docker build ( #29984 )
8e16421 server: reject partial media truncation ( #24076 )
806eee9 vulkan: fix stale prealloc_y reuse across flash attention and soft_max ( #29591 )
b3daa07 vulkan: sparse flash attention for quantized K/V ( #29639 )
c173a53 llama : fix unexpected graph reallocation in the k-pool models ( #29958 )
2107910 kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata ( #28498 )
e5983d6 ci : winget urls must be separate strings ( #29978 )
4ca6b76 ci : fix docker workflow permissions ( #29979 )
9f12cd4 ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 ( #28479 )
ebe18be vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON ( #29912 )
8216c84 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 ( #29483 )
1b43d31 cuda: tile the lightning indexer over keys and tokens for 4 heads ( #29901 )
a3a1c47 metal : few-row MMA mat-mul ( #29869 )
9d3aba6 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size ( #29633 )
d89651a CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels ( #29435 )
a7fb71f log, server: self contained colors, split child commands from logs in router mode ( #29895 )
0bb496d llama: support both embd + raw tokens in batch ( #29622 )
2ca15f5 CUDA: refactor swizzling code ( #29612 )
a7b94df ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 ( #29806 )
0eb6d9a cuda : move neu_padded to where it is used ( #29940 )
2e7c58c ci : windows llvm build requires ninja multi-config ( #29959 )
7f2dd88 ci : add windows arm64 vulkan release ( #29954 )
bf79dbb AGENTS.md : revamp ( #29656 )
dbe4c3e chat-peg-parser : clear current_tool when pending_tool_call is reset ( #29942 )
46847e6 ci : set default permissions ( #29945 )
2bc5635 cuda : move blocks_per_col to where it is used ( #29939 )
dd26678 CUDA: fix MMQ memory fault if n_expert >> n_ubatch ( #29941 )
16c163d vulkan: fix rdna4 mat_vec tuning ( #29934 )
0504396 imatrix: calculate activation-based statistics for new format (GGUF) imatrices ( #14891 )
8330e96 spec : fix n-gram drafts rejected at temp > 0 after truncation ( #29924 )
6716df6 common : prepare load_from_models_dir() for path conversion ( #29674 )
bf9a0cc server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list ( #29938 )
0faee50 ci : pushing tag needs deploy key ( #29937 )
f98b31c ci : improve release flow ( #29913 )
11fe021 webgpu: add f16 support to fill/set_rows ( #29897 )
836d571 mtmd : fix deprecated strdup warning on Windows ( #29863 )
eec18f5 vendor : update cpp-httplib to 0.59.0 ( #29886 )
1537a0a server : fix laya abort by limiting n_batch to n_ubatch ( #29903 )
edd6e2b common : add common_is_tty() helper and fix deprecated warnings on Windows ( #29860 )
9bf55f4 chat : honor json_schema in Ling 3.0 parser ( #29813 )
a55e952 ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance ( #29904 )
436f6f8 graph: gather the recurrent states once so the reserve covers every split ( #29856 )
b92761a ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. ( #29852 )
cb7934c model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M ( #29862 )
889edf4 qwen4exp : halve the indexer score memory ( #29825 )
99b9548 model: add support for clef decision model (text-only) ( #29831 )
bed0a85 CUDA: fuse shared experts into MMVQ ( #29184 )
4ebdf2c ci : use t4-medium for cuda jobs ( #29842 )
1fb7ef3 spec : add probabilistic sampling for simple draft and MTP ( #27694 )
134b2bb ggml-cuda : fix cpy transposed path corrupting non-contiguous dst ( #27663 )
2923cf2 ggml-quants : avoid invalid rounding in qkx3 scale search ( #29817 )
dd4c286 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 ( #27096 )
46ca246 model: support nimble decision model ( #29844 )
d8fbd25 readme : add cmd install commands ( #29850 )
926862e metal : add tensor API flash attention kernel for F16 KV ( #29570 )
a4cb4c6 llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) ( #29818 )
70849ee common : remove fs_open_ifstream() by using u8path() ( #29841 )
8d81559 llama : silence unused-result warnings ( #29839 )
6805ae3 llama : use GGML_ABORT instead of throw ( #29840 )
a8c9a4e opencl: use sigmoid f16 for bf16 ( #29787 )
392ded6 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load ( #29186 )
9e258a6 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory ( #28531 )
b933289 sycl: large register file for D=512 FA vec kernels ( #29062 )
c328acc sycl : do not use slow oneDNN reference matmul and fattn ( #28985 )
4e2713c qwen4exp : optimize mask constructions ( #29824 )
631109b ggml : add alloc_buffer_n to buffer type interface ( #23671 )
254b177 ci : fix missing zdnn backend check ( #29837 )
fb4b273 vulkan: add logging to pipeline compile issues ( #29794 )
207bdab pyproject : add linux platform marker to uv torch source ( #29177 )
5fc4f3c hexagon: install rebuilt HTP skels ( #29828 )
159c651 qwen4exp: fix tests ( #29819 )
a868c3e hexagon: add q2_k and q3_k quant type support ( #29717 )
ec7630a CUDA: fix 2 broken Volta FA cases ( #29803 )
78e2964 llama: refer to segment documentation [no ci] ( #29074 )
f1cee99 common,rpc : fix cache dir creation through symlinks on buggy libstdc++ ( #29816 )
68e79bd skill: note about model-specific CLI arguments + testings ( #29808 )
e358d59 ci: fix Fusion / metal by updating the qwen4exp baseline ( #29812 )
81e39ad llama : clamp kpool re-pool bound to existing pools ( #29805 )
dcd387a hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA ( #29685 )
d775ebf server: return HTTP 400 for invalid embedding requests ( #29060 )
2b36825 convert : write Gemma embedding scale for DFlash drafts ( #29802 )
42d9581 cuda : route sm70 to the Turing MMVQ nwarps table ( #29753 )
13b4d71 metal : release temporary private transfer buffers ( #29777 )
4b1622a webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 ( #29358 )
869034b llama : fix invalid assert in recurrent memory ( #29799 )
b56f34a CUDA: Handle compute type for NVFP4 on cublass path ( #29173 )
c061df1 Qwen4Exp: add MTP ( #29761 )
66e0c17 llama: fix qwen4exp ( #29751 )
7677678 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs ( #29792 )
552f18f mtmd: cap max_image to ubatch for non_causal models ( #29773 )
5503b04 meta: clear inactive AllReduce shards with FILL, not SCALE ( #29793 )
def4d40 jinja : skip copying loop scope unless a loop filter needs it ( #29776 )
32dd62e llama-mmap : avoid a second full-size copy of each tensor with direct-io ( #29749 )
f11d642 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 ( #29572 )
3aa0ce9 hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write ( #29785 )
b0aca3c BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS ( #29640 )
b8f96c3 common : add LLM-jp-4.1 Harmony dialect handler ( #29681 )
3ec4df4 opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler ( #29698 )
db33d3c vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 ( #29734 )
7dad6db llama-bench : fix verbosity filter to show GGML_LOG_ERROR ( #28229 )
2232bc8 metal : use bf16 math for mxfp4 mul-mat ( #29770 )
79625e0 llama-bench : fix docs ( #29464 )
66bcc27 docs : refresh CPU ops support matrix ( #29666 )
10f340d model : re-enable -sm tensor for qwen4exp ( #28569 )
0c1e570 webgpu: fix SSM_SCAN binding aliasing ( #29750 )
f7b384c ggml-opencl : replace alloca() with std::vector ( #29765 )
f872b59 cuda: guard the iq4_nl dequantize row kernel against short rows ( #29683 )
a4d880f Hexagon: optimize ALLREDUCE with support for safe scatter mode ( #29757 )
feb9a3d args: fix cli download mmproj arg ( #28977 )
4453b53 llama : preserve original batch order for speculative decoding layer inputs ( #29019 )
4f31296 test-llama-archs : toggle causal_attn to catch graph shape changes ( #29724 )
b016f46 convert : fix LoRA conversion crash for Qwen3.5 V-head reorder ( #28324 )
81ff93e llama: properly handle KV on training ( #28520 )
60e9cf7 batch: migrate the rest of examples to llama_batch_ext ( #29601 )
05af0d2 glm5-next: give dead indexer slots unique scatter rows ( #29745 )
2149c00 ggml/gguf : fix integer overflow ( #29384 )
876c75b codeowners : remove former ZenDNN owner ( #29747 )
b046420 cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast ( #29722 )
22bdcc4 mimo : support dflash (convert + feature extraction) ( #29650 )
ca2e203 jinja : support coerced array attributes ( #29574 )
bdeb855 ggml-et : remove useless alloca() ( #29663 )
3b3d022 ci : fix Models Backend Check by shortening the hrm_text fixture ( #29744 )
185103d llama: llama_prefetch_rows ( #29599 )
2090f60 ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) ( #29675 )
90c908d cpu: accept BF16 in src1 of mul_mat ( #28937 )
8df332d model-conversion : add --add-bos to run org model script ( #29558 )
4a096b8 ui : shared model display primitives ( #29644 )
8664eae ui : model download pipeline ( #27959 )
4cfb6d1 ui : model memory-fit estimation ( #27957 )
9b43336 ui : Hugging Face Hub data layer ( #27947 )
f653250 ui : model id grammar for sidecars, quants and capability parsing ( #27946 )
fa2bde5 ui : type-safe API types, fetch helpers and download-ready models store plumbing ( #29582 )
25747b0 openvino: serve GET_ROWS on a weight view from the base Constant ( #28381 )
db00347 ci : fix Fusion / metal by adding glm5-next to MTL.csv ( #29712 )
272aad8 musa : define CUDA_ARCH for device passes ( #29508 )
72db1e0 ci : add models backend check ( #29651 )
2a53ace SYCL: reduce tensor allreduce sync with pinned host buffers ( #29604 )
649dcb1 add GLM-5.3-Flash (GLM5-Next) support ( #27773 )
931351e vendor: update BoringSSL to 0.20260929.0 ( #29669 )
eae11d2 ggml-zdnn: impl buffer reset, fix memory leaks ( #29637 )
19e28a2 Hexagon f16 activation ops ( #29209 )
a6ea155 gguf : reject tensor size that wraps after padding ( #26979 )
d3954b9 ggml : check row bounds in get_rows_back ( #29575 )
48de2a1 model : support classifier_pooling for rerankers ( #29627 )
7fee178 hexagon: optimize concat op ( #29673 )
6a2743f CUDA: bitonic argsort handles rows wider than one block ( #28957 )
748d422 ggml-cuda: HIP: optimize packed byte subtraction ( __vsubss4 -> __vsub4 ) ( #29478 )
cee37ff ci: add zdnn backend build but not test ( #29541 )
6dbbac4 opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel ( #29555 )
5c200e0 vulkan: Tune GDN kernel, fix Intel performance ( #29476 )
83dd71f vulkan : Load F32 A matrix 2 at a time when its 2-aligned ( #29254 )
94a0ae3 vulkan: MOE aware mat_mul_id tile selection ( #29182 )
da89bb3 ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP ( #29504 )
a3f84fa vocab : keep NORMAL in PLaMo-2 and PLaMo-3 ( #29580 )
284153e ggml : accumulate f16 dot products in f32 on AVX512-FP16 ( #29545 )
b5cf8ce ggml : require input tensors to be GGML_OP_NONE ( #29647 )
e904318 hexagon: add FP32 GELU_ERF and GEGLU_ERF support ( #29631 )
d280808 common : stop accepting draft tokens at EOG ( #29638 )
ba0ba54 server : remove the built-in UI's service worker when the UI is not served ( #29565 )
00af635 common : use fs::path for config dir ( #29649 )
c85b92c tests : adjust server string regex to also match m2 utlra results ( #29648 )
31385c9 common : add fs_write_atomic() ( #29642 )
8019dc5 ggml : collect all input tensors into graph_inputs ( #29634 )
86ea01d ggml-zdnn: fix 0-row tensor crash ( #29636 )
18b74ff musa: build the docker image and CI container from the MUSA SDK images ( #29624 )
c13e04e ggml : speed up model loading ( #29598 )
c8cda8b ci: remove gpu-rocm keyed directory logs ( #28940 )
6d78fb0 llama : fix init in several tools/examples ( #29632 )
18bbc46 metal: FWHT perf optimizations ( #29602 )
0bc845d vulkan : reuse descriptor sets when bindings are constant ( #29280 )
139997d chat : fix Muse Glimmer ignoring response_format json_schema with --jinja ( #29615 )
76a5bc8 common : use fs::path for cache dirs ( #29595 )
46e17a6 tests : skip pytest workers when PYTEST_WORKERS=1 ( #29610 )
fc07d78 ci : update the oneAPI toolkit to 2026.1 ( #29273 )
526c43b mtmd: fix GCC 15 stringop-overflow in decode_embd_batch ( #29607 )
1c47294 hex-scripts: show trace events smaller than 100nsec in perfetto ( #29614 )
680a036 server : support typed content (vision/audio/video) input for /v1/embeddings endpoint ( #29556 )
66e665c vulkan: include functional header ( #29597 )
57b557c models: pad on the left with ggml_pad_ext ( #29567 )
14ebbd5 ggml-openvino: mark unaligned batch-stride views unsupported ( #29603 )
f1ea206 batch: migrate speculative, mtmd and server to batch_ext ( #29385 )
6c7a87f common : fix HF cache paths on Windows ( #29475 )
f00a64c webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor ( #29471 )
d77dd08 tests : refactor test-recurrent-state-rollback ( #29426 )
6f767fe ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 ( #29423 )
f916130 ci : ignore more vgpr spills in > 256 DQK fattn kernels ( #29571 )
03a667a vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain ( #29520 )
c2a9e16 HIP: fix template skip for DKQ > 256 mfma kernels ( #29559 )
4364bf7 metal: support left and circular padding in GGML_OP_PAD ( #29561 )
ed7ac35 context : do not re-reserve the scheduler when toggling causal_attn ( #28751 )
0c6a6a7 Enables Windows ARM64 build with MSVC cl.exe ( #28362 )
81ef10e tests : fix ggml init ( #29554 )
5262471 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache ( #28956 )
4da6337 server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) ( #28876 )
a97cce8 common : avoid side effects around params parsing ( #29537 )
136887b common : make string_split throw on invalid values ( #29518 )
9adc7f4 convert : export YaRN scaling parameters for PLaMo-3 ( #29528 )
6fd50a4 ci : bump ty to 0.0.84 ( #29529 )
33c923d jinja : add support for dict builtin ( #29477 )
c9064dd opencl: refine bin kernel loading condition ( #29503 )
c829670 sycl: FWHT kernels for block widths above 512 ( #29243 )
36d7b08 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 ( #26289 )
2ebd9ae HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes ( #28907 )
cea7462 vulkan: fix argsort kernel selection for Adreno ( #29469 )
da6c28e common : throw instead of abort on grammar without llguidance ( #29516 )
d7fb90e RPC: use RDMA completion channel to not spin ( #29440 )
7fb2b08 ci : enable GGML_SCHED_DEBUG_REALLOC=1 for ctest workflows ( #29514 )
187664b llama-bench : fix OOB access of hf_file ( #29515 )
85ca3b5 hrm : fix layer placement of z_l_init weight ( #29512 )
7ac59a6 hexagon: support tiled Q4_0 and Q8_0 GET_ROWS ( #29511 )
2b129cc hexagon: support for backend sampler ( #29502 )
9588757 cuda: support Nemotron 3 Puzzle state size 96 for ssm scan ( #28717 )
694ec23 musa: build the docker images from the PH1 MUSA SDK image ( #29481 )
6f856c7 cuda: add F16 input to the FWHT ( #29096 )
fcb3074 server : fix wake_fd warning on Windows ( #29479 )
2145525 Revert "Change max context length for auto-fitting with unified KV ( #28849 )" ( #29437 )
81bc6b8 jinja : implement sameas test ( #29448 )
86a24a1 jinja : fix compile error ( #29468 )
08618ff llama : fix K/V and recurrent state cleanup after failed restores ( #27530 )
a1de614 jinja : support noncall test statements with arg ( #29443 )
965f897 polished Readme and llama-bench ( #28968 )
d834d44 ggml-cpu: tiled mul_mat for k-quants ( #27851 )
9f70b2c opencl: add A8 Q8_0 non-MoE dp4a binary kernel ( #29439 )
4e74811 hexagon: find software divide calls using binary inspection tool ( #29449 )
171e884 vendor : update cpp-httplib to 0.58.0 ( #29407 )
4b1a27f common,rpc : simplify fs_create_directory_with_parents() ( #29432 )
fcc8915 mtmd: fix mel preprocessor in LFM2 audio ( #29403 )
a25c986 opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin , kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin ( #29401 )
e85e15c Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support ( #29373 ) ( #29409 )
b248f4a gguf-py : ByteLevel processing defaults bos/eos to False ( #29422 )
d81aef1 gguf-py : TemplateProcessing has final word on add_special_token ( #29417 )
27b20ba common : extract shared unicode path/string helpers ( #29415 )
e351231 metal: FWHT kernels for block widths above 512 ( #29095 )
5a75f14 metal : split fa kernels into per-dtype libraries ( #29329 )
e9f824d llama : add llama_prec_policy + model-driven W4A4 path ( #24364 )
d028c69 HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 ( #29231 )
66963a8 rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes ( #29283 )
cd74ef6 [SYCL] support sparse FA ( #28796 )
f9af9be musa: fix PH1 (MTT S5000) operator failures and build issues ( #29193 )
1ab7e5a CUDA: fuse RMS_NORM + SCALE into one kernel ( #29393 )
f805c57 llama : fix tensor split for fused qkv with uneven K/V head sizes ( #29294 )
4de0926 hexagon: add q5_k quant type support ( #29123 )
ed319fe hexagon: use DMA for contiguous dim1 CONCAT ( #29404 )
84e76d8 metal : fix graph capture and handle empty graphs ( #29390 )
cdc0642 metal : optimize sparse FA + clean-up ( #29377 )
bced459 sync : ggml ( #29396 )
a02c7f5 hexagon: handle multi-sequence in concat_2d ( #29344 )
5cf3a35 llama-grammar: fix numeric truncation for token_id parsing ( #29382 )
07fc586 hexagon: dynamic quantizer improvements ( #29395 )
97a418b hexagon: support I32 CPY and CONT ( #29379 )
a72e04a cuda : add F16 kernel support for CONV_2D_DW ( #29064 )
8212c78 test: flush status ( #28352 )
945064f ui : fix missing svg use and animation elements in preview and download ( #28962 )
fc343a8 llama: add llama_batch_ext ( #24669 )
308883b server : change default pytest workers to 4 ( #29376 )
70596c4 ci : use hf-jobs-cpu-performance, disable pytest workers ( #29369 )
70c4e15 vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 ( #27952 )
6b790a9 vulkan: handle misalignment in conv_2d and conv_3d ( #29365 )
3423f94 vulkan: tune KHR cooperative matrix support for Adreno GPUs ( #29328 )
53ed051 cuda : add conv3d with implicit GEMM ( #29137 )
f830688 model : add Ling 3.0 VL support ( #29151 )
2b70583 server,common : fix the GCC 12 stringop-overread false positive (again) ( #29325 )
4c5957c test-save-load-state : print a per-model results table in --models mode ( #29316 )
9710a32 hexagon: reject MUL_MAT_ID when src1 precision is F32 ( #29348 )
013b31c scripts : make-release-desc - link previous release in changelog title ( #29336 )
bd4f514 convert : allow vision target for DFlash/Dspark ( #29339 )
b9ae43a server: allow preset to set log file ( #29334 )
d2e5458 tests: add -b/--backend option to test-llama-archs for testing a specific backend ( #27372 )
6e60f35 ci : use hf-jobs-cpu-xl runner in server sanitize workflow ( #29297 )
fee39dd opencl: add A8 Q6_K non-MoE dp4a binary kernel ( #29057 )