llama.cpp 合併 K2 Horizon 支援,修復 Windows GGUF 無法載入
llama.cpp 合併了 K2 Horizon 稠密模型與 MoVA 支援,包含 GGUF 轉換、計算圖、tokenizer、chat template,以及推理與工具呼叫。
K2 Horizon 首次進入 llama.cpp 主線,並修掉 MSVC 正則不支援 \p{...} 導致該模型 GGUF 在 Windows 上完全無法載入的 bug,對在本地跑該模型的 Windows 使用者最直接。
原標題:b11454
閱讀原文
| 評分 | 75 / 73(平均 74,門檻 60) |
| 狀態 | 精選 |
|---|
原文
model : add K2 Horizon dense and MoVA support ( #29535 )
model: K2 Horizon gguf conversion code
model: loading hparams and tensors in k2-horizon.cpp
model: K2 Horizon compute graph
model: K2 Horizon compute graph adjustment and registering tokenizers
model: K2 Horizon chat template and accomodate safetensors naming
unicode : add the K2-Horizon pre-tokenizer splitter
The K2-Horizon regex had no arm in unicode_regex_split_custom and fell through to the
general std::regex fallback, which fails two ways.
On MSVC std::regex rejects \p{...}, so no K2-Horizon GGUF loads on Windows at all:
llama-quantize, llama-imatrix and llama-perplexity all abort with
regex_error(error_escape) before a token is produced.
Where the fallback does compile it is still wrong. unicode_regex_split collapses each
codepoint to a single byte naming its Unicode category before matching, and U+200C/U+200D
are category Control, which has no entry in k_ucat_cpt, so both become the 0xD0 fallback
byte. The literal and alternatives in K2's regex can then never match and
every ZWNJ or ZWJ ends a letter run.
The splitter is the existing llama3 one with a single rule widened, since K2's regex
differs from llama3's only in that a letter run also takes marks, ZWNJ and ZWJ.
tests/test-unicode.cpp gains a case for this: it fails before the change with
[Amy] [ZWNJ khaham] and passes after with the run intact.
tests: expand K2 Horizon unicode splitter coverage
unicode: handle K2 Horizon case folding and empty input
Assisted-by: Codex
jinja : support sequence indices in selectattr and rejectattr
Assisted-by: Codex
model : add K2 Horizon dense and MoVA support
Includes the K2 Horizon implementation from ifm-ai/llama.cpp with converter, tensor-parallel and model save/reload fixes.
Assisted-by: Codex
chat : support K2 Horizon reasoning and tool calls
Assisted-by: Codex
conversion: remove obsolete K2 Aurora alias
Assisted-by: Codex
k2-horizon: enforce response schemas and load YaRN betas
Constrain final JSON after reasoning, accept flexible JSON tool envelopes,
enforce XML dialects, and handle repeated or alternate thinking markers.
Load YaRN beta metadata instead of retaining the default values.
Add schema, streaming, continuation, and model reload regressions. Validate
CUDA and CPU builds and 0.9B, 4B, and MoVA conversation/tool round trips.
Assisted-by: Codex
renaming template fixture
adressing cisc follows ups
desloppify the parser / adress aldehir comments
clean test-chat
remove fallback : model trained mostly on high anyway
fix k2 attn_v_exp tn splitting and metal fusion baseline
k2-horizon : forward expand views before sums
k2-horizon: copy embds before group norm to fix TP
disable tesnor parallelism
Co-authored-by: Ryandito Diandaru ryandito.diandaru@mbzuai.ac.ae
Co-authored-by: WestWaters mario.papaleo2013@gmail.com
Co-authored-by: Natani L. Mayday 71436458+TaskPuppyNatani@users.noreply.github.com
Co-authored-by: West 100190545+WestWaters@users.noreply.github.com
Co-authored-by: aaryamonvikram aaryamonvikram@gmail.com
Co-authored-by: aaryamonvikram 96529820+aaryamonvikram@users.noreply.github.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53326308
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
X @ggerganov● 精選10/6 17:49AI 評分69
llama.cpp 已實現對 EmbeddingGemma 2 的 Day-0 支援,由作者 ggerganov 公布。對應 GGUF 權重由 ggml-org 發布在 Hugging Face,可直接下載用於本地推理。Day-0 意味著模型發布當天即可在 llama.cpp 上執行,無需等待社群自行轉換。
llama.cpp releases10/3 09:36AI 評分43
llama.cpp 合併 PR #29831,為 clef 決策模型加入初始支援,目前僅限純文字輸入。該改動同時更新了 gguf-py 常量並清理了靜態圖相關程式碼,隨附的建置產物覆蓋 macOS、Linux、Windows、Android 與 openEuler 等多個平台。
llama.cpp releases10/5 09:19AI 評分12
llama.cpp 發布 b11407 版本,修復了在開啟 -DGGML_VULKAN_RUN_TESTS=ON 時 Vulkan 後端出現的未宣告識別符號問題。該版本同步提供 macOS、Linux、Windows、Android 等多平台預編譯包,涵蓋 CPU、CUDA 12/13、Vulkan、ROCm 10.0、OpenVINO、SYCL 及 Snapdragon 的 Adreno GPU 與 Hexagon NPU 等後端。
llama.cpp releases10/4 09:06AI 評分20
llama.cpp 發布 b11384 版本,官網為 llama.app,並附 GitHub 建置證明。該版本提供 macOS、iOS、Linux、Android、Windows 及 openEuler 的預編譯包,涵蓋 Apple Silicon、CUDA 12/13、ROCm 10.0、Vulkan、SYCL、OpenVINO 與 Snapdragon 的 CPU、Adreno GPU、Hexagon NPU 等後端。
r/LocalLLaMA top of the day10/2 19:47AI 評分65
llama.cpp 的 PR #29184 提出在 CUDA 後端把共享專家融合進 MMVQ,為部分 MoE 架構帶來提速,例如 Qwen 35B A3B。有評論稱該改動使 TG 速度提升約 5%,並指出僅對部分 MoE 架構有效。