AISpot

llama.cpp 合併 K2 Horizon 支援,修復 Windows GGUF 無法載入

llama.cpp releases10/6 21:24版本更新模型地端推論開源開發工具

llama.cpp 合併了 K2 Horizon 稠密模型與 MoVA 支援,包含 GGUF 轉換、計算圖、tokenizer、chat template,以及推理與工具呼叫。

K2 Horizon 首次進入 llama.cpp 主線,並修掉 MSVC 正則不支援 \p{...} 導致該模型 GGUF 在 Windows 上完全無法載入的 bug,對在本地跑該模型的 Windows 使用者最直接。

原標題:b11454
閱讀原文

評分75 / 73(平均 74,門檻 60)
狀態精選

原文

model : add K2 Horizon dense and MoVA support ( #29535 ) model: K2 Horizon gguf conversion code model: loading hparams and tensors in k2-horizon.cpp model: K2 Horizon compute graph model: K2 Horizon compute graph adjustment and registering tokenizers model: K2 Horizon chat template and accomodate safetensors naming unicode : add the K2-Horizon pre-tokenizer splitter The K2-Horizon regex had no arm in unicode_regex_split_custom and fell through to the general std::regex fallback, which fails two ways. On MSVC std::regex rejects \p{...}, so no K2-Horizon GGUF loads on Windows at all: llama-quantize, llama-imatrix and llama-perplexity all abort with regex_error(error_escape) before a token is produced. Where the fallback does compile it is still wrong. unicode_regex_split collapses each codepoint to a single byte naming its Unicode category before matching, and U+200C/U+200D are category Control, which has no entry in k_ucat_cpt, so both become the 0xD0 fallback byte. The literal ‌ and ‍ alternatives in K2's regex can then never match and every ZWNJ or ZWJ ends a letter run. The splitter is the existing llama3 one with a single rule widened, since K2's regex differs from llama3's only in that a letter run also takes marks, ZWNJ and ZWJ. tests/test-unicode.cpp gains a case for this: it fails before the change with [Amy] [ZWNJ khaham] and passes after with the run intact. tests: expand K2 Horizon unicode splitter coverage unicode: handle K2 Horizon case folding and empty input Assisted-by: Codex jinja : support sequence indices in selectattr and rejectattr Assisted-by: Codex model : add K2 Horizon dense and MoVA support Includes the K2 Horizon implementation from ifm-ai/llama.cpp with converter, tensor-parallel and model save/reload fixes. Assisted-by: Codex chat : support K2 Horizon reasoning and tool calls Assisted-by: Codex conversion: remove obsolete K2 Aurora alias Assisted-by: Codex k2-horizon: enforce response schemas and load YaRN betas Constrain final JSON after reasoning, accept flexible JSON tool envelopes, enforce XML dialects, and handle repeated or alternate thinking markers. Load YaRN beta metadata instead of retaining the default values. Add schema, streaming, continuation, and model reload regressions. Validate CUDA and CPU builds and 0.9B, 4B, and MoVA conversation/tool round trips. Assisted-by: Codex renaming template fixture adressing cisc follows ups desloppify the parser / adress aldehir comments clean test-chat remove fallback : model trained mostly on high anyway fix k2 attn_v_exp tn splitting and metal fusion baseline k2-horizon : forward expand views before sums k2-horizon: copy embds before group norm to fix TP disable tesnor parallelism Co-authored-by: Ryandito Diandaru ryandito.diandaru@mbzuai.ac.ae Co-authored-by: WestWaters mario.papaleo2013@gmail.com Co-authored-by: Natani L. Mayday 71436458+TaskPuppyNatani@users.noreply.github.com Co-authored-by: West 100190545+WestWaters@users.noreply.github.com Co-authored-by: aaryamonvikram aaryamonvikram@gmail.com Co-authored-by: aaryamonvikram 96529820+aaryamonvikram@users.noreply.github.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/53326308 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows arm64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

相關報導

llama.cpp releases10/5 09:19AI 評分12

llama.cpp 發布 b11407,修復 Vulkan 測試編譯錯誤

llama.cpp 發布 b11407 版本,修復了在開啟 -DGGML_VULKAN_RUN_TESTS=ON 時 Vulkan 後端出現的未宣告識別符號問題。該版本同步提供 macOS、Linux、Windows、Android 等多平台預編譯包,涵蓋 CPU、CUDA 12/13、Vulkan、ROCm 10.0、OpenVINO、SYCL 及 Snapdragon 的 Adreno GPU 與 Hexagon NPU 等後端。