llama.cpp 修復 Ling 3.0 解析器對 json_schema 的支援
llama.cpp 合併 PR #29813,讓 Ling 3.0 解析器支援 inputs.json_schema,此前 response_format 請求不受約束。新邏輯為回應格式語法路徑優先於工具呼叫,啟用思考時要求 JSON 前有 think 塊,且 JSON 後不允許尾隨文字。該修復對應 issue #29652,由 Claude Opus 5.5 輔助完成。
你在本地跑開源模型並可能用結構化輸出,這個修復直接決定 Ling 3.0 在 llama.cpp 上的 JSON 約束是否可靠。
原標題:b11377
閱讀原文
| 評分 | 78 / 78(平均 78,門檻 60) |
| 狀態 | 精選 |
|---|
全文翻譯
chat:在 Ling 3.0 解析器中遵循 json_schema(#29813)
chat:在 Ling 3.0 解析器中遵循 json_schema
Ling 3.0 只為工具呼叫建置了語法,沒有處理 inputs.json_schema,因此 response_format 請求未受約束。
新增一條急切的回應格式語法路徑,其優先順序高於工具,遵循現有的解析器模式。在啟用思考時,要求在 JSON 之前,並且不允許在 JSON 回應之後出現尾隨文字。
修復 #29652。
協助者:Claude Opus 5.5
chat:要求 Ling 3.0 為回應格式提供 think 塊
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/52425527
macOS/iOS:
macOS Apple Silicon(arm64)
macOS Apple Silicon(arm64,啟用 KleidiAI)已停用
macOS Intel(x64)
iOS XCFramework
Linux:
Ubuntu x64(CPU)
Ubuntu arm64(CPU)
Ubuntu s390x(CPU)
Ubuntu x64(Vulkan)
Ubuntu arm64(Vulkan)
Ubuntu x64(CUDA 12)- CUDA 12.8 庫
Ubuntu x64(CUDA 13)- CUDA 13.4 庫
Ubuntu arm64(CUDA 13)- CUDA 13.4 庫
Ubuntu x64(ROCm 10.0)
Ubuntu x64(OpenVINO)
Ubuntu x64(SYCL FP32)
Ubuntu x64(SYCL FP16)
Linux arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU)- 設定指南
Android:
Android arm64(CPU)
Android arm64(Snapdragon:CPU、Adreno GPU、Hexagon NPU)- 設定指南
Windows:
Windows x64(CPU)
Windows arm64(CPU)
Windows arm64(OpenCL Adreno)
Windows x64(CUDA 12)- CUDA 12.4 DLL
Windows x64(CUDA 13)- CUDA 13.4 DLL
Windows arm64(CUDA 13)- CUDA 13.4 DLL
Windows x64(Vulkan)
Windows x64(OpenVINO)
Windows x64(SYCL)
Windows x64(ROCm 10.0)
openEuler:
已停用
openEuler x86(310p)
openEuler x86(910b,ACL Graph)
openEuler aarch64(310p)
openEuler aarch64(910b,ACL Graph)
UI:
UI
由 AI 翻譯,以原文為準。
原文
chat : honor json_schema in Ling 3.0 parser ( #29813 )
chat : honor json_schema in Ling 3.0 parser
Ling 3.0 only built a grammar for tool calls and did not handle inputs.json_schema, so response_format requests were left unconstrained.
Add an eager response-format grammar path with precedence over tools, following the existing parser patterns. Require before JSON when thinking is enabled and do not allow trailing prose after the JSON response.
Fixes #29652 .
Assisted-by: Claude Opus 5.5
chat : require Ling 3.0 think block for response formats
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52425527
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases● 精選10/3 09:36AI 評分72
llama.cpp 合併 PR #29831,為 clef 決策模型加入純文字推理支援,並更新 gguf-py 常量。該版本覆蓋 macOS Apple Silicon、Linux、Windows、Android 等平台,其中 Apple Silicon 的 KleidiAI 加速建置被停用。
llama.cpp releases● 精選10/3 03:38AI 評分80
llama.cpp 合併 PR #27694,把 simple draft 與 MTP 的投機解碼改為機率取樣,目標模型用拒絕取樣驗證草稿 token。新增開關控制機率草稿取樣,預設仍為貪心;語法約束請求回退到 argmax,並支援在拒絕取樣中處理語法約束。同時修復了草稿取樣器與目標模型共用 rng 流、掩碼後分布未重歸一化等問題,並隨草稿一起截斷候選。
X @ggerganov● 精選10/2 15:33AI 評分82
llama.cpp 最新建置已支援 Decision 模型,新增 /v1/systemone 端點,可在本地高效、私密地做 Jev 風格推理。該端點支援多個開源模型,後續還會增加。
llama.cpp releases10/3 02:06AI 評分58
llama.cpp 合併 PR #29844,新增對 nimble 決策模型的支援。該版本覆蓋 macOS Apple Silicon(arm64)、Ubuntu、Windows、Android 等多平台建置,其中 macOS Apple Silicon 的 KleidiAI 加速版本被標記為 DISABLED。官方發布頁為 llama.app,建置證明見 GitHub attestations。
llama.cpp releases10/3 02:58AI 評分58
llama.cpp 合併 PR #29817,在 qkx3 量化級別送入 nearest_int 前先鉗制到 [0, nmax],避免 imatrix 縮放搜尋產生無窮、NaN 或越界值時觸發 Debug 建置斷言。該問題由 #29804 報告,修復同時為 q2_K、q4_K、q5_K、q4_1、q5_1 的退化 imatrix 分組補充迴歸測試。合法範圍內的取值行為不變,macOS Apple Silicon(arm64)等預編譯包已隨本次發布更新。