AISpot

llama.cpp 為投機解碼加入機率取樣與拒絕取樣

llama.cpp releases10/3 03:38版本更新地端推論開源開發工具

llama.cpp 合併 PR #27694,把 simple draft 與 MTP 的投機解碼改為機率取樣,目標模型用拒絕取樣驗證草稿 token。新增開關控制機率草稿取樣,預設仍為貪心;語法約束請求回退到 argmax,並支援在拒絕取樣中處理語法約束。同時修復了草稿取樣器與目標模型共用 rng 流、掩碼後分布未重歸一化等問題,並隨草稿一起截斷候選。

你正為 Mac Studio 評估本地推理,這條改動直接影響 llama.cpp 在 Apple Silicon 上投機解碼的取樣正確性與速度,值得在選型時納入測試。

原標題:b11368
閱讀原文

評分82 / 78(平均 80,門檻 60)
狀態精選

全文翻譯

spec : 為簡單草稿和 MTP 新增機率取樣 ( #27694 ) 使草稿器具有機率性,目標通過拒絕取樣進行驗證 在起草前丟棄過期的 spec_draft_q 對語法約束請求回退到 argmax 取樣,並新增用於啟用機率草稿取樣的標誌。預設標誌值為貪心。 在拒絕取樣中支援語法約束請求 修復 - 在掩碼後重新歸一化分佈 在取樣器複製時複製 rng,並在重放時重新接受已起草的 token 修復草稿取樣器共享目標 rng 流的問題 簡化拒絕取樣器的輸入,並將重放移至伺服器 隨草稿一起截斷草稿候選 Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com Co-authored-by: Pranesh Gonegandla pgonegandla@nvidia.com 網站: https://llama.app 證明: https://github.com/ggml-org/llama.cpp/attestations/52346366 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI

由 AI 翻譯,以原文為準。

原文
spec : add probabilistic sampling for simple draft and MTP ( #27694 ) Make the drafter probabilistic and the target verify by rejection sampling Drop stale spec_draft_q before drafting Fallback to argmax sampling for grammar-constrained requests and adding flag for enabling probabilistic draft sampling. Default flag value is greedy. Support grammar-constrained requests in rejection sampling Fix - renormalize distribution after masking copy rng on sampler copy and re-accept drafted tokens on replay Fix draft sampler sharing the target's rng stream Simplify the rejection sampler's inputs and move replay to the server Truncate the draft candidates along with the draft Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com Co-authored-by: Pranesh Gonegandla pgonegandla@nvidia.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52346366 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.4 DLLs Windows arm64 (CUDA 13) - CUDA 13.4 DLLs Windows x64 (Vulkan) Windows x64 (OpenVINO) Windows x64 (SYCL) Windows x64 (ROCm 10.0) openEuler: DISABLED openEuler x86 (310p) openEuler x86 (910b, ACL Graph) openEuler aarch64 (310p) openEuler aarch64 (910b, ACL Graph) UI: UI