llama.cpp 為投機解碼加入機率取樣與拒絕取樣
llama.cpp 合併 PR #27694,把 simple draft 與 MTP 的投機解碼改為機率取樣,目標模型用拒絕取樣驗證草稿 token。新增開關控制機率草稿取樣,預設仍為貪心;語法約束請求回退到 argmax,並支援在拒絕取樣中處理語法約束。同時修復了草稿取樣器與目標模型共用 rng 流、掩碼後分布未重歸一化等問題,並隨草稿一起截斷候選。
你正為 Mac Studio 評估本地推理,這條改動直接影響 llama.cpp 在 Apple Silicon 上投機解碼的取樣正確性與速度,值得在選型時納入測試。
原標題:b11368
閱讀原文
| 評分 | 82 / 78(平均 80,門檻 60) |
|---|---|
| 狀態 | 精選 |
全文翻譯
spec : 為簡單草稿和 MTP 新增機率取樣 ( #27694 )
使草稿器具有機率性,目標通過拒絕取樣進行驗證
在起草前丟棄過期的 spec_draft_q
對語法約束請求回退到 argmax 取樣,並新增用於啟用機率草稿取樣的標誌。預設標誌值為貪心。
在拒絕取樣中支援語法約束請求
修復 - 在掩碼後重新歸一化分佈
在取樣器複製時複製 rng,並在重放時重新接受已起草的 token
修復草稿取樣器共享目標 rng 流的問題
簡化拒絕取樣器的輸入,並將重放移至伺服器
隨草稿一起截斷草稿候選
Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com
Co-authored-by: Pranesh Gonegandla pgonegandla@nvidia.com
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/52346366
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
由 AI 翻譯,以原文為準。
原文
spec : add probabilistic sampling for simple draft and MTP ( #27694 )
Make the drafter probabilistic and the target verify by rejection sampling
Drop stale spec_draft_q before drafting
Fallback to argmax sampling for grammar-constrained requests and adding flag for enabling probabilistic draft sampling. Default flag value is greedy.
Support grammar-constrained requests in rejection sampling
Fix - renormalize distribution after masking
copy rng on sampler copy and re-accept drafted tokens on replay
Fix draft sampler sharing the target's rng stream
Simplify the rejection sampler's inputs and move replay to the server
Truncate the draft candidates along with the draft
Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com
Co-authored-by: Pranesh Gonegandla pgonegandla@nvidia.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/52346366
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI