llama.cpp b11455 新增 PLaMo-3 分詞器預切分
llama.cpp 建置 b11455 新增 vocab 型別 plamo3,實現 PLaMo-3 分詞器的預切分。它在 Unigram DP 前插入硬邊界,針對形如 <|plamo:...|> 的文字,以及連續至少 4 個相同字元或 2 個空格的片段;否則 llama.cpp 對程式碼縮排和重複標點的分詞會與參考實現不同。
llama.cpp 新增 PLaMo-3 分詞器預切分規則,修復程式碼縮排與重複標點分詞偏差,對在 llama.cpp 中使用 PLaMo-3 的使用者最有用。
原標題:b11455
閱讀原文
| 評分 | 62 / 62(平均 62,門檻 60) |
| 狀態 | 精選 |
|---|
全文翻譯
vocab:實現 PLaMo-3 分詞器預分割(#30045)
vocab:實現 PLaMo-3 分詞器預分割
PLaMo-3 分詞器在執行 Unigram DP 之前插入硬邊界,在形如 <|plamo:...|> 的文字週圍,以及在至少 4 個相同字元或 2 個空格的連續段周圍。沒有它們,llama.cpp 對程式碼縮排和重複標點的分詞會與參考實現不同。通過獨立編碼每個片段,復現 llm_tokenizer_plamo2 中的兩遍 re.sub()。
新增 vocab 型別 "plamo3"
更新 src/llama-vocab.cpp
Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co
雜項更改
Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co
網站:
https://llama.app
證明:
https://github.com/ggml-org/llama.cpp/attestations/53332190
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI 已啟用) 已停用
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (CUDA 12) - CUDA 12.8 庫
- Ubuntu x64 (CUDA 13) - CUDA 13.4 庫
- Ubuntu arm64 (CUDA 13) - CUDA 13.4 庫
- Ubuntu x64 (ROCm 10.0)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
- Linux arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 設定指南
Android:
- Android arm64 (CPU)
- Android arm64 (Snapdragon:CPU、Adreno GPU、Hexagon NPU) - 設定指南
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLL
- Windows x64 (CUDA 13) - CUDA 13.4 DLL
- Windows arm64 (CUDA 13) - CUDA 13.4 DLL
- Windows x64 (Vulkan)
- Windows arm64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (ROCm 10.0)
openEuler:
- 已停用
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI:
- UI
由 AI 翻譯,以原文為準。
原文
vocab : implement PLaMo-3 tokenizer pre-segmentation ( #30045 )
vocab : implement PLaMo-3 tokenizer pre-segmentation
The PLaMo-3 tokenizer inserts hard boundaries before running the Unigram
DP, around <|plamo:...|>-looking text, and around runs of at least 4
identical characters or 2 spaces. Without them llama.cpp tokenizes code
indentation and repeated punctuation differently from the reference.
Reproduce the two re.sub() passes in llm_tokenizer_plamo2 by encoding
each segment independently.
add vocab type "plamo3"
Update src/llama-vocab.cpp
Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co
misc change
Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53332190
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/4 11:22AI 評分42
llama.cpp 發布 b11387 版本,修復了在截斷後溫度大於 0 時 n-gram 草稿被拒絕的問題(PR #29924),由 NVIDIA 的 Pranesh Gonegandla 參與提交。
llama.cpp releases10/6 19:16AI 評分40
llama.cpp 發布 b11447 建置版本,本次列出的功能變更是新增對 pplx-decider 模型的支援(PR #30044)。該版本一次性放出 macOS/iOS、Linux、Android、Windows、openEuler 五個平台的預編譯產物,計算後端覆蓋 CPU、CUDA 12.8 與 13.4、Vulkan、ROCm 10.0、OpenVINO、SYCL,以及驍龍平台的 Adreno GPU 與 Hexagon NPU;
llama.cpp releases10/4 20:50AI 評分23
llama.cpp 發布建置版本 b11397,主要變更是將 CUDA 後端的 neu_padded 移動到實際使用位置(PR #29940),作者為 Hugging Face 的 Adrien Gallouët。該版本繼續提供 macOS、Linux、Windows、Android 等多平台預編譯包,覆蓋 Apple Silicon、CUDA 12/13、ROCm 10.0、Vulkan、SYCL、OpenVINO 等後端;
llama.cpp releases10/4 15:12AI 評分16
llama.cpp 發布 b11391 建置版本,唯一程式碼改動是將 CUDA 後端的 blocks_per_col 變數移到實際使用位置(PR #29939),由 Hugging Face 的 Adrien Gallouët 提交。
llama.cpp releases10/3 09:36AI 評分43
llama.cpp 合併 PR #29831,為 clef 決策模型加入初始支援,目前僅限純文字輸入。該改動同時更新了 gguf-py 常量並清理了靜態圖相關程式碼,隨附的建置產物覆蓋 macOS、Linux、Windows、Android 與 openEuler 等多個平台。