llama.cpp 更新 Hexagon 後端,預設動態緩衝升至 512MB
llama.cpp b11490 把 Hexagon 後端的預設動態緩衝從 128MB 提升到 512MB,原因是 128MB 在大 MoE 模型上會導致效能退化。hex-bufs 新增 alloc_buffer_n 支援,並可將大張量拆分到獨立緩衝區;環境變數 GGML_HEXAGON_MBUF 改為接受 dyn、static、total 三個值。hex-run 還新增 --no-embd-offload 選項,用於簡化需要該設定的裝置上的命令列。
原標題:b11490
閱讀原文
| 評分 | 42 / 45(平均 43,門檻 60) |
| 狀態 | 未入選 |
|---|
原文
hexagon: enable alloc_buffer_n ( #30126 )
hex-bufs: add support for alloc_buffer_n
hex-bufs: add support for splitting large tensors into separate buffers
hex-bufs: update GGML_HEXAGON_MBUF to accept three values dyn,static,total
hex-bufs: bump dyn. default to 512MB since 128MB causes perf regressions with big MOEs
hex-run: add --no-embd-offload option to simplify command lines on devices that need it
Update scripts/snapdragon/run.py
Co-authored-by: Jhen-Jie Hong iainst0409@gmail.com
Co-authored-by: Jhen-Jie Hong iainst0409@gmail.com
Website:
https://llama.app
Attestations:
https://github.com/ggml-org/llama.cpp/attestations/53811782
macOS/iOS:
macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework
Linux:
Ubuntu x64 (CPU)
Ubuntu arm64 (CPU)
Ubuntu s390x (CPU)
Ubuntu x64 (Vulkan)
Ubuntu arm64 (Vulkan)
Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
Ubuntu x64 (ROCm 10.0)
Ubuntu x64 (OpenVINO)
Ubuntu x64 (SYCL FP32)
Ubuntu x64 (SYCL FP16)
Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Android arm64 (CPU)
Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Windows:
Windows x64 (CPU)
Windows arm64 (CPU)
Windows arm64 (OpenCL Adreno)
Windows x64 (CUDA 12) - CUDA 12.4 DLLs
Windows x64 (CUDA 13) - CUDA 13.4 DLLs
Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
Windows x64 (Vulkan)
Windows arm64 (Vulkan)
Windows x64 (OpenVINO)
Windows x64 (SYCL)
Windows x64 (ROCm 10.0)
openEuler:
DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)
UI:
UI
相關報導
llama.cpp releases10/7 23:40AI 評分41
llama.cpp 發布 b11486 建置,Hexagon 後端的 GELU 啟用函式精度得到改進(#30104)。該版本繼續為驍龍平台提供 Hexagon NPU、Adreno GPU 與 CPU 三類後端,覆蓋 Linux arm64、Android arm64 與 Windows arm64。
llama.cpp releases10/2 09:11AI 評分43
llama.cpp 合併 PR #29717,為 Hexagon 後端新增 q2_k 和 q3_k 兩種量化型別支援,並統一了 src1_row_size 的分配方式。該改動由高通工程師 Max Krasnyansky 參與提交,面向在驍龍 Hexagon NPU 上執行本地推理的使用者,相關建置覆蓋 Linux arm64 與 Android arm64 的 Snapdragon CPU、Adreno GPU、Hexagon NPU 組合。
llama.cpp releases10/8 08:41AI 評分40
llama.cpp 建置版本 b11500 發布,主要變更是 hex-mmadd 在融合 bias-add 時不再假設讀寫記憶體對齊(#30133)。
llama.cpp releases● 精選10/5 11:55AI 評分68
llama.cpp 的 Hexagon 後端在 row-split 多核模式下把 flash_attn 從按 Q token 切分改為按 KV head 切分,讓每個核心只讀自己負責的那部分 KV cache,不再重複讀取全部 KV cache。該行為由 GGML_HEXAGON_FA_HEAD_SPLIT 控制,預設開啟;當 n_kv_heads 不能被核心數整除時回退到原 token-block 切分(如 4 核上只有 2 個 KV head 的 Gemma-4)。
llama.cpp releases10/6 06:05AI 評分33
llama.cpp 發布 b11438,主要變更是 test-llama-archs 在生成模型前先初始化後端(#30034):開啟 GGML_BACKEND_DL 時,必須先顯式載入後端才能建立模型。