Issues / #1595
#1595 0.1.41 CPU share default: no gain, and ~5% slower at 1,000 tokens, on a V100 with all experts in VRAM
open · @christopherrobertbrooks-tech · 0 コメント · GitHub で見る
本文
Measured the new default (CPU share on for chunks below 1,024 tokens) against `STRATA_PREFILL_CPU_SHARE=0` on one Tesla V100-PCIe-32GB with an i7-13700KF (16 GB RAM, low-RAM mode), Coder IQ1_M with **all experts resident in VRAM** (`--resident-experts`), `v0.1.41` built for sm_70 (CUDA 12.9), same binary, arms alternating, 3 rounds x 3 runs, temperature 0. Prompt time (ms), mean of 3 runs per round: | prompt | share off (r1 / r2 / r3) | share on (r1 / r2 / r3) | | | ---: | --- | --- | ---: | | 256 | 793 / 789 / 795 | 793 / 796 / 792 | 0% | | 512 | 1,021 / 1,018 / 1,020 | 1,027 / 1,023 / 1,020 | +0.3% | | 1,000 | 1,468 / 1,464 / 1,469 | 1,538 / 1,538 / 1,535 | **+4.8%** | | 4,096 | 3,292 | 3,293 | 0% (share not used) | The engine printed the "CPU share is ON by default" line in the `on` arm, and the env check confirmed `=0` in the `off` arm. This fits the explanation in the notes: when every expert is already on the GPU, there's nothing for the CPU to save, so the share can only cost. The 1,000-token case is consistently slower in all 3 rounds, despite `auto`'s timing check. Maybe the default could skip the share when the expert cache holds everything (as with `--resident-experts`)? For now we run with `STRATA_PREFILL_CPU_SHARE=0`.
関連リンク
インストール・モデル・リリースへの站内リンク。