Issues / #1595

#1595 0.1.41 CPU share default: no gain, and ~5% slower at 1,000 tokens, on a V100 with all experts in VRAM

open · @christopherrobertbrooks-tech · 0 commentaires · Sur GitHub

NVIDIA / CUDAModels & quants

Description

Measured the new default (CPU share on for chunks below 1,024 tokens) against `STRATA_PREFILL_CPU_SHARE=0` on one Tesla
V100-PCIe-32GB with an i7-13700KF (16 GB RAM, low-RAM mode), Coder IQ1_M with **all experts resident in VRAM**
(`--resident-experts`), `v0.1.41` built for sm_70 (CUDA 12.9), same binary, arms alternating, 3 rounds x 3 runs,
temperature 0. Prompt time (ms), mean of 3 runs per round:

| prompt | share off (r1 / r2 / r3) | share on (r1 / r2 / r3) | |
| ---: | --- | --- | ---: |
| 256 | 793 / 789 / 795 | 793 / 796 / 792 | 0% |
| 512 | 1,021 / 1,018 / 1,020 | 1,027 / 1,023 / 1,020 | +0.3% |
| 1,000 | 1,468 / 1,464 / 1,469 | 1,538 / 1,538 / 1,535 | **+4.8%** |
| 4,096 | 3,292 | 3,293 | 0% (share not used) |

The engine printed the "CPU share is ON by default" line in the `on` arm, and the env check confirmed `=0` in the `off`
arm.

This fits the explanation in the notes: when every expert is already on the GPU, there's nothing for the CPU to save, so
the share can only cost. The 1,000-token case is consistently slower in all 3 rounds, despite `auto`'s timing check.
Maybe the default could skip the share when the expert cache holds everything (as with `--resident-experts`)? For now we
run with `STRATA_PREFILL_CPU_SHARE=0`.

Sur le site

Liens install, modèles, releases.