Issues / #588

#588 Reported expert-cache hit rate excludes --pcie-frac experts, so it rises as tok/s falls

closed · @yannickloth · 1 评论 · 在 GitHub 查看

BenchmarksServer & APINVIDIA / CUDAModels & quants

描述

The serve log's "decode expert cache hit rate" and the `/metrics` `hit_rate` field are the **VRAM-resident share** of the experts the pool looked up, not the share that avoided the CPU pool. Experts sent over PCIe for the GPU to read (`--pcie-frac`) are counted in **neither** `cache_hits` nor `cache_refused`, so raising `--pcie-frac` removes misses from the denominator and *raises* the reported hit rate even though those experts still have to be fetched and computed on the GPU:

- `src/core/expert_source.cpp`: in the pool dispatch, `if (kind[i] >= 0) { if (kind[i] == 0) ++d.cache_hits; ... continue; }` — `kind == 1` (PCIe) and `kind == 2` (peer) hit neither counter, and only the CPU path below does `++d.cache_refused`. `d.pcie_experts` tracks the offload separately (`:1965`).
- `src/program/generate.cpp` computes the rate over `hits + admitted + refused` and even documents the exclusion: *"the VRAM share of the experts the pool looked up while decoding; experts it sent over PCIe for the GPU to read (--pcie-frac) are in neither count"*.

So the metric is correct as a VRAM-residency figure, but it is presented (and shown in the Monitor tab) as a cache hit rate / health number, and it moves in the opposite direction from end-to-end speed when `--pcie-frac` changes. Measured on an RTX A3000 12 GB (Swift 1.5 IQ3_XXS, engine 0.1.35, a representative coding conversation):

| `--pcie-frac` | reported "hit rate" (fresh / warm) | decode tok/s |
| ---: | ---: | ---: |
| 0.00 | 0.393 / 0.546 | 41.0 |
| 0.35 (calibrated) | 0.468 / 0.619 | 40.3 |
| 0.55 | 0.534 / 0.689 | 36.2 |
| 0.75 | 0.644 / 0.777 | 34.7 |

0.75 reports a 0.78 hit rate while being the *slowest* arm. This makes the field hard to use as a calibration/health signal and easy to misread (e.g. "hit rate went up, why is it slower?").

### Suggestion

Surface the GPU-offloaded share so the resident rate is interpretable — e.g. append the `pcie`/`peer` delta to the serve line and expose it in `/metrics`, or rename the existing field to make the exclusion explicit (`vram_hit_rate`), or compute the rate over all routed lookups. Happy to implement whichever direction is preferred; I did not want to change the semantics unilaterally since the current behaviour is deliberate. Details/measurements: forks-Strata `ops/issues/p8-reduce-misses.md`.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。