Pull requests / #907
#907 calibrate: measure the adaptive expert tier (--adapt-every / --adapt-swaps / --adapt-decay)
closed · @1872183316 · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows
描述
`--calibrate` measures the adaptive expert tier too: after the worker count, the engine's default tier against two that swap more and remember longer, a restart each (as for the workers), kept only when it beats the default by more than `MIN_GAIN` (3%). | Candidate | `--adapt-every` | `--adapt-swaps` | `--adapt-decay` | | --- | ---: | ---: | ---: | | default | (engine: 4) | (96) | (0.7) | | busier | 1 | 80 | 0.97 | | busiest | 1 | 160 | 0.97 | `apply()` drops the three flags when the calibration keeps the default, so an older calibration's tier never lingers (the same rule as the other settings). ## Why On a PC whose CPU reads the missed experts slowly, the CPU's share is most of a verify window, and the default tier swaps about 24 experts per window. Measured in #906 (Xeon E5-2673 v3, DDR3, RTX 4060 Ti 16 GB on PCIe 3.0 x8, Q2_0, setup's default config, 4 prompts x 2 passes, interleaved): | Tier | with #764 | without #764 (lag 1) | | --- | ---: | ---: | | default | 52.61 tok/s | 48.71, 49.60, 49.55 | | every 1 / 80 / 0.97 | 54.67 - 55.25 | 53.19 | | every 1 / 160 / 0.97 | **56.90, 56.64** (+7.9%) | - | | every 1 / 240 / 0.97 | 56.36 | - | Hit rate on the code-edit prompt 0.62 -> 0.73, CPU time in those windows 40.1 -> 28.8 ms. A routing trace replayed offline (in #906) shows the misses follow the swaps per window more than the counting rule. A PC with a fast CPU and RAM may not gain, which is why this is a measurement and not a new default. ## Cost Three engine starts more: with the model already in the page cache (the first measurement loaded it) each start is seconds (16 s here), then one warm-up round and two measured rounds of the three calibration prompts with 512-token answers (the tier needs some windows to follow a text; with the calibration's usual 128 tokens it did not show, see the comment below). The whole calibration took 655 s on this PC; the setup message says about 10 minutes. ## Test `tools/test_calibrate.py`: `test_adaptive_tier` (a 10% faster candidate is kept, a 2% one is not, the candidates run with the worker count chosen before, `apply` keeps and removes the flags); `test_fewer_workers` counts the extra starts. All `tools/test_setup_*.py` pass. ## Measured end to end `tools/calibrate.py` on that PC (Q2_0, setup's default config): engine default 54.3, every 1 / 80 / 0.97 60.1, every 1 / 160 / 0.97 60.6 tok/s: **kept** (+11.6%); the other settings stayed at their defaults. Any PC other than this one is not measured. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。