Pull requests / #907

#907 calibrate: measure the adaptive expert tier (--adapt-every / --adapt-swaps / --adapt-decay)

closed · @1872183316 · 0 comentários · No GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows

Descrição

`--calibrate` measures the adaptive expert tier too: after the worker count, the engine's default tier against two
that swap more and remember longer, a restart each (as for the workers), kept only when it beats the default by more
than `MIN_GAIN` (3%).

| Candidate | `--adapt-every` | `--adapt-swaps` | `--adapt-decay` |
| --- | ---: | ---: | ---: |
| default | (engine: 4) | (96) | (0.7) |
| busier | 1 | 80 | 0.97 |
| busiest | 1 | 160 | 0.97 |

`apply()` drops the three flags when the calibration keeps the default, so an older calibration's tier never lingers
(the same rule as the other settings).

## Why

On a PC whose CPU reads the missed experts slowly, the CPU's share is most of a verify window, and the default tier
swaps about 24 experts per window. Measured in #906 (Xeon E5-2673 v3, DDR3, RTX 4060 Ti 16 GB on PCIe 3.0 x8, Q2_0,
setup's default config, 4 prompts x 2 passes, interleaved):

| Tier | with #764 | without #764 (lag 1) |
| --- | ---: | ---: |
| default | 52.61 tok/s | 48.71, 49.60, 49.55 |
| every 1 / 80 / 0.97 | 54.67 - 55.25 | 53.19 |
| every 1 / 160 / 0.97 | **56.90, 56.64** (+7.9%) | - |
| every 1 / 240 / 0.97 | 56.36 | - |

Hit rate on the code-edit prompt 0.62 -> 0.73, CPU time in those windows 40.1 -> 28.8 ms. A routing trace replayed
offline (in #906) shows the misses follow the swaps per window more than the counting rule. A PC with a fast CPU and
RAM may not gain, which is why this is a measurement and not a new default.

## Cost

Three engine starts more: with the model already in the page cache (the first measurement loaded it) each start is
seconds (16 s here), then one warm-up round and two measured rounds of the three calibration prompts with 512-token
answers (the tier needs some windows to follow a text; with the calibration's usual 128 tokens it did not show, see the
comment below). The whole calibration took 655 s on this PC; the setup message says about 10 minutes.

## Test

`tools/test_calibrate.py`: `test_adaptive_tier` (a 10% faster candidate is kept, a 2% one is not, the candidates run
with the worker count chosen before, `apply` keeps and removes the flags); `test_fewer_workers` counts the extra
starts. All `tools/test_setup_*.py` pass.

## Measured end to end

`tools/calibrate.py` on that PC (Q2_0, setup's default config): engine default 54.3, every 1 / 80 / 0.97 60.1,
every 1 / 160 / 0.97 60.6 tok/s: **kept** (+11.6%); the other settings stayed at their defaults. Any PC other than this
one is not measured.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.