Pull requests / #1332
#1332 calibrate: the PCIe share sweep reaches 1.0, where no expert goes to the CPU pool
open · @aly8246 · 0 Kommentare · Auf GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindows
Beschreibung
The PCIe share sweep stops at 0.75, so `setup --calibrate` never measures the one value that takes the CPU pool out of the decode window completely. That looks like the end of a grid rather than a decision, because 1.0 is not 0.75 carried further. The share is applied as `m = (nmiss * pcie_num) >> 8` in `src/core/expert_source.cpp:2473`, so at any value below 1.0 at least one miss still goes to the CPU pool, and only at 1.0 does every non-resident expert come over PCIe. A CPU expert sits on the layer's critical path while the GPU waits for it, which the verify profiler shows as the `waitA` / `waitB` stages. This adds 0.9 and 1.0. It cannot make anyone's calibration worse: `pick()` returns the default unless the best candidate beats it by `MIN_GAIN`, so a machine where 1.0 is the wrong answer keeps what it had, and 0.9 gives that machine an interior point next to 1.0. It costs one `s.rate()` per value, which is 3 prompts of 128 tokens, so about 8 seconds at the speeds below. What it does here. Intel Core Ultra 9 290HX Plus (8 P + 16 E, no SMT) and an RTX 5090 Laptop 24 GB on PCIe 5.0 x16, Windows 11, engine 0.1.40, the release binary, since I cannot build on this machine. The CPU is capped in frequency and power on purpose for laptop thermals, so this is the "slower CPU wants more" case from the docstring. Two rounds of three prompts, greedy, interleaved per request through `strata_tune`, so no engine restart is involved and the two arms see the same machine state: | `pcie_frac` | decode | `cpu_exp` per layer-window | `cpu_ms` per window | cache hit | | --- | --- | --- | --- | --- | | 0.55, the engine's default for this pack | 51.1 tok/s | 1.06 | 23.66 ms | 94.2% | | 1.00 | 71.3 tok/s | 0.00 | 0.03 ms | 100% | The window falls by 12.6 ms, of which 11.3 ms is less PCIe waiting and the rest is the CPU's expert time going away. `cpu_exp` reaches exactly 0.00 only at 1.0. A sequential sweep over the range, warmed, medians: 0.35 51.5, 0.55 51.1, 0.75 72.7, 0.85 69.0, 0.90 71.0, 0.95 71.9, 1.00 78.8. 0.75 through 0.95 sit inside each other's noise and 1.00 is above all of them. Sequential runs drift on this box, so the interleaved pair in the table is the part I would trust. At 257,682 prompt tokens with 1.0: 79.9, 91.3, 94.9 tok/s. Thirty requests in this machine's logs at 300K+ with the shipped 0.35 have a median of 47.5, range 32.5 to 58.3. #1245 is a second data point. It reports an RTX 5090 Laptop with a Core Ultra 9 275HX on PCIe 5.0 x16 running the same IQ3_S pack at 53-59 tok/s, on product defaults, because that run's `--calibrate` aborted part way through and the tuning never applied. 53-59 is what I measure at 0.55. Tests: `python -m unittest tools.test_calibrate` (14 tests, stand-in engine, no GPU), and `tools.test_setup_config` with `tools.test_setup_golden` (19 tests), all pass with the change. Two things I left alone. `SPEC_MIN_PS` already ends at 0.7, which is the best of the three here; twelve paired interleaved samples gave 0.70 -> 84.6 against 0.00 -> 70.4. And the sweep itself takes one measurement per value. On this box the same setting moved 8-28% between runs, from GPU boost state and CPU heat soak; the confirm step is interleaved so the winner survives, but a value can be dropped from the candidate set by one unlucky sample. An interleaved sweep would cost the same measurements and be much harder to fool. I can send that separately if you want it.
Mehr auf der Site
Links zu Install, Modellen, Releases.