Issues / #840

#840 --calibrate on Windows + AMD (HIP prebuilt 0.1.39) measures decode ~18x slower than the same engine through server.py, so the tuning is meaningless

open · @The-Dude-2020 · 0 comments · View on GitHub

BenchmarksSetup & installServer & APIAMD / HIPModels & quantsDocumentationWindows

Description

### Setup
- Strata 0.1.39, ready-made `strata-windows-x64-hip.zip` (ROCm 10.2.0a20260930), `--backend hip`
- AMD Radeon RX 7900 XTX 24 GB (gfx1100), Adrenalin driver 32.0.31041.1004
- Intel i5-14600K (6P + 8E, AVX2, no AVX-512), 48 GB DDR5, Windows 11 Pro 23H2 (22631)
- Model: Qwen3.8-Flash-Next IQ2_XS, 128K context, KV int8, `--pool-workers 9` (setup's hybrid-CPU rule), KV streaming off (48 GB)
- Model data on a PCIe 4.0 NVMe; PCIe probe: 22-27 GB/s host->device

Also a data point for #325: the Windows HIP zip **does** run a model on a discrete card. Through the server it is fast and correct (numbers below).

### What happens
`START-HERE.bat --calibrate` → model 1 (IQ2_XS). Every measurement comes out at ~4 tok/s:

```
    PCIe share 0.00: 4.0 tok/s
    PCIe share 0.20: 4.0 tok/s
    PCIe share 0.35: 4.1 tok/s
    PCIe share 0.55: 4.3 tok/s
    PCIe share 0.75: 4.2 tok/s
    draft floor 0.30: 4.0 tok/s
    draft floor 0.50: 4.3 tok/s
    draft floor 0.70: 4.4 tok/s
  Measuring with 13 CPU workers (restarts the engine) ...
```

The engine log (`strata-iq2_xs.log`) for the calibration requests. Note that prompt reading is also slow (7 tok/s for 22-34 tokens), not only decode:

```
strata serve: prompt 22 tokens = 0 reused + 22 read in 2925 ms (7.5 tok/s), 128 generated in 32233 ms (4.0 tok/s), drafts accepted 77 of 124, 1 checkpoints
strata serve: prompt 34 tokens = 0 reused + 34 read in 4423 ms (7.7 tok/s), 128 generated in 27744 ms (4.6 tok/s), drafts accepted 89 of 91, 1 checkpoints
strata serve: prompt 27 tokens = 0 reused + 27 read in 3750 ms (7.2 tok/s), 128 generated in 33106 ms (3.9 tok/s), drafts accepted 60 of 90, 1 checkpoints
strata serve: prompt 34 tokens = 0 reused + 34 read in 4765 ms (7.1 tok/s), 128 generated in 31229 ms (4.1 tok/s), drafts accepted 89 of 94, 1 checkpoints
strata serve: decode expert cache hit rate: 80.4% (47708 hits / 59335 lookups); 5465 more read by the GPU over PCIe or from another GPU (8.4% of all 64800 routed)
```

The same config and engine args, started through `serve/server.py --engine strata --config strata-iq2_xs.json` (as `run-iq2_xs.bat` does), a few minutes earlier, written to the same log:

```
strata serve: prompt 79 tokens = 0 reused + 79 read in 738 ms (107.0 tok/s), 512 generated in 6578 ms (77.8 tok/s), drafts accepted 292 of 446, 1 checkpoints
strata serve: prompt 79 tokens = 0 reused + 79 read in 770 ms (102.6 tok/s), 512 generated in 6906 ms (74.1 tok/s), drafts accepted 272 of 459, 1 checkpoints
strata serve: prompt 82904 tokens = 81920 reused + 984 read in 2965 ms (331.9 tok/s), 128 generated in 1931 ms (66.3 tok/s), drafts accepted 70 of 120, 6 checkpoints
```

So the calibration path runs at ~1/18 of the server path's decode speed and ~1/14 of its prompt speed. Draft acceptance is normal, so MTP is not the cause.

### Ruled out
- **A second engine holding memory:** only one `strata.exe` was running during the calibration.
- **Memory pressure:** 6.4 GB of RAM free, the pagefile at 158 MB, ~13 hard page-ins/s.
- **CPU saturation from something else:** ~53% total CPU during the calibration.
- **Experts not loaded:** the experts load at 5.1 GiB/s and the expert cache fills (11,573-11,661 experts, 15.5 GiB) just as with the server.

### Expected
The calibration measures the speed the server delivers (~74 tok/s here). Otherwise it should refuse to save: as it stands, any "winner" it picks is chosen from noise in a different regime and would be written into the config.

### Possibly relevant
`tools/calibrate.py` starts `StrataEngine` from the setup process (`START-HERE.bat` → `setup.py` → `calibrate.run`), not from `server.py`. I have not found what differs between the two launches (env from `child_env(cfg)`, priority, console/job object, power throttling of a child of the setup console?). The README's Task Scheduler note shows Windows can throttle the engine depending on how it is started. This may be another case of that, hitting compute rather than the expert load.

A sanity check could catch it: compare the calibration's default-settings rate with a short server-path measurement, or warn when the rate is far below the published figure for the GPU.

### Workaround
Skip `--calibrate`. I tune with A/B runs through the server (`--pcie-frac`, `--spec-min-p`, `--pool-workers` appended to the config's args, one server per arm). Happy to post those results here if useful.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.