Issues / #906

#906 Older PC (Haswell, DDR3, PCIe 3.0 x8): +15% decode on Q2_0 and IQ3_S from #706, #764 and a busier adaptive tier

open · @1872183316 · 1 Kommentare · Auf GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows

Beschreibung

Measured on an older PC (Haswell Xeon, DDR3, PCIe 3.0 x8), where the CPU's share of the experts is most of a verify
window. Three changes together give **+15.3% decode on Q2_0 and +14.4% on IQ3_S**, without changing any quantization:
PR #706's AVX2 Q2_0 kernel, PR #764 (the swaps take effect a window later), and an adaptive expert tier that swaps
more. All numbers below are from this one PC; the official RTX 5070 / Ryzen 5 7600 figures are on different hardware.

## The PC

- Xeon E5-2673 v3 (12 cores, AVX2, no AVX-512), 3 x 32 GB DDR3-1866 LRDIMM (see "Memory" below), RTX 4060 Ti 16 GB on
  PCIe 3.0 x8 (engine probe 6.2 GB/s), SATA SSD, Ubuntu 24.04, no display on the card.
- Engine: main at 6f32ec0 (0.1.39) built from source (CUDA 12.8, sm_89, `-DSTRATA_EXPERIMENTAL_SM60=ON` as setup does
  for CUDA 12), with #706's two commits and #764 ported by hand.
- Models: the original GSQ-RCO Q2_0 and IQ3_S, SHA-256 as published.
- Config: what setup writes for this PC (64K context, `--kv int8 --kv-resident 32768`, `--spec 4 --spec-min-p 0.5`,
  `--expert-cache auto` = 7,475 slots for Q2_0 and 3,531 for IQ3_S, `--prefill auto`).
- Method: the server, 4 prompts (a Chinese essay, a code edit, an English-to-Chinese translation, an explanation),
  greedy, 512 tokens, thinking off, 2 passes each; one binary, the arms switched by environment variables and
  interleaved A/B/A/B, the page cache dropped before each start.

## Result

"0.1.39 path" = `STRATA_Q2_LEGACY=1` (#706's switch back to the current kernel) + `STRATA_ADAPT_LAG=1` (#463's wait) +
the default tier. "Improved" = #706 + #764 + `--adapt-every 1 --adapt-swaps 160 --adapt-decay 0.97`.

| Model | 0.1.39 path (two runs) | Improved (two runs) | |
| --- | ---: | ---: | ---: |
| **Q2_0** | 48.08 / 48.33 tok/s | 55.40 / 55.80 tok/s | **+15.3%** |
| essay / code edit / translation / explanation | 46.4 / 46.6 / 44.7 / 55.2 | 53.1 / 59.7 / 46.1 / 63.6 | |
| **IQ3_S** | 30.52 / 30.51 tok/s | 34.79 / 35.05 tok/s | **+14.4%** |
| essay / code edit / translation / explanation | 31.4 / 29.4 / 25.9 / 35.5 | 38.2 / 34.4 / 27.8 / 39.5 | |

All 16 prompt pairs are faster. The translation gains least: its answer is 103 tokens, too short for the tier to follow
it. Prompt reading is unchanged (3,454 tokens: Q2_0 714 tok/s, IQ3_S 440 tok/s; it is bound by PCIe 3.0 x8).

How the three parts add up on Q2_0 (same method):

| Arms | Mean tok/s |
| --- | ---: |
| current kernel vs #706 (default tier, lag 1) | 48.86 → 50.95 (+4.3%) |
| #764 alone (default tier) | 49.29 → 52.61 (+6.7%; the lag-1 arm ran three times: 48.71, 49.60, 49.55) |
| + `--adapt-every 1 --adapt-swaps 160 --adapt-decay 0.97` | 52.61 → 56.77 (+7.9%) |

## Why the tier helps here

`STRATA_DECODE_TIMING=1`, Q2_0 code-edit windows: 71.4 ms per window, of which 40.1 ms is the CPU's experts (10.4 per
layer-window). A routing trace of the 4-prompt run (a measurement switch on our side: each verify window's routing and
residency per layer, 1,261 windows), replayed offline with the same capacity:

| Policy | Missed experts per layer-window |
| --- | ---: |
| as the engine ran (every 4 / 96 swaps / decay 0.7) | 5.61 |
| the same policy simulated, starting from `data/expert-profile.bin` | 5.47 |
| best fixed set for the whole run (knows the future) | 4.97 |
| every 4 / 200 / 0.9 | 4.27 (44 swaps per window) |
| every 4 / 400 / 0.9 | 3.63 (68 swaps per window) |

The misses follow the swaps per window far more than the counting rule. On the engine the hit rate went from 0.63 to
0.74 (Q2_0 code edit) and the CPU part of those windows from 40.1 to 28.8 ms. Without #764 the larger swap volume is
eaten by the wait for the copies (9.6 ms per window unaccounted with 80 swaps and lag 1); with it the wait is gone.
240 swaps was no better than 160 (56.36 tok/s).

A faster PC may not want this (its CPU reads the misses faster, and the copies cost the same), so rather than a new
default the proposal is to **let `--calibrate` measure the tier**: a restart per candidate, as for the worker count
(default, every 1 / 80 / 0.97, every 1 / 160 / 0.97), kept only above `MIN_GAIN`. We have it implemented with a test
and can open a PR.

## Same answers?

These change only where an expert is computed (VRAM tier or CPU) and #706's rounding, not the quantization. Checked with
`STRATA_LOGPOS` (teacher-forced log-probabilities) on a 3,166-token text read through the verify windows
(`--short-read 8192`, Q2_0), against a noise control that only changes which experts sit in VRAM:

| Against the 0.1.39 path | Mean NLL change (`<|im_end|>` excluded) | Mean \|Δ logprob\| | Argmax agreement |
| --- | ---: | ---: | ---: |
| improved | +0.013 | 0.232 | 96.75% |
| 0.1.39 path, `--expert-cache 7000` | +0.012 | 0.300 | 96.49% |
| 0.1.39 path, `--expert-cache 6500` | +0.016 | 0.262 | 96.68% |

Each configuration gave bit-identical files on two runs. The improved arm moves the numbers less than a smaller cache
does; #706's kernel alone is closer to an exact evaluation of the Q2_0 formula than the current one
(3.4e-8 vs 6.5e-8 relative, random blocks).

## Memory: odd DIMM counts

This board has 3 DIMMs, two on one home agent and one on the other (`/sys/devices/system/edac`): only the first 64 GB
are interleaved over two channels. The same CPU expert kernel reads **26.3 GB/s** from pages there and **16.6 GB/s**
from the single-channel rest. With the page cache full, the arena can land in the slow part. Used X99 boards with three
sticks are common; an even population (or dropping the page cache before a start) matters for the CPU's share.

## What did not help here

- `--calibrate`'s own settings: it chose `--pcie-frac 0.20 --spec-min-p 0.70` (+3% on its English prompts), but on
  the 4 prompts above it measured 49.29 vs 49.59 tok/s. The engine's PCIe probe already picks 0.17 (0.55 would be 18%
  slower here).
- `--spec` 2 / 3 / 5: 43.4 / 47.4 / 50.1 against 50.6 for the default 4.
- `--vram-reserve-mib 200`: 379 more slots, then the draft head no longer fits and the start fails (it needs about 400).
- Predicting the next layer's misses with its router (top-10 recall 54-62% at 42-55% precision) to fetch them over
  PCIe: the link carries ~4.7 experts per layer-window, so at best 5-8% of the CPU time; not pursued.

Mehr auf der Site

Links zu Install, Modellen, Releases.