Pull requests / #1345

#1345 calibrate: the sweep visits every value in every round instead of once

open · @aly8246 · 0 comentários · No GitHub

Benchmarks

Descrição

## What this fixes

Phase 1 measures each PCIe share **once** and takes the maximum. A `rate()` is a median of three 128-token generations, but all three run at the same setting in one block, so the whole measurement sits in one piece of the machine's state, and adjacent values cannot be ordered that way. From the sweep in #1332, on the same box:

| `pcie_frac` | 0.75 | 0.85 | 0.90 | 0.95 |
| --- | --- | --- | --- | --- |
| measured | 72.7 | 69.0 | 71.0 | 71.9 |

0.2% apart, while those same settings moved 8-28% between engine starts. @enkynakamura reports the same drift from the other side: +9.57 tok/s between two blocks at an unchanged setting on a 5080/Xeon box (p = 0.00004), larger than the effect being hunted. His interleaved run shows what it costs at the low end: 0.27 through 0.40 sit inside each other's noise (sd 2.3-2.7 tok/s) while 0.20 and 0.50 resolve cleanly against 0.33 (p = 0.0039 and p = 0.0005).

The consequence is not only a noisy number. Phase 1's winner is the share phase 2 then sweeps `spec_min_p` at, and the confirm step puts only that one resulting pair against the default, so a value that was worth a real gain can lose its single sample and never be measured again. Widening the grid, as #1332 does, widens that exposure, which is why this should land with it or before it.

## What it does

`sweep()` measures every value `SWEEP_ROUNDS` times (3 by default) inside the one engine start, one round visiting every value and every other round in reverse order, and the winner comes from the medians:

```
    round 1/3 PCIe share: 0 45.1  0.2 47.3  0.35 49.0  ...
    round 2/3 PCIe share: 0.55 46.9  0.5 48.8  ...  0 45.4     <- reversed
    round 3/3 PCIe share: 0 45.3  0.2 47.2  ...
    PCIe share, median of 3: 0: 45.2 (45.1-45.4), 0.2: 47.3 (47.2-47.4), ... -> best 0.35
```

Drift inside a sweep then lands on every value the same number of times, the median keeps one bad sample from deciding, and the per-value range is printed so a user can see whether the winner is real. Phase 2 uses the same function, and the confirm step uses the same round count instead of a hard-coded three.

## Cost

The sweep multiplies by `SWEEP_ROUNDS`. A `rate()` is three prompts of `MAX_NEW` 128 tokens, so at 95-100 tok/s and six values phase 1 goes from about 40 s to about 120 s, on a run that already loads the model more than once. It is a module constant, and a machine whose two leaders stay inside each other's noise can raise it.

## Not in this change

Phase 4 measures the CPU worker count, which needs a restart per value and cannot be interleaved, so each count is still one measurement. A value there would need a different shape.

Tests: `python -m unittest tools.test_calibrate` passes unchanged, because the stand-in engine is a pure function of the tune keys and the median of identical samples is that sample. `tools.test_setup_config` and `tools.test_setup_golden` pass as well.

Refs #1332, #1337.

No site

Links install, modelos, releases.