Issues / #587

#587 tools/make_profile.py --base cannot reorder a complete base (shipped profile reports 0 from the traces)

closed · @yannickloth · 1 comentarios · En GitHub

AMD / HIPNVIDIA / CUDAModels & quants

Descripción

`tools/make_profile.py` is documented to rank "the base profile's ranking, then the pairs your routing traces used, most frequent first" so a `--dump-routing` trace of your workload seeds the cache with the experts it actually routes. But the base is only ever *appended to*, and the shipped `data/expert-profile.bin` already ranks all 24,576 `(layer, expert)` pairs — so with it the trace cannot change anything:

```
$ python tools/make_profile.py coding-trace.bin --base data/expert-profile.bin --out /tmp/prof.bin
wrote /tmp/prof.bin: 24576 ranked pairs (24576 from the base, 0 from the traces, 0 filled in)
```

`/tmp/prof.bin` is byte-identical to `data/expert-profile.bin`. The only way to let a trace order the top is `--no-base`, which throws the shipped ranking away entirely.

In `rank_profile`/`main`, `take(read_profile(base))` fills all 24,576 slots before the trace pairs are considered, and `take()` skips anything already seen. This matters because `--expert-cache auto` fills the VRAM tier from the profile *in order*; a workload-specific profile therefore never gets to put its most-routed experts in the slots that fit. (Measured on an RTX A3000 12 GB / Swift 1.5 IQ3_XXS: the shipped base is 2,107 resident slots; a workload-ranked profile raised the fresh-conversation hit rate 0.468 → 0.528, but only via `--no-base`.)

### Proposed fix

Add `--reorder`: rank the trace's pairs first (most frequent first), then let the base fill the rest, then the fill — so a trace orders the top while the base still covers everything the trace did not see. Default behaviour is unchanged. PR: #589.

En el sitio

Enlaces a install, modelos, releases.