Issues / #597
#597 setup: a draft vocabulary for French, acceptance 0.52 to 0.64 and decode 11% faster on an RTX 5090
closed · @gputier · 4 comments · View on GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Description
Based on v0.1.38.
French answers draft poorly with the shipped subsets. `draft_vocab_en.bin` holds 40,525 ids, built from English and code. On 8 French Wikipedia articles, 23.8% of the token occurrences fall outside it. The draft head cannot propose those tokens, so the draft stops there. `cjk` and `cyrillic` add whole scripts, but no subset adds the accented Latin tokens French needs.
I built a subset for French and measured it against `en`. Both arms use the same model and settings. Only `draft_vocab.bin` in the `--mtp` folder differs.
## What the engine and setup do today
- The engine reads `<--mtp folder>/draft_vocab.bin` (int32 ids) at load: `src/core/mtp.cpp:433` (the list itself) and `src/core/mtp.cpp:309` (the VRAM estimate in `bind_bytes`). Any list of ids works, nothing in the engine is tied to a language.
- `setup.py` copies one of three shipped files into that folder: `DRAFT_VOCABS = {"cjk": ..., "en": ..., "cyrillic": ...}` at `setup.py:2740`, chosen by `--draft-vocab` (`setup.py:2968`), applied by `refresh_draft_vocab` (`setup.py:2824`).
- `tools/draft_vocab.py --add` accepts `han`, `kana`, `hangul`, `cjk_punct`, `cjk` or `cyrillic` (`tools/draft_vocab.py:11`, the `SCRIPTS` table at `:24`, the check at `:73`). There is no Latin block. A block by code point range would not help here anyway: the `en` list already holds the plain Latin letters, and what French lacks is specific words and word pieces.
## The list
The `en` list (40,525 ids) plus the 6,062 tokens that cover 99% of the occurrences in the 8 French articles, 46,587 ids in all. The corpus is plain text from French Wikipedia: 982,297 characters, 258,528 tokens.
- Occurrences outside the list on that corpus: 23.8% with `en`, 0.55% with the new list.
- Draft head, as the engine log reports it (`draft head over N tokens`): 81.2 MiB, then 93.3 MiB (+12 MiB).
The 99% cut and the 8 articles are my choice. I did not try other cut-offs or a larger corpus.
## Measurement
RTX 5090, Windows, IQ3_S, engine 0.1.38, `--spec 4`, temperature 0, 300 tokens per request, no reasoning. 12 French prompts written for this test and 4 code prompts. Two French prompts share a topic with corpus articles (the French Revolution, the Second World War); without them the median ratio below is 1.106 instead of 1.110. The arms alternate (`en`, new, `en`, new, `en`, new), so 3 passes each. Each server start runs 2 warm-up prompts that are not counted.
French, 12 prompts:
- Acceptance (drafts accepted over drafts offered): 0.516 with `en`, 0.638 with the new list. The 3 passes of each arm sit within 0.012 of each other.
- Decode: 141.0 tok/s with `en`, 157.5 tok/s with the new list (mean over the 12 prompts and 3 passes).
- Per prompt, ratio of the 3-pass means, new over `en`: median 1.110, from 1.063 to 1.216. The new list is ahead on 12 prompts out of 12, and on each of the 3 passes taken alone.
- Noise: the same prompt on the same arm differs by 2.3% (median) between two passes, and the pass means of one arm spread 2 to 3%.
Code, 4 prompts:
- Acceptance: 0.846 with `en`, 0.838 with the new list.
- Decode ratio, median per prompt: 0.986, with 4 to 6% spread between the passes of one arm. The matched passes disagree (the new list is ahead on 3, 1 and 0 prompts out of 4), so I read it as undecided. 4 prompts cannot separate a 1 to 2% loss from noise.
English was not measured in this run. An earlier, smaller run (2 English prompts) gave acceptance 0.591 with `en` and 0.599 with the new list.
The engine is not deterministic at temperature 0 on this machine: two passes of the same configuration gave different text on 7 prompts out of 8. I compared rates, not texts.
## What I did not measure
- Other GPUs, other quantizations, AMD, Linux.
- Other Latin-script languages (Spanish, German, Italian...). The list is built from French text and I make no claim for them.
- The effect of +12 MiB on a small card. It is the only cost I looked at.
- Prompt ingestion, and the quality of the answers.
## Ask
Two ways to ship this, and I would rather you pick:
1. A coverage mode in `tools/draft_vocab.py`: add to a base list the tokens that cover a given share of a text corpus, so a language can be added from its own text.
2. A fourth value `fr` in `--draft-vocab` and `DRAFT_VOCABS`, with a `data/draft_vocab_fr.bin` of about 182 KB (46,587 int32 ids) next to the other three.
I can send the script that builds the list (it takes a text file, the tokenizer files, the base list and a coverage threshold) and the list itself, as a PR in whichever shape you choose. If you want another corpus or another threshold first, tell me which.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.