Pull requests / #664
#664 setup, tools: a draft vocabulary for French, and draft_vocab.py builds one from a text (#597)
closed · @gputier · 0 comments · View on GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindowsLinux
Description
The PR you asked for in #597, in its two parts.
**1. `tools/draft_vocab.py --cover-text <text> [--cover 0.99]`**
French shares its script with English, so `--add` has no range for it. The new mode takes a UTF-8 text in the language and adds its most frequent tokens: the ones that make `--cover` of the text's token occurrences, less those the base already has. Most frequent first (the first seen first, among equals), added in id order after what `--add` brought, so the same text gives the same subset. It prints the share of the text's tokens outside the base and outside the new subset.
**2. `data/draft_vocab_fr.bin` and `--draft-vocab fr`**
`fr` is in `DRAFT_VOCABS` and `DRAFT_VOCAB_MIB` (153 MiB, by the token count like the others), in `--draft-vocab`'s help, in `refresh_draft_vocab` and in the small-card note. DETAILS.md has the paragraph.
The file is the output of this command, on the IQ3_S pack's first shard:
```
python tools/draft_vocab.py --gguf Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf \
--base data/draft_vocab_en.bin --cover-text corpus_fr.txt --out data/draft_vocab_fr.bin
corpus_fr.txt: 258528 tokens, 13746 distinct; outside the base 23.80%, outside the new subset 0.55%
wrote data/draft_vocab_fr.bin: 46587 ids (6062 added)
```
Its sha256 is `f236f21584a1483ad5f61f261dfb0239216028821baf045d11df25f851e9dcd0`, the same bytes as the list measured in #597.
**The corpus** is not in the PR (1 MB of CC BY-SA text). It is the plain-text extract (`action=query&prop=extracts&explaintext=1`) of 8 French Wikipedia articles, joined by a blank line in this order, fetched on 2026-10-03 at 06:04 UTC. The revisions at that time:
| Article | revid |
| --- | --- |
| Histoire de France | 240011002 |
| Photosynthèse | 238282939 |
| Révolution française | 239999162 |
| Économie de la France | 239479553 |
| Cuisine française | 239904005 |
| Changement climatique | 239281043 |
| Seconde Guerre mondiale | 239999269 |
| Paris | 239771294 |
I can attach the file if you want to regenerate the bytes.
**Tests**
- `tools/test_draft_vocab.py` (new, 2 tests): `covering()` on small counts, the share and the order among equal counts.
- `tools/test_setup_draft_vocab.py` (6 tests): `fr` starts with the English/code subset in its order, `refresh_draft_vocab` copies it, its MiB follows its token count, the small-card note stays silent when `fr` was chosen.
- Both pass on Linux (Python 3.13). The command above was run on Windows 11 (Python 3.13.2).
**Not in this PR**
The engine's hint when the draft head does not fit (`src/core/mtp.cpp`, the `smaller[]` table) and the server's `DRAFT_HEAD_HINT` still name `cyrillic` and `en` only. `fr` (46,587 ids) would sit between them. I left the engine out of a setup and tools PR; say so if you want it here.
The measurement behind the 11% is in #597: 12 French prompts, 300 tokens each, IQ3_S on an RTX 5090, acceptance 0.52 to 0.64, 141 to 158 tokens/s.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.