Pull requests / #664

#664 setup, tools: a draft vocabulary for French, and draft_vocab.py builds one from a text (#597)

closed · @gputier · 0 comentarios · En GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindowsLinux

Descripción

The PR you asked for in #597, in its two parts.

**1. `tools/draft_vocab.py --cover-text <text> [--cover 0.99]`**

French shares its script with English, so `--add` has no range for it. The new mode takes a UTF-8 text in the language and adds its most frequent tokens: the ones that make `--cover` of the text's token occurrences, less those the base already has. Most frequent first (the first seen first, among equals), added in id order after what `--add` brought, so the same text gives the same subset. It prints the share of the text's tokens outside the base and outside the new subset.

**2. `data/draft_vocab_fr.bin` and `--draft-vocab fr`**

`fr` is in `DRAFT_VOCABS` and `DRAFT_VOCAB_MIB` (153 MiB, by the token count like the others), in `--draft-vocab`'s help, in `refresh_draft_vocab` and in the small-card note. DETAILS.md has the paragraph.

The file is the output of this command, on the IQ3_S pack's first shard:

```
python tools/draft_vocab.py --gguf Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf \
    --base data/draft_vocab_en.bin --cover-text corpus_fr.txt --out data/draft_vocab_fr.bin
corpus_fr.txt: 258528 tokens, 13746 distinct; outside the base 23.80%, outside the new subset 0.55%
wrote data/draft_vocab_fr.bin: 46587 ids (6062 added)
```

Its sha256 is `f236f21584a1483ad5f61f261dfb0239216028821baf045d11df25f851e9dcd0`, the same bytes as the list measured in #597.

**The corpus** is not in the PR (1 MB of CC BY-SA text). It is the plain-text extract (`action=query&prop=extracts&explaintext=1`) of 8 French Wikipedia articles, joined by a blank line in this order, fetched on 2026-10-03 at 06:04 UTC. The revisions at that time:

| Article | revid |
| --- | --- |
| Histoire de France | 240011002 |
| Photosynthèse | 238282939 |
| Révolution française | 239999162 |
| Économie de la France | 239479553 |
| Cuisine française | 239904005 |
| Changement climatique | 239281043 |
| Seconde Guerre mondiale | 239999269 |
| Paris | 239771294 |

I can attach the file if you want to regenerate the bytes.

**Tests**

- `tools/test_draft_vocab.py` (new, 2 tests): `covering()` on small counts, the share and the order among equal counts.
- `tools/test_setup_draft_vocab.py` (6 tests): `fr` starts with the English/code subset in its order, `refresh_draft_vocab` copies it, its MiB follows its token count, the small-card note stays silent when `fr` was chosen.
- Both pass on Linux (Python 3.13). The command above was run on Windows 11 (Python 3.13.2).

**Not in this PR**

The engine's hint when the draft head does not fit (`src/core/mtp.cpp`, the `smaller[]` table) and the server's `DRAFT_HEAD_HINT` still name `cyrillic` and `en` only. `fr` (46,587 ids) would sit between them. I left the engine out of a setup and tools PR; say so if you want it here.

The measurement behind the 11% is in #597: 12 French prompts, 300 tokens each, IQ3_S on an RTX 5090, acceptance 0.52 to 0.64, 141 to 158 tokens/s.

En el sitio

Enlaces a install, modelos, releases.