Issues / #597

#597 setup: a draft vocabulary for French, acceptance 0.52 to 0.64 and decode 11% faster on an RTX 5090

closed · @gputier · 4 コメント · GitHub で見る

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

本文

Based on v0.1.38.

French answers draft poorly with the shipped subsets. `draft_vocab_en.bin` holds 40,525 ids, built from English and code. On 8 French Wikipedia articles, 23.8% of the token occurrences fall outside it. The draft head cannot propose those tokens, so the draft stops there. `cjk` and `cyrillic` add whole scripts, but no subset adds the accented Latin tokens French needs.

I built a subset for French and measured it against `en`. Both arms use the same model and settings. Only `draft_vocab.bin` in the `--mtp` folder differs.

## What the engine and setup do today

- The engine reads `<--mtp folder>/draft_vocab.bin` (int32 ids) at load: `src/core/mtp.cpp:433` (the list itself) and `src/core/mtp.cpp:309` (the VRAM estimate in `bind_bytes`). Any list of ids works, nothing in the engine is tied to a language.
- `setup.py` copies one of three shipped files into that folder: `DRAFT_VOCABS = {"cjk": ..., "en": ..., "cyrillic": ...}` at `setup.py:2740`, chosen by `--draft-vocab` (`setup.py:2968`), applied by `refresh_draft_vocab` (`setup.py:2824`).
- `tools/draft_vocab.py --add` accepts `han`, `kana`, `hangul`, `cjk_punct`, `cjk` or `cyrillic` (`tools/draft_vocab.py:11`, the `SCRIPTS` table at `:24`, the check at `:73`). There is no Latin block. A block by code point range would not help here anyway: the `en` list already holds the plain Latin letters, and what French lacks is specific words and word pieces.

## The list

The `en` list (40,525 ids) plus the 6,062 tokens that cover 99% of the occurrences in the 8 French articles, 46,587 ids in all. The corpus is plain text from French Wikipedia: 982,297 characters, 258,528 tokens.

- Occurrences outside the list on that corpus: 23.8% with `en`, 0.55% with the new list.
- Draft head, as the engine log reports it (`draft head over N tokens`): 81.2 MiB, then 93.3 MiB (+12 MiB).

The 99% cut and the 8 articles are my choice. I did not try other cut-offs or a larger corpus.

## Measurement

RTX 5090, Windows, IQ3_S, engine 0.1.38, `--spec 4`, temperature 0, 300 tokens per request, no reasoning. 12 French prompts written for this test and 4 code prompts. Two French prompts share a topic with corpus articles (the French Revolution, the Second World War); without them the median ratio below is 1.106 instead of 1.110. The arms alternate (`en`, new, `en`, new, `en`, new), so 3 passes each. Each server start runs 2 warm-up prompts that are not counted.

French, 12 prompts:

- Acceptance (drafts accepted over drafts offered): 0.516 with `en`, 0.638 with the new list. The 3 passes of each arm sit within 0.012 of each other.
- Decode: 141.0 tok/s with `en`, 157.5 tok/s with the new list (mean over the 12 prompts and 3 passes).
- Per prompt, ratio of the 3-pass means, new over `en`: median 1.110, from 1.063 to 1.216. The new list is ahead on 12 prompts out of 12, and on each of the 3 passes taken alone.
- Noise: the same prompt on the same arm differs by 2.3% (median) between two passes, and the pass means of one arm spread 2 to 3%.

Code, 4 prompts:

- Acceptance: 0.846 with `en`, 0.838 with the new list.
- Decode ratio, median per prompt: 0.986, with 4 to 6% spread between the passes of one arm. The matched passes disagree (the new list is ahead on 3, 1 and 0 prompts out of 4), so I read it as undecided. 4 prompts cannot separate a 1 to 2% loss from noise.

English was not measured in this run. An earlier, smaller run (2 English prompts) gave acceptance 0.591 with `en` and 0.599 with the new list.

The engine is not deterministic at temperature 0 on this machine: two passes of the same configuration gave different text on 7 prompts out of 8. I compared rates, not texts.

## What I did not measure

- Other GPUs, other quantizations, AMD, Linux.
- Other Latin-script languages (Spanish, German, Italian...). The list is built from French text and I make no claim for them.
- The effect of +12 MiB on a small card. It is the only cost I looked at.
- Prompt ingestion, and the quality of the answers.

## Ask

Two ways to ship this, and I would rather you pick:

1. A coverage mode in `tools/draft_vocab.py`: add to a base list the tokens that cover a given share of a text corpus, so a language can be added from its own text.
2. A fourth value `fr` in `--draft-vocab` and `DRAFT_VOCABS`, with a `data/draft_vocab_fr.bin` of about 182 KB (46,587 int32 ids) next to the other three.

I can send the script that builds the list (it takes a text file, the tokenizer files, the base list and a coverage threshold) and the list itself, as a PR in whichever shape you choose. If you want another corpus or another threshold first, tell me which.

関連リンク

インストール・モデル・リリースへの站内リンク。