Pull requests / #1186

#1186 setup: --draft-vocab es - the English/code subset plus the tokens of Spanish text

open · @txiki739 · 0 コメント · GitHub で見る

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

本文

## What

`--draft-vocab es`: the English/code subset plus the tokens of Spanish text, built exactly the way the French subset was (#597): `tools/draft_vocab.py --corpus`, coverage 0.99, on top of `data/draft_vocab_en.bin`.

Spanish answers draft poorly with the shipped subsets: **27.3%** of the token occurrences of Spanish text fall outside `draft_vocab_en.bin`, and the draft head cannot propose those.

- `data/draft_vocab_es.bin`: `draft_vocab_en.bin` plus 6,671 tokens (**47,196 ids**), from 10 Spanish Wikipedia articles (España, Madrid, Miguel de Cervantes, Idioma español, Unión Europea, Literatura española, Gastronomía de España, Economía de España, Guerra civil española, Segunda Guerra Mundial; 1.31 M characters, 326,738 tokens):
  ```
  python tools/draft_vocab.py --gguf <shard 1> --base data/draft_vocab_en.bin --corpus <the articles> \
      --coverage 0.99 --out data/draft_vocab_es.bin
  ```
- `setup.py`: `"es"` in `DRAFT_VOCABS` / `--draft-vocab` (~155 MiB head on IQ3_S, between `fr` and `cyrillic`), shipped like `fr`: a later `--draft-vocab cjk/en` replaces it, a subset made by hand stays.
- `docs/DETAILS.md`: the subset and what it measured, next to the French one.
- `tools/test_setup_draft_vocab.py`: `es` replaces a shipped subset and back, keeps the `en` ids first in their order, its MiB estimate, no small-card note when chosen.

The default subset (`cjk`) is unchanged.

## Coverage

| Spanish text | outside `en` | outside `es` |
| --- | ---: | ---: |
| the 10 articles (the corpus) | 27.25% | 0.86% |
| held out: México, Argentina, Gabriel García Márquez (366 K characters, Latin American Spanish) | 28.27% | 4.20% |

## Measured

RTX 3090 + Ryzen 7 5700X, UD-IQ4_XS, greedy, three Spanish prompts (a chat answer, a reasoning answer, an 18K-token document summary) x 2 passes. Only the draft subset file differs between the arms.

| | `en` | `es` |
| --- | ---: | ---: |
| drafts accepted, Spanish prompts | 0.695 | **0.754** |
| decode, Spanish prompts (mean) | 75.4 tok/s | **82.7 tok/s** |
| drafts accepted, a code prompt asked in Spanish | 0.730 | 0.819 |
| decode, that code prompt | 79.1 tok/s | 85.7 tok/s |

To be clear about the engine: these acceptance numbers come from a 0.1.30-based fork engine (eddoursul's) whose draft head reads the same `mtp/rt/draft_vocab.bin` subset file; I have not rerun them on 0.1.40. The coverage numbers are the tool's own and engine-independent.

## Tests

`python -m unittest tools.test_setup_draft_vocab tools.test_draft_vocab`: 13 tests, OK.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

関連リンク

インストール・モデル・リリースへの站内リンク。