Pull requests / #1186
#1186 setup: --draft-vocab es - the English/code subset plus the tokens of Spanish text
open · @txiki739 · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
描述
## What
`--draft-vocab es`: the English/code subset plus the tokens of Spanish text, built exactly the way the French subset was (#597): `tools/draft_vocab.py --corpus`, coverage 0.99, on top of `data/draft_vocab_en.bin`.
Spanish answers draft poorly with the shipped subsets: **27.3%** of the token occurrences of Spanish text fall outside `draft_vocab_en.bin`, and the draft head cannot propose those.
- `data/draft_vocab_es.bin`: `draft_vocab_en.bin` plus 6,671 tokens (**47,196 ids**), from 10 Spanish Wikipedia articles (España, Madrid, Miguel de Cervantes, Idioma español, Unión Europea, Literatura española, Gastronomía de España, Economía de España, Guerra civil española, Segunda Guerra Mundial; 1.31 M characters, 326,738 tokens):
```
python tools/draft_vocab.py --gguf <shard 1> --base data/draft_vocab_en.bin --corpus <the articles> \
--coverage 0.99 --out data/draft_vocab_es.bin
```
- `setup.py`: `"es"` in `DRAFT_VOCABS` / `--draft-vocab` (~155 MiB head on IQ3_S, between `fr` and `cyrillic`), shipped like `fr`: a later `--draft-vocab cjk/en` replaces it, a subset made by hand stays.
- `docs/DETAILS.md`: the subset and what it measured, next to the French one.
- `tools/test_setup_draft_vocab.py`: `es` replaces a shipped subset and back, keeps the `en` ids first in their order, its MiB estimate, no small-card note when chosen.
The default subset (`cjk`) is unchanged.
## Coverage
| Spanish text | outside `en` | outside `es` |
| --- | ---: | ---: |
| the 10 articles (the corpus) | 27.25% | 0.86% |
| held out: México, Argentina, Gabriel García Márquez (366 K characters, Latin American Spanish) | 28.27% | 4.20% |
## Measured
RTX 3090 + Ryzen 7 5700X, UD-IQ4_XS, greedy, three Spanish prompts (a chat answer, a reasoning answer, an 18K-token document summary) x 2 passes. Only the draft subset file differs between the arms.
| | `en` | `es` |
| --- | ---: | ---: |
| drafts accepted, Spanish prompts | 0.695 | **0.754** |
| decode, Spanish prompts (mean) | 75.4 tok/s | **82.7 tok/s** |
| drafts accepted, a code prompt asked in Spanish | 0.730 | 0.819 |
| decode, that code prompt | 79.1 tok/s | 85.7 tok/s |
To be clear about the engine: these acceptance numbers come from a 0.1.30-based fork engine (eddoursul's) whose draft head reads the same `mtp/rt/draft_vocab.bin` subset file; I have not rerun them on 0.1.40. The coverage numbers are the tool's own and engine-independent.
## Tests
`python -m unittest tools.test_setup_draft_vocab tools.test_draft_vocab`: 13 tests, OK.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。