Pull requests / #1259
#1259 setup: --draft-vocab it - the English/code subset plus the tokens of Italian text
open · @fioruccione · 0 commentaires · Sur GitHub
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Description
`--draft-vocab it`: the Italian counterpart of `fr` (#597), built the same way. It adds one data file, one entry in `setup.py`, the tests and a paragraph in DETAILS.md. The default subset (`cjk`) is unchanged. **Why:** 24.4% of the token occurrences in Italian text fall outside `draft_vocab_en.bin` (French: 23.6%), so the draft head cannot propose accented words and Italian word pieces. **The subset:** `draft_vocab_en.bin` plus the 14,150 tokens that cover 99% of an Italian Wikipedia sample, 54,675 ids in all. The sample is 3,000 articles of the HF dataset `wikimedia/wikipedia`, config `20231101.it` (30 blocks of 100 rows at random offsets, seed 42; 1.06 M words, 1.94 M tokens). It was built with the existing tool: ``` python tools/draft_vocab.py --gguf <UD-IQ4_XS shard 1> --base data/draft_vocab_en.bin --corpus itwiki.txt --coverage 0.99 --out data/draft_vocab_it.bin ``` On that corpus, 0.70% of the occurrences fall outside the new subset (24.38% with `en`). The head is 109.5 MiB on UD-IQ4_XS. The IQ3_S estimate in `DRAFT_VOCAB_MIB` (179) uses the same per-id ratio as `fr`. **Measured:** UD-IQ4_XS on 2x RTX 4060 Ti 16 GB (layer split), Threadripper PRO 3975WX (AVX2), engine 0.1.40, one request at a time. 8 prompts per language (Italian, the same prompts in English, and code in Python / JS / SQL / Bash / Rust / C), 2 passes each, temperature 0, reasoning off, 256 tokens. Each arm is a fresh server start with only `--mtp`'s draft subset changed. Engine metrics (`/metrics`): | draft subset | ids | drafts accepted: Italian | English | code | Italian decode tok/s (median) | | --- | ---: | ---: | ---: | ---: | ---: | | `cjk` (default) | 106,299 | 0.519 | 0.594 | 0.797 | 47.2 | | `en` | 40,525 | 0.508 | 0.604 | 0.797 | 48.0 | | **`it`** | 54,675 | **0.627** | **0.604** | **0.800** | **53.8** | Per prompt against `cjk`: Italian **+14.6%** on average, faster on 16 of 16 runs (+7.7% to +26.7%). English +1.8% and code +3.4%, so nothing is lost. For comparison, `cjk` plus the same Italian tokens (120,422 ids) gave Italian +12.2%, but English slightly below `cjk`. That is why the subset builds on `en`, like `fr`. Even at temperature 0, only 4-5 of the 24 texts per arm were identical between the two passes (the GPU/CPU rounding that BATCHING.md describes). So the comparison above is per-prompt means over 16 runs, not single outputs. **Tests:** `python -m unittest tools.test_setup_draft_vocab tools.test_draft_vocab` passes (11 tests). The `fr` checks (replaces a shipped subset and back, keeps a hand-made one, holds the `en` ids first and in order) now run for `fr` and `it`. Developed with an AI coding assistant; every number above was measured on this machine. Honestly, the AI assistant did the heavy lifting here — I ran the machine and asked the questions :) 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sur le site
Liens install, modèles, releases.