Pull requests / #1219

#1219 setup: a Spanish draft subset (--draft-vocab es); Spanish answers dra…

open · @elanonimo832 · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quants

描述

…ft 12% more

The project ships four draft subsets (cjk, en, cyrillic, fr) and none for Spanish, so a Spanish answer drafts against English and code tokens only. The French one, added in #597, is the model to follow: as tools/draft_vocab.py puts it, a language written in a script the base already holds (French, Spanish, German... in Latin letters) needs specific words and word pieces, not a whole script - and it names Spanish as the next one. This adds that file.

data/draft_vocab_es.bin (52,026 ids) is data/draft_vocab_en.bin plus the 11,501 tokens that cover 99% of a Spanish corpus:

  25 public-domain books of 1837-1921 from Project Gutenberg, ids
    15725 17340 17013 16961 17223 18005 58059 29506 14944 24127 31013 35407 10909 57781 30122 12848
    59852 46182 42440 52894 57982 60675 40066 30535 63509
    (Galdos, Valera, Pardo Bazan, Alarcon, Blasco Ibanez, Pereda, Palacio Valdes, Becquer, Hartzenbusch,
     Baroja, Unamuno, Valle-Inclan, Ruben Dario, Ortega y Gasset, Ramon y Cajal, ...; novel, theatre,
     poetry, history and memoirs)
  200 Spanish Wikipedia articles (wikimedia/wikipedia, config 20231101.es), sampled with 5 requests at
    offsets 0, 368231, 736462, 1104693, 1472924 (100 rows each), for the modern register

  tools/draft_vocab.py --gguf <model>-00001-of-00002.gguf --base data/draft_vocab_en.bin       --corpus <the 25 texts> <the 200 articles> --coverage 0.99 --out data/draft_vocab_es.bin

Measured (IQ3_XXS, RTX 5080 Laptop), on text left out of the corpus:
  - 4 books of the same period: 22.5% of their tokens were outside the English/code subset and 1.1% are outside this one; on Wikipedia text left out of it, 23.6% -> 3.4%. The default cjk subset: 22.5% and 23.6% (it holds 106,299 ids, this one 52,026).
  - one 900-token Spanish answer at temperature 0: drafts accepted 0.63 -> 0.71 and 64.8 -> 70.0 tok/s. With the default cjk subset on the same answer: 0.63 accepted, 64.8 tok/s.

tools/test_setup_draft_vocab.py gains the two checks the French subset has (a shipped subset is replaced and back, a hand-made one is kept; the file holds the en base first, unchanged, with the added ids sorted). test_sizes_follow_the_shipped_subsets now covers it too: DRAFT_VOCAB_MIB["es"] = 170, the map's own scaling (348 MiB at 106,299 ids).

sha256 of data/draft_vocab_es.bin: 9c445b6b75badb0d46d349c2feffa210b8117a35435c11e76d9f163670fa2388

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。