Pull requests / #287

#287 setup: --draft-vocab cyrillic (English/code + the whole Cyrillic script)

closed · @sergqwer · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installAMD / HIPModels & quants

描述

## Summary

The shipped draft subsets hold 142 of the vocabulary's 18,580 Cyrillic tokens, so an answer in Ukrainian, Russian,
Bulgarian or Serbian drafts almost nothing. This is #137 again, for another script.

`--draft-vocab cyrillic` takes `data/draft_vocab_cyrillic.bin`: the English/code subset plus every token whose text
is Cyrillic (58,963 ids). It was built with the repository's own tool:

```
python tools/draft_vocab.py --gguf <IQ2_XS shard 1> --base data/draft_vocab_en.bin --add cyrillic --out data/draft_vocab_cyrillic.bin
```

`tools/draft_vocab.py` gains the `cyrillic` script (U+0400-052F, 1C80-1C8F, 2DE0-2DFF, A640-A69F). setup.py and
DETAILS.md name the option. The default stays `cjk`.

## Measured

On the NVFP4 fork (same draft layer, same tokenizer): Ukrainian answers went from 1.4 to 2.1 tokens per round and
from 83 to 109 tok/s; English is unchanged. The file is byte-identical to the fork's.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。