Pull requests / #608
#608 setup: a 200K context between the 128K rule and 256K (#406)
closed · @MingShi350 · 0 commentaires · Sur GitHub
Setup & installServer & APINVIDIA / CUDAModels & quants
Description
## What
`CONTEXTS` jumps 131072 → 262144. #406 lets a PC below setup's 256K estimate keep 256K with a note, but there was nothing *between* the 128K recommendation and a size that PC may not hold — the only way to find out was to try 256K and watch the RAM.
204800 (200K) is inside the trained 262144, so it needs **no rope scaling**, and with IQ3_S it costs **~2.8 GB** of 8-bit KV against ~3.6 GB at 256K (13.7 KB per context token).
## Measured
RTX 4080 16 GB, Ryzen 9 9950X3D, 62 GB usable, **IQ3_XXS**, engine 0.1.37, 8-bit KV. At this size setup turns KV streaming **off** (the experts need 47 GB), so the KV sits in VRAM and a longer context costs the expert cache:
| | 128K | 200K |
|---|---:|---:|
| decode, shallow (median of 10, warm) | 94.1 t/s | **85.1 t/s** |
| decode, 30K deep | 87.2 t/s | **85.3 t/s** |
| decode, 124K deep | 94.1 t/s | 79.2 t/s |
| prefill | 3257–3284 t/s | 3189–3318 t/s |
| experts cached | 3661 | 2951 (−710) |
| a 203259-token prompt | — | 3093 t/s prefill, 68.6 t/s decode |
So 200K keeps prefill and shallow decode within 10%, and reaches a depth (124K–203K) that 128K does not — at 203K tokens the model still answers at 68.6 t/s.
## Files
- `setup.py` — `CONTEXTS` gains 204800, with these numbers in the comment.
- `tools/strata_mcp.py` — `FALLBACK_CONTEXTS` mirrors `CONTEXTS`, so a PC that cannot import setup sees the same menu.
- `tools/test_setup_risk.py` — the 256K menu pick is `"6"` now; a new test covers a 200K pick (kept, no rope scaling, the RAM note at 64 GB). `python -m unittest tools.test_setup_risk` → **35 tests, OK**.
## Notes
- `ram_ctx()` (`setup.py:2197`) now recommends 200K on a PC where the 200K estimate fits and the 256K one does not — the intended middle step. On a 64 GB PC nothing changes: `ctx_ram_need("IQ3_S", 204800)` is 77.1 GB, so the rule still recommends 128K.
- Every `tools/test_setup_*` suite passes on this branch. `tools/test_setup_golden.py` fails 46 times on this checkout **before and after** the change (identical failure list — an environment thing here), so this PR does not affect it.
Sur le site
Liens install, modèles, releases.