Pull requests / #608

#608 setup: a 200K context between the 128K rule and 256K (#406)

closed · @MingShi350 · 0 评论 · 在 GitHub 查看

Setup & installServer & APINVIDIA / CUDAModels & quants

描述

## What

`CONTEXTS` jumps 131072 → 262144. #406 lets a PC below setup's 256K estimate keep 256K with a note, but there was nothing *between* the 128K recommendation and a size that PC may not hold — the only way to find out was to try 256K and watch the RAM.

204800 (200K) is inside the trained 262144, so it needs **no rope scaling**, and with IQ3_S it costs **~2.8 GB** of 8-bit KV against ~3.6 GB at 256K (13.7 KB per context token).

## Measured

RTX 4080 16 GB, Ryzen 9 9950X3D, 62 GB usable, **IQ3_XXS**, engine 0.1.37, 8-bit KV. At this size setup turns KV streaming **off** (the experts need 47 GB), so the KV sits in VRAM and a longer context costs the expert cache:

| | 128K | 200K |
|---|---:|---:|
| decode, shallow (median of 10, warm) | 94.1 t/s | **85.1 t/s** |
| decode, 30K deep | 87.2 t/s | **85.3 t/s** |
| decode, 124K deep | 94.1 t/s | 79.2 t/s |
| prefill | 3257–3284 t/s | 3189–3318 t/s |
| experts cached | 3661 | 2951 (−710) |
| a 203259-token prompt | — | 3093 t/s prefill, 68.6 t/s decode |

So 200K keeps prefill and shallow decode within 10%, and reaches a depth (124K–203K) that 128K does not — at 203K tokens the model still answers at 68.6 t/s.

## Files

- `setup.py` — `CONTEXTS` gains 204800, with these numbers in the comment.
- `tools/strata_mcp.py` — `FALLBACK_CONTEXTS` mirrors `CONTEXTS`, so a PC that cannot import setup sees the same menu.
- `tools/test_setup_risk.py` — the 256K menu pick is `"6"` now; a new test covers a 200K pick (kept, no rope scaling, the RAM note at 64 GB). `python -m unittest tools.test_setup_risk` → **35 tests, OK**.

## Notes

- `ram_ctx()` (`setup.py:2197`) now recommends 200K on a PC where the 200K estimate fits and the 256K one does not — the intended middle step. On a 64 GB PC nothing changes: `ctx_ram_need("IQ3_S", 204800)` is 77.1 GB, so the rule still recommends 128K.
- Every `tools/test_setup_*` suite passes on this branch. `tools/test_setup_golden.py` fails 46 times on this checkout **before and after** the change (identical failure list — an environment thing here), so this PR does not affect it.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。