Pull requests / #834

#834 bench: community report, RTX 4090 + 32 GB RAM at 512K context (IQ2_XS)

closed · @T-Crypt · 0 comentários · No GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

Descrição

Community report: **512K context on one RTX 4090 + 32 GB of system RAM** (56 GB total memory), Flash-Next GSQ-RCO IQ2_XS, engine 0.1.39.

- `--max-context 524288 --rope-scaling yarn --rope-scale 2 --kv q4_0 --kv-resident 32768 --resident-experts --prefill auto`
- A 476,820-token prompt read in **135.5 s (3,411 tok/s)**, decode **122.0 tok/s at depth**, recall correct at all four sizes tested.
- On a 32 GB box the RAM floor decides it: with the default resident headroom, 0.1.39 left 1.41 GiB `MemAvailable` at 477K. `STRATA_RESIDENT_HEADROOM_GIB=6` left 3.53 GiB and kept the speed.
- `--prefill auto` read long prompts about 3x faster than `--prefill auto:32768` on this box.
- Also includes a short single-4090 A/B of #783 (+2 to +5% decode).

Limits are in the README: yarn past 262K is experimental, one run per size, recall is a single planted fact rather than needle_bench, and the HF revision of the model files wasn't recorded.

Data only: one new directory under `bench/results/`, no engine changes.

No site

Links install, modelos, releases.