Pull requests / #469
#469 bench: community results, RTX 2080 Ti 11 GB + Threadripper 3960X, IQ3_S at 262K with KV streaming
closed · @homeofe · 0 commentaires · Sur GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Description
Results-only PR, following `docs/COMMUNITY_BENCHMARKS.md`. It covers a configuration the reports so far don't: **an 11 GB Turing card (RTX 2080 Ti, sm_75) running IQ3_S at the full 262,144-token context**, with KV streaming and the image encoder on the CPU. It also fills in a gap from #389, which says "vision not measured (the ready-made encoder has no sm_75 code)". This encoder is a source build that runs on this PC. The PR changes only `bench/results/2026-10-02-community-rtx2080ti-11gb/` and adds one bullet to the "Community reports" list in `docs/COMMUNITY_BENCHMARKS.md`. It doesn't touch the engine or setup. ## Results (median [range] of 3 runs, 400 generated tokens each) | run | prompt tokens (reused) | prompt read tok/s | decode tok/s | time to first token | drafts accepted | |---|---|---|---|---|---| | fresh 4K | 4,096 (0) | 496.9 [483.7–498.6] | 39.6 [39.3–41.2] | 8.3 s | 65.6% | | fresh 32K | 32,768 (0) | 636.9 [635.8–638.5] | 38.0 [37.7–40.0] | 51.6 s | 63.7% | | fresh 128K | 128,000 (0) | 566.2 [566.0–566.9] | 34.1 [33.7–34.3] | 226.6 s | 55.7% | | fresh 250K (**1 run**) | 250,000 (0) | 466.6 | 31.8 | 536.8 s | 53.4% | | follow-up at 128K depth | 128,463 (128,400) | 63 new tokens in 1.28 s | 35.0 | 1.85 s | 59.5% | | follow-up at 250K depth | 250,463 (249,993) | 470 new tokens in 3.56 s | 34.6 | 4.55 s | 59.3% | - **Recall:** `tools/needle_bench.py` at 32k and 128k (depths 10/50/90) found 6 of 6. Code words at 40% depth in the 128K and 250K prompts were found by the follow-up turns. - **Images:** 3 of 3 answered exactly. The CPU encoder takes about 1.6 s per new 768×512 image, and the answer starts after about 5 s. - **Memory:** 10,525 of 11,264 MiB VRAM, constant across all runs. Engine RSS was 52.3–52.9 GiB, and the decode expert cache hit rate was 50–57%. KV block reads that hit VRAM fell from 99.6% at 4K to 94.3% at 250K. ## A finding for 11–12 GB cards: keep the image encoder off the GPU These are earlier measurements from the same day on the same PC, with log excerpts in `earlier-encoder-comparison.log`. With `strata-vision` resident on the GPU (1.26 GiB), `--prefill auto` dropped to 512-token prompt chunks (1,024 with KV streaming). Fresh 26–33K prompts then read at **~243 tok/s**. With the encoder on the CPU (`"gpu": false`, 24 threads) the chunks went back up to 3,072–4,096 tokens, and the same prompts read at **750–775 tok/s** at a 128K context, or 622–634 tok/s at 262K. Image answers still came in 2.6–10 s. Setup currently offers the GPU encoder on such cards. That may be worth a note in the docs, or a recommendation in setup for cards under 12 GB. ## Method and caveats (details in the README) - **Shared server:** everything was measured through this PC's own Strata server, which also serves an agent. It wasn't restarted for this report. Each run checked `/status` first, and afterwards the `/metrics` request counter and the engine log were checked for overlap; none overlapped. Every measured prompt starts with a unique prefix, so no fresh run reused a cache. - **Server version:** `main` at `aeb35be` (v0.1.33 plus a HIP-only commit), plus the server-side change from #436 (cancel on client disconnect). That change doesn't touch the engine or requests whose client stays connected. Engine 0.1.33 was built from source with CUDA 12.0, GCC 13.3, sm_75. - **Odd things, recorded but not investigated:** - The 8 GiB swap was full before and during the runs, but vmstat showed no paging during prompt reads. - The 250K follow-up reused 249,993 tokens and re-read the previous 400-token answer, where the 128K follow-up reused all of it. - One 128k needle case reused 55,296 tokens of its prompt, so its timing isn't a fresh read. - **Privacy:** the config is included with the API key removed and paths shown as `<workspace>`. - **GPU time:** about 45 minutes.
Sur le site
Liens install, modèles, releases.