Pull requests / #1496
#1496 Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11)
open · @gbwzzy218 · 0 commentaires · Sur GitHub
BenchmarksSetup & installNVIDIA / CUDADocumentationWindows
Description
## Title Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11). Results only; no issue. ## Summary Measurements of Flash-Next IQ2_XS on an 8 GB card with the release engine 0.1.40.3, comparing setup's 32K default with three 128K configurations. Main result: in setup's 128K k8v4 configuration, `--kv-resident 20480` instead of 32,768 raised median prompt throughput by 38.8–41.4% in three within-session comparisons (one in reversed order); decode did not change by a comparable amount. With 32,768 the automatic planner picks a 768-token chunk, below `stream_all_min()` (1,024), so prompts go through the 8-slot `STAGE` ring; with 20,480 it picks 1,024 tokens and a 384-slot ring. Two follow-up configurations with a fixed 768-token chunk are consistent with chunk size contributing most, but do not isolate the mechanisms. - **Hardware:** RTX 3070 Ti 8 GB (PCIe 4.0 x16, also drives the display), Ryzen 9 5900X (AVX2, no AVX-512), 64 GB DDR4-3600, NVMe, Windows 11 build 26200, driver 617.42. - **Software:** ready-made release engine 0.1.40.3 (SHA-256 in `system.txt`), tree at `d5ea713374` (identical to `main` when measured). - **Model:** Flash-Next IQ2_XS at setup's pinned revision, MTP on, greedy, `reasoning_effort "none"`. - **Method:** the unmodified community `benchmark.py`; 10 server starts in two sessions; 108 measured requests (5 per length, 3 at 128,000 tokens), all with zero reused prompt tokens; `needle_bench.py` 6/6 at 32k and 128k. ## What changed Adds `bench/results/2026-10-07-community-rtx3070ti-8gb/` only (28 files, no code changes): - `README.md`: the report (hardware and provenance, configurations, method, results, recall, limitations). - `results.json`, `summary.json`: every measured request with the engine's `/metrics` counters, and per-block medians and ranges. - `engine-start.txt`, `configs/`: what the engine planned at each start, and each block's effective arguments and `env`. - `needles.json`, `system.txt`, `orchestrator-session1.log`, `orchestrator-session2.log`, `bench_orchestrator.py`: recall results, hardware probes and hashes, execution timestamps, and the driver script (local paths replaced by placeholders). - `quality/`: single-run supplementary checks (task prompts, scorer and its outputs; a conversation-cache check). ## Extra Notes - Limitations: one machine and model; configurations measured as blocks rather than interleaved; a 3.2–4.2% unexplained difference between the two sessions (comparisons are made within a session); the display shares the GPU; the release binary does not log which staging path a chunk took. - Two suggestions for the maintainers are at the end of the README: a start-log hint when `--prefill auto` picks a chunk below the streaming threshold, and setup's `KV cache: 4-bit (Hadamard-rotated)` label printed for `--kv k8v4`. - The report was drafted with an AI assistant, and its numbers were checked against the raw data by two further AI reviews before submission. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sur le site
Liens install, modèles, releases.