Pull requests / #1496

#1496 Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11)

open · @gbwzzy218 · 0 Kommentare · Auf GitHub

BenchmarksSetup & installNVIDIA / CUDADocumentationWindows

Beschreibung

## Title
Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11). Results only; no issue.

## Summary
Measurements of Flash-Next IQ2_XS on an 8 GB card with the release engine 0.1.40.3, comparing setup's 32K default with three 128K configurations. Main result: in setup's 128K k8v4 configuration, `--kv-resident 20480` instead of 32,768 raised median prompt throughput by 38.8–41.4% in three within-session comparisons (one in reversed order); decode did not change by a comparable amount. With 32,768 the automatic planner picks a 768-token chunk, below `stream_all_min()` (1,024), so prompts go through the 8-slot `STAGE` ring; with 20,480 it picks 1,024 tokens and a 384-slot ring. Two follow-up configurations with a fixed 768-token chunk are consistent with chunk size contributing most, but do not isolate the mechanisms.

- **Hardware:** RTX 3070 Ti 8 GB (PCIe 4.0 x16, also drives the display), Ryzen 9 5900X (AVX2, no AVX-512), 64 GB DDR4-3600, NVMe, Windows 11 build 26200, driver 617.42.
- **Software:** ready-made release engine 0.1.40.3 (SHA-256 in `system.txt`), tree at `d5ea713374` (identical to `main` when measured).
- **Model:** Flash-Next IQ2_XS at setup's pinned revision, MTP on, greedy, `reasoning_effort "none"`.
- **Method:** the unmodified community `benchmark.py`; 10 server starts in two sessions; 108 measured requests (5 per length, 3 at 128,000 tokens), all with zero reused prompt tokens; `needle_bench.py` 6/6 at 32k and 128k.

## What changed
Adds `bench/results/2026-10-07-community-rtx3070ti-8gb/` only (28 files, no code changes):
- `README.md`: the report (hardware and provenance, configurations, method, results, recall, limitations).
- `results.json`, `summary.json`: every measured request with the engine's `/metrics` counters, and per-block medians and ranges.
- `engine-start.txt`, `configs/`: what the engine planned at each start, and each block's effective arguments and `env`.
- `needles.json`, `system.txt`, `orchestrator-session1.log`, `orchestrator-session2.log`, `bench_orchestrator.py`: recall results, hardware probes and hashes, execution timestamps, and the driver script (local paths replaced by placeholders).
- `quality/`: single-run supplementary checks (task prompts, scorer and its outputs; a conversation-cache check).

## Extra Notes
- Limitations: one machine and model; configurations measured as blocks rather than interleaved; a 3.2–4.2% unexplained difference between the two sessions (comparisons are made within a session); the display shares the GPU; the release binary does not log which staging path a chunk took.
- Two suggestions for the maintainers are at the end of the README: a start-log hint when `--prefill auto` picks a chunk below the streaming threshold, and setup's `KV cache: 4-bit (Hadamard-rotated)` label printed for `--kv k8v4`.
- The report was drafted with an AI assistant, and its numbers were checked against the raw data by two further AI reviews before submission.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Mehr auf der Site

Links zu Install, Modellen, Releases.