Pull requests / #1309
#1309 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)
open · @30crows · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quants
Beschreibung
Community speed and recall results from a notebook: the original Flash-Next **IQ3_S** on an RTX 4090 Laptop GPU (16 GB) with 64 GB of RAM, at 131,072- and 262,144-token context limits. Results only, no engine changes. This replaces #757, which GitHub closed when `main` was force-pushed; the commit is cherry-picked onto the new `main`. **Hardware and software:** Lenovo Legion 82WQ: RTX 4090 Laptop GPU (16 GB, 175 W limit, Dynamic Boost on), Core i9-13900HX, 64 GB RAM (61.28 GiB usable), WD SN850X NVMe. Ubuntu 26.04.1, kernel 7.0.0, driver 610.57.04. Strata at `09817da` (measured at `99f3dbd`, its counterpart before the history rewrite; identical file tree), engine 0.1.38 **built from source** with CUDA 13.4 / GCC 15.2, because setup got HTTP 404 for the ready-made engine. **Method:** the unchanged `benchmark.py` / `monitor.py` from the RTX 5090 report and `tools/needle_bench.py`. INT8 KV, reasoning off, temperature 0, 256-token output cap, three fresh-prompt runs per length (0 reused tokens). | Context limit | Prompt tokens | Prompt tok/s | Decode tok/s | TTFT s | | --- | ---: | ---: | ---: | ---: | | 131,072 | 4,096 | 1,345 | 51.2 | 3.1 | | 131,072 | 32,768 | 2,506 | 49.9 | 13.2 | | 131,072 | 128,000 | 2,386 | 48.1 | 53.8 | | 262,144 | 4,096 | 1,343 | 43.4 | 3.1 | | 262,144 | 32,768 | 2,458 | 43.3 | 13.4 | | 262,144 | 128,000 | 2,292 | 39.7 | 56.0 | | 262,144 | 250,000 | 2,116 | 40.3 | 118.5 | (medians; the ranges are in the report) - **Recall:** 15 of 15 needles found (32K and 128K at both context limits, plus 3 at about 249K tokens). The cases that reused a 16,384-token prefix are marked. - **Memory:** MemAvailable never fell below 6.17 GiB, swap did not grow, and VRAM peaked at 15.7 of 16.0 GB. - **Power:** prompt processing ran at the 175 W limit (median 170–172 W, up to 89 °C). A separate short check decoded at 116 W median, so decode did not reach the power limit. - **262,144 vs 131,072 tokens:** with KV streaming off (setup's choice here), the larger INT8 KV cache shrinks the expert cache from 3,865 to 2,856 slots, and decode is 13–17% slower at every prompt length. Prompt throughput stays within 4%. **Limitations:** one machine, one quantization, a synthetic workload, and the laptop's `performance` power mode only (the `max-power` and `custom` modes were not tried). No sustained thermal run. The `engine.log` files are force-added despite the global `*.log` ignore rule, as in the RTX 5090 report. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Mehr auf der Site
Links zu Install, Modellen, Releases.