Pull requests / #1309

#1309 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)

open · @30crows · 0 comments · View on GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quants

Description

Community speed and recall results from a notebook: the original Flash-Next **IQ3_S** on an RTX 4090 Laptop GPU (16 GB) with 64 GB of RAM, at 131,072- and 262,144-token context limits. Results only, no engine changes.

This replaces #757, which GitHub closed when `main` was force-pushed; the commit is cherry-picked onto the new `main`.

**Hardware and software:** Lenovo Legion 82WQ: RTX 4090 Laptop GPU (16 GB, 175 W limit, Dynamic Boost on), Core i9-13900HX, 64 GB RAM (61.28 GiB usable), WD SN850X NVMe. Ubuntu 26.04.1, kernel 7.0.0, driver 610.57.04. Strata at `09817da` (measured at `99f3dbd`, its counterpart before the history rewrite; identical file tree), engine 0.1.38 **built from source** with CUDA 13.4 / GCC 15.2, because setup got HTTP 404 for the ready-made engine.

**Method:** the unchanged `benchmark.py` / `monitor.py` from the RTX 5090 report and `tools/needle_bench.py`. INT8 KV, reasoning off, temperature 0, 256-token output cap, three fresh-prompt runs per length (0 reused tokens).

| Context limit | Prompt tokens | Prompt tok/s | Decode tok/s | TTFT s |
| --- | ---: | ---: | ---: | ---: |
| 131,072 | 4,096 | 1,345 | 51.2 | 3.1 |
| 131,072 | 32,768 | 2,506 | 49.9 | 13.2 |
| 131,072 | 128,000 | 2,386 | 48.1 | 53.8 |
| 262,144 | 4,096 | 1,343 | 43.4 | 3.1 |
| 262,144 | 32,768 | 2,458 | 43.3 | 13.4 |
| 262,144 | 128,000 | 2,292 | 39.7 | 56.0 |
| 262,144 | 250,000 | 2,116 | 40.3 | 118.5 |

(medians; the ranges are in the report)

- **Recall:** 15 of 15 needles found (32K and 128K at both context limits, plus 3 at about 249K tokens). The cases that reused a 16,384-token prefix are marked.
- **Memory:** MemAvailable never fell below 6.17 GiB, swap did not grow, and VRAM peaked at 15.7 of 16.0 GB.
- **Power:** prompt processing ran at the 175 W limit (median 170–172 W, up to 89 °C). A separate short check decoded at 116 W median, so decode did not reach the power limit.
- **262,144 vs 131,072 tokens:** with KV streaming off (setup's choice here), the larger INT8 KV cache shrinks the expert cache from 3,865 to 2,856 slots, and decode is 13–17% slower at every prompt length. Prompt throughput stays within 4%.

**Limitations:** one machine, one quantization, a synthetic workload, and the laptop's `performance` power mode only (the `max-power` and `custom` modes were not tried). No sustained thermal run.

The `engine.log` files are force-added despite the global `*.log` ignore rule, as in the RTX 5090 report.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.