Pull requests / #1460
#1460 Community benchmark: RTX 5080 in Docker Desktop on Windows (WSL2), Coder IQ1_M, 128K and 64K
open · @xWinIcex · 0 comments · View on GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
## What this is
A results-only community benchmark from an environment that had no report yet: the repository's **own Docker image on Docker Desktop for Windows (WSL2)**, built and launched exactly as `Dockerfile` and `docker-entrypoint.sh` document, with no hand-written launch command and nothing tuned by hand.
## Hardware and software
- NVIDIA RTX 5080, 16,303 MiB VRAM, driver 617.14. PCIe link **Gen 5 x8** while the card reports a max width of x16; the engine's startup probe measured **28.9 GB/s host-to-device** (`pcie_frac 0.55`). Power limit 360 W, clocks not fixed.
- AMD Ryzen 9 9950X (16C/32T, AVX-512), 64 GB RAM (the container sees 54.9 GiB), NVMe SSD.
- Windows 10 Pro 26H2 build 26300, WSL 2.7.10.0 (kernel 6.18.33.2-2), Docker Desktop 4.84.0.
- Commit `d5ea7133741e67743c0e886bb426c0ce8d69cf6c`, engine **0.1.40.3**, image built with `--build-arg CUDA_ARCHITECTURES=120`, base `nvidia/cuda:13.0.0-devel-ubuntu24.04`. CUDA 13.0.48, GCC 13.3.0, Python 3.12.3.
## Model
Coder **IQ1_M** from `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF` at the installer's pinned revision `5348543e0147355ac9cbcb031184a3546350988e`. All three files (both shards and `mmproj`) were SHA-256-verified against the repository's published LFS oids. MTP came from `Qwen/Qwen3.8-Flash-Next` via the installer (manifest with per-tensor hashes included).
## Configurations tested
Two context limits on the same loaded model and volume, only `CONTEXT`/`REINSTALL` differing:
| Limit | Prompt lengths | Runs |
| --- | --- | --- |
| 131,072 | 4,096 / 8,192 / 16,384 / 32,768 / 65,536 / 130,560 | 3 each |
| 65,536 | 4,096 / 8,192 / 16,384 / 32,768 / 65,024 | 3 each |
Greedy, `reasoning_effort: none`, 256-token output cap, one warm-up excluded. All 33 measured requests read their whole prompt: **zero reused tokens**. Method is the same script as `bench/results/2026-09-30-community-rtx-5090/benchmark.py`, changed only in four defaults.
Headline (median decode): **80.8 tok/s at 32,768 prompt tokens** and **72.2 tok/s at 130,560** with the 128K limit; prefill 2,193-3,304 tok/s.
## Three things this environment teaches
1. **Docker Desktop's default WSL2 memory cap silently selects the low-RAM mode.** With no `.wslconfig`, the container saw ~31 GB of the host's 64 GB, and because `setup.py` reads RAM from `/proc/meminfo` it did not error - it chose the low-RAM mode for the Coder. `[wsl2] memory=56GB` (included in the report) gives 54.9 GiB and the normal mode; both measured configurations used it.
2. **A request needs 8 tokens of headroom**: `prompt + max_tokens + 8 <= max_context`, and requests are never truncated. A 65,280-token prompt with `max_tokens: 256` on the 65,536 limit was refused with HTTP 400 ("at most 248 here"). `probe-request.py` reproduces it.
3. **At a 128K context on a 16 GB card the engine warns about its own configuration**: `152 MiB of VRAM free with everything loaded - LOW: requests may stall; add --vram-reserve-mib 1060 ... (or lower --max-context)`. Nothing failed in the measured runs, but the margin is real.
## One observation we could not explain
The 128K configuration measured equal or slightly **faster** than the 64K one at every shared prompt length (decode 80.8 vs 71.5 tok/s at 32,768; ranges 75.6-82.4 vs 70.3-73.6, so not obviously noise), even though it has the *smaller* expert cache (3,554 vs 4,060 experts, 6.78 vs 7.73 GiB) and only 152 MiB of free VRAM. With one session per configuration and no interleaving we did not isolate the cause, so it is reported as an observation to reproduce, not as a conclusion.
## Correctness
`tools/needle_bench.py` found **9 of 9** needles on the 128K configuration (depths 10/50/90% at 32k, 64k, 128k) and **6 of 6** on the 64K one (128k skipped for context, recorded in the output). A vision check with the encoder on the CPU read a drawn code word correctly.
## Limitations
One machine, one quantization, one GPU, one synthetic workload; long output, sampled decoding, thinking, concurrency, tool use and a thermal soak were not evaluated. The 128K-vs-64K comparison is two sessions, not interleaved settings. The PCIe link ran at x8 of a possible x16. Windows and WSL2 sit between Strata and the hardware. Full list in the report's README.
Results only - no engine changes.Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.