Pull requests / #1460

#1460 Community benchmark: RTX 5080 in Docker Desktop on Windows (WSL2), Coder IQ1_M, 128K and 64K

open · @xWinIcex · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

描述

## What this is

A results-only community benchmark from an environment that had no report yet: the repository's **own Docker image on Docker Desktop for Windows (WSL2)**, built and launched exactly as `Dockerfile` and `docker-entrypoint.sh` document, with no hand-written launch command and nothing tuned by hand.

## Hardware and software

- NVIDIA RTX 5080, 16,303 MiB VRAM, driver 617.14. PCIe link **Gen 5 x8** while the card reports a max width of x16; the engine's startup probe measured **28.9 GB/s host-to-device** (`pcie_frac 0.55`). Power limit 360 W, clocks not fixed.
- AMD Ryzen 9 9950X (16C/32T, AVX-512), 64 GB RAM (the container sees 54.9 GiB), NVMe SSD.
- Windows 10 Pro 26H2 build 26300, WSL 2.7.10.0 (kernel 6.18.33.2-2), Docker Desktop 4.84.0.
- Commit `d5ea7133741e67743c0e886bb426c0ce8d69cf6c`, engine **0.1.40.3**, image built with `--build-arg CUDA_ARCHITECTURES=120`, base `nvidia/cuda:13.0.0-devel-ubuntu24.04`. CUDA 13.0.48, GCC 13.3.0, Python 3.12.3.

## Model

Coder **IQ1_M** from `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF` at the installer's pinned revision `5348543e0147355ac9cbcb031184a3546350988e`. All three files (both shards and `mmproj`) were SHA-256-verified against the repository's published LFS oids. MTP came from `Qwen/Qwen3.8-Flash-Next` via the installer (manifest with per-tensor hashes included).

## Configurations tested

Two context limits on the same loaded model and volume, only `CONTEXT`/`REINSTALL` differing:

| Limit | Prompt lengths | Runs |
| --- | --- | --- |
| 131,072 | 4,096 / 8,192 / 16,384 / 32,768 / 65,536 / 130,560 | 3 each |
| 65,536 | 4,096 / 8,192 / 16,384 / 32,768 / 65,024 | 3 each |

Greedy, `reasoning_effort: none`, 256-token output cap, one warm-up excluded. All 33 measured requests read their whole prompt: **zero reused tokens**. Method is the same script as `bench/results/2026-09-30-community-rtx-5090/benchmark.py`, changed only in four defaults.

Headline (median decode): **80.8 tok/s at 32,768 prompt tokens** and **72.2 tok/s at 130,560** with the 128K limit; prefill 2,193-3,304 tok/s.

## Three things this environment teaches

1. **Docker Desktop's default WSL2 memory cap silently selects the low-RAM mode.** With no `.wslconfig`, the container saw ~31 GB of the host's 64 GB, and because `setup.py` reads RAM from `/proc/meminfo` it did not error - it chose the low-RAM mode for the Coder. `[wsl2] memory=56GB` (included in the report) gives 54.9 GiB and the normal mode; both measured configurations used it.
2. **A request needs 8 tokens of headroom**: `prompt + max_tokens + 8 <= max_context`, and requests are never truncated. A 65,280-token prompt with `max_tokens: 256` on the 65,536 limit was refused with HTTP 400 ("at most 248 here"). `probe-request.py` reproduces it.
3. **At a 128K context on a 16 GB card the engine warns about its own configuration**: `152 MiB of VRAM free with everything loaded - LOW: requests may stall; add --vram-reserve-mib 1060 ... (or lower --max-context)`. Nothing failed in the measured runs, but the margin is real.

## One observation we could not explain

The 128K configuration measured equal or slightly **faster** than the 64K one at every shared prompt length (decode 80.8 vs 71.5 tok/s at 32,768; ranges 75.6-82.4 vs 70.3-73.6, so not obviously noise), even though it has the *smaller* expert cache (3,554 vs 4,060 experts, 6.78 vs 7.73 GiB) and only 152 MiB of free VRAM. With one session per configuration and no interleaving we did not isolate the cause, so it is reported as an observation to reproduce, not as a conclusion.

## Correctness

`tools/needle_bench.py` found **9 of 9** needles on the 128K configuration (depths 10/50/90% at 32k, 64k, 128k) and **6 of 6** on the 64K one (128k skipped for context, recorded in the output). A vision check with the encoder on the CPU read a drawn code word correctly.

## Limitations

One machine, one quantization, one GPU, one synthetic workload; long output, sampled decoding, thinking, concurrency, tool use and a thermal soak were not evaluated. The 128K-vs-64K comparison is two sessions, not interleaved settings. The PCIe link ran at x8 of a possible x16. Windows and WSL2 sit between Strata and the hardware. Full list in the report's README.

Results only - no engine changes.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。