Pull requests / #1199

#1199 Community benchmark: RTX 3090 eGPU + 64GB, IQ2_XS vs IQ3_XXS

closed · @lucapug · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

Descripción

Community benchmark submission per docs/COMMUNITY_BENCHMARKS.md.

Hardware: single RTX 3090 (24 GB) in a Thunderbolt 3 eGPU enclosure, i7-11800H, 64 GB RAM, NVMe, Windows 11 Pro 26H2, NVIDIA driver 617.14. Strata commit 1678de3, prebuilt engine 0.1.34.

Model: Qwen3.8-Flash-Next GSQ-RCO, two quants compared on the same machine, same day, same settings (only the quantization changes): IQ2_XS and IQ3_XXS. Context 262144 (native), KV streaming on, calibrated per model (--pcie-frac 0.00, --spec-min-p 0.70, 7 CPU workers), MTP draft on, temperature 0.

Key numbers (3 runs per config, medians; all prompt tokens freshly processed — per-request nonce defeats prefix cache):

|                                             | IQ2_XS  | IQ3_XXS |
|-------------------------------|---------|---------|
| Prompt tok/s @ ~27K fresh     | 618     | 468     |
| Decode tok/s (short gens)     | 101-112 | 60-81   |
| TTFT s (reasoning 1st token)  | 2.17    | 3.59    |
| Needle @ ~25.7K fresh         | 3/3     | 3/3     |

Limitations: no dedicated warm-up pass (run 1 starts ~2 min after model load; ranges include it); decode measured on short generations via server-side per-request timings; single machine, single day. A 3090 + 3070-Laptop layer split was also measured on this machine and rejected (decode ~66 tok/s, prefill ~128-149 tok/s, context capped at 32K), consistent with MULTI_GPU.md guidance about much slower extra cards.

Raw per-run JSON, the bench script (API key via env var only, no machine-specific paths) and the full README are in bench/results/2026-10-06-community-rtx3090-egpu-64gb/.

En el sitio

Enlaces a install, modelos, releases.