Pull requests / #1199
#1199 Community benchmark: RTX 3090 eGPU + 64GB, IQ2_XS vs IQ3_XXS
closed · @lucapug · 0 コメント · GitHub で見る
BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows
本文
Community benchmark submission per docs/COMMUNITY_BENCHMARKS.md. Hardware: single RTX 3090 (24 GB) in a Thunderbolt 3 eGPU enclosure, i7-11800H, 64 GB RAM, NVMe, Windows 11 Pro 26H2, NVIDIA driver 617.14. Strata commit 1678de3, prebuilt engine 0.1.34. Model: Qwen3.8-Flash-Next GSQ-RCO, two quants compared on the same machine, same day, same settings (only the quantization changes): IQ2_XS and IQ3_XXS. Context 262144 (native), KV streaming on, calibrated per model (--pcie-frac 0.00, --spec-min-p 0.70, 7 CPU workers), MTP draft on, temperature 0. Key numbers (3 runs per config, medians; all prompt tokens freshly processed — per-request nonce defeats prefix cache): | | IQ2_XS | IQ3_XXS | |-------------------------------|---------|---------| | Prompt tok/s @ ~27K fresh | 618 | 468 | | Decode tok/s (short gens) | 101-112 | 60-81 | | TTFT s (reasoning 1st token) | 2.17 | 3.59 | | Needle @ ~25.7K fresh | 3/3 | 3/3 | Limitations: no dedicated warm-up pass (run 1 starts ~2 min after model load; ranges include it); decode measured on short generations via server-side per-request timings; single machine, single day. A 3090 + 3070-Laptop layer split was also measured on this machine and rejected (decode ~66 tok/s, prefill ~128-149 tok/s, context capped at 32K), consistent with MULTI_GPU.md guidance about much slower extra cards. Raw per-run JSON, the bench script (API key via env var only, no machine-specific paths) and the full README are in bench/results/2026-10-06-community-rtx3090-egpu-64gb/.
関連リンク
インストール・モデル・リリースへの站内リンク。