Issues / #1512
#1512 Community benchmark: 2x RTX (RTX 5090 Laptop 24GB + RTX 4090 Laptop 16GB)
open · @akiry09 · 2 コメント · GitHub で見る
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindows
本文
Measured on 2026-10-08 by akiry09. Two-GPU tensor-split run of qwen3.8-flash-next-iq3_s via the local Strata engine on :8080. Results report the six requested metrics: Prompt tokens, Reused, Prompt tok/s, Decode tok/s, TTFT seconds, Total seconds. Per-run values and per-config median (range) are shown. TL;DR - Two consumer laptop GPUs (RTX 5090 Laptop 24GB + RTX 4090 Laptop 16GB) run qwen3.8-flash-next-iq3_s together via tensor split; speculative decoding (MTP draft) is active. - Short prompt + 256-token generation: decode ~97 tok/s, TTFT ~0.86 s, total ~3.5 s. - Long prompt (424 tokens) + 128-token generation: prefill ~321 tok/s, decode ~107 tok/s, total ~2.5 s. - No prefix-cache reuse observed (Reused = 0) thanks to distinct prompts per run. Hardware and software - GPU and VRAM; CPU; installed RAM; storage; PCIe link if known: - GPU 0: NVIDIA RTX 5090 Laptop GPU, 24,463 MiB VRAM (~23.9 GiB), PCIe Gen5 x16 (bus 0000:01:00.0) - GPU 1: NVIDIA RTX 4090 Laptop GPU, 16,376 MiB VRAM (~16.0 GiB), PCIe Gen4 x16 (bus 0000:09:00.0) - Both selected for a dual-GPU tensor-split run (user-confirmed); ~40.8 GiB combined visible - CPU: Intel Core Ultra 9 275HX; installed RAM: 96 GB; storage: model on F: (NVMe) - OS; driver; CUDA or ROCm: - Host OS: Windows 11; inference via local Strata engine (OpenAI-compatible endpoint) - Driver: 616.92; CUDA toolkit bundled with the Strata engine (version not separately captured) - Strata commit; engine version; release binary or source build: - Strata 0.1.40.3 (inferred from running strata.exe 0.1.40.3 process); engine commit not measured; release binary - Background workloads and any power limits: - 5090 power.max_limit 175 W; 4090 power.max_limit 175 W but EC-firmware capped to ~95 W under load (known constraint, not re-measured here) - Other GPU apps present (Edge WebView, Strata UI) but negligible for compute Model and configuration - Model repository and revision; quantization; GGUF filenames: - Model id: qwen3.8-flash-next-iq3_s (Flash-Next family, Qwen3.8-based); quantization IQ3_S - Repository / revision / GGUF filename: not measured - Vision encoder; custom packs or profiles: - Endpoint reports input_modalities text+image, but vision not exercised in this run - Custom packs / profiles: not measured - Context; KV type and streaming; cache; prefill; low-RAM mode: - Context limit: 524,288 tokens (512K, as reported by server meta) - KV type / streaming window / prefill size / low-RAM: not measured - Prefix/expert cache: engine reuses conversation prefixes (observed); benchmark used distinct prompts to minimize reuse - MTP; reasoning; sampling; calibration; experimental speed projection: - MTP / draft speculative decoding: active (reported by engine; acceptance observed ~70–77% in earlier runs) - Reasoning: enabled (model returns reasoning_content) - Sampling: temperature 0 for all benchmark runs - Calibration / experimental speed projection: not measured Service: http://127.0.0.1:8080/v1 (OpenAI-compatible), API key "llama-local" Model: qwen3.8-flash-next-iq3_s Launched via Strata 0.1.40.3 (strata.exe); exact engine launch command not separately captured. Dual-GPU: tensor split across GPU 0 (RTX 5090 Laptop 24GB) and GPU 1 (RTX 4090 Laptop 16GB).
関連リンク
インストール・モデル・リリースへの站内リンク。