Issues / #1512

#1512 Community benchmark: 2x RTX (RTX 5090 Laptop 24GB + RTX 4090 Laptop 16GB)

open · @akiry09 · 2 comentários · No GitHub

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindows

Descrição

Measured on 2026-10-08 by akiry09. Two-GPU tensor-split run of qwen3.8-flash-next-iq3_s via the local
Strata engine on :8080. Results report the six requested metrics: Prompt tokens,
Reused, Prompt tok/s, Decode tok/s, TTFT seconds, Total seconds. Per-run values and
per-config median (range) are shown.
TL;DR
- Two consumer laptop GPUs (RTX 5090 Laptop 24GB + RTX 4090 Laptop 16GB) run qwen3.8-flash-next-iq3_s
together via tensor split; speculative decoding (MTP draft) is active.
- Short prompt + 256-token generation: decode ~97 tok/s, TTFT ~0.86 s, total ~3.5 s.
- Long prompt (424 tokens) + 128-token generation: prefill ~321 tok/s, decode ~107 tok/s,
total ~2.5 s.
- No prefix-cache reuse observed (Reused = 0) thanks to distinct prompts per run.
Hardware and software
- GPU and VRAM; CPU; installed RAM; storage; PCIe link if known:
  - GPU 0: NVIDIA RTX 5090 Laptop GPU, 24,463 MiB VRAM (~23.9 GiB), PCIe Gen5 x16 (bus 0000:01:00.0)
  - GPU 1: NVIDIA RTX 4090 Laptop GPU, 16,376 MiB VRAM (~16.0 GiB), PCIe Gen4 x16 (bus 0000:09:00.0)
  - Both selected for a dual-GPU tensor-split run (user-confirmed); ~40.8 GiB combined visible
  - CPU: Intel Core Ultra 9 275HX; installed RAM: 96 GB; storage: model on F: (NVMe)
- OS; driver; CUDA or ROCm:
  - Host OS: Windows 11; inference via local Strata engine (OpenAI-compatible endpoint)
  - Driver: 616.92; CUDA toolkit bundled with the Strata engine (version not separately captured)
- Strata commit; engine version; release binary or source build:
  - Strata 0.1.40.3 (inferred from running strata.exe 0.1.40.3 process); engine commit not measured; release binary
- Background workloads and any power limits:
  - 5090 power.max_limit 175 W; 4090 power.max_limit 175 W but EC-firmware capped to ~95 W under load (known constraint, not re-measured here)
  - Other GPU apps present (Edge WebView, Strata UI) but negligible for compute
Model and configuration
- Model repository and revision; quantization; GGUF filenames:
  - Model id: qwen3.8-flash-next-iq3_s (Flash-Next family, Qwen3.8-based); quantization IQ3_S
  - Repository / revision / GGUF filename: not measured
- Vision encoder; custom packs or profiles:
  - Endpoint reports input_modalities text+image, but vision not exercised in this run
  - Custom packs / profiles: not measured
- Context; KV type and streaming; cache; prefill; low-RAM mode:
  - Context limit: 524,288 tokens (512K, as reported by server meta)
  - KV type / streaming window / prefill size / low-RAM: not measured
  - Prefix/expert cache: engine reuses conversation prefixes (observed); benchmark used distinct prompts to minimize reuse
- MTP; reasoning; sampling; calibration; experimental speed projection:
  - MTP / draft speculative decoding: active (reported by engine; acceptance observed ~70–77% in earlier runs)
  - Reasoning: enabled (model returns reasoning_content)
  - Sampling: temperature 0 for all benchmark runs
  - Calibration / experimental speed projection: not measured
Service: http://127.0.0.1:8080/v1 (OpenAI-compatible), API key "llama-local"
Model:   qwen3.8-flash-next-iq3_s
Launched via Strata 0.1.40.3 (strata.exe); exact engine launch command not separately captured.
Dual-GPU: tensor split across GPU 0 (RTX 5090 Laptop 24GB) and GPU 1 (RTX 4090 Laptop 16GB).

No site

Links install, modelos, releases.