Issues / #922

#922 RTX 5070 Ti 16GB / Windows / IQ3_S: calibration ~14 tok/s and ~9–10 tok/s in controlled runs

RTX 5070 Ti 16GB / Windows / IQ3_S: calibration ~14 tok/s and ~9–10 tok/s in con

closed · @annD-annD · 23 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows

Description

### Summary

I am seeing much lower decode performance than expected with Qwen3.8-Flash-Next IQ3_S on a Windows system with an RTX 5070 Ti 16 GB.

Strata itself calibrates this system at only about 13–14.4 tok/s.

In controlled requests with ~2.85K uncached prompt tokens and 256 generated tokens, decode speed is consistently around 9–10.6 tok/s.

This seems much lower than the published IQ3_S numbers for an RTX 5070 12 GB.

I am not sure whether this is a configuration issue, a Windows/Blackwell-specific performance problem, or expected behavior with my setup.

---

### Hardware

- GPU: NVIDIA GeForce RTX 5070 Ti
- VRAM: 15.9 GB
- Compute capability: 12.0
- NVIDIA driver: 591.74
- CPU: AMD Ryzen 9 9950X 16-Core Processor
- AVX-512: available
- RAM: 126 GB usable
- OS: Windows
- PCIe host -> device probe: ~47.3 GB/s

---

### Strata

- Version: 0.1.39
- Engine: prebuilt Windows x64 CUDA 13.0
- GPU architecture included: sm_120
- Model: Qwen3.8-Flash-Next
- Quantization: IQ3_S
- Context: 131072
- KV cache: int8
- KV resident: 32768
- MTP: enabled
- Spec depth: 4
- spec-min-p: 0.5
- Vision: enabled, encoder on CPU
- Parallel requests: 1

The setup completed without errors.

---

### Calibration results

Command:
.\START-HERE.bat --calibrate

Results:
PCIe share 0.00: 13.0 tok/s
PCIe share 0.20: 13.4 tok/s
PCIe share 0.35: 13.4 tok/s
PCIe share 0.55: 13.2 tok/s
PCIe share 0.75: 12.5 tok/s

draft floor 0.30: 13.7 tok/s
draft floor 0.50: 14.4 tok/s
draft floor 0.70: 13.7 tok/s

15 workers: 13.1 tok/s
10 workers: 13.7 tok/s
8 workers: 13.6 tok/s

Strata selected:
--pool-workers 10
--spec 4
--spec-min-p 0.5

Startup / memory layout
Typical startup:
experts loaded: 46.84 GiB at ~6.7 GiB/s

filling the GPU's expert cache:
4222 experts
8.06 GiB of VRAM

KV streaming:
32768 of 131072 cells per QSA layer in VRAM
K/V in ~1.55 GiB pinned RAM

MTP draft layer:
835 MiB VRAM

VRAM free with everything loaded:
~401 MiB

Relevant expert arena messages:
expert arena: cudaHostRegister PORTABLE ok

large pages refused for 50295996416 B
(GetLargePageMinimum=2097152, VirtualAlloc error 1314)
using 4 KB pages

The main engine detects the GPU correctly:
GPU 0: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0

There is also an early:
ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected

but this appears to come from the CPU vision process. The main Strata engine detects and uses GPU 0 correctly.
Controlled decode test
I sent three fresh requests through the OpenAI-compatible API.
Each request had:
- ~2850 prompt tokens
- cached_tokens = 0
- max_tokens = 256
- temperature = 0
Results:
Run 1:
prompt_tokens: 2852
completion_tokens: 256
elapsed: 31.7 s

Run 2:
prompt_tokens: 2850
completion_tokens: 256
elapsed: 33.3 s

Run 3:
prompt_tokens: 2853
completion_tokens: 256
elapsed: 34.7 s

Strata log showed approximately:
9.2 tok/s
10.6 tok/s
10.3 tok/s

Example:
done: 256 tokens in 32 s (10.6 tok/s)
expert cache 84.7% hit
+10.8% of routed experts over PCIe

Other runs showed:
expert cache hit rate: ~80–85%
additional routed experts over PCIe: ~10–14%

Real chat runs
I also see large variation during normal use.
Examples:
794 tokens in 78 s = 13.8 tok/s
expert cache 83.1% hit
+11.6% over PCIe

and:
583 tokens in 178 s = 3.3 tok/s
expert cache 79.0% hit
+14.8% over PCIe

Another 64K-context test:
1305 tokens in 475 s = 2.9 tok/s
expert cache 83.4% hit
+11.3% over PCIe

Reducing the context from 128K to 64K did not materially improve performance.
I also tested a smaller resident KV area, which allowed a slightly larger expert cache:
4357 experts / 8.31 GiB VRAM

but performance remained around:
2.8 tok/s

with:
78.2% expert cache hit
15.1% routed over PCIe

Why this looks unusual
The published IQ3_S speed measurements for an RTX 5070 12 GB appear to be around:
~52 tok/s at short context
~53 tok/s around 4K
~45 tok/s around 128K

My RTX 5070 Ti 16 GB system is consistently far below that, including during Strata's own calibration.
Because the calibration itself reports only ~14 tok/s, this does not appear to be caused only by OpenWebUI or by a particular chat prompt.
Questions
1. Is ~13–14 tok/s during calibration expected for IQ3_S on a 5070 Ti 16 GB?
2. Does the ~80–85% expert-cache hit rate explain such a large difference from the published RTX 5070 results?
3. Could the Windows pinned-memory / 4 KB page path cause this level of slowdown?
4. Is there a recommended benchmark command or configuration I should run to reproduce the official IQ3_S speed measurements exactly?
5. Are there any known sm_120 / RTX 5070 Ti performance issues in 0.1.39?
I am happy to provide the full strata-iq3_s.log, config JSON, or additional benchmark results if useful.

Sur le site

Liens install, modèles, releases.