Issues / #922
#922 RTX 5070 Ti 16GB / Windows / IQ3_S: calibration ~14 tok/s and ~9–10 tok/s in controlled runs

closed · @annD-annD · 23 comentarios · En GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
Descripción
### Summary I am seeing much lower decode performance than expected with Qwen3.8-Flash-Next IQ3_S on a Windows system with an RTX 5070 Ti 16 GB. Strata itself calibrates this system at only about 13–14.4 tok/s. In controlled requests with ~2.85K uncached prompt tokens and 256 generated tokens, decode speed is consistently around 9–10.6 tok/s. This seems much lower than the published IQ3_S numbers for an RTX 5070 12 GB. I am not sure whether this is a configuration issue, a Windows/Blackwell-specific performance problem, or expected behavior with my setup. --- ### Hardware - GPU: NVIDIA GeForce RTX 5070 Ti - VRAM: 15.9 GB - Compute capability: 12.0 - NVIDIA driver: 591.74 - CPU: AMD Ryzen 9 9950X 16-Core Processor - AVX-512: available - RAM: 126 GB usable - OS: Windows - PCIe host -> device probe: ~47.3 GB/s --- ### Strata - Version: 0.1.39 - Engine: prebuilt Windows x64 CUDA 13.0 - GPU architecture included: sm_120 - Model: Qwen3.8-Flash-Next - Quantization: IQ3_S - Context: 131072 - KV cache: int8 - KV resident: 32768 - MTP: enabled - Spec depth: 4 - spec-min-p: 0.5 - Vision: enabled, encoder on CPU - Parallel requests: 1 The setup completed without errors. --- ### Calibration results Command: .\START-HERE.bat --calibrate Results: PCIe share 0.00: 13.0 tok/s PCIe share 0.20: 13.4 tok/s PCIe share 0.35: 13.4 tok/s PCIe share 0.55: 13.2 tok/s PCIe share 0.75: 12.5 tok/s draft floor 0.30: 13.7 tok/s draft floor 0.50: 14.4 tok/s draft floor 0.70: 13.7 tok/s 15 workers: 13.1 tok/s 10 workers: 13.7 tok/s 8 workers: 13.6 tok/s Strata selected: --pool-workers 10 --spec 4 --spec-min-p 0.5 Startup / memory layout Typical startup: experts loaded: 46.84 GiB at ~6.7 GiB/s filling the GPU's expert cache: 4222 experts 8.06 GiB of VRAM KV streaming: 32768 of 131072 cells per QSA layer in VRAM K/V in ~1.55 GiB pinned RAM MTP draft layer: 835 MiB VRAM VRAM free with everything loaded: ~401 MiB Relevant expert arena messages: expert arena: cudaHostRegister PORTABLE ok large pages refused for 50295996416 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314) using 4 KB pages The main engine detects the GPU correctly: GPU 0: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0 There is also an early: ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected but this appears to come from the CPU vision process. The main Strata engine detects and uses GPU 0 correctly. Controlled decode test I sent three fresh requests through the OpenAI-compatible API. Each request had: - ~2850 prompt tokens - cached_tokens = 0 - max_tokens = 256 - temperature = 0 Results: Run 1: prompt_tokens: 2852 completion_tokens: 256 elapsed: 31.7 s Run 2: prompt_tokens: 2850 completion_tokens: 256 elapsed: 33.3 s Run 3: prompt_tokens: 2853 completion_tokens: 256 elapsed: 34.7 s Strata log showed approximately: 9.2 tok/s 10.6 tok/s 10.3 tok/s Example: done: 256 tokens in 32 s (10.6 tok/s) expert cache 84.7% hit +10.8% of routed experts over PCIe Other runs showed: expert cache hit rate: ~80–85% additional routed experts over PCIe: ~10–14% Real chat runs I also see large variation during normal use. Examples: 794 tokens in 78 s = 13.8 tok/s expert cache 83.1% hit +11.6% over PCIe and: 583 tokens in 178 s = 3.3 tok/s expert cache 79.0% hit +14.8% over PCIe Another 64K-context test: 1305 tokens in 475 s = 2.9 tok/s expert cache 83.4% hit +11.3% over PCIe Reducing the context from 128K to 64K did not materially improve performance. I also tested a smaller resident KV area, which allowed a slightly larger expert cache: 4357 experts / 8.31 GiB VRAM but performance remained around: 2.8 tok/s with: 78.2% expert cache hit 15.1% routed over PCIe Why this looks unusual The published IQ3_S speed measurements for an RTX 5070 12 GB appear to be around: ~52 tok/s at short context ~53 tok/s around 4K ~45 tok/s around 128K My RTX 5070 Ti 16 GB system is consistently far below that, including during Strata's own calibration. Because the calibration itself reports only ~14 tok/s, this does not appear to be caused only by OpenWebUI or by a particular chat prompt. Questions 1. Is ~13–14 tok/s during calibration expected for IQ3_S on a 5070 Ti 16 GB? 2. Does the ~80–85% expert-cache hit rate explain such a large difference from the published RTX 5070 results? 3. Could the Windows pinned-memory / 4 KB page path cause this level of slowdown? 4. Is there a recommended benchmark command or configuration I should run to reproduce the official IQ3_S speed measurements exactly? 5. Are there any known sm_120 / RTX 5070 Ti performance issues in 0.1.39? I am happy to provide the full strata-iq3_s.log, config JSON, or additional benchmark results if useful.
En el sitio
Enlaces a install, modelos, releases.