Issues / #791
#791 Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer split
closed · @yy16432 · 2 commentaires · Sur GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quants
Description
# Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer split Measured on 2026-10-03 by a ZCode-assisted lab operator. Tests Qwen3.8-Flash-Next **IQ3_S** (GSQ-RCO) and the community **agentionai AP-Q4_K_M** GGUF on Strata 0.1.38, each on a single RTX 5070 Ti (calibrated) and on a 2× RTX 5060 Ti layer split (calibrated) — four configurations in total, same method throughout. Main limitations: 1–3 runs per point (decode of the native GSQ path proved stable — ±5 % across repeats — while the imported GGML-format AP quant varies ±20 %+ with MTP luck), and 262,144-token prompts are refused by design (prompt + output exceeds the 262,144 context capacity). ## Hardware and software - 1× RTX 5070 Ti 16 GB (comparison config: 2× RTX 5060 Ti 16 GB, layer split); Intel Core Ultra 7 265K (20 cores, **no AVX-512**, AVX2 expert kernels); 96 GB (2×48 GB DDR5); PCIe 3.0 NVMe for models and the PLE table; single card on the CPU-attached x16 slot - Ubuntu 24.04.4; driver 595.91.07; CUDA 13.0 - Strata commit `99f3dbd` (2026-10-03); engine 0.1.38, source build - Background workloads: none during measurements (single-user lab machine) ## Model and configuration - `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF` IQ3_S; community row: `agentionai/Qwen3.8-Flash-Next-AP-GGUF` AP-Q4_K_M, imported manually (`iq_pack.py --compat-bf16 --experts-bin` + a BF16 one-tensor `--embd-gguf` for its Q6_K token embedding + its own mmproj-F16 wired through the config's `vision` section) - Vision encoder on for all four configurations - Context 262,144; KV `int8` (VRAM-resident; `--kv-resident` off); expert cache `auto`; `--prefill auto` - MTP `--spec 4 --spec-min-p 0.70`; greedy sampling from the harness; **calibrated per configuration**: IQ3_S 5070 Ti `--pcie-frac 0.20`, IQ3_S split `0.00`, AP 5070 Ti `0.35`, AP split `0.00` (19 CPU workers each); speed projection off ```text # essence (IQ3_S, 5070 Ti; server 0.0.0.0:8080, gpu 0) --pack packs/iq3_s --native Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf --ple-gguf ...-00002-of-00002.gguf --expert-profile data/expert-profile.bin --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.70 --pcie-frac 0.20 --mtp <data>/mtp/rt --max-context 262144 --kv int8 --vision ``` ## Method Third-party OpenAI-compatible HTTP benchmark harness (GUI), one request at a time (the engine is single-request serial), greedy (temperature 0), output cap 256 tokens per request, harness timeout 600 s for the longest prompts. Prompts are generated by the harness per length bucket (content differs run to run within a bucket; re-running the same harness on a warm instance can inflate decode via prompt-lookup drafting — noted below where it applies). Service freshly (re)started and calibrated before each reported configuration; expert cache warm from one smoke request. Timing boundaries are the harness's TTFT/ITL fields, cross-checked against the engine's per-request `decode_tok_s` in `/metrics` (agrees within 2 %). Runs per point: IQ3_S split 2, IQ3_S 5070 Ti 1 (a second warm sweep landed within ±5 %), AP split 1 (after `--prefill auto`; the earlier 512-chunk sweep is excluded as a different configuration), AP 5070 Ti 3 (median reported; one of the three sweeps ran on a warm instance and skews high). RAM observed: ~57 GB resident experts + KV + runtime on IQ3_S; ~63 GB with AP-Q4_K_M. ## Results ### 1. IQ3_S on 2× RTX 5060 Ti (layer split, calibrated) | Actual prompt tokens | Reused tokens | Generated tokens | Runs | Prompt tok/s median and range | Decode tok/s median and range | TTFT seconds median and range | | ---: | ---: | ---: | ---: | --- | --- | --- | | 512 | 0 | 256 | 2 | 542 (452–631) | 73.0 (64.4–81.7) | 0.91 (0.87–0.94) | | 1024 | 0 | 256 | 2 | 798 (696–901) | 67.6 (61.7–73.5) | 1.11 (0.95–1.29) | | 2048 | 0 | 256 | 2 | 1083 (963–1204) | 69.4 (54.5–84.3) | 1.48 (1.33–1.78) | | 4096 | 0 | 256 | 2 | 1458 (1408–1508) | 70.4 (60.7–80.1) | 2.89 (2.80–2.99) | | 8192 | 0 | 256 | 2 | 1602 (1560–1644) | 66.5 (64.0–69.1) | 5.21 (5.06–5.37) | | 16384 | 0 | 256 | 2 | 2141 (2134–2149) | 65.2 (58.4–72.0) | 7.85 (7.72–7.98) | | 32768 | 0 | 256 | 2 | 2552 (2520–2585) | 65.2 (60.2–70.1) | 13.29 (13.14–13.48) | | 65536 | 0 | 256 | 2 | 2886 (2843–2929) | 68.5 (59.5–77.4) | 23.33 (23.23–23.42) | | 131072 | 0 | 256 | 2 | 2969 (2905–3032) | 62.7 (54.3–71.1) | 45.17 (44.80–45.54) | ### 2. IQ3_S on RTX 5070 Ti (single card, calibrated) | Actual prompt tokens | Reused tokens | Generated tokens | Runs | Prompt tok/s median and range | Decode tok/s median and range | TTFT seconds median and range | | ---: | ---: | ---: | ---: | --- | --- | --- | | 512 | 0 | 256 | 1 | 845 | 118.3 | 0.66 | | 1024 | 0 | 256 | 1 | 1325 | 100.2 | 0.83 | | 2048 | 0 | 256 | 1 | 1982 | 100.4 | 1.10 | | 4096 | 0 | 256 | 1 | 2822 | 100.3 | 1.52 | | 8192 | 0 | 256 | 1 | 3241 | 100.4 | 2.60 | | 16384 | 0 | 256 | 1 | 3373 | 99.4 | 4.97 | | 32768 | 0 | 256 | 1 | 3393 | 95.9 | 9.80 | | 65536 | 0 | 256 | 1 | 3333 | 99.5 | 19.9 | | 131072 | 0 | 256 | 1 | 3147 | 99.1 | 42.0 | **IQ3_S, 5070 Ti vs split:** decode +25–45 % at every length and essentially flat (95.9–100.4, ITL ≈ 10 ms) vs the split's 63–73 with ~10 tok/s droop toward 131K; prefill ~2× in the 2–8K range, +6 % at 131K; TTFT 45.2 → 42.0 s at 131K. The calibrated `--pcie-frac` flips 0.00 (split) → 0.20 (single). ### 3. AP-Q4_K_M on 2× RTX 5060 Ti (layer split, calibrated, `--prefill auto`) | Actual prompt tokens | Reused tokens | Generated tokens | Runs | Prompt tok/s median and range | Decode tok/s median and range | TTFT seconds median and range | | ---: | ---: | ---: | ---: | --- | --- | --- | | 512 | 0 | 256 | 1 | 452 | 64.4 | 1.21 | | 1024 | 0 | 256 | 1 | 696 | 61.7 | 1.54 | | 2048 | 0 | 256 | 1 | 963 | 54.5 | 2.19 | | 4096 | 0 | 256 | 1 | 1408 | 60.7 | 2.99 | | 8192 | 0 | 256 | 1 | 1560 | 64.0 | 5.34 | | 16384 | 0 | 256 | 1 | 2134 | 58.4 | 7.98 | | 32768 | 0 | 256 | 1 | 2585 | 60.2 | 13.44 | | 65536 | 0 | 256 | 1 | 2929 | 59.5 | 23.42 | | 131072 | 0 | 256 | 1 | 3032 | 54.3 | 43.57 | ### 4. AP-Q4_K_M on RTX 5070 Ti (single card, calibrated) | Actual prompt tokens | Reused tokens | Generated tokens | Runs | Prompt tok/s median and range | Decode tok/s median and range | TTFT seconds median and range | | ---: | ---: | ---: | ---: | --- | --- | --- | | 512 | 0 | 256 | 3 | 668 (636–679) | 92.1 (85.3–106.9) | 0.83 (0.81–0.86) | | 1024 | 0 | 256 | 3 | 1129 (1034–1139) | 79.0 (76.1–83.3) | 0.97 (0.95–1.04) | | 2048 | 0 | 256 | 3 | 1599 (1478–1606) | 77.1 (75.5–132.8) | 1.35 (1.33–1.44) | | 4096 | 0 | 256 | 3 | 2758 (2448–2788) | 76.7 (75.2–76.9) | 1.56 (1.53–1.73) | | 8192 | 0 | 256 | 3 | 2924 (2792–2917) | 114.4 (76.9–125.7) | 2.89 (2.89–3.01) | | 16384 | 0 | 256 | 3 | 2987 (2931–3013) | 77.8 (74.9–82.6) | 5.58 (5.53–5.68) | | 32768 | 0 | 256 | 3 | 2990 (2966–3016) | 84.0 (80.5–150.7) | 11.09 (10.99–11.17) | | 65536 | 0 | 256 | 3 | 2922 (2920–2932) | 74.1 (72.1–96.7) | 22.61 (22.55–22.62) | | 131072 | 0 | 256 | 3 | 2780 (2772–2795) | 76.9 (76.6–111.8) | 47.49 (47.31–47.63) | **AP-Q4_K_M, 5070 Ti vs split:** decode +26–43 % at every length (medians 74–114 vs 54–64); prefill +48–96 % below 16K, converging to ±0 % at 64K and −8 % at 131K (the only metric where the single card does not win). `--pcie-frac` 0.00 → 0.35 (the imported GGML experts lean on the CPU more, and the single fat card rewards a higher PCIe share). **Cross-model:** on identical hardware the AP quant decodes ~20 % below IQ3_S (100.4 vs 79–77 at mid lengths on the 5070 Ti) with ±20 %+ run-to-run variance vs the native path's ±5 %; per-request MTP acceptance 50–63 % on fresh prompts vs 87 % for IQ3_S. Prefill is format-insensitive within ±5 % except at very long lengths. Notes: 262,144-token prompts are refused by the server (prompt + 256 output exceeds the 262,144 capacity) — expected behavior, not an engine failure. The AP-Q4_K_M results are, to our knowledge, the first published measurements of that community quant on Strata. ## Correctness and limitations - Sanity checks: `/health` (context 262144, images on) before each sweep; one Chinese short-prompt generation per configuration (correct, coherent); OCR of a rendered test image correct on IQ3_S and on the imported AP quant (after wiring its own mmproj). - No needle/retrieval tests were run; quality is not assessed here, only speed. - The AP-Q4_K_M import required `--compat-bf16`, an `--embd-gguf` one-tensor BF16 file (its token embedding is Q6_K, which the native path cannot dequantize on the GPU) and an `experts.bin` repack for the full-RAM arena — described so others can reproduce. - Power limits: none applied; stock clocks. PCIe link idle-downgrades to Gen1 between requests (observed), irrelevant to the measurements.
Sur le site
Liens install, modèles, releases.