Pull requests / #1291
#1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K context
open · @NRK-SH · 0 comments · View on GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Description
This PR adds a community benchmark report under `bench/results/`, following `docs/COMMUNITY_BENCHMARKS.md`. ### What was measured Strata 0.1.34 (source build with a local Ivy Bridge compatibility port) serving the GSQ-RCO IQ3_S quantization of Qwen3.8-Flash-Next - a 125B-parameter MoE model with 6B active parameters per token - on an HP Z820 workstation: 2x Intel Xeon E5-2697 v2 (AVX only; no AVX2/FMA), 256 GB DDR3, one RTX 3090 24 GB, Windows 11 Pro for Workstations. Hybrid inference: GPU expert cache (7,237 slots / 13.73 GiB), CPU fallback (23 workers), INT8 KV streaming, mmap-based n-gram / PLE tables, MTP speculative decoding. **Keywords:** Qwen3.8-Flash-Next, GSQ-RCO, IQ3_S, Strata, MoE, 125B, 512K context, AVX-only, Xeon E5-2697 v2, HP Z820, RTX 3090, hybrid inference, expert cache, CPU fallback, MTP, GGUF, long-context, vision, tool calling, Windows 11. ### Model `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, revision `ed59f920...`, IQ3_S (54,817,524,224 + 28,800,138,432 bytes; both SHA-256 hashes match the repository's published LFS hashes). Local pack prepared with `compat_bf16: false`; BF16 vision projector; MTP draft pack. ### Method Three to four fresh-prompt runs at approximately 4K / 32K / 128K prompt tokens and two runs at approximately 508K, in increasing-length order on the same loaded engine. Greedy decoding, thinking disabled, 320-token output cap, and a unique nonce per request to prevent prefix reuse (`cache_n = 0` for every run). Engine timing lines for all runs are attached. Long-context recall was measured with the unmodified `tools/needle_bench.py` at 32k/128k, depths 10/50/90. ### Results (medians) | Size | Actual prompt tokens | Prompt tok/s | Decode tok/s | TTFT (s) | | --- | ---: | ---: | ---: | ---: | | ~4K | 4,139 | 995.4 | 87.4 | 4.22 | | ~32K | 32,646 | 1,659.8 | 85.3 | 19.85 | | ~128K | 130,777 | 1,648.0 | 74.9 | 80.08 | | ~508K | 508,255 | 1,071.9 | 59.0 | 477.1 | Recall: 6/6. Additional checks: vision code-word question, tool-call round trip, and a small objective task set (8/9). ### Attachments `README.md`, `benchmark.py`, `results.json`, `results.csv`, `summary.json`, `engine-timings.log`, `needles.json`, `memory-summary.json`, `strata-iq3s-512k-vision.json`, `pack-conversions.json`, `provenance-local.json`, `hashes.sha256`, and `assets/` (charts and system overview). ### Limitations One machine, a custom pack and a local engine build; synthetic workload; single-sequence operation. The 508K runs processed 508,255 tokens, slightly below the full 524,288-token window. No general quality claim is made. The report is results-only and does not touch engine code. Happy to adjust the structure or add measurements if that would make the report more useful.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.