Pull requests / #1291
#1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K context
open · @NRK-SH · 0 commentaires · Sur GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Description
This PR adds a community benchmark report under `bench/results/`, following `docs/COMMUNITY_BENCHMARKS.md`. ### What was measured Strata 0.1.34 (source build with a local Ivy Bridge compatibility port) serving the GSQ-RCO IQ3_S quantization of Qwen3.8-Flash-Next - a 125B-parameter MoE model with 6B active parameters per token - on an HP Z820 workstation: 2x Intel Xeon E5-2697 v2 (AVX only; no AVX2/FMA), 256 GB DDR3, one RTX 3090 24 GB, Windows 11 Pro for Workstations. Hybrid inference: GPU expert cache (7,237 slots / 13.73 GiB), CPU fallback (23 workers), INT8 KV streaming, mmap-based n-gram / PLE tables, MTP speculative decoding. **Keywords:** Qwen3.8-Flash-Next, GSQ-RCO, IQ3_S, Strata, MoE, 125B, 512K context, AVX-only, Xeon E5-2697 v2, HP Z820, RTX 3090, hybrid inference, expert cache, CPU fallback, MTP, GGUF, long-context, vision, tool calling, Windows 11. ### Model `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, revision `ed59f920...`, IQ3_S (54,817,524,224 + 28,800,138,432 bytes; both SHA-256 hashes match the repository's published LFS hashes). Local pack prepared with `compat_bf16: false`; BF16 vision projector; MTP draft pack. ### Method Three to four fresh-prompt runs at approximately 4K / 32K / 128K prompt tokens and two runs at approximately 508K, in increasing-length order on the same loaded engine. Greedy decoding, thinking disabled, 320-token output cap, and a unique nonce per request to prevent prefix reuse (`cache_n = 0` for every run). Engine timing lines for all runs are attached. Long-context recall was measured with the unmodified `tools/needle_bench.py` at 32k/128k, depths 10/50/90. ### Results (medians) | Size | Actual prompt tokens | Prompt tok/s | Decode tok/s | TTFT (s) | | --- | ---: | ---: | ---: | ---: | | ~4K | 4,139 | 995.4 | 87.4 | 4.22 | | ~32K | 32,646 | 1,659.8 | 85.3 | 19.85 | | ~128K | 130,777 | 1,648.0 | 74.9 | 80.08 | | ~508K | 508,255 | 1,071.9 | 59.0 | 477.1 | Recall: 6/6. Additional checks: vision code-word question, tool-call round trip, and a small objective task set (8/9). ### Attachments `README.md`, `benchmark.py`, `results.json`, `results.csv`, `summary.json`, `engine-timings.log`, `needles.json`, `memory-summary.json`, `strata-iq3s-512k-vision.json`, `pack-conversions.json`, `provenance-local.json`, `hashes.sha256`, and `assets/` (charts and system overview). ### Limitations One machine, a custom pack and a local engine build; synthetic workload; single-sequence operation. The 508K runs processed 508,255 tokens, slightly below the full 524,288-token window. No general quality claim is made. The report is results-only and does not touch engine code. Happy to adjust the structure or add measurements if that would make the report more useful.
Sur le site
Liens install, modèles, releases.