Pull requests / #1546

#1546 Community benchmark: RTX 3090 24 GB + 128 GB RAM

open · @butonic · 0 comments · View on GitHub

BenchmarksNVIDIA / CUDAModels & quantsLinux

Description

Adds a community benchmark report for one RTX 3090 24 GB + 128 GB RAM, Ryzen 9 5900X, Linux, Strata 0.1.40.3 source build.

| Configuration | Engine | Context | Decode tok/s at 4K / 32K / 128K | Needles |
|---|---|---:|---:|---:|
| Flash-Next IQ3_S | Strata | 262,144 | 88.0 / 92.7 / 90.4 | 6/6 |
| Flash-Next IQ3_S, 10 GiB reserved for Flux | Strata elastic | 262,144 | 50.5 / 48.7 / 51.9 | 6/6 |
| Unsloth UD-IQ4_XS | Strata | 262,144 | 53.1 / 54.3 / 56.9 | 6/6 |
| Unsloth UD-Q4_K_XL | Strata, experimental | 262,144 | 38.4 / 36.5 / 37.5 | 6/6 |
| Qwen3.8-27B UD-IQ4_XS | llama.cpp, non-Strata | 262,144 | 75.0 / 61.7 / 41.1 | 6/6 |

Method: loopback server, zero-foreign guard, three runs at 4,096 / 32,768 / 128,000 prompt tokens, 256-token output cap, and 32k/128k needle recall at depths 10/50/90.

Limits: synthetic workload, no answer-quality benchmark, no concurrent Flux generation, Unsloth packs use `--compat-bf16`, UD-Q4_K_XL is experimental, and the llama.cpp check uses q4_0 KV.

Let me know if I should add a run with different parameters or env vars. I am pretty happy with Flash-Next IQ3_S vs qwen3.8 27B. Nice Work! 🤗 

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.