Pull requests / #674

#674 bench: community report, RTX 3090 Ti + 2x Xeon E5-2699 v3 (HP Z840), IQ3_S at 262K

closed · @pertain99 · 0 commentaires · Sur GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

Description

Results-only community benchmark report per `docs/COMMUNITY_BENCHMARKS.md`, in `bench/results/2026-10-03-community-z840-rtx3090ti/` (README, `results.json`, the NUMA A/B script), plus one line in the Community reports list. No engine changes.

## Setup

- Strata **0.1.38**, source build (`CMAKE_CUDA_ARCHITECTURES=86`, CUDA 12.3 nvcc, gcc 12.3), Ubuntu 22.04.5 / kernel 6.8, driver 610.57.04 (open modules).
- **RTX 3090 Ti 24 GB on PCIe 3.0 x16** (engine probe 12.4 GB/s), power limit 320 W; GNOME desktop on the same card.
- **2x Xeon E5-2699 v3** (Haswell-EP, AVX2 only, no AVX-512), 384 GB DDR4-2133, two NUMA nodes; models on an Intel DC P4608 NVMe. Expert arena in 2 MB hugetlb pages, `LimitMEMLOCK=infinity`, `numa_balancing=0`.
- Original Flash-Next GSQ-RCO **IQ3_S**, 262,144 context, KV int8 streaming, vision encoder on; 7,911 experts / 15.0 GiB in VRAM.

## Results (engine-reported tok/s, greedy, thinking off)

| | Decode | Hit rate | Prompt |
|---|---:|---:|---:|
| Cold first request | 67.1 | 82.4% | — |
| Warm, before `--calibrate` | 87–91 | 92% | — |
| Warm, calibrated, `numactl --cpunodebind=0 --membind=0` (kept) | **94.9** (93.7–95.9) | 90.7% | **1,304** on a ~5K prompt |
| Warm, `numactl --interleave=all` (35 workers) | 98.8 (95.9–100.6) | 90.7% | 1,239 |
| Warm, no numactl | 90.1 (88.2–93.4) | 90.6% | 1,268 |

`--calibrate` kept `--pcie-frac 0.20` (the 0.55 default is 11% slower on Gen3), `--spec-min-p 0.70` (+9%, the largest single gain) and 17 pool workers (17 > 11 > 8). 46.84 GiB of experts load at 5.2 GiB/s; service ready in 32 s. Real agent traffic (Claude Code / OpenClaw / Hermes, thinking on): 56–80 tok/s at 76–85% hit rate; a 13.5K-token agent prompt read in 7 s.

## Notes for others on dual-socket / older Xeon hosts

- Binding the engine to the GPU's NUMA node was worth +5% over the default placement; interleaving both sockets gained +4% decode but cost 5% prompt speed, so it was not kept. The hugetlb pool has to follow the placement (all pages on node 0 for binding, split for interleave) or the comparison is silently unfair; the included script handles that.
- On Ubuntu 22.04 with the HWE 6.8 kernel the NVIDIA DKMS build needs gcc 12 (`-ftrivial-auto-var-init=zero`); not a Strata issue, but it blocks the driver install that Strata requires.
- This answers the question in #508 about whether a used Z840 with dual Xeons is a sensible Strata host: with a 24 GB card, yes — the CPU generation and PCIe 3.0 cost little at a 92% expert-cache hit rate.

Limitations: one short prompt and one ~5K prompt, output capped at 512 tokens, no needle/quality suite, single GPU, one request at a time.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Sur le site

Liens install, modèles, releases.