Issues / #1143

#1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all pass

open · @yusheng227507 · 0 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Description

Tested the four "Testers wanted" items on a dual-GPU Turing setup. All four pass — details below.

## Environment

| | |
|---|---|
| OS / kernel | Ubuntu 24.04.5 LTS, Linux 6.8.0-142-generic x86_64 |
| Build | source build (CMake Release, nvcc from CUDA 12.8, `setup.sh`) |
| Driver | 595.91.07 |
| Engine | v0.1.40, commit `1735d64` (matches origin tag) |
| CPU / RAM | Intel i7-10700K, 94 GB RAM |
| GPUs | **2× RTX 2080 Ti 22 GB** (`sm_75`, no BF16/FP8), PCIe 3.0 |
| Model | Qwen3.8-Flash-Next-GSQ-RCO-**IQ3_XXS** (2 shards), vision (mmproj) enabled |
| Production flags | `--layer-split 26 --kv int8 --kv-resident 32768 --expert-cache auto --prefill auto --spec 4 --mtp --max-context 131072`, `STRATA_STAGE_TRIM=1` |

Each test ran on its own config file (production config untouched) and its own server process on `:8080`. Prompts were built with the repo's own `tools/strata_tokenizer.py` against the repo's docs corpus (round-trip re-encode drift = 0 tokens), sent non-streaming to `/v1/chat/completions`, greedy, `reasoning_effort: none`, and read back from the response's `timings` field.

---

## 1. #743 — Turing prompt test (wide top-k kernel) at 90K / 131K / 155K

`STRATA_TOPK_STREAM` read from the config's `env` block, set **before** engine start (fresh process per session, as required by the `static const bool` in `qsa_select.cu`). All three sizes sit above the 90,112-cell kernel boundary. Distinct corpus offsets per size + `cache_n` verified `0` on every single request (no prefix-cache interference).

| Prompt tokens | New kernel (default) | Old kernel (`STRATA_TOPK_STREAM=0`) | Δ |
|---|---|---|---|
| 92,160 (90.0K, just above boundary) | **1,491.2 tok/s** prefill | 1,476.8 tok/s | **+1.0 %** |
| 131,072 | **1,435.8 tok/s** | 1,428.9 tok/s | **+0.5 %** |
| 158,720 | **1,388.1 tok/s** | 1,384.7 tok/s | **+0.2 %** |

- Re-run of 131,072 on the new kernel after the old-kernel session: **1,432.5 tok/s** (−0.2 % vs first run) → no thermal/drift artifact; the ordering holds.
- Decode at 158,720 tokens: 61–64 tok/s, no anomalies, finish reason `length` as expected.
- Wall time for the full request (155K in + 32 out): 115 s.

**Verdict: no regression on Turing; the new wide top-k kernel is equal or slightly faster at every length above the boundary.**

## 2. #776 — `--batch 4` on the all-resident layer-split config

- Config: production flags + `--batch 4`. Server reported `concurrency: {serving: 4, requested: 4}`; engine log confirms `batch windows of up to 4 sequences` on both stages.
- Workload: **52 rounds × 4 concurrent greedy requests** (80-token prompt, `max_tokens: 200`) over ~8 minutes → **208 requests total**.
- Result: **0 failures, engine never exited, server never unreachable, no loops** (the #776 symptom did not reproduce on 0.1.40).
- Per-sequence decode: 24–57 tok/s; aggregate 102–221 tok/s per round. Prompt path: 91–100 tok/s.
- Observation worth noting: with 4 batch slots, expert-cache residency dropped from **89 % → 80 %** (slots take cache); host RAM stayed ~60 GB. So the zero-doorbell/100 %-resident path is *not* what a `--batch 4` run hits on 22 GB cards.

**Verdict: #776 fix confirmed, stable on Turing.**

## 3. #859 — `--pipeline-windows` with interleaved conversations

Engine accepted `--pipeline-windows 2` (clamped from any larger value as documented):

```
strata serve: --pipeline-windows 2: two verifiers per stage, the stages overlap across windows (decode too; the GDN snapshots beside the window)
```

Tested both requested combinations, **4 conversations × 4 rounds round-robin interleaved = 16 interleaved pairs** each (exceeds the "10 or more pairs" ask), prompts growing 72 → 147 tokens, 300-token completions:

| Config | Pairs | Errors | Loops | Decode tok/s (min/mean/max) |
|---|---|---|---|---|
| `--pipeline-windows 2` + `--resident-experts` (#848) | 16 | 0 | 0 | 76.7 / 83.9 / 96.8 |
| `--pipeline-windows 2` only | 16 | 0 | 0 | 75.7 / 83.5 / 89.7 |

Second verifier costs 284 MiB of CUDA0 expert cache when resident mode is off (per engine log). No prompt-path or decode anomalies in either combination.

**Verdict: both combinations stable; no hangs, no crashes.**

## 4. #848 — resident RAM mode soak (30 minutes, solo serving)

Config: production flags + `--resident-experts`. Engine log confirms the mode is live:

```
strata generate: resident RAM mode: 5.46 GiB of experts in RAM (page-locked), 12094 in the GPU cache; adaptive swaps exchange them with the GPU cache (no file reads)
strata serve: layer split: 87% of the experts resident, the prompt path's streamed ring 96 slots
```

Workload: **30 minutes, 176 requests** — one growing conversation (4 short rounds) plus a **20K-token prompt every 5th request** (30 of them) to force expert-set churn through the adaptive swaps.

| Metric | Result |
|---|---|
| Requests / errors / loops / HTTP failures | **176 / 0 / 0 / 0** |
| 20K-token prefill | 1,168.6 – 1,197.7 tok/s, flat for the whole 30 min (no degradation) |
| Decode | 65 – 87 tok/s (mean 76.1) |
| Host RAM (steady state, last 5 min) | 18.9 – 19.1 GB |
| GPU memory | 21.6 – 21.7 GB per card (stable, no creep) |
| GPU temperature | 50 – 61 °C |

**Verdict: stable for the full 30-minute soak; adaptive swaps show no slowdown over time and no memory creep.**

---

## Summary

| Item | Result |
|---|---|
| #743 Turing prompts 90K/131K/155K A/B | ✅ no regression; new kernel +0.2 % … +1.0 % |
| #776 `--batch` on resident split | ✅ 208 concurrent requests, no engine exit |
| #859 `--pipeline-windows` ≥10 interleaved pairs | ✅ both with and without #848 |
| #848 resident RAM mode | ✅ 30-min soak, flat perf, no creep |

Happy to rerun with different flags, add `STRATA_STAGE_TRIM=0` variants, or share the raw timing JSONs (all recorded under `test-results/`).

Sur le site

Liens install, modèles, releases.