Issues / #617

#617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box

closed · @noahark · 3 コメント · GitHub で見る

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

本文

# V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box

Follow-up to #585 (the sm_70 link fix that made this install possible). Three things below that may be useful
to other V100 / older-GPU users and maybe to the project: measured throughput, a thinking-budget pitfall with
a one-field fix, and a capability comparison against Qwen3.8-27B (llama.cpp) on identical hardware.

## Setup

- Tesla V100-PCIE-32GB (TCC, PCIe Gen3), dual Xeon E5-2696 v3, 128 GB RAM, Windows 10 LTSC
- Strata v0.1.38, engine compiled from source for sm_70 (CUDA 12.4, MSVC 14.44)
- Qwen3.8-Flash-Next IQ3_XXS, 131,072 context, KV streamed to RAM, MTP draft (q2_0)
- 15,553 / 24,576 experts (63%) cached in VRAM (25.2 GiB); rest served from RAM by the CPU

## Throughput (server-side timings)

| Scenario | Output | Speed | Expert cache hit |
|---|---|---|---|
| First request ever (JIT warm-up) | 175 tok | 7.5 tok/s | 81.0% |
| Short answer | 39 tok | 10.1 tok/s | 94.8% |
| Medium answer | 289 tok | 13.7 tok/s | 95.5% |
| ~400-word essay (thinking + prose) | 322 tok | 14.2 tok/s | 95.6% |
| Long-form writing | 1,492 tok | 15.8 tok/s | 96.6% |
| Long multi-method math reasoning | 1,761 tok | 16.8 tok/s | 97.3% |
| 31k-token interactive session (hardest problem, web UI) | 31,796 tok | 16.7 tok/s steady | 96.6% |

Prompt processing 61.7 tok/s warm. MTP acceptance 97.8% on repetitive output, ~65% on free prose.
Decode is memory-bound on this card: GPU util ~19%, 42-48 W of 250 W, 49 C - the GPU mostly waits
while the CPU computes RAM-resident experts. Notably **speed does not degrade with context**: flat
16.7-16.9 tok/s through a 31k-token session (KV streaming presumably).

## The thinking-budget pitfall (with a one-field fix)

Hard problems can burn the *entire* `max_tokens` on thinking and emit an empty answer, silently.

- Example (factorial-trailing-zeros olympiad-style prompt): with `max_tokens=5000` AND with
  `max_tokens=15000`, completion_tokens hit the cap exactly and `content` was empty - 15,000 tokens
  for nothing. Same temperature/top_p/seed both times (near-deterministic replay of a spiraling path).
- In the web UI (fresh sampling) the same prompt solves correctly in 6,699 tokens.
- **Fix: send `reasoning_budget_tokens: 4000` per request.** The server already supports it (serve/server.py,
  `reasoning_budget()`). With the cap, the *same spiraling seed-42 request* stopped thinking and answered
  correctly (8090). Across three hard problems, capped Flash matched a Qwen3.8-27B running under
  `--reasoning-budget 4096` token-for-token in correctness (correct / partial / correct) at 14,369 total
  tokens vs the 27B's 14,306.

Suggestions, if useful: mention `reasoning_budget_tokens` in the README/API docs for programmatic users,
and consider exposing a thinking cap in the web UI (the UI currently only has all-or-nothing skip).

## Flash-125B (Strata) vs Qwen3.8-27B (llama.cpp thinking build), same hardware

Six-question boundary battery + the exam above, both under thinking caps:

| Dimension | Flash-125B | Qwen3.8-27B |
|---|---|---|
| Mainstream cross-domain knowledge (10 sub-questions) | 10/10 | 9/10 |
| Long-tail knowledge (name all four Wallfacers from "The Three-Body Problem") | 4/4 exact | 2/4, two hallucinated names |
| Long-doc retrieval (2.6k-token synthetic log, 12 planted cross-referenced facts) | 12/12 | 12/12 |
| Exact code-output prediction (metaclass/descriptor/closure; try-else-finally x continue) | 2/2 char-exact | 2/2 char-exact |
| 5-constraint JSON output | perfect | perfect |
| Character-count constraint (write 3 sentences with exact counts of a specific character) | failed | passed |
| Proof completeness (optimal bracket-flip algorithm), unlimited budget | full optimality proof | truncated |

Takeaway for our use: the 125B wins on long-tail knowledge and proof depth; the 27B wins on speed
(1.7-2x wall clock), fine-grained format obedience, and stability. Their hallucination surfaces are
disjoint, so we cross-check important facts between them.

Thanks for the project - happy to contribute the numbers to docs/COMMUNITY_BENCHMARKS.md in PR form
if the table there wants a V100 row.

関連リンク

インストール・モデル・リリースへの站内リンク。