Issues / #617
#617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box
closed · @noahark · 3 comments · View on GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Description
# V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box Follow-up to #585 (the sm_70 link fix that made this install possible). Three things below that may be useful to other V100 / older-GPU users and maybe to the project: measured throughput, a thinking-budget pitfall with a one-field fix, and a capability comparison against Qwen3.8-27B (llama.cpp) on identical hardware. ## Setup - Tesla V100-PCIE-32GB (TCC, PCIe Gen3), dual Xeon E5-2696 v3, 128 GB RAM, Windows 10 LTSC - Strata v0.1.38, engine compiled from source for sm_70 (CUDA 12.4, MSVC 14.44) - Qwen3.8-Flash-Next IQ3_XXS, 131,072 context, KV streamed to RAM, MTP draft (q2_0) - 15,553 / 24,576 experts (63%) cached in VRAM (25.2 GiB); rest served from RAM by the CPU ## Throughput (server-side timings) | Scenario | Output | Speed | Expert cache hit | |---|---|---|---| | First request ever (JIT warm-up) | 175 tok | 7.5 tok/s | 81.0% | | Short answer | 39 tok | 10.1 tok/s | 94.8% | | Medium answer | 289 tok | 13.7 tok/s | 95.5% | | ~400-word essay (thinking + prose) | 322 tok | 14.2 tok/s | 95.6% | | Long-form writing | 1,492 tok | 15.8 tok/s | 96.6% | | Long multi-method math reasoning | 1,761 tok | 16.8 tok/s | 97.3% | | 31k-token interactive session (hardest problem, web UI) | 31,796 tok | 16.7 tok/s steady | 96.6% | Prompt processing 61.7 tok/s warm. MTP acceptance 97.8% on repetitive output, ~65% on free prose. Decode is memory-bound on this card: GPU util ~19%, 42-48 W of 250 W, 49 C - the GPU mostly waits while the CPU computes RAM-resident experts. Notably **speed does not degrade with context**: flat 16.7-16.9 tok/s through a 31k-token session (KV streaming presumably). ## The thinking-budget pitfall (with a one-field fix) Hard problems can burn the *entire* `max_tokens` on thinking and emit an empty answer, silently. - Example (factorial-trailing-zeros olympiad-style prompt): with `max_tokens=5000` AND with `max_tokens=15000`, completion_tokens hit the cap exactly and `content` was empty - 15,000 tokens for nothing. Same temperature/top_p/seed both times (near-deterministic replay of a spiraling path). - In the web UI (fresh sampling) the same prompt solves correctly in 6,699 tokens. - **Fix: send `reasoning_budget_tokens: 4000` per request.** The server already supports it (serve/server.py, `reasoning_budget()`). With the cap, the *same spiraling seed-42 request* stopped thinking and answered correctly (8090). Across three hard problems, capped Flash matched a Qwen3.8-27B running under `--reasoning-budget 4096` token-for-token in correctness (correct / partial / correct) at 14,369 total tokens vs the 27B's 14,306. Suggestions, if useful: mention `reasoning_budget_tokens` in the README/API docs for programmatic users, and consider exposing a thinking cap in the web UI (the UI currently only has all-or-nothing skip). ## Flash-125B (Strata) vs Qwen3.8-27B (llama.cpp thinking build), same hardware Six-question boundary battery + the exam above, both under thinking caps: | Dimension | Flash-125B | Qwen3.8-27B | |---|---|---| | Mainstream cross-domain knowledge (10 sub-questions) | 10/10 | 9/10 | | Long-tail knowledge (name all four Wallfacers from "The Three-Body Problem") | 4/4 exact | 2/4, two hallucinated names | | Long-doc retrieval (2.6k-token synthetic log, 12 planted cross-referenced facts) | 12/12 | 12/12 | | Exact code-output prediction (metaclass/descriptor/closure; try-else-finally x continue) | 2/2 char-exact | 2/2 char-exact | | 5-constraint JSON output | perfect | perfect | | Character-count constraint (write 3 sentences with exact counts of a specific character) | failed | passed | | Proof completeness (optimal bracket-flip algorithm), unlimited budget | full optimality proof | truncated | Takeaway for our use: the 125B wins on long-tail knowledge and proof depth; the 27B wins on speed (1.7-2x wall clock), fine-grained format obedience, and stability. Their hallucination surfaces are disjoint, so we cross-check important facts between them. Thanks for the project - happy to contribute the numbers to docs/COMMUNITY_BENCHMARKS.md in PR form if the table there wants a V100 row.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.