Pull requests / #1355

#1355 docs(bench): V100 follow-ups: 0.1.40 engine, UD-IQ4_XS tier, context decay, --parallel 4

open · @noahark · 0 commentaires · Sur GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Description

Rebuilt and slimmed after the history cleanup. The 0.1.38-0.1.39 material from the
original PR #707 has landed in 0.1.40.2 as bench/results/2026-10-03-community-v100 -
thank you. This PR is now rebased on current main and carries the net-new data only:

- Update (2026-10-04): --parallel 4 concurrent decode on the calibrated XXS server.
  Group throughput 84.4 tok/s = 1.43x solo; cost 2.9 GiB of expert cache.
- Update (2026-10-05): UD-IQ4_XS tier on the same box. Install notes (1.38 GiB pack,
  experts read in place from the GGUF, --resident-budget-gib 55 fully RAM-resident),
  calibration keeping only --pool-workers 18 (35/23/18 -> 31.9/30.4/38.0 tok/s; the
  worker spread is 20%+ here vs 11% on IQ3_S and <2% on XXS), hand overrides losing
  to the calibrate picks on real loads, a same-prompt table against XXS and S, 3/3 on
  the H1/H3/H5 quality suite, and two quirks (cold-start warm-up; enable_thinking:
  false with an image drops the image).
- Update (2026-10-07): v0.1.40 ready-made engine, recalibrated (--pcie-frac 0.33):
  calibrate bench 76.9 -> 81.8 tok/s, API decode +8%, 4.3K prefill 1,175 tok/s, H1
  re-checked. This is the Volta (sm_70) data point the 0.1.40 release notes ask for
  under "Testers wanted".
- Update (2026-10-07): decode speed vs context length on all three tiers (0.1.39):
  XXS and IQ3_S show no decay to 128K; UD-IQ4_XS is flat to 96K then about -20% at
  128K (draft acceptance 0.73 -> 0.58). Plus a deployment note about expert files on
  an HDD under page-cache pressure.

Also updates the row in bench/results/COMMUNITY.md (tiers, headline numbers, engine
range to 0.1.40).

Hardware: Tesla V100-PCIE-32GB (sm_70, TCC), 2x Xeon E5-2696 v3, 128 GB DDR3L-1600,
Windows 10, driver 581.80. All numbers are server-side timings; scripts and raw JSON
dumps available on request.

Sur le site

Liens install, modèles, releases.