Pull requests / #1158

#1158 bench: community reports, RTX 4090, IQ3_XXS 204800 and IQ3_S 143360 - 0.1.38 to 0.1.40.1

closed · @Dmitry-B · 0 comentarios · En GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quants

Descripción

…t 143360 - 0.1.38 through 0.1.40.1

Four arms of one configuration on one PC (Ryzen 9 7950X, 48 GB RAM): Strata 0.1.38, 0.1.39, 0.1.40 and 0.1.40.1, each with two prompt arms - code-explanation text and Russian prose - three runs each at 4,096, 32,768 and 131,072 prompt tokens plus a generation-only case, and six recall checks per arm (6/6 in every arm). Prompt throughput 1844/2976/2962 -> 1979/3136/3118 tok/s over 0.1.38 -> 0.1.40; 0.1.40 -> 0.1.40.1 moved nothing, which is the expected result - the engine binary is byte-identical in both reports. The report says which columns are comparable and which are not, names the contaminated runs, and publishes the discarded cold-start run as its own arm.

Plus the IQ3_S 143,360 resident-experts arm of the same PC (0.1.38); two opt-in checks measured on 0.1.40 (STRATA_KV_PREFETCH=1 without --kv-resident: -4.3% at 32K, -2.1% at 128K, nothing to overlap; "parallel": 2: a slot costs 2.88 GiB at this context, the expert cache 6946 -> 3329 slots, one client -28% decode while 3-4 clients get 3-4x lower wall time; --batch-mtp cannot admit with an MMVQ draft pack - #1012, #1063); and a results-only verification of the 0.1.40.1 hotfix: the release corpus serve/fixtures/rcall_specimens.json replayed through the server's own parser (37/37, 444 runs), eight live tool-call requests, three requests waiting for a killed engine (16.6 / 40.5 / 40.4 s, no 300 s hang), and 527 unit tests on this machine.

Supersedes the closed submissions of 2026-10-03 (#624) and 2026-10-06, whose branches were built on the history before main was rewritten. Results-only: no engine or server code is touched.

En el sitio

Enlaces a install, modelos, releases.