Pull requests / #1722
#1722 bench: RTX 5090 Swift IQ3_XXS 1M tuning and native NVMe restore
open · @mgeldi · 0 コメント · GitHub で見る
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation
本文
## Summary Results-only community report measured on 2026-10-09: RTX 5090 (480 W, PCIe 5 x16), Ryzen 7 9800X3D, 91.9 GiB usable RAM, Swift IQ3_XXS, Strata 0.1.41 at `fb58e0d`, 1,048,576-position YaRN context. - Same-day prefill ceiling comparison at 32K/131K, greedy and sampled/xhigh. Larger chunks improve fresh prefill, but decode changes vary; the local configuration retained `auto`. - Native NVMe SAVE/restart/RESTORE at 32K/128K/1M. The 14.445 GiB million-token file restored in 6.107 s; continuation reused 1,000,166 tokens and read 51, returning all eight exact audit values. - Separate external-supervisor supplement: three paired rounds, worker plus durable checkpoint median 17.208 → 14.584 s. This is custom Hermes orchestration around native SAVE and NInfer loading, not an upstream async-SAVE or simultaneous-inference feature. Related discussion: #669 (different hardware/model/version; no cross-report speedup claim). ## What changed - One dated report folder: README, original per-run outputs/timings and native timing lines, sanitized provenance/configuration, 39 unique gzip payloads with 96 request entries, portable stdlib replay helper and fixtures. - One row in `bench/results/COMMUNITY.md`. - No engine, server, setup, model, default or build changes. Model weights, packs, session images, binaries, private conversations and learned user profiles are excluded. ## Validation - Eight local HTTP/SSE fixture tests passed. - All frozen payload hashes and 66 file checksums verified; published request records pass the replay eligibility checks and README links resolve. - Independent review checked tables, sample sizes, cache reuse, native/external attribution and private-data removal. - Native engine binary remains unchanged. No extra inference was run just to prepare the contribution. ## Limits and disclosure Fresh confirmation is **two** requests per cell, initial screen/disk cases **one**, cached decode grid **three** rounds per point/mode; counts are not pooled. Synthetic recall does not establish general 1M quality. The additional arithmetic JSON smoke failed and its actual output is retained. Learned-profile speed comparison is observational, not a general gain claim. The portable helper replays native requests/slot operations; the external supervisor implementation is not included, so the NInfer-overlap supplement is documented evidence rather than a standalone replay. Prepared with an AI assistant from locally captured engine/client measurements.
関連リンク
インストール・モデル・リリースへの站内リンク。