Pull requests / #1722

#1722 bench: RTX 5090 Swift IQ3_XXS 1M tuning and native NVMe restore

open · @mgeldi · 0 コメント · GitHub で見る

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentation

本文

## Summary

Results-only community report measured on 2026-10-09: RTX 5090 (480 W, PCIe 5 x16), Ryzen 7 9800X3D, 91.9 GiB usable RAM, Swift IQ3_XXS, Strata 0.1.41 at `fb58e0d`, 1,048,576-position YaRN context.

- Same-day prefill ceiling comparison at 32K/131K, greedy and sampled/xhigh. Larger chunks improve fresh prefill, but decode changes vary; the local configuration retained `auto`.
- Native NVMe SAVE/restart/RESTORE at 32K/128K/1M. The 14.445 GiB million-token file restored in 6.107 s; continuation reused 1,000,166 tokens and read 51, returning all eight exact audit values.
- Separate external-supervisor supplement: three paired rounds, worker plus durable checkpoint median 17.208 → 14.584 s. This is custom Hermes orchestration around native SAVE and NInfer loading, not an upstream async-SAVE or simultaneous-inference feature.

Related discussion: #669 (different hardware/model/version; no cross-report speedup claim).

## What changed

- One dated report folder: README, original per-run outputs/timings and native timing lines, sanitized provenance/configuration, 39 unique gzip payloads with 96 request entries, portable stdlib replay helper and fixtures.
- One row in `bench/results/COMMUNITY.md`.
- No engine, server, setup, model, default or build changes. Model weights, packs, session images, binaries, private conversations and learned user profiles are excluded.

## Validation

- Eight local HTTP/SSE fixture tests passed.
- All frozen payload hashes and 66 file checksums verified; published request records pass the replay eligibility checks and README links resolve.
- Independent review checked tables, sample sizes, cache reuse, native/external attribution and private-data removal.
- Native engine binary remains unchanged. No extra inference was run just to prepare the contribution.

## Limits and disclosure

Fresh confirmation is **two** requests per cell, initial screen/disk cases **one**, cached decode grid **three** rounds per point/mode; counts are not pooled. Synthetic recall does not establish general 1M quality. The additional arithmetic JSON smoke failed and its actual output is retained. Learned-profile speed comparison is observational, not a general gain claim.

The portable helper replays native requests/slot operations; the external supervisor implementation is not included, so the NInfer-overlap supplement is documented evidence rather than a standalone replay.

Prepared with an AI assistant from locally captured engine/client measurements.

関連リンク

インストール・モデル・リリースへの站内リンク。