Issues / #1312

#1312 [Windows][RTX 3090 24 GB, 64 GB RAM, IQ3_S] Feature request: expose the QSA indexer budget (512 blocks / 2048 tokens)

open · @eakkawat · 0 comments · View on GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindows

Description

### What I'd like
An engine-side override for the QSA indexer's selection budget — 512 blocks / 2048 tokens — like the flag SGLang
exposes. A field report raised it to 1024 blocks / 4096 tokens and reported that a long-context retrieval answer
stopped being hallucinated, at about a 6% decode cost.

### Why it matters
12 of the model's 48 layers are QSA. Each keeps the best 512 blocks of 4 tokens (2048 positions, plus the tail of the
current incomplete block), so the final attention sees at most 2051 positions no matter how long the context is. If
the sentence that answers a question sits in a block that doesn't make the top-512, that layer never reads it — even
at 131K context. That is a plausible mechanism for the "invents a plausible answer instead of saying it can't find
it" failures, which look like the same shape DeepSeek-V4-Flash shows with its own sparse attention (2048 individual
tokens, per Sebastian Raschka's comparison).

Reference (not my measurement): a field report on YouTube,
<https://www.youtube.com/watch?v=33MjW1aw2tU> — the flag is described at 9:17 (<https://www.youtube.com/watch?v=33MjW1aw2tU&t=557s>) and the result at 10:19
(<https://www.youtube.com/watch?v=33MjW1aw2tU&t=619s>). He raised the budget (512 -> 1024 blocks, 2048 -> 4096 tokens) with SGLang on a cyber-CTF
prompt where the answer is genuinely absent, and the model then stated the value was not in the logs and pointed at
the right link instead of fabricating, in two runs. Caveat: 2 runs, 1 benchmark, different engine and quant, so I am
not claiming it generalises — only that the knob is worth having.

### Current state on 0.1.39 (Windows)
- `engine/strata.exe`: 154 flags, none of them sets the indexer selection size. `--native-qsa-indexer` is a
  pinned-kernel toggle ("experimental pinned indexer key cache and pooling") with no value.
- `serve/runconfig.py`: no key for it. Its only budget is `reasoning_budget_tokens` (per-request thinking cap).
- The ~140 `STRATA_*` knobs include QSA/top-k kernel switches (`STRATA_TOPK_OLD`, `STRATA_TOPK_STREAM`,
  `STRATA_QSA_WARP`, `STRATA_INDEXER_PER_TOKEN`, `STRATA_IDX_FP16_CHECK`, `STRATA_SELECT_OLD`) but nothing that
  changes the selection size.
- The load log prints no line with the effective budget, so I cannot confirm which value is live.

### The ask
1. Is the 512-block budget a constant in the loader/kernels, or read from the model/config (the checkpoint describes
   it as "indexer MQA ... budget 512 blocks or 2048 tokens")? If it is a model attribute, a runtime override looks
   cheap to expose.
2. If it is cheap: a knob (`--qsa-indexer-blocks N`, or `STRATA_QSA_BLOCKS`) — and if you would rather not carry the
   feature, at least one load-time log line with the effective budget, so it is visible.
3. Expected cost: 2x blocks means 2x indexer scoring work and up to 2x sparse-attention KV reads per QSA layer; the
   report above measured about 6% overall on a 3090-class setup.

### What I can test and report
Windows 11, RTX 3090 24 GB, 64 GB RAM, IQ3_S (GSQ-RCO), `--kv k8v4`, `--max-context 131072`,
`--vram-reserve-mib 1015`, no RAM budget. Greedy baseline at a 3053-token prompt on 0.1.39: 1018 tok/s prompt read,
62.1 t/s decode cold-read, 55.8 t/s cached, 94.4% expert-cache hit.

If a knob appears, I will A/B greedy at 512 vs 1024 blocks on retrieval prompts (answer located at the far end of a
long prompt, and a prompt where the answer is genuinely absent) and report decode/prefill for both.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.