Issues / #620

#620 IQ3_S: "native head upload: out of memory" at startup when the desktop runs on a second GPU (more free VRAM on the card)

closed · @gu-feng · 1 comentários · No GitHub

Server & APINVIDIA / CUDAModels & quantsWindows

Descrição

<html>
<body>
<h2 id="environment" class="atx">Environment</h2>
<ul>
<li>Engine 0.1.38 ready-made release (<code>BUILD.json</code>: version 0.1.38, source &quot;release&quot;, cuda 13.0, archs include 120), Windows 11 Pro build 26200</li>
<li>NVIDIA RTX 5060 Ti 16 GB (sm_120) — the only CUDA device; <strong>the desktop runs on the iGPU</strong> (Intel UHD 770), so the RTX card has no display attached and sits at 0 MiB used before the model starts</li>
<li>i5-14600K, 64 GB RAM</li>
</ul>
<h2 id="model-and-args" class="atx">Model and args</h2>
<p>IQ3_S (GSQ-RCO):</p>
<pre><code class="fenced-code-block">--expert-cache auto --prefill auto --spec 4 --max-context 131072 --kv int8 --vision --vram-reserve-mib 700</code></pre>
<p>Notably <strong>without <code>--kv-resident</code></strong> — a config regeneration had dropped it, so the full 131072-cell KV lived in VRAM.</p>
<h2 id="symptom" class="atx">Symptom</h2>
<p>Startup dies right after the arena load:</p>
<pre><code class="fenced-code-block">strata generate: expert arena read unbuffered (0 of 16 probe reads from the file cache; 6.3 GiB available, 51.1 GiB of files)
strata generate: expert arena: VirtualLock stopped at 0 of 1302 MiB (error 87); cudaHostRegister of the whole arena
                 FAILED (out of memory); 47 slices pinned (45 GiB); large pages (2097152 B)
strata generate: loaded 46.84 GiB at 4.80 GiB/s
strata generate: native head upload: out of memory</code></pre>
<p>The <code>experimental native Q5_K head, 521472000 bytes</code> line never prints — the cudaMalloc of the head (521 MB) fails, and the process exits before ready.</p>
<h2 id="the-strange-part" class="atx">The strange part</h2>
<p>It reproduces when the card has <strong>more</strong> free VRAM, not less:</p>

Display cable | dGPU VRAM before start | Result
-- | -- | --
Monitor on the iGPU (desktop off the card) | 0 MiB used | "native head upload: out of memory"
Monitor on the RTX card (desktop takes ~1 GB) | ~1 GB used | starts fine


<p>With <code>--kv-resident 32768</code> added (the KV streamed to pinned RAM, ~2 GB of VRAM freed), the same start with the display on the iGPU succeeds:</p>
<pre><code class="fenced-code-block">strata generate: experimental native Q5_K head, 521472000 bytes
strata generate: expert cache auto: 9.12 GiB free, 700 MiB reserved (+218 MiB for the draft head) -&gt; 3318 slots
strata generate: expert cache 3771 slots, 7.23 GiB of VRAM; policy is
strata serve: 514 MiB of VRAM free with everything loaded</code></pre>
<p>So the head&#39;s allocation OOMs only in the combination with more free VRAM and no KV streaming — which reads like the ~521 MiB head is not accounted for by the expert-cache sizing in this configuration (&quot;THE HEAD BEFORE THE CACHE&quot; from 0.1.36&#39;s notes says the head is loaded before the sizing, but on this card it OOMs anyway).</p>
<p>Happy to attach the full <code>strata-iq3_s.log</code> of the failing start or test any candidate fix.</p>
<!--EndFragment-->
</body>
</html>

No site

Links install, modelos, releases.