Pull requests / #126

#126 Windows: the expert load runs 24x slower from Task Scheduler (0.05 vs 1.42 GiB/s); a hint names the cause

closed · merged 2026-09-29 · @Zauberio · 0 评论 · 在 GitHub 查看

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

描述

## The symptom

When the serve is started from **Task Scheduler** (a common way to have Strata up at logon), the ~40 GB expert
load crawls at **~0.05 GiB/s** (13-14 min until `ready`). The same binary, same args and same cache state
started from a terminal or SSH loads at **1.4-1.5 GiB/s (~35 s)**.

## The cause (measured)

The Task Scheduler registration defaults: **Priority 7 (Below normal)** and no `<RunLevel>` (a
**least-privilege token**). Re-registering the task with `Priority 4` + `<RunLevel>HighestAvailable</RunLevel>`
makes the same task-spawned chain run at **1.42 GiB/s (35 s)**.

Elimination matrix (RTX 5070 Ti, Ryzen 7 9800X3D, NVMe, 64 GB RAM, Swift-1.5 IQ3_XXS, the v0.1.20 engine):

| launch context / variable | expert load |
| --- | ---: |
| Task Scheduler chain with the defaults | 0.05 GiB/s (821-841 s) |
| SSH -> the *identical* chain (same bat/python/engine/args) | **1.51 GiB/s** |
| engine alone (`--serve --vision`), SSH | 1.43 GiB/s |
| engine alone, cold file cache, 53 GB RAM free | 1.50 GiB/s |
| engine alone + 25 GB memory squeeze (commit 41 GB) | 1.52 GiB/s |
| Task chain at Normal priority (`start /normal`) | 0.05 GiB/s (821 s) |
| **Task chain, Priority 4 + HighestAvailable** | **1.42 GiB/s (35 s)** |

So it is the launch context - not the disk, the cache state, memory pressure, or the engine's serve/vision
mode. Note the priority *class* alone did not help (processes verified at Normal, still 0.05 GiB/s), so the
run-level/limited token or the scheduler's I/O treatment is implicated. Both registration settings were
changed together, so the isolated effect of each is not measured. (Possible connection: a limited token
strips `SeLockMemoryPrivilege`, which the large-pages path wants - not proven.)

## This PR

1. **docs/DETAILS.md** - a "Running it at startup (Task Scheduler)" section: the measurements and the two
   registration settings that restore full speed.
2. **src/program/generate.cpp** - when the reported load rate is < 0.2 GiB/s, print a hint naming this cause
   under the `loaded ... GiB at ...` line, so users do not blame their disk.

Verified: the tree compiles (MSVC + CUDA sm_120, the setup.py recipe), and the hint path was exercised at
runtime - observed printing under the load line (smoke with a temporarily raised threshold; the shipped
threshold is 0.2 GiB/s and the task-context rate measured above is 0.05 GiB/s).

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。