Issues / #1483
#1483 Local-adaptation tuning directions (method, not values) — plus the hardware-specialization layer (model → GPU → CPU) that I'm building
open · @1314521gjy · 2 Kommentare · Auf GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
Beschreibung
Hi — I'm `@1314521gjy`. I adapted this engine to my machine (RTX 4080S 32 GB + 96 GB), and that work's report is already merged (**#780**); **#1351** (a report) and **#1354** (docs) are still open. Before this I did the same kind of work on another local stack — a hardware-adaptation fusion build at `github.com/1314521gjy/ninfer-fusion-kvmem` — so **what I do is make an engine fit a specific card and CPU.** This issue brings two things: **a set of tuning directions** (method, not values, from roughly a hundred controlled runs on one machine), and **the hardware-specialization layer I am working on now** — this issue is one of that work's outputs, not a suggestion I'm handing over. ## What I'd like to add Your `--calibrate` and `DETAILS.md` already say *which* knobs a PC should tune. I'd like to contribute the other half: **how to tune without lying to yourself.** It deliberately carries **directions, not values** — on my own machine one unchanged config measured 111 → 134 tok/s across sessions, so values are not portable. **The rules that changed our results:** - **Self-witness first.** Every knob needs a field the engine prints *itself*, that can go red. Without it, "no difference" is unreadable — the switch may simply not have engaged. We wasted a full sweep of arms on a switch nobody could prove had engaged, which is what made this a rule. - **Discard a warm-up arm, and mirror the order (A B B A).** The first arm of a session is systematically cold, and absolute speed drifts; only same-batch comparisons are evidence. - **The VRAM-headroom cliff.** When free VRAM approaches zero the driver pages, the engine reports nothing, and timing stops being *predictable* rather than merely slower: we measured the **same deterministic counters 45% apart**. The field to watch is the one the engine prints when it captures each window — not the cache size you configured. - **Some knobs are machine-specific, and we have a second machine to prove it.** A member of our group ran the same directions on a 24 GB laptop card with a hybrid-core CPU and found the **opposite** optima for the CPU expert thread count and for the CPU/GPU miss split. His write-up (shared with his permission): context **256K → 512K at almost no cost**, expert slots **+26%**, cold prefill **+13–19%**, and he dodged the headroom cliff — **22.8 → 66.0 tok/s** once free VRAM went from 262 MiB to 1069 MiB. He labels that explicitly as *dodging a trap, not tuning*. Same method, different machine, opposite values: **the method transfers, the values do not.** - **A feature that is a big win elsewhere can be worth nothing here.** The elastic-KV route that gave him +29% expert slots gave us **no slot gain at all** — because we already run KV streaming, which solves the same problem. Same mechanism, opposite conclusion. - **Don't use the acceptance rate as a verdict.** Tightening the draft floor *raised* acceptance (0.60 → 0.80) and still got slower; what matters is tokens per verification, and then wall-clock. ## What I'm working on: completing the specialization at the hardware layer (model → GPU → CPU) **This part is current work on my side, and this issue is one of its outputs** — not a suggestion I'm handing over. This engine is a **model-specific weapon**: one model's geometry is hard-wired, and that hard-wiring is where its order-of-magnitude advantage comes from. What I'm working on is **completing the same idea in the other direction** — giving the *hardware* the same treatment: - **GPU-specific**: how much VRAM is left on *this* card, how wide is *this* link, and what its bandwidth and latency actually are. - **CPU-specific**: how many cores, which instruction set, how many memory channels, and how the host thread and the worker pool are laid out. - Both treated as **first-class inputs** rather than one-off calibration — and both **verifiable**: every profile binds to a field the engine prints. The model layer answers *for whom*; the hardware layers answer *on what*. The light route I'm using: **probe → pick a profile → bind each profile to a self-witness field → keep profile + readings + rollback as one portable record.** Then the third layer is not guesswork — it is checkable the way the first layer is. The direction guide above is the first product of that work; the profile-and-record side is next. **I'd rather show findings first, and touch your tree only when you say so.** ## Three questions — any answer is fine, including "no" 1. **Would you like this as a document in the repo? And if so, where** — `docs/`, a section of `DETAILS.md`, or `bench/results/`? If you'd rather not, say so and I'll keep it out of the repo. 2. Is the **hardware-specialization layer** useful to you — and would you like the outputs of that work to keep coming back to you in this shape (findings first, doc/code changes only on your word)? 3. May I add **one line** to the tuning part of `DETAILS.md` — *"always measure, and bind each knob to a field the engine prints"* — as a minimal separate PR? No AI-assisted wording anywhere; I sign as `1314521gjy`.
Mehr auf der Site
Links zu Install, Modellen, Releases.