Issues / #1483

#1483 Local-adaptation tuning directions (method, not values) — plus the hardware-specialization layer (model → GPU → CPU) that I'm building

open · @1314521gjy · 2 评论 · 在 GitHub 查看

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

描述

Hi — I'm `@1314521gjy`. I adapted this engine to my machine (RTX 4080S 32 GB + 96 GB), and that work's report is already merged (**#780**); **#1351** (a report) and **#1354** (docs) are still open. Before this I did the same kind of work on another local stack — a hardware-adaptation fusion build at `github.com/1314521gjy/ninfer-fusion-kvmem` — so **what I do is make an engine fit a specific card and CPU.**

This issue brings two things: **a set of tuning directions** (method, not values, from roughly a hundred controlled runs on one machine), and **the hardware-specialization layer I am working on now** — this issue is one of that work's outputs, not a suggestion I'm handing over.

## What I'd like to add

Your `--calibrate` and `DETAILS.md` already say *which* knobs a PC should tune. I'd like to contribute the other half: **how to tune without lying to yourself.** It deliberately carries **directions, not values** — on my own machine one unchanged config measured 111 → 134 tok/s across sessions, so values are not portable.

**The rules that changed our results:**

- **Self-witness first.** Every knob needs a field the engine prints *itself*, that can go red. Without it, "no difference" is unreadable — the switch may simply not have engaged. We wasted a full sweep of arms on a switch nobody could prove had engaged, which is what made this a rule.
- **Discard a warm-up arm, and mirror the order (A B B A).** The first arm of a session is systematically cold, and absolute speed drifts; only same-batch comparisons are evidence.
- **The VRAM-headroom cliff.** When free VRAM approaches zero the driver pages, the engine reports nothing, and timing stops being *predictable* rather than merely slower: we measured the **same deterministic counters 45% apart**. The field to watch is the one the engine prints when it captures each window — not the cache size you configured.
- **Some knobs are machine-specific, and we have a second machine to prove it.** A member of our group ran the same directions on a 24 GB laptop card with a hybrid-core CPU and found the **opposite** optima for the CPU expert thread count and for the CPU/GPU miss split. His write-up (shared with his permission): context **256K → 512K at almost no cost**, expert slots **+26%**, cold prefill **+13–19%**, and he dodged the headroom cliff — **22.8 → 66.0 tok/s** once free VRAM went from 262 MiB to 1069 MiB. He labels that explicitly as *dodging a trap, not tuning*. Same method, different machine, opposite values: **the method transfers, the values do not.**
- **A feature that is a big win elsewhere can be worth nothing here.** The elastic-KV route that gave him +29% expert slots gave us **no slot gain at all** — because we already run KV streaming, which solves the same problem. Same mechanism, opposite conclusion.
- **Don't use the acceptance rate as a verdict.** Tightening the draft floor *raised* acceptance (0.60 → 0.80) and still got slower; what matters is tokens per verification, and then wall-clock.

## What I'm working on: completing the specialization at the hardware layer (model → GPU → CPU)

**This part is current work on my side, and this issue is one of its outputs** — not a suggestion I'm handing over.

This engine is a **model-specific weapon**: one model's geometry is hard-wired, and that hard-wiring is where its order-of-magnitude advantage comes from. What I'm working on is **completing the same idea in the other direction** — giving the *hardware* the same treatment:

- **GPU-specific**: how much VRAM is left on *this* card, how wide is *this* link, and what its bandwidth and latency actually are.
- **CPU-specific**: how many cores, which instruction set, how many memory channels, and how the host thread and the worker pool are laid out.
- Both treated as **first-class inputs** rather than one-off calibration — and both **verifiable**: every profile binds to a field the engine prints.

The model layer answers *for whom*; the hardware layers answer *on what*. The light route I'm using: **probe → pick a profile → bind each profile to a self-witness field → keep profile + readings + rollback as one portable record.** Then the third layer is not guesswork — it is checkable the way the first layer is.

The direction guide above is the first product of that work; the profile-and-record side is next. **I'd rather show findings first, and touch your tree only when you say so.**

## Three questions — any answer is fine, including "no"

1. **Would you like this as a document in the repo? And if so, where** — `docs/`, a section of `DETAILS.md`, or `bench/results/`? If you'd rather not, say so and I'll keep it out of the repo.
2. Is the **hardware-specialization layer** useful to you — and would you like the outputs of that work to keep coming back to you in this shape (findings first, doc/code changes only on your word)?
3. May I add **one line** to the tuning part of `DETAILS.md` — *"always measure, and bind each knob to a field the engine prints"* — as a minimal separate PR?

No AI-assisted wording anywhere; I sign as `1314521gjy`.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。