Pull requests / #794

#794 split: the auto placement sees each stage's PCIe link

closed · @chengshenyangtai · 0 comentários · No GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDA

Descrição

## What this fixes

    `--layer-split auto` picks the split by a predicted window time: layer compute plus a miss cost for experts no cache 
    holds. That miss cost (`190 ms per unit of routed mass`) is **one constant for every stage** — it was fitted on a pair 
    of equal x16 cards, so the search cannot see an asymmetric pair and spreads layers by compute and cache size alone.

    Measured case: 2x RTX 4090 D, the second card on a **physical x1 link** (probed: 23.9 GB/s on the x16 card vs 1.6 GB/s
    on the x1 card). The prompt path streams every missed expert over *its stage's* link, so the layers given to the x1 card
     dominate every chunk:

    | | placement | 29K-token prompt, fresh reads |
    |---|---|---|
    | search as-is | K=22 (26 layers on the x1 card) | ~350 tok/s |
    | with this change | K=38 (the x1 card's cache holds all of its layers' pairs — zero streaming there) | ~1410 tok/s |

    Same machine, same model (Qwen3.8-Flash-Next UD-Q4_K_XL), 3 fresh real-text prompts per arm, engine-reported timings. 
    The old placement's per-chunk streaming on x1 is ~26 GB ÷ 1.6 GB/s ≈ 17 s; the observed chunk time matched.
 
    ## How

    - Probe each stage's link inside the search (`probe_pcie_h2d_gbps`). The stage-setup probes are skipped when 
    `--pcie-frac` is given, and the search needs the reading either way.
    - Scale each stage's miss term by `clamp(20 GB/s / bw, 1, 64)`. 20 GB/s is `pcie_frac_for_gbps`' x16 reference; the 
    lower clamp keeps a symmetric machine on the exact formula the sweep's K was fitted with (all scales 1.0 → the objective
     is unchanged).
    - The objective's miss term becomes per-stage: `Σ_i miss_ms × scale_i × (missed mass of stage i)`, with the per-stage 
    held mass now accumulated in the placement fill.

    Log lines on that machine:
   layer split auto: CUDA0 link 23.9 GB/s -> miss cost x1.00
   layer split auto: CUDA1 link 1.6 GB/s -> miss cost x12.15
   layer split auto: K=38 - predicted 17.1 ms per decode window; the caches hold 10451 of 24576 profiled pairs (~97.4% of
   the routed mass)

    ## Limits / notes

    - The 190 ms constant mixes the decode CPU-pool cost with the prompt link cost; scaling the whole per-stage term is the
    first-order fix. Splitting the constant into its two parts would need a fresh fit on both a narrow and a wide link —
    happy to do that if you want it.
    - Symmetric pairs are untouched: an x16 gen4/5 link reads 23-28 GB/s, so its scale is exactly 1.0.
    - `--pcie-frac` still overrides the decode share exactly as before; this only changes the *placement search*'s cost
    model.

No site

Links install, modelos, releases.