Pull requests / #773
#773 file tier: unbuffered reads on Linux too (O_DIRECT + the kernel's asynchronous reads)
closed · @cakescats · 0 comments · View on GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindowsLinux
Description
#286's unbuffered file tier exists on Windows only. On Linux the experts outside the RAM copy are always read through the page cache, as page faults of the mapped GGUF: small requests on the critical path, and pages that a 32 GB PC cannot keep beside a 16 GiB RAM copy anyway, so the same experts are read from the drive again and again. This adds the Linux side of the same mechanism. `open_direct` opens the files with `O_DIRECT`. `read_direct` queues every 4 KiB-aligned window of a batch with one `io_submit` and waits for all of them. That is the kernel's own asynchronous read (`linux/aio_abi.h`): no library, and Docker's default seccomp profile allows it, unlike io_uring. Each thread has its own context and aligned buffer, so concurrent batches never see each other's completions. Windows merges windows up to 1 MiB apart because of NTFS; here only windows that touch are merged. A refused submit or a failed read falls back to the mapping, as on Windows. The decision is Windows' `cache_counts = false` arm: unbuffered when the file cache could not keep the expert bytes outside the RAM copy (`platform::file_cache_keeps`, from MemAvailable and the tightest cgroup limit via `host_available_memory`). It is re-checked once the copy is built (#577), and `STRATA_UNBUFFERED_LOAD=1/0` forces it either way. A machine whose RAM can cache those bytes keeps the buffered path. A pack's `experts.bin` works too: its Linux open now records the path. `direct_fallbacks()` counts the blobs an unbuffered read could not deliver. The Windows code paths are unchanged; I had no Windows machine to build on. The new `#elif defined(__linux__)` branches sit beside them. **Measured.** RTX 3080 Ti Laptop 16 GB, i9-12900H, 30 GB RAM, KIOXIA PCIe 4 NVMe, Linux 7.0, ext4. Setup: - Qwen3.8-Flash-Next Abliterated Q4_K_M in place (`--native`), `--resident-budget-gib 16`, `--expert-cache auto`, 8K context, greedy. - Three prompts: code, Russian prose, a 7.8K-token summary. Each was run with and without MTP. - Two rounds of buffered (`STRATA_UNBUFFERED_LOAD=0`) against unbuffered, with the same binary. | decode, tok/s | buffered | unbuffered | | |---|---|---|---| | code, no MTP | 3.77 | 7.58 | ×2.01 | | prose, no MTP | 5.12 | 9.93 | ×1.94 | | 7.8K summary, no MTP | 4.00 | 7.22 | ×1.81 | | code, MTP | 4.42 | 7.97 | ×1.80 | | prose, MTP | 4.47 | 8.00 | ×1.79 | | 7.8K summary, MTP | 4.53 | 6.75 | ×1.49 | | **all** | **4.38** | **7.90** | **×1.80** | - The 7.8K prompt read at 154 and 172 tok/s buffered, 184 and 290 unbuffered. - Startup went from 36 s to 18–21 s. - The code prompt's answer is the same text in all 16 runs. The other two differ between runs of either arm, because the adaptive GPU cache changes which experts the CPU computes; that was the case before this change too. **Tests.** `file_expert_source_test` gains `test_unbuffered_reads`. It writes a synthetic pack of random bytes and reads it unbuffered three ways: one blob at a time, a batch of adjacent experts, and `copy_blob`. It requires the file's bytes and `direct_fallbacks() == 0`. A mutation that fails every read is caught. The pack is written in the working directory, not `/tmp`, so a file system without `O_DIRECT` skips the test with a message instead of failing it. `ctest`: 63 of 66 pass. The other three fail on this machine before and after the change, for these reasons: - `ple_parity` needs the Q2_0 pack; - `platform_memory_test` needs a higher `ulimit -l`; - `expert_multi_test` needs AVX-512. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.