Catálogo

Model size and VRAM
OpenAI SDK pointed at Strata
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="local")
resp = client.chat.completions.create(
    model="strata",
    messages=[{"role": "user", "content": "Estimate VRAM for IQ3_XXS"}],
)
print(resp.choices[0].message.content)

Base: Qwen3.8-Flash-Next

VariantRAMVRAMDownloadRTX 5070 tok/smodels.actions
Q2_0
Fastest quant; great for 48 GB RAM systems.
48 GB+12 GB+~70 GB94
IQ2_XS
Recommended default for 64 GB RAM; balance of speed and quality.
64 GB+12 GB+~72 GB79
IQ3_XXS
Smarter than IQ2; still fits most 64 GB builds.
64 GB+12 GB+~74 GB62
IQ3_S
Best quality that still runs on 64 GB; slowest of the IQ line.
64 GB+12 GB+~76 GB53
Coder
Coding-focused variant; fits 32 GB RAM (weaker on CJK general text).
32 GB+12 GB+~68 GB55
Unsloth UD-IQ4_XS
~4-bit Unsloth pack; needs ~80 GB RAM or partial SSD offload.
80 GB+12 GB+~94 GB45

Full matrix: docs/MODELS.md