モデル一覧

from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="local")
resp = client.chat.completions.create(
model="strata",
messages=[{"role": "user", "content": "Estimate VRAM for IQ3_XXS"}],
)
print(resp.choices[0].message.content)Base: Qwen3.8-Flash-Next
| Variant | RAM | VRAM | サイズ | RTX 5070 tok/s | models.actions |
|---|---|---|---|---|---|
| Q2_0 Fastest quant; great for 48 GB RAM systems. | 48 GB+ | 12 GB+ | ~70 GB | 94 | |
| IQ2_XS Recommended default for 64 GB RAM; balance of speed and quality. | 64 GB+ | 12 GB+ | ~72 GB | 79 | |
| IQ3_XXS Smarter than IQ2; still fits most 64 GB builds. | 64 GB+ | 12 GB+ | ~74 GB | 62 | |
| IQ3_S Best quality that still runs on 64 GB; slowest of the IQ line. | 64 GB+ | 12 GB+ | ~76 GB | 53 | |
| Coder Coding-focused variant; fits 32 GB RAM (weaker on CJK general text). | 32 GB+ | 12 GB+ | ~68 GB | 55 | |
| Unsloth UD-IQ4_XS ~4-bit Unsloth pack; needs ~80 GB RAM or partial SSD offload. | 80 GB+ | 12 GB+ | ~94 GB | 45 |
Full matrix: docs/MODELS.md