Rows are measured configurations, not a league table. GPU VRAM and PCIe generation are separate because host-to-device bandwidth can materially change large-model decode. Prompt length, model and cache choices also affect throughput. Peak system RAM is process max RSS unless the row says otherwise.
Public benchmark ledger
Explore what changes.
Keep the context.
Compare measured speed by GPU, PCIe link and model, inspect quality evidence, or see every supported model ordered by its persistent system RAM floor. Every row keeps the configuration details that make the number meaningful.
Speed results
RTX PRO 6000 · 96 GB · PCIe 4.0 ×16
Prefill is prompt processing speed. Decode is Krasis’s internal engine measurement; HTTP round trip includes the local client/server path.
Scroll sideways to inspect every column →
| Model | Params | PCIe | Attention + KV | Prefill | Decode | HTTP round trip | Peak system RAM | HCS | Min free VRAM |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash-0731 | 284B main · 13B active | 4.0 ×16 | INT4/HQQ6/Native cache | 2,620.6 tok/s | 38.16 tok/s | 70.30 tok/s | 152.3 GB process RAM (benchmark report) | 6660/11008 (60.5%) | 1,189 MB |
| Qwen3-Coder-Next | 80B main · 3B active | 4.0 ×16 | INT4/HQQ4/k4v4 | 11,211.1 tok/s | 91.34 tok/s | 161.82 tok/s | 79.2 GB process RAM (benchmark report) | 24576/24576 (100.0%) | 53,156 MB |
| Ornith-1.0-397B | 397B main · 17B active | 4.0 ×16 | INT4/HQQ4/k4v4 | 2,354.5 tok/s | 23.58 tok/s | 41.73 tok/s | 201.8 GB process RAM (benchmark report) | 13161/30720 (42.8%) | 1,346 MB |
| Step-3.7-Flash | 199.4B main · 13.9B active | 4.0 ×16 | INT4/HQQ4/k4v4 | 5,261.0 tok/s | 55.40 tok/s | 112.82 tok/s | 103.4 GB process RAM (benchmark report) | 11121/12096 (91.9%) | 1,274 MB |
Quality compares quantised runs with a BF16 reference. Lower drift is closer to BF16. A blocked row means that exact configuration still awaits the required comparison; it is not presented as a pass.
This is the persistent model floor for the default INT4 launcher profile: routed expert storage plus dual HQQ staging. Add RAM for the operating system and any optional conversation cache. GPU VRAM is separate and Krasis measures its budget at runtime. Active parameters approximate the portion of each MoE model used per token; they help explain decode work, not the full storage needed.
Base data mirrors the public Krasis ledgers at revision 7ab1cea. Supplemental timing-disabled runs link their reports separately.Speed source ↗Quality source ↗DeepSeek PCIe 5.0 run ↗Ornith-397B PCIe 5.0 run ↗DeepSeek RTX A6000 run ↗
Floors were calculated for every pinned launcher checkpoint at revision ff96d91 from real model geometry, not stored per-model guesses.Model catalog ↗RAM calculation ↗