Public benchmark ledger

Explore what changes.
Keep the context.

Compare measured speed by GPU, PCIe link and model, inspect quality evidence, or see every supported model ordered by its persistent system RAM floor. Every row keeps the configuration details that make the number meaningful.

4measured configurations

Speed results

RTX PRO 6000 · 96 GB · PCIe 4.0 ×16

Prefill is prompt processing speed. Decode is Krasis’s internal engine measurement; HTTP round trip includes the local client/server path.

Scroll sideways to inspect every column →

ModelParamsPCIeAttention + KVPrefillDecodeHTTP round tripPeak system RAMHCSMin free VRAM
DeepSeek-V4-Flash-0731284B main · 13B active4.0 ×16INT4/HQQ6/Native cache2,620.6 tok/s38.16 tok/s70.30 tok/s152.3 GB process RAM (benchmark report)6660/11008 (60.5%)1,189 MB
Qwen3-Coder-Next80B main · 3B active4.0 ×16INT4/HQQ4/k4v411,211.1 tok/s91.34 tok/s161.82 tok/s79.2 GB process RAM (benchmark report)24576/24576 (100.0%)53,156 MB
Ornith-1.0-397B397B main · 17B active4.0 ×16INT4/HQQ4/k4v42,354.5 tok/s23.58 tok/s41.73 tok/s201.8 GB process RAM (benchmark report)13161/30720 (42.8%)1,346 MB
Step-3.7-Flash199.4B main · 13.9B active4.0 ×16INT4/HQQ4/k4v45,261.0 tok/s55.40 tok/s112.82 tok/s103.4 GB process RAM (benchmark report)11121/12096 (91.9%)1,274 MB
How to read speed

Rows are measured configurations, not a league table. GPU VRAM and PCIe generation are separate because host-to-device bandwidth can materially change large-model decode. Prompt length, model and cache choices also affect throughput. Peak system RAM is process max RSS unless the row says otherwise.