Public benchmark ledger

Explore what changes.
Keep the context.

Compare measured speed by GPU or model, then inspect the quality evidence for each supported model and quantisation profile. Every row keeps the configuration details that make the number meaningful.

4measured configurations

Speed results

RTX PRO 6000 Blackwell · 96 GB

Prefill is prompt processing speed. Decode is Krasis’s internal engine measurement; HTTP round trip includes the local client/server path.

Scroll sideways to inspect every column →

HardwareModelParamsActive paramsAttention + KVPrefillDecodeHTTP round tripPeak system RAMHCSMin free VRAM
1x RTX PRO 6000 Blackwell 96GB, AMD EPYC 7742DeepSeek-V4-Flash-0731304.2B checkpoint / 284B main13B mainINT4/BF16/BF16 KV152.2 @1K / 320.6 @2K / 906.3 @8.6K / 1,328.2 @23K / 1,204.3 @62K tok/s29.38/28.20/28.49 @1K (50/100/250 outputs); 19.41 @62K tok/s54.05/35.97/30.60 @1K; 19.41 @62K tok/s145.5 GB process RAM (benchmark report)6440/11008 (58.5%)650 MB
1x RTX PRO 6000 Blackwell 96GB, AMD EPYC 7742Qwen3-Coder-Next80B total3B activeINT4/HQQ4/k4v411,211.1 tok/s91.34 tok/s161.82 tok/s79.2 GB process RAM (benchmark report)24576/24576 (100.0%)53,156 MB
1x RTX PRO 6000 Blackwell 96GB, AMD EPYC 7742Ornith-1.0-397B397B total17B activeINT4/HQQ4/k4v42,354.5 tok/s23.58 tok/s41.73 tok/s201.8 GB process RAM (benchmark report)13161/30720 (42.8%)1,346 MB
1x RTX PRO 6000 Blackwell 96GB, AMD EPYC 7742Step-3.7-Flash201.4B total / 199.4B text13.9BINT4/HQQ4/k4v45,261.0 tok/s55.40 tok/s112.82 tok/s103.4 GB process RAM (benchmark report)11121/12096 (91.9%)1,274 MB
How to read speed

Rows are measured configurations, not a league table. Prompt length, model, hardware and cache choices all affect throughput. Peak system RAM is process max RSS unless the row says otherwise.