Plan inference & fine-tuning — hardware, VRAM and cost.
Compare open-source LLMs across a catalog of hundreds of models; live cloud prices, TCO and per-token cost analysis. No sign-up, free.
Start Calculating →Real scenarios, real numbers
These cards are computed server-side by the same calculator engine; each is editable in the app.
Meta Llama 3.3 70B
1× NVIDIA H100 SXM5 (80GB HBM3)
~30 tok/s
$0.36 / 1M tok
74 GB VRAM
Qwen 3 32B
1× NVIDIA GeForce RTX 4090 (24GB GDDR6X)
~39 tok/s
$0.08 / 1M tok
18 GB VRAM
Meta Llama 3.1 8B
1× NVIDIA GeForce RTX 4090 (24GB GDDR6X)
~109 tok/s
$0.08 / 1M tok
7 GB VRAM
Qwen 3 30B-A3B (MoE)
1× NVIDIA GeForce RTX 5090 (32GB GDDR7)
~575 tok/s
$0.11 / 1M tok
17 GB VRAM
Google Gemma 3 27B (Multimodal)
1× NVIDIA RTX 6000 Ada Generation (48GB ECC)
~22 tok/s
$0.09 / 1M tok
30 GB VRAM
Qwen 3 235B-A22B (MoE)
2× NVIDIA H200 SXM (141GB HBM3e)
~289 tok/s
$0.99 / 1M tok
243 GB VRAM
Mistral Small 3 (24B Instruct)
1× NVIDIA GeForce RTX 4090 (24GB GDDR6X)
~52 tok/s
$0.08 / 1M tok
14 GB VRAM
Qwen 3 8B
1× NVIDIA GeForce RTX 3090 (24GB GDDR6X)
~101 tok/s
$0.05 / 1M tok
7 GB VRAM
Inference
TTFT, TPOT, tokens/s and VRAM — from 8B to 671B MoE.
Fine-Tuning
GPU hours, VRAM and platform cost for QLoRA, LoRA and full fine-tuning.
Live Cloud Prices
Up-to-date GPU prices from RunPod, Lambda and Modal; on-prem TCO comparison.
125 open-source models in the catalog