Best Open-Source AI Models to Self-Host
The best open-weight AI models you can run yourself for zero per-token cost.
Llama 3.1 70B
Meta's flagship open-weight model. Near-frontier quality on most tasks. Runs on 2x A100 GPUs or equivalent. Available via Groq, Together, and others for cheap inference without self-hosting.
Top Picks
Llama 3.1 70B
Best General Open-SourceMeta (self-hosted)
Meta's flagship open-weight model. Near-frontier quality on most tasks. Runs on 2x A100 GPUs or equivalent. Available via Groq, Together, and others for cheap inference without self-hosting.
$0 self-hosted / $0.59/1M via Groq
Calculate your cost with Llama 3.1 70B →DeepSeek V3
Best Open-Source for CodingDeepSeek
Frontier-competitive coding and reasoning at open-weight. Significant GPU requirements but zero licensing cost. Via API it's already among the cheapest capable models.
$0 self-hosted / $0.27/1M via API
Try DeepSeek V3 →Mistral Large
Best EU Open-SourceMistral
Open-weight, EU-based, GDPR-friendly. Good balance of quality and size. Runs on single-node setups with sufficient VRAM.
$0 self-hosted / $2/1M via API
Try Mistral Large →Frequently Asked Questions
What hardware do I need to self-host LLMs?
For 70B parameter models: 2x A100 (80GB) or equivalent. For smaller models (7B-13B): a single RTX 4090 or A10G. Cloud options (Lambda Labs, RunPod) are often cheaper than owned hardware for intermittent use.
Is self-hosting actually cheaper?
For consistent high-volume usage: yes, dramatically. For intermittent or low-volume use: usually no — managed APIs win on cost when amortizing GPU rental over light usage.
Not sure which is right for you?
Use the calculator to estimate your real cost, or take the decision quiz.
Related
Pricing based on publicly available rates. Check current provider pricing before subscribing. Some links may be affiliate links — see our affiliate disclosure.