← all articles

The best open-source LLMs you can self-host in 2026

In January 2025, DeepSeek released R1 and within about a week the news had wiped close to $600 billion off Nvidia’s market value in a single trading day, one of the largest one-day losses for any company in US stock market history. Reuters and half the financial press covered it as an AI bubble story. I read it as something else: proof that a lab most people hadn’t heard of a year earlier could put out weights good enough to make a $3 trillion company’s investors nervous, and you could download those weights and run them yourself.

This list is for people who actually plan to run the model, not just read about it. If you’re comparing API pricing between OpenAI and Anthropic, that’s a different article, and I wrote about what changes when you switch providers over in what changes when you move from one provider to another. This one is for the case where the answer to “which model” also has to answer “on whose GPU, under what license, and what happens when it breaks at 2am with no vendor support line to call.” I keep a running set of these comparisons on the blog if this is the kind of thing you come back for.

I picked these seven after spending the past few months running quantized versions of most of them through Ollama and vLLM on rented A100 and H100 boxes, for content pipelines and a couple of client projects where sending text to a third-party API wasn’t an option. Some of these I’d deploy again tomorrow. One I’d only touch for a narrow reason I’ll explain at the end.

how I picked

  • the license actually allows commercial self-hosting without a revenue cap or a call to a sales team, not just an “open” name on the tin
  • weights are downloadable from Hugging Face or the vendor’s own repo, no waitlist, no research-access application form
  • it runs on hardware a small team can rent or own: I drew the line around what fits on one or two 80GB GPUs at 4-bit or 8-bit quantization
  • there’s been a real release or fine-tune in the last 12 months, not an abandoned 2023 checkpoint
  • it works out of the box with at least one of Ollama, vLLM, llama.cpp, or Hugging Face’s TGI
  • I could find real benchmark numbers or my own test results, not just a vendor blog post claiming state of the art

the picks

Llama 4 (Meta)

Meta’s Llama family is still the default answer when someone asks which open model to run, and Llama 4 (Scout and Maverick, both released April 2025) is why. Both use a mixture-of-experts architecture, so you’re not paying full compute for every token, and Scout fits on a single H100 at native precision if you’re patient with context length. I’ve run Maverick through vLLM for a document classification job and it held up fine against GPT-4o class output on anything that wasn’t reasoning-heavy.

The catch is the license. The Llama Community License is generous until you cross 700 million monthly active users, at which point you need a separate agreement from Meta. Almost nobody reading this will hit that number, but it means Llama isn’t OSI-approved open source, whatever Meta’s marketing calls it.

  • huge ecosystem: every serving framework, quantization tool, and fine-tuning library supports Llama on day one
  • mixture-of-experts variants keep inference cost down relative to dense models of similar quality
  • extensive community fine-tunes on Hugging Face if the base instruct model doesn’t fit your use case

  • license isn’t OSI-recognized open source, and the 700M MAU clause is a real, if rare, ceiling

  • Maverick’s full weights are large enough that most self-hosters end up quantizing, which costs some quality

Pricing: free to download and run under the Llama Community License. Your cost is compute, figure on an H100 or two for anything past the 8B variant.

Link: Meta’s Llama page

DeepSeek V3 and R1

DeepSeek is the one that started the conversation above. V3 is a 671 billion parameter mixture-of-experts model with only about 37 billion parameters active per token, released under an MIT license, no usage cap, no revenue clause. R1 is the reasoning-tuned sibling that made headlines. Both are genuinely open: code, weights, and a technical paper explaining the training recipe, all public.

I haven’t run the full 671B weights myself. Almost nobody self-hosts that; the realistic path is a distilled version (DeepSeek published Qwen and Llama distillations at sizes from 7B to 70B) or renting inference from a host like Fireworks or Together rather than standing up the full model on your own hardware. The distilled 32B is the one I’ve actually put through a workload, and it’s a genuinely strong open reasoning model for its size.

  • MIT license end to end, the least restrictive terms on this list
  • distilled variants from 7B to 70B make the reasoning gains usable on hardware smaller teams actually have
  • training details published in enough depth that you can sanity check the claims instead of trusting a blog post

  • the full-size model is impractical to self-host for almost anyone, and some data governance teams still balk at a Chinese lab’s model for regulated data

  • R1’s reasoning traces before the final answer can run long, which shows up directly in your inference bill if you’re paying per token

Pricing: free under MIT license. Full V3/R1 self-hosting needs enterprise-grade multi-GPU clusters; the distilled models run on a single 80GB GPU.

Link: DeepSeek-V3 on GitHub

Qwen2.5 (Alibaba)

Qwen is Alibaba’s open model family, and it’s quietly become the one I reach for first on multilingual and coding tasks, ahead of Llama. Qwen2.5 shipped in sizes from 0.5B up to 72B under Apache 2.0, and the coder-specific variants beat comparably sized Llama models on the internal benchmarks I ran for a code-review assistant.

  • Apache 2.0 across the whole family, no revenue caps, no usage clauses to read twice
  • widest size range on this list, from phone-sized 0.5B models to 72B, so you match hardware instead of overbuying
  • strong multilingual coverage, including Chinese, which matters if any of your users aren’t English-first

  • documentation and community support skew toward Chinese-language forums, so English troubleshooting threads are thinner than Llama’s

Pricing: free, Apache 2.0. The 7B and 14B variants run comfortably on a single consumer GPU with 24GB of VRAM at 4-bit quantization.

Link: Qwen on GitHub

Mistral Small 3 and Mistral Large 2 (Mistral AI)

Mistral is the European answer, and for anyone whose self-hosting decision is driven by data residency rather than cost, that matters more than the benchmark scores. Mistral Small 3, released January 2025, is a 24B model Mistral specifically tuned to be fast on a single GPU, and it ships under Apache 2.0. Mistral Large 2 is the flagship, but it sits under a more restrictive license that requires a separate commercial agreement.

  • Small 3 is genuinely fast, and I’ve found it holds up for summarization and extraction tasks at a fraction of Large 2’s compute
  • an EU-based company simplifies the GDPR conversation if that’s the reason you’re self-hosting in the first place
  • Apache 2.0 on the smaller models means no licensing back-and-forth for commercial use

  • Large 2’s commercial license requires a separate agreement with Mistral, unlike the Apache-licensed smaller models

  • falls a step behind Qwen and Llama on raw performance at equivalent parameter counts in my own side-by-side tests

Pricing: Small 3 and the Mistral 7B/8x7B line are free under Apache 2.0. Large 2 needs a commercial license from Mistral for production use beyond evaluation.

Link: Mistral’s model docs

Gemma 3 (Google)

Gemma is Google’s open-weight line, built off the same research as Gemini but released under Google’s own usage terms rather than Apache or MIT. Gemma 3 comes in 1B, 4B, 12B, and 27B sizes, and the smaller ones are aimed squarely at running on a laptop or a single consumer GPU rather than a data center.

The 27B model is the one worth self-hosting seriously. It’s dense, not mixture-of-experts, so the VRAM math is simpler: expect to fit it on a single 24GB card at 4-bit quantization.

  • smallest models on this list (1B, 4B) actually run on a laptop, useful for edge or offline deployments
  • Google’s own instruction tuning is solid out of the box, less fine-tuning needed for straightforward chat or extraction tasks
  • good documentation and first-party support in Ollama’s model library

  • the Gemma license isn’t Apache or MIT, it carries Google-specific usage restrictions worth reading before you build a product on top of it

  • less of a track record on agentic or tool-calling workloads compared to Llama or Qwen in my testing

Pricing: free to download under Google’s Gemma license terms; compute cost scales down fast given the small model sizes.

Link: Google’s Gemma page

Phi-4 (Microsoft)

Phi-4 is Microsoft’s bet that a smaller model trained on carefully filtered, textbook-quality data can compete with models several times its size. It’s 14 billion parameters, MIT licensed, and in my testing on structured reasoning tasks like math word problems it punches well above its parameter count.

It’s not a generalist. I wouldn’t put Phi-4 in front of open-ended creative writing or long multi-turn conversations, that’s not what it was trained for.

  • MIT license, no restrictions at all
  • small enough to run on a single consumer GPU with room to spare, cheap to serve at scale
  • strong on math and structured reasoning specifically, which is a real gap in some larger general models

  • narrower training focus means it underperforms bigger generalist models on open-ended tasks

Pricing: free, MIT license. Runs on 16GB of VRAM or less at reasonable quantization.

Link: Phi-4 on Hugging Face

OLMo 2 (Allen Institute for AI)

OLMo is the odd one out on this list because it isn’t trying to win a benchmark race, it’s trying to be the model where nothing is hidden. Ai2, the Allen Institute for AI, publishes the weights, the training code, and the training data itself under Apache 2.0. Every other model on this list gives you weights and a paper describing the data. OLMo gives you the data.

That matters if your self-hosting decision is driven by compliance rather than cost or performance, if you need to answer “what exactly was this trained on” for a legal or academic reason. I wouldn’t pick OLMo 2 for raw output quality over Qwen or Llama at similar sizes. I’d pick it specifically when full auditability is the requirement, which is a smaller audience than the rest of this list, but a real one.

  • full training data release alongside weights and code, unmatched transparency on this list
  • Apache 2.0, non-profit steward, no commercial license negotiation ever
  • useful reference point for research or audit work where “what was this trained on” needs a real answer

  • trails Qwen, Llama, and DeepSeek on general benchmark performance at comparable sizes

Pricing: free, Apache 2.0, including the training data itself.

Link: OLMo at the Allen Institute for AI

comparison table

model price primary strength primary weakness
Llama 4 (Meta) free, Llama Community License biggest ecosystem and tooling support license caps out at 700M MAU, not OSI open source
DeepSeek V3 / R1 free, MIT best reasoning per dollar, fully open license full model impractical to self-host, distillation needed
Qwen2.5 (Alibaba) free, Apache 2.0 widest size range, strong coding and multilingual thinner English-language community support
Mistral Small 3 / Large 2 free (Small 3), commercial license (Large 2) fast inference, EU data residency Large 2 not free for production use
Gemma 3 (Google) free, Google usage terms smallest, laptop and edge-friendly non-permissive license, weaker on agentic tasks
Phi-4 (Microsoft) free, MIT best-in-class math and reasoning at 14B narrow, not a generalist model
OLMo 2 (Ai2) free, Apache 2.0 full training data transparency behind on raw benchmark performance

how to choose

Start with the reason you’re self-hosting at all, because it changes the answer. If it’s cost at scale, run the numbers on GPU rental against API pricing before you commit. A 70B model at 4-bit quantization needs roughly 40 to 48GB of VRAM, which is one A100 or two 3090s, and at current cloud rental prices that’s usually only cheaper than an API if you’re running enough volume to keep the box busy most of the day. If you’re only running it three hours a day, you’re paying for the other twenty-one hours of idle GPU too.

If the reason is data residency or a client contract that won’t let text leave your infrastructure, license terms matter more than benchmark scores. Mistral and OLMo are the cleanest stories here, an EU vendor on one side and full transparency on the other. That’s also the exact conversation covered in more depth over at theprivacywire.com/blog/, if data sovereignty is the actual driver behind the project rather than a side benefit.

Whatever you pick, budget time for the parts that aren’t the model. Quantization format, serving framework choice (vLLM for throughput, llama.cpp for a single box with no GPU, Ollama for getting something running today), and monitoring for drift and latency once it’s in production are where most of the real work goes. I’ve written separately about the observability side in the best LLM observability tools in 2026, and if you’re pairing any of these with retrieval, the embedding model matters as much as the LLM, which I cover in the best open source embedding models in 2026.

One more thing worth saying plainly: don’t self-host because it feels more serious than calling an API. If your volume is low and nobody on the team wants to own GPU capacity planning, an API is the correct choice and self-hosting is a distraction. Self-host when the math or the compliance requirement actually forces it, not because it sounds like the grown-up option.

verdict / top pick

If I had to run one model on one box tomorrow with no other context, I’d take Qwen2.5-72B. The license is clean, the size range means I’m not locked into one hardware tier, and it’s the model that’s held up best across the mixed bag of tasks, extraction, summarization, some code review, that actually make up most self-hosted workloads rather than the demos.

For pure reasoning per GPU-dollar, DeepSeek’s distilled 32B is the one I’d reach for next, and it’s the model I’m most likely to still be running a year from now given how fast that lab ships. Llama still wins if you need the ecosystem, the fine-tunes, and the safety of the biggest community having already solved your problem. And if an auditor is ever going to ask what your model was trained on, OLMo 2 is the only honest answer on this list.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-28.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →