Skip to content

Open-Weight LLMs for Business: Mistral, Gemma, gpt-oss, Llama

Which open-weight LLMs can a business actually run in 2026? Mistral, Gemma 4, gpt-oss and Llama compared on licenses, sizes and when self-hosting makes sense.

Guides5 min read
By the AI App Hunters editors · Updated 1 Oct 2026

For most businesses, a hosted AI service is still the right default. Open-weight models make sense when your data cannot leave your environment, when volume is high and steady, or when you need to fine-tune. In 2026 the best-licensed options come from Mistral, Google (Gemma 4) and OpenAI (gpt-oss), all under Apache 2.0 for at least some models.

"Open source" is used loosely here. Strictly, most of these are open-weight: you can download and run the model, but you do not get the training data.

Key takeaways

  • Apache 2.0 licensing is now common among open models. That is good news for commercial use.
  • Model names move fast. Check the vendor page for the exact version and license before you build on it.
  • Self-hosting adds GPUs, engineering and monitoring. It is not free just because the weights are.
  • A hybrid is common: a hosted frontier model for hard tasks, a small open model for private or high-volume ones.
  • Meta's Llama is still widely used, but I could not confirm a newer release than Llama 4 from the pages I checked.

What counts as an open LLM today?

Three levels, from most to least open:

  1. Open source in full: weights, code and data. Rare.
  2. Open-weight, permissive license: download, run, fine-tune, sell products. Apache 2.0 is the example.
  3. Open-weight, custom license: you can download, but the vendor's own terms apply. Read them.

Your lawyer cares about this difference. Your engineers care about whether it runs on the hardware you have.

Which open models are worth looking at?

Family Recent models (per vendor pages) License noted Notes
Mistral Large 3, Small 4, Ministral 3 (14B, 8B, 3B) Apache 2.0 Medium 3.5 is listed under a modified MIT license
Gemma (Google) Gemma 4: E2B, E4B, 12B, 26B, 31B Apache 2.0 for Gemma 4 Earlier Gemma versions use Gemma Terms of Use
gpt-oss (OpenAI) gpt-oss-120b, gpt-oss-20b Apache 2.0 Released August 5, 2025
Llama (Meta) Llama 4 Scout and Maverick are the latest I could confirm Meta's Llama license Requires accepting terms to download
DeepSeek DeepSeek-Flash, DeepSeek-V4-Pro Not confirmed on pages I read Mostly used through its API

Mistral

Mistral's documentation lists Mistral Large 3 (a multimodal open-weight model, Apache 2.0), Mistral Small 4 (instruct, reasoning and coding in one model, Apache 2.0) and the Ministral 3 series in 14B, 8B and 3B sizes. Mistral Medium 3.5 is listed under a modified MIT license. Specialist models cover audio, OCR, code and safety, and some carry different licenses, so check each.

Mistral also sells a hosted assistant. As of October 2026 its pricing page lists a Free plan, Pro at $14.99 per month, Team at $24.99 per user per month with a $50 monthly minimum, and Enterprise with custom pricing. API prices are per million tokens, and the page gives Mistral Large at $0.5 input and $1.5 output.

That makes Mistral a useful middle path: use the API now, move to self-hosting later if volume justifies it.

Gemma 4

Google DeepMind describes Gemma 4 as its most capable open family. Sizes run from tiny E2B and E4B for mobile and edge devices to 12B, 26B and 31B for personal computers. Google's terms page says Gemma 4 has its own Apache 2.0 license, unlike earlier Gemma models. Good news if you assumed all Gemma versions share one license.

gpt-oss

OpenAI released gpt-oss-120b and gpt-oss-20b on August 5, 2025 under Apache 2.0. OpenAI says the 120b model performs near o4-mini and runs on an 80GB GPU. The 20b model matches o3-mini on its tests and runs on a 16GB device. Both support 128k context. These are OpenAI's own benchmark claims. Run your own evaluation.

Llama

The Meta Llama GitHub repository lists Llama 4 (Scout and Maverick, released April 5, 2025) as its newest entry, and the Hugging Face organization shows the same. Access requires accepting Meta's license. Meta's current developer site leads with its Muse models, and the page I read did not say whether those are open-weight. If Llama is on your shortlist, read the license text for the exact version.

DeepSeek

DeepSeek is mainly consumed as an API. Its docs list DeepSeek-Flash and DeepSeek-V4-Pro, with prices that vary by cache status and time of day, and context up to 1M tokens. The pages I read did not state weights or license terms, so check before treating it as self-hostable. Our DeepSeek directory page has more.

When does self-hosting make sense?

Good reasons:

  • Data control: regulated data, client confidentiality, or customers who forbid third-party processing.
  • Steady volume: predictable, high token usage where per-token fees add up.
  • Customization: fine-tuning on your own formats and terminology.
  • Offline or edge use: devices that cannot rely on a connection.

Weak reasons:

  • "It's free." The weights are. The GPUs, on-call time and security reviews are not.
  • "It will be as good as the best hosted model." Sometimes close, often not. Test it.

For everyday work, hosted assistants such as Claude and ChatGPT are simpler and usually the better first step.

How do you evaluate an open model for your business?

  1. Write 20 real tasks from your own work, with known good answers.
  2. Run them on a hosted API version of the candidate before you buy hardware.
  3. Score quality, speed and cost per task.
  4. Read the license for the exact version, including any naming, attribution or usage-size clauses.
  5. Price the full stack: GPUs or cloud rental, serving software, monitoring, and the engineer who owns it.
  6. Plan for model churn. Names and versions change quickly, so keep your prompts and evaluation set portable.

If you plan to wire a model into workflows, see n8n vs Zapier vs Make for the automation layer, and AI coding tools if engineers will do the integration.

Bottom line

Start hosted. Move to an open-weight model when you have a clear reason: privacy, volume or fine-tuning. Shortlist Mistral, Gemma 4 and gpt-oss first, because their Apache 2.0 licensing keeps legal review short. Treat Llama as a "read the license" option. Whatever you pick, test on your own tasks, not on leaderboards.

Frequently asked questions

What is the difference between open source and open weight?+

Open-weight means you can download the trained model and run it yourself. Open source in the strict sense also includes training code and data, which most releases do not share.

Which open LLM has the most permissive license?+

Apache 2.0 is the cleanest. Mistral Large 3, Mistral Small 4, Gemma 4 and OpenAI's gpt-oss models are listed under Apache 2.0. Always read the license file for the exact version you deploy.

Is self-hosting an open model cheaper than using an API?+

Only at steady, high volume or when data cannot leave your environment. For most SMBs, a hosted API is cheaper once you count GPUs, engineering time and monitoring.

Can a business use Llama commercially?+

Llama is released under Meta's own license, and access requires accepting its terms. Read those terms for the version you plan to use. I could not confirm which Llama release is newest today.

What hardware does a small open model need?+

OpenAI says gpt-oss-20b runs on a 16GB device and gpt-oss-120b on an 80GB GPU. Smaller models like Gemma 4 E2B and E4B target phones and edge devices.

Sources, checked 1 Oct 2026

  1. Mistral models overview
  2. Mistral pricing
  3. OpenAI: Introducing gpt-oss
  4. Google DeepMind: Gemma
  5. Gemma terms of use and Gemma 4 license note
  6. Meta Llama models on GitHub
  7. Meta Llama on Hugging Face
  8. Meta AI developer platform
  9. DeepSeek API docs
  10. DeepSeek pricing

Keep reading

App of the Week6 min read

App of the Week: Attio, the AI-native CRM

Attio is an AI-native CRM with agents, workflows and an MCP server. What it does, who it suits, October 2026 pricing, limits and the best alternatives.

Attio logo
App of the Week5 min read

App of the Week: HeyGen, AI Avatar Video

HeyGen makes avatar videos from text and translates video into 175+ languages. What it does, October 2026 pricing, limits and the best alternatives.

HeyGen logo
App of the Week4 min read

App of the Week: Profound, AI Search Visibility

Profound tracks how your brand shows up in ChatGPT, Claude, Perplexity and Gemini, then runs agents to improve it. Features, pricing, limits, alternatives.

Profound logo