Open-weight AI models: which licences let a business use them freely?
Yes, you can run open-weight AI models commercially and even sell products built on them. But the licence decides whether that is routine or a contract-law minefield. Here is how to tell the difference before you commit your GPUs, your roadmap and your customer contracts to a model family.
Apache 2.0 and MIT models, which now dominate new releases, can be run, fine-tuned, embedded in products and sold on with almost no strings. The trouble is the rest: Llama, Gemma 3 and a cluster of custom community licences carry naming rules, use policies, pass-through obligations and user or revenue thresholds that your engineering team will never notice until a customer, auditor or acquirer asks about them.
The distinction matters most for the organisations doing exactly what these licences affect: running AI on your own infrastructure and building products on top of it. If you only ever call a hosted API, the model licence is the provider's problem. The moment you download weights onto your own GPUs, or wrap a model in your own software, you are the licensee, and the terms bind you directly.
This article sorts today's common open-weight models into three buckets (safe, check-first, avoid) for a company planning to self-host or productise AI, and lists what to record so the next audit, enterprise sale or acquisition passes without surprises.
Why "open" is doing a lot of work
"Open weights" means one thing only: the trained model files are downloadable and runnable by you. It says nothing about what you may do with them. The Open Source Initiative has been explicit that Meta's Llama licence, the most famous "open" model family, is not open source at all, and the same is true of several other widely used families.
The market has been moving in the right direction. The 2026 licence landscape analysis from Presenc AI puts Apache 2.0 at roughly 38 percent of new open-weight releases on Hugging Face, with MIT at around 18 percent. That means more than half of new releases are genuinely permissive. But the restricted licences punch above their weight in downloads: Llama's community licence covers about 14 percent of new releases but a much larger share of actual use, and Gemma, Tongyi Qianwen, OpenRAIL and research-only variants fill out the rest.
There is also a structural difference in how the two categories work, which WCR Legal's comparison of OSS and AI model licences lays out well. Apache 2.0 and MIT are copyright licences: permanent, irrevocable, with terms fixed forever on publication. The custom community licences are contracts: they can be updated unilaterally, they can terminate on breach, they restrict what your product may do, and they can oblige you to pass restrictions on to your own customers. Two models carrying the same "open" label can sit in entirely different legal categories.
The two licences you can treat as free
Apache 2.0 and MIT are the only two widely used model licences that behave like the open-source licences your legal team already has a workflow for. Both permit commercial use, fine-tuning, redistribution and sublicensing with no field-of-use restrictions, no scale thresholds and no termination rights. Apache 2.0 adds an explicit patent grant and a patent-retaliation clause, which is why procurement teams generally prefer it over MIT for model weights, where training and inference methods may be patent-encumbered. You can read the full Apache 2.0 text on Google's Gemma licence page; the conditions are just: keep the copyright notices, include the licence text when you distribute, state that you modified files, and include any NOTICE file.
For a business, the practical upshot is: these models require no legal review beyond your standard OSS process. Run them, fine-tune them, embed them in a paid product, rebrand the output, keep your modifications closed. Nothing to pass through to your customers, nothing to disclose as a contingent liability to investors.
Today's mainstream safe models
The 2026 self-hosting landscape gives you real choice within these two licences (Context Studios' open-weight LLM guide has the current picture):
| Model family | Licence | Notable facts |
|---|---|---|
| Qwen3.5 (Alibaba) | Apache 2.0 | One architecture from 0.8B to 122B-A10B, 262K context extensible to about 1M |
| gpt-oss-120b / 20b (OpenAI) | Apache 2.0 | 117B params, 5.1B active; the 120b fits on a single 80GB GPU, the 20b in 16GB |
| Gemma 4 (Google DeepMind) | Apache 2.0 | Five sizes from E2B to 31B, 256K context. Note: this is a change from Gemma 1 to 3 |
| Mistral Small 4 / Large 3 | Apache 2.0 | 119B/6.5B active and 675B respectively; Mistral's flagship line is now permissively licensed |
| GLM-5.2 (Z.ai) | MIT | Strongest self-hostable coding model class, 1M-token context |
| DeepSeek-V4 | MIT | 1.6T/49B-active Pro and 284B/13B Flash, 1M context |
One nuance: these are family-level licences, and some labs mix licences within a family. Alibaba's small Qwen checkpoints are mostly Apache 2.0, but some larger or specialised variants historically shipped under the Tongyi Qianwen licence with a 100 million monthly active user threshold. Always check the licence tag on the exact model repository you are downloading, not the family blog post.
The middle ground: usable, but read the strings
These licences permit commercial use in most cases, but each carries conditions that belong in a legal review, not a Slack thread.
Llama Community Licence (Llama 3.x and 4)
The Llama 4 Community License Agreement (effective 5 April 2025) is the most consequential of the restricted licences. Its conditions:
- The 700 million MAU clause. If the products or services you make available had more than 700 million monthly active users in the preceding calendar month, you must request a separate licence from Meta, which Meta may grant "in its sole discretion". Almost no company will ever hit this. The problem is that it converts the licence into a conditional grant: investors and acquirers flag it as an unquantified dependency on Meta's goodwill, and if you are ever acquired by a platform that already exceeds the threshold, your product's right to use Llama evaporates with it.
- Attribution and naming. You must prominently display "Built with Llama" on websites, interfaces and product documentation. If you use Llama or its outputs to create, train or fine-tune another AI model that you distribute, the name must begin with "Llama".
- An Acceptable Use Policy incorporated by reference. The AUP is part of the licence, and it prohibits uses with no open-source equivalent. The one that catches commercial teams most often: you may not use Llama for "the unauthorised or unlicensed practice of any profession", explicitly including financial, legal and medical practice. An AI assistant that dispenses financial, legal or health guidance, built by a company without the relevant professional licences, is arguably outside the licence entirely.
- The no-improvement clause. You may not use Llama's outputs or fine-tunes to improve any model that is not itself a Llama derivative. If your roadmap includes distilling one model's outputs into a smaller in-house model, Llama is off the table for that pipeline.
- Termination on breach. Meta can terminate the agreement if you breach any term, and on termination you must delete the model.
Also note the governing structure: EEA-based entities contract with Meta Platforms Ireland, everyone else with Meta Platforms Inc., under Californian law.
Gemma Terms of Use (Gemma 1 through 3)
Google's Gemma Terms of Use applies to Gemma 1, 2, 3 and their variants. Two provisions make it a genuinely different category from Apache 2.0:
- Flow-down obligations. "Distribution" under the terms includes hosting the model via API or web access, and you must include the use restrictions as an enforceable provision in any agreement governing downstream use. If you build a product on Gemma 3, your customers' terms of service must carry Google's Prohibited Use Policy forward. That is a contractual monitoring obligation, not an attribution checkbox.
- Remote restriction and termination rights. Google "reserves the right to restrict (remotely or otherwise) usage" of Gemma services it believes violate the agreement, and may terminate the licence on breach, after which you must delete the model and derivatives. The terms can also be updated unilaterally, with continued use constituting acceptance.
The Model Derivatives definition is unusually broad: it captures any model created by "transfer of patterns of the weights, parameters, operations, or Output of Gemma", including distillation, so the restrictions follow your derivative works, not just the original checkpoint. On the plus side, Google expressly claims no rights in outputs you generate.
The important 2026 development: Gemma 4 moved to Apache 2.0. Google's own terms page now excludes Gemma 4 from the Gemma ToU and points to its own Apache 2.0 licence. If you standardised on Gemma 3 and assumed Gemma 4 was the same legal animal, it is not; the reverse mistake is worse, since old Gemma 3 checkpoints in your estate are still under the restrictive terms.
Mistral's modified MIT
Mistral's own licensing FAQ is admirably clear: most models are Apache 2.0, but certain models ship under a modified MIT licence that adds a single condition. Companies with monthly revenue exceeding $20 million must either obtain a commercial licence from Mistral or use the models via Mistral Studio. That is a scale cliff worth knowing about: your use is free and clear until a specific month's revenue crosses the line, and then it is not. Check which licence governs the specific checkpoint before you build on it.
Other check-first cases
- Tongyi Qianwen Licence (larger Qwen variants). Commercial use permitted with a 100 million MAU threshold and restrictions on building competitive AI services. If your product is itself an AI service platform, get counsel involved before relying on a Tongyi-licensed checkpoint.
- NVIDIA Open Model License. Hardware-agnostic and output-friendly, but no patent grant for model methods and a restriction on training competing foundation models.
- Kimi K2.7 Code (Modified MIT) and MiniMax-M3 (vendor licence): both usable commercially in the typical case, but they carry conditions that require reading before you commit a product to them.
- OpenRAIL-M (Stable Diffusion family): commercial use permitted, but use restrictions propagate contractually to every downstream recipient, similar in spirit to Gemma's flow-down.
The no-go list for commercial use
CC-BY-NC and other non-commercial variants are research-only. Roughly 9 percent of new releases still carry them, and they are common among embedding models (NV-Embed-v2, SFR-Embedding-Mistral, Linq-Embed-Mistral). Deploying one behind a commercial RAG product is unlicensed use. The fix is usually substitution: for embeddings, an Apache 2.0 or MIT alternative nearly always exists (the BGE family is MIT, for instance). Non-commercial checkpoints inside an AI bill of materials are a common compliance gap when auditing model stacks.
A quick worked example: screening models for a UAE product
Scenario: a Dubai property-management group with 200,000 tenants wants two AI services on its own GPU servers: an internal assistant for staff, and a tenant-facing assistant that answers questions about leases, service charges and maintenance. Three model families are shortlisted.
Step 1: inventory the actual checkpoints. The engineering team's shortlist is Llama 4 (for quality), Gemma 3 12B (for the tenant app, cheap to serve), and Qwen3.5-35B-A3B (a middle option). Each checkpoint's licence is pulled from its repository, not from memory: Llama 4 Community License, Gemma Terms of Use, Apache 2.0.
Step 2: map the use cases against the terms.
- Llama 4: the 700M MAU clause is irrelevant at their scale, but the naming rule ("Built with Llama" displayed, "Llama" prefix on any fine-tune they distribute) applies to the tenant-facing product, and the AUP bites in one place: staff also want the internal assistant to draft replies on service-charge disputes, which drifts toward unlicensed financial advice. The AUP makes that use case a licence risk, not just a regulatory one.
- Gemma 3: the flow-down obligation means the tenant app's terms of use must incorporate Google's Prohibited Use Policy as an enforceable provision, and the app's roadmap includes third-party access, which widens the compliance surface. The termination and remote-restriction rights sit badly with a service that tenants depend on for maintenance requests.
- Qwen3.5-35B-A3B: Apache 2.0. No use restrictions, no attribution beyond the NOTICE file, no pass-through, no termination. Fine-tune it, embed it, rebrand the product, say nothing on your website if you do not want to.
Step 3: decide. Qwen3.5-35B-A3B wins for both services on licence grounds, with gpt-oss-120b (also Apache 2.0) qualified as the fallback on the single 80GB GPU already owned. Llama 4 is cleared for the internal assistant only, with the AUP boundary documented: no legal, financial or health guidance without licensed professionals in the loop. Gemma 3 is dropped: the flow-down and termination terms cost more legal effort than the model saves.
Step 4: record it. The screen takes an afternoon. The record is what makes it worth something two years later in due diligence.
What to record for the next auditor, buyer or investor
The compliance burden of restricted licences is mostly record-keeping. Five items cover it:
- A model licence registry, maintained like a software bill of materials. For each checkpoint: model name, exact version, licence and licence URL, download date, and any conditions flagged.
- Snapshots of licence text at download time. Custom licences can change unilaterally (the Gemma terms page carries a "last modified" date for exactly this reason). What binds you is the version you accepted, and the only way to prove that is to have saved it.
- AUP and prohibited-use mapping for every custom-licensed model: each current and roadmap use case checked against the policy, with the analysis signed off.
- Flow-down confirmation: for Gemma-style licences, evidence that customer agreements carry the required restrictions.
- Threshold register: Llama's 700M MAU, Mistral's $20M monthly revenue, Qwen's 100M MAU, each listed as a contingent item with a note on reachability.
A minimal registry entry, kept in version control alongside your deployment code, looks like this:
# model-registry.yaml
models:
- name: Qwen3.5-35B-A3B
version: "2026.03"
license: Apache-2.0
license_url: https://huggingface.co/Qwen/Qwen3.5-35B-A3B/blob/main/LICENSE
downloaded: 2026-09-14
license_snapshot: registry/licenses/qwen3.5-35b-apache2.txt
conditions: [notice-file-on-distribution]
approved_uses: [internal-assistant, tenant-assistant]
reviewed_by: platform-team
review_date: 2026-09-15
- name: Llama-4-Scout
version: "17B-16E"
license: Llama-4-Community-License
license_url: https://www.llama.com/llama4/license/
downloaded: 2026-09-14
license_snapshot: registry/licenses/llama4-community-2025-04-05.txt
conditions:
- display-built-with-llama
- llama-prefix-on-distributed-finetunes
- aup-incorporated: no-unlicensed-professional-practice
- threshold: 700M-MAU
approved_uses: [internal-assistant]
excluded_uses: [financial-guidance, legal-guidance]
reviewed_by: platform-team
review_date: 2026-09-15
Two engineering practices make the registry real rather than aspirational. First, make the licence tag a deployment gate: your serving layer should refuse to load a model whose registry entry is missing or expired. Second, review the registry whenever you adopt a new checkpoint of the same family, because families change licences between generations, as Gemma 3 to Gemma 4 demonstrates.
Where data residency and regulation slot in
For GCC organisations, self-hosting open-weight models is usually driven by data residency first and licence second, but the two interact. Running weights on infrastructure you control, in a UAE data centre or fully air-gapped, keeps prompts and documents inside your legal boundary, which is the main reason regulated sectors here evaluate open weights before hosted APIs.
On the regulatory side, one point from the EU AI Act is worth knowing even outside Europe, because it applies to what you do with the model rather than where you are: if you fine-tune a model and distribute it, or make it available via API to EU users, you may step into the "provider" role for general-purpose AI, which carries documentation, copyright-policy and transparency obligations that the model's licence never mentioned. Those GPAI obligations have applied since 2 August 2025. Merely deploying a model internally does not. The licence gives you the right to build; the regulation tells you what you owe once you ship.
What to do next
- Audit the model stack you already run. Pull the actual licence file for every checkpoint in production, including embeddings and rerankers, which is where non-commercial licences hide. Tag each entry: Apache/MIT, custom-licensed, or non-commercial.
- Screen every custom licence against your product roadmap, not just today's use cases. The three questions that matter: does the AUP exclude any planned feature, do the terms bind your customers, and does any threshold or termination right create business-continuity risk?
- Prefer Apache 2.0 and MIT at equal capability. In 2026 there is a permissively licensed model at every size tier, from Qwen3.5 small variants on a laptop to DeepSeek-V4 on a cluster, so accepting custom-licence terms should be a deliberate choice justified by capability, not a default.
- Put a gateway in front of the models. Keeping a stable OpenAI-compatible endpoint between your applications and the model servers makes swapping families, including swapping out a licence problem, a configuration change rather than a rewrite.
- Formalise the registry with the five record types above, and assign an owner.
If you are planning the platform itself, GPUs, serving, routing and the governance around it, that is the work our AI infrastructure team does: designing, building and running the model platform on your own infrastructure, with one gateway routing every model and keeping access control in a single place. The licence audit in this article is a good first artefact to bring to that conversation, and the registry format above will already be halfway to what a platform needs.
Link to this article
Citing this in your own writing? Use the permanent link below.https://www.azrty.com/blog/open-weight-ai-models-which-licences-let-a-business-use-them-freely
<a href="https://www.azrty.com/blog/open-weight-ai-models-which-licences-let-a-business-use-them-freely">Open-weight AI models: which licences let a business use them freely?</a> (Azrty)
