Open-weight AI models: which licences let a business use them freely?

Yes, you can run open-weight AI models commercially and even sell products built on them. But the licence decides whether that is routine or a contract-law minefield. Here is how to tell the difference before you commit your GPUs, your roadmap and your customer contracts to a model family.

Apache 2.0 and MIT models, which now dominate new releases, can be run, fine-tuned, embedded in products and sold on with almost no strings. The trouble is the rest: Llama, Gemma 3 and a cluster of custom community licences carry naming rules, use policies, pass-through obligations and user or revenue thresholds that your engineering team will never notice until a customer, auditor or acquirer asks about them.

The distinction matters most for the organisations doing exactly what these licences affect: running AI on your own infrastructure and building products on top of it. If you only ever call a hosted API, the model licence is the provider's problem. The moment you download weights onto your own GPUs, or wrap a model in your own software, you are the licensee, and the terms bind you directly.

This article sorts today's common open-weight models into three buckets (safe, check-first, avoid) for a company planning to self-host or productise AI, and lists what to record so the next audit, enterprise sale or acquisition passes without surprises.

Why "open" is doing a lot of work

"Open weights" means one thing only: the trained model files are downloadable and runnable by you. It says nothing about what you may do with them. The Open Source Initiative has been explicit that Meta's Llama licence, the most famous "open" model family, is not open source at all, and the same is true of several other widely used families.

The market has been moving in the right direction. The 2026 licence landscape analysis from Presenc AI puts Apache 2.0 at roughly 38 percent of new open-weight releases on Hugging Face, with MIT at around 18 percent. That means more than half of new releases are genuinely permissive. But the restricted licences punch above their weight in downloads: Llama's community licence covers about 14 percent of new releases but a much larger share of actual use, and Gemma, Tongyi Qianwen, OpenRAIL and research-only variants fill out the rest.

There is also a structural difference in how the two categories work, which WCR Legal's comparison of OSS and AI model licences lays out well. Apache 2.0 and MIT are copyright licences: permanent, irrevocable, with terms fixed forever on publication. The custom community licences are contracts: they can be updated unilaterally, they can terminate on breach, they restrict what your product may do, and they can oblige you to pass restrictions on to your own customers. Two models carrying the same "open" label can sit in entirely different legal categories.

The two licences you can treat as free

Apache 2.0 and MIT are the only two widely used model licences that behave like the open-source licences your legal team already has a workflow for. Both permit commercial use, fine-tuning, redistribution and sublicensing with no field-of-use restrictions, no scale thresholds and no termination rights. Apache 2.0 adds an explicit patent grant and a patent-retaliation clause, which is why procurement teams generally prefer it over MIT for model weights, where training and inference methods may be patent-encumbered. You can read the full Apache 2.0 text on Google's Gemma licence page; the conditions are just: keep the copyright notices, include the licence text when you distribute, state that you modified files, and include any NOTICE file.

For a business, the practical upshot is: these models require no legal review beyond your standard OSS process. Run them, fine-tune them, embed them in a paid product, rebrand the output, keep your modifications closed. Nothing to pass through to your customers, nothing to disclose as a contingent liability to investors.

Today's mainstream safe models

The 2026 self-hosting landscape gives you real choice within these two licences (Context Studios' open-weight LLM guide has the current picture):

Model familyLicenceNotable facts
Qwen3.5 (Alibaba)Apache 2.0One architecture from 0.8B to 122B-A10B, 262K context extensible to about 1M
gpt-oss-120b / 20b (OpenAI)Apache 2.0117B params, 5.1B active; the 120b fits on a single 80GB GPU, the 20b in 16GB
Gemma 4 (Google DeepMind)Apache 2.0Five sizes from E2B to 31B, 256K context. Note: this is a change from Gemma 1 to 3
Mistral Small 4 / Large 3Apache 2.0119B/6.5B active and 675B respectively; Mistral's flagship line is now permissively licensed
GLM-5.2 (Z.ai)MITStrongest self-hostable coding model class, 1M-token context
DeepSeek-V4MIT1.6T/49B-active Pro and 284B/13B Flash, 1M context

One nuance: these are family-level licences, and some labs mix licences within a family. Alibaba's small Qwen checkpoints are mostly Apache 2.0, but some larger or specialised variants historically shipped under the Tongyi Qianwen licence with a 100 million monthly active user threshold. Always check the licence tag on the exact model repository you are downloading, not the family blog post.

The middle ground: usable, but read the strings

These licences permit commercial use in most cases, but each carries conditions that belong in a legal review, not a Slack thread.

Llama Community Licence (Llama 3.x and 4)

The Llama 4 Community License Agreement (effective 5 April 2025) is the most consequential of the restricted licences. Its conditions:

Also note the governing structure: EEA-based entities contract with Meta Platforms Ireland, everyone else with Meta Platforms Inc., under Californian law.

Gemma Terms of Use (Gemma 1 through 3)

Google's Gemma Terms of Use applies to Gemma 1, 2, 3 and their variants. Two provisions make it a genuinely different category from Apache 2.0:

The Model Derivatives definition is unusually broad: it captures any model created by "transfer of patterns of the weights, parameters, operations, or Output of Gemma", including distillation, so the restrictions follow your derivative works, not just the original checkpoint. On the plus side, Google expressly claims no rights in outputs you generate.

The important 2026 development: Gemma 4 moved to Apache 2.0. Google's own terms page now excludes Gemma 4 from the Gemma ToU and points to its own Apache 2.0 licence. If you standardised on Gemma 3 and assumed Gemma 4 was the same legal animal, it is not; the reverse mistake is worse, since old Gemma 3 checkpoints in your estate are still under the restrictive terms.

Mistral's modified MIT

Mistral's own licensing FAQ is admirably clear: most models are Apache 2.0, but certain models ship under a modified MIT licence that adds a single condition. Companies with monthly revenue exceeding $20 million must either obtain a commercial licence from Mistral or use the models via Mistral Studio. That is a scale cliff worth knowing about: your use is free and clear until a specific month's revenue crosses the line, and then it is not. Check which licence governs the specific checkpoint before you build on it.

Other check-first cases

The no-go list for commercial use

CC-BY-NC and other non-commercial variants are research-only. Roughly 9 percent of new releases still carry them, and they are common among embedding models (NV-Embed-v2, SFR-Embedding-Mistral, Linq-Embed-Mistral). Deploying one behind a commercial RAG product is unlicensed use. The fix is usually substitution: for embeddings, an Apache 2.0 or MIT alternative nearly always exists (the BGE family is MIT, for instance). Non-commercial checkpoints inside an AI bill of materials are a common compliance gap when auditing model stacks.

A quick worked example: screening models for a UAE product

Scenario: a Dubai property-management group with 200,000 tenants wants two AI services on its own GPU servers: an internal assistant for staff, and a tenant-facing assistant that answers questions about leases, service charges and maintenance. Three model families are shortlisted.

Step 1: inventory the actual checkpoints. The engineering team's shortlist is Llama 4 (for quality), Gemma 3 12B (for the tenant app, cheap to serve), and Qwen3.5-35B-A3B (a middle option). Each checkpoint's licence is pulled from its repository, not from memory: Llama 4 Community License, Gemma Terms of Use, Apache 2.0.

Step 2: map the use cases against the terms.

Step 3: decide. Qwen3.5-35B-A3B wins for both services on licence grounds, with gpt-oss-120b (also Apache 2.0) qualified as the fallback on the single 80GB GPU already owned. Llama 4 is cleared for the internal assistant only, with the AUP boundary documented: no legal, financial or health guidance without licensed professionals in the loop. Gemma 3 is dropped: the flow-down and termination terms cost more legal effort than the model saves.

Step 4: record it. The screen takes an afternoon. The record is what makes it worth something two years later in due diligence.

What to record for the next auditor, buyer or investor

The compliance burden of restricted licences is mostly record-keeping. Five items cover it:

  1. A model licence registry, maintained like a software bill of materials. For each checkpoint: model name, exact version, licence and licence URL, download date, and any conditions flagged.
  2. Snapshots of licence text at download time. Custom licences can change unilaterally (the Gemma terms page carries a "last modified" date for exactly this reason). What binds you is the version you accepted, and the only way to prove that is to have saved it.
  3. AUP and prohibited-use mapping for every custom-licensed model: each current and roadmap use case checked against the policy, with the analysis signed off.
  4. Flow-down confirmation: for Gemma-style licences, evidence that customer agreements carry the required restrictions.
  5. Threshold register: Llama's 700M MAU, Mistral's $20M monthly revenue, Qwen's 100M MAU, each listed as a contingent item with a note on reachability.

A minimal registry entry, kept in version control alongside your deployment code, looks like this:

# model-registry.yaml
models:
  - name: Qwen3.5-35B-A3B
    version: "2026.03"
    license: Apache-2.0
    license_url: https://huggingface.co/Qwen/Qwen3.5-35B-A3B/blob/main/LICENSE
    downloaded: 2026-09-14
    license_snapshot: registry/licenses/qwen3.5-35b-apache2.txt
    conditions: [notice-file-on-distribution]
    approved_uses: [internal-assistant, tenant-assistant]
    reviewed_by: platform-team
    review_date: 2026-09-15

  - name: Llama-4-Scout
    version: "17B-16E"
    license: Llama-4-Community-License
    license_url: https://www.llama.com/llama4/license/
    downloaded: 2026-09-14
    license_snapshot: registry/licenses/llama4-community-2025-04-05.txt
    conditions:
      - display-built-with-llama
      - llama-prefix-on-distributed-finetunes
      - aup-incorporated: no-unlicensed-professional-practice
      - threshold: 700M-MAU
    approved_uses: [internal-assistant]
    excluded_uses: [financial-guidance, legal-guidance]
    reviewed_by: platform-team
    review_date: 2026-09-15

Two engineering practices make the registry real rather than aspirational. First, make the licence tag a deployment gate: your serving layer should refuse to load a model whose registry entry is missing or expired. Second, review the registry whenever you adopt a new checkpoint of the same family, because families change licences between generations, as Gemma 3 to Gemma 4 demonstrates.

Where data residency and regulation slot in

For GCC organisations, self-hosting open-weight models is usually driven by data residency first and licence second, but the two interact. Running weights on infrastructure you control, in a UAE data centre or fully air-gapped, keeps prompts and documents inside your legal boundary, which is the main reason regulated sectors here evaluate open weights before hosted APIs.

On the regulatory side, one point from the EU AI Act is worth knowing even outside Europe, because it applies to what you do with the model rather than where you are: if you fine-tune a model and distribute it, or make it available via API to EU users, you may step into the "provider" role for general-purpose AI, which carries documentation, copyright-policy and transparency obligations that the model's licence never mentioned. Those GPAI obligations have applied since 2 August 2025. Merely deploying a model internally does not. The licence gives you the right to build; the regulation tells you what you owe once you ship.

What to do next

  1. Audit the model stack you already run. Pull the actual licence file for every checkpoint in production, including embeddings and rerankers, which is where non-commercial licences hide. Tag each entry: Apache/MIT, custom-licensed, or non-commercial.
  2. Screen every custom licence against your product roadmap, not just today's use cases. The three questions that matter: does the AUP exclude any planned feature, do the terms bind your customers, and does any threshold or termination right create business-continuity risk?
  3. Prefer Apache 2.0 and MIT at equal capability. In 2026 there is a permissively licensed model at every size tier, from Qwen3.5 small variants on a laptop to DeepSeek-V4 on a cluster, so accepting custom-licence terms should be a deliberate choice justified by capability, not a default.
  4. Put a gateway in front of the models. Keeping a stable OpenAI-compatible endpoint between your applications and the model servers makes swapping families, including swapping out a licence problem, a configuration change rather than a rewrite.
  5. Formalise the registry with the five record types above, and assign an owner.

If you are planning the platform itself, GPUs, serving, routing and the governance around it, that is the work our AI infrastructure team does: designing, building and running the model platform on your own infrastructure, with one gateway routing every model and keeping access control in a single place. The licence audit in this article is a good first artefact to bring to that conversation, and the registry format above will already be halfway to what a platform needs.

AI infrastructureopen-weight modelslicensingcommercial useself-hosting
Found this useful? Share it.

Link to this article

Citing this in your own writing? Use the permanent link below.
Permalink
https://www.azrty.com/blog/open-weight-ai-models-which-licences-let-a-business-use-them-freely
HTML
<a href="https://www.azrty.com/blog/open-weight-ai-models-which-licences-let-a-business-use-them-freely">Open-weight AI models: which licences let a business use them freely?</a> (Azrty)
Get a readiness assessmentOne call to find where AI will pay off in your business.
Related
Open-weight AI models: which licences let a business use them freely? | Azrty