Leadership ┬╖ 5 minute read
Open-Weight Models Explained for Executives
Open-weight models are models whose parameters can be downloaded and run in your own environment. "Open" describes availability rather than a single licence, and terms vary in ways that matter commercially. They win on control, data residency, and cost at steady high volume, and they need in-house operating capability.
Open-weight models changed what is possible for companies whose data cannot leave their environment, and the word "open" conceals differences that matter commercially. This explainer sets out what you actually get, the licensing questions to ask, where open weights win, and what they cost to operate.
What does open-weight actually mean?
That the model's trained parameters are published and can be downloaded and run on infrastructure you control. It usually does not mean the training data or the full training process are public, which is why "open-weight" is a more accurate term than "open source" for most of these models.
The practical benefit is control: the model runs where you decide, the data never leaves, and nobody can deprecate it out from under you. The open-weight models for regulated industries guide covers regulated use specifically.
What do the licences actually say?
They vary, and the differences are commercially significant:
| Licence question | Why it matters |
|---|---|
| Commercial use permitted? | Some licences restrict commercial use, or restrict it above a size threshold |
| Redistribution | Whether you may ship the model inside a product |
| Modification and fine-tuning | Whether derivatives are permitted and under what terms |
| Attribution | Whether you must name the model publicly |
| Acceptable-use restrictions | Prohibited applications, which may cover your sector |
| Output ownership | What the licence says about generated output |
Review the specific licence before building on a model, particularly if you intend to redistribute it in a product or operate in a restricted sector. Consult counsel; this is general guidance, not legal advice.
Where do open weights win?
Data that cannot leave. The decisive case. Regulated health, defense, certain financial and legal material, and data under customer contractual restrictions can make external APIs impossible regardless of terms. An open-weight model inside the environment makes the workload feasible.
Air-gapped operation. Environments with no external connectivity have no alternative.
Steady high volume. At sufficient sustained volume, infrastructure can cost less than per-token fees. Model this honestly including engineering time.
Specialized fine-tuning. Fine-tuning an open-weight model for a narrow, high-volume task can deliver the required quality at low per-task cost.
Longevity. A model you host cannot be deprecated by a provider, which matters for systems with long validation cycles, such as regulated manufacturing or medical contexts. The model deprecation risk management guide covers the hosted-model risk this avoids.
What does it cost?
Infrastructure, operations, and currency work. GPU capacity, platform engineering, monitoring, capacity planning, and the continuing job of evaluating and adopting newer models as they appear. At low or bursty volume this usually exceeds API fees; at steady high volume it can be substantially lower.
The estimate companies get wrong is staff time. A self-hosted model is an operated service, and someone must own it. Include that in the comparison. The private AI explained for executives piece covers the deployment spectrum.
Do they perform well enough?
For many enterprise tasks, yes. Classification, extraction, retrieval-grounded answering, and summarization of standard material are handled well by current open-weight models, particularly when the task is focused. Frontier hosted models generally lead on complex multi-step reasoning, breadth of knowledge, and difficult code.
The decision should be made by evaluation on your own tasks rather than by general benchmarks, which measure things that may not resemble your workload. Run the same evaluation set against hosted and open-weight candidates and compare pass rate, cost per task, and latency. The AI evaluation explained for executives piece covers the method.
How do they fit a mixed strategy?
Most enterprises end up mixed: hosted frontier models for complex work, hosted efficient models for moderate work, and open-weight models for sensitive data or high-volume specialized tasks. A gateway makes this manageable by routing per task and data class. The model routing explained for executives and AI gateways explained for executives pieces cover the mechanism.
What changes as the field moves?
Two things worth planning for. Capability gaps narrow: tasks that required a frontier model last year are handled by open weights this year, so the evaluation that justified a hosted model should be repeated annually rather than treated as settled. And operating cost falls as inference tooling improves and hardware becomes more available, which shifts the volume threshold at which self-hosting pays. Companies that treat the hosted-versus-self-hosted decision as permanent tend to be on the wrong side of it within two years, in one direction or the other.
What should executives ask?
- Which of our workloads genuinely cannot use hosted models, and on what basis?
- What does the licence of the model we are considering permit for our use?
- What is the total cost including engineering time at our real volume?
- What does evaluation show on our tasks, not on public benchmarks?
- Who operates and keeps current any model we host?
How can FISTA Solutions help?
FISTA Solutions deploys open-weight models in client environments where data residency, air-gapping, or volume justify it, reviews licensing fit, benchmarks candidates on clients' own cases, and integrates them behind a governed gateway alongside hosted models, through its AI enablement practice and its AI agents builds. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries.
To compare hosted and open-weight options on your own workload, talk to FISTA on WhatsApp, or read when to self-host LLMs.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What does open-weight mean?
That the model's trained parameters are published and can be downloaded and run on your own infrastructure. It does not necessarily mean the training data, code, or process are public, and it does not imply an unrestricted licence. The practical benefit is that you can run the model where you choose.
02Are open-weight models free to use commercially?
It depends on the licence. Some permit broad commercial use, others restrict it by company size or use case, and most impose acceptable-use terms. Review the specific licence for commercial use, redistribution, modification, attribution, and any restrictions before building on it. Consult counsel.
03When do open-weight models make sense?
When data cannot leave the environment under any contractual arrangement, when air-gapped or restricted operation is required, when regulatory or customer commitments demand it, or when volume is high and steady enough that infrastructure costs less than per-token fees. Also when a specialized fine-tune is needed.
04What do open-weight models cost to run?
Infrastructure (often GPU capacity), platform engineering, monitoring, and the ongoing work of evaluating and adopting newer models, in place of per-token fees. At low or bursty volume this is usually more expensive; at steady high volume it can be substantially cheaper.
05Do open-weight models perform well enough?
For many focused enterprise tasks, yes: classification, extraction, retrieval-grounded answering, and summarization of standard material. Frontier hosted models generally lead on complex reasoning and breadth. Decide by evaluation on your own tasks rather than by general benchmarks.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.