AI
Why Open-Weight Models Are Now Enterprise-Ready
What changed
Open-weight models are now close enough to the frontier that, for most enterprise work, capability is no longer the actual deciding question.
In May 2023, the best closed model led the best open-weight model by about 15 percent on the Arena human-preference leaderboard. As of March 2026, that lead was 3.4 percent, according to Stanford's AI Index 2026. The gap is not zero, and it moves each time a new frontier model ships. But it is small.
The more telling numbers come from professional tasks. On the AI Index's benchmarks for tax, corporate finance and legal reasoning, the top 15 models sit within a few percentage points of each other. On the corporate finance benchmark, which tests reading credit agreements of 200 pages and more, an open-weight model came first.
The report draws the conclusion itself: with capability no longer a clear differentiator, the competition is shifting to cost, reliability and real-world usefulness. For a CTO, that is the point where the question changes from "is the model good enough?" to "can we run it the way our organisation needs?".
Serving is no longer the hard part
A few years ago, running a capable model on your own hardware was a research project. Today it is an installation.
Mature open-source inference engines handle the hard engineering: batching requests, managing GPU memory, serving long contexts. They expose an OpenAI-compatible API, so applications built against a cloud endpoint can point at an internal one instead. Model weights are published with documentation, and hardware that fits in a server room can run models that would have needed a data centre not long ago.
For an infrastructure team, getting a model to answer a request on internal hardware is now a matter of days.
What enterprise-ready actually means
Enterprise-ready is not a benchmark score. In a bank, a hospital or a law firm, it means four things:
Predictable behaviour. The same model version answers today and next quarter, until you decide to change it.
Control over data. Prompts and answers stay where your policies say they stay.
Accountability. You can say who used which model, through which application, and when.
Clear limits. Each team reaches the models it is meant to use, within the limits set for it.
Open weights already deliver the first two by design: you hold the version, and the model runs where you put it. The last two do not come with the model.
What the model does not bring
A model file contains weights, not an organisation chart. An inference engine serves requests quickly and reliably, which is exactly its job. Neither is meant to answer the questions your DPO or auditor will ask.
That work sits in a separate layer around the model:
Question from the organisation | What has to exist |
|---|---|
Who is allowed to use this? | Role-based access, connected to your existing login (SSO/OIDC) |
Which application sent this request? | Keys issued per application, revocable one by one |
What was asked, and when? | An audit trail of requests, stored on your side |
Can one team overload the system? | Rate limits and routing across models and GPUs |
How do we switch models safely? | One stable API in front, so applications do not change when the model does |
This is scope, not a shortcoming. The model and the engine do their part well. The governance layer is simply a different part.
How teams close the gap
There are two ways to get there. Some teams build the layer themselves: an API gateway, an identity integration, a logging pipeline, a dashboard for keys and limits. It works, and it becomes a product your team now maintains alongside its actual job.
The other way is to put a governance layer on top of the models you choose. That is what Xinity is. It runs entirely on your own servers and uses proven open-source inference engines underneath, then adds what the organisation needs around them: an OpenAI-compatible API, role-based access with SSO/OIDC, keys per application, routing and rate limiting, usage tracking and an audit trail. The core is open source, so your security team can read exactly what runs inside your network.
The model is ready. The question now is whether the layer around it is. Read the code on GitHub or start a 30-day pilot on your own hardware at xinity.ai.
Sources
Stanford HAI, AI Index Report 2026, Chapter 2: Technical Performance: open vs. closed Arena gap (15.2 percent in May 2023, 3.4 percent in March 2026), professional-domain benchmarks, shift toward cost and reliability
Xinity, open-source platform repository: platform capabilities named in the article