AI
How to evaluate on-premise AI infrastructure for regulated enterprises
If you work in a bank, a hospital, a law firm, a newsroom or a public authority, the question is no longer whether your teams will use AI. They already do. The question is whether you can run it in a way that holds up when someone checks.
On-premise AI promises exactly that: models running on hardware you control, with data that never leaves your network. But "on-premise" on its own says little about whether a platform is ready for a regulated environment. Two offerings can both run inside your data centre and still differ completely in what they let you prove.
This guide walks through what to evaluate, in the order that matters, and ends with ten questions to take into every vendor call.
Start with the question your auditor will ask
Most evaluations start with the model: which one, how big, how fast. That matters, but it's not what your auditor, your data protection officer or your regulator will ask first.
They'll ask three things. Who used the AI? For what? And where did the data go?
If a platform can't answer those questions clearly, the model it runs doesn't matter. So start your evaluation where the scrutiny will come from, and work back to the technology from there.
1. Where it runs
The first check is the simplest, and it's easy to get wrong. "On-premise" should mean that inference runs on hardware you own or have chosen, inside a perimeter you control.
Check the details:
Prompts and responses. Do they stay inside your network at every step, or does any part of the request pass through an external service?
Logs and telemetry. Where are they written, and does anything get sent back to the vendor by default?
Dependencies. Does the platform need an internet connection to work, or can it run fully air-gapped if you need it to?
For organisations bound by professional secrecy, banking rules or patient confidentiality, this is the foundation. Under GDPR Article 28, every external processor you add is another contract to negotiate and another party to assess. Keeping inference inside your own perimeter means prompts and answers are never sent to an external model provider, which removes a whole category of processors to assess.
2. Access control
Once AI runs in your building, the next question is who can reach it. In most organisations the answer should differ by team. Your legal department and your marketing team probably shouldn't have the same access, to the same data, under the same rules.
Look for:
Role-based access control (RBAC), so permissions follow roles rather than individual arrangements.
Single sign-on (SSO), connected to the identity provider you already use, so leavers lose access automatically.
Keys per application, so every tool that calls a model can be identified and switched off on its own.
The test is simple: can you control access per team and per application, as a setting? If the answer involves emails and spreadsheets, access control isn't really in place.
3. Audit trail
This is where many platforms fall short, and where regulated organisations have the least room to compromise.
A usable audit trail records every request: when it was made, through which application, with which key and to which model. It should be stored on your side, and your data protection officer should be able to read it without opening a ticket with the vendor.
Ask how the audit trail is produced. If logging is something you have to switch on, configure per application or remember to maintain, it will have gaps. It should happen by default, for every request, without anyone having to think about it.
The GDPR documentation obligations, including records of processing and the security measures under Article 32, stay with you regardless of where the model runs. A good platform makes them easier to meet. It can't take them off your hands.
4. The policy layer
Access control decides who can reach the AI. The policy layer decides what happens once they do.
Think of it as the set of rules that sits above the models: which requests go to which model, how usage is tracked, which applications are allowed in at all. Without it, every application has to carry its own rules, and they drift apart over time.
This is the layer most often missing when organisations assemble on-premise AI themselves. Running a model on your own hardware has become straightforward. Governing how dozens of teams and applications use it is the harder part. At Xinity, this is what we build: a gateway and dashboard that sit in front of your models and handle routing, access control, request recording and usage tracking in one place.
5. Deployment and model choice
Models change quickly. The one that fits best today may not be the one you want in twelve months. Your infrastructure should let you switch without rebuilding everything around it.
The practical test is compatibility. A platform with an OpenAI-compatible endpoint lets existing applications point to a new base URL with a new key, without a rewrite. Many coding tools, chat interfaces, automation platforms and developer frameworks already support this.
Also check:
Which models you can run, and whether you can bring your own.
How new models are deployed, and how long it realistically takes.
Whether several models can run side by side, so different workloads can use different models.
6. Operations
Finally, ask the unglamorous questions. Who installs it? Who updates it? Who's on call when something breaks at nine on a Monday morning?
On-premise gives you control, and it also gives you responsibility. Be clear about which parts you'll operate yourself and which the vendor supports, and get that split in writing. A clear shared-responsibility model is a good sign. A vague one is a warning.
Ten questions to take into every vendor call
Do prompts and responses stay inside our network at every step?
Does anything, including logs or telemetry, leave our environment by default?
Can the platform run fully air-gapped?
Can we control access per team and per application?
Does it connect to our existing SSO?
Is every request logged by default, and where is that log stored?
Can our DPO read the audit trail without contacting you?
Where do the rules above the models live: routing, access, usage tracking?
Is the endpoint OpenAI-compatible, so our applications don't need rewriting?
Which parts do we operate, which do you support, and is that written down?
A platform that answers all ten clearly is ready for a regulated environment. One that hesitates on the audit trail or the policy layer probably isn't yet, however good its models are.
See it on your own hardware
If you want to test these questions against real workloads rather than a slide deck, Xinity runs a 30-day pilot on your own infrastructure. You'll see how AI is used, through which applications, and where every request goes, and then you decide.
Start the pilot or explore the open-source core on GitHub.
This article was drafted with the help of AI and naturally reviewed by the Xinity team.