Sovereign AI
Private AI Deployment for Regulated Industries
Drafted with AI, reviewed by the Xinity team.
Regulated organisations are caught between two facts. AI is now useful enough that teams will use it whether or not IT provides it. And the default way to use it, sending prompts to a public API, means sensitive data is processed on someone else's infrastructure under someone else's terms. For a bank, a hospital or a law firm, that is often where the conversation with compliance ends.
This guide sets out what private AI deployment means, what it requires beyond the model itself, and how organisations in banking, healthcare, law and the public sector can evaluate their options.
What "private AI deployment" means
The phrase covers three different architectures, and the differences matter when compliance asks where the data goes.
Public API with a data processing agreement. Data goes to a third-party model provider, which commits by contract to how it is handled. Compliance rests on the contract, not on the architecture. GDPR allows this through the processor construct, but the data still leaves the organisation. For information covered by professional secrecy, such as banking secrecy, medical confidentiality or legal privilege, a processor contract might not be enough.
Dedicated cloud or VPC deployment. The model runs in an environment isolated from other tenants, but the provider still owns and operates the hardware, the hypervisor and the software supply chain. Data can stay in a chosen region, but control over the stack does not sit with you.
On-premise or private infrastructure. The model runs on hardware the organisation owns or directly controls, inside its own network. No data crosses an external network during inference. This is the architecture that meets strict residency, air-gap and supply-chain requirements.
Organisations with the strictest requirements usually end up at the third option. The question then is what the platform around the model needs to provide.
What regulated organisations need beyond the model
Running an open-weight model on your own server is a starting point, not a solution. The model does no access control, keeps no record of who asked what, and enforces no policy. The platform around it has to.
Access control and identity
Every request should be attributable to a person, a service account or a department. When an auditor asks who had access to the model that processed a client file, "anyone with the endpoint URL" is not an answer. Role-based access control at the API level, connected to the identity provider the organisation already uses, is the baseline for production.
Policy enforcement
Different teams need different rules. A legal team working with confidential deal documents needs different constraints than a service team summarising public product information. Those rules belong in the platform layer, applied the same way for every application, rather than rebuilt by each development team.
Audit trail
Each request should leave a record of who made it, which model and configuration answered it, and when. That record is how the organisation shows compliance afterwards, to an internal audit committee, a data protection authority or a court. Logging that lives only in application code can be skipped by any application that leaves it out, so it belongs in the serving layer.
Validated, stable models
Regulated organisations cannot run models whose behaviour changes without notice. Testing a specific model version against the organisation's own use cases before it reaches production, and keeping it fixed until the next version has passed the same tests, is a governance step that belongs in the deployment workflow from day one.
Sector considerations
Banking and financial services
DORA makes ICT third-party risk a supervisory matter, and supervisors such as the FMA in Austria and BaFin in Germany expect outsourcing arrangements to be documented, controlled and exit-ready. Banking secrecy, in Austria under §38 BWG, protects client information on top of GDPR. The infrastructure that serves models should be treated like any other critical IT system: change-controlled, access-logged and able to produce evidence on request.
Healthcare
Health data is a special category under Article 9 GDPR, and medical confidentiality applies on top. Hospital IT teams using AI for documentation, administration or clinical support need patient data to stay inside the facility's infrastructure during processing. Auditability also matters for clinical governance: the hospital needs to be able to show what information the AI had when it produced a given output.
Legal
Law firms handle privileged communications and confidential commercial information, and the duty of confidentiality extends to the systems used to process them. Sending client material to a public API, even under a data processing agreement, puts a third party in the chain. Many firms' confidentiality obligations and client terms make that hard to justify.
Public sector
Public bodies in Austria and across the EU work under national data protection law as well as GDPR, and often under their own procurement and data localisation rules. Running AI on infrastructure controlled by a non-EU provider raises questions beyond data protection: dependency on a foreign supply chain, and the ability to keep public services running independently.
Evaluating a private AI deployment platform
Five questions cut through vendor positioning:
Does inference run entirely inside my network, with no outbound calls? A platform that contacts a licensing server, a telemetry endpoint or an external model repository during inference is not fully private.
Is access control enforced at the API level and connected to my directory? Integration with LDAP, SAML or OIDC is the practical test.
Is there a per-request audit log in the platform layer? It should record identity, model, configuration and timestamp, be queryable, and be protected against alteration.
Can I fix a model version and promote new ones on my own schedule? Production governance means freezing, testing and promoting versions when the organisation decides, not when the vendor ships.
How is the platform supported on-premise? If support needs remote access to the production environment, check that against your network segmentation policy before you sign.
The shadow AI problem
Organisations that don't offer a compliant AI option don't stop AI use. They push it into personal accounts, browser extensions and unsanctioned SaaS tools, where IT cannot see it and no policy applies. A governed private option addresses this directly: teams get the tools they want, under controls the organisation can defend.
Starting a deployment
For most organisations the practical first step is a scoped pilot: a defined set of use cases, a small user group and a short evaluation window. It tests the architecture against the organisation's own IT and compliance requirements before anyone commits to a rollout.
Before the pilot starts:
Define the use cases in scope and classify the data they will touch
Confirm the environment (on-premise servers, private cloud or hybrid) and check the hardware meets the model's requirements
Map access control to the existing directory
Agree the audit log format and destination with the compliance or data protection team
Set the process for validating and promoting model versions
A pilot that skips these steps often produces a deployment that works technically but cannot be signed off by compliance or legal, and so never reaches production.
Summary
Private AI deployment is an architecture question before it is a vendor question. Data must stay under the organisation's control during inference, every request must be attributable and logged, and model behaviour must follow policy the organisation can enforce.
Get the architecture right first and the shortlist narrows quickly to platforms that actually run on-premise, control access and keep an audit trail.
Xinity is the layer around the model. It deploys, governs and audits open-weight models on infrastructure you control, with SSO/LDAP, role-based access control and audit logging built in. If your organisation is evaluating private AI deployment, the 30-day pilot programme is the structured way to start.