Sovereign AI
Enterprise AI: 7 Things IT Leaders Must Check | Xinity
The decision to deploy AI in the enterprise stopped being a strategic option some time ago. It is an operational necessity. But taking that step raises a question that goes well beyond technology: who actually controls the infrastructure your AI runs on?
The EU AI Act imposes binding transparency obligations, GDPR fines reach up to four percent of global annual turnover, and uncontrolled shadow AI use quietly undermines any compliance approach. At the same time, migration to sovereign AI infrastructure takes days in practice, not months.
This article gives IT decision-makers seven concrete reference points: from the legal duties around data sovereignty, through the real threat of shadow AI, to open-weight models, predictable infrastructure costs and audit-proof logging.
Data sovereignty is not an option, it is a legal duty
The legal framework is already binding, and it moved in July 2026. Under Regulation (EU) 2026/1744, the Digital Omnibus on AI, standalone high-risk systems under Annex III now fall due on 2 December 2027 rather than the original 2 August 2026. AI embedded in regulated products under Annex I moves to 2 August 2028.
More important than the dates that moved are the ones that did not:
Since 2 February 2025, the prohibited practices under Article 5 and the AI literacy duty under Article 4 have applied. The latter was reworded on 27 July 2026.
Since 2 August 2026, the transparency obligations under Article 50 have applied.
From 2 December 2026, a few weeks away, machine-readable marking under Article 50(2) applies to systems already on the market, alongside new prohibitions.
Anyone building a compliance plan around December 2027 is planning past the nearest deadline.
As long as inference runs on servers outside your own control, data sovereignty remains structurally unsecured, regardless of what an external provider commits to contractually.
The financial consequences are concrete. GDPR breaches cost up to 4 % of global annual turnover. Under the AI Act, up to 7 % applies to prohibited practices and up to 3 % to breaches of the high-risk obligations. Compliance is no longer a contractual question. It is an architectural one. See the full compliance overview.
Technically, data sovereignty means this: inference runs exclusively on servers inside your own network. No prompt, no response, no metadata leaves your environment for an external API.
Shadow AI is the biggest compliance problem in the building
In practice, the largest compliance risk does not come from bad strategic decisions. It comes from everyday improvisation by individual teams.
Shadow AI is the unapproved use of public LLM APIs by employees who have no compliant alternative. No approval process, no audit trail, no data protection impact assessment. The scale is documented: in the PagerDuty 2026 Shadow AI Survey, conducted by Wakefield Research among 1,250 office professionals at companies with revenue above $500 million, 66 % said they had used AI tools at work despite believing this was against company policy, and more than a third had entered customer data into public AI models. Verizon's 2026 Data Breach Investigations Report found shadow AI detections rose fourfold in a year.
The structural problem is loss of control. Every request to an external provider can transmit customer data, contract contents or patient records. That data can feed training pipelines, which makes remediation difficult or impossible. The same question applies to storage as to processing: where your AI data actually lives.
Bans do not solve it. Leaving teams without a compliant alternative drives usage underground. The only effective answer is substitution: an internal offering productive enough to make external services unnecessary. Organisations that provide sovereign inference internally remove the incentive for shadow AI structurally, not through policy.
Open-weight models make control real and verifiable
The compliant internal offering has to be auditable itself. This is where proprietary API models fail structurally: neither model weights nor training process nor inference logic are open to inspection. For security teams in banks, hospitals and public authorities, that leaves a black-box problem no vendor assurance can close.
Open-weight models solve it at the root. Security teams examine the model itself rather than relying on vendor documentation. This is not a theoretical advantage: open models now match proprietary systems in many production use cases.
For regulated environments, Xinity provides specific models:
Qwen3.8 27B FP8 (dense, 27B, vision, 262K context, Apache 2.0) for general inference workloads
Qwen3.6 35B A3B (MoE, 35B total / 3B active, 262K context, Apache 2.0) for mixed enterprise workloads
Ministral 3 3B (dense, 3B, vision, Apache 2.0) for high request volumes with low latency requirements
EmbeddingGemma (CPU-capable, Gemma Terms of Use) for RAG applications with structured document access
Existing applications need no code changes in most cases, because the API stays OpenAI-compatible. The switch is invisible at the application layer and decisive at the regulatory one.
Migration to sovereign inference takes days, not months
The move follows three phases: hardware assessment, runtime installation and live operation. None of them requires cloud provisioning or external dependencies.
Phase 1: hardware assessment establishes which GPU infrastructure already sits in your data centre and can be used directly. Existing servers are inventoried and capacity requirements matched against target models. Anyone already running their own server infrastructure skips every cloud provisioning step.
Phase 2: runtime installation happens on those existing servers, entirely inside your own network. No data leaves during setup, and the installation process itself has no external dependency.
Phase 3: live operation does not start with a full cutover. The recommended approach is A/B routing with 5 to 10 % of real traffic before switching over completely. That lets you verify response quality and latency against your existing API baseline before production-critical workloads migrate.
The decisive difference is not the technology. It is the starting point: control your own hardware and you control the timeline.
Understanding the AI stack: where control actually forms
The AI stack covers hardware, operating system and drivers, model serving and orchestration, application logic and user interfaces. Most organisations already control the hardware, application and interface layers. The serving layer is what makes the difference.
The serving layer is the only point where access rights, policies and audit logging take effect across every model and every application. Whatever model runs beneath it and whatever application sits above it, controlling this layer means controlling the entire inference process.
One practical benefit: intelligent routing at this level enables load distribution across multiple GPUs, automatic failover and switching between models without adapting existing applications. New model versions are swapped in the stack and the application layer notices nothing.
Giving up that control has concrete regulatory consequences. Operating only at the model or application layer makes it impossible to demonstrate without gaps which data was processed, when, by which model. And the Article 50 transparency obligations have applied since August 2026, not from the high-risk deadline onward. More on where the layers sit: Understanding the AI Stack.
Governance does not start at the model. It starts at the layer holding everything else together.
From variable token costs to predictable infrastructure
Per-token APIs scale costs directly with usage volume, and the rates move. On 30 July 2026, OpenAI cut the price of GPT-5.6 Luna by 80 % and Terra by 20 %, three weeks after those models launched. Luna went from $1 / $6 to $0.20 / $1.20 per million input / output tokens; Terra from $2.50 / $15 to $2 / $12. A budget approved against spring rates was wrong by August, in that case in the customer's favour, but the direction is not the point. The point is that the number is set by someone else and changes without your involvement.
The headline rate is also not what you pay. Reasoning models generate internal thinking tokens that are billed at output rates whether or not you ever see them, which puts effective cost on reasoning-heavy work at roughly three to ten times the base rate. Output tokens themselves run three to six times input on most providers.
For hospitals, public authorities and banks that plan and get budgets approved annually, that volatility is structurally incompatible with the normal budget process.
On-premise infrastructure converts variable operating expense into predictable capital expenditure. Hardware, energy, operations and licensing can be calculated once. Above a certain request volume, total cost tips clearly in favour of owned infrastructure, because there are no volume discount negotiations, no routing overhead and no price changes imposed by an external provider.
Procurement teams and CFOs can set their own token costs against on-premise alternatives and find the break-even point for their specific request volume with the Xinity ROI calculator.
Audit trails as regulatory proof
For high-risk systems, the EU AI Act requires complete records under Articles 12 and 16 from 2 December 2027. The central question: which model produced which outputs, when, from which inputs? In an inspection, authorities do not ask for statements of intent. They ask for complete inference logs.
A compliant audit trail covers at least the following:
Timestamp of every request
Model version at the time of inference
Input hash
Generated output
Calling application
Identity of the requesting user or system
With public APIs, technical control over log data sits with the provider. Whether and how much of it reaches the customer varies by provider and is not contractually standardised.
Compliant audit trails cannot be retrofitted. They form at the serving layer that every inference request passes through, or they do not form at all. That decision is made when the platform is built.
Conclusion: seven points, one architectural decision
Adopting AI in regulated industries is not a model or tooling decision. It is an infrastructure and compliance decision.
The 16 additional months to December 2027 are not a reprieve. They are preparation time. The next binding date falls in December 2026, not 2027.
The concrete next step is a hardware assessment of your existing infrastructure. It quantifies which GPU capacity is already usable, which migration path is realistic, and at what request volume owned infrastructure becomes more economical than external APIs.
Anyone who wants to evaluate sovereign inference without a long-term commitment can do so with a 30-day pilot on their own hardware. Real workloads, your own data, your own infrastructure. Apply for a pilot spot — the programme is open to companies with 15 or more employees.
For IT decision-makers in the DACH region, the Sovereign AI Lab Meetup in Vienna offers a structured setting to compare architectural decisions, regulatory requirements and operating models with peers. Details at Sovereign AI Lab.
Sovereign AI is not a question of trusting a provider. It is a question of architecture.
A first draft of this article was generated with an AI writing tool. It was then reviewed, rewritten and fact-checked by Xinity's editorial team, which holds editorial responsibility for the published text. Regulatory dates were verified against Regulation (EU) 2026/1744 and the EU AI Act; figures and quotes were checked against the primary sources linked in the text.