Our Customers

Who Runs AI on Their Own Hardware

Xinity is used by organizations that have a working AI use case and a legal or contractual reason their data cannot leave the building. In practice that means regulated sectors in the DACH region: banking, healthcare, legal, media, public administration and industrial production. Inference runs on hardware the organization owns, so no prompt, document or output reaches an external provider.

Three reasons companies run AI on their own hardware

Obligation

Regulated and public institutions

For these organizations, on-premise inference is not a preference. It is the only configuration their regulator accepts.

Healthcare, legal, financial and public sector organizations operating under strict data protection law, where full control over the place of processing is a documented requirement.

✓Aligned with GDPR and the EU AI Act by architecture
✓Audit trail on every inference request
✓No data leaves your infrastructure
✓On-premise deployment, no cloud dependency
Secrecy

Industry, media and production

Here the risk is not a fine. It is a competitor, or a source, learning something they should not.

Manufacturing, engineering and media companies protecting trade secrets, proprietary workflows and confidential sources from exposure to external systems.

✓Protect proprietary models and training data
✓Run AI on air-gapped or isolated networks
✓Source protection for editorial teams
✓No vendor lock-in, open standard APIs
Strategy

Data-first companies

No regulator is asking. These companies keep AI in-house because the data is the business.

Organizations that choose to keep AI and data internal for competitive or economic reasons, independent of any regulatory requirement.

✓Fixed capacity cost, no per-token surprises
✓Competitive advantage through data ownership
✓Cost falls as internal adoption rises
✓Independence from US hyperscalers

What regulated organizations need beyond inference

Most technical teams already know the open source engines are good. An inference engine answers whether you can serve a model on a GPU. It was never built to answer the questions that decide whether the system is allowed to reach production. Those are the questions Xinity is built around.

Who can use this, and who approved it?

Role-based access control and SSO or LDAP integration, so access follows the identity system you already run instead of a shared API key in a config file.

What was processed, by whom, and when?

An audit trail on every inference request, with filtering, export and per-instance review in the dashboard. Built for the person who has to answer the question months later.

Who patches it when a vulnerability is published?

We do, and you decide when to apply it. Because the deployment is yours, updates go through your own change management rather than arriving unannounced.

Who is liable when something breaks?

A company in Vienna with a signed contract, a data processing agreement and support commitments in writing. Self-built infrastructure has no counterparty.

Who answers before the board demo?

Named support with agreed response times, escalating by tier. Not a community forum and not the engineer who set it up two years ago.

What do we actually show the auditor?

The deployment itself. It runs on your servers, the core is open source so your security team reads the code rather than trusting a description, and the evidence lives on your side.

In production today

We evaluated every option. Only architectural sovereignty delivers what both journalistic source protection and the integrity of public information platforms truly require: infrastructure no one else can access.
Martin Mair
CIO and Member of the Board, Mediengruppe Wiener Zeitung
Confare CIO of the Year 2026
97%
lower cost per request
MeinDienstplan, after moving inference on-premise

Cost per request fell from EUR 2.00 to EUR 0.06. Owning the inference removes the per-token meter, so the cost per request falls as internal usage rises rather than climbing with it.

Find your industry

Each page names the specific regulation that applies and what a deployment looks like in that sector.

Not on the list? Book a demo and we will work out whether an on-premise deployment fits your situation.

Questions we get asked

For some teams that is the right answer, and we say so publicly. vLLM is excellent inference software. What it does not provide is access control, an audit trail an auditor accepts, a patch process, a data processing agreement, or a counterparty who is liable for the software. The realistic alternative to buying a platform is not vLLM on its own, it is vLLM plus the layer you build and maintain around it.

Read the honest version, including when you should not buy from us