The Platform

Xinity Runtime

dashboard.xinity.ai
Xinity Dashboard
01
Integration

Plugs into the applications you already run

Xinity Runtime exposes an OpenAI-compatible API on your local network. Your applications keep their SDKs, prompts and request formats. What changes is the base URL and the API key.

OpenAI-compatible endpoint inside your network
Streaming, JSON mode and function calling unchanged
No new tooling, no retraining your teams
02
Deployment

Deploy models with the guesswork removed

Pick from validated open-weight models with license terms and release dates up front. Before anything runs, the dashboard shows what your hardware can carry.

Driver, supported features and VRAM per node
Capacity and throughput estimates before deployment
A named constraint when a model cannot deploy
03
Control

One control plane for every AI request

Access follows your identity system, every inference request lands in the audit trail, and usage and cost are visible across the fleet.

Role-based access with SSO or LDAP
Audit trail with filtering and export
Prometheus-compatible metrics for your monitoring

Key Benefits

Why Choose Xinity AI

Sovereignty

Data Sovereignty

Inference runs on hardware you own. There is no external API call in the path, so prompts, documents and outputs stay inside your network.

Enterprise

Enterprise-Ready

Role-based access through SSO or LDAP, multi-node operation, and metrics that plug into Prometheus and Grafana. Built to be run by your ops team.

Cost

Cost Predictability

A fixed infrastructure cost instead of a per-token meter. Capacity estimates per node show what a model will handle before you commit hardware to it.

Migration

Fast Migration

A new base URL and API key. SDKs, prompts, streaming and function calling stay as they are, so existing applications move in days.

Compliance

Regulatory Control

An audit trail on every inference request, with filtering and export for whoever has to answer questions later. Built for GDPR and EU AI Act requirements.

Open Source

Open-Core Foundation

The engine is open source. Your security team reads the code instead of trusting a description, and you can start without talking to us.

Current Models

Validated Models

NEW · GENERAL WORKHORSE
NEW

Qwen3.8 27B FP8

qwen3.8-27b-fp8

The most capable Qwen generation to date. Coding, agentic workflows, documents, and native image and video understanding in one dense model.

LicenseApache 2.0
ArchitectureDense, 27B, vision
Context262K tokens
HuggingFaceQwen/Qwen3.8-27B-FP8
MoE WORKHORSE

Qwen3.6 35B A3B

qwen3.6-35b-a3b-fp8

Chat, documents, agentic workflows, coding. Sparse MoE efficiency for mixed enterprise workloads.

LicenseApache 2.0
ArchitectureMoE, 35B total / 3B active
Context262K tokens
HuggingFaceQwen/Qwen3.6-35B-A3B-FP8
HIGH-VOLUME TASKS

Ministral 3 3B

ministral-3b

Classification, extraction, routing, summarization at volume. Image input for document and scan processing.

LicenseApache 2.0
ArchitectureDense, 3B, vision
Context262K tokens
HuggingFacemistralai/Ministral-3-3B-Instruct-2512
RETRIEVAL

EmbeddingGemma

embedding-gemma

Multilingual document embeddings for fully on-premise RAG. Pairs with any generation model above.

LicenseGemma Terms of Use
TypeEmbedding model
FootprintCPU-capable
HuggingFacegoogle/embeddinggemma-300m
YOUR MODEL

Not on the list?

Xinity serves any open-weight model with an OpenAI-compatible serving path. Validation for regulated environments takes days, not months.

Contact Us