Private AI Infrastructure

Managed AI Runtime

Building a private AI environment is one thing — operating it securely, up to date and economically is another. The Managed AI Runtime is our operations offering: a containerised, private AI environment with your models and your data, operated by us — for a predictable monthly fee plus transparently passed-through compute costs.

More services

Managed AI Runtime at a glance

The scope of services is divided into four areas. Operations: containerised runtime, private model endpoints on modern inference servers (e.g. vLLM), updates, patches and backups. Security: single sign-on via Entra ID/OIDC, role-based permissions, secrets management and private networks. Compliance: data processing agreement, technical and organisational measures, subprocessor list, logging and deletion concept, plus a documentation package that supports you with obligations under GDPR and the AI Act. Cost control: continuous monitoring of token and GPU consumption with monthly reporting.

On compute we take an unusual approach: we pass GPU and infrastructure costs through to you transparently at cost. We earn on operations — not by reselling compute time at a markup. That makes your costs predictable and our recommendations independent: if a cheaper provider fits your workload better, we switch to it.

Our portability promise also applies to operations: your AI can leave us. Deployments are containerised, configuration and data paths are documented, and exit documentation is part of the service. On request, we test once a year in the AI Exit Drill whether your stack can actually be restored on alternative infrastructure — with us, digital sovereignty is not claimed but tested. Our support commitments are honestly calculated: we agree on response times we can organisationally keep, instead of making enterprise promises that only exist on paper.

Managed AI Runtime: what we deliver

Containerised runtime and private model endpoints (e.g. vLLM)
SSO via Entra ID/OIDC, RBAC and private networks
Monitoring, logging, backup and patch management
Compliance package: DPA, TOMs, subprocessor list, AI Act documentation
Token and GPU cost control with monthly reporting
Exit documentation and optional annual AI Exit Drill
Use cases

Managed AI Runtime in practice

// 01

Private LLM for sensitive documents

Contracts, HR documents or engineering data are processed with a private model instance — without content going to external model providers.

// 02

Knowledge search across company data

RAG across technical documentation, knowledge base and file shares: employees find answers instead of documents — operated in your controlled environment.

// 03

Internal AI APIs for business applications

Your applications use private inference endpoints for classification, extraction or text generation — with stable monthly costs instead of surprising API bills.

How we proceed

Managed AI Runtime: step by step

1

Takeover or setup of the environment according to target architecture

2

Hardening and integration: SSO, network, permissions

3

Operations start with monitoring, backup and runbooks

4

Regular operations with monthly cost and status report

Support and response times are agreed individually and calculated so that they can be reliably met organisationally. All prices are net; the individual quotation is binding.

FAQ

Managed AI Runtime: frequently asked questions

Still have open questions? We're happy to clarify them in an initial call.

What does managed operation cost?
The monthly operations fee depends on scope and support level; you receive the exact price as a transparent quotation after the initial call. Compute, storage and traffic costs are added — passed through transparently at cost, with no hidden markup.
Do we need our own GPUs for this?
No. We operate your environment on suitable rented infrastructure with German and European providers — from smaller inference systems to high-end GPUs. Owning hardware only pays off with consistently high utilisation; we calculate exactly that for you honestly.
Does this run in our Azure environment or with you?
Both are possible. We follow a hybrid approach: Microsoft 365, Entra ID and Business Central stay where they are — the private AI environment runs where control, cost and data protection require it. The target architecture defines this per workload.
How do we get out again?
In an orderly, documented way: deployments are containerised, configuration and data paths are available as exit documentation. Your stack can be migrated to another operator or your own team — on request we test this annually in the AI Exit Drill.

Let's talk about your project.

Free initial consultation, 30–45 minutes, remote. An honest assessment — even if the answer is that you don't actually need it.