On-premise AI without giving up the platform.
Plenty of tutorials show how to run a model on a server. The hard part is everything around it: governed data access, agent orchestration, policy-gated actions, human approvals, monitoring, and audit evidence — operated reliably inside your own boundary. That is what NeuroCluster deploys on your infrastructure.
Who actually needs on-premise AI
For most organizations an EU-hosted dedicated tenant is enough. On-premises and air-gapped deployment is for the cases where data cannot leave the boundary at all:
- Defense, national security, and classified environments where external connectivity is prohibited.
- Critical infrastructure operators under NIS2 with strict OT-adjacent data rules.
- Healthcare and government workloads where policy or procurement mandates on-site processing.
- Organizations with existing GPU estates that want AI workloads where the data already lives.
What runs on your hardware
The full platform, not a stripped-down agent runner: open-weight LLM inference on your GPUs, embeddings and vector search, governed RAG over your data sources, the agent orchestration layer with policy gates and human-in-the-loop approvals, and the evidence layer that records every action.
Agent-generated code executes in sandboxed micro-VMs, so even in an air-gapped environment untrusted code never runs directly on cluster nodes.
Hardware and operations
NeuroCluster deploys on standard Kubernetes with NVIDIA GPUs — from a handful of nodes for departmental workloads to multi-node GPU clusters for full inference estates. Sizing depends on model choice and concurrency; a readiness call includes a concrete hardware recommendation.
Operations are GitOps-based: declarative configuration, versioned changes, and updates you review and apply inside your boundary. In air-gapped environments, updates arrive as signed bundles you transfer and apply on your schedule. No phone-home, no silent changes.
The enterprise-controls difference
Self-hosting a model gets you privacy; it does not get you governance. The controls that make AI defensible in front of a security review — identity-bound execution, row-level data policies, approval workflows, deterministic logging, exportable evidence packs — are the platform's core, and they work identically on-premises and air-gapped.
This matters for the EU AI Act as much as for security: high-risk obligations around logging, human oversight, and technical documentation do not disappear because the deployment is private. On-premise NeuroCluster generates the same evidence packs as every other profile.
Frequently asked questions
What is the difference between on-premises and air-gapped deployment?
On-premises runs on your infrastructure but can retain controlled external connectivity (for example, for model updates or policy-allowed external APIs). Air-gapped is fully disconnected: updates arrive as signed bundles you transfer manually, and nothing leaves the boundary.
Which LLMs can we run on-premise?
Open-weight models — Qwen, Mistral, and other open models — run on your GPUs. The AI gateway routes between them per workflow, so you can serve different models for different sensitivity levels inside the same boundary.
What hardware do we need?
Standard Kubernetes with NVIDIA GPUs. The footprint depends on model sizes and concurrency — departmental workloads run on a few GPU nodes; a readiness call includes a concrete sizing recommendation for your use cases.
How do updates work in an air-gapped environment?
Updates ship as signed, versioned bundles that your team transfers into the environment and applies through the same GitOps flow — reviewed, scheduled, and logged by you.
Keep evaluating
Private AI cloud
The full deployment-model spectrum, from EU shared cloud to dedicated tenant.
Platform: private runtime
Sandboxed execution, tenant isolation, and runtime security architecture.
Trust Center
Security controls, residency documentation, and the procurement pack.
Case study: energy grid forecasting
Critical-infrastructure AI with governed agents and human review.