Home · Glossary · On-premise AI
Enterprise AI glossary · Sovereign & On-Premise AI

On-premise AI

Deployment of AI workloads — model serving, embeddings, RAG, fine-tuning — entirely on hardware physically located in customer facilities.

Definition

What On-premise AI means in practice

On-premise AI is the broader category that sovereign AI sits inside. A workload is on-prem when every component — GPU inference, embedding generation, vector storage, orchestration, logging — runs on hardware physically located in customer facilities (or in a colocation facility the customer leases). On-prem is necessary but not sufficient for sovereign: a deployment can be on-prem and still phone home to a telemetry endpoint, fetch model weights from a vendor registry, or call out to a third-party API for one feature. Sovereign deployments close those loopholes by blocking all outbound network egress at the cluster namespace level.

Go deeper
Sovereign AI: architecture and deployment →

All 62 terms, in plain language

Sovereign AI, RAG, agentic AI, IDP, MLOps and the regulations that shape enterprise AI.

Browse the glossary →Talk to an engineer →