Engineering
What we learned helping teams self-host AI
Teams choose self-hosting for a simple reason: prompts, files, and conversation history never leave their environment. The decision is usually right. What surprises them is where the actual work lands, rarely in the deployment itself, almost always in the operational seams around it. Notes from deployments we have supported.
Egress rules, not servers, are the first blocker
Standing up the platform inside a private cloud is the easy week. The hard conversation is with the network team: a self-hosted workspace still needs controlled egress to model providers, web search, and any SaaS integrations the business wants. The successful pattern is an explicit allowlist (named endpoints for each provider, nothing else), agreed with security before deployment, not discovered as a series of firewall tickets after. If truly zero egress is the requirement, that constrains you to locally hosted models, and that trade-off should be chosen consciously, not stumbled into.
Model access is a procurement problem in disguise
Multi-model access means API agreements with several providers, each with its own terms, quotas, and regional availability. Enterprises with existing cloud contracts often route through their cloud vendor’s model marketplace for some providers and go direct for others. Sorting this out early matters because quota approvals can take longer than the deployment itself, and a workspace with one working model on launch day undermines the multi-model pitch internally.
Secrets and keys deserve first-class treatment
A self-hosted AI workspace concentrates valuable credentials: provider API keys, OAuth tokens for connected drives, encryption keys for sensitive fields. The deployments that age well put these in a proper secrets manager with rotation from day one, rather than in environment files that quietly become permanent. This is also what makes key rotation after a personnel change a non-event instead of an incident.
Plan the second deploy before the first
The riskiest moment in self-hosting is not the launch; it is the first upgrade. Teams that treat the deployment as a one-off end up frozen on an old version because nobody wants to touch it. Teams that set up a repeatable pipeline (staged environments, database migrations as code, a rollback path), upgrade monthly without drama and keep pace with model and feature releases. If you adopt one habit from this post, make it this one: the first deployment is a rehearsal for every deployment after it.
It is worth it, for the right reasons
None of this argues against self-hosting. Data residency, regulatory requirements, and the simple assurance that conversations stay inside your perimeter are real, durable benefits. The lesson is narrower: budget the effort where it actually goes (network policy, model procurement, secrets, upgrade discipline), and the deployment itself will be the least interesting part of the project. That is how it should be.