Some work cannot go to a hosted AI endpoint. An insurer says so, a client contract says so, or the documents are simply not yours to send. The answer to that is not a better vendor policy. It is a model that runs where the data already is.
Where it runs#
Inside your own infrastructure, or a private, isolated instance you control. Both count. Which one fits is a question about your approvals, your infrastructure and who your security team will sign off on — not a preference of mine.
Why open-weight models are part of this#
Where data genuinely cannot leave your environment, the model has to run inside it. That rules out hosted endpoints and rules in open-weight models you host yourself, running on infrastructure you control. Which family fits is a question for your workflow, your hardware and your approvals — it is answered per engagement, not picked in advance.
The families I work with today include DeepSeek, Qwen, Kimi and GLM. Open-weight releases move quickly, so read that as a snapshot rather than a shortlist: part of the work is benchmarking candidates on your own documents instead of trusting a leaderboard.
Residency is the obvious reason. Three others matter more once a workflow is in production:
- The model does not change underneath you. A hosted endpoint can be updated or retired on someone else’s schedule. A report that was defensible last quarter should be reproducible this quarter, and pinned weights are what make that possible.
- Nothing is transmitted. No prompt, no document, no client name leaves the building — not as an assurance about a vendor’s policy, but because there is no outbound request.
- No lock-in on the most expensive component. Weights you hold can be swapped, benchmarked against each other, or kept in place for years. Inference cost at report volume is a running cost, not a one-off.
Open weights are not always the right answer — a hosted model is often faster to prove a workflow with, and I will say so. The point is that the choice stays yours, and it is made on the workflow’s constraints rather than on what I happen to resell.
What open weights actually protect#
Data residency is the reason most people give first. It is not the most valuable one.
What makes a firm’s AI useful is rarely the model. It is everything assembled around it: the documents you decide are authoritative, the prompts you refine over months, the examples that teach it your house style, the corrections your people make and feed back. That accumulated judgement is your domain knowledge, and it is the part a competitor cannot buy.
Every hosted request carries a piece of it outside your walls. The protection you have then is contractual: a vendor’s policy, which may well be excellent, and which you rely on because you cannot inspect the alternative. Weights you hold remove the question rather than answer it. There is no outbound request to govern.
To be exact about the limit: this protects that knowledge from one specific exposure, which is transmission to a third party. It does nothing about a badly governed internal deployment, an over-broad permission, or a corpus nobody curated. Those stay your problem, and designing for them is part of the work.
You own it, and you run it#
Handover is part of the scope, not an extra invoice: documentation, a runbook, and a system your team can operate without me. Running on your own hardware does not change that. If anything it matters more, because the hardware is yours too.
Your team runs it, not me. Administering the machines — the operating system, the drivers, the patching, the backups — stays with the people who already do that work for you. I build it, document it and hand it over. I would rather say that plainly here than leave the question open until contracting.
What I run today#
I run self-hosted open-weight models on hardware I own: multiple 48GB workstation GPUs and 256GB or more of host RAM, which serves models above 70B parameters at a precision chosen per model. I have not yet run one inside a client’s production environment, and this page will say so until that changes.
What I would need from your side#
An environment to deploy into, and someone who can approve changes to it. Your infrastructure and security people involved early rather than at sign-off. A real workflow with real documents, because a demo set will not surface the constraints that decide whether this works.
What it costs#
Case by case, and deliberately so. This is not one of the packaged engagements with a published band, because the number moves on things you control rather than things I do: the environment I deploy into, how many approvals a change needs, what your security review asks for, and what hardware you already own. Scoping settles it before either of us commits.
What is not published yet#
Two things a careful buyer will ask for, and my reasons for not having them here:
- No client outcome for a private deployment. There is no engagement to describe. I would rather say that than imply one.
- No hardware sizing for your environment. What I run is above; what you would need depends on model size, document volume and what you already own.
Each line here is deleted the day it stops being true.
Common questions#
What is private AI deployment?
Running open-weight AI models inside your own infrastructure, or a private, isolated instance you control, so documents and prompts never leave your environment.
Why use open-weight models?
Where data cannot leave your environment, the model has to run inside it. Weights you hold also do not change underneath you, nothing is transmitted to a third party, and the most expensive component is not locked in.
Who runs the system after delivery?
Your team. Handover is part of the scope: documentation, a runbook and a system your team can operate. Administering the machines stays with the people who already do that work.
How much does private AI deployment cost?
Case by case. It depends on the environment, the approvals a change needs, what your security review asks for and the hardware you already own, and scoping settles it before either side commits.
Start where it is cheapest#
The value check estimates what one workflow is worth to you. That is a different question from what this costs, and it is still the first thing worth doing: it turns a discovery call from a conversation about AI in general into one about a specific workflow with a number attached. If the honest answer is that a hosted model would prove it faster and cheaper, I will say so.
Request a discovery call · Run the value check · All services
Model and product names are referenced for description only and do not imply partnership, endorsement, or certification by any third party.