Language models and company data: where the boundary lies
When a company decides to build a language model into a product or a process, the first question is not “which model is best” but “what data will it see”. The answer shapes the architecture, the cost and what can be promised to customers and lawyers. We settle it before estimating the project, not after launch.
Three ways to connect a model
An external provider. Requests go to the model through an API — directly or through a gateway that routes them between several providers. It is the fastest route and gives access to the strongest models, but the data in each request is passed to a third party. So we fix up front which classes of data may be sent, how the provider stores requests and whether it trains on them, which region processes them and what must be removed or anonymised before sending. With an external API, nobody can promise that “data never leaves the company”.
A regional provider. When jurisdiction, the contract with the provider or an infrastructure boundary matters, we choose a model available in the right country. Where a provider comes from does not replace the checks: we read its storage and processing terms as carefully as anyone else’s.
A model on your own servers. An open-weight model runs on the company’s GPUs, and requests never leave its infrastructure. That takes your own hardware and operations, and the choice is limited to models you can deploy yourself. A prototype can run on a workstation, but that is no substitute for a production deployment.
Six questions before the estimate
- Which data may be sent to an external provider, and which may not.
- What has to be anonymised, aggregated or removed from a request.
- Whether the model must run inside your infrastructure.
- Who stores the model’s requests and answers, and for how long.
- Which metrics can be collected without keeping the content of requests.
- How data is deleted once the project ends.
The answers determine the architecture. If, say, contracts must not be sent to external services, the knowledge base over them is built with vector search and a model inside your perimeter — even when an external API would be cheaper.
Quality is measured, not promised
A language model’s quality depends on the data and on how the task is framed, so we do not guarantee metrics in advance. Instead we assemble a set of examples from your process — questions with correct answers, typical documents, difficult cases — and measure every change against it: another model, a new prompt, a different way of searching the documents. Iterations are part of the project, and we name the risks before work starts.
We also test behaviour at the edges: what the model says when the documents hold no answer, how it handles off-topic questions, and whether it reveals anything a user should not see.
An agent gets the least it needs
When a model does not just answer but acts — creates tasks, writes to the CRM, sends emails — a boundary of authority joins the boundary of data. The agent receives only the permissions its scenario needs, a person confirms important actions, and every step is logged. That way a mistake can be found and corrected rather than discovered through its consequences.
Where to start
With one process whose rules are clear and whose result is easy to check. On it we agree how data is handled, assemble the examples and measure whether there is an effect. Only then does it make sense to extend the rollout.

