Skip to content
Applied AI

Private AI and private LLMs that run where your data already lives.

Some data can't be sent to a public AI service: patient records, client files, contracts, source code. Our private LLM development work puts assistants and agents in your own cloud account or on your own servers, where they answer from your own documents and respect who is allowed to see what.

A glass bell jar over a steel gear on a base with a blue ring.
On this page

Three levels of private

"Private AI" means different things to different vendors. These are the three options, from least to most contained:

OptionWhere your data goesGood forTrade-off
A commercial model API on business termsTo the model vendor, under a contract that rules out training on it and limits retentionMost businesses, when the contract and, for health data, a BAA cover the data involvedYour data still leaves your network
A commercial model hosted in your own cloud accountTo your cloud provider, inside your account and regionCompanies already on Azure, AWS or Google Cloud with data-residency rulesModel choice is limited to what your cloud offers, and you manage more of it
An open-weight model on your own servers or private GPUsNowhere. Prompts, documents and answers stay on your hardwareRegulated data, contracts that restrict processing, or networks with no internet accessYou pay for hardware and someone has to run it

The first option is enough for more companies than you might expect. If it's enough for you, we will say so, because it is cheaper to run and gives you the strongest models.

What open-weight models are good for

Open-weight models can be downloaded and run on your own hardware under licenses that allow commercial use. OpenAI released two in August 2025 under the Apache 2.0 license, and says the larger one runs on a single 80 GB GPU. Meta, Mistral, Google and Alibaba publish open-weight families too.

They are not the strongest models available. For narrow, well-defined work, such as answering from your own policies, drafting from templates, classifying documents or pulling fields from forms, they are often good enough. We measure that on your own examples before you buy any hardware.

What we build

  • Assistants that answer from your documents, with citations back to the source page, so staff can check the answer.
  • Retrieval that respects your permissions: a user only gets answers drawn from documents they are allowed to open.
  • Document processing that reads contracts, claims, applications or invoices and writes the fields into your systems, with a person checking anything uncertain.
  • Agents that act inside your systems, under the approval rules described on our AI agent cost guide.
  • The serving stack itself: model hosting, access control, logging, monitoring and updates.

What we built for a dental group

For a dental group with 26 providers across four locations, we built a voice agent that answers the phones without sending patient audio to any outside service. Speech recognition and an open-weight language model run on a single GPU server inside the practice's own network. Nothing crosses the network boundary, and there is no per-token fee. Read the case study.

How we keep it safe

  • Your data is not used to train anyone else's model, and that is written into the contract.
  • Every question and answer can be logged for audit, and the log stays with you.
  • Access follows your existing identity provider and document permissions.
  • For health data, a Business Associate Agreement is available.
  • Our ISO/IEC 27001:2022 certificate covers the integration, training and deployment of machine learning and generative AI models by name. Read the scope.

What it costs

An assistant over one document collection, on a commercial API under business terms, is a small project. An on-premise deployment with document permissions, several data sources and agents that act on records is a larger one, plus the hardware. We fix the price after the audit, which includes a test of candidate models on your own examples. For the running costs of each option, see what AI agent development costs.

// FIXED-FEE EVALUATION

The Automation Audit

One workflow you name, mapped end to end, with the arithmetic done before anyone writes code.

Scope
One workflow you choose, traced end to end, including the steps nobody documented.
Duration
Two weeks, fixed.
Fee
Fixed, and quoted in full before we start. No hourly drift.
You get
A written map: what can be automated, what it would save in hours, what building it would cost, and what we would leave alone.
You keep it
The map is yours whether or not you hire us to build anything.
The TechLand Commitment

And if the audit shows the automation will not pay for itself inside twelve months, we will tell you, and we will not quote the build.

Start with the audit

Questions we get asked about private AI

Will a private model be as good as ChatGPT?

On general questions, usually not. On a narrow task with good examples and access to the right documents, often close enough that the difference doesn't matter. We measure it on your own cases before you commit, and if it isn't good enough, we tell you.

Do you fine-tune models on our data?

Only when testing shows it helps more than better retrieval would, which is less often than people expect. A fine-tuned model is yours, and it stays on your infrastructure.

Can it run with no internet connection?

Yes. An open-weight model on your own servers can run on a network with no route to the internet. Updates are then delivered and installed like any other software in that environment.

What hardware do we need?

That depends on the model size, the number of users and how fast answers need to be. The audit sizes it, and we give you the options before you buy anything.

Sources