Private LLM deployment

Private LLM deployment: on your own servers or cloud account

Easital Technologies Ltd. sets up language models that run inside infrastructure you control: a server in your office or data center, or your own cloud account. Prompts and documents stay in your network, because no outside model API is called. Easital has set up local and on-premise model environments and has trained and fine-tuned models.

What a private LLM is

A private LLM is a language model that runs on infrastructure you control, so that prompts, documents and outputs are not sent to a model provider.

Several terms overlap. A local LLM usually means a model running on a single machine, such as a workstation, for one person. An on-premise LLM runs on your organization’s own servers and serves many users. A self-hosted LLM covers both, and also a model you run in your own cloud account. All of them use open-weight models: models whose weights are published for download under a license that lets you run them yourself. License terms differ by model, and we check them for your use.

Who needs a private LLM

Four situations justify running your own model: data residency, confidentiality, cost at volume, and the need for a model that is close by or offline.

  • Data residency

    A law, a regulator or a customer contract requires that data stays in a given country or facility.

  • Confidentiality

    Client files, contracts, source code, health or financial records that you may not send to an outside processor.

  • Cost at volume

    Steady, high token volume, where a fixed hardware cost works out lower than the per-token bill. See LLM cost optimization for the break-even reasoning.

  • Latency and availability

    The model sits next to the application, with no dependence on an internet link or a provider’s rate limits, including at sites with no outside connection.

A hosted API is the better choice when volume is low or irregular, when the task needs the largest models available, or when nobody on your side can operate a server.

Private LLM vs hosted API

A hosted API is quicker to start and gives access to larger models. A private LLM keeps data in your network and turns a usage-based bill into a fixed one. The two can be combined.

Hosted model API and private LLM compared
Hosted model APIPrivate LLM
Where data goesTo the provider, under its termsStays in your network
Cost shapePer token; rises with useFixed hardware and operations, at any level of use
Effort to startAn API keyHardware, setup and testing
Model choiceThe provider’s models, including the largestOpen-weight models that fit your hardware
CapacityThe provider’s rate limitsWhat your hardware can serve
Model changesThe provider updates and retires models on its scheduleYou decide when a model changes
Who operates itThe providerYour team, or Easital on your behalf
Fits bestLow or variable volume; tasks that need the largest modelsSensitive data; steady high volume; offline sites

What a deployment includes

A deployment has six parts. Together they turn a model that runs on one engineer’s machine into a service a team can rely on.

  • Hardware sizing

    GPU memory and the number of GPUs, or whether a CPU is enough, worked out from the model size, the context length, the number of simultaneous users and the response time you expect.

  • Model choice

    A shortlist of open-weight models that fit the hardware and the license terms, tested on your own tasks. This includes quantized versions, which use less memory at some cost in quality.

  • Serving

    An inference server that exposes the model to your applications through an API, handles several requests at once, streams responses and queues requests under load.

  • Access control

    Sign-in or keys per user and per team, network rules that limit who can reach the server, and a log of usage. Outbound connections are closed and checked.

  • Monitoring

    GPU and memory use, response time, queue length and errors, with alerts, and usage figures per team for capacity planning.

  • Updates

    A procedure for testing a new model version against your evaluation set, rolling it out and rolling it back.

Where the model has to answer from your documents, the deployment is combined with RAG development. Where a general model is not accurate enough for your domain, we fine-tune one.

How a private LLM deployment runs

A deployment runs in seven steps. Models are tested before any hardware is bought, because the model decides the hardware.

  1. Define the workload

    The tasks, the number of users, peak simultaneous requests, the response time needed, the data rules and where the system must run.

  2. Build the evaluation set

    Real examples of each task with acceptable outputs. This set decides which model is good enough.

  3. Shortlist and test models

    Candidates are run against the evaluation set on rented or existing hardware, and quality, speed and memory use are recorded.

  4. Size and provision hardware

    With a model chosen, we specify the hardware for your peak load, and you buy, rent or reuse it.

  5. Install serving and security

    The inference server, access control, logging and network isolation, followed by a check that nothing leaves the network.

  6. Connect applications and load test

    Your applications are connected and the service is tested at the expected number of simultaneous users.

  7. Hand over or operate

    Your administrators receive a runbook, the dashboards and the update procedure, or Easital operates the service for you.

Proof from our own work

A private model environment is by its nature not public, so there is no live link to show for one. The two products below show related work that can be checked.

  • Easital product

    Manob.ai

    Built and run by Easital

    Manob.ai gives each project a cloud sandbox with live preview and one-click deployment. Provisioning and isolating compute for many users is the same kind of infrastructure work that a model server needs.

  • Easital product

    StepVideo

    Built and run by Easital

    StepVideo redacts secrets on the user’s own device before a recording is uploaded: private keys, access tokens, API keys, card numbers and email addresses. What may leave a machine is also the central question of a private deployment.

See all of our work

How to choose a local LLM

There is no single best local LLM. The right one is the largest model that fits your hardware, passes your own test set and carries a license that allows your use.

New open-weight models are released often, so a named recommendation goes out of date quickly, and this page names none. The decision follows five checks:

  • Task. Chat, extraction, summarization, code and embeddings have different needs, and a smaller model can be enough for a narrow task.
  • Memory. The weights have to fit in GPU memory, or in system memory on a CPU, with room left for the context. Quantization reduces the memory needed.
  • License. Check that commercial use, and your specific use, are permitted.
  • Languages and context length. Check both against your real documents and users.
  • Your own test. Public leaderboards measure general benchmarks. Your documents and questions are the test that counts.

To run an LLM locally on one machine, you need a model file, a runtime that loads it and enough memory. That is a good way to compare models. It is not a production service: it has no access control, no handling of many users at once and no monitoring.

What a private deployment does not solve

Moving the model inside your network removes the outside processor. It does not make the application around the model secure, and it does not make you compliant by itself.

The OWASP Top 10 for LLM Applications 2025 lists risks such as prompt injection and sensitive information disclosure. Both depend on how the application is built: what the model is allowed to read, which users can ask what, and how its output is used. They apply to a private model as much as to a hosted one, and we design access rules and output checks for them. Whether a deployment meets a regulation is a decision for your compliance team.

Technology we use for private LLMs

Everything here is listed by category. The model and the serving software are chosen by testing at the time of the project, and the options change quickly.

Models
  • Open-weight language models
  • Embedding models
  • Fine-tuned models
  • Quantized versions
Hardware
  • GPU servers on your premises
  • GPU instances in your cloud account
  • CPU-only servers for small models
Serving
  • Inference server with request batching
  • API gateway
  • Load balancing across GPUs
Security
  • Network isolation
  • Single sign-on or API keys
  • Role-based access
  • Usage logs

Ways to work with us

There are three ways to engage Easital, depending on how far along you are.

  • Feasibility study

    We define the workload, test a shortlist of models on your examples and deliver a hardware estimate with a cost comparison against a hosted API.

    Best for: deciding whether a private LLM is worth it, before any hardware is bought.

  • Full deployment

    The seven steps above, from workload to a monitored service, with a handover to your administrators.

    Best for: organizations with a clear requirement to keep data in their own network.

  • Work on an existing setup

    For a model you already run: more throughput, access control, monitoring, an update procedure, or fine-tuning on your data. Related work is described under LLM development.

    Best for: a local model that worked as a trial and now has to serve a team.

Private and local LLMs: questions and answers

What is the difference between a private LLM, a local LLM and an on-premise LLM?

A private LLM is any language model that runs on infrastructure you control. A local LLM usually means one running on a single machine for one person. An on-premise LLM runs on your organization’s own servers for many users. Self-hosted LLM is the general term and also covers a model in your own cloud account.

What is the best local LLM?

There is no single answer, and a named model would soon be out of date. The right local LLM is the largest one that fits your hardware’s memory, passes a test on your own tasks and has a license that permits your use. We run that comparison for you and do not recommend a model before testing it.

How do we run an LLM locally?

On one machine you need three things: a downloaded open-weight model, a runtime that loads it, and enough GPU or system memory to hold it. For a team you also need an inference server, access control, monitoring and an update procedure, which is what a deployment adds.

What hardware does a private LLM need?

It depends on the model’s size, the context length and how many people use it at once. GPU memory is usually the limiting factor. Small models can run on a CPU. We specify hardware after testing models on your tasks.

Are open-weight models good enough to replace a hosted model?

For many well-defined tasks, such as extraction, classification, summarization and answering from your own documents, a tested open-weight model can meet the bar. For the hardest reasoning tasks, hosted providers offer larger models than most organizations can run. Your evaluation set decides. A mixed setup is also possible, with sensitive work on the private model and the rest on an API.

Does a private LLM make us compliant with data protection rules?

Not by itself. It removes one outside processor and lets you show where data is processed, which helps with residency and confidentiality requirements. Compliance also depends on access control, retention, logging and your own policies, and the judgment belongs to your compliance team.

Tell us what has to stay inside your network

Describe the data, the tasks and the number of people who would use the model. We write back with questions and a proposal for a feasibility study.

Easital is an AI and SaaS engineering company that takes AI software to production, and runs AI products of its own. Founded in 2019.