Private LLM Guide: When and How to Self-Host a Model

Private LLM Guide: When and How to Self-Host a Model

A private LLM is a language model that runs on infrastructure you control, such as your own servers or your own cloud account, so prompts and documents are not sent to an outside model API. It is worth doing when confidentiality, data residency, offline operation, latency or sustained high volume make a hosted API a poor fit. It takes GPU capacity sized to the model, an open-weight model whose license you have read, serving software, access control, monitoring and a plan for updates.

What is a private LLM, and how does it differ from a local LLM?

The terms overlap. All of them describe running an open-weight model yourself instead of calling a provider's API.

  • Local LLM: a model running on one machine, often a laptop or workstation.
  • On-premise LLM: a model running on servers in your office or data center.
  • Private or self-hosted LLM: the general case, which also covers GPU servers in your own cloud account.

Running locally keeps the data on the machine. Ollama's FAQ (opened October 3, 2026) states: "Ollama runs locally. We don’t see your prompts or data when you run locally."

Who needs a private LLM deployment?

Organizations whose requirements a hosted API cannot meet by contract or by design. Five reasons recur.

Confidentiality. Hosted APIs offer contractual protection. OpenAI's data controls documentation (opened October 3, 2026) states that "data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)." The same page says abuse monitoring logs "may contain certain customer content, such as prompts and responses" and are "retained for up to 30 days" by default. Zero Data Retention exists but is "subject to prior approval by OpenAI". A private deployment replaces those terms with a technical fact: the data does not leave your network.

Data residency. Under Article 44 of the GDPR, a transfer of personal data to a third country "shall take place only if" the conditions of that chapter are met. Providers now sell regional processing, with limits. OpenAI notes that "Data residency does not apply to system data" and charges "a 10% uplift" on eligible newer models. Anthropic's pricing page (opened October 3, 2026) lists "a 1.1x multiplier" for US-only inference. A server in your own facility settles the location question.

Cost at volume. A GPU costs the same per hour whether it serves one request or thousands. The break-even section below shows when that undercuts per-token pricing.

Latency. A model on your network removes the round trip to a provider. Response time then depends on your hardware. Knoop and Holtmann, posted January 14, 2026, measured an RTX 5090 delivering "3.5-4.6x higher throughput than the 5060 Ti with 21x lower latency for RAG" on the same models.

Offline operation. Sites without reliable internet access can still run a model. Ollama's FAQ notes that "Ollama can run in local only mode by disabling Ollama’s cloud features."

Teams with low or uneven volume, no one to operate servers, or a need for the largest hosted models are usually better served by an API.

What does a private LLM deployment take?

Six things: hardware, a model, serving software, access control, monitoring and an update process.

Hardware sizing

The model's weights have to fit in GPU memory, with room left for the context being processed. File size depends on parameter count and on quantization, which stores each weight in fewer bits. Hugging Face's GGUF documentation (opened October 3, 2026) lists the common 4-bit format at "4.5 bits-per-weight".

The Ollama model library, opened October 3, 2026, lists these download sizes:

Model tag Size listed Context window listed
qwen3.5:4b 3.4GB 256K
gemma4:12b 7.7GB - 8.0GB 256K
phi4 (14b) 9.1GB 16K
gpt-oss:20b 14GB 128K
qwen3.5:27b 17GB 256K
gemma4:31b 19GB - 20GB 256K
llama3.3 (70b) 43GB 128K
gpt-oss:120b 65GB 128K
qwen3.5:122b 81GB 256K

Ollama's gpt-oss page gives a concrete mapping to hardware: quantization "enables the smaller model to run on systems with as little as 16GB memory, and the larger model to fit on a single 80GB GPU."

A model that does not fit can still run with part of it in system memory. llama.cpp supports "CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity". Size the GPU for the model, the longest context you will send and the number of simultaneous users.

Open-weight model choice

There is no single best local LLM. The right model is the one that passes your own test set on hardware you can afford, under a license you can accept.

The Ollama library, opened October 3, 2026, lists these open-weight families, each with more than a million pulls:

  • Gemma from Google DeepMind (Gemma 4, in sizes up to 31b)
  • Qwen from Alibaba (Qwen 3.5, from 0.8b to 122b, and newer 3.6 and 3.8 releases)
  • Llama from Meta (Llama 3.3 at 70b, and Llama 4)
  • DeepSeek (DeepSeek-R1, from 1.5b to 671b)
  • Mistral (Mistral Small 3.2 at 24b)
  • Phi from Microsoft (Phi-4 at 14b)
  • gpt-oss from OpenAI (20b and 120b)

Choose with five checks.

  1. Task results. Run each candidate on your evaluation set. Public benchmarks and download counts measure general ability and popularity.
  2. Fit. Compare the size of the quantized model with your GPU memory.
  3. License. Terms differ. Ollama's pages describe gpt-oss as under a "Permissive Apache 2.0 license" and DeepSeek-R1 weights as "licensed under the MIT License", while Llama 3.3 is governed by the "Llama 3.3 Community License" and an acceptable use policy.
  4. Features. Check context length, tool calling, image input and language coverage against your use case.
  5. Upkeep. Prefer a family with regular releases and wide support in serving software.

Serving software

Three open-source tools cover most deployments.

Tool What it is Typical use
Ollama A model runner with a model library and "a REST API for running and managing models" One machine, small teams, prototypes
llama.cpp A "Plain C/C++ implementation without any dependencies" with an OpenAI-compatible server, where "Apple silicon is a first-class citizen" Laptops, CPUs, edge devices, embedded use
vLLM "a fast and easy-to-use library for LLM inference and serving" with "Continuous batching of incoming requests" and an OpenAI-compatible API server Many concurrent users on GPU servers

An OpenAI-compatible endpoint lets existing application code switch between a hosted API and your own server by changing the base URL.

Access control

Access control is your job. Ollama "binds 127.0.0.1 port 11434 by default", which keeps it private until someone changes the bind address to share it.

Measurements show that many people do. Cisco researchers reported on September 1, 2025: "we identified 1,139 vulnerable Ollama instances", of which 214 were answering requests with live models. A year-long measurement by Xu et al., posted September 7, 2026, counted "152,137 cumulative IPs" with exposed Ollama endpoints.

Built-in keys are not enough. vLLM's security documentation (opened October 3, 2026) says its API key covers only some path prefixes and warns: "Do not rely exclusively on --api-key for securing access to vLLM." Its recommendation is to "deploy vLLM behind a reverse proxy" that allowlists endpoints. Add network isolation and per-user authentication.

Monitoring

Track latency, load and output quality. vLLM's metrics documentation describes Prometheus metrics exposed "via the /metrics endpoint", including time to first token, end-to-end request latency and the number of running requests. Rerun the evaluation set on a schedule to catch quality changes.

Updates

Plan for model and software updates from the first day. The Ollama library showed Gemma 4 updated two days before this article was checked and a Qwen release one day before. Pin the model version in production, and test a new version against the evaluation set before switching.

Software patches matter as much. Xu et al. found that across five known vulnerabilities, "only 0.43-2.90% of below-fix IPs upgraded in place".

API vs self-hosted LLM: how do they compare?

A hosted API is simpler and cheaper at low volume. Self-hosting gives control over data and cost at high steady volume, and transfers the operating work to you.

Hosted API Self-hosted
Where prompts are processed The provider's infrastructure Your servers or your cloud account
Data terms Contract and provider settings Your own network and policies
Model choice The provider's proprietary models Open-weight models
Up-front cost None Hardware purchase or reserved GPU capacity, plus setup
Pricing Per token Per GPU hour or owned hardware, regardless of use
Cost at low volume Low High, because idle capacity is still paid for
Cost at high steady volume Grows with usage Flat until capacity is reached, then you add GPUs
Operations The provider's Yours: patches, monitoring, on-call
Works offline No Yes
Model updates The provider retires versions on its schedule You decide when to change

How do you work out break-even?

Divide the fixed monthly cost of self-hosting by the API price per token. The result is the monthly volume at which the two cost the same.

A worked example with prices opened on October 3, 2026:

  • GPU. Lambda's pricing page lists a single NVIDIA H100 PCIe with 80 GB at $3.29 per GPU per hour before tax. Running it for all 730 hours of an average month costs $2,401.70.
  • API. OpenAI's pricing page lists gpt-6.1-sol at $2.00 per 1M input tokens and $10.00 per 1M output tokens. Anthropic lists Claude Sonnet 5.5 at the same $2 and $10.
  • Break-even. $2,401.70 buys about 1.2 billion input tokens or about 240 million output tokens at those prices. Break-even for one rented GPU therefore sits between roughly 240 million and 1.2 billion tokens a month, depending on the mix.

Against a low-priced model the bar is far higher. OpenAI lists gpt-6-luna at $0.10 input and $0.50 output per 1M tokens, which puts break-even between about 4.8 billion and 24 billion tokens a month.

Four things move the result.

  1. Throughput. The GPU has to be able to serve the break-even volume. Measure tokens per second for your model and hardware under realistic load.
  2. Utilization. An idle GPU costs the same as a busy one.
  3. API discounts. Anthropic's Batch API gives "a 50% discount on both input and output tokens", and cache hits are billed at "0.1x base input price". Both raise the break-even point.
  4. People. Rented or owned, the system needs engineering time for setup, monitoring and upgrades.

Owned hardware changes the arithmetic. Knoop and Holtmann estimated that on consumer GPUs, "Self-hosted inference costs $0.001-0.04 per million tokens (electricity only)", with "hardware breaking even in under four months at moderate volume (30M tokens/day)." That figure excludes staff time and compares against budget-tier APIs.

The comparison is fair only when both sides do the job equally well. Check answer quality on your evaluation set before comparing prices.

How do you run an LLM locally to try it?

Install a runner, download a model that fits your memory and call it on localhost. It is a sound first test before any hardware decision.

  1. Install Ollama or llama.cpp on a machine with a supported GPU or an Apple silicon Mac.
  2. Pull a model smaller than your available memory, starting from the size table above.
  3. Send a request to the local API. Ollama's README shows curl http://localhost:11434/api/chat.
  4. Run your evaluation questions and note quality and speed.
  5. Keep the server bound to localhost until access control is in place.

Key takeaways

  • A private LLM keeps prompts and documents inside infrastructure you control.
  • The reasons to self-host are confidentiality, data residency, cost at steady high volume, latency and offline use.
  • Model size sets the hardware. The examples above range from 3.4GB to 81GB.
  • No model is best for everyone. Test candidates from current families on your own tasks and read the license.
  • Exposure is the common security failure: researchers counted 1,139 exposed Ollama servers in 2025 and 152,137 exposed IPs over a later one-year study.
  • One rented H100 at $3.29 per hour breaks even against a $2 and $10 per 1M token API at roughly 240 million to 1.2 billion tokens a month, before staff time.

Frequently asked questions

What is the best local LLM?

It depends on the task, the hardware and the license you can accept. Current open-weight families listed by Ollama include Gemma, Qwen, Llama, DeepSeek, Mistral, Phi and gpt-oss. Test two or three that fit your memory on your own evaluation questions and choose on results.

Can you run an LLM on a laptop?

Yes, with a small or quantized model. Ollama lists qwen3.5:4b at 3.4GB, and its gpt-oss page says the 20b model can run "on systems with as little as 16GB memory". Larger models need a workstation or server GPU.

Is a private LLM more secure than an API?

It gives you control over where data goes, which settles confidentiality and residency questions. It also makes you responsible for securing the server. The exposed Ollama instances found by Cisco and by Xu et al. were self-hosted systems left open to the internet.

Is self-hosting an LLM cheaper than using an API?

Only at steady, high volume. With the October 2026 prices above, one rented H100 needs hundreds of millions of tokens a month to match a mid-priced API, and billions to match a low-priced one. Below that, the API costs less.

What hardware do you need for a private LLM?

A GPU with more memory than the model file, plus headroom for context and concurrent users. As a reference, Ollama says gpt-oss:120b fits "on a single 80GB GPU", and Lambda lists single 80 GB H100 instances and 48 GB A6000 instances.

Easital Technologies Ltd. sets up language models inside infrastructure the client controls, on local machines, on-premise servers or the client's own cloud account, and has trained and fine-tuned models. Related pages: private LLM deployment, LLM development, LLM cost optimization and RAG development.

Sources

All sources were opened and checked on October 3, 2026.

  1. Ollama, "FAQ", documentation, undated, opened October 3, 2026. https://docs.ollama.com/faq
  2. OpenAI, "Data controls in the OpenAI platform", API documentation, undated, opened October 3, 2026. https://developers.openai.com/api/docs/guides/your-data
  3. Regulation (EU) 2016/679 (GDPR), Article 44, "General principle for transfers", text as published at gdpr-info.eu. https://gdpr-info.eu/art-44-gdpr/
  4. Anthropic, "Pricing", Claude Platform documentation, undated, opened October 3, 2026. https://docs.claude.com/en/docs/about-claude/pricing
  5. Knoop, Holtmann, "Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs", arXiv 2601.09527, January 14, 2026. https://arxiv.org/abs/2601.09527
  6. Hugging Face, "GGUF", Hub documentation, undated, opened October 3, 2026. https://huggingface.co/docs/hub/gguf
  7. Ollama, model library and model pages (gemma4, qwen3.5, phi4, gpt-oss, llama3.3, deepseek-r1), as listed on October 3, 2026. https://ollama.com/library
  8. Ollama, README, GitHub, opened October 3, 2026. https://github.com/ollama/ollama
  9. ggml-org, llama.cpp README, GitHub, opened October 3, 2026. https://github.com/ggml-org/llama.cpp
  10. vLLM, documentation home, "Security" and "Production Metrics", undated, opened October 3, 2026. https://docs.vllm.ai/en/latest/
  11. Cisco, "Detecting Exposed LLM Servers: Shodan Case Study on Ollama", Cisco Blogs, September 1, 2025. https://blogs.cisco.com/security/detecting-exposed-llm-servers-shodan-case-study-on-ollama
  12. Xu, Li, Qiu, Sun, "Ollama in the Wild: A Longitudinal Measurement of Exposed Ollama LLM Endpoints at Internet Scale", arXiv 2609.07115, September 7, 2026. https://arxiv.org/abs/2609.07115
  13. Lambda, "GPU cloud pricing", opened October 3, 2026. https://lambda.ai/pricing
  14. OpenAI, "Pricing", API documentation, opened October 3, 2026. https://developers.openai.com/api/docs/pricing

Further reading

Work with Easital on an LLM feature

Send a short description of the product or the problem you want solved. We reply by email with questions and a proposed next step.

Discuss your project