There is no single best local LLM. The right one is the largest model that fits your hardware, passes your own test set and carries a license that allows your use.
New open-weight models are released often, so a named recommendation goes out of date quickly, and this page names none. The decision follows five checks:
- Task. Chat, extraction, summarization, code and embeddings have different needs, and a smaller model can be enough for a narrow task.
- Memory. The weights have to fit in GPU memory, or in system memory on a CPU, with room left for the context. Quantization reduces the memory needed.
- License. Check that commercial use, and your specific use, are permitted.
- Languages and context length. Check both against your real documents and users.
- Your own test. Public leaderboards measure general benchmarks. Your documents and questions are the test that counts.
To run an LLM locally on one machine, you need a model file, a runtime that loads it and enough memory. That is a good way to compare models. It is not a production service: it has no access control, no handling of many users at once and no monitoring.