LLM cost optimization

LLM cost optimization: measured per request, feature and user

Easital Technologies Ltd. reduces what a product spends on language models. We measure where the tokens go, then change what is sent, which model answers and how often it is called. Every change is checked against an evaluation set, so that a cheaper setup does not become a worse product. Easital has done this work on its own platform, Manob.ai.

What LLM cost optimization is

LLM cost optimization is the work of lowering what an application pays for language-model usage per request, per user or per completed task, while keeping the quality of the output at a level you have agreed and can measure.

Hosted model APIs bill by the token, a unit of roughly a word fragment. Input tokens (everything you send) and output tokens (everything the model writes) are priced separately, and each model has its own prices. The LLM token cost of one call is therefore input tokens times the input price plus output tokens times the output price.

A product’s bill is that figure multiplied by the number of calls, and four numbers decide it: tokens in, tokens out, the price of the model and the number of calls per task. Every technique on this page changes one of the four. The same reasoning applies whether you want to reduce OpenAI API costs or the bill from any other provider.

What an LLM cost calculator can and cannot tell you

An LLM cost calculator multiplies token counts by a model’s published prices. That gives the price of one call. It does not tell you what your product will spend.

An LLM cost comparison table is a reasonable way to shortlist models, and provider pricing pages are the place to check current prices, which change often. A calculator leaves out four things:

  • Your real input size. The system prompt, the conversation history, retrieved documents and tool definitions are all sent, and billed, on every call.
  • Calls per task. An agent may call the model many times to finish one job, and a failed step may be retried.
  • Discounts you are not using. Cached input and batch processing are billed at lower rates by several providers.
  • Quality. A cheaper model that needs two attempts, or gives answers users reject, is not cheaper.

So we build the calculator from your own logs: cost per request, per feature, per user and per successful task.

What the service covers

The work has six parts. The first two are measurement, and they come before any change.

  • Usage instrumentation

    Every model call is logged with its feature, user or tenant, model, input tokens, cached tokens, output tokens, duration and outcome.

  • Evaluation set

    Real inputs with acceptable outputs for the features that cost the most, so that “still good enough” becomes a test that can be run.

  • Model routing

    A rule or a small classifier sends simple requests to a smaller model and keeps the larger model for requests that need it.

  • Prompt and context reduction

    Shorter instructions, trimmed or summarized history, fewer and better-chosen retrieved passages, and only the tool definitions a step can use.

  • Caching and batching

    Prompts are ordered so the repeated part can be cached, repeated questions are served from stored answers, and work that can wait is sent as batch jobs.

  • Metering, quotas and alerts

    Usage limits per user or tenant, credit systems for usage-based plans, caps on spend per run and alerts when spending departs from the normal pattern.

How a cost optimization project runs

A project runs in six steps: measurement first, cheap changes before structural ones, and a quality check after every change.

  1. Measure

    We add logging to every model call, or read the logs you have, until cost can be shown per feature, per user and per task.

  2. Rank where the money goes

    Features and call patterns are sorted by spend. The list shows which changes are worth making.

  3. Build the evaluation set

    For each expensive feature we collect real inputs and agree with you what an acceptable output is. The current system’s score on this set is the baseline.

  4. Apply the levers in order of effort

    First the changes that leave the model alone: remove unused context, cap output length, fix retry loops, turn on caching. Then routing and batching, and changes to the architecture only if they are still needed.

  5. Verify each change

    Each change is run against the evaluation set and compared on cost per successful task. A change that lowers quality below the agreed level is reverted.

  6. Set budgets and hand over

    Quotas, spend caps and alerts go live, and you receive a written before-and-after report taken from your own logs, with the dashboards and a runbook.

Proof from our own work

Each of these products runs language or speech models on every use, which only works as a business when the cost of one task is known.

  • Easital product

    Manob.ai

    Built and run by Easital

    On Manob.ai a multi-step agent task uses more of the model than a simple prompt, so cost has to be measured and controlled per task. Easital carried out the token-cost optimization on this platform.

  • Easital product

    StepVideo

    Built and run by Easital

    An AI pipeline in which language and speech models from more than one provider handle each recording. It is sold by subscription, so the cost of processing one recording has to be known.

  • Client project

    Calldone

    Built by Easital for a client in the United States

    A voice platform where three costs run at once during a call: speech-to-text, a language model and text-to-speech. Cost per minute of conversation is the figure to control.

See all of our work

The levers: what each one changes and what it risks

There are nine levers. Each has a risk that has to be managed, which is why evaluation is on the list.

LLM cost levers, what each changes and the risk to manage
LeverWhat it changesRisk to manage
Measure firstNothing yet. It shows cost per feature, user and task.Skipping it means optimizing by guesswork.
Model routingSimple requests go to a smaller, cheaper model.A hard request sent to the small model gets a weaker answer.
Prompt and context sizeFewer input tokens on every call.Cutting context the model needed.
CachingRepeated input is billed at a lower rate, or a stored answer is reused.Stale answers, or an answer shown to the wrong user.
BatchingWork that can wait runs as discounted batch jobs.Results arrive later.
Output limitsFewer output tokens: length caps and structured formats.Answers cut off before they are complete.
Retries and loopsFewer calls per task: step caps and stop conditions.A task stopped before it is finished.
EvaluationLets a cheaper model or prompt be trusted.A test set that does not match real traffic.
Self-hostingPer-token fees become a fixed hardware cost.Idle capacity below break-even, plus operations work.

Retries, loops and spend caps

An agent that never decides it is finished, a retry without a limit or a public endpoint without a quota keeps spending until someone notices. Cost control is also a security matter.

The OWASP Top 10 for LLM Applications 2025 lists this as Unbounded Consumption. It describes “Denial of Wallet” as an attack in which a high volume of operations is used to “exploit the cost-per-use model of cloud-based AI services, leading to unsustainable financial burdens”.

The controls are simple: a cap on steps and spend per run, a limit on retries, quotas per user and per tenant, rate limits on public endpoints and an alert when spend departs from the usual pattern.

When self-hosting breaks even

Self-hosting a model pays off when your steady token volume is high enough that a fixed monthly hardware cost is lower than the per-token bill for the same work, at the quality you need.

Three things decide it. Utilization: a GPU that sits idle most of the day costs the same as a busy one. Quality: the open-weight model that fits your hardware has to pass your evaluation set. Operations: someone has to run, monitor and update the server. Data residency or confidentiality can justify self-hosting below the break-even volume. See private LLM deployment.

Technology we work with

The work is done in your stack and with your providers. The techniques apply to any provider, hosted or self-hosted.

Model providers
  • Commercial model APIs
  • Open-weight models, hosted or on-premise

We have worked with all major commercial model providers and with open-weight models, so an audit starts from whichever ones you already pay for.

Cost levers
  • Prompt caching
  • Response caching
  • Batch processing
  • Model routing
Measurement
  • Token and cost logs per call
  • Cost dashboards
  • Evaluation sets
  • Budget alerts

Ways to work with us

There are three ways to engage Easital on cost.

  • Cost audit

    With read access to your usage logs and the code that calls the model, we deliver a written breakdown of spend by feature and call pattern, and a list of changes ranked by effect and risk.

    Best for: a bill that has grown faster than usage, with no clear explanation.

  • Implementation

    We make the changes in your codebase, verify each against the evaluation set and report the before-and-after figures from your logs.

    Best for: teams that know what should change and lack the time to do it.

  • Cost control in a new build

    Metering, quotas, routing and caching designed in from the start, as part of an LLM development or AI agent project.

    Best for: products about to launch usage-based or credit-based pricing.

LLM cost: questions and answers

How do we reduce OpenAI API costs?

Start by logging cost per feature, then work in this order: remove input the model does not need (long system prompts, full history, extra retrieved passages, unused tool definitions), cap output length, turn on prompt caching, send simple requests to a smaller model, move non-urgent work to batch processing, and limit retries and agent loops. Check quality against an evaluation set after each change.

How is LLM token cost calculated?

For one call: input tokens times the model’s input price, plus output tokens times its output price. Input includes the system prompt, the conversation history, retrieved documents and tool definitions, and it is billed again on every call. A product’s LLM cost is that figure multiplied by calls per task and by the number of tasks.

Will a cheaper model make our product worse?

It can, and that is what the evaluation set is for. We run the cheaper model or the shorter prompt against real examples with agreed acceptable outputs. If it fails on part of the traffic, routing sends it only the requests it handles well.

How much can we save?

We do not quote a percentage before measuring. The answer depends on how much unused context is sent, how many calls a task takes and how much of the traffic a smaller model can handle. The cost audit produces figures from your own logs.

When is self-hosting an LLM cheaper than an API?

When usage is steady and high enough to keep the hardware busy, and an open-weight model passes your quality tests. With low or irregular volume, per-token pricing is usually cheaper because you pay nothing while idle. See private LLM deployment.

What access do you need to our systems?

Read access to usage logs or the provider’s usage export, the prompts, and the code that calls the model. Where logs contain customer data, we can work from redacted or sampled records. Access and confidentiality terms are agreed in the contract.

Tell us what your model bill looks like today

Share which features call a model, which providers you use and how the bill has moved. We write back with questions and a proposal for an audit.

Easital is an AI and SaaS engineering company that takes AI software to production, and runs AI products of its own. Founded in 2019.