Easital product
Manob.ai
Built and run by Easital
On Manob.ai a multi-step agent task uses more of the model than a simple prompt, so cost has to be measured and controlled per task. Easital carried out the token-cost optimization on this platform.
LLM cost optimization
Easital Technologies Ltd. reduces what a product spends on language models. We measure where the tokens go, then change what is sent, which model answers and how often it is called. Every change is checked against an evaluation set, so that a cheaper setup does not become a worse product. Easital has done this work on its own platform, Manob.ai.
LLM cost optimization is the work of lowering what an application pays for language-model usage per request, per user or per completed task, while keeping the quality of the output at a level you have agreed and can measure.
Hosted model APIs bill by the token, a unit of roughly a word fragment. Input tokens (everything you send) and output tokens (everything the model writes) are priced separately, and each model has its own prices. The LLM token cost of one call is therefore input tokens times the input price plus output tokens times the output price.
A product’s bill is that figure multiplied by the number of calls, and four numbers decide it: tokens in, tokens out, the price of the model and the number of calls per task. Every technique on this page changes one of the four. The same reasoning applies whether you want to reduce OpenAI API costs or the bill from any other provider.
An LLM cost calculator multiplies token counts by a model’s published prices. That gives the price of one call. It does not tell you what your product will spend.
An LLM cost comparison table is a reasonable way to shortlist models, and provider pricing pages are the place to check current prices, which change often. A calculator leaves out four things:
So we build the calculator from your own logs: cost per request, per feature, per user and per successful task.
The work has six parts. The first two are measurement, and they come before any change.
Every model call is logged with its feature, user or tenant, model, input tokens, cached tokens, output tokens, duration and outcome.
Real inputs with acceptable outputs for the features that cost the most, so that “still good enough” becomes a test that can be run.
A rule or a small classifier sends simple requests to a smaller model and keeps the larger model for requests that need it.
Shorter instructions, trimmed or summarized history, fewer and better-chosen retrieved passages, and only the tool definitions a step can use.
Prompts are ordered so the repeated part can be cached, repeated questions are served from stored answers, and work that can wait is sent as batch jobs.
Usage limits per user or tenant, credit systems for usage-based plans, caps on spend per run and alerts when spending departs from the normal pattern.
A project runs in six steps: measurement first, cheap changes before structural ones, and a quality check after every change.
We add logging to every model call, or read the logs you have, until cost can be shown per feature, per user and per task.
Features and call patterns are sorted by spend. The list shows which changes are worth making.
For each expensive feature we collect real inputs and agree with you what an acceptable output is. The current system’s score on this set is the baseline.
First the changes that leave the model alone: remove unused context, cap output length, fix retry loops, turn on caching. Then routing and batching, and changes to the architecture only if they are still needed.
Each change is run against the evaluation set and compared on cost per successful task. A change that lowers quality below the agreed level is reverted.
Quotas, spend caps and alerts go live, and you receive a written before-and-after report taken from your own logs, with the dashboards and a runbook.
Each of these products runs language or speech models on every use, which only works as a business when the cost of one task is known.
Easital product
Built and run by Easital
On Manob.ai a multi-step agent task uses more of the model than a simple prompt, so cost has to be measured and controlled per task. Easital carried out the token-cost optimization on this platform.
Easital product
Built and run by Easital
An AI pipeline in which language and speech models from more than one provider handle each recording. It is sold by subscription, so the cost of processing one recording has to be known.
Client project
Built by Easital for a client in the United States
A voice platform where three costs run at once during a call: speech-to-text, a language model and text-to-speech. Cost per minute of conversation is the figure to control.
There are nine levers. Each has a risk that has to be managed, which is why evaluation is on the list.
| Lever | What it changes | Risk to manage |
|---|---|---|
| Measure first | Nothing yet. It shows cost per feature, user and task. | Skipping it means optimizing by guesswork. |
| Model routing | Simple requests go to a smaller, cheaper model. | A hard request sent to the small model gets a weaker answer. |
| Prompt and context size | Fewer input tokens on every call. | Cutting context the model needed. |
| Caching | Repeated input is billed at a lower rate, or a stored answer is reused. | Stale answers, or an answer shown to the wrong user. |
| Batching | Work that can wait runs as discounted batch jobs. | Results arrive later. |
| Output limits | Fewer output tokens: length caps and structured formats. | Answers cut off before they are complete. |
| Retries and loops | Fewer calls per task: step caps and stop conditions. | A task stopped before it is finished. |
| Evaluation | Lets a cheaper model or prompt be trusted. | A test set that does not match real traffic. |
| Self-hosting | Per-token fees become a fixed hardware cost. | Idle capacity below break-even, plus operations work. |
An agent that never decides it is finished, a retry without a limit or a public endpoint without a quota keeps spending until someone notices. Cost control is also a security matter.
The OWASP Top 10 for LLM Applications 2025 lists this as Unbounded Consumption. It describes “Denial of Wallet” as an attack in which a high volume of operations is used to “exploit the cost-per-use model of cloud-based AI services, leading to unsustainable financial burdens”.
The controls are simple: a cap on steps and spend per run, a limit on retries, quotas per user and per tenant, rate limits on public endpoints and an alert when spend departs from the usual pattern.
Self-hosting a model pays off when your steady token volume is high enough that a fixed monthly hardware cost is lower than the per-token bill for the same work, at the quality you need.
Three things decide it. Utilization: a GPU that sits idle most of the day costs the same as a busy one. Quality: the open-weight model that fits your hardware has to pass your evaluation set. Operations: someone has to run, monitor and update the server. Data residency or confidentiality can justify self-hosting below the break-even volume. See private LLM deployment.
The work is done in your stack and with your providers. The techniques apply to any provider, hosted or self-hosted.
We have worked with all major commercial model providers and with open-weight models, so an audit starts from whichever ones you already pay for.
There are three ways to engage Easital on cost.
With read access to your usage logs and the code that calls the model, we deliver a written breakdown of spend by feature and call pattern, and a list of changes ranked by effect and risk.
Best for: a bill that has grown faster than usage, with no clear explanation.
We make the changes in your codebase, verify each against the evaluation set and report the before-and-after figures from your logs.
Best for: teams that know what should change and lack the time to do it.
Metering, quotas, routing and caching designed in from the start, as part of an LLM development or AI agent project.
Best for: products about to launch usage-based or credit-based pricing.
Start by logging cost per feature, then work in this order: remove input the model does not need (long system prompts, full history, extra retrieved passages, unused tool definitions), cap output length, turn on prompt caching, send simple requests to a smaller model, move non-urgent work to batch processing, and limit retries and agent loops. Check quality against an evaluation set after each change.
For one call: input tokens times the model’s input price, plus output tokens times its output price. Input includes the system prompt, the conversation history, retrieved documents and tool definitions, and it is billed again on every call. A product’s LLM cost is that figure multiplied by calls per task and by the number of tasks.
It can, and that is what the evaluation set is for. We run the cheaper model or the shorter prompt against real examples with agreed acceptable outputs. If it fails on part of the traffic, routing sends it only the requests it handles well.
We do not quote a percentage before measuring. The answer depends on how much unused context is sent, how many calls a task takes and how much of the traffic a smaller model can handle. The cost audit produces figures from your own logs.
When usage is steady and high enough to keep the hardware busy, and an open-weight model passes your quality tests. With low or irregular volume, per-token pricing is usually cheaper because you pay nothing while idle. See private LLM deployment.
Read access to usage logs or the provider’s usage export, the prompts, and the code that calls the model. Where logs contain customer data, we can work from redacted or sampled records. Access and confidentiality terms are agreed in the contract.
Share which features call a model, which providers you use and how the bill has moved. We write back with questions and a proposal for an audit.
Easital is an AI and SaaS engineering company that takes AI software to production, and runs AI products of its own. Founded in 2019.