How to Reduce LLM API Costs: 9 Levers in Order

How to Reduce LLM API Costs: 9 Levers in Order

The three changes that usually cut an LLM bill the most are using a cheaper model where quality allows, caching the repeated part of every prompt, and moving non-urgent work to batch processing. The providers publish the discounts: cached input is billed at about 10% of the normal input price or less, and batch requests at 50%. Measure cost per feature first, because each lever only pays where the spend is.

Every price and documentation quote below was read on October 3, 2026. Where a page carries its own date, that date is given.

What do LLM APIs cost per token in 2026?

LLM APIs are priced per million tokens, and output costs four to six times as much as input on the models below. Within one provider, the largest model can cost 100 times the smallest.

Provider Model Input Cached input Output Batch input / output
OpenAI gpt-6-astra $10.00 $1.00 $50.00 $5.00 / $25.00
OpenAI gpt-6.1-sol $2.00 $0.10 $10.00 $1.00 / $5.00
OpenAI gpt-6-luna $0.10 $0.01 $0.50 $0.05 / $0.25
Anthropic Claude Fable 5.1 $10 $0.25 $50 $5 / $25
Anthropic Claude Opus 5.5 $4 $0.20 $20 $2 / $10
Anthropic Claude Sonnet 5.5 $2 $0.20 $10 $1 / $5
Anthropic Claude Haiku 4.5 $1 $0.10 $5 $0.50 / $2.50
Google Gemini 3.1 Pro Preview $2.00 $0.20 $12.00 $1.00 / $6.00
Google Gemini 3.8 Flash $0.75 $0.075 $3.75 $0.375 / $1.875
Google Gemini 2.5 Flash-Lite $0.10 $0.01 $0.40 $0.05 / $0.20

Prices are in US dollars per 1 million tokens, from the OpenAI API pricing page, the Anthropic pricing page and the Gemini API pricing page, last updated October 1, 2026. OpenAI rows are standard short-context rates. Gemini 3.1 Pro Preview rows apply to prompts up to 200k tokens. Google lists the Gemini 3.8 Flash prices "through December 31, 2026" and double those figures from January 1, 2027.

To estimate a bill, use this formula:

Monthly cost = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000

A feature that handles 100,000 requests a month at 2,000 input tokens and 300 output tokens uses 200 million input tokens and 30 million output tokens. At the prices above, that is $35 on gpt-6-luna, $700 on Claude Sonnet 5.5 and $3,500 on gpt-6-astra.

For a rough conversion from text, Anthropic's pricing page says "1 token is approximately 4 characters or 0.75 words in English."

Which levers reduce LLM cost the most?

Model choice, prompt caching and batch processing carry the largest published price differences. The other six levers either make those three possible or stop waste they cannot reach.

# Lever Published effect
1 Measure per-feature spend Shows where the other levers apply
2 Right-size the model Up to 100 times between models of one provider
3 Shorten prompts and context Proportional to the tokens removed
4 Prompt caching Cached input billed at 0.1 times the input price, or less
5 Batch processing 50% off input and output
6 Output limits Output costs four to six times as much as input
7 Stop loops and retries Agents use about 4 times the tokens of chat
8 Evaluate cheaper models Makes lever 2 safe to use
9 Self-hosting Pays only above a volume threshold

The order follows the size of the published price differences. Your own usage data may reorder it.

How do you measure LLM spend per feature?

Tag every request with the feature that caused it, and record input, cached and output tokens from each response. A single monthly invoice cannot tell you which feature to fix.

Give each feature its own API key or workspace so the provider's reports split the same way. Anthropic's Usage and Cost API lets you "Filter by API key, workspace, model, service tier" and group results by those dimensions.

The result is a table of cost per feature and cost per request. Work on the largest rows first.

Which model should each task use?

Each task should use the cheapest model that passes your tests for that task. Classification, extraction and routing rarely need the largest model.

Anthropic's pricing page gives the same advice for its own lineup: "Choose Haiku for simple tasks, Sonnet for most production workloads, and Opus for the most complex reasoning".

Routing research supports the approach. The FrugalGPT paper, submitted May 9, 2023, found that a cascade of models "can match the performance of the best individual LLM (e.g. GPT-4) with up to 98% cost reduction". RouteLLM, submitted June 26, 2024, reported routers that reduce costs "by over 2 times in certain cases" without compromising the quality of responses. Both used benchmark tasks and models that are now outdated, so treat the percentages as upper bounds and test on your own traffic.

How do shorter prompts and context reduce cost?

Every call resends the system prompt, tool definitions, conversation history and retrieved documents, and all of it is billed as input. Removing tokens that do not change the answer reduces cost in direct proportion.

Tool definitions and fetched content are often larger than expected. Anthropic's pricing page lists its browser toolset at "about 6,600 input tokens" per request and a fetched "Large documentation page (100 kB)" at "~25,000 tokens".

Practical steps:

  • Summarize or truncate old conversation turns.
  • Retrieve fewer, smaller passages in retrieval-augmented generation.
  • Load tool definitions only when needed. OpenAI's prompt caching guide recommends tool search with defer_loading: true "to reduce input tokens spent on tool definitions".
  • Count tokens per model. Anthropic notes that the tokenizer in Claude 4.7 and later "produces approximately 30% more tokens for the same text", and its token counting endpoint is "free to use".

How much does prompt caching save?

Prompt caching bills the repeated start of a prompt at 10% of the input price or less. It helps most when a long system prompt, tool list or document is sent on every call.

Provider Cache read Cache write Lifetime Minimum prefix
OpenAI (GPT-5.6 and later) 0.1× input; 0.05× on GPT-6.1 Sol 1.25× input 30 minutes after last write or reuse 1,024 tokens
Anthropic 0.1× input; 0.05× on Opus 5.5; 0.025× on Fable 5.1 1.25× (5-minute) or 2× (1-hour) 5 minutes, refreshed on each use, or 1 hour 512 to 4,096 tokens by model
Google Gemini Context caching price, 0.1× input on the models above Hourly storage fee for stored caches Not specified for implicit caching 2,048 or 4,096 tokens by model

OpenAI's guide says cached tokens are "discounted up to 95%" and that "Prompt caching is enabled by default for supported OpenAI models." It also gives the arithmetic: "Across ten requests, one write and nine full reads cost 2.15× at that rate, compared with 10× without caching."

Anthropic's pricing page states that "caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write)." Its prompt caching documentation adds that "The cache is refreshed for no additional cost each time the cached content is used."

Google's context caching page, last updated September 2, 2026, says "Implicit caching is enabled by default for all Gemini 2.5 and newer models."

Caches match on the exact start of the prompt. Put stable content first and anything that changes last. OpenAI's wording is "Keep the prefix stable. Put stable developer instructions and shared reference material first." A timestamp or user name near the top of a system prompt breaks every cache hit after it.

How much does batch processing save?

Batch processing costs half the standard price on all three providers, in exchange for results that arrive within hours.

  • OpenAI's Batch API guide lists a "50% cost discount compared to synchronous APIs" and says "Each batch completes within 24 hours (and often more quickly)".
  • Anthropic's batch processing documentation says "All usage is charged at 50% of the standard API prices", with "most batches finishing in less than 1 hour".
  • Google's Batch API page, last updated September 17, 2026, describes processing "at 50% of the standard cost" with a target turnaround of 24 hours.

Good candidates are evaluations, bulk classification, document processing and nightly summaries. Anthropic adds that "The pricing discounts from prompt caching and Message Batches can stack".

OpenAI also offers Flex processing, in beta, where "Tokens are priced at Batch API rates" in exchange for "slower response times and occasional resource unavailability."

Do output limits reduce cost?

Yes. Output tokens cost four to six times as much as input tokens on the models above, so shorter answers save more per token than shorter prompts.

Reasoning adds output that users never see. OpenAI's reasoning guide says reasoning tokens "are billed as output tokens", and Google's pricing page lists its output price as "including thinking tokens".

Ask for the format you need, such as a label, a JSON object or three sentences. Set a maximum output length and use lower reasoning effort for simple tasks. Set the limit with care: OpenAI warns that when the limit is reached early "you could incur costs for input and reasoning tokens without receiving a visible response."

How do loops and retries inflate the bill?

Agents call the model many times per task, and each call resends the growing context. Anthropic wrote on June 13, 2025 that "agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats."

The controls are simple:

  • A maximum number of steps per task. Anthropic's agent guidance of December 19, 2024 says it is common to include "stopping conditions (such as a maximum number of iterations) to maintain control."
  • Retries with backoff and a low cap, and no retry for errors that will fail the same way again.
  • Per-user rate limits and quotas. OWASP's 2025 list calls the abuse case "Denial of Wallet" and recommends "rate limiting and user quotas".
  • A monthly spend limit at the provider. Anthropic's rate limits documentation describes spend limits that "set a maximum monthly cost an organization can incur for API usage."

How do you know a cheaper model is good enough?

You know by running both models on a fixed test set drawn from real traffic and comparing the results. Without that, a model swap is a guess, and teams keep the expensive model to be safe.

A useful evaluation set has real inputs, expected outputs or scoring rules, and the cases that have failed before. Run it on every prompt change and every model change.

Evaluation runs are cheap. Google names "running evaluations" as a Batch API use, and OpenAI lists "model evaluations" for Flex processing, so both run at half price.

When does self-hosting an LLM break even?

Self-hosting breaks even only when a GPU stays busy enough that its fixed monthly cost is lower than the API bill for the same tokens.

Lambda's pricing page lists a single on-demand NVIDIA H100 at $3.29 to $4.29 per GPU-hour. At $4.29, one GPU running all month (730 hours) costs about $3,132.

That amount buys different volumes depending on which API model you would replace. With ten input tokens for every output token:

API model replaced Tokens that $3,132 buys per month
Claude Sonnet 5.5 About 1.04 billion input and 104 million output
gpt-6-luna About 20.9 billion input and 2.09 billion output

Below those volumes, the API is cheaper before engineering time is counted. Whether one GPU can serve that many tokens depends on the model and the serving software, so measure throughput before committing. Privacy or data-location rules can justify self-hosting at lower volumes.

Key takeaways

  • Model choice is the largest lever: published prices differ by up to 100 times within one provider.
  • Cached input is billed at 10% of the input price or less on OpenAI, Anthropic and Google.
  • Batch processing is 50% off on all three and can be combined with caching.
  • Output and reasoning tokens cost four to six times as much as input, so cap them.
  • Agents multiply token use, so set step limits, retry caps and spend limits.
  • An evaluation set is what makes a cheaper model safe to adopt.

Frequently asked questions

How do I reduce OpenAI API costs?

Move each task to the smallest model that passes your tests, keep prompt prefixes stable so cached input applies, and send non-urgent work through the Batch API or Flex processing at half price. Then cap output length and reasoning effort.

How do I calculate LLM cost per request?

Multiply input tokens by the input price and output tokens by the output price, add the two, and divide by 1,000,000. A request with 2,000 input and 300 output tokens costs $0.007 on Claude Sonnet 5.5 at $2 and $10 per million tokens.

Is prompt caching automatic?

On OpenAI and Google it is on by default for supported models. On Anthropic you add a cache_control field to the request. In all three cases the prompt must be longer than a minimum and must start with the same content each time.

Can caching and batch discounts be combined?

Yes on Anthropic, whose documentation says the discounts "can stack". Cache hits inside batches are best effort because requests run concurrently.

Is self-hosting cheaper than an LLM API?

Only at high, steady volume. One on-demand H100 at $4.29 per hour costs about $3,132 a month. At a ten-to-one input-to-output mix, that pays for roughly 1 billion input tokens on Claude Sonnet 5.5 and about 20 billion on gpt-6-luna.

Easital Technologies Ltd. measures where a product's tokens go and then changes what is sent, which model answers and how often it is called, with each change checked against an evaluation set. See LLM cost optimization, private LLM deployment and our work.

Sources

All sources were opened and checked on October 3, 2026.

  1. OpenAI, "Pricing", OpenAI API documentation, read October 3, 2026. https://developers.openai.com/api/docs/pricing
  2. Anthropic, "Pricing", Claude API documentation, read October 3, 2026. https://platform.claude.com/docs/en/about-claude/pricing
  3. Google, "Gemini Developer API pricing", last updated October 1, 2026. https://ai.google.dev/gemini-api/docs/pricing
  4. OpenAI, "Prompt caching", OpenAI API documentation, read October 3, 2026. https://developers.openai.com/api/docs/guides/prompt-caching
  5. Anthropic, "Prompt caching", Claude API documentation, read October 3, 2026. https://platform.claude.com/docs/en/build-with-claude/prompt-caching
  6. Google, "Context caching", Gemini API documentation, last updated September 2, 2026. https://ai.google.dev/gemini-api/docs/caching
  7. OpenAI, "Batch API", OpenAI API documentation, read October 3, 2026. https://developers.openai.com/api/docs/guides/batch
  8. Anthropic, "Batch processing", Claude API documentation, read October 3, 2026. https://platform.claude.com/docs/en/build-with-claude/batch-processing
  9. Google, "Batch API", Gemini API documentation, last updated September 17, 2026. https://ai.google.dev/gemini-api/docs/batch-api
  10. OpenAI, "Flex processing", OpenAI API documentation, read October 3, 2026. https://developers.openai.com/api/docs/guides/flex-processing
  11. OpenAI, "Reasoning models", OpenAI API documentation, read October 3, 2026. https://developers.openai.com/api/docs/guides/reasoning
  12. Anthropic, "Usage and Cost API", Claude API documentation, read October 3, 2026. https://platform.claude.com/docs/en/manage-claude/usage-cost-api
  13. Anthropic, "Token counting", Claude API documentation, read October 3, 2026. https://platform.claude.com/docs/en/build-with-claude/token-counting
  14. Anthropic, "Rate limits", Claude API documentation, read October 3, 2026. https://platform.claude.com/docs/en/api/rate-limits
  15. Chen, Zaharia, Zou, "FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance", arXiv 2305.05176, May 9, 2023. https://arxiv.org/abs/2305.05176
  16. Ong et al., "RouteLLM: Learning to Route LLMs with Preference Data", arXiv 2406.18665, June 26, 2024. https://arxiv.org/abs/2406.18665
  17. Anthropic, "How we built our multi-agent research system", June 13, 2025. https://www.anthropic.com/engineering/multi-agent-research-system
  18. Anthropic, "Building effective agents", December 19, 2024. https://www.anthropic.com/engineering/building-effective-agents
  19. OWASP Gen AI Security Project, "LLM10:2025 Unbounded Consumption", 2025. https://genai.owasp.org/llmrisk/llm102025-unbounded-consumption/
  20. Lambda, "GPU cloud pricing", read October 3, 2026. https://lambda.ai/pricing

Further reading

Work with Easital on an LLM feature

Send a short description of the product or the problem you want solved. We reply by email with questions and a proposed next step.

Discuss your project