AgentCostAI

LLM Model Routing Savings Calculator

Estimate how much your team could save by keeping complex requests on the current LLM while routing eligible work to a lower-cost alternative. Enter monthly request volume, average tokens per request, model prices and the share of traffic that can be rerouted. The calculator returns baseline cost, projected routed cost, absolute savings and savings percentage.

LLM routing savings calculator

Estimate the monthly savings from routing eligible requests to a lower-cost model.

What this model-routing calculator measures

The calculator compares two scenarios using the same monthly workload. The baseline scenario sends every request to the current model. The routed scenario sends the eligible share to an alternative model and leaves the remaining traffic on the current model. This isolates the financial effect of model selection. It is useful for evaluating repetitive or lower-complexity tasks such as classification, extraction, short summaries, query rewriting and structured responses. It should not be treated as evidence that two models provide equal quality; model suitability must be tested separately.

Inputs and how to choose them

Monthly requests: Use completed LLM requests from a representative month. If traffic is growing quickly, calculate a current-volume case and a forecast case rather than mixing both into one estimate. Average tokens per request: Include the tokens covered by the price used in your calculation. Averages should reflect production traffic, not a single prompt. Long system prompts, conversation history and generated output can materially change the result. Current and alternative model prices: Enter comparable prices in the same currency and pricing unit. If your provider publishes prices per one million tokens, both model rates must use that unit. When input and output tokens have different prices, derive a workload-weighted blended rate or calculate the two components separately before combining them. Eligible traffic percentage: Enter the share of requests that can move to the alternative model without violating your quality, latency, safety or contractual requirements. Do not assume every request is eligible simply because the alternative model is cheaper.

Formulas used by the calculator

Let R be monthly requests, T be average tokens per request, Pc be the current-model price per one million tokens, Pa be the alternative-model price per one million tokens, and E be the eligible routing share expressed as a decimal. Baseline monthly cost = R × T ÷ 1,000,000 × Pc. Routed monthly cost = (R × (1 − E) × T ÷ 1,000,000 × Pc) + (R × E × T ÷ 1,000,000 × Pa). Monthly savings = baseline monthly cost − routed monthly cost. Savings percentage = monthly savings ÷ baseline monthly cost × 100. Projected annual savings = monthly savings × 12. The annual figure assumes request volume, token usage, prices and routing eligibility remain constant for twelve months.

Worked model-routing example

Consider a hypothetical workload with 1,000,000 monthly requests, 1,500 average tokens per request, a current-model rate of €10 per one million tokens and an alternative-model rate of €2. Assume testing shows that 60% of requests are eligible for rerouting. The baseline is 1,000,000 × 1,500 ÷ 1,000,000 × €10 = €15,000 per month. Under routing, 40% remains on the current model and costs €6,000. The 60% alternative-model share costs €1,800. Total routed cost is therefore €7,800 per month. Projected savings are €7,200 per month, or 48% of the baseline. If the workload remains stable, the annualized estimate is €86,400. These rates are illustrative inputs, not claims about current provider pricing or guaranteed savings.

How to interpret the result

Baseline cost answers: “What would this workload cost if every request continued using the current model?” Routed cost estimates the model spend after applying the selected traffic split. Absolute savings shows the potential budget difference, while savings percentage makes scenarios with different workload sizes easier to compare. A large theoretical saving is not automatically the best routing policy. Start with requests that have clear acceptance criteria and low failure costs. Keep ambiguous, high-stakes or tool-intensive requests on the stronger model until evaluation data supports a change. If the alternative model causes retries, longer outputs or escalation to a premium model, include those effects in the average token and eligibility assumptions.

A practical decision process for setting eligible traffic

First, group requests by task rather than routing a random percentage of all traffic. For example, separate extraction, summarization, support drafting and complex reasoning. Next, build an evaluation set containing common cases, difficult cases and known failures. Compare output accuracy, format compliance, latency and safety for the current and alternative models. Route only the task groups that meet your acceptance threshold. Begin with a limited share, monitor failures and calculate the effective cost per successful result—not only cost per request. Increase eligibility when production evidence supports it. Reduce or disable routing for task groups that produce excessive retries, manual review or customer-facing errors.

Assumptions and limitations

This is a planning estimate, not a provider invoice forecast or a savings guarantee. Actual spend can differ because of separate input and output rates, cached-token discounts, reasoning tokens, batch pricing, minimum charges, currency conversion, changing provider prices and taxes. The estimate also assumes one average token count for routed and non-routed requests. Traffic may vary by day, client or agent, and the alternative model may generate a different number of output tokens. Provider availability and rate limits can also change the final routing mix. For a more conservative business case, test lower eligible-traffic percentages and higher token averages. Recalculate whenever workload composition or model pricing changes.

Move from a scenario to controlled production routing

A spreadsheet-style estimate helps identify whether routing is worth testing, but production control requires request-level rules and ongoing measurement. AgentCost can route AI requests, apply daily and monthly budget limits, separate usage with client API keys, and report requests, spend and estimated savings by agent or client. Its OpenAI-compatible chat completions endpoint is designed to keep the request pattern familiar while adding cost controls between your application and configured providers.

Frequently asked questions

What percentage of LLM traffic should I reroute?

Use the percentage that passes task-specific quality and safety tests. Begin with well-defined, lower-complexity requests and a limited production share. Do not use a universal percentage across unrelated agents or workflows.

How do I handle separate input and output token prices?

Calculate input and output costs separately using representative token averages, then add them. Alternatively, derive a blended effective price per one million total tokens from your actual input-output mix and use that same method for both models.

Does the annual savings estimate account for growth?

No. It multiplies the monthly estimate by twelve and therefore assumes stable volume, token usage, prices and routing eligibility. Run separate scenarios for expected growth or seasonal traffic.

Can a cheaper model increase total cost?

Yes. Total cost can rise if the alternative produces longer outputs, causes retries, requires more manual review or frequently escalates to the original model. Evaluate cost per successful task as well as price per token.

Are the calculated savings guaranteed?

No. The result depends entirely on the supplied workload, pricing and routing assumptions. Actual savings vary with model performance, provider pricing, token usage and production routing behavior.

Apply your routing scenario with spend controls

Use AgentCost to route configured AI workloads, enforce daily and monthly budgets, and review spend and estimated savings by client or agent. Connect through the OpenAI-compatible chat completions endpoint and turn the calculator’s assumptions into a measurable production policy.

Explore AgentCost →