AgentCostAI

LLM API Pricing Comparison by Workload

Compare model prices using the workload that actually drives your bill: average input tokens, generated output, cached input and request volume. The calculator normalizes the dated, source-linked pricing data into estimated cost per request and monthly API spend, helping you identify models worth testing rather than choosing on headline price alone.

Interactive LLM cost calculator

Estimate monthly model cost from request volume, token usage, and your provider's per-million-token prices. Enter the prices from the provider you actually use.

Compare normalized LLM API prices

Provider price pages do not always present costs in the same way. Input tokens, output tokens and cached input may have separate rates, while a low input price can be offset by expensive generation. This comparison converts the displayed rates to a common cost-per-million-token basis and applies them to one workload. Review the source link and last-checked date attached to each price before making a purchasing or routing decision, because providers can change prices, model names and billing rules.

Calculator inputs and what they mean

Enter average input tokens per request, average output tokens per request, the percentage of input served from cache, and expected monthly requests. Input tokens include instructions, conversation history, retrieved context and other text sent to the model. Output tokens are the tokens the model generates, not the maximum output limit in your request. Cached percentage should reflect the reusable input that qualifies for the provider’s discounted cached-token rate. Monthly requests should include retries and background calls if those requests are billed.

How cost per request is calculated

For a model with prices quoted per one million tokens, uncached input cost equals input tokens × (1 − cache rate) × input price ÷ 1,000,000. Cached input cost equals input tokens × cache rate × cached-input price ÷ 1,000,000. Output cost equals output tokens × output price ÷ 1,000,000. Estimated cost per request is the sum of those three amounts. Estimated monthly cost equals cost per request × monthly requests. Where a pricing category is not available or does not apply, use the calculator’s displayed assumptions and verify the provider’s billing documentation.

Worked workload example

Suppose an application averages 2,000 input tokens and 500 output tokens per request, processes 100,000 requests each month, and expects 25% of its input to qualify for cached pricing. The monthly workload contains 200 million input tokens and 50 million output tokens. Of the input volume, 50 million tokens are treated as cached and 150 million as uncached. The calculator multiplies each volume by the corresponding model rate, then adds the components. This makes it possible to compare models without relying on a misleading single-token price.

How to choose a lower-cost model

Start by sorting for the lowest estimated monthly cost, but treat that result as a shortlist rather than an automatic winner. Test candidates on representative production tasks and reject any model that misses required accuracy, structured-output reliability, latency, context length or tool-use behavior. For short classification and extraction jobs, input price may dominate. For content generation or long answers, output price can matter more. Prompt-heavy workloads with stable instructions may benefit from cached-input pricing, but only when the provider’s cache eligibility rules match the request pattern.

Model mixed workloads separately

A single average can hide expensive traffic. Separate high-volume simple tasks from low-volume complex tasks—for example, classify support tickets with one workload and draft detailed responses with another. Also model agent loops independently: one user action may trigger planning, retrieval, tool calls, retries and a final response. Add the monthly estimates across workload classes to produce a more defensible forecast. If quality testing shows that only some requests need a premium model, route those cases selectively instead of pricing every request at the premium rate.

Assumptions and limitations

Results are estimates, not invoices. Actual charges can differ because of tokenization, hidden or reasoning tokens, cache write rules, minimum charges, batch discounts, volume tiers, regional pricing, image or audio inputs, tool execution, failed requests, taxes and currency conversion. Average token counts also conceal variance and traffic spikes. Use observed API usage where possible, test low, typical and high-volume scenarios, and confirm current prices from the linked provider source before committing a budget.

Turn a pricing comparison into ongoing cost control

A spreadsheet-style comparison shows which models may cost less for a defined workload; production control requires applying that decision to real traffic. AgentCost can route AI requests toward cost-effective configured models, enforce daily and monthly budgets, and record requests, spend and estimated savings per client or agent. Its OpenAI-compatible chat completions endpoint lets teams use a familiar request shape while adding a controlled path for AI-agent workloads. Savings are not guaranteed and depend on workload, provider pricing, model selection and routing configuration.

Frequently asked questions

What is the best way to compare LLM API pricing?

Use the same input-token count, output-token count, cache ratio and request volume for every model. Compare estimated workload cost, then test the lower-cost candidates for quality, latency and operational requirements.

Why are input and output token prices compared separately?

Models often charge different rates for tokens sent to the API and tokens generated in the response. Two models with similar input pricing can have materially different costs for output-heavy workloads.

Does the cheapest model always produce the lowest total cost?

No. A cheaper model may require longer prompts, more retries or additional calls to reach acceptable quality. Include the complete workflow and measure successful-task cost, not just the price of one request.

How should cached tokens be estimated?

Use the share of input that repeatedly qualifies for the provider’s cache discount. If you do not have production evidence, run a conservative scenario with no cached input and compare it with a realistic expected-cache scenario.

How often should LLM price comparisons be updated?

Recheck prices before procurement, major routing changes or budget planning. Always review the source link and last-checked date in the dataset because provider rates and billing conditions can change.

Can AgentCost prevent an AI agent from exceeding its budget?

AgentCost supports daily and monthly budget limits. Configured controls can reject requests after a limit is reached, while the portal records usage, cost and estimated savings by client or agent.

Apply cost-aware routing to production AI traffic

Use the comparison to identify models worth testing, then connect workloads through AgentCost to route requests, enforce client or agent budgets, and monitor actual spend and estimated savings.

Explore AgentCost →