AgentCostAI Home Resources Documentation Start for €29

LLM Cost Calculator for AI Workloads

Calculate a transparent baseline for an AI agent, SaaS feature or LLM workflow. Enter average input and output units, separate unit prices, monthly request volume, retry rate and expected growth to estimate current and projected spend. The calculator uses your figures rather than an unverified static model-price table.

Interactive LLM cost calculator

Estimate monthly model cost from request volume, token usage, and your provider's per-million-token prices. Enter the prices from the provider you actually use.

Estimate LLM spend from your own workload data

Use this calculator when you know—or can reasonably estimate—how much context each request sends, how much output it generates and what your provider charges. It is suitable for chat features, document processing, support agents, extraction pipelines and other request-based AI workloads. Calculate each materially different workflow separately: combining a short classification task with a long document-analysis agent can hide the real cost of both.

Inputs and how to choose them

Input units are the average billable units sent with each request, including relevant instructions, conversation history, retrieved context and user content. Output units are the average generated per response. Enter the corresponding input and output prices using the pricing basis shown in the calculator. Request volume should cover the same period as the result—normally one month. Retry rate represents extra attempts caused by timeouts, provider errors or application retry logic. Growth represents the expected increase in request volume over the projection period. Use measured averages or a representative sample where possible; maximum token limits are not the same as actual usage.

Formulas used by the calculator

When prices are quoted per one million units, input cost per request = average input units × input price ÷ 1,000,000. Output cost per request = average output units × output price ÷ 1,000,000. Base cost per request = input cost + output cost. Effective requests = initial requests × (1 + retry rate as a decimal). Estimated period cost = base cost per request × effective requests. For compounding growth, projected requests after n periods = current requests × (1 + growth rate)^n, and projected cost uses that request count with the same retry and per-request assumptions. For example, a 5% retry rate means 1.05 effective attempts per initial request under this model.

How to interpret the result

The baseline estimate answers a planning question: what would this workload cost if its average usage, selected prices and retry behavior remained consistent? Compare baseline and projected totals to a product budget, customer allowance or expected gross margin. Also inspect the input/output split. A high input share may point to large prompts, repeated history or retrieved context; a high output share may indicate verbose responses or loose generation limits. Run low, expected and high scenarios instead of treating one estimate as a guaranteed invoice.

Practical scenario and decision logic

Suppose a support agent has short user questions but repeatedly sends a long policy document and full conversation history. If the result is dominated by input cost, reducing unnecessary context may matter more than shortening answers. If retries add a meaningful amount, investigate request timeouts, error handling and duplicate submissions before changing models. If projected growth pushes spend beyond the allowable budget, consider a per-agent ceiling or routing simpler requests toward a lower-cost configured model. Recalculate after each proposed change so the comparison uses the same request volume and growth assumptions.

Export and document your assumptions

Export the calculation to CSV to preserve the inputs and results used in a forecast, client proposal or internal review. Record the provider price date, model, workload, currency and source of the usage averages alongside the export. A CSV is a planning artifact, not a substitute for provider billing data. Re-run the estimate when prices, prompts, routing rules, response lengths or traffic patterns change.

Assumptions and limitations

The estimate assumes the entered unit prices apply uniformly and that average input and output usage remains stable. It does not automatically account for provider-specific rounding, cached-input pricing, batch discounts, tiered rates, minimum charges, taxes, currency conversion, tool-call billing or other fees unless you incorporate them into your inputs. The retry calculation treats the retry rate as additional attempts relative to original requests; complex recursive retry policies may behave differently. Growth projections hold per-request usage and price constant, so they are scenarios rather than guarantees.

Turn a forecast into ongoing cost control

A calculator creates a baseline at one point in time. AgentCost can provide the operational layer after that estimate: route AI requests, apply daily and monthly budget limits, and track requests, spend and estimated savings by client or agent. Its OpenAI-compatible chat-completions endpoint lets supported applications connect through AgentCost using a server-side client key. Monitoring actual activity helps reveal when usage, retries or model selection diverge from the assumptions in the forecast.

Frequently asked questions

How do I calculate LLM cost per request?

Multiply average input units by the input unit price and average output units by the output unit price, normalizing both to the provider’s pricing basis. Add the two amounts. If prices are per one million units, divide each multiplication by 1,000,000 before adding them.

Should system prompts and conversation history count as input usage?

Include any content billed as input by the provider. Depending on the request, that can include system instructions, prior messages, retrieved documents, tool-related content and the current user message.

How should I estimate retry rate?

Divide extra retry attempts by original requests for a representative period. For example, 50 additional attempts across 1,000 original requests is a 5% retry rate. Check that duplicate submissions and application-level loops are counted consistently.

Why can the estimate differ from the provider invoice?

Actual bills may reflect changing request sizes, model prices, caching rules, rounding, discounts, failed requests, tool calls, taxes or fees not represented by the calculator. Validate estimates against provider usage and billing records.

Can I use the result to set an AI-agent budget?

Yes, as a starting point. Use an expected scenario plus an appropriate operating margin, then compare actual usage with the estimate. AgentCost supports daily and monthly budget controls for clients or agents, including rejecting requests when a configured limit is reached.

Does the calculator guarantee savings?

No. It estimates cost from the data you enter. Any savings depend on actual workload behavior, provider pricing, model choice and routing configuration.

Operationalize your LLM cost baseline

Use AgentCost to route AI requests, separate usage by client or agent, enforce daily and monthly budget limits, and review spend and estimated savings as workloads run.

Explore AgentCost