AgentCostAI

AI Agent Cost per Task Calculator

Estimate what an AI-agent task really costs after model calls, input and output tokens, retries, and failed outcomes are included. Enter your workload assumptions to compare cost per run, cost per successful task, and projected model spend for 1,000 completed tasks.

Interactive LLM cost calculator

Estimate monthly model cost from request volume, token usage, and your provider's per-million-token prices. Enter the prices from the provider you actually use.

What this calculator measures

A provider’s token charge is only the starting point for agent economics. One user-visible task may trigger several model calls, and an unreliable workflow may consume additional tokens through retries or runs that never produce an accepted result. This calculator converts those factors into three practical figures: the model cost of one normal run, the expected cost of obtaining one successful task, and the expected cost of 1,000 successful tasks. Use the result for workflow design, customer pricing, model comparisons, budget planning, or deciding whether an agent is economical enough to operate at scale.

Inputs and how to enter them

Calls per run: Count every LLM request normally made during one complete agent run. A planner call, three tool-related reasoning calls, and a final answer call would equal five calls. Input and output tokens: Enter the average tokens consumed by each call. Input tokens may include instructions, conversation history, retrieved documents, tool results, and repeated context. Output tokens are generated by the model. Use measured averages where possible rather than configured maximums. Model prices: Enter the provider’s current input and output prices in the unit shown by the calculator. Model prices change, so confirm them against the provider’s pricing page or invoice. Keep the currency consistent across both fields. Retry rate: Enter the expected share of runs that require another attempt. For example, 15% means roughly 15 additional runs for every 100 initial runs under a one-retry average assumption. Task success rate: Enter the percentage of attempted tasks that produce an acceptable result. This should reflect business success, not merely an HTTP 200 response. A syntactically valid answer that fails evaluation is still an unsuccessful task.

Formula used to estimate cost per successful task

Let C be calls per run, I and O be average input and output tokens per call, PI and PO be input and output prices per million tokens, R be the retry rate as a decimal, and S be the task success rate as a decimal. Cost per model call = (I × PI ÷ 1,000,000) + (O × PO ÷ 1,000,000). Base cost per run = C × cost per model call. Retry-adjusted attempted-task cost = base cost per run × (1 + R). Estimated cost per successful task = retry-adjusted attempted-task cost ÷ S. Estimated cost for 1,000 successful tasks = cost per successful task × 1,000. Dividing by the success rate matters because failed attempts still consume tokens. At an 80% success rate, the expected workload requires 1.25 attempted tasks for each successful outcome before separately accounting for retries.

Worked example

Consider an agent that makes four calls per run. Each call averages 2,000 input tokens and 500 output tokens. Suppose input tokens cost 1 currency unit per million and output tokens cost 4 currency units per million. The cost per call is (2,000 × 1 ÷ 1,000,000) + (500 × 4 ÷ 1,000,000), or 0.004. Four calls produce a base run cost of 0.016. With a 20% retry rate, the expected attempted-task cost becomes 0.0192. If only 80% of attempted tasks succeed, the estimated cost per successful task is 0.024. At the same averages, 1,000 successful tasks would cost approximately 24 currency units in model usage. This example is illustrative rather than a statement of any provider’s current pricing. Enter the prices and observed usage for your own models.

How to interpret the result

Use cost per run when comparing the model footprint of two workflow designs under identical reliability assumptions. Use cost per successful task when evaluating unit economics, because customers and internal teams usually value completed outcomes rather than model attempts. Use the 1,000-task projection to translate unit cost into an operating budget. If the cost per successful task is unexpectedly high, check the drivers in this order: success rate, retry frequency, number of calls, repeated input context, output length, and model price. Improving reliability can be more valuable than making one call slightly cheaper. Conversely, a reliable workflow may still be unnecessarily expensive if every simple step uses a premium model. For a mixed-model agent, calculate each distinct call type separately and add the run-level costs, or enter weighted-average token prices. A single average can hide an expensive planner, evaluator, or fallback call, so preserve a separate scenario for important model changes.

Planning scenarios and budgets

Create a baseline from production traces, then change one assumption at a time. A model-routing scenario should update the relevant token prices and any expected change in success rate. A prompt-compression scenario should reduce input tokens without assuming output usage also falls. A reliability scenario should change retries or success rate while leaving model prices fixed. The shareable scenario summary can document the assumptions behind an estimate for engineering, finance, or a client. Include the observation period, model names, pricing date, average token usage, and your definition of success. For budgeting, multiply the successful-task cost by forecast volume and add a separate allowance for traffic spikes and costs that the calculator does not cover.

Assumptions and limitations

The estimate assumes average token usage and a stable workload. Real agents often have long-tail runs with much larger contexts, extra tool loops, or fallback models. The retry calculation uses an average one-extra-run interpretation; recursive retries or multiple retry stages require a higher effective multiplier. Avoid double counting. If your measured success rate already incorporates every retry and your token averages cover the complete retry sequence, do not add the same retry cost again. Decide whether a “run” means one execution attempt or the entire retry chain, then keep that definition consistent. The calculator estimates model-token cost, not total cost of ownership. It does not automatically include vector databases, search APIs, tool calls, storage, network transfer, observability, human review, engineering time, taxes, provider minimums, cached-token pricing, batch discounts, or rate-specific contract terms. Treat the output as a planning estimate and reconcile it with actual usage data.

Move from an estimate to live cost control

A calculator is useful for validating unit economics before deployment, but production traffic changes with prompts, customers, models, and failure patterns. AgentCost provides an OpenAI-compatible chat completions endpoint and helps teams route AI requests, apply daily or monthly budget limits, and review requests, spend, limits, and estimated savings. Teams can use dedicated client or agent keys to separate usage and make costs accountable. After estimating a sensible cost per successful task here, translate the expected volume into a budget, monitor live per-agent spend, and revisit the calculation when token usage, routing rules, or success rates change.

Frequently asked questions

What is AI agent cost per task?

It is the expected model cost required to produce one task outcome. A useful calculation includes all LLM calls, input and output tokens, retry attempts, and the cost of unsuccessful runs—not just the price of the final response.

Why is cost per successful task higher than cost per run?

A run can fail, require a retry, or produce an unacceptable result while still consuming tokens. Cost per successful task spreads the expense of those attempts across the outcomes that actually succeed.

Should I use average or maximum token usage?

Use observed averages for expected unit cost and a high-percentile or conservative value for budget stress testing. Maximum token settings are limits, not necessarily actual consumption.

How should I calculate a workflow that uses several models?

Calculate the token cost of each call type with its own model price and add those costs to obtain the base run cost. If the calculator accepts one price pair, use a carefully calculated weighted average or create separate scenarios for each model stage.

Does this include infrastructure and tool costs?

No. Unless you manually incorporate them into a separate budget, the estimate covers the entered model-token prices. Add search, databases, third-party APIs, storage, compute, monitoring, and human review separately.

How often should I update the calculation?

Recalculate when provider prices, model routing, prompts, context size, retry rules, success criteria, or workload mix changes. Production averages should also be reviewed periodically because token usage can drift over time.

Control the spend behind every agent

Use your calculated cost per successful task to set operating expectations, then manage live requests with routing, per-agent or per-client visibility, and daily or monthly budget limits through AgentCost.

Explore AgentCost →