Make LLM cost allocation accountable
Turn a combined model-provider bill into a defensible monthly showback. Define agents, teams, customers and projects; separate directly attributable spend from shared costs; choose an allocation key; and preview what each business unit should own. AgentCost can support this workflow with dedicated client or agent keys, usage and cost analytics, estimated savings, and daily or monthly budget controls.
Build a monthly LLM cost allocation worksheet
Use the interactive worksheet to model the hierarchy behind your AI workloads. Add each agent, assign it to a team, customer and project, then enter direct spend and any shared cost pool that still needs to be distributed. The essential inputs are: • Reporting month and one consistent currency. • Agent or workload name. • Accountable team, customer and project. • Direct LLM cost already traceable to that workload. • Shared pool name and amount, such as orchestration or untagged provider spend. • Allocation key, such as requests, tokens, active seats, revenue or equal share. • Allocation units for every destination included in the pool. The preview produces a monthly showback: direct cost, allocated shared cost and total attributed cost for each row. Roll the rows up by team, customer or project to answer different management and billing questions without changing the underlying workload data.
Use direct attribution before allocating shared spend
Allocation is a fallback for costs that cannot be traced directly. First assign spend generated by a known agent or customer to that owner. Only place genuinely shared or untagged amounts into an allocation pool. A practical order of operations is: 1. Attribute requests to a dedicated agent or client identity where possible. 2. Map that identity to its team, customer and project. 3. Reconcile attributed costs with the reporting-period total. 4. Put the remaining amount into a named shared pool. 5. Allocate that pool using a documented business driver. This avoids making a high-volume customer subsidize another customer's premium-model workload merely because both appear on the same provider invoice. AgentCost supports direct attribution by allowing a dedicated API key for each client or agent and recording requests, spend and estimated savings.
How the allocation formula works
For each shared pool, calculate a destination's allocation weight as: Allocation weight = destination allocation units ÷ total allocation units in the pool Allocated shared cost = shared pool amount × allocation weight Total attributed cost = direct cost + allocated shared cost If a €1,200 pool is divided by request volume and one project generated 50,000 of 100,000 eligible requests, its weight is 0.50 and its allocation is €600. The worksheet should preserve full precision during calculation and round only displayed currency values. Any rounding remainder should be assigned consistently and disclosed so the final table still reconciles to the source total. Reconciliation check: Sum of total attributed costs = sum of direct costs + sum of shared pools Do not combine currencies in one pool unless you first convert every amount using the same documented exchange rate and date.
Choose an allocation key that reflects cost causation
Request count is useful when requests are relatively similar in size and model choice. Token volume is usually more representative when prompts and outputs vary substantially. Active seats can support internal showback when usage telemetry is incomplete, but it measures access rather than actual consumption. Revenue can be appropriate for distributing a general platform overhead, although it should not be presented as usage-based attribution. Equal share is simple and auditable, but it is best reserved for genuinely common costs or an early-stage estimate. Decision logic: • Use direct measured cost when agent-level records are available. • Use tokens when variable context and output length drive the bill. • Use requests when request profiles and model mix are comparable. • Use seats for access-oriented internal services with limited telemetry. • Use revenue only when the policy is intentionally commercial rather than consumption-based. • Use equal share when no better causal measure exists, and label the result as an estimate. Keep separate pools when different costs require different drivers. For example, allocate shared inference by tokens but divide a common test environment equally across the participating projects.
Worked SaaS allocation example
Consider a SaaS company with €4,800 of directly traced LLM spend and a €1,200 shared orchestration pool for the month. Customer A's support agent has €1,800 of direct cost and 50,000 eligible requests. Customer B's copilot has €2,400 and 30,000 requests. An internal research agent has €600 and 20,000 requests. The total allocation base is 100,000 requests. Customer A receives 50% of the shared pool, or €600, for a total attributed cost of €2,400. Customer B receives 30%, or €360, for a total of €2,760. Internal research receives 20%, or €240, for a total of €840. The resulting showback reconciles to €6,000: €4,800 direct plus €1,200 shared. The same rows can be grouped by customer for margin review, by team for budget ownership or by project for product investment decisions. These figures are illustrative; an actual policy should use your recorded costs and an allocation key appropriate to the workload.
CSV template for repeatable monthly showback
Download the template from the worksheet or structure your import with these columns: month,currency,agent,team,customer,project,direct_cost,pool_name,pool_amount,allocation_key,allocation_units 2026-01,EUR,support-agent,customer-success,Customer A,Helpdesk,1800,shared-orchestration,1200,requests,50000 2026-01,EUR,copilot-agent,product,Customer B,Copilot,2400,shared-orchestration,1200,requests,30000 2026-01,EUR,research-agent,operations,Internal,Research,600,shared-orchestration,1200,requests,20000 Use stable names from month to month so historical comparisons do not fragment. If the same pool appears on several destination rows, repeat an identical pool amount and allocation key; treat the pool as one grouped amount rather than adding the repeated values together. Validate that every pool has a positive allocation base and that direct costs plus distinct shared pools reconcile with the reporting total.
Connect allocation with routing, savings and spend controls
A showback explains where spend went; controls influence what happens next. Route each client or agent through a dedicated AgentCost key, apply daily or monthly budget limits, and review requests, costs and estimated savings in the portal. Configured limits can reject requests once the relevant ceiling is reached. Keep actual cost and estimated savings as separate measures. Actual attributed cost belongs in financial reconciliation. Estimated savings can be rolled up by agent, customer or team to evaluate routing decisions, but it is not an invoice credit and depends on available usage, pricing data, model choice and routing configuration. For example, if a customer-facing agent repeatedly approaches its monthly ceiling, the accountable team can inspect its request pattern and model selection before changing the budget. A lower-cost routing option may reduce future spend, while a justified workload increase may require a higher limit. Allocation provides ownership; routing and limits provide operational action.
Interpret the output—and know its limits
Use showback when the goal is visibility and accountability without transferring money between departments. Use chargeback only after finance and workload owners agree on the allocation policy, data source, treatment of shared pools, rounding rules and dispute process. An allocation table is not automatically a customer invoice. Taxes, contracted markups, prepaid commitments, provider credits, currency conversion and non-LLM service costs may require separate treatment. Results are only as reliable as the mapping and allocation units supplied. Changes in model pricing, token accounting or workload design can also make one month's driver unsuitable for the next. Review the policy periodically, retain the source totals, and version any mapping changes. A defensible result should be reproducible, reconciled and understandable to the team or customer receiving it.
Frequently asked questions
What is LLM cost allocation?
LLM cost allocation is the process of assigning model and AI-agent spend to accountable agents, teams, customers or projects. Directly traceable costs are assigned first; remaining shared costs are distributed using a documented allocation key.
Should LLM spend be allocated by requests or tokens?
Use requests when calls have similar size and model mix. Use tokens when prompt length, output length or context size varies materially. If model prices differ substantially, direct measured cost is preferable to either proxy.
How do I allocate one provider invoice across customers?
Map directly identified usage to each customer, reconcile it with the invoice-period total, and place only the unresolved amount into a shared pool. Allocate that pool using a consistent driver, then verify that all customer totals add back to the source amount.
What is the difference between showback and chargeback?
Showback reports attributed cost to create visibility but does not necessarily bill the recipient. Chargeback uses the allocation to transfer or invoice costs. Chargeback requires stronger governance around data, contracts, adjustments and disputes.
How does AgentCost support an allocation workflow?
AgentCost can give each client or agent a dedicated API key, record requests and spend, report estimated savings, and enforce daily or monthly budgets. Teams can map those agent or client identities to their own team, project and customer structure for reporting.
Does AgentCost guarantee savings?
No. Savings vary with workload, provider pricing, model choice and routing configuration. Estimated savings should be reviewed separately from actual cost and should not be treated as a guaranteed reduction or invoice credit.
Evaluate AgentCost with your workload structure
Map your current agents to teams, customers and projects, identify where direct attribution can replace estimates, and decide which workloads need daily or monthly limits. Then use AgentCost's per-agent cost and savings analytics to make the monthly allocation operational.