RAG Cost Calculator
Estimate the complete operating cost of a retrieval-augmented generation workload before launch or scaling. Model document ingestion and re-indexing separately from query-time retrieval, reranking and LLM generation, then inspect the monthly cost waterfall, per-query economics and sensitivity to traffic or token changes.
Interactive LLM cost calculator
Estimate monthly model cost from request volume, token usage, and your provider's per-million-token prices. Enter the prices from the provider you actually use.
Calculate the full monthly cost of a RAG pipeline
A RAG system creates costs at two different stages. Index-time costs come from processing documents, updating changed content and generating embeddings. Query-time costs come from vector retrieval, optional reranking and the language model tokens used to produce an answer. Enter your workload assumptions and provider rates in consistent currency units. The calculator separates each component so a low embedding bill does not hide a much larger generation bill—or a frequent re-indexing schedule does not get mistaken for a one-time setup cost.
Inputs to collect before building your estimate
For ingestion, record the number of documents, average tokens per document, new documents added each month and the share of existing content re-indexed monthly. Use measured token counts where possible; file size and page count are unreliable proxies after parsing and chunking. For query traffic, enter monthly requests and any retrieval or reranking activity required per request. For generation, use average LLM input and output tokens per request. Input tokens should include the user prompt, system instructions, conversation history and retrieved context sent to the model. Finally, enter current embedding, retrieval, reranking and LLM rates using the billing units shown by each provider.
Formulas used by the RAG cost calculator
Monthly embedding volume is estimated as: new-content tokens + re-indexed-content tokens. New-content tokens equal new documents × average tokens per document. Re-indexed-content tokens equal existing documents × monthly re-indexing rate × average tokens per document. Embedding cost equals embedded tokens divided by the provider’s token billing unit, multiplied by its embedding rate. LLM input cost equals monthly requests × average input tokens per request ÷ token billing unit × input-token rate. Output cost uses the same formula with output tokens and the output-token rate. Retrieval and reranking costs are calculated from monthly query activity using the provider’s applicable billing unit. Total monthly RAG cost is the sum of ingestion or re-indexing charges, embeddings, retrieval, reranking, LLM input and LLM output.
Read the monthly cost waterfall correctly
The waterfall shows how each stage contributes to projected monthly spend. Start with the largest bar rather than optimizing every component equally. If LLM input dominates, inspect chunk size, the number of retrieved passages, prompt length and conversation history. If output dominates, review answer-length limits and whether every workflow needs a long generated response. If reranking dominates, test fewer candidates or reserve reranking for ambiguous searches. If re-indexing dominates, check whether the entire corpus is being rebuilt when only changed documents need new embeddings. These are diagnostic directions, not guaranteed savings: relevance and answer quality must be evaluated alongside cost.
Use per-query cost for pricing and capacity decisions
Per-query cost is calculated by dividing projected recurring monthly cost by monthly request volume. Decide whether to include monthly ingestion and re-indexing in that figure. Including them produces a fully allocated unit cost suitable for product margins or client pricing. Excluding them isolates the marginal cost of serving one more query. For example, an AI agency quoting a fixed monthly client plan should generally allocate recurring indexing expense across expected requests, while an engineering team comparing two retrieval configurations may prefer query-time cost alone. Low-volume workloads can have a high fully allocated cost per query even when model calls are inexpensive.
Run sensitivity analysis before choosing a budget
A single forecast is rarely sufficient because traffic, context length and output length change after launch. Compare at least three cases: expected volume, a higher-traffic case and a high-token case. Doubling request volume approximately doubles retrieval, reranking and generation costs when per-request behavior remains constant. Increasing retrieved context raises LLM input spend even if request volume does not change. A larger corpus does not necessarily increase generation cost, but it can increase embedding and re-indexing expense. Use the sensitivity view to identify which assumption produces the largest movement, then validate that assumption with a pilot or production sample.
Assumptions, exclusions and estimate limitations
The result is a planning estimate, not a provider invoice. It assumes the entered averages represent the full month and that rates remain unchanged. Actual spend can differ because of retries, failed requests, agent loops, cache behavior, minimum charges, batch discounts, storage, data transfer, parsing, OCR, vector database capacity, observability and other infrastructure. Include those items separately if they apply. Avoid mixing decimal and binary units or rates quoted per 1,000 operations with rates quoted per one million tokens. Keep every component in the same currency. Download the calculator assumptions with the estimate so reviewers can see the traffic, token, re-indexing and pricing inputs behind the total.
Turn the projection into an operating control
A forecast is most useful when it becomes a limit rather than remaining in a spreadsheet. Use the projected monthly total as a starting budget, with an explicit allowance for forecast error and traffic growth. AgentCost can give a client or agent a dedicated API key, apply daily and monthly budget limits, and track requests, spend and estimated savings. Its OpenAI-compatible chat completions endpoint lets teams send supported model traffic through one controlled path. Compare actual agent costs with the calculator estimate, then revise request volume, token averages or routing decisions using observed usage.
Frequently asked questions
What costs should a RAG estimate include?
Include document ingestion and re-indexing, embedding generation, vector retrieval, reranking when used, LLM input tokens and LLM output tokens. Also account separately for storage, parsing, OCR, data transfer and operational infrastructure if your providers charge for them.
Is embedding a one-time RAG cost?
Not always. The initial corpus creates an upfront embedding workload, while new or changed documents create recurring embedding costs. Full re-indexes, chunking changes or a new embedding model can cause additional embedding volume.
How do I estimate RAG input tokens?
Measure the complete payload sent to the language model: system instructions, the user request, conversation history, retrieved passages and formatting. Sample real requests and use an average or percentile appropriate to the budget risk you want to manage.
Why can RAG cost increase when request volume stays flat?
Longer retrieved context, more reranking candidates, larger generated answers, frequent re-indexing, retries or agent loops can raise spend without increasing user-visible request count. Provider rate changes can also affect the total.
Should ingestion be included in cost per query?
Include recurring ingestion and re-indexing when calculating a fully allocated unit cost for pricing or margins. Exclude them when you need only the marginal query-serving cost. Reporting both figures usually gives the clearest view.
How should I set a monthly RAG budget?
Start with the calculator’s expected monthly total, review the sensitivity scenarios and add a deliberate allowance for uncertainty. Set daily and monthly controls, then compare actual request and spend data with the assumptions rather than treating the first estimate as fixed.
Put your RAG estimate behind an enforceable budget
Transfer the projected monthly spend into an AgentCost budget, separate usage by client or agent, and monitor actual request costs against the assumptions used in your forecast.