Queries per month
5,000
RAG cost planning
Plan retrieval-augmented generation costs by modeling query volume, context tokens, and answer length. No documents or private text are collected.
Start with these editable numeric assumptions, then adjust the calculator fields to match your own traffic and token usage.
Queries per month
5,000
Retrieved context tokens
2,500
Question and instructions
500 tokens
Output tokens per answer
700
Run the calculator to see projected cost and usage volume.
Enter your usage details, then select Calculate estimate to see your projected cost.
Estimated cost = input usage cost + output usage cost + supported optional charges.
RAG spend is often driven by the size of retrieved context sent into the model for every query.
Run low and high context scenarios to see how chunk count and context length affect monthly cost.
If your provider charges separately for embeddings, estimate indexing and query embedding costs alongside this generation estimate.
Anthropic
Text model to evaluate for retrieval answer generation.
Input $2.00 / 1M, output $10.00 / 1M.
Open modelAnthropic
Text model to evaluate for retrieval answer generation.
Input $3.00 / 1M, output $15.00 / 1M.
Open modelGemini
Text model to evaluate for retrieval answer generation.
Input $1.50 / 1M, output $9.00 / 1M.
Open modelDeepSeek
Text model to evaluate for retrieval answer generation.
Input $0.22 / 1M, output $0.66 / 1M.
Open modelRead ai api pricing guide before refining calculator assumptions.
Read input vs output tokens guide before refining calculator assumptions.
Estimate provider, model, token, and monthly AI API cost.
Forecast spend from known input and output token volumes.
The biggest drivers are query volume, retrieved context size, answer length, selected model price, and whether embeddings are billed separately.
No. This calculator only uses numeric token assumptions and safe model/provider slugs, never raw document text.
Start with average chunk size multiplied by the number of chunks sent to the model, then add question and instruction tokens.