The Cheapest AI Model for RAG (Document Q&A)
RAG bills are driven by input and cached-input prices. Every model ranked for 50,000 document questions a month, plus the long-context threshold to watch.
Updated September 19, 2026
Retrieval-augmented generation (RAG) has a pricing shape all its own: long inputs, short outputs. Each question ships thousands of tokens of retrieved documents and gets back a paragraph or two. That makes the input price, and the cached input price, the numbers that decide your bill — not the output rate the marketing pages lead with.
Cost for 50,000 questions a month
50,000 questions a month with ~8,000 tokens of retrieved context and ~600 output tokens each.
| Model | Input | Cached input | Output | Per month |
|---|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $116.00 |
| DeepSeek V4.1 Flash | $0.30 | $0.006 | $1.20 | $156.00 |
| Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | $195.00 |
| Gemini 2.5 Flash | $0.30 | $0.03 | $2.50 | $195.00 |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | $412.50 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | $550.00 |
| Grok 4.3 | $1.25 | $0.20 | $2.50 | $575.00 |
| DeepSeek V4 Pro | $1.32 | $0.044 | $3.96 | $646.80 |
| Gemini 2.5 Pro | $1.25 | $0.125 | $10.00 | $800.00 |
| Gemini 3.5 Flash | $1.50 | $0.15 | $9.00 | $870.00 |
GPT-5.6 Luna is the cheapest model we track for this workload at $116.00 a month.
Caching changes the ranking
If your users ask about the same documents repeatedly, a cache hit turns the most expensive part of the request into the cheapest. With 60% of the retrieved context served from cache, the same 50,000 questions cost $72.80 on GPT-5.6 Luna instead of $116.00. Models with a deep cache discount move up the table; models without a published cached rate do not move at all.
Watch the long-context threshold
Several providers charge a higher rate once a prompt crosses a size threshold — often 128K or 200K tokens. Retrieval pipelines drift over that line as they add more chunks, and the bill can jump without any code change. Each model page lists the threshold and the higher rate.
Cheaper retrieval beats a cheaper model
- Retrieve less. Going from 8,000 to 4,000 tokens of context halves the input bill outright. Rerank your chunks and send only the top ones.
- Deduplicate chunks. Overlapping passages are paid for twice.
- Batch the offline work. Nightly summarization or indexing runs qualify for batch discounts, usually 50% off.
- Keep the answer short. Output is billed several times higher than input on every provider.
Run your own numbers in the calculator, or paste a real prompt into the token counter to see how many tokens your context actually uses.