Inference is not your main cost
Teams budget obsessively for tokens and then get surprised by the real bill. On a typical production system, inference is 10–20% of first-year cost.
The rest is engineering, data work and operations — the same things that have always dominated software budgets.
Where it actually goes
Data preparation, 25–35%. Getting documents out of the systems they are trapped in, cleaning them, handling the PDF that is actually a scan, resolving contradictions between sources. Always larger than estimated.
Engineering, 30–40%. The retrieval pipeline, integrations, interface, permissions, error handling. Ordinary software engineering, which AI projects still require.
Evaluation, 10–15%. Building and maintaining the measurement infrastructure. Cutting this is tempting and consistently regretted.
Inference, 10–20%. The part everyone models in advance.
Operations, 5–10%. Monitoring, incident response, keeping content current.
Controlling inference anyway
It is the easiest line to reduce once you look:
- Right-size the model per task. Query rewriting and reranking do not need your most expensive model.
- Cache aggressively. Repeated questions are extremely common in support workloads.
- Budget context tightly. Retrieving twenty passages instead of five triples prompt cost and often reduces accuracy.
- Stream responses. It does not reduce cost but it changes perceived latency enough to let you use a cheaper model.
The honest planning number
For a production RAG system over a real corpus with evaluation and monitoring, expect three to six months and a team of three to five. Anyone promising a fortnight is describing a demo.
