RAG inverts the usual quality argument
For generic chat, the frontier model wins on quality and the argument for local is cost/privacy. RAG scrambles that: answer quality in retrieval-augmented systems is dominated by whether the right chunks reach the prompt — chunking strategy, embedding quality, reranking, query rewriting — and only secondarily by the generator model's tier. A 30B local model reading the right three paragraphs beats GPT-4o reading the wrong ones, every time. Which means the part of the stack worth obsessing over runs identically well on your own hardware, and the part where the frontier is better (generation) is the part RAG makes matter less. Our local RAG guide covers the tuning that actually moves quality.
The token math is brutal for metered RAG
RAG prompts are stuffed: system prompt + query + five or ten retrieved chunks, often 3,000–8,000 input tokens per question answered in 200. That 90%+ input weighting makes RAG the most expensive-per-interaction chat workload on metered APIs — and the cheapest locally, where context is just VRAM you already own. At ~2M tokens/day (a moderately busy internal knowledge assistant), GPT-4o runs ~$285/month; the same pipeline on a used RTX 3090 build costs ~$24/month of electricity, breaking even on the $1,240 hardware in about 5 months. Embeddings compound this: cloud embedding APIs bill re-indexing your corpus every time you improve chunking (and you will, repeatedly, while tuning), whereas open embedders like bge or nomic-embed re-index nightly for free.
The honest counterweight: mini-class models are cheap even at RAG volumes (~$17/month at our preset), and if your generator can be Mini-class and your documents aren't sensitive, cloud stays the budget option. RAG's local case is strongest when either quality needs exceed Mini-class or the corpus itself is the constraint — which brings us to the real issue.
The corpus is the crown jewels
Ask what's actually in the knowledge base: contracts, internal wikis, customer tickets, research notes, the accumulated institutional memory of your organization. A cloud RAG deployment doesn't send the occasional prompt to a third party — it uploads the entire corpus (to the embedding provider, often a vector-DB SaaS, and chunks of it to the LLM on every query). That's three vendors holding your documents to run one Q&A system. A fully local pipeline — open embedder, local vector store (Chroma/Qdrant on your disk), local generator — holds the corpus at exactly one location: yours. For anything under NDA, GDPR-sensitive, or simply strategic, this is the argument that ends meetings; the GDPR comparison makes the compliance version of it.
Where cloud legitimately wins RAG
Three real cases. Million-token contexts: Gemini-class models can sometimes skip RAG entirely — paste the whole corpus per query. Elegant for small corpora (and billed accordingly; it makes economic sense only at low query volume). Managed pipelines: the assistants-with-file-search products get a prototype running in an afternoon with zero retrieval engineering — the right first step to prove value before building anything. Very sparse usage: a knowledge base queried ten times a day doesn't justify hardware; a $5/month cloud bill does.
A reference local stack
One used RTX 3090 (24 GB) carries the whole thing: Qwen 3 32B Q4 as generator (~19 GB), nomic-embed on the side, Qdrant or Chroma for vectors, orchestrated by LlamaIndex or a few hundred lines of your own Python. That's the $1,240 build; a 16 GB card with a 14B generator is the budget version and loses less quality than you'd guess (retrieval dominates, again). Prototype in the cloud if you like — then bring the corpus home.