home / blog / building-a-rag-system

A RAG Stack That Actually Ships in 2026

The RAG stack I actually ship for clients in 2026 (vector store, embeddings, orchestration) plus the chunk-size, cost, and citation gotchas that bit me along the way.

A GPU used for Retrieval Augmentation Generation

Most RAG “how-to” posts are just the textbook diagram: ingest → chunk → embed → vector store → retrieve → prompt → LLM. That diagram is correct and useless. It tells you nothing about which vector store, which embedding model, or which decisions actually move the needle on quality and cost.

This post is the stack I’ve actually shipped for a couple of small-business RAG projects in 2025–2026, the decisions I’d make differently next time, and rough numbers on what it costs to run.

The default stack

LayerWhat I useWhy
Document storageDigitalOcean Spaces (S3-compatible)Cheap, no egress surprises, easy to swap to R2/S3 later
ChunkingCustom Python script, paragraph-based with ~400 token windows + 50 token overlapBetter than fixed-size splits; less worse than full semantic chunking
EmbeddingsDigitalOcean’s hosted GTE Large EN v1.5 endpointStays inside the DO stack (no extra vendor, no data leaving the platform) with solid retrieval quality at 1024 dimensions
Vector databasepgvector on a DigitalOcean Managed Postgres instanceI’m already running Postgres for app data; one fewer vendor, and it scales fine into the hundreds of thousands of chunks
RetrievalTop-K=8, with an optional re-ranking passTop-K alone gives generic answers; a re-ranker is the cheapest quality lever when you need it
LLMA mix: Llama 4 Maverick and Llama 3.3 70B for most answers, Claude Sonnet 4 / Opus for the hard onesOpen models on DO handle the bulk cheaply; escalate to Claude only when the workload demands it
OrchestrationPlain Python; no frameworkLangChain and friends add more complexity than they remove at this scale

The decisions that actually matter

Chunk size and overlap

The most common mistake is fixed-size chunks of 1,000 tokens with zero overlap, because that’s the LangChain default. You get fast indexing and bad answers. My current default: ~400 token chunks with ~50 token overlap, split on paragraph boundaries when available. For Q&A over documentation, this is consistently the cheapest quality win.

Re-ranking is the cheapest quality lever

Top-K vector retrieval alone produces semantically-similar-but-not-actually-relevant chunks all the time. A re-ranker takes your top 20 and gives you back the actual top 8 for a rounding-error cost per query. When retrieval quality is the bottleneck, it’s the cheapest lever to pull, a near-10x improvement in answer relevance. It’s optional, not load-bearing: on smaller, well-chunked corpora, plain Top-K is often good enough to ship without it.

Two-tier prompting

The user’s query usually needs rewriting before retrieval (resolve pronouns, expand abbreviations, decompose multi-part questions). Have a small, fast model (Llama 3.3 70B, or Claude Haiku) do the rewrite, then send the rewritten query to retrieval, then send the retrieved context + original question to your expensive model for the final answer. Easily 30–50% cost savings on the LLM line item.

Citations are non-negotiable

Every answer should cite which chunk(s) produced it. If you don’t do this, you can’t debug hallucinations and your users can’t verify the answer. Easy way: ask the LLM to wrap citations in [chunk_id] markers in its response, then resolve those server-side.

What I’d do differently next time

  • Hybrid search. Pure vector search misses queries that are keyword-heavy (names, codes, product IDs). BM25 + vector combined is consistently better. Most vector DBs now have built-in support; use it.
  • Right-size the embeddings. GTE Large EN v1.5 at 1024 dimensions is plenty for tens of thousands of chunks. If storage or query latency starts to bite, a smaller embedding model (or truncating the vectors) trims cost with barely any recall loss at this scale.
  • Don’t use a framework on day one. LangChain and LlamaIndex are useful if you have the size and team that justify their abstractions. Solo or small team, plain Python wins for a long time before that’s true.

Rough cost model (10K chunks, 1K queries/day)

ComponentCost
Document storage (Spaces)$5/mo
Embeddings (GTE Large on DO: one-time + monthly updates)~$1 one-time + $0.50/mo
Vector DB (pgvector on existing DO Postgres)folded into your ~$15/mo Postgres bill
Re-ranking (optional)~$30/mo
LLM (mostly Llama 4 / 3.3 on DO, Claude for hard queries)~$120/mo
Query rewriting (small model)~$5/mo
Total~$175/mo (~$145 without re-ranking)

That gets you a production RAG system answering ~30K questions a month for well under $0.01 per answer. Leaning on open models hosted on DigitalOcean for the bulk of the work, and escalating to Claude only for the hard queries, is what keeps the LLM line so low. Scale up or down by changing how often you reach for the heavier model and whether you run the re-rank pass.

If you want help building one

I’ve shipped RAG for a couple of clients now. Documentation search, internal knowledge bases, customer support pre-answer drafting. Drop me a line if you have a corpus and a use case.

← Back to all posts Reply to this post →