The Lab — RAG Systems

Grounding AI in What's Actually True.

A language model without retrieval is a brilliant amnesiac — it knows a lot, but only up to a point in time, and it can't access your data, your documents, or your domain knowledge. Retrieval-Augmented Generation changes that. RAG gives AI systems a memory that can be updated, queried, and trusted. It's the difference between an AI that hallucinates confidently and one that answers from evidence. I think it's one of the most practically important techniques in applied AI, and I build RAG pipelines into almost everything I deploy.

What Is RAG?

The core idea

"RAG is the difference between asking someone what they remember and handing them the relevant documents before they answer."

Retrieval-Augmented Generation (RAG) is a technique that combines a language model's reasoning capabilities with a retrieval system's ability to find relevant information at query time. Instead of relying solely on what the model learned during training, RAG retrieves relevant documents or chunks from an external knowledge base and injects them into the model's context before generating a response.

The pipeline has three core stages: at index time, documents are chunked, embedded into vectors, and stored in a vector database. At query time, the user's question is embedded, the most semantically similar chunks are retrieved, and those chunks are passed to the language model as context. The model then generates a response grounded in the retrieved evidence.

The result is an AI system that can answer questions about your specific data, stay current without retraining, cite its sources, and fail gracefully when it doesn't know something — rather than inventing a plausible-sounding answer.

Why RAG Matters for Agent Systems

For agents that need to act in the real world, hallucination isn't just an annoyance — it's a failure mode. RAG is how you give agents grounded, verifiable knowledge.

01

Eliminates Hallucination on Domain Knowledge

Language models hallucinate when asked about things outside their training data. RAG replaces guessing with retrieval — the model answers from documents you control, not from statistical patterns in its weights.

02

Keeps Knowledge Current

Training cutoffs make models stale. RAG knowledge bases can be updated continuously — new documents, new data, new policies — without retraining or fine-tuning the model. The knowledge layer and the reasoning layer evolve independently.

03

Enables Source Attribution

When an agent answers from retrieved chunks, it can cite exactly which document, page, or section it drew from. This makes AI outputs auditable and verifiable — critical for any production use case where trust matters.

04

Scales to Arbitrary Knowledge

A model's context window is finite. A vector database is not. RAG lets agents reason over millions of documents by retrieving only the relevant subset at query time — the model sees what it needs, not everything.

05

Composable with MCP

RAG pipelines exposed through MCP servers become reusable tools any agent can call. An agent doesn't need to know how retrieval works — it just calls the knowledge tool and gets grounded context back.

06

Graceful Degradation

A well-built RAG system knows when it doesn't know. When retrieval returns low-confidence results, the system can say so — rather than generating a confident but wrong answer. Calibrated uncertainty is a feature, not a bug.

Systems I'm Building

Active RAG projects in the lab — each one solving a real knowledge retrieval problem.

Active

Lab Knowledge Base

A RAG system over my own lab documentation — architecture decisions, runbooks, experiment notes, and research papers. Exposed as an MCP server so any agent in the lab can query it. The system ingests markdown, PDFs, and structured notes, and returns answers with source citations.

Use case

Internal agent knowledge layer

Stack

pgvector · PostgreSQL · OpenAI text-embedding-3-large · TypeScript

Active

Technical Documentation RAG

A retrieval system over technical documentation for the tools and frameworks I use — API references, SDK docs, changelog entries. Agents can query it when they need accurate, up-to-date technical details rather than relying on potentially stale training knowledge.

Use case

Agent tool-use grounding

Stack

pgvector · PostgreSQL · OpenAI Embeddings · Cheerio scraper

Experimental

Multi-Collection Federated RAG

An experimental system that routes queries across multiple specialized knowledge collections — routing the query to the most relevant collection before retrieval, then merging and reranking results. Designed for cases where a single flat index isn't the right structure.

Use case

Cross-domain agent knowledge

Stack

pgvector · LangChain · Cohere Rerank · TypeScript

How I Structure RAG Pipelines

Every RAG system I build follows the same two-phase architecture — index time and query time — with deliberate decisions at each stage.

Index Time
01

Document Ingestion

Source documents are loaded from their origin — file system, S3, database, web scraper, or API. Each document is normalized to a consistent internal format with metadata preserved.

02

Chunking

Documents are split into chunks sized for the embedding model and retrieval context. Chunk boundaries respect semantic units — paragraphs, sections, code blocks — not arbitrary character counts.

03

Embedding

Each chunk is embedded into a high-dimensional vector using an embedding model. The vector captures the semantic meaning of the chunk, enabling similarity search rather than keyword matching.

04

Storage

Vectors and their source chunks are stored in pgvector. Metadata — source document, page, section, timestamp — is stored alongside for filtering and attribution.

Query Time
05

Query Embedding

The user's query is embedded using the same model used at index time. This ensures the query vector lives in the same semantic space as the document vectors.

06

Retrieval

The top-k most similar chunks are retrieved via approximate nearest-neighbor search. Metadata filters can narrow the search to specific collections, date ranges, or document types.

07

Reranking

Retrieved chunks are reranked using a cross-encoder model that scores each chunk against the full query. This catches cases where vector similarity missed the most relevant result.

08

Generation

The top reranked chunks are injected into the language model's context as grounding evidence. The model generates a response that cites its sources and stays within the retrieved evidence.

Chunking & Embedding Strategy

Chunking is where most RAG systems fail. The right strategy depends on the document type, the query patterns, and the embedding model.

Semantic Chunking

Split on semantic boundaries — paragraphs, sections, headings — rather than fixed character counts. A 512-token chunk that cuts mid-sentence retrieves worse than a 600-token chunk that ends at a natural boundary.

Best for: Prose documents, documentation, articles

Hierarchical Chunking

Store chunks at multiple granularities — document summary, section summary, paragraph. Retrieve at the paragraph level but include section context in the generation prompt. Gives the model both precision and context.

Best for: Long documents, technical manuals, research papers

Sliding Window with Overlap

Chunks overlap by 10–20% to avoid splitting relevant content across chunk boundaries. The overlap ensures that a sentence at the end of one chunk also appears at the start of the next.

Best for: Dense technical content, code documentation

Structured Extraction

For structured data — tables, JSON, code — extract and embed the structured representation directly rather than treating it as prose. A table row embeds better as 'field: value' pairs than as raw HTML.

Best for: Tables, code, structured data, API references

I use OpenAI text-embedding-3-large as the default embedding model — 3072 dimensions, strong multilingual performance, and good separation between semantically distinct chunks. For latency-sensitive applications I drop to text-embedding-3-small with minimal quality loss.

Retrieval & Reranking

Vector similarity gets you close. Reranking gets you accurate. I treat them as two distinct stages with different jobs.

Stage 1: Vector Retrieval

Approximate nearest-neighbor search over pgvector using cosine similarity. Fast, scalable, and good enough to get the right documents into the candidate set. I retrieve top-20 to top-50 candidates — more than I'll use, but enough to give the reranker good material to work with.

pgvector HNSW indexCosine similarityMetadata pre-filtering

Stage 2: Cross-Encoder Reranking

A cross-encoder model scores each candidate chunk against the full query — not just the query vector, but the actual query text. Cross-encoders are slower than vector search but dramatically more accurate at identifying the most relevant chunks. I use Cohere Rerank for production and a local cross-encoder for offline use.

Cohere Rerank APILocal cross-encoder (sentence-transformers)Top-5 selection

Stage 3: Context Assembly

The top reranked chunks are assembled into a context block with source citations. I include the chunk text, the source document name, and the section heading. The generation prompt instructs the model to answer only from the provided context and to cite sources explicitly.

Source attributionContext window budgetingCitation formatting

Evaluation & Quality Metrics

A RAG system you can't measure is a RAG system you can't improve. I track these metrics across every system I build.

01

Retrieval Recall@k

What fraction of the time does the correct document appear in the top-k retrieved results? This measures the retrieval stage in isolation, before reranking. A low recall@k means the problem is in the index or the embedding model, not the generator.

02

MRR (Mean Reciprocal Rank)

How high does the most relevant chunk rank in the retrieved results? MRR penalizes systems that find the right document but bury it at position 8. A good reranker should push the most relevant chunk to position 1 or 2.

03

Answer Faithfulness

Does the generated answer stay within the retrieved context, or does the model add information from its training weights? I use an LLM-as-judge approach to score faithfulness — asking a separate model to verify each claim in the answer against the retrieved chunks.

04

Answer Relevance

Does the generated answer actually address the question asked? A faithful answer that doesn't answer the question is still a failure. Relevance is scored separately from faithfulness.

05

Context Precision

Of the chunks retrieved, what fraction were actually useful for answering the question? High context precision means the retrieval is tight. Low precision means the model is wading through noise to find the signal.

06

Latency P95

End-to-end latency at the 95th percentile — from query received to response returned. RAG adds latency at every stage. I track P95 rather than average because tail latency is what users actually experience.

Challenges & Lessons Learned

Building RAG systems that work in production is harder than building ones that work in demos. These are the lessons that cost me the most time.

01

Chunking Strategy Is the Highest-Leverage Decision

I spent weeks tuning embedding models and retrieval parameters before realizing the real problem was chunking. Bad chunk boundaries — splitting mid-sentence, mid-table, mid-code-block — degraded retrieval quality more than any other factor. Fix chunking first.

02

Vector Search Alone Is Not Enough

Pure vector similarity retrieval has a ceiling. It finds semantically similar chunks, but 'semantically similar' isn't always the same as 'most relevant to this specific question.' Adding a reranking stage consistently improved answer quality by 15–30% on my evaluation sets.

03

Metadata Filtering Is Underrated

Filtering by document type, date range, or collection before vector search dramatically improves precision — especially in multi-collection systems. Don't make the retrieval system search everything when you know the answer is in a specific subset.

04

Evaluation Is Not Optional

Without a test set and metrics, you're flying blind. Every change I made to chunking, embedding, or retrieval that I thought would help sometimes hurt. You can't know without measuring. Building an evaluation harness early is the best investment in a RAG project.

05

The Generation Prompt Matters as Much as Retrieval

A good retrieval stage can be undone by a bad generation prompt. The prompt needs to explicitly instruct the model to stay within the retrieved context, cite sources, and express uncertainty when the context doesn't contain the answer. Without this, models drift back to their training weights.

Tech Stack

The tools and frameworks I use to build and run RAG systems.

Storage & Indexing

pgvector — Vector similarity search on PostgreSQL
PostgreSQL — Primary data store for chunks and metadata
HNSW index — Approximate nearest-neighbor for fast retrieval

Embedding & Reranking

OpenAI text-embedding-3-large — Primary embedding model
OpenAI text-embedding-3-small — Latency-sensitive applications
Cohere Rerank — Cross-encoder reranking for precision
sentence-transformers — Local cross-encoder for offline use

Pipeline & Orchestration

TypeScript — Primary pipeline implementation language
LangChain — Experimental multi-collection routing
Drizzle ORM — Database access layer
MCP SDK — Exposing RAG as agent-callable tools

Explore More of the Lab

RAG systems are the knowledge layer that grounds agent reasoning. See how they connect to the rest of the stack.

JoeCairns.AI

Building AI agents, automation workflows, and MCP servers — and documenting every lesson learned along the way.

Connect

© 2026 AI with Joe. All rights reserved.

Building AI that actually works

Admin