System design

How I architect AI systems

A practitioner's view of the patterns, layers, and decisions that go into building production-grade AI — from data ingestion to agent orchestration.

Core design principles

Every system I build is guided by a small set of principles that keep complexity manageable and outcomes predictable.

Composability over monoliths

Break systems into discrete, testable components that can be swapped, upgraded, or replaced independently. A retrieval layer should not know about the generation layer.

Observability from day one

Instrument everything — latency, token counts, retrieval scores, tool call outcomes. You can't improve what you can't measure, and AI systems fail in subtle ways.

Fail gracefully, not silently

LLMs hallucinate. Retrievers miss. Agents loop. Design every layer with explicit fallback paths and surface failures to the operator before they reach the user.

Human-in-the-loop by default

Autonomous agents are powerful but risky. I default to checkpoints, approval gates, and audit trails — especially for actions with real-world side effects.

The stack, layer by layer

Most of my AI systems share a common layered architecture. Each layer has a clear responsibility and a well-defined interface to the layers above and below it.

01

Data & ingestion

Raw data sources — documents, APIs, databases, event streams — normalised and chunked for downstream use. Quality here determines quality everywhere.

02

Retrieval & memory

Vector stores, keyword indexes, and hybrid search. Short-term working memory for agents, long-term episodic memory for personalisation.

03

Reasoning & generation

The LLM layer — prompt engineering, context assembly, structured output, and model routing. Often the smallest part of the system by line count.

04

Orchestration & agents

Task decomposition, tool use, multi-agent coordination, and loop control. This is where most of the interesting (and dangerous) behaviour lives.

05

Evaluation & feedback

Automated evals, human review pipelines, and continuous improvement loops. A system without evals is a system you can't trust.

06

Delivery & integration

APIs, webhooks, UI surfaces, and enterprise connectors. The best AI system is useless if it can't be reached by the people who need it.

Patterns I use regularly

RAG (Retrieval-Augmented Generation)

Ground LLM responses in verified, up-to-date documents. Reduces hallucination and enables domain-specific knowledge without fine-tuning.

retrievalgroundingknowledge

ReAct agent loops

Interleave reasoning traces with tool calls so the agent can observe, plan, and act iteratively rather than in a single shot.

agentstool usereasoning

Multi-agent orchestration

Decompose complex tasks across specialised sub-agents — a planner, a researcher, a writer, a critic — coordinated by a supervisor.

agentscoordinationswarms

Structured output pipelines

Use JSON schema constraints and function calling to get deterministic, parseable outputs from LLMs rather than free-form text.

reliabilityparsingschemas

MCP server integration

Expose tools and context to LLMs via the Model Context Protocol — a clean, standardised interface for giving models access to external systems.

MCPtoolscontext

Evaluation-driven iteration

Define success metrics before building, run automated evals on every change, and use human review to catch what automation misses.

evalsqualityiteration

See these patterns in practice

The Projects section shows real systems built on these foundations — with notes on what worked, what didn't, and what I'd do differently.

JoeCairns.AI

Building AI agents, automation workflows, and MCP servers — and documenting every lesson learned along the way.

Connect

© 2026 AI with Joe. All rights reserved.

Building AI that actually works

Admin