System design
How I architect AI systems
A practitioner's view of the patterns, layers, and decisions that go into building production-grade AI — from data ingestion to agent orchestration.
Core design principles
Every system I build is guided by a small set of principles that keep complexity manageable and outcomes predictable.
Composability over monoliths
Break systems into discrete, testable components that can be swapped, upgraded, or replaced independently. A retrieval layer should not know about the generation layer.
Observability from day one
Instrument everything — latency, token counts, retrieval scores, tool call outcomes. You can't improve what you can't measure, and AI systems fail in subtle ways.
Fail gracefully, not silently
LLMs hallucinate. Retrievers miss. Agents loop. Design every layer with explicit fallback paths and surface failures to the operator before they reach the user.
Human-in-the-loop by default
Autonomous agents are powerful but risky. I default to checkpoints, approval gates, and audit trails — especially for actions with real-world side effects.
The stack, layer by layer
Most of my AI systems share a common layered architecture. Each layer has a clear responsibility and a well-defined interface to the layers above and below it.
Data & ingestion
Raw data sources — documents, APIs, databases, event streams — normalised and chunked for downstream use. Quality here determines quality everywhere.
Retrieval & memory
Vector stores, keyword indexes, and hybrid search. Short-term working memory for agents, long-term episodic memory for personalisation.
Reasoning & generation
The LLM layer — prompt engineering, context assembly, structured output, and model routing. Often the smallest part of the system by line count.
Orchestration & agents
Task decomposition, tool use, multi-agent coordination, and loop control. This is where most of the interesting (and dangerous) behaviour lives.
Evaluation & feedback
Automated evals, human review pipelines, and continuous improvement loops. A system without evals is a system you can't trust.
Delivery & integration
APIs, webhooks, UI surfaces, and enterprise connectors. The best AI system is useless if it can't be reached by the people who need it.
Patterns I use regularly
RAG (Retrieval-Augmented Generation)
Ground LLM responses in verified, up-to-date documents. Reduces hallucination and enables domain-specific knowledge without fine-tuning.
ReAct agent loops
Interleave reasoning traces with tool calls so the agent can observe, plan, and act iteratively rather than in a single shot.
Multi-agent orchestration
Decompose complex tasks across specialised sub-agents — a planner, a researcher, a writer, a critic — coordinated by a supervisor.
Structured output pipelines
Use JSON schema constraints and function calling to get deterministic, parseable outputs from LLMs rather than free-form text.
MCP server integration
Expose tools and context to LLMs via the Model Context Protocol — a clean, standardised interface for giving models access to external systems.
Evaluation-driven iteration
Define success metrics before building, run automated evals on every change, and use human review to catch what automation misses.
See these patterns in practice
The Projects section shows real systems built on these foundations — with notes on what worked, what didn't, and what I'd do differently.