Skip to slide
Chapter 18 · Reference Architecture and a Design Checklist
177 / 191

CHAPTER 18 · Reference Architecture and a Design Checklist · 1 / 15

The reference architecture

A complete, production-grade agent, the kind the earlier chapters describe, has this shape:

                          ┌─────────────────────────────────────────┐
                          │                 Client                  │
                          │   (renders streamed event timeline)     │
                          └───────────────┬─────────────────────────┘
                                          │  request + SSE stream
                                          ▼
┌──────────────────────────────────────────────────────────────────────────────┐
│                               API / Web tier                                   │
│  • Edge hardening: security headers, CORS, rate limiting                       │
│  • Auth middleware (identity)            ─────────────► [Auth system]           │
│  • Authorization helpers (access checks)                                       │
│  • Input validation                                                            │
│  • Routes per resource (chat, documents, projects, …)                          │
└───────────────┬────────────────────────────────────────────────┬──────────────┘
                │                                                  │
                ▼                                                  ▼
┌───────────────────────────────────────┐         ┌──────────────────────────────┐
│            Agent core                  │         │      Document pipeline        │
│  • Context builder (assemble prompt)   │         │  ingest → store → convert →   │
│  • Orchestrator (the loop, streaming,  │         │  extract → version → ready    │
│    event timeline)                     │         └───────────────┬──────────────┘
│  • Tool executor (dispatch + results)  │                         │
│  • System prompt + emitted protocols   │                         │
└───────┬───────────────────────┬────────┘                        │
        │                       │                                  │
        ▼                       ▼                                  ▼
┌────────────────┐   ┌────────────────────────┐         ┌──────────────────────┐
│ Provider layer │   │        Tools           │         │   Object storage     │
│ (adapter +     │   │ read/find/list/fetch,  │◄───────►│   (file bytes)       │
│  registry +    │   │ generate, edit,        │         └──────────────────────┘
│  tiering)      │   │ research/integrations  │                  ▲
└──────┬─────────┘   └───────────┬────────────┘                  │
       │                         │                               │
       ▼                         ▼                               │
┌────────────────┐   ┌────────────────────────┐         ┌────────┴─────────────┐
│ Model providers│   │   External APIs        │         │   Relational DB      │
│ (Claude/Gemini/│   │   (+ cache table)      │◄───────►│  users, convos(events),│
│  OpenAI/…)     │   └────────────────────────┘         │  artifacts, versions, │
└────────────────┘                                      │  edits, sharing,      │
                                                        │  secrets(encrypted)   │
   [Secrets manager / env] ──► operator + user keys     └──────────────────────┘
   [Observability/tracing]  ◄── traces, tokens, events
   [Queue + workers]        ◄── heavy/async work (as scale demands)

The data flows: a request is authenticated, authorized, and validated at the API tier; the agent core builds context (loading durable state), runs the loop (calling the provider layer and tools), and streams an event timeline back; tools read/write object storage, the database, and external APIs; everything durable lives in the database and object storage; secrets come from the environment/manager; traces flow to observability; heavy work goes to workers via a queue when scale requires.

← → arrow keys work too