Skip to slide
Chapter 3 · Building the Prompt
14 / 142

CHAPTER 03 · Building the Prompt

Building the Prompt

In Chapter 2 we hand-waved "the harness builds a prompt and calls the model." This chapter opens that step up, because what goes into the prompt, and in what order, is one of the most consequential design decisions in the whole harness. The "Seven Pillars" breakdown calls this context engineering, and it lists it as pillar number one for a reason.

When you query a modern agent API (OpenAI's Responses API, Anthropic's Messages API), you do not hand it one big string. You hand it a structured request with a few distinct fields. Codex's request, for example, has three you really need to understand:

  • instructions: the system and developer guidance, the model's "personality" and rules.
  • tools: a list of tool definitions, each one a JSON schema describing what the tool does and what arguments it takes.
  • input: the conversation itself, as a list of typed items (user messages, assistant messages, tool calls, tool results).

The server then assembles these into the actual token sequence the model sees. Conceptually, think of the final prompt as a single ordered list of items, each tagged with a role that signals how much authority it carries. In decreasing order of priority: system, developer, user, assistant. System rules outrank user requests, which outrank the assistant's own past words.

Layered instructions

The instructions are not one block of text written by one person. They are layered, and the layering is deliberate. Reading across the Codex write-up and the "Inside the Agent Harness" deep-dive, a typical stack looks like this:

LayerSourcePurpose
Base instructionsBundled with the model (e.g. gpt-5.2-codex_prompt.md)The model's core behavior and "personality"
Developer instructionsThe user's config fileProject- or developer-level overrides
User instructionsAGENTS.md / CLAUDE.md, aggregated from several filesConventions, test commands, house style
Skills metadataConfigured skillsShort descriptions of available procedures (Chapter 10)
Environment contextComputed at runtimeCurrent working directory, shell, sandbox rules

The rule for aggregating user instructions is "more specific wins." Codex walks from the project root down to your current working directory, and instructions closer to where you are working take precedence, up to a size limit (32 KiB by default). This is layering you will rebuild in Chapter 9.

The conversation as structured items, not a transcript

Here is a detail that separates a toy from a real harness. The history is not stored as a flat string like "User said X, then I ran Y, then I saw Z." It is stored as typed, structured items, and crucially, a tool call and its output are linked by a shared call_id.

Why does that linkage matter? Because the model needs to understand cause and effect. If you ran three shell commands in parallel and dumped three blobs of output, the model would not know which output came from which command. The call_id ties result to request, so the model reads the causal relationship correctly. The "Inside the Agent Harness" analysis flags this as one of the things that makes structured tool outputs critical: the model needs to know what happened, not just what text appeared.

Tool definitions as JSON schemas

The tools field is a list of JSON schemas. Here is Codex's shell tool, lightly trimmed:

{
  "type": "function",
  "name": "shell",
  "description": "Run a shell command and return its output",
  "parameters": {
    "type": "object",
    "properties": {
      "command":    {"type": "array", "items": {"type": "string"},
                     "description": "The command to execute"},
      "workdir":    {"type": "string", "description": "Working directory"},
      "timeout_ms": {"type": "number", "description": "Timeout in milliseconds"}
    },
    "required": ["command"]
  }
}

The model reads these schemas to know what it is allowed to ask for and how to format the request. The harness dynamically decides which tools to include based on permissions, sandbox mode, and context. For example, Codex only adds a view_image tool when an image is actually present. Fewer, well-described tools beat a giant undifferentiated pile.

← → arrow keys work too