CHAPTER 03 · The Case Against: Cognition's "Don't Build Multi-Agents" · 7 / 9
Real-world examples Yan offers
The article grounds the theory in two examples worth remembering:
Claude Code subagents. As of mid-2025, Claude Code does spawn subtasks, but it never runs work in parallel with the subtask agent, and the subagent is usually only asked to answer a question, not to write code. Why? The subagent lacks the main agent's full context, so it can only safely handle a well-defined question. Running multiple parallel subagents would risk the conflicting-decisions problem. The benefit it does capture is that the subagent's investigative work stays out of the main agent's history, allowing longer traces before running out of context. This is a sub-agent used for context isolation, not for parallel labor, and the design is deliberately simple.
Edit-apply models. In 2024, many models were bad at editing code, so a common practice was to have a large model write a markdown explanation of the changes, then feed that to a small "edit apply" model to rewrite the file. These systems were faulty: the small model often misread the large model's instructions over the slightest ambiguity. Today the editing decision and the applying are more often done by a single model in one action. The lesson mirrors the principles: splitting a decision across two models introduced a communication gap that caused errors.