Skip to slide
Chapter 14 · Autonomous Task Agents: Manus and CodeAct
101 / 142

CHAPTER 14 · Autonomous Task Agents: Manus and CodeAct · 2 / 10

CodeAct: code as the action language

Here is Manus's signature idea. Most agents act through structured tool calls: the model emits something like {"action": "search", "query": "..."} and the harness runs it. CodeAct instead lets the model write a short Python program as its action. Need the weather? It writes Python that imports a client, calls the API, and prints the result. The sandbox runs that code and returns the output (or the error) as the observation.

Why is this powerful? Because code is far more expressive than a fixed menu of tool calls. One snippet can chain several operations, add conditionals, loop, and use any library. The CodeAct paper (ICML 2024) found that agents which produce code for actions have significantly higher success rates on complex tool-use tasks than those limited to text or JSON tool calls. The action space is the whole language, not a handful of functions. And the agent can debug itself: if the code throws, it reads the error and rewrites the code, exactly like a developer at a REPL.

The flip side, and Manus handles this carefully, is safety. Arbitrary code execution is the most dangerous thing an agent can do, which is why Manus runs it in a locked sandbox and follows strict rules (one action per step, prefer non-interactive flags, never run irreversible operations without permission). Everything from Chapter 7 applies, doubly.

← → arrow keys work too