CHAPTER 04 · Practical Lessons from Building Manus · 4 / 7
Verify with computable, binary checks: the Intern Test
The articles are skeptical of soft, subjective evaluation. Static benchmarks like GAIA saturated quickly and did not match real user satisfaction. The recommendation is to focus on tasks that are computationally verifiable, with clear yes-or-no outcomes:
- Did the code compile?
- Did the file exist after the command ran?
- Can the sub-agent verify the output of the parent?
Part 2 calls this the "Intern Test": prefer binary success or failure metrics on real environments over subjective LLM-as-a-judge scores. The point is not that LLM-as-a-judge is useless, but that hard, checkable signals are more trustworthy when you can get them.