CHAPTER 08 · Parallelism, the Responses API, and the App Server · 6 / 8
Code MVP: a parallel tool executor
"""
chapter 08: parallel tool execution.
Run independent (read-only) tool calls concurrently, but force mutating
calls to run sequentially to avoid races. Results come back in order.
Builds on Chapter 6's registry and Chapter 7's gated_call.
"""
from concurrent.futures import ThreadPoolExecutor
# Which tools are safe to run in parallel (read-only) vs must be serialized.
READ_ONLY = {"shell_read", "read_file", "search", "list_files"}
def is_parallel_safe(call: dict) -> bool:
"""A toy classifier. Real harnesses also inspect the command itself."""
return call["name"] in READ_ONLY
def run_tools(calls: list, execute, max_workers: int = 5) -> list:
"""
calls: [{"call_id": ..., "name": ..., "args": ...}, ...]
execute: a function(call) -> result string (e.g. wraps gated_call).
Returns results in the SAME order as calls (like Codex's FuturesOrdered).
"""
parallel = [c for c in calls if is_parallel_safe(c)]
serial = [c for c in calls if not is_parallel_safe(c)]
results = {}
# Read-only calls run concurrently.
if parallel:
with ThreadPoolExecutor(max_workers=max_workers) as pool:
futures = {pool.submit(execute, c): c["call_id"] for c in parallel}
for fut in futures:
results[futures[fut]] = fut.result()
# Mutating calls run one after another, in order.
for c in serial:
results[c["call_id"]] = execute(c)
# Return in the original order so call_id linkage (Chapter 3) stays clean.
return [{"call_id": c["call_id"], "output": results[c["call_id"]]} for c in calls]
if __name__ == "__main__":
import time
def fake_execute(call):
time.sleep(0.5) # pretend each call takes 0.5s
return f"result of {call['name']}({call['args']})"
calls = [
{"call_id": "a", "name": "read_file", "args": {"path": "x.py"}},
{"call_id": "b", "name": "read_file", "args": {"path": "y.py"}},
{"call_id": "c", "name": "read_file", "args": {"path": "z.py"}},
{"call_id": "d", "name": "write_file", "args": {"path": "out.py"}}, # serial
]
start = time.time()
for r in run_tools(calls, fake_execute):
print(r["call_id"], "->", r["output"])
print(f"elapsed: {time.time() - start:.2f}s (3 reads in parallel + 1 serial write)")
Run it: the three reads finish together in about half a second instead of one and a half, and the write runs after them. Total time is roughly 1 second, not 2. That is the parallelism win, with the write safely serialized so it cannot race the reads.