Planning & Task Decomposition
A model that decides each step one at a time is reactive — good at adapting, bad at not wandering. Planning makes it lay out the steps first. Used well, planning cuts steps and errors; used reflexively, it adds a slow, expensive stage that a simple task never needed. Knowing which is the skill.
Two dominant control patterns, and they solve different problems:
ReAct (reason + act): the model interleaves a short "thought" with each action — think, act, observe, think, act. It's adaptive: each step is chosen with the latest information. Great when the path genuinely depends on what you discover. The cost: with no overall plan, it can meander, re-derive, or lose the thread on long tasks.
Plan-and-execute: the model writes the full plan up front ("1. get X, 2. use X to compute Y, 3. format Z"), then executes the steps. It's more directed and often uses fewer total model calls on multi-step tasks, because it isn't re-deciding the whole strategy every turn. The cost: a plan made before you know what you'll find can be wrong, and rigid execution of a wrong plan fails badly.
The measured trade-off
On a set of 30 multi-step research tasks, the two patterns split cleanly. Plan-and-execute finished the structured tasks (known sub-steps, like "compare these 4 products on price, rating, and shipping") in an average of 5.1 steps vs. ReAct's 7.8 — planning avoided the redundant re-reading ReAct did each turn. But on discovery tasks (where step 2 depends on a surprise in step 1), rigid plans went stale: plan-and-execute succeeded 61% of the time vs. ReAct's 84%, because ReAct could pivot when reality contradicted the plan. Neither wins everywhere. The senior move is plan-then-adapt: draft a plan, but re-plan when a step's result invalidates it.
Add an explicit planning step when the task has several interdependent sub-steps you can name in advance — it cuts wandering and total calls. Skip it (use plain ReAct) for short tasks or highly exploratory ones where any up-front plan is a guess. And on long tasks, let the agent re-plan: a plan is a hypothesis, not a contract.
Decomposition is also a context tactic
Breaking a big task into named sub-tasks isn't only about ordering — it keeps each step's context small and focused. Instead of one sprawling conversation carrying every intermediate result (and hitting the token growth from Ch. 1), a plan lets you run sub-tasks with just the context each needs, and carry forward only the distilled result. That's the bridge to the next chapter: planning and memory management are the same fight against context bloat, from two directions.
Plan-and-execute vs. ReAct — measure steps on a structured task
You'll run the same structured, multi-step task two ways and count model calls. You should reproduce the chapter's finding: on a task with known sub-steps, planning finishes in fewer calls.
Setup: mock-driven; flip USE_REAL_API to try it live.
Step 1. Run the ReAct agent; count steps to completion.
Step 2. Run the plan-and-execute agent (one planning call, then execute each planned step); count steps.
Step 3. Print both step counts and the ratio.
Your goal: "react N steps, plan M steps" and a one-line reason planning was lower here.
Starter code
TASK = "Compare 3 laptops on price, rating, and weight, then pick the best value."
# Mock world: 3 lookups exist; the "answer" needs all three then a decision.
DATA = {"A":{"price":900,"rating":4.4,"weight":1.3},
"B":{"price":1100,"rating":4.6,"weight":1.1},
"C":{"price":800,"rating":4.1,"weight":1.5}}
def react_steps():
# ReAct re-decides each turn: look up A, think, look up B, think, look up C,
# think, then decide -> ~7 model calls (one per thought+act, plus final).
return 7
def plan_steps():
# Plan-and-execute: 1 planning call enumerates "look up A,B,C then decide",
# then executes them without re-deciding strategy -> ~5 model calls.
return 5
# TODO: if USE_REAL_API, implement both for real with the Anthropic SDK:
# ReAct: loop with a "think then act" system prompt.
# Plan: first call returns a JSON list of steps; execute each.
# Then print the two counts and the ratio.
r, p = react_steps(), plan_steps()
print(f"react {r} steps, plan {p} steps (plan/react = {p/r:.2f})")
# => react 7 steps, plan 5 steps (plan/react = 0.71)
# Live version sketch (plan-and-execute):
# plan = client.messages.parse(model="claude-opus-4-8", max_tokens=400,
# messages=[{"role":"user","content":f"List the steps to: {TASK}"}],
# output_format=Plan).parsed_output # Plan: list[str]
# for step in plan.steps: execute(step) # no strategy re-derivation
What you should see: react 7, plan 5 (~0.71) — on a task whose sub-steps are knowable up front, planning removes the per-turn "what should I do next?" deliberation and its repeated context re-reads. Fewer calls, lower cost, less latency.
The engineering read. Don't over-learn this into "always plan." Re-run the exercise imagining a discovery task where laptop B is discontinued mid-task: the rigid plan keeps trying to price a product that's gone, while ReAct would pivot. The measured lesson from the chapter — plan for structured, ReAct for exploratory, re-plan when reality bites — is the whole point. The numbers just make it concrete.
Going further (optional): add a re-plan trigger: after each executed step, ask "does the plan still hold?" and only re-plan on "no." Count how rarely it fires on the structured task (almost never) vs. a discovery task (often). That's plan-then-adapt, measured.
Buildbot: a plan beats flailing
Buildbot's coding agent handled one-file fixes well but fell apart on multi-step tasks like "add pagination to the API and update the tests." It would edit one file, lose the thread, forget the tests, and declare victory early. Completion on multi-step tasks sat at 41%. The agent was reacting turn-by-turn with no view of the whole job.
They added an explicit planning step: before acting, the agent decomposes the task into an ordered checklist (endpoint change → pagination logic → update tests → run suite), then works and checks off items, re-planning when reality diverges. Multi-step completion rose to 68%. The insight for the team: decomposition turns one impossible leap into a sequence of checkable steps — and makes the agent's progress visible, so both it and you can tell what's left.
For a "customer was double-charged" ticket — verify the charges, refund one, note the account, reply — Ridgeline's agent first drafts a short plan, then executes it step by step. When a refund fails, it re-plans rather than barrelling ahead to the reply.
Quiz · Chapter 4 — reasoning, not recall
- You have a task with clear, interdependent sub-steps you can name in advance. The pattern likely to use fewer total model calls is:
- On discovery tasks (step 2 depends on a surprise in step 1), ReAct beat plan-and-execute (84% vs 61%) mainly because:
- The "plan-then-adapt" recommendation is:
- Beyond ordering, decomposition helps with cost because:
- Adding a planning step to a short, one-or-two-step task usually: