Naming the Workflow Patterns in the Pipeline
A colleague sketches an "ask the web" agent as: rewrite the query, call a search API, fetch and extract the top 8 pages one after another, then feed everything to an LLM to write the answer. They call the whole thing "the agent."
- Using the named vocabulary from
agent-design-patterns-workflows-to- autonomous-agents, identify which pattern(s) this design actually uses, and say precisely why "the agent" is the wrong word for it. - Identify the one change to this design that would most improve latency without changing cost, and name the pattern it introduces.
- Would adding "if the sources are weak, search again with a refined query, up to once" change your answer to (2)? Why or why not?
1. Naming it correctly
This is a prompt chain — rewrite → search → fetch/extract →
synthesize, each step's output feeding the next, in a fixed order
decided by code, not by a model deciding its own next step at run
time. Calling it "an agent" is the exact mistake
agent-design-patterns-workflows-to-autonomous-agents warns against:
the defining feature of Level 2 autonomy is that the model decides its
own next step, and nothing in this design does — every stage always
runs, in the same order, for every query. It's Level 1, full stop.
2. The latency fix
Fetching the 8 pages serially is the bottleneck: at roughly 1–1.5s per fetch, that step alone costs 8–12s before synthesis even starts. Fetching them concurrently, bounded by a per-page timeout, cuts that step to roughly the time of the slowest single fetch (with a timeout) instead of the sum — often a 5–8x reduction on that step alone, at zero extra dollar cost since the same 8 fetches happen either way. This is parallelization-sectioning: the 8 fetches are different work (different pages), run concurrently, which is exactly the distinguishing question that pattern gives — different work run concurrently is sectioning; the same work run redundantly would be voting, which this isn't.
3. The bounded retry
No — it doesn't change the Level 1 classification, provided the retry is capped at a fixed number (here, once) and triggered by a simple threshold check (confidence below a bar) rather than the model deciding, in an open-ended way, whether and how many times to retry. A single bounded retry from a fixed rule is still control flow decided in code, just with one conditional branch — a workflow can branch. The moment the number of retries becomes something the model decides per query rather than a fixed cap, that's the point it climbs to Level 2, and it should be a deliberate design decision to cross that line, not something that happens by accident because "let it try again" sounded harmless.
Share this question