The supervisor agent pattern is one of the most widely deployed multi-agent architectures. It organises agents into a two-tier hierarchy: a supervisor (also called a manager, orchestrator, or coordinator) that plans and delegates, and worker agents that execute specific tasks within a defined scope. The supervisor receives the high-level goal, decides which workers to engage and in what order, reviews their outputs, and synthesises a final result. This structure provides centralised control, natural quality checkpointing, and a clean separation between planning and execution. Understanding when the supervisor pattern is the right choice — and how to implement it well — is essential for anyone building multi-agent systems.
The Core Structure
A supervisor agent system has three main components: the supervisor, a set of workers, and a communication protocol between them. The supervisor is typically the most capable model in the system — it needs to understand the full task, know what each worker can do, plan a sequence of delegations, and evaluate whether worker outputs meet the required standard. Workers are specialised: each has a defined domain (web search, code execution, document analysis, API calls) and executes tasks within that domain reliably. Communication flows primarily through the supervisor: workers receive tasks from the supervisor and return results to the supervisor; workers typically do not communicate directly with each other. The supervisor uses the accumulated results to decide next steps until the overall goal is achieved or a stopping condition is met.
Why Centralise Control in a Supervisor
The alternative to a supervisor is decentralised coordination — agents handing off to each other directly based on local decisions. Decentralised approaches (like the swarm pattern) are more flexible but harder to control: it is difficult to enforce quality standards, prevent redundant work, or maintain a coherent overall strategy when no agent has the full picture. The supervisor pattern trades some flexibility for control: the supervisor always knows the current state of the overall task, can catch errors before they compound, can redirect the strategy based on what workers discover, and can decide when the result is good enough to return. For tasks where quality control and strategic coherence matter — complex software engineering, multi-step research, business process automation — the supervisor’s centralised view justifies the overhead.
Supervisor as Quality Gate
One of the supervisor pattern’s most valuable properties is natural quality gating: the supervisor reviews every worker output before deciding to proceed. If a worker’s output is incorrect, incomplete, or below standard, the supervisor can request revision (send the same worker a follow-up task with specific feedback), re-delegate to a different worker, or adjust the overall strategy. This quality loop is implicit in the architecture — it does not require a separate evaluator agent, because the supervisor’s review step serves that role. In practice, the supervisor’s review quality determines the system’s overall quality ceiling: a supervisor that cannot reliably assess whether a worker’s output is correct will propagate errors rather than catch them. This means the supervisor role needs the most capable model, and the additional cost is justified by the quality improvement it provides.
Figure 1 — Supervisor agent pattern: information flow
Implementing the Supervisor in LangGraph
LangGraph is the most natural framework for implementing the supervisor pattern, because its graph model makes the supervisor-worker relationship explicit. The supervisor is a node that receives state, decides which worker to call next (or whether to finish), and routes to the appropriate worker node. Each worker is a node that executes its task and returns results to the supervisor. LangGraph’s conditional edges implement the routing logic: after each worker completes, the graph returns to the supervisor node, which inspects the full state and decides the next step. This structure is directly inspectable in the graph definition, making the system easy to reason about and debug. State in LangGraph’s supervisor implementation typically includes the original goal, the list of completed worker outputs, and any intermediate findings that the supervisor needs for routing decisions.
Scaling the Supervisor: When It Becomes a Bottleneck
The supervisor pattern’s centralised control becomes a bottleneck as the number of workers and the complexity of the task grows. If the task requires 50 parallel operations, having every one of them report back to a single supervisor creates latency and context window pressure — the supervisor accumulates the outputs of all 50 workers in its context, which may overflow the window. Solutions for scaling: hierarchical supervision (a top-level supervisor delegates to mid-level supervisors who manage smaller groups of workers), asynchronous result collection (the supervisor dispatches many workers and collects results in batches rather than one by one), and summarisation checkpoints (the supervisor summarises accumulated results periodically to keep the context manageable). For most practical agent systems with 2–10 workers on tasks requiring 5–20 steps, the simple supervisor pattern scales adequately without these extensions.
Supervisor Pattern vs Direct Orchestration
A common question is whether a supervisor agent is necessary, or whether the orchestration logic can be implemented as regular code: a Python function that calls worker tools in the appropriate order based on hardcoded logic. The answer depends on how dynamic the task is. If the sequence of worker calls is always the same regardless of intermediate results, hardcoded orchestration is simpler and more reliable — no LLM inference needed for routing. If the sequence of worker calls depends on what workers discover (the search results determine which analysis to run; the code execution output determines whether to revise the code), then the supervisor agent’s ability to make routing decisions based on content is essential. The supervisor pattern adds value precisely when the orchestration requires understanding the meaning of intermediate results, not just their structure.
Failure Modes to Design For
The supervisor pattern has predictable failure modes worth designing against from the start. Supervisor context overflow: as the task accumulates worker outputs, the supervisor’s context fills — implement summarisation or selective retention to keep it manageable. Supervisor disagreement with worker quality: the supervisor may accept outputs that are subtly incorrect, especially if the supervisor and worker use different quality standards — validate the supervisor’s quality judgments on representative test cases. Delegation loops: the supervisor delegates to a worker, the worker returns inadequate output, the supervisor re-delegates with more specific instructions, the worker still fails — implement a maximum delegation count per subtask. Worker unavailability: design the supervisor to have fallback workers or graceful degradation paths when a specific worker fails consistently, rather than retrying the same worker indefinitely.
The supervisor pattern is the right architectural choice when you need centralised control over a multi-step task, quality gating between steps, and the ability to route dynamically based on intermediate results. Its overhead — the supervisor’s inference cost, the two-tier communication structure, the context management requirements — is justified by the quality improvements that centralised review provides. Build the supervisor as the most capable agent in the system, design explicit stopping conditions, test the quality gating on representative inputs, and implement summarisation early to prevent context overflow on longer tasks.
Defining Worker Scope Precisely
The quality of a supervisor agent system is heavily determined by how precisely each worker’s scope is defined. Overlapping scopes — two workers that both could handle a given subtask — cause the supervisor to make inconsistent routing decisions and can lead to duplicated work. Gaps in scope — subtasks that no worker covers — cause the supervisor to either attempt the task itself (consuming reasoning capacity it should be using for coordination) or fail the task entirely. The best practice is to define each worker’s scope as a precise, non-overlapping domain: “Worker A handles web searches only, returning raw search results without interpretation,” “Worker B handles Python code execution only, returning output and errors without analysis,” “Worker C handles document analysis only, returning structured summaries.” Tight scope definitions also make it easier to test each worker independently and to replace a worker implementation without changing the supervisor’s routing logic.
The supervisor pattern scales from simple two-worker systems to complex hierarchical architectures with many layers of delegation. Starting simple — one supervisor, two or three workers — and adding complexity only as the task requires it is the most reliable development path. The core value of the pattern is always the same: centralised quality control and strategic oversight that no individual worker has the scope to provide on its own.