Mohammed Mutahar
← Back to Projects

Conductor

Four specialized agents coordinated by a state graph, running as durable background jobs behind a REST API, with state checkpointed after every step so a crashed worker resumes instead of starting over.

Multi-Agent SystemsAgent OrchestrationLangGraphDistributed SystemsFastAPICeleryState ManagementBackend ArchitectureDockerPython
Conductor

Conductor

Multi-agent research orchestration, built to be read

Conductor takes a topic, breaks it into research questions, searches the web for each one, finds the patterns across what it gathered, and writes you a report. Four specialized agents, coordinated by a state graph, running as a background job behind a REST API.

The four-agent research pipeline is not the interesting part. Everyone builds that. The interesting part is everything around it that makes it survive contact with reality.


The gap this was built to close

Almost every agent orchestration example you'll find is a Python script. You run it, it blocks your terminal for three minutes while making a dozen LLM calls, it prints a report, and if anything fails at minute two — a rate limit, a search timeout, a dropped connection — you lose all of it and start over. That's fine for a blog post. It's not a system.

The distance between that script and something you'd actually deploy is almost entirely unglamorous infrastructure: how requests are accepted without blocking, where work executes, what happens to state when a process dies, how a caller finds out whether their job finished. None of that is about agents. All of it is what makes agents usable.

I built Conductor to close that gap and to be legible while doing it. It's written as a reference implementation — the kind of codebase where someone learning orchestration can open workflow.py and see the entire control flow in one screen. Optimizing for readability over cleverness was a constraint I set deliberately, and it shaped a lot of the decisions below.


The shape of a request

Submitting a topic returns immediately with a run_id and a URL to poll. Nothing about the actual work happens on the request thread.

POST /api/v1/runs  →  { run_id, status: "pending", poll_url }

Behind that, FastAPI enqueues a task to Redis, a Celery worker picks it up, and the worker executes the LangGraph workflow. The caller polls their poll_url and watches status move from pending through running to complete, at which point the final report is in the response body.

This is a boring pattern and that's the point — it's the pattern, because a research run takes minutes and no HTTP request should be held open for minutes. Getting the async boundary in the right place from the start is much easier than retrofitting it once someone's built a frontend against a blocking endpoint.


Why a graph and not a chain

The agents don't run in a straight line. The researcher needs to loop — one pass per research question, with a decision after each one about whether there's more to do.

LangGraph models this as a directed graph where each agent is a node taking state in and returning a partial state update, and the loop is expressed as a conditional edge with a router function deciding where to go next:

workflow.add_conditional_edges(
    "researcher",
    should_continue_research,
    {"continue_research": "researcher", "synthesize": "synthesizer"}
)

The value of writing it this way isn't that it's shorter than a while loop — it isn't. It's that the control flow becomes a data structure you can inspect, checkpoint, and reason about, rather than a call stack you can only observe by watching it run. That property is what everything below depends on.

The state is the contract

All four agents share one AgentState TypedDict. Each node reads the fields it needs and writes back its own outputs. The planner writes questions; the researcher appends findings; the synthesizer adds themes; the writer produces the report.

I like this for the same reason I'm wary of it. It's radically legible — one type definition tells you everything that flows through the system, and adding an agent means adding a field and an edge. It's also a shared mutable blackboard, which is a design that gets worse with every agent you add. At four agents it's the clearest possible choice. At fifteen it would be a source of invisible coupling, where agents depend on fields written by agents they know nothing about. It's the right call for this system and I wouldn't defend it as the right call for a much larger one.


Checkpointing is the feature

State is written to Postgres after every single node. If the worker crashes mid-run, the run resumes from the last completed node instead of starting over.

This is the design decision I'd defend hardest, because it addresses the thing that actually makes agent systems painful to operate: the work is expensive and the failure rate is nonzero. Every node is one or more LLM calls plus web requests. A run that dies three nodes deep has already spent real money and real minutes, and the failure is very often something transient that would succeed on retry — an API hiccup, a rate limit, a worker restart during a deploy.

Without durable state, the correct response to a transient failure is to throw away all the successful work and redo it. That's the default in most agent code, and it's why people don't trust these pipelines with anything long-running. Persisting state per node turns a class of catastrophic failures into a class of annoyances.


Two ways to run it

There's a Docker path that brings up Postgres, Redis, the API, a worker, and Flower for task monitoring in one command, and a Docker-free path that points at free cloud-hosted Postgres (Neon) and Redis (Upstash) instead.

The second path exists because "install Docker first" is where a meaningful fraction of readers stop, and a reference implementation nobody can start isn't a reference. The same reasoning covers the Windows-specific notes scattered through the docs — Celery's default multiprocessing pool fails on Windows and needs --pool=solo, make isn't there by default. Small, unglamorous, and the difference between a project that runs on someone's actual machine and one that runs on mine.


What's honestly missing

There's no evaluation. This is the real gap. The system reliably produces a report; nothing in it establishes whether the report is any good. Are the claims grounded in the sources the researcher actually retrieved, or did the writer smooth over gaps with plausible filler? Did the planner's decomposition cover the topic or miss an obvious dimension? Does the synthesizer find real patterns or manufacture themes? A multi-agent pipeline without an eval harness is a system whose quality you're taking on faith, and per-component evaluation — planner coverage, researcher groundedness, writer citation fidelity — is the obvious and necessary next build.

The multi-agent structure hasn't been forced to justify itself. Two of the four agents have no tools; they're LLM calls over accumulated state. It's entirely possible that a single well-prompted call could match a synthesizer-then-writer handoff on output quality at a fraction of the cost and latency. I believe the separation helps — narrower jobs generally produce better output — but belief isn't measurement, and the right way to settle it is a single-prompt baseline to compare against. That's the same missing eval harness from a different angle.

Cost and latency are unbounded by design. The research loop runs until it decides it's done. There's no budget ceiling, no wall-clock cap, no cap on total tokens spent per run. Fine for a local reference implementation, unacceptable for anything with users.

Research quality is capped by one search tool. Everything the system knows comes through DuckDuckGo results. No source ranking, no primary-source preference, no cross-checking between results.


Stack

LangGraph for the workflow graph, Gemini for agent reasoning, FastAPI for the API layer, Celery and Redis for background execution, Postgres for run records and state checkpoints, Docker Compose to tie it together, and pytest for tests.

Related Projects