← Back to Blog
How to Manage 50 AI Agents with Codex

How to Manage 50 AI Agents with Codex

I have tried to watch several Codex conversations at once. With two or three, I felt capable. By the fifth window, the questions were lining up. One agent needed me to remember why we changed an algorithm last week. Another wanted the boundary of a code change. A third had finished an experiment and wanted to know whose call came next. The agents were working in parallel; I was paying the toll at every context switch.

Ten conversations are already hard to follow. Every new agent adds capacity and another line for questions to reach me. If I want to work with 20 or 50 of them, I have to change who answers those questions. I want to speak with a few long-term owners who manage the work below them. Fifty is the design goal, not a productivity number I have measured.

Why start with a Chief of Staff

I start with the Chief of Staff because setting up each new owner would otherwise consume my attention. A client project, an internal engineering effort, and a learning track all need duties, communication rules, authority, and a place for records. I tell the Chief of Staff what I want. It drafts the new director's role, and I decide whether to create it.

Its job resembles hiring and developing directors. A director's role, decisions, and history live in Git. If a conversation goes off track, the Chief of Staff can rebuild it from those records. It also checks that scheduled work is running and reviews direct exchanges with one director at a time. It educates every director and coaches the one whose exchange reveals a problem.

A mistake in one project can expose an instruction that every director needs. If one director loses a decision during a handoff, correcting only that conversation leaves the next director open to the same failure. The Chief of Staff explains what went wrong, turns the correction into a clear instruction, and passes it to every role. Each director then knows what to record or ask before its own next handoff.

The Chief of Staff does no project work. A health check looks at schedules, run status, and blockers; it does not turn into a review of business reports or code. Keeping this layer in shape saves me from training every new owner myself.

Why each long-term project needs a director

The two names matter. In my earlier workflow, I opened a full Codex session and gave it a task. It could read files, change code, and run checks itself. It could also call other agents for help. I called that session a manager because it could both work and manage.

A director is also a full Codex session, but I give it a long-term project instead of an execution task. It keeps goals, preferences, and decisions. When work arrives, it starts a manager. That manager is the director's sub-agent. The manager can do the task itself or spawn its own sub-agent. The director supplies context, handles questions, reviews the result, and reports to me. I use the new name to distinguish this role from the original manager who still does hands-on work.

One director owns one long-term project or clear area of work. Two workstreams for the same client can have different directors if their goals and decision cycles differ. I speak with a few directors who know their projects instead of briefing each new manager from the beginning.

The director stays in its parent conversation while several managers run. I can give it more work before they finish. That is how I can move from managing a few conversations toward managing dozens of agents. Dependencies, review capacity, and cost still set the practical limit.

The team: directors manage managers

Illustrated Codex team architecture: human, Chief of Staff, project directors, managers, and their sub-agents

The counts in the diagram are illustrative. Two parent-child relationships matter: a manager is the director's sub-agent, and a manager may spawn sub-agents of its own.

If every problem still comes to me, the boxes in the diagram have achieved nothing. Each layer needs to know what it owns and when to pass a decision upward. I still set the goals, priorities, and tradeoffs that need a person. The rest of the boundaries fit in one table:

RoleOwnsDoes not do
MeSet goals and priorities; make decisions that require a personWatch every manager and sub-agent; resolve each routine blocker for a director
Chief of StaffEstablish and educate every director; check that scheduled work runs; pass corrections to every roleDo project work; review business reports or code during health checks
Project directorKeep project context; start managers; handle questions; review resultsExecute concrete tasks itself; forward every small question to me; act beyond its authority
Authorized managerTake a task as the director's sub-agent; do it or spawn another sub-agentChange project goals or permissions on its own; pass unchecked results to the director
Manager's sub-agentComplete a task from its manager and report backBypass the manager to ask me or the director; change the task goal on its own

Questions move up this line: sub-agent to manager, manager to director, director to me only when a tradeoff or new authority needs my decision.

The director sets each manager's authority through the task it assigns. A code quality manager may fix a scoped issue directly. Another manager may only direct its own agents. The title alone does not expand either role's authority.

What happens when I start an experiment

I used an experiment as an example in the meeting: clean the data, change the algorithm, then validate the model. I could open three conversations and keep track of who is waiting for whom. Instead, I give the experiment to its project director. It starts one manager for data cleaning and another for the algorithm change. Validation needs the cleaned data, so it starts when that input is ready.

Each manager is the director's sub-agent and a full Codex session. The data cleaning manager can read files, write code, and run checks itself. If the work merits another split, it can spawn its own sub-agent. The algorithm manager has the same choice. When Claude Code fits the task, a manager can call it within its scope or ask its own agent to call it. The director does not write the code for them. It checks whether their results fit together as one experiment.

Want more practical breakdowns?

AI, engineering, and experiments. One or two useful emails a month.

No spam. Unsubscribe anytime.

Most of the time, the manager works quietly in the background. It either does the task or directs its own agents. An agent asks its manager first; the manager asks the director only when needed. The director remains in the parent conversation, so I can discuss the next task while work continues. It comes to me for a new permission, a change of goal, or a tradeoff that needs my judgment.

Bring an important manager conversation into view

For some work, I need to see the reasoning, not just hear the director's summary. If an algorithm choice is disputed, I may want to ask the person doing the work why another approach was rejected. The director can bring that manager into a visible Codex conversation, and I can talk with it directly. Afterward, the manager summarizes the new decisions, open questions, and results back to the director. The director can read the exchange too, update the project record, and keep following the work. I open this view when my judgment is useful; the manager otherwise works quietly in the background.

How Codex works with Claude Code

Someone asked whether a coding task suited to Claude Code would force us to manage the same project in two applications. I do not want to explain the project again each time I switch coding tools. Codex remains the project entry point: I talk to the director, the director gives the task to a manager, and an authorized manager can use Codex tools, call Claude Code CLI in a separate working directory, or ask its own sub-agent to call it.

Claude Code supports non-interactive calls with claude -p. The --output-format json option returns a result that another tool can read. A schematic call looks like this:

claude -p "Make the scoped code change in this working directory. Run the relevant checks. Return changed files, check results, and open questions. Do not commit." --output-format json

The manager adds the file scope, acceptance criteria, and authorized actions. Claude Code still uses its own authentication and tool permission settings. When several coding agents run at once, they can use separate working directories or non-overlapping files. The Claude Code CLI reference and non-interactive mode guide describe these calls.

Claude Code returns its output to the manager or the lower-level agent that called it. The result then follows the manager → director path. The director reviews changes and tests before giving me a project-level conclusion. I do not maintain a second project conversation in another app.

How projects share context

Projects run into similar problems. If every director researches the same question from scratch, the earlier work is wasted. Copying a whole conversation creates a different problem: the next director must dig through it to find the answer. I use two ways to share context.

For a lesson that needs to last, a director writes the conclusion, result, and source into its Git project record. Another director reads the relevant record when needed instead of loading the whole project history.

To inspect the original discussion, copy the Codex conversation's deep link and give it to a director with access. It can open the conversation, read the parts relevant to its project, and save a distilled lesson in its own Git record if the lesson will be reused. For example, Project B can read Project A's note about a data validation method. If the note lacks an important detail, B follows the deep link to the original discussion and checks when the method applies.

When a project ends, its director can hand the record or conversation link to the next owner. The recipient condenses and saves the experience that should carry forward.

Scheduled checks keep work moving without flooding me

A director owns a project for the long term, but an open conversation does not tell it when to revisit a task. I set a schedule for each role. One director may check every hour during the workday; another may need one check a day. Each trigger starts one check, and a new run should not overlap one still in progress.

The point is not to send me an “all good” note every hour. If nothing material changes, it stays quiet. It reports important progress, blockers, risks, or a decision that needs me.

That is the structure. To keep it working, I also need three small habits: a way to remember decisions, a way to question a task before starting it, and a way to pass an instruction down without changing its meaning.

Tip 1: Keep project memory in Git

A director is meant to stay with a project. Its chat history keeps growing. Loading all of it for every task is expensive and buries the current question; relying only on the chat makes a replacement conversation hard to rebuild. I keep the durable parts in Git.

Each director has a name, a job description, and a Git folder. The folder holds its role, working preferences, current state, important decisions, and a chronological history. One possible layout is:

projects/forecasting/
  role.md         # Role and authority
  current.md      # Current goals, tasks, blockers
  decisions.md    # Key decisions and reasons
  history/        # Older handoffs and work records

The file names can differ. A director normally reads its role, current state, and records relevant to today's task. Older history stays in files until it is needed. If the director drifts or its conversation must be replaced, the Chief of Staff has a record from which to rebuild the role.

Tip 2: Use a Socratic question to check task understanding

Socratic questioning is a way to test an idea through probing questions. What exactly do you mean? What evidence supports it? Which assumption is hiding inside it? What follows if that assumption is wrong? The questioner does not rush to supply the answer. The questions make the other person's reasoning visible. The University of Connecticut's introduction groups these questions around clarification, assumptions, evidence, consequences, and alternative views.

I borrow the move, not the entire philosophical dialogue. An agent can hear “improve the forecasting model” and start experiments immediately. It may reduce prediction error only to learn that runtime was the constraint I cared about. In its first substantive reply, I want it to restate the goal in one sentence and ask a question that exposes the tradeoff: “If lower error doubles runtime, which measure should this experiment protect first?” That is more useful than “May I begin?” One question finds a gap in the goal; the other merely adds a confirmation step.

The question should not freeze authorized work. A manager can clean the data and inspect the current baseline while the tradeoff is being resolved. Only a choice that truly needs my judgment moves up through manager and director. Scheduled checks do not ask the question again. Its value is at the task's entrance: find a bad assumption before three agents spend a day running toward different finish lines.

Tip 3: Borrow ASD-STE100 principles for clearer writing

ASD-STE100 is Simplified Technical English, a controlled form of English for technical documentation that grew out of aerospace maintenance writing. It has two parts: writing rules and a controlled dictionary. The aim is to make an instruction mean the same thing to readers with different backgrounds. The official ASD-STE100 standard contains the full rules and dictionary.

With 50 agents, the question is not only who reports to whom. It is also what remains of a sentence after several handoffs. “Process the data, run it, and let me know if anything looks wrong” sounds ordinary. For a manager, it leaves too much open: process how, run what, and what counts as wrong? I would write: “Manager A removes duplicate records and records how many it removed. Manager B runs the baseline experiment on the cleaned data. If a required input field is missing, report it to the director before filling any value.” The actor, action, output, and escalation condition are in the instructions.

In Chinese conversations, I borrow the underlying habits: concrete words, short sentences, and one name for each concept. For English technical handoffs, I can go further and use ASD-STE100's rules and controlled vocabulary. The Chinese example borrows a writing approach; it does not claim compliance with an English standard. The more layers a message crosses, the more I am willing to spend a few seconds making it precise.

Fifty agents is not a scorecard. If I still have to patrol 50 windows, I have only turned the mess into an org chart. The design earns its keep when work moves forward with directors and managers, and the decisions that matter still find their way back to me intact.

New ideas, straight to your inbox.

AI, engineering, and experiments. One or two useful emails a month.

No spam. Unsubscribe anytime.