AI Coding Agents Can Now Work for Days. So Why Aren’t Teams 10x Faster?

Software engineer orchestrating multiple AI coding agents across parallel tasks, dependencies and review checkpoints
Aamer Rasheed
Aamer Rasheed
Founder & SaaS/AI Solutions Architect, Digital Sensei Technologies

AI coding agents are becoming capable of working for hours, and in some environments, days. They can inspect repositories, modify multiple files, run tests, use development tools, create subagents, save intermediate work and continue after interruptions.

That sounds like the beginning of a huge productivity jump. So if I can run five capable AI coding agents at the same time, why should I not become five times faster?

Because coding capacity is no longer the only bottleneck. Orchestration is becoming the harder problem.

What has actually changed?

AI coding started with autocomplete. Then developers began using chat assistants to generate functions, explain errors and review code. Agentic coding is different. An AI coding agent can now receive a goal, inspect the existing project, decide which files are relevant, execute commands, change code, run tests and continue working through several steps.

OpenAI's Agents API, introduced in September 2026, is specifically designed for agents that can work reliably for days while preserving progress, using tools and coordinating subagents.

Google has reached a similar conclusion. Its Agent Executor was built because workflows lasting hours or days become difficult to operate reliably unless they can recover from interruptions and resume from saved state.

The interesting engineering problem is therefore shifting. It is becoming less about:

Can AI write this code?

And more about: Which agent should work on what, what context should it receive, which tasks can safely run together, and when should a human intervene?

That is orchestration.

More agents do not automatically mean more productivity

When I think about my own work, this is perhaps the biggest misconception around multi-agent development. If somebody gave me five highly capable coding agents tomorrow, I would not immediately become five times faster. My first problem would be deciding which five tasks can safely happen at the same time.

Some work is naturally parallel. One agent may investigate a frontend performance issue while another reviews an unrelated backend module. But many software tasks have dependencies.

A frontend screen may depend on an API contract that has not yet been finalized. An API may depend on a database structure that is still changing. Two agents editing the same module may both produce perfectly reasonable changes that conflict when combined.

Current OpenAI guidance makes the same distinction. It recommends multi-agent execution for independent, bounded work, but suggests keeping dependent steps together and coordinating agents that touch the same files.

The ability to create more workers does not remove dependency management. It makes dependency management more important.

My approach has been to design first, then let AI accelerate execution

In a recent AI-enabled SaaS project, I first designed the overall system before asking AI coding agents to implement features. I worked out the major user types, what each portal needed to do, how onboarding should work, what information had to be stored and how the important workflows connected. Only after that did implementation begin.

I did not give an agent a vague instruction such as: "Build the backend." I started with a bounded piece of work.

Set up the backend foundation. Then authentication. Then the basic SaaS portal structure. Then connect the completed backend capability with its relevant frontend flow. Then test it before moving forward.

This pattern continued through the project. I found that AI agents performed extremely well when the work was converted into clear executable steps.

The best results often happen before the agent writes any code

Before implementing an important feature, I normally discuss it with the AI coding agent first. I treat that discussion almost like a technical meeting with an experienced team member. I explain the business requirement and the existing architectural boundaries. The agent may suggest several implementation approaches.

I review those options, choose the one that fits the project, refine the flow, define the required API behaviour and remove ambiguity. Then I ask an important question:

Is anything still unclear or missing before implementation starts?

Only when the requirement is sufficiently clear do I let the agent proceed. For something as simple as a change-password feature, that can mean agreeing on the UI flow, API contract, request parameters, validations, expected responses and error behaviour first.

That preparation may look slower than immediately generating code. In practice, I have found the opposite.

More thinking before execution saves significant time during implementation.

An AI agent can move very fast once it knows exactly where it is going. It can also move very fast in the wrong direction when it does not.

Long-running agents create a context problem

Another limitation becomes obvious when you work with AI coding agents over many days. Context is not unlimited. The longer one conversation grows, the easier it becomes for earlier decisions to become less prominent or for unrelated information to compete for attention.

Anthropic has written about this directly in its research on long-running agents. Complex work often spans multiple context windows, so agents need mechanisms that allow later sessions to understand what earlier sessions accomplished.

My practical response has been not to force one agent to know everything. I have used different agents for different responsibilities and kept each one focused.

When a longer task continued for several days, I deliberately re-established context before continuing. That included asking the agent to review the earlier discussion and existing implementation before taking the next step. One principle emerged from this:

An agent does not need the whole project in context. It needs the right context for the task it owns.

Handoffs matter more when agents specialize

Separating agents solves one problem but creates another. How does context move between them? I handled this through explicit handoffs.

When a backend-focused agent completed functionality that the UI agent needed, I asked it to prepare a handoff containing the relevant information. The frontend agent therefore did not need to rediscover how the backend worked.

If frontend work later required something from the backend, I created the handoff in the opposite direction. This became a surprisingly effective pattern.

Context travelled with the work instead of depending on one endlessly growing conversation.

This is not very different from managing a good human engineering team. People work better when responsibility is clear and handoffs contain the information the next person actually needs. AI agents are beginning to require the same discipline.

Orchestration is not new, AI simply makes it more important

Some of the best lessons for agent orchestration come from software systems that were built before today's AI agents existed.

In one field operations platform I worked on, different sites had different rainfall thresholds. Connected rain gauges continuously sent readings into the system. When a site's configured threshold was reached, the software automatically created an inspection task for the relevant inspectors so that conditions could be checked.

The important pattern was: event → evaluate rule → trigger workflow → assign responsibility

In a healthcare platform, connected devices followed a similar pattern: reading → threshold check → alert → care staff review

And in an auction platform, a user could define the maximum price they were willing to pay. The system could automatically place another bid when someone outbid them, but only until the user's chosen limit was reached. Relevant users were notified when they were outbid.

These systems teach an important lesson for today's AI discussion. Automation needs boundaries. AI does not change that.

Not every decision needs an AI agent

There is also a temptation today to put an LLM into every workflow. I do not think that is good system design. If a rule is exact, deterministic software is often the better solution.

If rainfall exceeds a configured value, trigger an inspection. If an auction bid exceeds the user's maximum amount, stop bidding. Those decisions do not need a language model.

AI becomes valuable when the work requires interpretation, reasoning, dealing with uncertainty or understanding information that cannot easily be expressed as fixed rules. A well-designed agentic system therefore may combine:

deterministic rules + AI reasoning + human judgment

Choosing which one should make each decision is itself an orchestration problem.

More autonomy requires stronger boundaries

Long-running AI agents also create a security and control question. An agent that can read source code has one level of risk. An agent that can modify files has another. An agent that can deploy code, delete production data, issue refunds or alter customer records has a much larger potential impact.

OpenAI's current agent guidance includes approval controls specifically so sensitive actions can pause for human review. Anthropic describes the same challenge in terms of limiting an agent's potential "blast radius".

Google recently introduced anomaly detection for enterprise agents because a session can appear successful while an agent has taken an inappropriate action through one of its tools.

This suggests a simple principle:

The more autonomy an agent receives, the stronger its boundaries, permissions and review points should become.

An AI agent should not receive access simply because it is technically capable of using it.

Long-running agents also need a different kind of monitoring

There is another operational problem. An agent can technically still be running without making useful progress. That means simply checking whether the process is alive is not enough.

A better question is: What has it actually completed?

Stack Overflow for Agents describes this as distinguishing between agent "liveness" and real progress. For a long-running coding agent, useful monitoring might include:

  • Which task is active?
  • What files changed?
  • Which tests passed?
  • What is currently blocking progress?
  • When was the last meaningful result produced?

This will become increasingly important as agents work longer without constant human observation.

So what does good orchestration look like?

Based on my experience so far, I would reduce it to five practical principles.

1. Give agents bounded work

Avoid vague instructions. Define an executable outcome.

2. Remove ambiguity before implementation

Discuss the approach first. Let the agent ask questions before it starts changing the system.

3. Keep context focused

Do not expect one agent to remember everything indefinitely. Give each agent the context relevant to its responsibility.

4. Make handoffs explicit

When work moves between agents, pass the important decisions and contracts with it. Do not force the next agent to reconstruct them.

5. Parallelize only truly independent work

More agents do not automatically mean more speed. Run work in parallel only when dependencies and shared resources are understood.

The bottleneck is moving

Coding agents will continue to improve. They will work longer. They will use more tools. They will create and coordinate more subagents. And they will probably require less human involvement in many individual implementation tasks.

But that does not eliminate engineering. It changes where engineering effort is spent. The valuable questions increasingly become:

  • What should be built?
  • How should the system be divided?
  • Which work can happen independently?
  • Which context does each agent need?
  • Where are the boundaries?
  • Which actions require human approval?
  • How do we know the agents are actually making progress?
  • And who is responsible when all those pieces come together?

That is why I believe the next bottleneck in agentic software development will not simply be coding ability. It will be orchestration. When AI agents can work for days, keeping them busy is no longer the hardest problem.

Making sure they are working on the right thing, in the right order, with the right context and within the right boundaries may become the real engineering work.

Share this insight:LinkedInFacebookView Case Studies