Weekly Articles

Agent-First Work Allocation: An Experiment for Project Teams

· 5 min read

Agent-First Work Allocation: An Experiment for Project Teams

Most project teams approach AI the same way: they look at who’s already doing what, then ask where AI could help. Agent-first work allocation reverses that order, starting not from your existing team but from the recurring, rule-based work on a project stream. You assign that work to AI agents by default and bring people in specifically where judgement is required. It’s a small change in sequencing, but running it changes what you notice: it’s the tasks that break, rather than the ones that go smoothly, that show you where human judgement is needed, and what kind.

Where the Idea Comes From

The idea came from an unlikely source: Cofounder 2, a platform built for solopreneurs running an entire company with AI agents. It gives you an org chart, engineering agents, sales agents, marketing agents, all reporting to a single human who makes the calls that matter. The product itself is a niche tool for solo founders, but the structure behind it is worth borrowing for a project team: an org chart built agent-first, with people slotted in only where judgement is required.

Most of us have never drawn that kind of chart for our own projects. We’re used to starting from the team we already have and asking what tasks AI could take off their plate, a framing that tends to produce small, incremental answers because it protects the team’s existing shape rather than questioning it. Start from a blank page instead: map your project’s recurring work, assign it to agents first, and see what’s left over for the people on your team. Flip the starting point like this and the picture changes: some tasks you assumed only a person could do turn out to be handled cleanly by an agent, while other tasks that looked purely mechanical turn out to need more judgement than anyone realised.

How to Apply It to a Project Stream

Putting this into practice doesn’t mean restructuring your whole programme overnight, which would be a fast way to create chaos if anything went wrong. It works better as a contained experiment on a single project stream, where the blast radius of a mistake stays small and you can watch closely what happens.

Start by listing the recurring, rule-based work on that stream: things like dependency tracking, meeting prep, risk register updates and stakeholder briefing drafts, the tasks that come around every week regardless of what else is happening on the project. Assign each of these to a dedicated agent by default, before asking whether a person should be doing it instead. Then, rather than judging the experiment after the fact, track what happens as you go: what the agent produced, whether a judgement call came up, and who ended up owning the task once it did.

A Worked Example

A short worked example makes this concrete. On a 12-week platform migration stream, two recurring tasks were assigned to an agent by default from week four:

Task: Dependency tracking Assigned to: Agent, by default What happened: Flagged three at-risk dependencies correctly, two days earlier than the weekly stand-up would have Judgement call surfaced: No, clean automation Kept with: Agent

Task: Stakeholder briefing drafts Assigned to: Agent, by default What happened: First-pass drafts were strong once the agent had the stakeholder’s history; edits shrank from a rewrite to a tone check within two weeks Judgement call surfaced: Yes, early on, but it turned out to be habit rather than necessity Kept with: Agent, with PM review

Try It on Your Own Stream

Once you’ve seen how the mapping works, you can run the same exercise against your own stream using this template:

Task: [dependency tracking / meeting prep / risk register updates / stakeholder briefing drafts / your own recurring item] Assigned to: Agent, by default What happened: [what the agent produced or caught] Judgement call surfaced: [yes or no, and what kind if yes] Kept with: [agent, person, or agent plus review]

Every row where a judgement call surfaces on your own map is raw material for an Agent Governance Charter’s escalation boundaries, which formalise, after the fact, what your map already found. Run the map first and let the charter follow, rather than trying to write escalation rules before you’ve seen what needs escalating.

What Breaks Tells You More Than What Works

Once you’re running the experiment, pay closer attention to what breaks than to what goes smoothly. When an agent handles something well, that’s reassuring but not particularly informative, since these are usually tasks you’d already have guessed AI could manage. The more useful information comes from the failures, because they show you where human judgement was needed, and what kind of judgement it was.

In practice, the results run in both directions. Some tasks you assumed needed your judgement turn out not to, once the agent has enough context to work with; the human step in that process was habit rather than necessity. Other tasks that looked purely routine throw up an edge case that exposes complexity and judgement calls nobody realised they were making.

Where This Leaves Us, For Now

None of this is a case for restructuring your team around an agent org chart by next sprint. It’s a case for running a smaller, contained version of the same idea: one stream, a handful of recurring tasks, agents assigned by default, and close attention paid to what breaks along the way.

What breaks is, in the end, the real output of the exercise. It shows you, more clearly than any role description could, where your judgement is doing something an agent can’t, and where it’s been doing something an agent could have done all along, simply because that’s how the work has always been divided up.

If you ran this experiment on one of your current project streams, which task do you think would surprise you most, by working better than expected, or by exposing something you didn’t realise was a judgement call?

Frequently Asked Questions

What does “agent-first work allocation” mean for a project team? It means mapping a project stream’s recurring, rule-based tasks and assigning them to AI agents by default, then involving people specifically where judgement is required, inverting the usual approach of assigning to people first and considering AI second.

Is agent-first work allocation safe to apply across a whole programme? Not as a starting point. It works better as a contained experiment on a single project stream first, where you can observe what breaks and what that reveals before considering wider application.

What’s the real value of running this kind of experiment? The value is in what fails, not what succeeds. Failures reveal specifically where human judgement is needed and what kind, which often produces a more accurate picture of where a PM’s time creates value than a standard task audit.


If you ran this experiment on one of your current project streams, which task do you think would surprise you most, by working better than expected, or by exposing something you didn’t realise was a judgement call?

Yes, AI helped me to write this :)