← All Production Notes

PRODUCTION NOTE · 11 MIN READ

From faster agents to productive teams

25.9.2026 ApocoAI productivityTeams

Agents in our Slack review pull requests, investigate incidents and turn bug reports into issues. Our software-factory note describes how the team building Deploy Agents Massively (DAM), an open-source agent platform, got there. [1]

The working day did not get much lighter. More work is in motion. Answers arrive faster and still need checking. People return to an agent session and have to recover what they were trying to do. Nearly finished work waits while newer messages pull attention elsewhere.

The work did not vanish when agents took it on. It moved.

Where the work went

Verification. A generated answer is often almost right, and “almost” is where the work is. Someone has to check it against the source, fix the details and say what is still uncertain. Skip that, and the work lands on the next person. METR’s early-2025 study followed experienced developers on open-source projects they had contributed to for years. With AI tools, they took 19% longer, while believing they had been 20% faster. Feeling faster is not the same as being faster. The study is a warning, not an estimate of today’s tools. [2]

Attention. Each parallel agent can return with findings or a decision to make. Boris Cherny, who leads Claude Code at Anthropic, says running five agents at once comes down to “how good I am at context switching”. [3] For a lead juggling projects, choosing where to put attention becomes part of the work.

Noise. In one deployment, we let an agent comment in a shared channel without being mentioned, a setting we call ambient mode. It commented so often that the team switched the behaviour off. Telling it to be quieter did not help. We expect better models and tuning to fix this, but we have not shown it yet. Until then, an agent that speaks too often spends the team’s attention.

Follow-through. An agent can flag work nobody has picked up. The flag is not the result. Someone still has to take responsibility, change the plan or decide to drop the work.

Adoption, autonomy, productivity

The software-factory note separates adoption, how much AI a team uses, from autonomy, how much work AI can finish without a human. Neither says whether that work was worth doing, or whether the team is better off for it. Productivity covers both: how much worthwhile work the team gets done for the full effort it takes.

Full effort means time, attention and cost. It includes the work above (verification, attention, noise, follow-through) plus rework and whatever lands on the person who receives the result. A gain counts only if it still holds once all of that is counted.

We built an AI-native collective productivity framework around this idea. [4] It asks where AI improves a team’s work, what more the team could achieve and what to change next. Here, we apply its six levels to our experience building and running DAM.

Choose one team, department or company and keep that scope constant. The examples illustrate what can make a level possible, not a required technology stack.

LevelWhat gets betterWhat it looks like
0No useful AI-enabled gain yet.AI is not used in this work, or examining its use shows no useful gain after full costs.
1Parts of the work get faster or better, and the gain survives review.An assistant reads project sources, follows shared instructions and runs an analysis. Individual tasks improve; people still carry the work through handoffs and follow-up.
2Complete, worthwhile results arrive sooner or better, or become practical at all, at justified total cost.An agent carries a task through to agreed use: a validated change, a report used in a decision, an accepted handover. People sign off on consequential choices; routine steps run within an agreed scope.
3The same team accomplishes more valuable or demanding work together, after counting costs to everyone affected.People and shared agents combine knowledge, challenge assumptions and develop solutions. Decisions and reasoning stay available for colleagues to build on, including across projects.
4AI helps the organization keep improving how it works.AI helps people investigate recurring problems, test changes and develop skills. Comparisons across repeated work show worthwhile gains after the cost of learning and change. Failed experiments stay in the record.
5Strategy and delivery improve each other.AI helps changed priorities reach affected work and delivery evidence inform strategic choices. That two-way connection repeatedly improves choices and results. People set direction and decide what to pursue or stop.

The ambient agent had shared access and permission to speak. It still made the team’s day worse. Tools can enable a gain; they do not prove one.

Every level above 0 needs evidence and retains the useful gains below it. If AI’s effect is unknown, leave the level open: “not enough evidence yet.” A company assessment has to cover its important work and dependencies, including where it is stronger and where it has gaps. One successful team cannot establish the company’s level. Here we focus on team work.

Moving up: what it takes

Choose the move that addresses the constraint you actually have. It need not be the next number up.

From 1 to 2: finish the job, including the checking

In our DAM work, the support agent turns bug reports into issues with steps to reproduce. [1] An accepted issue may finish a triage task; it does not finish a bug fix. Agree what result someone needs to use before claiming the whole job got better.

The engineering is mostly around the agent. It needs access to the code, the conversation where the request arrives and the tool where the result belongs. It also needs sources, tests and acceptance criteria to check against. Our code-review agent’s “issue-fit” check asks whether the change does what the issue requested. [1] DAM enforces approvals outside the agent’s sandbox. [5] The approval request must still explain what will change and why, so the person approving can judge it.

Count checking alongside generation. Compare the time and effort to reach a usable result; faster drafting can leave the next person with more work.

From 2 to 3: connect the work, and protect people’s attention

Shared agents let colleagues question findings and build on one another’s work. In a deployment we support, designers and developers discuss a proposed change with a code-aware agent in Slack. They examine its findings and shape an issue together; the agent files it after approval. In our Friday learning channel, agents track presenters, receive transcripts and ask for updates. We are still tuning what they add and how often they speak.

For a lead running several projects, each agent’s response can mean another context switch. We want agents to explain what changed, which decision unblocks others and what can wait. A short brief should let the lead return without reconstructing the project.

DAM’s schedule Precheck can skip a run when a configured check finds nothing to do, before calling a model. [5] That handles predictable scheduled work; deciding whether to interrupt a live discussion remains harder.

Agents also need current priorities and an agreed remit: which records they may update, which actions they may finish, and which decisions need a person. Asking at every step consumes attention; acting on an obsolete plan creates rework.

The test is whether shared results improve once you count the effort of rebuilding context, relaying updates and chasing follow-up. If triage and context switching outweigh the better results, shared access has not made the team more productive.

From 3 to 4: make improvement repeatable

DAM records agent traces, logs and spend. [6] These help explain failures and costs. To establish improvement, we need records of the team’s work too.

Measuring does not wait for level 4; the levels below need evidence too. Level 4 adds repeated improvement through learning. Consider a possible loop: an agent groups returned pull requests by reason, people confirm that missing tests recur, and they change the test guidance. A few weeks later, they compare similar work for review effort, repeat returns and defects reaching production. Fewer returns alone could mean weaker review.

The engineering needs to connect a process change to later outcomes, preserve failed experiments and let people compare one cycle with the next. Track attention too: how long decisions wait and how often people rebuild context before acting.

Framing tasks, spotting weak answers and knowing when to stop an agent matter at every level. Here, look for evidence that people handle later work better, including when the agent fails. Improvement should hold across repeated work, allowing for unsuccessful cycles.

From 4 to 5: connect strategy and delivery

At level 4, the team improves how it works. Level 5 connects that to what it works on. When direction changes, agents find affected commitments, explain the implications and carry agreed changes into plans and tasks. When delivery challenges an assumption, agents bring the evidence to the people who set direction, early enough to change a choice. People decide what to pursue, change or stop.

Engineering this means linking priorities, decisions and delivery records, with authority to update the affected work after a decision. An agent reliably pursuing an obsolete priority shows why execution alone is insufficient.

Start with one live choice, such as taking on a more demanding kind of project. Agents gather delivery evidence, constraints and capability gaps. Run a bounded test, agreeing in advance what would justify continuing, changing or stopping. One useful decision does not establish level 5: the two-way connection has to improve choices and results across the work being assessed.

Where we are

The DAM team’s working self-assessment is level 3 on this productivity framework: AI-enabled collaboration increases what the team can accomplish together. Shared agents review code, investigate incidents and answer users in Slack; the code-review agent also chases stale pull requests. [1] By our own rule, those agents make the collaboration possible but do not prove the gain. That is why this is the team’s assessment of its working experience, not a measured productivity rating.

The team does not yet claim level 4. It has not yet shown that experience keeps improving its methods, skills and later results. The next step is to compare recurring work across cycles: quality, time to a usable result, repeat problems and the full effort of checking and follow-through. DAM’s agent telemetry helps investigate causes; the work itself has to show the improvement.

The software-factory note also places the team at roughly L3, on its autonomy scale. The matching number is a coincidence: that scale asks how far agents get without a human; this one asks what the team gets out of it.

DAM and Brainio

DAM is open source and supplies the execution, access controls, schedules and shared agents described above. The platform alone does not maintain what the team decided, why it decided it or what changed afterwards.

In our own product team, turning a discussion into plans, issues and follow-up still runs through one person. That is the constraint we want to address next.

We are building Brainio, a shared platform for people and teachable agents to work from team knowledge, decisions and methods. The design connects familiar places such as Slack to existing task and knowledge tools. It gives teams control over what agents can access, what they can do and where information is processed.

Meeting transcripts and summaries run today. The connected flow below does not run yet. Apoco is the first test team.

After a discussion, agents would record decisions and reasons, carry agreed actions into work tools, inform absent colleagues and follow up. A colleague could examine the reasoning with an agent and bring back an objection or better idea. Agents would use team priorities to explain which decision needs attention and what can wait.

Teams would also be able to inspect and correct the agents’ methods. If a handover omits why another approach was rejected, an agent could propose a change to its handover method for the team to adopt. Recording later handovers would help test whether the correction made work better.

Our test is fewer repeat problems, less context reconstruction and less dependence on one person, after counting checking and upkeep. Brainio earns no level by being installed; the work has to improve.

What to do with this

Pick one team and follow a recurring request through to a result someone uses. Count checking, chasing, upkeep and recipient effort. One case can identify a useful change; it cannot establish a whole-team level.

Record the current outcome and full effort. Agree what better would mean: a usable result sooner, less rework, harder work within the same resources, or a colleague needing less help. Try one change across comparable cases and check that quality holds. Keep, revise or stop it according to what happens. Progress can happen within a level.

If you lead several projects, start with your own week: where did your attention go, and which interruptions helped the work?

If your agents got faster and your team’s day did not get lighter, bring us one recurring situation. We would like to hear how you handle it today.

Notes and references

[1] Apoco, SW factories, production note, 25 September 2026. The code-review, incident and support agents, and the autonomy levels.

[2] METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. Randomized study of 16 experienced developers working on 246 issues. Its February 2026 update reports that selection effects and difficulties measuring parallel work make the later experiment an unreliable estimate of current speedup.

[3] Gergely Orosz, Building Claude Code with Boris Cherny, The Pragmatic Engineer, 4 March 2026. Cherny’s account of running parallel Claude Code sessions and switching between them.

[4] Apoco, AI-native collective productivity. The framework behind the levels in this note.

[5] Deploy Agents Massively, public repository, with documentation for schedules and Precheck and security and approval controls. The deployment and channel examples are our experience of using the platform, not behaviour every installation provides automatically.

[6] Apoco, Sovereign AI is the right to change your mind, production note, 25 September 2026. The account of making traces and spend visible to agent owners.