PRODUCTION NOTE · 11 MIN READ
SW factories

Software engineering is one of the most rapidly transforming industries with the rise of AI. We already accepted that writing code by hand is now completely obsolete. Next comes the unattended era where code is shipped to production without a human reviewing the code at all.
What changed
Early 2026 the code writing tools from the top-tier AI labs got so good that most developers started to adopt them for work. This still meant that a developer was sitting next to the agent and interactively solving the issue on their laptop, but it already led to massive productivity gains, at least when it comes to writing new code.
The trouble came one step later. With so much new code being generated, the volume of pull requests to be reviewed quickly outgrew the capacity of teams and especially the senior engineers who were swarmed with thousands of lines that nobody really understood. LinearB looked at 8.1 million pull requests and found that AI-assisted ones are 2.6× larger than human ones, and that pull requests opened by agents wait 5× longer for a reviewer. [1]
The bottleneck shifted completely from writing new code to proving that it is correct and safe to ship to production. And that it’s the right thing to ship. Now that developers can produce vast amounts of code, the business alignment is more important than ever.
What it breaks
When people are not engaged with the code from the beginning, it’s harder to reason about it and judge the agent’s confident decisions, the DORA study states this explicitly "verification is a fundamentally different cognitive task than creation" [3]. And arguably less enjoyable. Combined with the rising pressure and expectations to ship more features faster this leads to alarming consequences. The Faros study found that teams merge unreviewed code 31% more often, while monthly production incidents rise 58% and bugs per developer rise 54% [2].
What it opens
There is no going back, AI adoption is not slowing down anytime soon and the current assistant powered trend is not sustainable. That is where the unattended era comes to rescue. Instead of using humans to meticulously review thousands of lines of code we can better utilise their ingenuity by having them focus on the process itself. As the Warp’s CEO, Zach Lloyd said, “software engineering will become factory engineering.” [4]
The term software factory is quickly gaining popularity as developers around the world are striving to reach the holy grail of AI written code - not having to review it. But it is much more than just about the review, with AI tools we can build powerful feedback loops. Imagine agents that are monitoring production, reporting incidents and feeding bugs back into the factory to be implemented, reviewed and perhaps even shipped automatically.
How we see it
To make this bright future possible humans need to focus hard on everything else but the code. That’s the process around it, the code quality checks and tests, tracing incidents and bugs back to agent decisions and making sure that the system improves. Another very important dimension is the alignment of the tasks and verification methods with the business needs. And most importantly measuring all of it. That’s what factory engineering is all about, it brings engineering back into what now feels like uncontrolled experimentation. Luckily we can have agents help with every step of the way.
People still have their place in a software factory and not just around the process, some decisions are still so important that they deserve an expert point of view. Humans can also save a lot of money in tokens by answering just a few questions that the agent presents to steer the architecture or scope of the new feature and prune the branches of thought that the agent would need to explore. This is called grilling and it is what leads to less declined pull requests and more code actually aligned with the software and product design.
When the process is well defined, the test suite is robust, all questions are answered and tasks are clearly defined and aligned with the business needs, there is nothing stopping you to run the agents overnight. You wake up to a high quality pull requests which are automatically reviewed by other agents according to the project guidelines and you even have a beautiful website generated with diagrams and visual presentation of the code changes and a video recording of the new feature in action. Once you trust the process enough, low-risk items and bug fixes can be merged automatically.
Levels of autonomy
Not all teams using AI work the same way. For some, AI is a better autocomplete. Others already hand complete tasks to agents and come back to finished work. This is the difference between AI adoption and AI autonomy. Adoption tells you how much AI a team uses. Autonomy tells you how much work AI can complete without a human. A team can have high adoption and low autonomy if people still review every change.
We think about autonomy as how much work you can give to an agent and trust it to finish without you. Or more practically: what work can I give to the factory and trust it to complete without me?
A normal change goes through something like this:
Issue → Code → Tests → PR → Merge → Deploy → Monitor → Recover
The autonomy level tells us how far through this loop the factory can go before it needs a human.
| Level | What the factory owns | What the human does | Example |
|---|---|---|---|
| L0 - Manual | Nothing | Does the work | Engineer implements an issue |
| L1 - Assisted | Parts of Code → Tests | Drives the work | AI writes code or helps debug |
| L2 - Task | Code → Tests | Gives the task, steers the agent and reviews the result | Agent implements a defined task with human supervision |
| L3 - Change | Issue → Code → Tests → PR | Approves the result | Agent takes a ready issue and returns a validated PR unattended |
| L4 - Merge | Issue → Code → Tests → PR → Merge | Handles work the factory cannot verify | Proven low-risk changes merge automatically |
| L5 - Production | Issue → Code → Tests → PR → Merge → Deploy → Monitor → Recover | Handles exceptions the factory cannot recover from | Factory ships, monitors and rolls back a change |
| L6 - Factory | Issue → Code → Tests → PR → Merge → Deploy → Monitor → Recover → Improve | Governs the factory | Factory uses failures to improve future verification. |
There is no single autonomy level for a company. We might be L5 for documentation, L4 for simple bug fixes and L2 for a risky database migration. The company decides which types of work can reach which level and can set a maximum level for sensitive changes.
Moving up means making the next part of the loop safe to run without a human. Most of the work is not in the agent itself, but in the process around it: the environment, tests, checks, permissions, monitoring and recovery.
| Move | What needs to change |
|---|---|
| L0 → L1 | Bring AI into the normal development workflow and give it enough project context to be useful. |
| L1 → L2 | Make tasks clear enough for an agent to own. Give it the codebase, tools and environment it needs to implement and test the work. |
| L2 → L3 | Make unattended work possible. The agent needs a sandbox where it can run the full project, reliable tests and checks, clear acceptance criteria and a way to stop and ask when something is missing. |
| L3 → L4 | Make verification strong enough to replace human approval. Define which changes qualify, what evidence they must produce and which checks must pass before they can merge automatically. |
| L4 → L5 | Extend verification into production. Add deployment checks, telemetry, failure detection, safe permissions and recovery so the factory can tell whether a change actually worked and recover when it did not. |
| L5 → L6 | Close the feedback loop. Trace production failures back to the work that caused them and use what escaped to improve the checks and process for future work. |
Autonomy works in both directions. If the factory can no longer handle a step reliably, the human gate moves back and the level goes down.
A company might also have multiple factories. Similar projects can share one, while projects with very different designs and processes might need their own.
Today we are roughly at L3, with some types of work already reaching L4 and L5. The next section shows what that looks like today.
What we did
We looked at off the shelf factories and found none that would fit our workflows and tools we use without locking us in to their system and opinions and often at a price-point.
That’s why we build DAM (Deploy Agents Massively), an open-source platform for running agents in the cloud with enterprise guarantees. We focus on the infrastructure and building blocks to help humans make their factories without the headache of managing servers and automation pipelines. DAM gives us the freedom to assemble the factory in a way that’s compatible with how we work, with no vendor lock-in and our own opinionated process that AI agents help to enforce and a cost control.
The platform does the hardest things for us:
- a reliable sandbox where the agents never see the credentials to external services,
- artifacts that allow us to create visual outputs for us simple humans to digest complex topics
- and the Slack integration combined with periodic tasks make the agents ambient and omnipresent.
It brings many of the pieces you’d need to build a factory tailored to your needs and it’s constantly evolving, but the best example is how we use it to build and monitor itself.
Code reviews
We stopped doing human code reviews, instead we are building a suite of static checks and tests that are systematically preventing some critical problems like backwards incompatibility of an API schema. On top of that we have a “code-guardian” agent running in DAM which carefully judges the code against a set of hand-designed rules we’re constantly improving. One of the less obvious one is an “issue-fit”, which checks that the code actually implements the task and satisfies all acceptance criteria. This agent will also ping us on slack if some pull request is running stale.
Incident investigation
We stopped investigating incidents. Instead we are gathering telemetry and logs and we have an agent to investigate. We call it “buginator” and it is actively watching the deployment state, CI pipeline health and for some of our non-production deployments we gave it direct access to the kubernetes cluster to poke around. It is so powerful that we can reliably trace down any out of memory crash, LLM provider issues, failed release or a failing CI pipeline. Asking buginator “what went wrong” on slack is now the default way to trace down an issue and “file an issue” is typically the next message.
User support and issue creation
Another area agents are really good at is answering questions of our users and filing issues for missing features and bug reports. We gave our support agent full context of the codebase, existing issues and known bugs so it is well informed to answer most questions on Slack and our users love it! If someone finds a bug in the platform, the agent transforms the conversation into a detailed issue with steps to reproduce and often even the concrete pieces of code that are to blame.
What next
We believe that quality software can be built autonomously overnight without burning too much money. For this we need to be able to spin the a full software stack inside the sandboxed agent without loosing the security guarantees. This requires running docker and kubernetes just like you’d run it on a developer laptop and brings a set of challenges we are excited to solve.
Next thing we are already building is cost monitoring and caps, which we need to prevent runaway bills in case an agent decides to rewrite the entire codebase due to a poorly defined feature. Once we can sleep at night it gives us the opportunity to build benchmarking across harnesses and models potentially including open-weight ones. With all this flexibility, it’s not only cost to quality ratio, but also sovereignty, and speed. An agent on slack needs to answer faster than an agent doing a PR review. Furthermore, by having a central platform, we can give each agent access to it’s own telemetry and cost metrics so it can self-optimize.
We think the whole factory should end up as a file in the repository formalising the agreed processes. The agents than help teams apply them in the flow of work. Sometimes that means answering a question such as “what should I work on next?” using the team’s priorities. Sometimes it means prompting for a missing review or owner. For stricter steps, the system can require an approval before work moves forward. The agents and by extension the processes and entire factory can be self-improving but it is important that all modifications to the workflows are scrutinised like any other change. You don’t want your agents to behave differently just because your colleague told them on slack to write in Chinese from now on.
If there is anything common in science, software and engineering it’s that we’re standing in the shoulders of giants. It should be possible to collectively learn, improve as an industry and learn from the work of others. That’s the reason we’re also building an open-source catalog of agents teams start from, that already do a job well and tailor them to their needs and contribute their findings back to the ecosystem.
References
- LinearB, 2026 Software Engineering Benchmarks Report: The AI Productivity Edition, May 2026. 8.1M pull requests, 4,800 engineering teams. linearb.io
- Faros, AI Engineering Report 2026: The Acceleration Whiplash, March 2026. 22,000 developers, 4,000 teams. faros.ai
- DORA, Balancing AI Tensions, March 2026. Survey of 1,110 Google engineers. dora.dev
- Zach Lloyd, “AI Engineer World’s Fair: Daily Dispatch (Loops)”, Latent Space. latent.space