PRODUCTION NOTE · 7 MIN READ
The research loop is going AI-native

Many researchers already use AI to summarize papers, explore ideas, and write code. Now, systems are beginning to take on more of the experimental process itself. Projects such as Andrej Karpathy's autoresearch and Google's AlphaEvolve offer early examples: AI systems that can propose changes, test them, and use the results to guide another round of experiments with limited human intervention.
This points toward a more AI-native research process, where AI participates in the ongoing cycle of investigation. A researcher could set a direction, leave an agent exploring possible solutions overnight, and return to findings that shape the next day's work. The prospect is compelling because each experiment can inform what the system tries next, allowing an investigation to advance without a person directing every step.
Practical use remains limited, and deciding which questions matter or what a result means still calls for human judgment. The immediate challenge is turning these early experimental loops into a dependable way of working. What would researchers need to confidently delegate more of an investigation?
From reading papers to running experiments
AI-native research spans a range of ways to integrate AI into the research process, from assisting with individual tasks to potentially directing entire investigations. The starting point is often augmentation: researchers use AI to summarize literature, write experiment code, or keep records while directing each step themselves. As teams delegate more of the process, AI can connect these tasks into cycles of proposing experiments, running them, evaluating results, and choosing what to try next. A possible long-term goal is fully AI-directed research, where systems formulate questions, pursue investigations, and improve their own research methods through experience. This progression is gradual.
As researchers begin delegating more of the process, their role shifts from executing every step to steering the work. Domain expertise, human relationships, and collaboration remain central to that role. With less time spent on routine execution, researchers can devote more attention to deciding what to investigate, interpreting the results, and working with others to move the research forward.
One practical example is the use of evolutionary algorithms alongside AI agents to explore possible solutions. Agents can generate and test variations, compare their performance, and use the most promising results to guide the next round of experiments.
The design-build-test-learn (DBTL) cycle provides a useful structure for this process. An agent can help design an experiment, build the proposed solution, test it, and use the findings to inform the next iteration. In principle, an agent could carry out the entire cycle, while researchers set its direction and decide how to act on the results.
Progress toward greater autonomy depends on how reliably AI can carry out experiments, evaluate the results, and use those findings to decide what to try next. Where those decisions still require human judgment, researchers remain involved. Different research activities can therefore support different degrees of delegation.
AI-native research in practice: optimizing FPGA designs
A concrete example is optimizing hardware designs for field-programmable gate arrays (FPGAs), chips whose circuitry can be configured after manufacturing. High-level synthesis (HLS) converts code written in a high-level programming language, typically C or C++, into a hardware description in a language such as Verilog or VHDL. This is a step toward translating software code into a hardware configuration for the FPGA.
An FPGA has a fixed amount of logic, memory, and arithmetic resources, making hardware design an optimization problem. Performing more operations in parallel can make a design faster, but can also consume more resources. The search aims to approximate a Pareto frontier: a set of designs where improving one objective, such as performance, requires a trade-off in another, such as resource usage. Researchers can then choose a design that meets their performance requirements and fits within the chip's limits while preserving the required behavior.
This optimization problem fits naturally into the design-build-test-learn (DBTL) cycle. Each iteration starts with a hypothesis about how changing the C or C++ code or synthesis directives could improve the design. The modified code is synthesized into a hardware description, simulations check its behavior against the expected results, and synthesis reports provide estimates of performance and resource usage. Candidates offering promising trade-offs can be retained as starting points for further iterations.
AI can test different hypotheses and use the results to refine its proposals. It can also consult existing studies to identify optimization techniques and formulate new hypotheses to test.
We build the infrastructure. Researchers run the experiments.
As researchers put AI into their design-build-test-learn loops, we have observed a recurring set of practical challenges. Experiments need to keep running for long periods, agents need controlled access to tools and data, and repeated iterations can become expensive. Researchers also need to understand what happened during a run and retain the freedom to choose the models and agent tools that suit their work. Handling these concerns adds work around the investigation itself.
Consider the FPGA researcher above. They have a baseline design, synthesis tools, correctness checks, and a workflow for generating and evaluating candidates. Their goal is to reduce latency while preserving behavior and staying within the chip's resource limits. To leave that workflow running overnight, they also need confidence that the environment will remain available, the agent will stay within its permitted access, and the investigation will stay within budget.
The tools to support that researcher need to make these controls straightforward: a secure sandbox with access to the required files and tools, a persistent environment that retains experiment state and results, and enforceable limits on individual attempts and overall spending. These let the researcher step away from a long synthesis run and return to review its progress, deciding whether the findings justify extending the budget or changing direction.
Andrej Karpathy's autoresearch provides a working example of an autonomous experimental loop in model training. An agent modifies training code, evaluates the result, and keeps or discards the change before trying again. Each training run has a fixed five-minute budget, excluding startup and compilation, making candidates comparable on the same hardware. It demonstrates the practical value of bounded experiments. Researchers running repeated experiments also need a way to limit the total cost of the investigation, including agent and tool costs.
During long-running investigations, researchers need visibility into what the agent is doing, where experiments are failing, and when their judgment is needed. In the FPGA example, repeated synthesis failures should be visible with enough context for the researcher to diagnose the problem and steer the next attempts. Telemetry also supports smart checks at cost gates: repeated failures or spending without progress can trigger a pause or human review before more resources are committed. This gives researchers a way to intervene while the investigation is running and helps the harness enforce cost controls informed by the run's behavior.
Researchers also need room to change the model or agent tools used within their DBTL workflow. Testing whether a different model produces more useful candidates at an acceptable cost should not require rebuilding the surrounding execution environment and controls. Model interoperability makes that choice part of the investigation.
These common needs guide how we are trying to support researchers. We are building a platform and execution harness that brings these tools together, so researchers can run their own DBTL workflows with less infrastructure work. They define the scientific questions, experimental methods, and evaluation criteria; our focus is helping them run that work securely, persistently, and within the limits they choose.
What comes next: teams of researchers and AI agents
Much of the current effort focuses on enabling AI to carry out the mechanical work of research while people provide direction and judgment. A further challenge is extending this approach to teams: creating an environment where researchers and AI agents can coordinate their work, share findings, and pursue a common research goal with increasing autonomy.
AI-native research often takes place within isolated workflows. Supporting collaboration means making experiments, results, and the reasoning behind decisions visible across the team. Researchers and agents may focus on different parts of an investigation, but each needs enough context to build on others' findings and formulate the next hypothesis.
That shared view needs to preserve the controls that make individual workflows reliable. Secure access, persistent experiment state, cost limits, and useful telemetry remain essential as work spans multiple researchers and agents. The challenge is to make the broader investigation transparent and coordinated while allowing each contributor to work independently within agreed boundaries.
An open question is how much progress will come from more capable models and how much will depend on harnesses designed around collaboration. Understanding how these two approaches work together will help determine how AI-native research grows from individual experimental loops into coordinated team investigations.