Google Cloud AI Research has built a system that decides which scientific ideas an AI research agent should spend its limited experiments on, then checks that each experiment tested what it claimed to. The system, called AIM (Agentic Idea Management), is described on a project page from lead author Hyeong Kyu Choi and nine co-authors at Google and the University of Wisconsin-Madison.

Most automated research tools are good at generating ideas and running them. The harder job is bookkeeping: which of hundreds of candidates deserve scarce GPU hours, and which results can be trusted. AIM treats that as its own problem. The authors say they drew on Bayesian optimization, an old statistical method for choosing what to try next, and keep two jobs apart: mapping the pool of ideas, and deciding which search branches get experimental resources and which go without.

The first part, which the authors call a surrogate, groups ideas by research direction and ranks them using scores so far, gaps in the evidence, novelty, and lessons from earlier attempts. The ranks are ordinal. They say which idea looks more promising, and the page states they are not calibrated predictions of reward. The second part picks between exploring and exploiting, at the level of whole groups and of single ideas, and sends chosen ideas to parallel solvers. Whatever comes back is used to refine, combine, repair, or add ideas.

The third part is the one worth copying. A solution auditor checks that the task was valid and that the code actually built matches the idea it was meant to test. Invalid evidence is discarded. When the implementation drifted, the idea is rewritten to describe what was really evaluated, so scores and lessons attach to the method that ran. That guards against a quiet failure in automated research: crediting an idea for results produced by something else. A fourth component, a resource planner, sets how many solver branches run per round, trading wider exploration against more rounds in which new evidence can steer the next experiment.

On results, the authors tested AIM on ten AutoLab tasks. Their comparison point is ScientistOne, which the page identifies as the strongest reported baseline by task-group average. Across three runs, AIM averaged 67.0 against 65.4 for ScientistOne on the system optimization group, a 1.6-point lead. On the model development and CUDA group it averaged 55.8 against 50.9, a gap of 4.9 points. These figures come from the paper’s own experiments, not an outside test.

The speed claim is narrower. On the Flash Attention task, AIM matched the best result ScientistOne ever posted as much as 3.1 times sooner. AIM’s own peak there, 90.5 percent, took 3.3 hours to reach. The authors stress that the 3.1x figure applies to that one task. They also concede AIM does not win everywhere: another method, AdaEvolve, leads on the Data Selection IFEval task.

The page also carries an explorer holding recordings of 27 runs spread over nine tasks, which shows the idea map, the rankings, and the actions chosen. Its own caveat is blunt: the rationale text and lessons are written by the agents, they are not independent verification, and the explorer does not rerun experiments. Inspectable here means a readable log of the system’s reasoning, not proof the reasoning was right.

The gains are modest on one group and larger on the other, over only three runs per task, and the page publishes no independent replication. The market angle: running an experiment is now the easy part of automated research, and deciding which experiment deserves the compute is where budgets get won or wasted. For any team running agent-driven experiments on a fixed GPU allowance, the audit step is the cheapest piece to borrow: it is a check on whether the code matched the idea, not a new capability.

Reported by the AIM project page (Google Cloud AI Research and the University of Wisconsin-Madison). The page carries no publication date.