CoreWeave has built an AI agent that does not just answer questions about a machine learning experiment. It designs the next one, launches it, and tells you what it found.

The agent, called ARIA, lives inside Weights & Biases, the experiment-tracking platform CoreWeave owns (the company describes it as “Weights & Biases by CoreWeave” in its own announcement). A researcher can ask something like whether adding dropout after the attention layers would help with overfitting, and ARIA writes the training configuration, launches the run through the platform’s W&B Launch tool on the team’s own infrastructure, and waits for it to finish.

Once the run completes, ARIA compares the result against the baseline, builds comparison charts inside the workspace, and drafts a written report. CoreWeave says that when the outcome does not settle anything, ARIA does not stop there: it recommends what to try next, whether that means testing another dropout value, running the same setup for longer, or sweeping across several settings at once. The company frames this as compressing “the gap between run finished and next run configured” from hours to minutes, a claim that reflects CoreWeave’s own telemetry rather than an independent benchmark.

The pitch is specificity. CoreWeave contrasts ARIA with an ordinary coding assistant, which the company says can technically hit the Weights & Biases API but was never built to understand it: such a tool, in CoreWeave’s telling, cannot tell a run’s label apart from its underlying ID, and it tends to pull an entire training history when a short summary would answer the question. ARIA is trained specifically on the platform’s data model, which CoreWeave says lets it answer just as fast whether a project has 20 logged runs or 20,000.

The tool also decides how to present what it finds rather than just answering in text. For a two-dimensional parameter sweep it reaches for a heat map; when several hyperparameters are interacting, it builds a parallel-coordinates plot; and when the comparison is between a handful of discrete configurations, a bar chart does the job instead. Those become live, shared dashboards that keep updating as new runs come in, rather than a one-off answer that goes stale.

CoreWeave’s roadmap points toward turning ARIA into a shared research collaborator that an entire team consults rather than a single researcher’s assistant, plus enterprise controls, custom MCP connections to internal tools, and support for outside model providers instead of a single built-in one. On the infrastructure side, CoreWeave says ARIA will eventually connect to its own Mission Control product through MCP, letting a team analyze compute usage and experiment results in the same thread.

The move matters beyond machine learning research teams. CoreWeave built its business renting Nvidia GPU capacity to AI labs, and W&B, the tool now hosting ARIA, was itself an acquisition. Wrapping an autonomous research agent around that stack is a bet that the company can climb from selling raw compute into selling the software layer that decides how that compute gets used, the same territory Databricks and Google’s Vertex AI already compete for.

CoreWeave opened ARIA to public preview on July 29, on the Weights & Biases blog, not as news announced this week; it remains in that preview today, live inside any Weights & Biases project, with no pricing or general-availability date disclosed. Teams already running experiments on the platform should treat the preview as a chance to test whether an agent-drafted hypothesis holds up against a human researcher’s judgment before handing it more of the research loop.

Reported by CoreWeave in its own blog post published July 29, 2026.