xAI released Grok 4.6 today, a model built to hold together across dozens of sequential steps rather than to win a single exchange. The company says its priority this cycle was endurance: research runs, full codebase work, and product builds that continue across many turns without losing the thread. Grok 4.6 follows Grok 4.5, xAI’s prior flagship release.
On xAI’s own composite benchmark, the Artificial Analysis Intelligence Index, which averages nine separate evaluations, Grok 4.6 scored 61, tying OpenAI’s GPT-5.6 Sol Max. That parity claim is narrower than it sounds. Break the same chart into its component tests and the two models split unevenly. GPT-5.6 Sol Max leads on DeepSWE 1.1, scoring 73 percent against Grok 4.6’s 65.9 percent, and on Terminal-Bench 3.0, at 34.6 percent versus 26 percent. Grok 4.6 leads on CursorBench 3.2, FrontierCode 1.1, APEX-Agents, APEX-SWE, and a legal-reasoning benchmark from Harvey. A third model in xAI’s own chart, Fable 5 Max, posted the highest index score of the four: 62.
Every number in that comparison, including the rival scores, comes from xAI’s own chart, built from what the company calls the best of competitors’ self-reported or public results. Matching one composite index is not the same as leading on the underlying skills it bundles together. xAI did not publish independent, third-party verification of any of these figures.
Behind the model, xAI describes a training pipeline leaning more on model-generated data than earlier Grok versions did. Grok 4.5 was used to regenerate supervised fine-tuning examples across reasoning, software engineering, and general knowledge tasks, with automated filtering to remove flawed traces. A follow-on reinforcement-learning stage covered agentic tasks that include kernel optimization, front-end development, and computer-aided design.
xAI also revised its safety stack for this release, describing its broadest pre-deployment evaluation effort to date, paired with post-deployment and outside testing. The company says the goal is to let Grok 4.6 operate safely in higher-stakes technical territory, including patching security vulnerabilities and supporting AI research itself, without loosening guardrails elsewhere.
Grok 4.6 is live inside Cursor, the AI-assisted code editor, and Grok Build, xAI’s own app-building tool, along with the API and third-party platforms including OpenRouter, Vercel, and Cloudflare. On xAI’s API, input tokens run $2 per million and output tokens run $6 per million, with a faster response variant available at twice those rates.
xAI is pairing the launch with a distribution push: developers on Cursor or Grok Build get double their normal usage allowance for the first week. That is not a lasting price cut. It is a short window built to pull developer attention toward Grok 4.6 at the moment a competing release could otherwise win it, and to gather fresh usage data while the model is new.
Teams evaluating coding agents should treat the Intelligence Index tie as a starting point rather than a verdict. Run the specific sub-benchmarks, like Terminal-Bench or DeepSWE, that most resemble their actual workload before shifting budget toward either model.
xAI announced Grok 4.6 on its news blog, x.ai, in August 2026.