xAI released Grok 4.7 on Monday, calling it the company’s strongest model yet for coding and knowledge work, and held pricing at Grok 4.6’s level: two dollars per million tokens in, six dollars per million out. That pricing decision matters more than the launch itself. Holding the price line while claiming a capability jump is a bet that xAI can absorb higher compute costs per query rather than pass them to developers, at a moment when rivals such as GPT-5.6 Sol charge double xAI’s input rate and Fable 5.1 charges five times as much.
According to xAI’s own release, Grok 4.7 sits on a larger base model than Grok 4.6, trained through a longer reinforcement learning run weighted toward tasks that take many hours rather than quick single-turn answers. xAI says the result is a model that verifies its answers more thoroughly and handles longer context windows. The company also built in native support for Grok Bot, its coding and conversation harness, during training.
The benchmark table xAI published tells a split story, and all of the numbers in it are xAI’s own. Grok 4.7 leads on CursorBench 4.0 (46.3 percent versus 40.4 for Grok 4.6), on EEBench (64.0 percent), and on the Harvey Legal Agent Benchmark (19.6 percent). But Fable 5.1 beats it on AA Briefcase v1.1 (1,678 versus 1,657), crushes it on Terminal-Bench 4.0 (57.9 percent versus 37.6), and edges it on HealthBench Professional (62.1 versus 56.7). On the GDPval professional-work leaderboard, Fable 5.1 also sits well ahead in Elo terms, at 1,735 against Grok 4.7’s 1,695. xAI is not claiming an outright win here. It is claiming competitiveness at a lower price, and the release does not include any independent evaluation to check that framing against.
Safety is the section where xAI makes its boldest unqualified claim: that Grok 4.7 is the most resistant model the company has tested against refusals and jailbreak attempts. The company says the model tops LatchBio’s biosafety benchmark at 62.4 percent and blocks all but 3.3 percent of risky prompts on HackerBench v0.3, a cybersecurity-misuse test, while rarely refusing legitimate security work. xAI has also opened the model’s red-team capabilities to a limited set of cybersecurity partners on an invite-only basis, though it has not disclosed how many partners are involved or what governs the access.
Grok 4.7 is live now inside Cursor and Grok Build. Developers can also reach it through the Grok API, and it is rolling out to the third-party coding tools, routers, and cloud platforms that already carry Grok models. A faster variant runs at twice the output speed for twice the price, a straightforward latency-for-cost tradeoff rather than a capability change.
The real signal here is what xAI chose not to contest. Rather than claim a clean sweep, it published a scorecard where a competitor visibly wins on three of seven benchmarks, betting that flat pricing and coding-specific strength matter more to developers than topping every leaderboard. Teams picking a coding model on price alone should still run their own workload against Terminal-Bench-style multi-hour tasks before switching, since that is the category where the gap to Fable 5.1 is largest.
Per xAI’s own launch announcement for Grok 4.7, which carries no publication date.