
Users now have a low-cost frontier option for coding and knowledge work with xAI’s Grok 4.7, priced at $2 per million input tokens and $6 per million output tokens. The model improves on its predecessors with a larger base and longer reinforcement learning, and it is built to verify its own output more carefully. Independent testing, however, places it well behind the leading Western models on combined intelligence and agentic coding benchmarks.
What Grok 4.7 Brings to Coding and Knowledge Work
xAI describes Grok 4.7 as its most capable model yet for coding and knowledge tasks. The company built it on a larger base model and trained it with longer reinforcement learning, a process designed to help the model check its own answers before responding.
Pricing is the headline differentiator. At $2 per million input tokens and $6 per million output tokens, Grok 4.7 sits far below typical Western frontier pricing and closer to the rates offered by Chinese model providers. That positions the model as an economical choice for developers and businesses that need capable reasoning without top-tier costs.
How Does Grok 4.7 Score on Independent Benchmarks?
The independent Artificial Analysis Intelligence Index, version 4.3.2, combines ten benchmarks into a single score. Grok 4.7 lands at 46, placing it mid-pack. By comparison, Claude Fable 5.1 and GPT-6 each lead with a score of 53.
One notable detail from the index: Grok 4.7’s two highest reasoning levels appear to perform about the same. The added compute or reasoning depth that typically separates tiers on other models does not produce a measurable difference here.
Where the Gap Widens: Agentic Coding
The difference becomes starker in agentic coding, where a model must work through multi-step software tasks rather than answer isolated questions. On Terminal-Bench 4.0, Grok 4.7 scores 26 percent.
That figure sits far behind GPT-6 Astra at 60 percent and Claude Fable 5.1 at 55 percent. Even DeepSeek V4.1 Flash, a cheaper model aimed at reducing memory needs for AI agents, edges past Grok 4.7 at 27 percent.
Agentic coding is one of the clearest measures of whether a model can reliably operate in real development workflows, where a single broken step cascades into failures. The 26 percent score indicates that Grok 4.7, despite its knowledge-work strengths, struggles to complete the kind of long-horizon coding tasks the benchmark measures.
Where Is Grok 4.7 Available?
Developers can access Grok 4.7 through several channels:
- The Grok API
- Cursor, the AI code editor
- Grok Build, xAI’s own building environment
That distribution covers both direct API usage and the editing workflows where agentic coding performance matters most.
What the Benchmark Picture Means for Buyers
The data tells a straightforward story: Grok 4.7 offers frontier-style capabilities at a fraction of the usual price, but the top of the market remains occupied by Claude Fable 5.1 and GPT-6. A business choosing a model today is weighing token cost against measured reliability on complex, multi-step tasks.
For simple knowledge work and general reasoning, the 46 index score is competitive enough for many use cases. For agentic coding, the 26 percent Terminal-Bench result is a caution flag. Teams shipping AI agents that write and modify code will want to benchmark against that gap directly.
FAQ
What is Grok 4.7?
Grok 4.7 is xAI’s latest model, described by the company as its most capable yet for coding and knowledge work. It is built on a larger base model, trained with longer reinforcement learning, and designed to better verify its own output.
How much does Grok 4.7 cost?
Grok 4.7 costs $2 per million input tokens and $6 per million output tokens. These rates are closer to Chinese model pricing than to typical Western frontier model pricing.
How does Grok 4.7 compare to Claude and GPT-6?
On the Artificial Analysis Intelligence Index v4.3.2, Grok 4.7 scores 46 overall, while Claude Fable 5.1 and GPT-6 each score 53. On Terminal-Bench 4.0 agentic coding, Grok 4.7 scores 26 percent, versus 60 percent for GPT-6 Astra and 55 percent for Claude Fable 5.1.
Related coverage
This article summarizes reporting from the-decoder.com. See our editorial disclaimer for how our articles are produced.
Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.