News

Releases & Models

CONFIRMED: xAI launches Grok 4.7 for long-running tasks at $2 per million input tokens

xAI announced Grok 4.7 on September 21, with a larger base model, longer reinforcement training, and a focus on coding and knowledge work. Independent evaluations published soon after place the model mid-pack on the Artificial Analysis Intelligence Index, behind Claude Fable 5.1 and GPT-6 on agentic coding benchmarks such as Terminal-Bench.

[{"h":"CONFIRMED: xAI finally puts Grok 4.7 on the market","p":"xAI announced Grok 4.7 on September 21, ending a cycle of promises, missed dates, and infrastructure clues that had kept the model in the realm of speculation. The release presents a system aimed at coding and knowledge work, with a larger base than Grok 4.6, a longer reinforcement-learning run, and more attention to tasks that can take hours. The central change is not simply a new version. It is xAI's attempt to turn persistence and verification into a commercial advantage."},{"h":"What changes for people using AI at work","p":"According to xAI, Grok 4.7 was trained to sustain long contexts, check its own work more carefully, and handle problems that require multiple steps. The company also says it improves document and presentation creation, along with professional tasks measured by benchmarks such as GDPval and AA Briefcase. That matters more to teams that need to leave a conversation with a usable artifact than to users seeking only quick answers."},{"h":"Aggressive pricing, but do not confuse the headline with total cost","p":"xAI lists a starting price of $2 per million input tokens and $6 per million output tokens, plus a fast variant at twice the speed for twice the price. That is competitive for text-heavy agents, but real cost still depends on context size, retries, tool calls, and output volume during verification."},{"h":"The counterpoint now has numbers","p":"The open question at launch, whether the persistence and verification claims would translate into real performance, now has a partial answer. On the Artificial Analysis Intelligence Index, which combines ten benchmarks, Grok 4.7 scored 46 and landed mid pack, behind Claude Fable 5.1 and GPT-6, which tied at 53. The gap widens in agentic coding: on Terminal-Bench 4.0, an independent measurement put Grok 4.7 at just 26 percent, versus 60 percent for GPT-6 Astra and 55 percent for Claude Fable 5.1. Even the cheaper DeepSeek V4.1 Flash edged past it on that test. xAI's own published benchmarks, like CursorBench 4.0 and AA Briefcase, do show real gains over Grok 4.6 and competitive results on office tasks, so this is not a story of overall failure, but of a model that competes well on price and office work while losing ground precisely on the front xAI chose to highlight: long-running agentic coding."},{"h":"Our read","p":"Grok 4.7 confirms xAI's aggressive pricing thesis, but the launch's central promise, working longer without giving up early on coding tasks, does not yet hold up in the independent numbers published so far. For small and midsize Brazilian companies evaluating coding agents, Grok 4.7 may be worth it for cost on office and document tasks, but it still trails Claude and GPT-6 on terminal and real software engineering benchmarks. Still worth testing, with expectations calibrated by price, not technical leadership."},{"h":"Sources","p":"xAI, official Grok 4.7 announcement: https://x.ai/news/grok-4-7\\nThe Decoder, independent Artificial Analysis Intelligence Index benchmark: https://the-decoder.com/xai-launches-grok-4-7-at-bargain-prices-but-benchmarks-reveal-a-wide-gap-to-claude-and-gpt-6/\\nxAI, public model documentation: https://docs.x.ai/docs/models"}]
WA