Grok 4.7 keeps the same standard token pricing as Grok 4.6. SpaceXAI positions that as a key part of the upgrade, arguing that users get improved performance without a higher base price.
The biggest changes come from the model’s training and architecture. SpaceXAI says Grok 4.7 was built on a larger base model and went through a longer reinforcement-learning process using a more difficult mix of tasks, with greater emphasis on work that can take hours to complete. The company says that training improved the model’s ability to verify its own output and manage extended context.
SpaceXAI also trained Grok 4.7 to understand the Grok Bot harness natively. Grok Bot is designed around persistent agents that can use software tools and continue working for long periods, and the company says the added training improves Grok 4.7 on conversational work and general knowledge tasks.
The company is also emphasizing document and presentation creation. On professional-work benchmarks, SpaceXAI says Grok 4.7 improved over Grok 4.6 while remaining competitive with other frontier models.
On CursorBench 4.0, which measures longer-running coding tasks, Grok 4.7 scored 46.3%, compared with 40.4% for Grok 4.6, 41.7% for GPT-5.6 Sol and 51.8% for Fable 5.1.
The model posted a 71.0% score on DeepSWE v1.1 at high effort, versus 65.2% for Grok 4.6, 72.7% for GPT-5.6 Sol and 70.0% for Fable 5.1. On Terminal-Bench 4.0, Grok 4.7 reached 38.0%, compared with 20.3% for Grok 4.6, 37.3% for GPT-5.6 Sol and 57.9% for Fable 5.1.
Results varied across other professional evaluations. Grok 4.7 scored 1,657 on AA Briefcase v1.1, compared with 1,546 for Grok 4.6, 1,487 for GPT-5.6 Sol and 1,678 for Fable 5.1. It reached 64.0% on EEBench, 19.6% on the Harvey Legal Agent Benchmark and 56.7% on HealthBench Professional.
A separate GDPval comparison gave Grok 4.7 an Elo score of 1,695. Fable 5.1 scored 1,735, Grok 4.6 reached 1,605 and GPT-6 Astra was listed at 1,542.
Safety is another focus of the release. SpaceXAI says Grok 4.7 uses an entirely new safeguard stack and is the strongest model it has tested for jailbreak resistance and refusal behavior.
In biological safety testing, the company says Grok 4.7 reached 62.4% on LatchBio’s biosafety benchmark. On HackerBench v0.3, SpaceXAI’s test for risky and malicious cybersecurity prompts, the model allowed 3.3% of risky dual-use requests through while rarely rejecting legitimate security work.
SpaceXAI has also begun giving selected cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defensive research.
The release follows Grok 4.6, which SpaceXAI introduced in August with an emphasis on long-running agents and more ambitious interactive and visual tasks. With Grok 4.7, the company is pushing further toward extended coding sessions, professional knowledge work and agent-based workflows while holding standard pricing at $2 per million input tokens and $6 per million output tokens.
This analysis is based on reporting from x.ai.
Image courtesy of SpaceXAI.
This article was generated with AI assistance and reviewed for accuracy and quality.