The company is also positioning Grok 4.6 as a stronger coding and agent model. On the Artificial Analysis Intelligence Index, it scores 61, matching GPT-5.6 Sol Max and improving from 56 for Grok 4.5 High. Fable 5 Max scores 62 on the same composite benchmark.

The gains are more pronounced on several individual evaluations. Grok 4.6 reaches 69.9% on CursorBench v3.2, compared with 66.7% for Grok 4.5 High. On DeepSWE v1.1, it improves from 54% to 65.9%, while FrontierCode v1.1 Extended rises from 56.6% to 61.3%.
Agent-oriented tests show a similar pattern. Grok 4.6 scores 57.5% on APEX-Agents, up from 47.1% for its predecessor, and reaches 56.4% on APEX-SWE, compared with 53.6% for Grok 4.5 High. On Terminal-Bench v3.0, the new model improves to 26% from 15.7%, although GPT-5.6 Sol Max and Fable 5 Max score higher on that evaluation. SpaceXAI says Grok 4.6 also performs strongly on longer-form professional work. It posts a score of 1,577 on AA-Briefcase, above GPT-5.6 Sol Max at 1,502 and narrowly ahead of Fable 5 Max at 1,574. On Harvey LAB, Grok 4.6 reaches 15.8%, compared with 12.9% for Grok 4.5 High.
The company cautions that third-party comparison figures use the best self-reported or publicly available results, so the benchmark table is not a single controlled evaluation across every model. The results more clearly establish a substantial step up from Grok 4.5 than an across-the-board lead over competing frontier systems.
SpaceXAI attributes the improvement to a longer supplemental training run that used curated model-generated reasoning data, technical material and engineering data, along with changes to the optimizer and overall training process. Grok 4.5 was then used to regenerate supervised fine-tuning trajectories spanning reasoning, agent environments, STEM, software engineering and knowledge work.
Reinforcement learning also targeted a broader set of agentic tasks. The company says those environments included general coding, knowledge work, kernel optimization, web development and computer-aided design. That training appears aimed at making the model more effective once a task extends beyond a single prompt. SpaceXAI says Grok 4.6 showed more self-testing and verification during longer runs, including checking its own work before proceeding. It also produced stronger initial results on visual and interactive projects than Grok 4.5 in the company's testing.
SpaceXAI describes one of Grok 4.6's strengths as taking a broad product idea and turning it into an early working application. The model can research an unfamiliar topic, organize the project, build the main interactions and continue refining the result after feedback.
The model's pricing is another major part of the launch. Standard API access begins at $2 per million input tokens and $6 per million output tokens. A faster version is available at twice those rates. Long-context use carries higher costs. Grok 4.6 supports a 500,000-token context window, but prompts at or above 200,000 tokens are priced at $4 per million input tokens and $12 per million output tokens. Cached-input pricing also rises from $0.50 to $1 per million tokens at that threshold.
Artificial Analysis reported a cost of roughly $0.84 per task for Grok 4.6 in its testing. It also found that the model completed AA-Briefcase workloads in about 53 turns and around 0.5 billion input tokens on average, compared with roughly 103 turns and 2 billion input tokens for Claude Opus 5 Max. Those results depend on the benchmark setup and do not guarantee equivalent efficiency in production systems.
The model supports text and image inputs with text output, along with function calling, structured outputs and reasoning. SpaceXAI is also offering twice the included Grok 4.6 usage in Cursor and Grok Build during the first week.
Safety work was expanded alongside the model's capabilities. SpaceXAI says Grok 4.6 underwent its broadest pre-deployment evaluation effort so far, followed by additional post-deployment and third-party testing. The company says its safeguards are intended to support legitimate uses including vulnerability remediation, engineering design and AI research.
Grok 4.6 does not lead every benchmark shown in SpaceXAI's own comparison. GPT-5.6 Sol Max remains ahead on DeepSWE and Terminal-Bench, while Fable 5 Max leads CursorBench, FrontierCode and several other evaluations. The more consistent story is the improvement over Grok 4.5 across coding, agentic and knowledge-work tasks.
With Grok 4.6, SpaceXAI is putting more weight on sustained task execution rather than isolated responses. Its combination of stronger agent benchmarks, expanded training for multi-step environments and relatively low starting API prices makes the model's ability to complete longer workflows a central part of the release.
This analysis is based on reporting from xAI.
Images courtesy of SpaceXAI.
This article was generated with AI assistance and reviewed for accuracy and quality.