xAI Unveils Grok 4.7 With a 500,000-Token Context Window and Image Input

xAI Unveils Grok 4.7 With a 500,000-Token Context Window and Image Input

Image to help understand the article

Grok 4.7 targets coding and agent tasks

xAI unveiled Grok 4.7 on Sept. 21, 2026, positioning the artificial intelligence model as an upgrade for coding and other complex tasks. The company said the model outperformed its predecessor, Grok 4.6, across several evaluations.

The model, identified in xAI’s documentation as grok-4.7, has a 500,000-token context window and can accept text and images as input, though it produces text-only responses. Users can select low, medium, high or xhigh reasoning effort, with high set as the default.

Grok 4.7 is available through Cursor and Grok Build, as well as the Grok API, third-party coding tools, model routers and cloud platforms.

Pricing starts at $2 per million input tokens

xAI set standard pricing at $2 per million input tokens and $6 per million output tokens. A Fast version, which the company says doubles output speed, costs twice as much at $4 and $12, respectively.

VentureBeat reported that Cursor charges $6 per million input tokens and $18 per million output tokens for Fast usage beyond 256,000 tokens. Artificial Analysis listed a 75% discount for cached input.

Company tests show gains, but rivals lead in some areas

In xAI’s published results, Grok 4.7 scored 46.3% on CursorBench 4.0, 64% on EEBench, 1,657 on AA Briefcase v1.1, 37.6% on Terminal-Bench 4.0 and 19.6% on Harvey Legal Agent. Grok 4.6 posted 40.4%, 53%, 1,546, 20.3% and 15.8%, respectively.

Fable 5.1 led Grok 4.7 on CursorBench, AA Briefcase and Terminal-Bench in the same company table, but trailed it on EEBench and Harvey Legal Agent.

Under the high-effort setting on DeepSWE v1.1, xAI reported a 71% score for Grok 4.7, ahead of Grok 4.6 at 65.2% and Fable 5.1 at 70%, but behind GPT-5.6 Sol at 72.7%. On HealthBench Professional, Grok 4.7 scored 56.7%, compared with 60.5% for GPT-5.6 Sol and 62.1% for Fable 5.1.

Independent testing finds slower output and mixed results

Artificial Analysis gave Grok 4.7 an Intelligence Index score of 46 using the xhigh setting, ranking it 21st among 212 models. The evaluator measured output at 39.2 tokens per second, below the comparison-group median of 71, with 0.87 seconds before the first token appeared.

The group calculated a weighted cost of $3.74 per Intelligence Index task. Citing its data, VentureBeat reported that Grok 4.7 generated 81,000 output tokens per task under xhigh, 196% more than GPT-6 Astra. Its estimated cost per task was also higher than GPT-5.6 Sol’s $1.99.

Vals AI ranked Grok 4.7 10th among 59 models with a Vals Index score of 60.22%, compared with 59.17% for Grok 4.6 and 51.53% for Grok 4.5. It estimated a cost of $11.92 per test. Grok 4.7 ranked fourth on MedScribe and Terminal-Bench 4.0 and fifth on Vibe Code Bench, while posting comparatively weaker results on ProofBench and SAGE.

Terminal-Bench results varied substantially by evaluator and test configuration. XAI reported 37.6%, while Vals AI measured 28.28%. VentureBeat cited figures of 38% and 26% under different conditions, and reported that GPT-6 Astra reached 59.6% under xhigh. The Decoder similarly reported 26% for Grok 4.7, 60% for GPT-6 Astra and 55% for Claude Fable 5.1.

GitHub brings Grok 4.7 to Copilot users

GitHub is gradually rolling out Grok 4.7 to its Copilot Pro, Pro+, Max, Business and Enterprise plans. The deployment covers Visual Studio Code, Visual Studio, Copilot CLI, GitHub’s cloud agents, JetBrains development environments, Xcode and Eclipse, giving American developers access to the model through widely used coding platforms.

Source: Original Korean article - Trendy News Korea

Comments