xAI Launches Grok 4.7 With More Parameters but Trails Rivals on Benchmarks
What happened
xAI, the artificial intelligence company backed by Elon Musk, released its newest model, Grok 4.7, on Monday, September 21, 2026. The company called it "a notable improvement over Grok 4.6 at the same price and speed." The launch came after at least five apparent delays stretching back to late July.
Key points
- Grok 4.7 has 2.1 trillion parameters, a 40% increase from the 1.5 trillion in Grok 4.6.
- It costs $2 per million input tokens and $6 per million output tokens, the same pricing as Grok 4.6.
- The model is live immediately in the Grok app, Cursor, Grok Build, and the xAI API, with no waitlist.
- Musk posted on X that Grok 4.7 represents "a strong combination of intelligence, speed & low cost."
What xAI says about Grok 4.7
xAI says Grok 4.7 spends more time working through difficult problems and double-checks its own answers more often than Grok 4.6 did. The company also claims it has deployed its strongest safety guardrails yet. Beyond that, xAI incorporated supplemental training data from SpaceX, Musk's rocket company, including Starlink satellite telemetry, manufacturing records, and engineering failure logs. The goal, according to xAI, is a model that reasons better about hardware and physical systems than models trained purely on internet text.
How it ranks on benchmarks
Despite the improvements, benchmark scores place Grok 4.7 in second place behind leading rivals. On GDPval, which measures model performance on real-world knowledge tasks like legal memos and spreadsheets, Grok 4.7 scored 1,695. Claude Fable 5.1 led with 1,735. On AA-Briefcase, a test of multi-hour office work built by Artificial Analysis, Grok 4.7 also ranked second. On EEBench, an electrical-engineering benchmark, Grok 4.7 trailed GPT-6 Astra.
What is confirmed
Grok 4.7 launched on September 21, 2026. It runs on 2.1 trillion parameters. Pricing is $2 per million input tokens and $6 per million output tokens. The model is available in the Grok app, Cursor, Grok Build, and the xAI API. Benchmark results place it behind Claude Fable 5.1 on GDPval and AA-Briefcase, and behind GPT-6 Astra on EEBench. xAI trained the model partly on SpaceX data.
What is still unclear
The article does not provide full benchmark details for AA-Briefcase, as the source text was cut off before completing that section. The exact reasons for the multiple delays between late July and the September 21 launch are not officially explained beyond Musk's shifting public timeline statements. It is also unclear how the SpaceX training data specifically affects real-world performance beyond xAI's claims.
Why it matters
Grok 4.7 shows that xAI continues to push its models toward greater scale and capability, but the benchmark results suggest it still trails competitors like Anthropic's Claude and OpenAI's GPT-6 Astra on widely used tests. For users, the immediate availability without a waitlist and unchanged pricing means access to a larger model at no extra cost, even if it does not top the leaderboards.