This Grok 4.6 review starts with the short version: SpaceXAI (formerly xAI) released Grok 4.6 on August 12, 2026, holding headline pricing flat at $2 per million input tokens and $6 per million output tokens while claiming it matches OpenAI’s GPT-5.6 Sol on a composite benchmark. That claim is true on exactly one measure, xAI’s own ten-row evaluation table shows Claude models winning more categories, and “flat pricing” hides a 200,000-token cliff that reprices an entire request rather than just the overage. Here’s what that actually means for anyone deciding whether to switch.

What Actually Changed

Two corporate details sit behind this release that most model-number coverage skips. xAI merged into SpaceX in an all-stock deal that closed on February 2, 2026, creating a combined entity reported at roughly $1.25 trillion in valuation, and the combined company rebranded to SpaceXAI on July 6, 2026 β€” the Grok product name didn’t change, only the parent company’s. On the model itself, xAI describes a longer supplemental training pass than Grok 4.5 used, built on curated model-generated reasoning data, engineering data, an improved optimizer, and a supervised fine-tuning stage where Grok 4.5 itself regenerated the training trajectories. No parameter count, mixture-of-experts disclosure, or system card was published β€” competing reports guess anywhere from the same 1.5-trillion-parameter base as Grok 4.5 to a newer 2-trillion-parameter model, and neither is confirmed by xAI. The most concrete coding number xAI has published is a DeepSWE 1.1 jump from 54% to 65.9%.

The “Matches GPT-5.6 Sol” Claim, Read Carefully

Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, a nine-benchmark composite, tying GPT-5.6 Sol and trailing Claude Fable 5 and Claude Opus 5 by one and two points. That headline is accurate for what it measures β€” but xAI’s own published ten-row evaluation table, covering individual benchmarks rather than the blended composite, shows Claude models winning the most rows, with Grok 4.6 strongest on knowledge work and legal reasoning and weakest on terminal use. There’s a second layer worth flagging: xAI itself notes that these are vendor-reported launch results using the best of self-reported or publicly available third-party scores, and that harnesses, reasoning settings and tool access may differ between runs. This site has already covered why that caveat matters more than it sounds β€” controlled testing has found that harness choice alone can swing coding results by as much as a full model generation, and the same confound showed up in Qwen3.8-Max-0902’s benchmark comparison against Claude Opus 5, where the evaluation ran through the competing vendor’s own harness. Treat “matches GPT-5.6 Sol” as true for one composite score under conditions xAI controlled, not as a settled verdict across real workloads.

What “Flat Pricing” Doesn’t Cover

The $2/$6 headline rate really didn’t change from Grok 4.5, and multiple independent trackers confirm it. What sits underneath that headline is less flat.

What you might assumeWhat actually happens
A long prompt only pays extra for the tokens past 200,000Once a request crosses the 200K-token threshold, xAI bills every token in that entire request at the doubled long-context rate ($4/$12), not just the excess
Batch discounts apply the same way across the Grok lineupGrok 4.6 is excluded from xAI’s 20% batch discount entirely; only the older grok-4.3 and grok-4.20 family qualify β€” a notably smaller discount than the 50% batch rate Anthropic and OpenAI offer on their own flagships
Cached input costs stayed flat like everything elseCached input actually rose on 4.6, to $0.50 per million tokens versus $0.30 on 4.5 β€” the one line item that got more expensive, not less

The Detail Nobody Headlines: Smaller Context Than Its Own Predecessor

Grok 4.6 isn’t the widest context window xAI sells, and it’s easy to assume the newest model has the largest window. It carries 500,000 tokens at the standard rate; grok-4.3 and the entire grok-4.20 family carry a full 1 million tokens. For a workload built around one enormous document in a single call, the older, cheaper model is both roomier and less expensive β€” a genuinely counterintuitive result worth checking before assuming the newest release is the better fit by default.

A Reported Reliability Flag

At least one independent review reports that Grok 4.6 shows a meaningfully higher confident-fabrication rate than its benchmark parity with GPT-5.6 Sol would suggest. That’s a single-source finding, not confirmed independently here, but it’s the kind of gap between “scores the same on a composite index” and “behaves the same in practice” that a benchmark table alone won’t surface β€” worth testing directly on tasks where a confidently wrong answer is costly before trusting the model at the same level its index score implies.

Where It’s Actually Positioned

Grok 4.6 inherits Grok 4.5’s core positioning as a model co-trained with Cursor specifically β€” IDE patterns like tab completion, multi-file edits and project-aware context are part of its base training, not added afterward through fine-tuning. If Cursor integration was the reason to pick Grok 4.5, that reason carries over unchanged. Beyond coding, its stronger categories are knowledge work and legal reasoning; terminal-heavy agentic work is its weakest published category, which lines up with where Claude and GPT-5.6 Sol both hold an edge.

Who Should (and Shouldn’t) Switch

Grok 4.6 makes sense for teams already inside the Cursor and xAI ecosystem doing knowledge-work or legal-reasoning-heavy tasks, where the flat headline pricing genuinely undercuts GPT-5.6 Sol and comparable frontier models β€” as long as prompts stay under the 200,000-token cliff and the workload doesn’t depend on batch discounts. It’s a weaker fit for terminal-heavy agentic work, for single-call jobs against very large documents where the older grok-4.3’s larger context window is actually the cheaper and roomier option, and for any use case where a wrong-but-confident answer is costly, given the single reported fabrication-rate flag above. Model your actual token pattern against the 200K cliff and current batch-discount rules before committing β€” the same “read past the headline number” habit this site recommends when weighing free versus paid AI tools generally, and the same open-disclosure question raised when covering GLM-5.3’s undisclosed parameter count and license terms.

Frequently Asked Questions

Does Grok 4.6 really match GPT-5.6 Sol?

On the Artificial Analysis Intelligence Index composite score, yes, both land at 61. On xAI’s own individual ten-row benchmark table, Claude models win the most categories, so the “matches” claim holds for one blended metric, not across the board.

Did Grok 4.6’s pricing actually stay flat?

The headline per-token rate did, at $2/$6 per million tokens. But cached input rose from $0.30 to $0.50 per million tokens, batch discounts were dropped entirely for this model, and prompts over 200,000 tokens get billed at double the rate for the whole request, not just the portion over the threshold.

Does Grok 4.6 have a bigger context window than Grok 4.5?

No β€” it carries 500,000 tokens, smaller than the 1-million-token window on the older grok-4.3 and grok-4.20 family, which remain the better choice for single very large documents.

What company makes Grok 4.6?

SpaceXAI, formerly known as xAI. xAI merged into SpaceX in an all-stock deal that closed February 2, 2026, and the combined company rebranded to SpaceXAI on July 6, 2026. The Grok product name didn’t change.

How many parameters does Grok 4.6 have?

xAI has not disclosed a parameter count, architecture details, or a system card for Grok 4.6. Reports outside the official announcement disagree, with some describing the same 1.5-trillion-parameter base as Grok 4.5 and others describing a newer 2-trillion-parameter model.

Is Grok 4.6 reliable for factual accuracy?

At least one independent review reports a meaningfully higher confident-fabrication rate than its benchmark score parity with GPT-5.6 Sol would suggest. This is a single-source finding worth testing directly on your own tasks rather than assuming benchmark parity implies equal reliability.