Grok 4.5

RawGraph

Grok 4.5 is a proprietary multimodal large language model and reasoning model in the Grok family. It was developed by SpaceXAI in collaboration with Cursor and released through the xAI API on July 8, 2026. The model accepts text and images, produces text, and has a 500,000-token context window. Its launch emphasized software engineering, tool-using agents, and professional knowledge work rather than a general consumer-chat release.[1][3][4]

Two things set Grok 4.5 apart from the rest of the mid-2026 frontier cohort. The first is how it was trained: Cursor states that trillions of tokens of real developer session data from its editor went into the model, an arrangement that arrived alongside SpaceX's agreement to buy Cursor's parent company.[7][11] The second is its price-to-capability position: at $2 per million input tokens and $6 per million output tokens it sits well below the flagship tiers it is benchmarked against, and both SpaceXAI and independent evaluators report that it reaches comparable results using far fewer output tokens.[2][5][13]

Naming and corporate identity

Coverage of this model uses "xAI" and "SpaceXAI" interchangeably, and both are correct. SpaceX announced the acquisition of xAI in early February 2026, an all-stock deal that Elon Musk described as creating a vertically integrated company spanning AI, rockets, satellite internet, and X.[20] The SpaceX-xAI merger folded the AI business into the launch company, and SpaceXAI became the public-facing brand for the model line.

The Grok 4.5 model card settles the legal question directly. A footnote on its first page states that "SpaceXAI is a doing-business-as (dba) name of XAI LLC," and that the two names "may be used interchangeably throughout this card."[6] The underlying corporate entity is therefore still XAI LLC; SpaceXAI is the trade name. This article uses SpaceXAI for the company as it presents itself in 2026, and xAI where the historical entity or the still-current x.ai domain and API are meant.

The Cursor relationship has its own corporate layer. Cursor is the product; Anysphere is the company that makes it. SpaceX said on June 16, 2026 that it would buy Anysphere in an all-stock deal worth $60 billion, and Reuters reported that arrangement as the backdrop to the Grok 4.5 launch three weeks later.[9][11] Cursor's own site carried an "Anysphere, Inc." copyright notice at the time of the model's release and still carried it in late July, with the deal not yet closed; Cursor blog posts were published under the Cursor Team byline rather than a SpaceXAI one.[7]

Release and positioning

SpaceXAI's release notes place the initial API release on July 8, 2026. Cursor published a joint-release post the same day, and Reuters reported that access began immediately through the API, Grok Build, and Cursor.[1][7][9] The current SpaceXAI announcement page displays July 16 and describes the model as launching that day, but the official release notes separately document both the July 8 API launch and July 17 availability for users of the EU API console.[1][2] The announcement page itself carried the date "Jul 8, 2026" when it was archived on the afternoon of the launch, so the July 16 stamp was applied later, alongside edits that added an SWE-Marathon chart and removed the launch-day note about the European Union.[22] July 8 is therefore the initial release date.

The launch landed in a crowded week. TechCrunch noted that OpenAI was scheduled to release GPT-5.6 the following day, and Meta Superintelligence Labs introduced Muse Spark 1.1 on July 9.[14][21] Grok 4.5 was therefore measured against brand-new competitors almost immediately.

Grok 4.5 was positioned as a model for coding, agentic tasks, and knowledge work. Cursor made it available in its desktop, web, iOS, command-line, and software-development-kit products at launch. SpaceXAI offered it through Grok Build and its developer platform, while the model card also listed Microsoft Office add-ins and third-party API gateways. The model card said availability on consumer platforms, including the Grok website, mobile applications, and X, was planned for later.[3][6][7]

Cursor framed the model as a departure from its previous strategy. Where Composer 2.5 had been trained as a coding specialist, Cursor described Grok 4.5 as "the first we've built for more than software engineering," with a deliberately broader data mixture drawing on STEM tasks, research papers, and other knowledge work. Cursor also said the two models occupy different weight classes and that Composer 2.5 would continue to be offered alongside it.[7]

The Cursor training arrangement

The most distinctive claim about Grok 4.5 is that it learned from what developers actually did inside a working code editor, not only from static repositories. It is also the claim that carries the most unresolved questions, so the documented facts and the contested ones are separated below.

How the partnership came about

The relationship predates the acquisition. On April 21, 2026, Cursor announced a partnership with SpaceX to accelerate model training, saying its own efforts had been "bottlenecked by compute" and that it would use SpaceXAI's Colossus infrastructure to scale up.[10] Two months later, on June 16, SpaceX announced it would acquire Anysphere outright for $60 billion in stock.[9] Grok 4.5 shipped on July 8, roughly three weeks after that announcement and while the acquisition was still pending. Cursor's launch post describes the model as one "we trained jointly with SpaceXAI," language that reflects a co-development partnership rather than a straightforward vendor relationship.[7]

What Cursor and SpaceXAI say about the data

Cursor's description is the more expansive of the two. Its launch post states that training "included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools," and that "this dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments."[7] Cursor separately says it used reinforcement learning on difficult problems in realistic environments, built at scale by a distributed system of agents that construct, test, and refine each training environment.[7]

SpaceXAI's own documentation is narrower. The model card carries a single footnote on the point: "Grok 4.5 was also subject to supplemental training using anonymized Cursor workflow data to improve coding and agentic performance."[6] Three qualifiers there matter. The data is described as anonymized; it is described as workflow data rather than raw source code; and it is described as supplemental training rather than part of the base pretraining mixture, which the model card lists separately as publicly available data, internally generated data, and other data for which SpaceXAI "has secured the necessary rights."[6] That framing constrains how far the Cursor-derived gains would be expected to generalize beyond agentic coding.

Neither the SpaceXAI announcement nor the Cursor launch post discusses user consent, opt-out mechanisms, or privacy controls for the training data.[2][7] The relevant policy lives elsewhere, on Cursor's data-use page, which was last updated July 15, 2026, a week after the model shipped.[12]

That page describes a binary setting. With Privacy Mode enabled, "Customer Data will not be used for training by Cursor," and Cursor says it maintains zero-data-retention agreements with all model providers so that those providers will not store or train on the data either. With Privacy Mode turned off, Cursor says it "may use and store codebase data, prompts, editor actions, code snippets, and other code data and actions to improve our AI features and train our models."[12] The categories in that second sentence line up closely with what Cursor says went into Grok 4.5.

Two limits on the public record should be stated plainly. First, the consent mechanism is a settings toggle rather than a specific, affirmative opt-in for this model or this partner, and neither company has published the proportion of users who had Privacy Mode enabled during the collection period. Second, no primary source published by either company states how the anonymization described in the model card was performed or verified. Commentary raising data-provenance and GDPR questions about a coding tool's session data being used to train a model owned by the same parent company circulated widely after the acquisition was announced, but that commentary appeared largely on aggregator and analysis blogs rather than in reported coverage, and several such posts described a change to Privacy Mode's training behavior that Cursor's own policy page does not support. Those specific claims are not repeated here.

Disclosed benchmark contamination

The clearest documented problem with the training arrangement came from Cursor itself. A footnote to the launch post states that "Grok 4.5 has an advantage on CursorBench because an earlier snapshot of the Cursor codebase was accidentally included in training," that "the exact impact is unclear," and that the data "has been removed for future models."[7] Cursor excluded CursorBench from its release comparison as a result.

The disclosure is narrow: it concerns one internal benchmark, and it does not invalidate the unrelated evaluations in the model card. It does, however, illustrate the structural risk in a training pipeline that ingests a company's own development activity, since the same repositories that generate training data can also generate the evaluations. Cursor said it was separately working on a larger update to CursorBench.[7]

Model and training

SpaceXAI describes Grok 4.5 as a proprietary text model with image understanding. It accepts text and JPEG or PNG images, but its documented output modality is text rather than generated images. The API's knowledge cutoff is listed as February 1, 2026, while the model card gives a pretraining cutoff of January 2026; the model catalog says it has no real-time knowledge unless a search tool is enabled.[3][4][6]

Cursor identifies the architecture as a mixture-of-experts model.[7] SpaceXAI's own documentation does not state a parameter count. Artificial Analysis reported the model at 1.5 trillion parameters, roughly three times the size of Grok 4.3, attributing the figure to a disclosure by Elon Musk rather than to independent verification or to a formal SpaceXAI publication.[13] The figure should be read on that basis: it is a founder's public statement, not a documented specification.

The model card says its training data included public data, internally generated material, and other licensed or rights-secured sources. Post-training used supervised fine-tuning and reinforcement learning with human and synthetic reward signals. It also received the supplemental training on anonymized Cursor workflow data described above.[6]

SpaceXAI reported that training ran across tens of thousands of Nvidia GB300 graphics processors. Its announcement describes data deduplication, quality scoring, and domain-focused selection, followed by reinforcement learning on hundreds of thousands of tasks. These tasks centered on multistep software engineering and technical work, using automated and model-based grading and long-running asynchronous agent rollouts, with the stack built so that "agentic rollouts can run for many hours while learning continues across tens of thousands of GPUs."[2] SpaceXAI has not published enough architectural detail to reproduce the model.

PropertyDocumented value
DevelopersSpaceXAI and Cursor
Corporate entityXAI LLC, doing business as SpaceXAI
Initial releaseJuly 8, 2026
Model identifiersgrok-4.5, grok-4.5-latest
ArchitectureMixture of experts (Cursor); parameter count not formally published
Input and outputText and image input; text output
Context window500,000 tokens
Knowledge cutoffFebruary 1, 2026 (API docs); January 2026 pretraining cutoff (model card)
Reasoning controlLow, medium, or high effort; high by default; cannot be disabled
Serving speed80 tokens per second (developer-reported)
Access at launchxAI API, Grok Build, and Cursor
LicenseProprietary

Capabilities and API

The model works with both the xAI Responses API and Chat Completions API. It supports client-defined function calling and SpaceXAI's server-side web search, X search, and code-execution tools. These interfaces allow an application to combine the model's generated text with external information or executable actions, but they do not make the underlying model's stored knowledge current.[3][4]

Reasoning effort is configurable as low, medium, or high, with high used by default; SpaceXAI's documentation states that reasoning cannot be disabled on this model.[19] This setting lets developers trade latency and token use against additional inference-time reasoning. Published benchmark figures for Grok 4.5 are almost always reported at the high setting, and the model card labels them accordingly, which matters when comparing against competitors evaluated at their own maximum settings.[6][19]

Grok 4.5 also supports image inputs for tasks such as interpreting screenshots or visual documents, while professional integrations described by SpaceXAI include drafting and editing work in Word, PowerPoint, Excel, and Outlook. The model card says Grok 4.5 is the default model in the Word, PowerPoint, and Excel add-ins, while the launch post credits Grok Build, the terminal agent that also defaults to Grok 4.5, with building multi-sheet Excel models that incorporate web research.[2][6]

The model's principal advertised capability is long-horizon software engineering. SpaceXAI and Cursor describe it as able to inspect codebases, edit multiple files, use terminals, run tests, and iterate on failures. Those operations depend on an agent host that supplies the relevant tools and permissions. The distinction matters: the model generates decisions and tool calls, while Grok Build, Cursor, or another application controls execution.[2][7]

Access and pricing

The API model ID is grok-4.5; grok-4.5-latest tracks the current version. Standard pricing for requests with fewer than 200,000 input tokens is $2 per million input tokens, $0.30 per million cached input tokens, and $6 per million output tokens. When a request contains 200,000 or more input tokens, the respective prices double to $4, $0.60, and $12.[4][5]

Cursor prices the model separately inside its own product. Its launch post lists a base model at $2 per million input tokens and $6 per million output tokens, matching the API, plus a faster variant at $4 and $18. Cursor included the model in individual and team subscription plans and doubled the included usage for the first week, and SpaceXAI offered free Grok 4.5 usage for a limited period in Grok Build and Cursor.[2][7]

Server-side tool use is billed separately from tokens. SpaceXAI lists web search, X search, and code execution at $5 per 1,000 calls, collections search at $2.50 per 1,000, and file-attachment search at $10 per 1,000. Consequently, the advertised $2 and $6 token rates do not represent the full cost of every agentic workload.[5]

SpaceXAI also distributed the model through several gateways, including OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic. Availability was geographically staggered. The launch post carried a note that Grok 4.5 was "not yet available in the EU in any SpaceXAI products or the API console" and that EU availability was "expected in mid-July"; the company then recorded EU API-console access on July 17, nine days after the initial API release.[1][3][22]

Evaluation

SpaceXAI published results on a wide range of benchmarks in its model card and launch post. Except where a third party is named, these are developer-reported figures produced with different benchmark harnesses, so they are not interchangeable measures of one general capability. The model card states that competitor figures are drawn from those developers' published system cards or benchmark leaderboards, which means the comparisons are assembled rather than run head to head under identical conditions. Cursor announced the model card's publication on July 14, 2026, six days after the model shipped, and SpaceXAI revised it on July 20.[2][6][17]

Coding and software engineering

Grok 4.5 is evaluated at high reasoning effort throughout; competitor settings are given where the model card specifies them.

EvaluationGrok 4.5Best listed competitorWhat it measures
DeepSWE v1.062.0%Claude Fable 5 66.1%Repository issue resolution, run by Artificial Analysis in each provider's harness
DeepSWE v1.153.0%Claude Fable 5 70.0%Same suite under the mini-swe-agent harness
APEX-SWE51.2%Claude Fable 5 54.8%Integration and observability work, run by Mercor
SWE-Bench Pro64.7%Claude Fable 5 80.4%Hard multi-file issues from maintained repositories
SWE-Bench Multilingual78.0%Claude Opus 4.8 84.4%Issue resolution across programming languages
SWE-Marathon29.0%Claude Opus 4.8 26.0%Ultra-long-horizon engineering tasks
FrontierSWE (dominance)78%Claude Fable 5 89%Win rate on expert-scope technical challenges
ProgramBench57.2%GPT-5.5 60.2%Rebuilding program behavior from a compiled binary
Terminal-Bench 2.183.3%Claude Fable 5 84.3%Agent performance in containerized terminal environments
SWE-Atlas-QnA84.0%tied with GPT-5.6 Sol 84.0%Answering questions about a codebase using exploration tools

The pattern is consistent and worth stating plainly: Grok 4.5 is competitive across this suite but leads it only rarely. The launch post as it now stands displays five coding charts, and another model scores higher on four of them; SWE-Marathon is the exception, where Grok 4.5 placed first at 29.0%. That chart was not on the page at launch. The version archived on July 8 carried four charts, DeepSWE v1.0, DeepSWE v1.1, Terminal-Bench 2.1 and SWE-Bench Pro, and Claude Fable 5 led all four.[2][22] The model card presents that result as the strongest evidence for the model's long-horizon claim, describing it as "exceptionally capable in solving long-horizon engineering tasks, beating all other frontier models tested."[6] Results also depend on tool configuration, inference budget, task selection, and harness design.

The model card additionally reports an internal suite called FalseClaimBench, which checks whether an agent reports having done work it never performed. Grok 4.5 scored 84.0% fully true claims against 59.0% for Claude Opus 4.8 and 38.0% for GLM-5.2. This is a SpaceXAI-designed evaluation with no external replication, but it measures a failure mode, agents claiming completed edits that do not exist, that has practical consequences in long agentic sessions.[6]

Knowledge work, tool use, and engineering domains

EvaluationGrok 4.5Context
GDPval-AA v21535Fifth of eight listed; Claude Fable 5 led at 1760, GPT-5.6 Sol 1743, Claude Opus 4.8 1600, Grok 4.3 1085
Tau-3 Banking32.6%Second of eight listed, behind GPT-5.6 Sol at 33.0%
3DCodeBench49.8%Highest listed, ahead of Claude Opus 4.8 at 47.0%
CAD-Bench88.1%Third, within 0.4 points of Claude Fable 5 at 88.5%
RelBench38.4%Third of three, behind GPT-5.6 Sol 41.8% and Claude Opus 4.8 40.7%
DeepSearchQA38.4%Behind Claude Opus 4.8 at 40.7%, on a SpaceXAI implementation of the benchmark

The jump against the previous generation is large on the knowledge-work measures. On GDPval-AA v2 the model card puts Grok 4.3 at 1085 and Grok 4.5 at 1535, a gain of 450 points on Artificial Analysis' harness.[6]

Token efficiency

Token efficiency is the claim SpaceXAI leaned on hardest, and it is unusually concrete. The launch post reports that Grok 4.5 resolves SWE-Bench Pro tasks with 15,954 output tokens on average against 67,020 for Claude Opus 4.8 at maximum effort, a 4.2-fold difference, and states more generally that the model achieves "roughly 2x the token efficiency of comparable leading models, solving tasks in under half the number of steps."[2] SpaceXAI says the model is served at 80 tokens per second.[2]

Artificial Analysis independently corroborated the direction of that claim. It measured Grok 4.5 consuming 1.9 million tokens on its Coding Agent Index against 6.2 million for GPT-5.5, and put the cost at $2.49 per coding task against $5.07 for GPT-5.5, and $0.31 per Intelligence Index task, which it described as five times cheaper than Claude Sonnet 5 at maximum effort while scoring higher.[13] Because output tokens are billed at three times the input rate, efficiency compounds with the low headline price rather than merely adding to it.

Independent evaluation

Artificial Analysis provides the most-cited independent measurement. On its Intelligence Index v4.1, Grok 4.5 at high reasoning effort scores 54. At launch that placed it fourth overall, behind only Claude Fable 5, GPT-5.5, and Claude Opus 4.8, and above every open-weights model and every Gemini model then measured; Artificial Analysis characterized the release as bringing SpaceXAI to the intelligence frontier, noting a 16-point gain over Grok 4.3's score of 38.[13][8]

That placement has since moved, which is the expected behavior of a versioned index in a fast-moving field rather than a change in the model. As accessed on July 27, 2026, Artificial Analysis' leaderboard showed a dozen model configurations at or above Grok 4.5's score of 54, including several effort variants of Claude Opus 5 at 56 to 61, GPT-5.6 Sol at 56 to 59, Kimi K3 at 57, and Claude Opus 4.8 at maximum effort at 56.[16] The model's own page reported a rank of 13 out of 190 tracked models on the same date, with a median output speed of 54.8 tokens per second, a time to first answer token of 8.31 seconds, and a blended price of $1.35 per million tokens.[8] These are dated observations from one provider and change as the evaluator revises its suite, adds models, or as serving configurations shift. The serving measurement has drifted a long way in particular: the same page recorded 89.5 tokens per second when it was archived on July 9 and 61.0 on July 26, so SpaceXAI's advertised 80 tokens per second describes a moving quantity rather than a steady state.[23]

Artificial Analysis also recorded a regression that the vendor material does not surface. On its AA-Omniscience benchmark, Grok 4.5's factual accuracy improved from Grok 4.3's 35% to 52%, but its hallucination rate more than doubled over the same comparison, from 25% to 54%. Artificial Analysis describes this as a common pattern for larger models, which answer more questions rather than declining them.[13] That finding sits in direct tension with the model card's own factuality section, which reports a single-turn hallucination rate of 0.98% for Grok 4.5 against 1.14% for GPT-5.5 and 3.35% for Claude Opus 4.8.[6] The two measurements are not directly comparable, since they use different task sets, different grading, and different definitions of an unsupported claim, but the gap between a sub-1% vendor figure and a 54% independent one is large enough that neither number should be quoted without its methodology.

On its Coding Agent Index, Artificial Analysis placed Grok 4.5 in the Grok Build harness third with a score of 76, on par with GPT-5.5 in Codex and below Claude Fable 5 in Claude Code, at substantially lower cost.[13]

A second independent evaluation came from Snorkel AI, which tested Grok 4.5 against GPT-5.5 and Claude Opus 4.8 on GDPval+, a set of roughly 2,000 professional workplace tasks, using the open-source Harbor framework. Grok 4.5 recorded the strongest overall result at a 29% mean pass rate against 22% for GPT-5.5 and 21% for Claude Opus 4.8, with the gains concentrated in education, quality-assurance analysis, legal, and healthcare tasks. Snorkel flagged an important caveat: at SpaceXAI's request Grok 4.5 ran in its own Grok Build harness while the other two used Snorkel's Stirrup agent, which limits direct comparability. Snorkel also noted that all three models pass fewer than a third of expert criteria, so "expert-level work remains an open frontier."[15]

The "Opus-class" characterisation

Elon Musk's summary of the model became its most repeated description, and it is worth attributing precisely rather than restating as fact. Musk wrote that Grok 4.5 "is an Opus-class model, but faster, more token-efficient and lower cost," and separately that "our internal assessment is that Grok 4.5 is roughly comparable to Opus 4.7, but much faster."[14]

The comparison target in the second, more specific statement is Claude Opus 4.7, not the then-current Claude Opus 4.8, and it is described as an internal assessment. TechCrunch reported Opus 4.7 as priced at $5 per million input tokens and $25 per million output tokens at the time, against $2 and $6 for Grok 4.5.[14] SpaceXAI's own model card is consistent with the narrower reading: on the coding suites where both appear, Grok 4.5 generally sits above Opus 4.7 and below Opus 4.8, leading Opus 4.7 by 21.9 points on DeepSWE v1.0, 13.0 points on SWE-Marathon, and 4.4 points on Terminal-Bench 2.1, while trailing Opus 4.8 on SWE-Bench Pro, SWE-Bench Multilingual and DeepSWE v1.1. The pattern is not uniform: on DeepSWE v1.0 Grok 4.5 scores 62.0% against Opus 4.8's 55.8%.[6] The phrase "Opus-class" is therefore defensible as a capability tier and not as a claim of parity with Anthropic's newest model at the time.

Competitive position

Grok 4.5 entered a market where four other frontier releases were live or arriving within days.

ModelDeveloperPosition relative to Grok 4.5
Claude Opus 5AnthropicAhead on the Artificial Analysis Intelligence Index as of late July 2026, at 56 to 61 depending on effort setting, against 54[16]
Claude Fable 5AnthropicLed most coding benchmarks in SpaceXAI's own model card, including SWE-Bench Pro, Terminal-Bench 2.1, and both DeepSWE versions[6]
Claude Opus 4.8AnthropicSplit results: ahead on SWE-Bench Multilingual, SWE-Bench Pro, and DeepSWE v1.1; behind on DeepSWE v1.0, SWE-Marathon, APEX-SWE, and token efficiency[6]
GPT-5.6OpenAIReleased the day after Grok 4.5; the Sol variant leads on Tau-3 Banking, RelBench, and GDPval-AA v2 in SpaceXAI's own comparisons, and ties on SWE-Atlas-QnA[6][14]
Gemini 3.1 ProGoogleScored below Grok 4.5 on EEBench in the model card's electrical-engineering comparison; Artificial Analysis reported Grok 4.5 above all Gemini models on its index at launch[6][13]
Muse Spark 1.1MetaIntroduced July 9, 2026, one day after Grok 4.5; a multimodal agentic model launched with Meta's first public model API, with no head-to-head figures in either vendor's published comparisons[21]
GLM-5.2Zhipu AIThe open-weights comparison point in the model card; trails Grok 4.5 on every suite where both appear[6]

The summary that survives cross-checking is narrower than the launch messaging but still substantial. Grok 4.5 is not the most capable model in its cohort on most published measures. It is unusually cheap for the capability it does deliver, unusually economical with output tokens, and, on the specific axis of very long-horizon engineering tasks, ahead of the models it was compared against.

Safety

The model card reports testing for disallowed content, jailbreaks, cybersecurity behavior, and chemical, biological, radiological, and nuclear risks. It characterizes the model as showing "solid but sub-threshold dual-use knowledge" in biology and chemistry, "indicating limited actionable uplift for an already-trained actor," while also documenting substantial cybersecurity capability. These are developer-run evaluations rather than independent safety audits.[6]

Safety measureGrok 4.5 resultDirection
Jailbreak compliance on should-refuse prompts0.73%Lower is better
General refusal compliance1.1%Lower is better
Child-safety compliance0.0%Lower is better
CBRN refusal accuracy (bio / chem / radiological-nuclear)97.9% / 96.7% / 97.9%Higher is better
Self-harm compliance0.5%Lower is better
Sycophancy0.01%Lower is better
Epistemic bias rate20.4%Lower is better
MASK-Rectified dishonesty under pressure0.67%Lower is better

On cyber capability the model card reports the unsafeguarded score deliberately, to measure ability rather than refusal behavior. Grok 4.5 reproduced 80.4% of tasks on CyberGym, above Claude Opus 4.8 at 78.1% and below Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. With safeguards enabled, SpaceXAI's internal HackerBench v0.2 reports Grok 4.5 complying with 7.8% of harmful or dual-use cyber requests, against 25.0% for GPT-5.5 and 22.8% for Claude Opus 4.8, while refusing 0.0% of benign tasks.[6] Cursor said separately that it had added "new safeguards reflecting the model's cybersecurity capabilities" at launch.[7]

The epistemic-bias figure is the least flattering number SpaceXAI published about itself. A 20.4% bias rate means that on roughly one in five contested questions the model framed its answer more favorably toward one side or expressed an opinion where the company's own standard says it should stay neutral. The model card presents this as a tracked release indicator rather than a pass or fail result.[6]

SpaceXAI states that Grok 4.5 is not intended to make autonomous high-stakes decisions in medicine, law, finance, or safety-critical systems without human oversight and validation by domain experts. More generally, its long context window and access to tools do not remove familiar language-model failure modes such as incorrect factual claims, flawed code, missed constraints, or inappropriate actions. Applications using terminal or search access therefore need permission boundaries, output review, and task-specific testing.[4][6]

The model is also closed rather than openly reproducible. Its weights, parameter count, complete data mixture, and full training process are not public. Evaluation claims should therefore be read together with the model card, the disclosed CursorBench overlap, and independent measurements rather than treated as a complete account of performance.[6][7][8]

Reception and criticism

Reception split along a consistent line: the pricing and efficiency claims held up under independent testing, while the capability framing did not fully survive it.

Coverage generally accepted the cost argument. Artificial Analysis' per-task cost measurements and Snorkel's professional-work results were both favorable and were produced outside SpaceXAI.[13][15] The token-efficiency claim, which is the mechanism behind most of the cost advantage, was corroborated by an independent evaluator rather than resting on vendor charts alone.[13]

The capability framing drew more scepticism. TechCrunch wrote that the launch-day benchmark charts "appeared to show Grok's competitiveness with other top models from SpaceXAI competitors, although just short of best-in-class."[14] The charts bear that out: of the four coding benchmarks SpaceXAI displayed on July 8, Grok 4.5 beat Claude Opus 4.8 on two and lost the other two, and Claude Fable 5 led all four.[22] Artificial Analysis, measuring independently, placed the model fourth on its Intelligence Index rather than level with the Opus tier.[13] The divergence between the model card's sub-1% hallucination figure and Artificial Analysis' finding that hallucination more than doubled from the previous generation remains the single largest unreconciled discrepancy in the public record for this model.[6][13]

A separate and more serious controversy attached to the client software rather than to the model. On July 12, 2026, a security researcher publishing as cereblab released a wire-level analysis of Grok Build version 0.2.93, in which Grok 4.5 is the default model, reporting that ordinary sessions uploaded the entire tracked Git repository and its history to cloud storage regardless of which files the agent actually opened, including committed secrets and files the agent had been told not to read. The analysis put the upload volume at roughly 27,800 times what the coding task required, and reported that the in-client privacy toggle did not stop the transfers.[18] The finding contradicted SpaceXAI's own marketing language for Grok Build. Elon Musk confirmed that the uploads had occurred and said SpaceXAI would delete the data collected before the fix; the company documented zero-data-retention behavior and a privacy endpoint.[18] SpaceXAI then published the Grok Build client as open source under the Apache 2.0 licence, in a repository created on July 14, two days after the analysis.[24] As of late July 2026 no independent audit had verified the deletion.[18]

The incident concerns Grok Build's file-handling behavior, not the Grok 4.5 model weights or the Cursor training pipeline, and the two should not be conflated. It nonetheless bears on the same question raised by the Cursor arrangement, which is how much developer code moves to SpaceXAI's infrastructure and under what controls, and it arrived four days after a launch whose central selling point was putting the model inside working codebases.

Grok models have a longer history of content and moderation controversy, including the MechaHitler incident in 2025 and a child-safety controversy in 2026. No comparable incident has been documented for Grok 4.5 specifically as of late July 2026, and its model card reports 0.0% compliance on child-safety prompts and 1.1% on the broad disallowed-content suite; those are vendor-run measurements on a model that had been public for under three weeks.[6]

References

  1. ^SpaceXAI. "Release Notes." SpaceXAI Docs. Updated July 23, 2026. docs.x.ai/...release-notes
  2. ^SpaceXAI. "Introducing Grok 4.5." July 16, 2026. x.ai/...grok-4-5
  3. ^SpaceXAI. "Grok 4.5." SpaceXAI Docs. Updated July 17, 2026. docs.x.ai/...grok-4-5
  4. ^SpaceXAI. "Models and Pricing." SpaceXAI Docs. Updated July 9, 2026. docs.x.ai/...models
  5. ^SpaceXAI. "Pricing." SpaceXAI Docs. Accessed July 27, 2026. docs.x.ai/...pricing
  6. ^SpaceXAI. "Model Card: Grok 4.5." July 14, 2026, revised July 20, 2026. media.x.ai/...4p5-5184fdf9.pdf
  7. ^Cursor. "Introducing Grok 4.5." July 8, 2026. cursor.com/...grok-4-5
  8. ^Artificial Analysis. "Grok 4.5 (high): Intelligence, Performance and Price Analysis." Accessed July 27, 2026. artificialanalysis.ai/...grok-4-5
  9. ^Reuters. "SpaceXAI launches Grok 4.5 model for coding, agentic tasks." July 8, 2026. investing.com/...-for-coding-agentic-tasks-4782511
  10. ^Cursor. "Cursor partners with SpaceX on model training." April 21, 2026. cursor.com/...spacex-model-training
  11. ^CNBC. "SpaceX to acquire the AI coding startup Cursor for $60 billion." June 16, 2026. cnbc.com/...spacex-spcx-cursor-acquisition-ipo
  12. ^Cursor. "Data Use and Privacy Overview." Updated July 15, 2026. cursor.com/data-use
  13. ^Artificial Analysis. "Grok 4.5 brings SpaceXAI to the intelligence frontier." July 8, 2026. artificialanalysis.ai/...the-intelligence-frontier
  14. ^TechCrunch. "SpaceXAI releases Grok 4.5, which Elon describes as an 'Opus-class model'." July 8, 2026. techcrunch.com/...describes-as-an-opus-class-model
  15. ^Snorkel AI. "Grok 4.5 testing results: how SpaceXAI's new model performs on real professional work." July 8, 2026. snorkel.ai/...l-performs-on-real-professional-work
  16. ^Artificial Analysis. "Model Leaderboard." Accessed July 27, 2026. artificialanalysis.ai/...models
  17. ^Cursor. "Grok 4.5 Model Card." July 14, 2026. cursor.com/...grok-4-5-model-card
  18. ^The Next Web. "Grok Build was uploading entire Git repositories to xAI's cloud, including committed secrets." July 14, 2026. thenextweb.com/...-entire-git-repositories-secrets
  19. ^SpaceXAI. "Reasoning." SpaceXAI Docs. Accessed July 27, 2026. docs.x.ai/...reasoning
  20. ^Yahoo Finance. "Elon Musk confirms SpaceX merger with xAI ahead of IPO." February 3, 2026. finance.yahoo.com/...th-xai-ahead-of-ipo-220616795
  21. ^Meta AI. "Introducing Muse Spark 1.1." July 9, 2026. ai.meta.com/...introducing-muse-spark-meta-model-api
  22. ^SpaceXAI. "Introducing Grok 4.5." Archived copy of the launch-day page, captured July 8, 2026. web.archive.org/...grok-4-5
  23. ^Artificial Analysis. "Grok 4.5 (high)." Archived copy captured July 9, 2026. web.archive.org/...grok-4-5
  24. ^SpaceXAI. "xai-org/grok-build." GitHub. Repository created July 14, 2026, Apache 2.0. github.com/...grok-build

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

2 revisions by 1 contributors · v3 · 5,675 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Reviewer note: Every number checked against xAI docs, the Grok 4.5 model card PDF, the launch post (via archive, x.ai blocks direct fetch), Cursor, Artificial Analysis raw data and Snorkel. 12 fixes applied.

Suggest edit