'Grok 4.6' has arrived, making it more suitable for long hours of agent work and bringing its performance to par with GPT-5.6 Sol and Claude Fable 5.



SpaceXAI announced its AI model ' Grok 4.6 ' on August 12, 2026. Based on Grok 4.5, it offers significant improvements at the same price, focusing on long-running AI agent tasks and more advanced interactive visual processing.

Introducing Grok 4.6 | SpaceXAI

https://x.ai/news/grok-4-6




SpaceXAI announced that Grok 4.6 'achieves cutting-edge intelligence across multiple agent-based coding and knowledge work benchmarks.' It scored 61 points on the Artificial Intelligence Analysis Index (AAII), which comprehensively evaluates nine benchmarks, achieving a score equivalent to OpenAI's flagship model, ' GPT-5.6 Sol .'



Looking at the benchmark details, Grok 4.6 achieved a particularly high score of 1753 in '

GDPVal-AA v2, ' which measures the practical capabilities of a model. It also achieved 69.9% in ' CursorBench 3.2 ,' which evaluates ambiguous and multi-file coding agent tasks, and 61.3% in ' FrontierCode v1.1 ,' which shows how well a model meets the standards of a high-quality production environment codebase, second only to Anthropic's ' Claude Fable 5 Max .' Compared to Grok 4.5, performance has improved, especially in long-duration agent tasks and coding-related evaluations. On the other hand, in ' DeepSWE 1.1, ' which measures pollution tolerance for AI coding agents, it scored 65.9%, a significant improvement over Grok 4.6 High, but still inferior to GPT-5.6 Sol and Claude Fable 5 Max.



Furthermore, an analysis by 'Artificial Analysis,' a website that compares and analyzes the performance and characteristics of AI models, assessed that 'SpaceX AI has returned to the intelligent frontier alongside OpenAI, and is now positioned second only to Anthropic.'




During training, Grok 4.6 underwent longer-term additional training than Grok 4.5, utilizing inference, advanced technical concepts, high-quality engineering data, and improved optimizers and training recipes. This reportedly built a stronger foundation for subsequent supervised fine-tuning (SFT) and reinforcement learning (RL). Subsequently, Grok 4.5 was used to regenerate SFT trajectories across various inference levels, agent harnesses, and domains such as STEM , software engineering, and knowledge work, filtering out problematic data through model-based checks.

Grok 4.6 places particular emphasis on maximizing its applicability and 'work continuity capability' across multiple steps. As a result of this training, SpaceXAI says that Grok 4.6 has proven to be particularly good at 'turning vague product ideas into a working first version,' enabling it to explore unfamiliar areas, design application structures, implement core interactions, and continuously improve results through multiple feedback iterations.

Grok 4.6 is designed not only for simple text generation, but also for transforming ideas into working applications and deliverables. Because it can establish the application's structure and visual language in a single process if a concrete product idea is provided, it's particularly useful in projects that adopt a method of starting concretely and iterating to achieve the fastest possible results.

Furthermore, security measures have been strengthened, making it useful and safe for applications such as vulnerability patching, accelerating the engineering design cycle, and enhancing AI research. SpaceXAI has also built a security flow to support high-performance AI like Grok 4.6, conducting the most extensive pre-deployment testing to date on functionality and security measures, as well as extensive post-deployment testing and third-party testing.

As of August 13, 2026, Grok 4.6 will be available through the coding agents ' Cursor ' and ' Grok Build ,' as well as via API. It will also be available through partners such as OpenRouter , Vercel , and Cloudflare . Pricing starts at $2 (approximately 318 yen) per 1 million input tokens and $6 (approximately 956 yen) per 1 million output tokens, with a high-speed version available at twice the price.

in AI, Posted by log1e_dh