Grok 4.6 Launches, Rivals GPT-5.6 in Performance at Over 60% Lower Cost

xAI has launched Grok 4.6, achieving a score of 61 on the AI Intelligence Index, matching GPT-5.6 Sol (Max) and trailing slightly behind Claude Opus 5 (63 points) and Claude Fable 5 (62 points). This performance positions Grok 4.6 among the world’s top-tier models, directly challenging flagship offerings from OpenAI and Anthropic. Built on the foundation of Grok 4.5, the update introduces core enhancements in reasoning and advanced technical concept training, leveraging curated data generated by the model itself, alongside high-quality engineering datasets and an optimized training algorithm.
A pivotal shift in the training methodology involves a self-reinforcing data loop: Grok 4.5 was utilized to regenerate supervised fine-tuning trajectories, focusing on reasoning, agent tool utilization, and domains such as STEM, software engineering, and knowledge work. The model also underwent extensive training on agent-based reinforcement learning tasks, including knowledge work, general coding, kernel optimization, web development, and computer-aided design—a strategy clearly designed to address the primary workloads of the emerging “agent era.”
Agent benchmarks have seen dramatic improvements, substantially lowering the effort required for complex, long-term tasks.
In agent-related benchmark evaluations, Grok 4.6 demonstrated exceptional performance. It achieved an Elo score of 1753 on GDPval-AA v2, ranking second only to Claude Opus 5. DeepSWE v1.1 scores rose to 65.9%, a notable increase from 54% in Grok 4.5, while APEX-Agents jumped from 47.1% to 57.5%. Terminal-Bench v3.0 improved from 15.7% to 26%, with additional strong results of 69.9% on CursorBench v3.2 and 61.3% on FrontierCode v1.1, marking significant breakthroughs in both coding and agent capabilities.
Notably, Grok 4.6 offers a distinct cost-efficiency advantage. Priced at $2 per million input tokens and $6 per million output tokens, it is over 60% cheaper than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). On the AA-Briefcase long-term knowledge work benchmark, Grok 4.6 completed tasks in an average of 53 rounds using approximately 500 million input tokens, whereas Claude Opus 5 required roughly 103 rounds and 2 billion tokens. This means Grok 4.6 achieves comparable results using less than a quarter of the resources, highlighting a clear cost-performance edge. The model is now accessible via Cursor, Grok Build, API, OpenRouter, Vercel, and Cloudflare. During its first week, users receive double usage credits on Grok Build and Cursor. With a context window of 500,000 tokens, a competitive landscape focused on “equal intelligence at lower prices” has officially begun.
Related article
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
Related Special Topic Recommendations
Comments (0)
0/500

xAI has launched Grok 4.6, achieving a score of 61 on the AI Intelligence Index, matching GPT-5.6 Sol (Max) and trailing slightly behind Claude Opus 5 (63 points) and Claude Fable 5 (62 points). This performance positions Grok 4.6 among the world’s top-tier models, directly challenging flagship offerings from OpenAI and Anthropic. Built on the foundation of Grok 4.5, the update introduces core enhancements in reasoning and advanced technical concept training, leveraging curated data generated by the model itself, alongside high-quality engineering datasets and an optimized training algorithm.
A pivotal shift in the training methodology involves a self-reinforcing data loop: Grok 4.5 was utilized to regenerate supervised fine-tuning trajectories, focusing on reasoning, agent tool utilization, and domains such as STEM, software engineering, and knowledge work. The model also underwent extensive training on agent-based reinforcement learning tasks, including knowledge work, general coding, kernel optimization, web development, and computer-aided design—a strategy clearly designed to address the primary workloads of the emerging “agent era.”
Agent benchmarks have seen dramatic improvements, substantially lowering the effort required for complex, long-term tasks.
In agent-related benchmark evaluations, Grok 4.6 demonstrated exceptional performance. It achieved an Elo score of 1753 on GDPval-AA v2, ranking second only to Claude Opus 5. DeepSWE v1.1 scores rose to 65.9%, a notable increase from 54% in Grok 4.5, while APEX-Agents jumped from 47.1% to 57.5%. Terminal-Bench v3.0 improved from 15.7% to 26%, with additional strong results of 69.9% on CursorBench v3.2 and 61.3% on FrontierCode v1.1, marking significant breakthroughs in both coding and agent capabilities.
Notably, Grok 4.6 offers a distinct cost-efficiency advantage. Priced at $2 per million input tokens and $6 per million output tokens, it is over 60% cheaper than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). On the AA-Briefcase long-term knowledge work benchmark, Grok 4.6 completed tasks in an average of 53 rounds using approximately 500 million input tokens, whereas Claude Opus 5 required roughly 103 rounds and 2 billion tokens. This means Grok 4.6 achieves comparable results using less than a quarter of the resources, highlighting a clear cost-performance edge. The model is now accessible via Cursor, Grok Build, API, OpenRouter, Vercel, and Cloudflare. During its first week, users receive double usage credits on Grok Build and Cursor. With a context window of 500,000 tokens, a competitive landscape focused on “equal intelligence at lower prices” has officially begun.
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a





Home






