xAI released Grok 4.6, a 1.5 trillion-parameter model with extended supplemental training focused on long-running agentic tasks, described as a major step up from 4.5 at the same price point.
Benchmarks show Grok 4.6 hitting 88.4% on Terminal-Bench v2.1 and competitive performance on Code Arena with GPT-5.6 Sol, while Independent evals from Artificial Analysis place it at #61 on the Intelligence Index, roughly in line with GPT-5.6 Sol Max but behind Claude Opus/Fable.
Pricing is a central differentiator: Grok 4.6 costs $2/$6 per 1M input/output tokens, materially below frontier peers, with practitioners immediately framing it as the new default for coding and bug-finding workloads.
Grok 4.6 training emphasized agentic RL across coding, web development, computer-aided design, and kernel optimization tasks, with reported improvements in long-task self-testing behavior and more stable performance on reasoning-heavy benchmarks.
Architects care: at 37% of frontier pricing (Grok 4.6 $2/$6 vs. GPT-5.6 Sol at $5+/$15 typical), Grok 4.6 shifts the cost-capability tradeoff for autonomous coding and knowledge-work agents; coupled with the Grok Bot team-member product launch, this signals xAI/SpaceX entering the 'agentic coding' category seriously.