xAI just dropped Grok 4.7 — their new frontier model built for coding and knowledge work that actually runs for hours without breaking.
Bigger base model than 4.6. Longer RL training on harder tasks. Better at verifying its own work, handling long context, and running inside Grok Bot.
Benchmark flex:
• CursorBench 4.0: 46.3% (vs 40.4% on 4.6)
• Terminal Bench 4.0: 38% (up from 20.3%)
• DeepSWE 1.1: 71% at high reasoning
• AA Briefcase: 1,657
• EEBench: 64%
Beats GPT 5.6 Sol on some evals. Still behind on others vs GPT 5.6 Sol and Fable 5.1.
Pricing stays at $2/M input tokens, $6/M output tokens. 2x speed version at 2x price.
New safeguard stack: 3.3% risky dual-use prompt leak on HackerBench 0.3, 62.4% on LatchBio biosafety benchmark. Red team access for select cybersec partners.
Live now on Cursor, Grok Build, Grok API, plus third-party tools and cloud platforms.
If you're building agents or need long-running code execution, this is the upgrade.
Bigger base model than 4.6. Longer RL training on harder tasks. Better at verifying its own work, handling long context, and running inside Grok Bot.
Benchmark flex:
• CursorBench 4.0: 46.3% (vs 40.4% on 4.6)
• Terminal Bench 4.0: 38% (up from 20.3%)
• DeepSWE 1.1: 71% at high reasoning
• AA Briefcase: 1,657
• EEBench: 64%
Beats GPT 5.6 Sol on some evals. Still behind on others vs GPT 5.6 Sol and Fable 5.1.
Pricing stays at $2/M input tokens, $6/M output tokens. 2x speed version at 2x price.
New safeguard stack: 3.3% risky dual-use prompt leak on HackerBench 0.3, 62.4% on LatchBio biosafety benchmark. Red team access for select cybersec partners.
Live now on Cursor, Grok Build, Grok API, plus third-party tools and cloud platforms.
If you're building agents or need long-running code execution, this is the upgrade.