Grok 4.6 just took the #1 spot on MedAgentBench....one of the most interesting benchmarks for real-world agentic healthcare tasks

and Grok now has two generations sitting in the top 3

• #1 Grok 4.6 — ~95.9%
• #2 GPT-5.6 Sol — ~94.7%
• #3 Grok 4.5 — ~93.4%

MedAgentBench goes far beyond answering medical questions

It tests AI agents on 300 clinically derived tasks across 10 categories inside a realistic electronic health-record environment, requiring the model to reason, use tools and actually execute multi-step clinical workflows

Grok 4.5 was already one of the strongest medical agents tested. Grok 4.6 just pushed it to another level

Grok is rapidly becoming one of the strongest AI model families for agentic healthcare work