Grok 4.6 just took the #1 spot on MedAgentBench....one of the most interesting benchmarks for real-world agentic healthcare tasks
and Grok now has two generations sitting in the top 3
• #1 Grok 4.6 — ~95.9%
• #2 GPT-5.6 Sol — ~94.7%
• #3 Grok 4.5 — ~93.4%
MedAgentBench goes far beyond answering medical questions
It tests AI agents on 300 clinically derived tasks across 10 categories inside a realistic electronic health-record environment, requiring the model to reason, use tools and actually execute multi-step clinical workflows
Grok 4.5 was already one of the strongest medical agents tested. Grok 4.6 just pushed it to another level
Grok is rapidly becoming one of the strongest AI model families for agentic healthcare work
and Grok now has two generations sitting in the top 3
• #1 Grok 4.6 — ~95.9%
• #2 GPT-5.6 Sol — ~94.7%
• #3 Grok 4.5 — ~93.4%
MedAgentBench goes far beyond answering medical questions
It tests AI agents on 300 clinically derived tasks across 10 categories inside a realistic electronic health-record environment, requiring the model to reason, use tools and actually execute multi-step clinical workflows
Grok 4.5 was already one of the strongest medical agents tested. Grok 4.6 just pushed it to another level
Grok is rapidly becoming one of the strongest AI model families for agentic healthcare work