Grok 4.7 vs GPT-6 Astra, Claude, and Chinese models — a fresh comparison
· Source: original
Grok 4.7, GPT-6 Astra, Claude and Qwen3.8 Max: a fresh comparison ⚖️
Four releases in a row — and the freshest, Grok 4.7 from SpaceXAI, came out literally yesterday relative to the publication of this breakdown. GPT-6 Astra from OpenAI and Claude Fable 5.1 from Anthropic came out three weeks before that, and Claude Opus 5 lined up alongside them in July. All the releases fit within a single month.
Inside Anthropic, the pairing was split by task: for most things you should enable Claude Opus 5, and bring out Fable 5.1 when Opus can no longer handle the benchmarks.
On price, the Chinese are right on their heels. Qwen3.8 Max costs $2 and $6 per million tokens — exactly the same figures as Grok. On the overall index, Qwen3.8 Max trails Grok by just one point. Kimi K3 on day-to-day engineering tasks sits in the same corridor as the fresh Grok.
The set of tasks for comparison — long code, terminal, schemas, legal cases and billing. And here's where the numbers diverge: the labs' benchmarks and the independent Artificial Analysis run sit separately, and on Terminal-Bench the discrepancy is almost one and a half times.
Full comparison — in the breakdown on Habr
🤖 Interested in AI agents and automation?
Prompts for building AI agents and automations — read up on the topic:
🔗 The entire prompt library · "AI Agents" category
A ready-made product on the topic: AI Prompts for Productivity — grab it and apply it right away.