Four models from the same family on five routine tasks — 40 runs with honest measurements
· Source: original
⚖️ 4 models from one family, 5 routine tasks, 40 runs: the expensive one turned out to be faster
Analysis on Habr — four models from the same family, at different price points, on five identical routine tasks. Each task was run twice, for a total of 40 runs. The tests were hidden — the models were never shown them.
The occasion was an argument on Habr: does a programmer need the smartest model, or are cheap ones enough for routine work. That argument racked up 300 comments and not a single measurement on identical tasks.
⏱️ What actually differs is price and time. The most expensive model turned out to be the fastest: 341 seconds versus 395 for the cheapest one. The reason — fewer turns and half as much text.
📊 On quality the picture is more even: four out of five tasks were solved by all the models — 32 runs out of 32 successful. Then came the fifth task. The junior, cheapest model wrote a bug and reported back with two checkmarks in a row, the second of which contradicts the first
🤖 Interested in AI agents and automation?
Prompts for building AI agents and automations — read up on the topic:
🔗 The entire prompt library · "AI Agents" category
A ready-made product on the topic: Prompts for Programmers — grab it and apply it right away.