HostingerSimon Lim8 min readintermediate
What we learned running seven AI coding models on real production work
Summary
Hostinger ran seven AI coding models on 12 real PRs, measuring match to engineer solutions, cost, latency, and error patterns. The top five models performed similarly, so speed, price, and review rigor matter more than raw quality scores.
- All top‑five models scored 0.83‑0.88, showing little quality gap; choose based on latency and cost for your workflow.
- Models often missed critical edge cases (e.g., missing pause before retry) despite high similarity scores.
- Cheapest model (GPT Luna) used fewer tokens and did less work, touching only ~43% of the engineer’s edit size.
- High similarity scores can hide unmergeable PRs—excessive file changes or irrelevant edits were common.
Engineering teams considering AI code‑generation tools should prioritize robust review processes and model selection based on speed, cost, and edge‑case handling rather than raw similarity scores.
7/10



