proomt

Search

Search posts, papers, and topics

All posts

HostingerSimon Lim8 min readintermediate

What we learned running seven AI coding models on real production work

Summary

Hostinger ran seven AI coding models on 12 real PRs, measuring match to engineer solutions, cost, latency, and error patterns. The top five models performed similarly, so speed, price, and review rigor matter more than raw quality scores.

  • All top‑five models scored 0.83‑0.88, showing little quality gap; choose based on latency and cost for your workflow.
  • Models often missed critical edge cases (e.g., missing pause before retry) despite high similarity scores.
  • Cheapest model (GPT Luna) used fewer tokens and did less work, touching only ~43% of the engineer’s edit size.
  • High similarity scores can hide unmergeable PRs—excessive file changes or irrelevant edits were common.

Engineering teams considering AI code‑generation tools should prioritize robust review processes and model selection based on speed, cost, and edge‑case handling rather than raw similarity scores.

7/10

Related reading

  1. The Agent Said It Was Done. The Database Disagreed.

    ThinkingBox benchmarks AI agents by checking the final database state after tool calls, revealing that many LLM‑driven agents succeed on a single attempt but fail to repeat the correct outcome. Across 507 tasks run 20 times, models differ widely in consistency and cost per successful attempt, with Kimi‑K3 being broad but inconsistent and Claude Opus models being more reliable.

    Hugging Facehuggingface.co13 minHN1
  2. 63% of developers have more work since non-devs began coding with AI, but most say it's good for the industry

    A Zapier survey of 797 U.S. developers finds 63% see more work as non‑technical coworkers use AI coding tools, yet two‑thirds view the shift positively. Developers are moving toward higher‑value strategy, governance, and cross‑functional work, and the top future‑proof skills are security, AI/ML, low‑code QA, strategy, and data engineering.

    Zapier Engineeringzapier.com5 min
  3. The 2026 State of Code Abundance Report: Exposing the Enterprise AI Readiness Gap

    A report highlights a significant gap between enterprise confidence in AI-generated code and operational reality, with 81% of leaders reporting increased production issues despite high readiness scores. This "code abundance" means code is generated faster than organizations can effectively test, govern, and manage it, leading to challenges in cost attribution and governance.

    Codeshipcloudbees.com5 min
  4. What four people at Hostinger actually do with AI all day (and what happens when you have an agentic beef)

    Hostinger staff use custom AI agents to automate daily tasks—from code reviews to influencer lead sourcing—shifting their work from doing the work to managing the agents. Building reliable “harnesses” (prompt contexts, constraints) consumes most of the engineering effort, and agents still hallucinate, repeat work, or suggest unsafe fixes, so human oversight remains essential.

    Hostingerhostinger.com8 min