proomt

Search

Search posts, papers, and topics

All posts

Simon Willison21 min readintermediate

2026 in LLMs (so far)

Summary

The post recaps 2026 LLM milestones, noting that Claude Opus 4.5 and GPT‑5.1 made coding agents reliable enough for daily use, sparking AI‑driven side projects and a surge of sandboxing and agent‑security discussions. The author reflects on "AI mania", personal experiments, and the cultural impact on engineers.

  • Claude Opus 4.5 and GPT‑5.1 released in Nov 2025; their coding agents shifted from error‑prone to reliable for everyday use.
  • Reliability enabled engineers to rely on agents for real projects, prompting the author to build a pure‑Python JavaScript interpreter and a Python‑based WebAssembly runtime.
  • Sandbox and agent‑security concerns exploded, with ~40 of 277 conference sessions covering them, highlighting the need for robust isolation.
  • The term "Deep Blue" was coined to describe engineer ennui when AI can do everything, signaling a cultural shift in software work.

Engineers should know LLM coding agents are now production‑ready and that security/sandboxing is a critical, active area of work.

5/10

Related reading

  1. Note on 24th September 2026

    The author argues that LLM coding agents make software engineering harder and demand extraordinary discipline and expertise. A sponsor note adds that focusing on quality over quantity yields better security vulnerability detection.

    Simon Willisonsimonwillison.net1 min
  2. The Agent Said It Was Done. The Database Disagreed.

    ThinkingBox benchmarks AI agents by checking the final database state after tool calls, revealing that many LLM‑driven agents succeed on a single attempt but fail to repeat the correct outcome. Across 507 tasks run 20 times, models differ widely in consistency and cost per successful attempt, with Kimi‑K3 being broad but inconsistent and Claude Opus models being more reliable.

    Hugging Facehuggingface.co13 minHN1
  3. Claude Opus 5.5

    Claude Opus 5.5 is Anthropic’s latest LLM, positioned to match Claude Fable 5.1 on most tasks while cutting compute cost by ~40% and latency by >30%. It ships with the same safety guardrails as the top‑tier models, scores best on Anthropic’s internal alignment audit, and shows measurable gains on coding‑heavy benchmarks (e.g., 66.4% on Agentic‑coding Terminal‑Bench vs 52.3% for Opus 5). Pricing d…

    Hacker News front pageanthropic.com15 minreleaseHN17921117
  4. Agentic DevOps World 2026: Key Takeaways

    Enterprise AI code generation is outpacing governance: 92% of leaders feel confident but 81% see more production issues; post‑generation stages (review, testing, deployment) are now the bottleneck. Successful adoption requires up‑skilling, end‑to‑end metrics, and tooling (e.g., CloudBees DevOps Agent Kit, PR‑auto‑approval agents).

    Codeshipcloudbees.com8 min