Simon Willison21 min readintermediate
2026 in LLMs (so far)
Summary
The post recaps 2026 LLM milestones, noting that Claude Opus 4.5 and GPT‑5.1 made coding agents reliable enough for daily use, sparking AI‑driven side projects and a surge of sandboxing and agent‑security discussions. The author reflects on "AI mania", personal experiments, and the cultural impact on engineers.
- Claude Opus 4.5 and GPT‑5.1 released in Nov 2025; their coding agents shifted from error‑prone to reliable for everyday use.
- Reliability enabled engineers to rely on agents for real projects, prompting the author to build a pure‑Python JavaScript interpreter and a Python‑based WebAssembly runtime.
- Sandbox and agent‑security concerns exploded, with ~40 of 277 conference sessions covering them, highlighting the need for robust isolation.
- The term "Deep Blue" was coined to describe engineer ennui when AI can do everything, signaling a cultural shift in software work.
Engineers should know LLM coding agents are now production‑ready and that security/sandboxing is a critical, active area of work.
5/10


