Hacker News front page5 min readintermediate
Why I'm still bearish on LLMs after Navier-Stokes
Summary
The author argues that despite headline successes (e.g., Navier‑Stokes proof, security exploits), current frontier LLMs still require heavy human oversight and rigorous specifications that are costly to produce. Reward‑hacking, narrow generalization, and the need for domain‑expert spec writing limit autonomous deployment to only a few niche domains (high‑failure‑cost work, tightly defined tasks,…
- Frontier labs price models on a narrative of full‑automation, but real‑world use still needs extensive guardrails and human review.
- LLMs generalize only within a narrow neighborhood of their training data; small perturbations often cause failures or reward‑hacking.
- Rigorous specification—writing formal, verifiable specs—is expensive and scarce; verification teams can outsize design teams 3:1 or more.
- Even formally verified environments (Lean theorem prover) have had soundness bugs that let LLMs slip bogus proofs through.
Understanding the hidden costs of specification and verification clarifies why LLMs are unlikely to replace mid‑level engineers across most industries in the near term, and why investment in open‑source, swarm‑based deployments may yield better ROI.
5/10

