proomt

Search

Search posts, papers, and topics

All posts

AntithesisRohan Padhye7 min readintermediate

Why we taught agents to break distributed safety properties

Summary

The post introduces an Antithesis skill that lets AI agents perform mutation testing on distributed systems, injecting subtle bugs to check whether generated test suites can falsify safety properties. Using rqlite, the agents ran ~24 hours of tests, falsified 11 of 13 properties, and uncovered three real upstream bugs, demonstrating the technique’s practical value.

  • AI agents can infer safety invariants from code and auto‑generate Antithesis harnesses for distributed systems like rqlite.
  • Mutation testing with targeted bugs (e.g., swapping to an older snapshot) exposed gaps; 11 of 13 safety properties were falsified after 19 mutants across 46 runs.
  • ~24 hours of fault‑injection runs uncovered three upstream rqlite bugs, proving the approach catches real‑world issues.
  • When a property isn’t falsified, the workflow suggests refining test configuration or revisiting the property definition.

Distributed‑systems engineers and test‑automation teams should care because AI‑driven mutation testing can reveal subtle safety bugs that ordinary fault injection misses, increasing confidence in system correctness.

8/10

Related reading

  1. Changing the game: How Google uses agentic AI to secure hundreds of millions of lines of code

    Google’s AI & Infrastructure team built an agentic pipeline (Mantis) that runs pre‑submit AI‑driven scans on every code check‑in, validates findings with a fast triage agent (AST + call‑graph analysis) achieving >92% precision in <1 min, then auto‑generates fixes via a bug‑fix agent. Localized threat models and a two‑step scan cut false‑positives to ~3% and prevent hundreds of vulnerabilities eac…

    Google Cloud Bloggoogle.com4 min
  2. Optimizing GitHub Actions for Agent PRs: Speculative Test Slicing and AST Impact Analysis

    A step‑by‑step guide for handling the flood of pull requests generated by code‑generation agents. It builds a TypeScript CLI that uses ts‑morph to do AST‑level change‑impact analysis, maps affected symbols to tests, and runs only those tests in a “speculative” GitHub Actions job while a full‑suite verification runs in the background. The article includes concrete CLI code, dependency choices, con…

    SitePointsitepoint.com17 min
  3. How to operate shared platforms safely at agent scale

    Datadog explains how scaling AI agents turns isolated executions into shared‑platform risk and outlines a systematic approach to model agent trajectories, monitor per‑dependency constraints, and enforce workload‑specific capacity policies. The result is proactive detection of bottlenecks and protection against noisy‑neighbor failures.

    Datadogdatadoghq.com11 min
  4. Towards Self-Driving Codebases

    The post argues that AI agents could eventually handle low‑level engineering tasks—bug fixing, debugging, UI consistency, growth experiments—if the dev toolchain is made “agent‑legible”. It outlines missing primitives (global memory, code‑base rot prevention, better dev environments) and proposes a bootstrapping process to measure and improve a repo’s “agent readiness”. The piece is largely specu…

    Hacker News front pagedetail.dev9 minHN12099