Related reading
RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation
RiskChainBench is a new benchmark that pairs synthetic obfuscated message restoration inputs with human‑labeled local web environments, requiring models to both decode malicious instructions and investigate the linked site. Across ten models, restoration accuracy varies widely and web‑agent failures dominate the error budget.
Hugging Face Daily Papersarxiv.org1 minpaperDefending Against Active Exploitation of Citrix NetScaler ADC and Gateway Appliances
Mandiant and Google identified active exploitation of two zero‑day bugs in Citrix NetScaler ADC/Gateway that let attackers gain root via malformed DTLS packets, then persist with custom PHP web shells and a Python tunneler. The blog details the exploit mechanics, persistence tricks, detection signatures, and remediation steps.
Google Cloud Bloggoogle.com22 minSecuring AI Pull Requests: Building a Deterministic AST Audit Harness in GitHub Actions
A step‑by‑step tutorial for building a deterministic AST‑based security scanner that runs in GitHub Actions to catch prototype‑pollution, high‑entropy secrets, and unauthorized network egress in AI‑generated pull requests.
SitePointsitepoint.com19 minAI Agents Are Disrupting Open Source Security Disclosure
AI agents can turn minimal public hints about software bugs into working exploits, rendering traditional embargoes ineffective. The article cites a study where a GPT‑4 agent exploited 87% of a 15‑vulnerability benchmark from CVE descriptions and discusses faster releases and revocable capabilities as mitigations.



