Related reading
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
Feyospace‑v1 presents a data‑centric training pipeline for cyber‑security agents, combining five systems (Choulea, SkyReal, Hongzwang, PSBreakup, Kreator) to generate and verify 164 k long‑context trajectories across diverse exploit environments. The resulting checkpoints improve baseline performance by ~24% on CyberGym and achieve a 63% verified success rate, ranking top among similarly‑sized op…
Hugging Face Daily Papersarxiv.org1 minpaperCADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
CADWorld is a new benchmark suite of 200 long‑horizon mechanical CAD tasks in FreeCAD, covering sketching, part modeling, assembly, CAM, FEM, and more. Agents interact via screenshots and GUI actions; success is checked by executable validation of the saved CAD artifacts. Seven existing agents achieve at most 17.5 % success versus an 87 % expert baseline, highlighting the gap between GUI competen…
Hugging Face Daily Papersarxiv.org1 minpaperModernizing the Trade Lifecycle With Governed Data and AI
Databricks argues that modernizing the trade lifecycle now hinges on building a governed, real‑time data foundation that spans research, trading, risk, ops and compliance, rather than isolated AI pilots. Starting with a few high‑value questions—execution cost, shock risk, exception rates—and using Unity Catalog and Agent Bricks lets firms achieve measurable speed and auditability gains before sca…
Databricksdatabricks.com5 minMathematical Billiards (2024)
Hacker News front pageuni-heidelberg.deHN51

