Can AI be trusted with infrastructure?Measure it, don't assume it.
PhareauBench is a research benchmark measuring how AI models actually perform on infrastructure work — infrastructure-as-code, incident response, terminal operations, agentic tool use — and whether they exercise the judgement that operating real systems demands. It is in active design; the methodology will be published as an open technical report.
Profiles, not a leaderboard
v1 publishes a per-model radar profile across capability and safety axes — no single aggregate score, no headline ranking. Composite scoring comes later, only once the evidence shows which axes actually matter.
Safety measured as judgement
The flagship probe asks whether a model can tell a destructive infrastructure change from a safe one — and whether it proceeds when properly authorised. Blanket refusal is a failure mode too; the benchmark is designed so it cannot be gamed by saying no to everything.
Grounded in real incidents
The design is driven by documented production incidents and the academic literature, not intuition. A production wipe by an AI agent running terraform destroy is the class of failure this benchmark exists to measure and prevent.
Vendor-neutral by charter
Subjects are the most-used current models regardless of vendor. The lab's own tooling preferences play no role in what gets evaluated or how results are narrated.
The benchmark is being designed in the open.