The instrument
Point this at your agents. See what they can actually prove.
A free, open probe that grades any agent system against three laws and returns a level, 1 to 4, with the specific failure patterns you’re running. Ground truth is your database and your record — never the logs’ self-report. No signup to read your result. No sales call attached.
Three laws. Ten probes. One verdict.
Can an agent act outside its charter? Rules must live in code paths and DB constraints — a guardrail in a prompt is a preference.
Is any “done” self-reported? Delivery must be verified in the substrate, never asserted by the agent. This is the law most fleets fail first.
Could a duty go silent without anyone noticing? The roster is diffed against reality daily, so silence itself raises the alarm.
The probes are the standard’s ten anti-patterns — the ways a system fails while looking correct — or start with the plain-English six. The verdict is the minimum of the load-bearing pillars: one weak law caps the whole score.
Three ways in
Your agent runs it
Connect the GOVENANT MCP to the coding agent you already use, sign in, and say “run a trust sweep.” Read-only, aggregate-only — your prompts, data, and secrets never leave your environment. Result in minutes.
Install guides for 8 tools →Can’t run code yet?
Leave your details below with “just exploring” or your fleet size — we’ll send the self-check kit: the questions the probe asks, in plain English, so you can scope the answer before committing engineering time.
Get the self-check kit →Have us run it with you
The expert audit: two weeks, every agent you run graded against the standard, findings ranked by real risk, remediation mapped to the three laws. Fixed fee, from $9,500.
The expert audit →What a result looks like
Finding 1 — the self-reported outcome. In the sample window, the overwhelming majority of completions were asserted by the same execution that did the work; no separate check confirmed the outcome existed. The record cannot distinguish a task that succeeded from one that claimed to.
Finding 2 — the silent duty. Two scheduled responsibilities show no real outcome in weeks while their status remained green; nothing distinguishes “no work to do” from “not working.”
What moves this fleet to GOVENANT-2: verified-outcome rows at the terminal edge and a validation gate at the chokepoint — the specific remediation ships with every report.
Every real run gets its own private report URL you can share internally — anomalies first, findings ranked, each with the pattern name, the evidence, and the specific fix.
Get the self-check kit
Three fields. We send the kit and, if you ask, a human reads your result with you. No sequence-blast — you can reply to any mail and reach a person.
Not a certification. Nobody hands out a stamp — including us. Levels are self-assessed against the open standard, published with their probe logs, and open to anyone’s challenge, in the manner of web accessibility guidelines. No independent GOVENANT certification has been issued to any product, ours included. The probe’s verdict is only as good as the record it reads — which is exactly why it reads the record and not the logs.
Ran it? Join the registry.
Free and self-serve: publish your level with its probe log and wear the badge. Every entry is a public claim anyone can challenge — that’s what makes it worth something.
The registry →Want it done with you?
The expert audit: two weeks, fixed fee from $9,500, findings ranked by exposure, and a board-ready one-pager your CISO can actually forward.
The expert audit →