Why
Four tools, four different questions.
These get conflated in the same budget line, and they answer questions that do not substitute for each other. Most teams need more than one; almost nobody needs this instead of evaluation.
| Answers | Acts when | What it cannot do | |
|---|---|---|---|
| Evaluation | How does this model score on a benchmark? | Before deployment | Constrain what the agent does on Tuesday |
| Gateway / proxy | What tokens went where, at what cost? | At the call | Know whether a human authorised this action |
| Observability | What happened, and where did it get slow? | After | Refuse anything |
| Agent governance | Is this action within authority a named person signed for? | Before the action | Tell you whether the model is any good |
The category is forming, not empty
Every major platform vendor is building agent governance — for their own estate. That is rational and it is also the opening: governing a competitor's runtime well is a feature that helps a customer leave, so none of them will build it across. An enterprise runs agents on several.
We would rather say that plainly than claim to be alone in a category. A comparison page that pretends otherwise reads as either naive or dishonest to exactly the reader it is written for.
Neutral by construction
Authority, evidence and refusal live outside any one runtime. Each new entrant that governs only its own estate makes a layer that spans them worth more, not less.
assertedThe vocabulary is published
The entity model and authority grammar are Apache 2.0 and deliberately unbranded. RotaGrant is the reference implementation and the document says so — a specification with one implementation is a description with ambitions.
observed →Several jurisdictions, one workload
A product built for one country resolves one regime. This one resolves a profile per decision and seals it into the record, which is what makes "global" a mechanism rather than an ambition.
observed →What we would not claim
This does not make an agent safe, correct, or well designed. It does not evaluate a model, and it does not know whether the answer your agent gave was right. It states what an agent may do, refuses what exceeds it, and produces a record a stranger can check — and a vendor telling you it does more than that is selling you a feeling.