Understand the system
Read what an agent did and why, well enough to say whether it should have. Assessed against real output from the candidate's own domain, not a generic scenario.
Drawn from: the Evidence and Composition cells.
The canvas is free and deliberately uncalibrated. Where the line falls between thin and solid depends on your regulatory regime, your workflow, and what it costs you to be wrong, so we do not set it centrally. That is not an omission with a paywall behind it. It is the reason the rungs below exist at all: calibration is how the method completes, and it is local work that somebody has to do.
Everything on this site. Free to use, teach, translate, and adapt under Creative Commons, including by people who will compete with us using it.
Anyone
Live · available now
PriceFree, permanently
A multi-year, co-created project rather than a launch. Written in public, with contributors named on the entries they authored.
Practitioners
Building · in progress
PriceOne-off, and mostly it is marketing
Self-paced: how to set anchors for your own regime, score twelve cells against one workflow, and run the disagreement rather than average it away.
Team and workflow leads
Planned · not built
PricePer seat, low
For people who will run this inside organizations that are not their own. Contributors go first when this opens, at a published revenue share.
Consultants, internal capability teams
Planned · not built
PricePer seat, with renewal
Proof that a named person is qualified to oversee agents doing regulated work. This is the business. Everything above exists to make this the obvious standard rather than one vendor's certificate.
Regulated employers, defense suppliers
Building · substrate exists
PricePer candidate, proctored
Other training organizations delivering the credential under license. A credential only one body can issue is a product; one many can issue is a standard.
Training providers, assessment bodies
Planned · not built
PriceLicense and revenue share
Setting the anchors with you, running the canvas inside an enterprise, designing the gate into an existing control environment, and mapping the Evidence cell to the regimes you actually live under. This is the local work the uncalibrated scale requires.
Enterprises, primes, regulated operators
Live · available now
PriceEngagement
Rung zero is real and free. Rung six is real, because facilitation is a thing people already do. Rung four has a substrate rather than a product: the judgment taxonomy is built and reviewed, and the Euler Center already holds a Pearson VUE Test Center designation with over 1,500 tests administered, so proctoring is not a dependency.
Rungs one, two, three, and five are not built. Saying so costs a little credibility today and saves a lot later, and it is the same standard the rest of this site applies to Klarna and Bayer.
Three capabilities, assessed by simulation rather than by quiz, and mapped to the same requirement several regimes now impose: a person who can understand the system, intervene in it, and halt it. The fourth item below is what is not solved yet.
Read what an agent did and why, well enough to say whether it should have. Assessed against real output from the candidate's own domain, not a generic scenario.
Drawn from: the Evidence and Composition cells.
Recognize the exception, override without waiting for permission, and know the difference between a bad output and a bad frame. This is where automation bias is tested directly.
Drawn from: the Authority cell, and the failure in the Klarna entry.
Stop a running process and put the prior state back. Assessed by doing it, not by describing it, because a rollback nobody has executed is a hypothesis.
Drawn from: the gate's sixth item.
There is no assessor cohort and no defined pass mark. The taxonomy and the test center exist; the people qualified to judge oversight competence do not, because nobody has said what good looks like at what level.
This is the open gap, and it is the one thing that cannot be borrowed from an existing standard.
Nothing here assumes you have read the books. Every underlined name is a link to an entry that explains it in plain language and says where the claim came from. These are claims about agents, not first principles about work; those are here.
It perceives, plans, acts, observes, and goes again. That is the same shape as the improvement cycle and the military decision loop. Composing work out of agents means composing loops, not filling boxes on an org chart.
Take the human out of the critical path and throughput stops binding. What binds instead is judgment and accountability at the edge: who framed the problem, who catches the exception, who is answerable. That is what the twelve cells are for.
Ashby's Law says a controller needs at least as many responses as the situation has surprises. An agent can own a loop only where its variety matches. Whatever is left has to land on a person deliberately, or it lands on one by accident.
Jaques measured roles by how far ahead the longest task reaches before feedback arrives. Agents take seconds to days. People keep the framing and the accountability for outcomes nobody sees for a year, which is a different job from the one most managers were promoted into.
The canvas and the gate, and calibration engagements.
The book, and the assessed credential on an existing substrate.
The audit product, facilitator training, and the assessor network.