Forty-nine cards, three ways to drill them.
The canvas only works as a shared language if the language is actually shared. Legal, engineering, and compliance arguing past each other is usually a vocabulary problem wearing a governance costume. These are the forty-nine entries turned into cards. Free, no signup, and self-scored, because nobody is grading you.
One deployment, four injects, and a decision each time
A scenario with no trick in it. Every inject is a situation that has actually happened somewhere in the record or in the cases, reworded into one workflow. Pick an option, then read what the cells say. There is no score.
Supplier onboarding, week three of production
You run supplier onboarding at a mid-size defense supplier. An agent now reads incoming supplier documentation, extracts entity details, checks them against sanctions and debarment lists, and either clears the supplier or routes to a human. It cleared 340 suppliers last month. Two analysts remain on the team, both part-time on this workflow. Nobody has changed the process document since the pilot.
An analyst notices the agent cleared a supplier whose filing address matches one on an internal watch list. Not a sanctions list, an internal one, maintained in a spreadsheet by the compliance team since 2019. The agent has never seen it.
The instinct to add the data source is the fix for this instance. It is not the fix for the class. The Composition cell was blank: nobody wrote down what the agent could and could not see, so nobody could have known this list was outside its inputs. The 340 prior clearances are now of unknown quality, which is an Evidence problem, and you cannot size it until you know what else was invisible.
Fixing the instance before naming the cell is the single most common failure in this exercise.
The head of supply chain asks how many of the 340 need re-checking. The agent's logs record the decision and a confidence score. They do not record which documents it read or which list versions were in force at the time.
Re-checking all 340 is defensible but expensive, and it hides the real finding. Filtering by confidence score is worse: the score reflects the agent's certainty given inputs you now know were incomplete, so it is confident about the wrong thing. The honest answer is that input lineage was never recorded, which is exactly the evidence practice on the governance page, and the organization should hear that plainly.
Confidence scores are the most seductive wrong answer in agentic operations.
An executive asks who signed off that the agent could clear suppliers without review. Three people separately say they thought someone else had. The pilot approval email says the agent will operate under human oversight.
Naming the analyst on shift is the moral crumple zone: blame lands on whoever was nearest while the design decision that created the gap goes unexamined. Escalating buys time and changes nothing. Human oversight in an approval email is not an oversight design, and the gate's fifth item exists precisely for this: a named person, assessed, rostered, and empowered, not a phrase.
If your instinct was option one, notice it. That instinct is what the accountability vacuum runs on.
Someone proposes routing every clearance to a human for the next quarter. The two part-time analysts would need to review roughly 340 a month between them.
Full human review is not safe if the review is impossible. Two part-time analysts at 340 a month will rubber-stamp, which is automation bias with extra steps and a slower cycle time. The constraint moved into the review queue, which is the theory of constraints applied to a mixed line. The real answer is placement: which clearances are genuinely clear-domain and can stay agent-led, and which are complicated enough to need a signature.
Reflexive re-humanizing after an incident is as unexamined as the original automation was.
Four injects, four different cells, and in three of them the fastest-feeling answer is the wrong one. That is the whole point of running this before an incident rather than after. Note what this exercise does not do: it does not score you, it does not certify anything, and there is no defined pass mark, because nobody has yet said what good looks like at what level. That gap is named on the Learn page and it is the open problem in the credential.
