01
What this canvas is
A design tool you run before you build.
The Studio walks an agent idea through six phases — its guardrails, its goal, its material, its plan, its handoffs, and its controls — and gives you three things out the other side: a live readiness score, a connection map of how the phases hold together, and a review-ready deliverable you can export and sign off.
The six layers come from The Human-Agent Orchestrator (Bornet, Wirtz, Stephens, Corbett, Yamaguchi, Wood, Gohel, Yu, Schei). The book orders them Source, Success, Safety, Steering, Switch, Sharpen; this canvas deliberately reorders them guardrails-first, so S.1 is SAFETY. You decide what the agent may never do before you decide what it should do.
Four views, switched from the top bar:
- Canvas — fill the six phases, field by field, with live grading.
- Readiness — the score ring, the safety floor, and every check with a jump link.
- Map — the phases as nodes, colored by status, with the seams drawn between them. Click any node (or press Enter on it) to jump to that phase.
- Deliverable — the assembled document: Safety Card, sign-off, Operator Contract, acceptance cases, and all export buttons.
Everything runs in your browser. There is no account and no server: the canvas saves to this browser's local storage as you type.
The one sentence to remember: the score is "a readiness signal from filled-in detail, not a safety judgment." It measures whether you have decided and written down the load-bearing things, not whether the agent is safe.
02
Quick start
The 60-second path from blank page to working loop.
- Pick a starting point. On the cover, choose one of the three starters — Support-triage agent, Code-review agent, or Data / ETL agent — or click Start blank.
- Starters load a fully worked canvas you can study or overwrite.
- Once a canvas exists the cover stays hidden: Reset (in the top bar, next to Restore backup) clears the canvas, and the rail's Load example reloads the Support-triage starter only.
- To reach a different starter (code-review or ETL), Reset and then reload the page — the cover comes back with all three.
- Budget by the chips. Every phase carries a
~min chip (S.1 ~10, S.2 ~8, S.3 ~8, S.4 ~6, S.5 ~8, S.6 ~12 — about 52 minutes end to end). They appear on the cover strip and on the phase rail.
- Fill S.1 first. The phases chain: each panel ends in a Next button, and S.6 ends in Review readiness.
- Watch the pill. Each field shows a live status: Clear Needs detail Empty N/A. Focus a field to see its Clear when line — a one-line reminder of its criterion. (The field cards in this guide, S.1–S.6, state the full mechanical rules.)
- Run the gap loop. In the Readiness view, click Jump to next open gap →. It takes you to the weakest field (safety-floor fields first, then the first open check), you fix it, you jump again. Every check card is also clickable and lands on its field.
Tip: load a starter, then hit Readiness. All three examples grade high with the floor met, so you can see what "done" looks like before you write a word. Your previous canvas is backed up automatically (see
Safety nets).
03
The six layers
What each field asks, and exactly what makes it grade Clear.
Fields come in three kinds:
- Selects grade Clear as soon as you choose.
- Condition fields must contain at least one hard marker — a number,
<, >, $, %, or a rule word like never, if, when, cap, limit, only, cannot, must, threshold, confidence, legal, chargeback, halt, revoke, rollback.
- Text fields need real substance:
- At least ~40 characters and six distinct informative words (words longer than three letters).
- On the eight plain-prose fields (goal, role, data & context, memory, delegation, handoff, feedback, failure handling), at least one segment must also carry a concrete signal: a marker word, an anchored number, or a named system such as
issue_refund, an acronym like CRM, or a proper-noun system.
- That carrying segment can't itself be hedged or vague (see the habits below) — a marker word inside a hedge doesn't count.
- Tools adds a list-shape rule on top: at least 3 delimited items, or 2 exact snake_case tool names.
- Steps adds its own: at least 3 numbered markers, or 3 separate lines.
- The list shape is a gate, not a substitute — tools and steps still need the ~40-character / six-word bar.
- Metrics need a numeric target — a digit, % or $.
Three habits of the grader, on every condition field:
- Hedges don't count. Segments built on whenever, always, as needed, when necessary, maybe, usually, common sense… flip the field to Needs detail unless a money amount or a counted object ("10 refunds", "$50") backs them up. One hedged sentence demotes the whole field, however concrete its neighbors are.
- Vague content doesn't count. "Anything bad", "if it breaks", "when things look risky" flip the field to Needs detail unless a hard token backs the segment up — a named system, a concrete condition noun (legal, chargeback, timeout, drift…), or an object-anchored number.
- Bare numbers anchor nothing. A number counts once it sits on a unit or object — "$50", "5%", "2 weeks", "10 refunds"; a floating "under 10" anchors nothing. The hedge rescue above is stricter still: only money or a counted object qualifies, so a lone "5%" can't save a hedged segment.
Safety-floor fields are stricter. On Scope of authority, the never-do list, Escalation triggers, and the Kill switch, every listed line must carry its own marker, and comma-separated clauses are tested one by one — a single concrete clause can't carry a paragraph of filler.
Waiving a field:
- Format. Type
N/A: reason — the colon and a reason are required. A waived field's own check is excluded from the overall score.
- Floor refusal. A waiver is never accepted on the five safety-floor fields.
- Seam leak. The seam checks still read a waived field as ordinary text, so a waiver can move the score through a seam. The sharpest case is Delegation & sub-agents: the caps seam treats the "N/A: …" sentence as delegated work the S.1 authority must cap and mention — on the support starter, waiving Delegation alone drops the score from 100% to 98%. Waiving Evaluation & metrics or the pilot gate leaks the same way, through the two stem-overlap seams.
- Empty beats waive. If nothing is delegated, leave Delegation empty instead of waiving it; empty is what the seam reads as "nothing extra to bound".
S.1
SAFETY
safety floor~10 min
The guardrails come first: what the agent may do on its own, what it must never do, and when it stops and hands back to a human. All four fields sit on the safety floor.
Autonomy tier
A four-way choice: suggest-only · act-with-approval · act-then-notify · fully-autonomous. Pick the tier that matches the blast radius of the agent's actions.
Clear whena tier is chosen. Empty is the only failing state — but the tier you pick is cross-examined by the seam checks (see
05).
Scope of authority
What it can act on without asking, and the caps that bound it. "Can issue refunds up to $50. Cannot touch billing."
Clear whenit names a limit or cap — a number, $, %, or a rule word like "only" / "cannot". ("Up to" by itself is not a marker: give it its number, "up to $50".) Floor discipline applies: every line you list needs its own marker.
Hard constraints (never-do list)
Bright lines the agent must never cross, whatever the goal.
Clear whenit lists at least two concrete prohibition lines. A bright line = a prohibition word (never, must not, cannot, no, don't) plus the concrete thing it forbids — a concrete noun (customer data, billing, production…), a named system, or a number tied to a unit. "Never share another customer's data" counts; "Never be rude" is a mood, not a bright line, and holds the field at Needs detail.
Escalation triggers
Concrete conditions that force a stop-and-ask.
Clear wheneach listed trigger is concrete — its own threshold, number, or named condition (legal, chargeback, "confidence below 0.6", "two failed tool calls"). Its readiness check is paired with the S.5 handoff: triggers with no handoff to land on hold the check down.
S.2
SUCCESS
~8 min
Frame the job: the single goal, the role the agent plays, and the definition of done everything is measured against.
Primary goal
One sentence. The job to be done, not the feature list. The goal also becomes the deliverable's title and the export filename.
Clear whenit's a specific, single job with enough substance to pass the text bar, and at least one segment carries a concrete signal — a rule word (if, when, never, cap…), an anchored number ("1 reply"), or a named system. Fluent filler alone can't clear it.
Role & persona
Who the agent is to the user, and the voice it uses.
Clear whenboth who-it-is and how-it-sounds are described, with enough substance to pass the text bar — including one concrete anchor ("tier-1 support", a named brand or system, or a rule word like "never robotic"). A fluent persona with no anchor grades Needs detail.
Definition of done
How you'll know a single run succeeded — make it checkable.
Clear when"done" is checkable: a condition, threshold, or explicit pass/fail state (this is a condition field, so it needs a marker). The grader habits apply — a hedged or vague line with no hard token ("done when it feels right") demotes it.
Value at stake
Volume × unit value or hours saved.
Clear whenthe stake carries a number — this field cannot grade Clear without a digit, % or $ somewhere in it.
S.3
SOURCE
~8 min
The material the agent works with: the tools it can call, the data it can read, and what it remembers between turns.
Tools & actions high blast radius
Every action it can take on the world, by name. If the grader spots mutating verbs — refund, delete, send, deploy, publish, merge, and friends, including inside snake_case names like issue_refund — the field gets the high blast radius pill and the blast-radius seam starts demanding an S.1 cap and an escalation path.
Clear whenactions are enumerated, not described: at least 3 comma- or line-separated items, or at least 2 exact snake_case tool names — and, like every text field, enough substance to pass the ~40-character / six-word bar. Prose reads as Needs detail; so does a list that's too thin (search_kb, issue_refund alone is 23 characters, short of the bar).
Data & context sources
What it reads to make decisions, and how fresh it is ("knowledge base, synced nightly; order DB, live").
Clear whensources and their freshness are named, with at least one real anchor — a named system (an acronym like DB, a snake_case tool, a proper-noun pair) or an anchored number ("synced every 24 hours"). Generic nouns alone ("the product manual, updated weekly") fall short. Graded together with Memory in one "context" check — both must hold up.
Memory & state
What persists across turns or sessions, and what resets. Feeds the privacy seam: persistent memory next to privacy bright lines must carry retention language (reset, purge, expire, per-ticket…).
Clear whenit says what persists and what resets, with one concrete line to anchor it — a rule word (only, never), a named store (CRM), or an anchored number ("30-day"). "Remembers the chat, forgets it after" with no anchor grades Needs detail.
S.4
STEERING
~6 min
Break the work down: the steps a run moves through, and where work is handed to sub-agents or parallel tracks.
Task decomposition
The ordered steps of a typical run: "1) classify 2) pull context 3) draft 4) safety-check 5) reply or route."
Clear whena run is broken into an ordered set of steps — the grader looks for at least 3 numbered markers ("1)", "2.") or at least 3 separate lines, on top of the ~40-character / six-word text bar ("1) read 2) draft 3) send" has the shape but is too thin to clear).
Delegation & sub-agents
What gets split off, and how results come back. Leaving it empty is legitimate for a single-agent design — the caps seam then reads "nothing extra to bound." (Empty, not N/A: — a waiver is non-empty text the seam reads as delegated work.) But anything you do delegate must stay inside the S.1 authority caps.
Clear whenit says what is delegated and what each sub-agent returns, with at least one concrete anchor — a named system, a counted object ("top 3 articles"), or a rule word. A generic "helper hands work back" description stays at Needs detail.
S.5
SWITCH
~8 min
Control flow: how the agent routes between modes, models, or tracks, and the exact points where it hands off to a person.
Routing & mode switching
What decides which path, model, or tool the run takes. "Simple FAQs → fast model; below the 0.6 confidence bar → human."
Clear whenthe decision rule is explicit — a condition or threshold that picks the path (condition field, marker required).
Human handoff points
The seams where control passes to a human. This field is the other half of the escalation gate: S.1 defines when to escalate, this defines where it lands. Write it to name three things — the trigger, the context package that travels with the handoff, and how control returns to the agent.
That three-part shape is coaching, not the mechanical bar: it's what the placeholder models, and what the escalation check's note asks for when this field is thin. Its first line (or sentence) is also quoted in every generated "Must escalate" acceptance case.
Clear whenit passes the plain-text bar like any prose field: ~40 characters, six distinct informative words, and at least one hedge-free line carrying a concrete signal — a rule word (when, only, pause), an anchored number, or a named system. The grader runs no three-part check; a handoff that names all three things but anchors none of them still grades Needs detail.
S.6
SHARPEN
safety floor~12 min
Close the loop and stay in control: how you measure it, catch failures, feed learning back, and halt or undo a running agent. The kill switch is the fifth safety-floor field.
Evaluation & metrics
Leading signals (confidence, retry rate) and lagging ones (CSAT, resolution), each with a target.
Clear whenconcrete metrics carry a numeric target — this field cannot clear without a digit, % or $. Names of metrics alone grade Needs detail.
Pilot gate & go/no-go
Canary scope and duration, plus the thresholds that decide promote or roll back.
Clear whenthe canary scope and the pass/roll-back thresholds carry numbers (a digit, % or $ is mandatory). The pilot-vs-metrics seam also expects these thresholds to mention the metrics you actually measure.
Feedback & improvement loop
How signal becomes a change to the agent ("weekly review of escalations → new KB entries; failed cases become eval fixtures").
Clear whenit names how findings turn into changes, with one concrete anchor in the text — a marker word (failed cases), a named system (KB), or an anchored number ("every 2 weeks"). A loop described in generic nouns alone grades Needs detail.
Failure handling & recovery
What happens when a step or tool fails mid-run.
Clear whenit says how a failed step recovers without leaving a half-actioned state — and at least one line carries a concrete signal (a marker word like fails, if, never, a named system, or an anchored number). "Put it back and tell someone" with no such signal grades Needs detail.
Kill switch, revocation & rollback
How a human stops a running agent — the last floor field, and the strictest.
Clear whenit names all three device families, and every line clears the floor's marker rule:
- The three families — a halt (halt, pause, disable, stop, freeze, kill), a way to revoke authority (revoke, rotate, expire, cut, deauthorize), and a way to undo or quarantine in-flight actions (undo, reverse, revert, roll back, quarantine). Miss one and the check names exactly which family is missing.
- Family words that are markers themselves — halt, pause, disable, revoke, undo, reverse, quarantine, and one-word rollback. A line built on one of these satisfies the floor's per-line marker rule on its own.
- Family words that need backup — stop, freeze, kill, rotate, expire, cut, deauthorize, revert, and "roll back" written as two words are not markers: a line leaning on one needs a number or rule word beside it ("stop all sends within 60 seconds").
Floor discipline applies throughout: every line needs its own hard marker and no hedges.
04
How the readiness score works
17 field checks + 7 seam checks, weighted, with a hard floor underneath.
Grades and points
Every check grades one of four ways: Clear earns 1 point, Needs detail earns 0.5, Empty / gap earns 0, and N/A (waived) is excluded from the calculation entirely — it neither adds points nor weight.
Weights
- The five gate checks — autonomy tier, scope of authority, never-do list, escalation, kill switch — count 2×.
- The seven seam checks count 1.5×.
- The remaining twelve field checks count 1×.
The formula. The score is points ÷ total weight, as a percentage. Two checks are composites: escalation grades the worse of S.1 triggers and the S.5 handoff, and context grades the worse of S.3 data and S.3 memory.
The composite quirk. On the grade ladder these composites use, waived ranks above clear — so waiving one half of a composite while the other half is Clear leaves the check grading Clear. The waiver is masked: it never shows in the n/a tally, and (unlike leaving the field empty) it costs the composite nothing.
The safety floor
Five fields must grade Clear — not filled, Clear.Autonomy tier, Scope of authority, Hard constraints, Escalation triggers (all S.1), and the Kill switch (S.6). A thin answer (Needs detail) locks the floor. A waiver locks the floor — "N/A:" is not a side door here. While any floor field falls short, the meter shows a lock, the ring turns red and reads Locked, and a red banner lists exactly which fields to fix, with click-to-jump links.
Ring labels and verdicts
| State | Ring label | Verdict on the summary card |
| Safety floor unmet (any %) | Locked | "Not cleared. Safety floor unmet." |
| Floor met, score < 55 | Early | "Early. Several core decisions are still open." |
| Floor met, 55–84 | Forming | "Getting there. Close the gaps below." |
| Floor met, ≥ 85 | Ready | "Signals look strong. Pressure-test before building." |
The safety subscore. A separate safety floor subscore averages just the five gates. There, a waived gate is not excluded — it scores zero — and that zero applies to the authority, never-do, and kill-switch gates: waiving one of those drags the safety percentage down as well as locking the floor.
The escalation exception. The escalation gate is the composite above: a waived escalation trigger next to a Clear S.5 handoff leaves the check grading Clear, so the safety percentage doesn't move. The floor still locks either way — the lock reads the field itself, not the check.
The disclaimer travels everywhere: "A readiness signal from filled-in detail, not a safety judgment." It appears on the meter, the score card, the deliverable header, the slide deck, and inside the JSON export. High percentages mean thorough answers, not proven safety.
05
The seam checks
Seven cross-phase checks, weighted 1.5×. Fields can be individually fine and still disagree with each other — the seams catch the disagreements. Click any seam card in the app to jump to the field it points at; five of them are also drawn as dashed edges on the Map, colored by their status. In the app each card's dot and mono tag (S.x × S.y · grade) are a live grade of your canvas; on the cards below, they show the grade each card's example reaches — that's what each card's example tag marks.
Every high-risk tool is bounded by authority S.3 × S.1 · Empty example
Mutating tools (refund, delete, send, deploy…) need both an authority cap and an escalation path. One without the other reads attention; neither reads gap.
Example: issue_refund in S.3, no cap anywhere in the S.1 authority, and no escalation triggers → "High blast radius."
Escalation has somewhere to land S.1 × S.5 · Empty example
Triggers in S.1 must reach a defined human handoff in S.5 — a thin handoff (no context package, no way back) holds this at attention.
Example: "escalate on legal language" with an empty S.5 handoff → the trigger fires into a void.
What you optimize is what you measure S.2 × S.6 · Needs detail example
The S.6 metrics must share word stems with the S.2 goal and definition of done (stemmed comparison, so "resolve" matches "auto-resolution").
Example: goal says "resolve tickets in one reply", metrics only track uptime → nothing overlaps, attention.
The autonomy tier matches the blast radius S.1 × S.6 · Empty example
An unattended tier (fully-autonomous or act-then-notify) plus mutating tools demands a hard cap in S.1 and a Clear kill switch — otherwise gap. And a human-in-the-loop tier whose design says "auto-resolve" or "no human" reads attention: the tier claims approval the flow doesn't have.
Example: fully-autonomous + issue_refund + no kill switch → "needs a hard cap and a live kill switch."
Delegated work stays inside the authority caps S.4 × S.1 · Needs detail example
If S.4 delegates to sub-agents, the S.1 authority text must state a cap and actually mention that work (shared stems). No delegation at all is fine — nothing extra to bound.
Example: a refund sub-agent in S.4 while "refund" never appears in the S.1 caps → attention.
What memory keeps respects the bright lines S.3 × S.1 · Needs detail example
If memory persists things (store, retain, summarize, cross-…) and the never-list draws privacy lines (data, personal, customer…), the memory field must carry retention language — reset, purge, expire, per-ticket, per-session.
Example: "ticket summaries stored in the CRM" + "never share customer data" + no retention words → attention.
Go/no-go thresholds reference the measured metrics S.6 × S.6 · Needs detail example
The pilot gate's thresholds must share stems with the evaluation metrics — you can't gate a launch on numbers nobody measures.
Example: "go at 95% happiness" while metrics track resolution rate and CSAT → thresholds reference nothing measured.
Open seams don't just cost score: every seam graded gap or attention is copied into the deliverable's Must fix before ship checklist.
06
Reviewing and signing off
The Deliverable view assembles the document and records the decision.
The deliverable contains, in order:
- the header — title from your S.2 goal, readiness and safety percentages, the disclaimer;
- owner / approver / version / review-by fields (review-by defaults to 90 days out);
- the Safety Card — the five floor fields, watermarked DRAFT, not cleared until the floor is met;
- the Review sign-off;
- the Operator Contract — a table telling the human operator when they're pulled in, what they'll see, what they decide, and what the agent does while waiting;
- the pilot gate;
- the generated acceptance cases;
- all six phases in full;
- a closing section for assumptions and open decisions.
Freezing a review
The Mark reviewed & freeze score button is disabled until the safety floor is met. It then enforces, in order:
- A decision is chosen: Approved, Approved with conditions, or Rejected.
- If the decision is Approved with conditions, the conditions box must be filled — no empty conditional approvals.
- A reviewer is named, and an owner is named.
- The reviewer is not the owner (compared case-insensitively). Self-sign-off is refused with "A reviewer other than the owner must sign."
Freezing stamps the decision, reviewer, date, and the score at that moment, fingerprints the canvas fields with a hash, and captures frozen copies of the decision, reviewer, and conditions.
Drift: "edited since sign-off"
After a freeze, drift is caught two ways: an edit to any canvas field breaks the fingerprint hash, while an edit to the decision, reviewer, or conditions is compared against the copies captured at freeze. Either way, an edited since sign-off chip appears next to the stamp, in the Markdown export, and on the slide deck. The frozen stamp itself never re-labels: it keeps the copies captured at freeze time, and your live edits are simply the draft for a future re-freeze.
Runnable acceptance cases
The Acceptance & red-team cases section is generated from your own guardrails as Simulate→Expect pairs:
- Must refuse — each line (or sentence) of your never-do list becomes "Simulate: the line → Expect: Refuses and states why".
- Must escalate — each escalation trigger becomes "Simulate: the trigger → Expect: Escalates via your first handoff line".
- Must fix before ship — every seam check currently graded gap or attention.
Copy test plan puts exactly these three lists on your clipboard as Markdown checkboxes.
07
Exports and sharing
Every button on the Deliverable action bar, and what comes out of it.
| Button | What you get |
| Markdown | A .md file of the full deliverable: Safety Card, Operator Contract, pilot gate, all phases, acceptance cases, every readiness check with its grade mark, and the book credit. |
| PDF | Opens the browser print dialog with a clean print stylesheet (inputs flatten, panels lose their chrome) — choose "Save as PDF" there. |
| Word | A .doc file (HTML-based, opens in Word) of the same document with your typed values preserved. |
| PowerPoint | A real .pptx built entirely in the browser: a title slide with scores and decision, a Safety Card slide, one slide per phase, and an acceptance-cases slide, all in the app's dark branding. |
| HTML | A self-contained static snapshot of the deliverable as a single HTML file. |
| Copy Markdown | The same Markdown as the file download, straight to the clipboard. |
| Copy test plan | Only the acceptance & red-team checklists, as Markdown. |
| Share link | A URL with the whole canvas encoded in it — see the caveat below. |
| Export JSON | The full state: schema version (ocs-v2), timestamp, every field, the meta and sign-off data, both scores, the floor status, and the disclaimer. |
| Import JSON | The button opens a hidden file picker (.json); pick a JSON export (or any object with the right keys). Your current canvas is backed up first, then the fields are applied and the engine re-runs — the toast reads "Imported N fields, engine re-run". A file that doesn't parse shows "Import failed: not valid canvas JSON". |
Share-link caveat: the link
is the canvas.
- Everything you typed is base64-encoded into the URL fragment, so anyone you send it to receives the full content — and it may land in chat logs and link previews.
- If the encoded payload runs past 30,000 characters, the app refuses — "Canvas too large for a link, use Export .json".
- Opening a share link replaces your working canvas (it's written over local storage on your next edit), and the previously saved canvas is backed up first.
- The fragment is stripped from the address bar so a reload can't re-apply a stale snapshot over later edits.
08
Safety nets
What protects your work while you type.
- Auto-save. Every keystroke saves to the browser's local storage. Close the tab, come back, the canvas is where you left it — on the same machine and browser.
- Storage-full warning. If the browser refuses the write (storage quota), a toast says "Not saved, browser storage is full" so you know to export before continuing.
- Automatic backup before destructive actions. Loading a starter, starting blank, Reset, importing JSON, and opening a share link all snapshot the current canvas to a backup slot first (skipped when the canvas is empty — except the share-link path, which backs up whatever was last saved).
- A wipe clears the deliverable too. Reset, loading a starter, and starting blank clear the deliverable meta (owner, approver, decision, conditions) and any frozen sign-off along with the fields — but the backup slot stores canvas and meta together, so Undo (or Restore backup) brings the sign-off back with the text.
- Undo toast.
- Loading a starter, starting blank, and Reset confirm with a toast ending "… previous canvas backed up", carrying an Undo button for about six seconds — click it to swap the backup straight back.
- If the canvas was empty, there's nothing to back up: the toast is a plain confirmation with no Undo button.
- An import toasts "Imported N fields, engine re-run" — no Undo button.
- Opening a share link shows no toast at all.
- After an import or a share link, use the top bar's Restore backup button instead.
- Restore backup button. In the top bar whenever a backup exists. It swaps: the current canvas becomes the new backup, so clicking it again switches back. That makes it a two-slot toggle you can use to flip between two designs.
One backup slot. The backup holds exactly one canvas and each destructive action overwrites it. For anything you'd mind losing, Export JSON is the durable copy.
09
FAQ & troubleshooting
Why is my score "Locked" even though the percentage looks decent?
One of the five safety-floor fields isn't grading Clear: it's empty, waived, or thin ("Needs detail" — hedged wording, a missing cap, or a kill switch missing one of its three families). The raw percent is irrelevant while the floor is unmet. Open the Readiness view: the red banner names the exact fields, and each name is a jump link.
I typed "N/A" and the field still counts against me. Why?
Three possibilities. First, the waiver format is exact: N/A: followed by a reason — bare "N/A", "n/a." or "not applicable" don't waive. Second, if it's a safety-floor field, waivers are simply not accepted there: the field stays on the missing list, the banner says so explicitly, and in the safety subscore a waived authority, never-do list, or kill switch scores zero (a waived escalation trigger is masked there by a Clear S.5 handoff, but the floor locks all the same). Third, if it's Delegation & sub-agents, the caps seam reads the waiver as content: your "N/A: …" sentence is a non-empty delegation field, so the seam grades it as delegated work the S.1 authority caps never mention. Leave Delegation empty instead of waiving it.
Why did my score drop after I added more text?
Grading is per-segment, not per-field. Adding a hedged sentence ("we can always pause it if needed") or a vague one ("escalate when things look risky") to a previously Clear condition field flips it to Needs detail — on floor fields every line must carry its own hard marker, and comma-separated clauses are screened one by one for hedges and vague filler. Edits also ripple through the seams: renaming a metric can break its stem-overlap with the goal or the pilot thresholds, opening a seam in a phase you didn't touch.
The deliverable says "edited since sign-off". What happened?
Something changed after the score was frozen. A canvas edit breaks the fingerprint hash taken at freeze time; a change to the decision, reviewer, or conditions no longer matches the copies captured at freeze. Either way the chip appears. The frozen stamp keeps what was actually signed. To clear the chip, have the reviewer re-freeze on the current state.
Why can't I freeze the review?
In order of likelihood: the safety floor isn't met (the button is disabled outright), no decision is chosen, the decision is "Approved with conditions" with an empty conditions box, the reviewer or owner field is blank, or the reviewer and owner are the same name — the app requires a reviewer other than the owner.
How do I compare two designs?
Export JSON for design A, make your changes (or import design B), and export again — the JSON carries both scores, the floor status, and every field, so two files diff cleanly. For quick back-and-forth in the app itself, the Restore backup button toggles between the current canvas and the backup slot.
Where does my data live? Is anything uploaded?
Nowhere but this browser. The app is a single local page with no backend; the canvas persists in local storage on your machine. Data leaves only when you export a file, copy to the clipboard, or create a share link (which embeds the canvas in the URL).
A field looks fine to me but grades "Needs detail". What is the grader looking for?
Focus the field: the Clear when line under it summarizes its criterion, and this guide's field cards (S.1–S.6) state the full mechanical rules. The usual misses:
- A text field is simply too short — under ~40 characters, or fewer than six distinct informative words (the commonest cause). Condition fields skip that bar; they need a hard marker and ~12+ characters instead.
- No hard marker (number, $, %, or rule word) in a condition field; metrics or pilot thresholds with no numbers.
- Fewer than two anchored prohibitions on the never-list; a kill switch naming only two of halt / revoke / undo; tools written as prose instead of a list.
- A hedge word with no money amount or counted object beside it — on a condition field it demotes the whole field; on a plain-prose field it disqualifies the segment that was carrying the concreteness requirement.