The failure pattern is now familiar. An agent capability is scoped as "assistive," ships as "semi-autonomous," and by the second quarter is executing transactions that the original PRD never contemplated. There is no malice in this drift, only the absence of a classification schema that would have made the change visible.
Enterprise-ready organizations classify agents on a four-tier scale before any deployment: retrieval-only, suggestion, supervised action, and delegated action. Each tier carries a distinct evaluation surface, escalation path, and audit expectation. The tier is declared in the PRD, reviewed by security and legal, and re-attested on any material change to tools, prompts, or context.
The classification matters most at the transitions. Moving from suggestion to supervised action introduces a new class of reversibility question: can the human reviewer meaningfully understand what the agent proposes to do, and can the action be undone without customer-visible harm? Moving from supervised to delegated action introduces liability questions that require executive sponsorship, not product management sign-off.
In practice, the tier taxonomy also disciplines the roadmap. It becomes far harder to promise autonomous behavior in a sales conversation when the organization has agreed, in writing, that delegated action requires named executive sponsorship and a documented kill-switch. That friction is the point.