Assess the agent or workflow, not the vendor as a whole. One sheet per agent.
What it does
- Purpose and the workflows it runs
- Inputs it reads, outputs it produces, actions it executes in which systems
- Autonomy level as documented: assistance, copilot, action, workflow, autonomous
- Human approval: where required, where optional, where absent
Control
- Audit log of actions: exists, exportable, retention
- Explainability: can a user see why an action was taken
- Guardrails: limits, allow-lists, rollback
Data
- Which of our data is sent to the model, and to whose model
- Training on our data: yes/no, contractual basis
- Data residency and retention
Evidence
- Documentation links for 1 to 10
- What we were able to verify ourselves, and how (demo, sandbox, reference)
Outcome
Fit for <use case>: yes / with conditions / no, with the evidence that decides it.