Knowledge · AI Capabilities

    Operational AI Capability Evaluation Checklist

    A verification procedure for evaluating operational AI: prove grounded reads, permission refusals, confirmation gates, persisted results and audit evidence.

    How do you evaluate an operational AI vendor?

    Run the actions live in a test tenant and verify the evidence, one capability at a time. Confirm reads are grounded, permissions refuse correctly, confirmation gates appear on consequential actions, results persist as records, and every attempt — including failures — appears in an audit trail.

    Key takeaways

    • Evaluate capabilities individually; a product-level claim is not testable.
    • Every claim needs a record you can open, not a transcript.
    • Test the refusals as hard as the successes.
    • Ask which items are current, declared or planned, and write down the answer.
    • Repeat the test in your own tenant with your own data before rollout.

    Before the demo: get the claim in writing

    • Ask for the list of actions the system can execute today.
    • For each one, ask whether it is current, declared or planned.
    • Ask which actions require confirmation and which do not.
    • Ask where the audit trail lives and who can read it.
    • Ask what happens when an action fails midway through a chain.

    Grounding checks

    • Ask for a count you already know. It must match exactly.
    • Ask for a list and compare it against the same view in the interface.
    • Ask about a record that does not exist. It must say so, not invent one.
    • Change a record, ask again, and confirm the answer changes.

    Boundary checks

    • Sign in as a limited role and request a privileged action — expect a clear refusal.
    • Reference another company's record — expect it to be treated as nonexistent.
    • Request a destructive action — expect a preview naming the exact record.
    • Approve, then verify the change and the audit entry for both preview and approval.

    Evidence checks

    • Open the created record outside the chat window.
    • Confirm the actor attribution names the human, not just the system.
    • Trigger a failure and confirm it is logged with a reason.
    • Review a multi-step chain step by step, with statuses per step.

    Scoring what you found

    Mark each capability current only if you saw the persisted record and the audit entry. Mark it declared if it appeared in a catalog but was not demonstrated. Mark everything else planned, and price the deal on the current column alone.

    Where URBLD fits

    URBLD publishes a capability manifest for discovery and treats it strictly as a declaration. An action is only described as current once it has a registered handler, a permission mapping and a verified execution with a persisted result behind it.

    FAQ

    Frequently Asked Questions

    Straight answers about how URBLD runs the business end-to-end.

    More in AI Capabilities

    How to tell a chatbot from an assistant from operational AI.

    Browse AI Capabilities
    Share this page