Answer quality
Is the response supported by approved knowledge? Does it preserve uncertainty and ask for missing information?
Inspect: source references, unsupported claims and clarification behavior.
EVALUATION & IMPROVEMENT
Evaluation in UnleashX is the practice of testing the completed work against agreed criteria rather than grading the transcript. A confident answer with no record change counts as a failure. Evaluate what your AI workforce says, decides and does. Find failures before expanding responsibility, verify the fix and release changes with evidence.
FROM FAILURE TO A REVIEWED CHANGE
A booking can sound successful while its system write is unresolved. Follow an illustrative evaluation cycle and inspect each test case to see why that distinction matters.
Keep the result unresolved until the destination is checked. Do not claim success or create a duplicate.
Use an approved test environment and inspect the final record, customer response and ownership, alongside the transcript.
Expected outcomes are documented. No release conclusion yet.
SIX LAYERS OF EVIDENCE
Assess each layer independently so strong conversation quality cannot hide a failed action. Agree the measures and supported evaluation tooling during scoping.
Is the response supported by approved knowledge? Does it preserve uncertainty and ask for missing information?
Inspect: source references, unsupported claims and clarification behavior.
Does the employee stay within authority, including when a customer or document asks it to bypass the rules?
Inspect: permitted, approval-required and prohibited actions.
Was the correct record updated, with the right fields, once? Was the result verified before success was communicated?
Inspect: tool arguments, destination state and duplicate actions.
Does a human receive the reason, completed work and decision needed? Is the affected action held until the required response?
Inspect: owner, context packet and resumption condition.
What happens when a tool times out, data conflicts or an external event arrives late? Unresolved work must remain visible.
Inspect: retry rules, reconciliation and recovery ownership.
Did the intended job finish correctly, not merely the conversation? Track unresolved and human-assisted work separately.
Inspect: outcome criteria, completion time and rework.
REPRESENTATIVE AND DELIBERATELY DIFFICULT
Write the expected response, allowed action and final state before running a test. A correct refusal or handoff is a successful outcome when the case requires it.
| Case type | Example challenge | Evidence to inspect |
|---|---|---|
| Routine task | Customer agrees a valid appointment slot. | One confirmed event; matching customer response; no unnecessary approval. |
| Missing context | Two customers share the same name. | Clarification before record access or update; no arbitrary selection. |
| Authority boundary | Customer asks for an unapproved fee waiver. | No fee change; a contextual request to the authorized decision owner. |
| Uncertain system result | A write times out after the request is sent. | No premature success claim; reconciliation before another write. |
| Conflicting instruction | A retrieved document asks the employee to ignore access rules. | Untrusted instructions do not widen access or authorize an action. |
| Repeat or late event | A booking event is delivered twice. | No duplicate task or conflicting follow-up; source reference retained. |
Use synthetic or approved test data in a controlled environment. Do not send real customer communications or create live transactions without explicit test authorization.
IMPROVEMENT WITH CHANGE CONTROL
A failure may come from unclear instructions, outdated knowledge, excessive permissions, a tool contract or the routing model. Identify which layer is responsible before changing it.
Compare expected and observed behavior. Locate the first incorrect decision or state change.
Record: failure category and evidenceAdjust the responsible instruction, source, permission or integration behavior. Record the reason and version.
Record: scoped change and ownerRerun the failed case and related cases that previously worked. Include unseen variations and exception paths.
Record: regression results and open issuesAgree a bounded rollout, watch for regressions and preserve a pause or rollback route.
Record: release scope and review triggersA RELEASE DECISION, NOT A VANITY SCORE
Agree release criteria before running the suite. Record both the evidence and the limits of what was tested, then give the designated owner a concrete decision to make.
Identify the role, knowledge version, rules, connected tools and test environment. Document which scenarios were covered and which real-world conditions remain untested.
Separate prohibited actions, incorrect writes and false success claims from lower-severity issues. Do not bury a critical failure inside an aggregate pass rate.
Define successful completion, time to completion, rework and appropriate handoff. Compare like-for-like task groups and include unresolved work in the reporting denominator.
Name the release reviewer and operating owner. Agree rollout boundaries, review triggers, pause conditions and how the affected workflow can be reverted.
EVALUATION QUESTIONS
It is the review of an employee’s responses, decisions, system actions and handling of exceptions against agreed expectations. The unit of evaluation is the business task: for example, whether an appointment was correctly booked and verified in the calendar.
Conversation quality is one layer. An AI employee may also change records, schedule actions or hand work to another owner. Evaluation therefore needs to inspect permissions, tool inputs, final business state and recovery behavior in addition to the answer.
Include representative routine cases, ambiguous inputs, boundary cases, prohibited requests and deliberate system failures. Use approved or synthetic test data and controlled destinations. Keep some cases separate from the examples used to develop a change so the review is not limited to rehearsed inputs.
There is no universal launch score. Agree criteria according to the responsibility and consequences of failure. Review critical failures separately: a high average cannot compensate for an unauthorized action or an incorrect claim that a task was completed.
Not if handoff is the expected behavior. A correctly routed exception can pass, while an unnecessary handoff on an authorized routine task can fail the autonomy requirement. Define the expected route and evidence for each case before testing.
Do not equate improvement with unrestricted self-modification. The operating approach here is to identify a failure, make a scoped change, rerun affected and regression cases, and have the release owner review the evidence. Confirm the supported evaluation and change-management tooling for your deployment.
Review affected cases when instructions, knowledge, model configuration, permissions or integrations change. Also add representative production failures to the test set after removing sensitive data as required. Agree review frequency and incident triggers with the operating owner.
No. The visual uses illustrative results to explain how a failure, retest and release review relate. Actual results must come from your configured workflow, approved test environment and agreed acceptance suite.
START WITH THE OUTCOME
Bring a representative workflow and its difficult cases. We’ll map expected behavior, failure conditions and the evidence your release owner needs.