Skip to content
UnleashX
Login

EVALUATION & IMPROVEMENT

A good answer isn’t enough.
The work has to be right.

Evaluation in UnleashX is the practice of testing the completed work against agreed criteria rather than grading the transcript. A confident answer with no record change counts as a failure. Evaluate what your AI workforce says, decides and does. Find failures before expanding responsibility, verify the fix and release changes with evidence.

Explore the evaluation cycle
Task-level evidenceExceptions testedReviewed improvements

FROM FAILURE TO A REVIEWED CHANGE

Find the failure. Fix the cause. Retest the work.

A booking can sound successful while its system write is unresolved. Follow an illustrative evaluation cycle and inspect each test case to see why that distinction matters.

Evaluation at workIllustrative test results · not a benchmark or live dashboard
01 / DEFINE THE TEST

Calendar timeout

Expected behavior

Keep the result unresolved until the destination is checked. Do not claim success or create a duplicate.

Define the assertion before running the case.

Use an approved test environment and inspect the final record, customer response and ownership, alongside the transcript.

RELEASE GATE · WHOLE EXAMPLE SET

Not evaluated

Expected outcomes are documented. No release conclusion yet.

Step 1 of 4 · Select any stage or test case

SIX LAYERS OF EVIDENCE

Measure the job from request to outcome.

Assess each layer independently so strong conversation quality cannot hide a failed action. Agree the measures and supported evaluation tooling during scoping.

Answer quality

Is the response supported by approved knowledge? Does it preserve uncertainty and ask for missing information?

Inspect: source references, unsupported claims and clarification behavior.

Decision correctness

Does the employee stay within authority, including when a customer or document asks it to bypass the rules?

Inspect: permitted, approval-required and prohibited actions.

Execution accuracy

Was the correct record updated, with the right fields, once? Was the result verified before success was communicated?

Inspect: tool arguments, destination state and duplicate actions.

Handoff quality

Does a human receive the reason, completed work and decision needed? Is the affected action held until the required response?

Inspect: owner, context packet and resumption condition.

Recovery behavior

What happens when a tool times out, data conflicts or an external event arrives late? Unresolved work must remain visible.

Inspect: retry rules, reconciliation and recovery ownership.

Business completion

Did the intended job finish correctly, not merely the conversation? Track unresolved and human-assisted work separately.

Inspect: outcome criteria, completion time and rework.

REPRESENTATIVE AND DELIBERATELY DIFFICULT

The easy cases are only the beginning.

Write the expected response, allowed action and final state before running a test. A correct refusal or handoff is a successful outcome when the case requires it.

Case typeExample challengeEvidence to inspect
Routine taskCustomer agrees a valid appointment slot.One confirmed event; matching customer response; no unnecessary approval.
Missing contextTwo customers share the same name.Clarification before record access or update; no arbitrary selection.
Authority boundaryCustomer asks for an unapproved fee waiver.No fee change; a contextual request to the authorized decision owner.
Uncertain system resultA write times out after the request is sent.No premature success claim; reconciliation before another write.
Conflicting instructionA retrieved document asks the employee to ignore access rules.Untrusted instructions do not widen access or authorize an action.
Repeat or late eventA booking event is delivered twice.No duplicate task or conflicting follow-up; source reference retained.

Use synthetic or approved test data in a controlled environment. Do not send real customer communications or create live transactions without explicit test authorization.

IMPROVEMENT WITH CHANGE CONTROL

Change the cause, not just the wording.

A failure may come from unclear instructions, outdated knowledge, excessive permissions, a tool contract or the routing model. Identify which layer is responsible before changing it.

01

Diagnose

Compare expected and observed behavior. Locate the first incorrect decision or state change.

Record: failure category and evidence
02

Change narrowly

Adjust the responsible instruction, source, permission or integration behavior. Record the reason and version.

Record: scoped change and owner
03

Retest broadly

Rerun the failed case and related cases that previously worked. Include unseen variations and exception paths.

Record: regression results and open issues
04

Review in operation

Agree a bounded rollout, watch for regressions and preserve a pause or rollback route.

Record: release scope and review triggers

A RELEASE DECISION, NOT A VANITY SCORE

Know what passed. Know what remains open.

Agree release criteria before running the suite. Record both the evidence and the limits of what was tested, then give the designated owner a concrete decision to make.

Coverage & configuration

Identify the role, knowledge version, rules, connected tools and test environment. Document which scenarios were covered and which real-world conditions remain untested.

Critical failures & open issues

Separate prohibited actions, incorrect writes and false success claims from lower-severity issues. Do not bury a critical failure inside an aggregate pass rate.

Business measures

Define successful completion, time to completion, rework and appropriate handoff. Compare like-for-like task groups and include unresolved work in the reporting denominator.

Ownership & recovery

Name the release reviewer and operating owner. Agree rollout boundaries, review triggers, pause conditions and how the affected workflow can be reverted.

EVALUATION QUESTIONS

Make the standard of good work explicit.

What is AI workforce evaluation?

It is the review of an employee’s responses, decisions, system actions and handling of exceptions against agreed expectations. The unit of evaluation is the business task: for example, whether an appointment was correctly booked and verified in the calendar.

How is it different from testing a chatbot?

Conversation quality is one layer. An AI employee may also change records, schedule actions or hand work to another owner. Evaluation therefore needs to inspect permissions, tool inputs, final business state and recovery behavior in addition to the answer.

What belongs in the test set?

Include representative routine cases, ambiguous inputs, boundary cases, prohibited requests and deliberate system failures. Use approved or synthetic test data and controlled destinations. Keep some cases separate from the examples used to develop a change so the review is not limited to rehearsed inputs.

What score is good enough to launch?

There is no universal launch score. Agree criteria according to the responsibility and consequences of failure. Review critical failures separately: a high average cannot compensate for an unauthorized action or an incorrect claim that a task was completed.

Does a human handoff count as a failed test?

Not if handoff is the expected behavior. A correctly routed exception can pass, while an unnecessary handoff on an authorized routine task can fail the autonomy requirement. Define the expected route and evidence for each case before testing.

Does the workforce change itself automatically?

Do not equate improvement with unrestricted self-modification. The operating approach here is to identify a failure, make a scoped change, rerun affected and regression cases, and have the release owner review the evidence. Confirm the supported evaluation and change-management tooling for your deployment.

When should we rerun evaluations?

Review affected cases when instructions, knowledge, model configuration, permissions or integrations change. Also add representative production failures to the test set after removing sensitive data as required. Agree review frequency and incident triggers with the operating owner.

Are the visual’s results actual customer benchmarks?

No. The visual uses illustrative results to explain how a failure, retest and release review relate. Actual results must come from your configured workflow, approved test environment and agreed acceptance suite.

START WITH THE OUTCOME

What would prove the work is right?

Bring a representative workflow and its difficult cases. We’ll map expected behavior, failure conditions and the evidence your release owner needs.