Overview
Manually reviewing AI Agent conversations is time-consuming, difficult to scale, and can produce inconsistent results between reviewers. Eva provides organizations with an automated and consistent way to evaluate conversation quality and identify interactions that require attention.
Each evaluated conversation is assessed across five quality metrics: Customer Understanding, Task Execution, Compliance and Safety, Customer Experience, and Customer Sentiment. Based on these metric evaluations, Eva also assigns an Overall Score that reflects the quality of the conversation as a whole.
Eva provides a written summary explaining the evaluation and the reasoning behind the results. By applying the same evaluation criteria across conversations, organizations can monitor AI Agent performance at scale, identify quality issues, and better understand where improvements may be needed.

Key Considerations
- Eva requires a replenishable AI token account. Contact your Customer Success account manager for more information.
- Each evaluated conversation consumes a fixed number of AI credits, regardless of conversation length.
- Conversations shorter than 30 seconds are automatically excluded.
- Evaluation begins after Eva is enabled and does not apply retroactively to previous conversations. You can change the evaluation percentage or disable Eva at any time.
- Eva operates in the brand's configured language, regardless of the language selected for the individual AI Agent or the language used in the conversation.
Setting Up Eva
Enable Eva in the AI Agent Settings, then select the percentage of conversations you want Eva to evaluate.

Selecting the Evaluation Percentage
Choose the percentage based on your monitoring needs and conversation volume. Consider using higher coverage when launching a new AI Agent or after making significant changes, and a lower percentage (e.g., 10–25%) for ongoing monitoring of stable agents. For high-volume agents, even a relatively low percentage (e.g. 5% of 1000 daily conversations) can provide a meaningful sample of conversations.
Note: A higher evaluation percentage provides broader quality coverage but also increases token consumption.
Evaluation Components
Each Eva evaluation includes:
- Scores for five evaluation dimensions
- An Overall Score
- An evaluation Summary

The Five Evaluation Dimensions
Eva evaluates each of the five dimensions independently, using the conversation and available system records as evidence. A high or low score in one dimension does not directly affect the others.
What is scored:
- Customer Understanding – Did the agent correctly grasp what the customer asked for, including follow-ups and context?
- Task Execution – Did the customer get what they came for — resolved, correctly handed off, or correctly declined?
- Compliance & Safety – Did the agent stay within its role and rules, without inventing facts or making unauthorized promises?
- Customer Experience – How much effort and friction did it take the customer to reach an outcome?
- Customer Sentiment – How did the customer's own expressed mood evolve from the start of the conversation to the end?
Score Values:
- 5 – Flawless in this dimension. Everything understood/resolved/compliant on the first attempt, no issues.
- 4 – Good with minor friction. Solid performance with one small imperfection, such as an extra clarifying exchange, a secondary request only partially handled, slight overpromising.
- 3 – Mixed. The agent succeeded partially or eventually, but with meaningful problems such as a misunderstood request, unresolved needs, an invented detail, or avoidable back-and-forth.
- 2 – Mostly failed. Repeated problems dominated the interaction, such as persistent misunderstanding, an unresolved main request, fabrications, or a frustrating dead-end experience.
- 1 – Complete failure in this dimension. The agent never understood the request, nothing was achieved, unsafe or off-role behavior occurred, or the customer left angry or gave up.
Overall Score
The Overall Score is not a calculated average of the five dimension scores. Instead, Eva evaluates the conversation as a whole and assigns an overall quality score based on all relevant aspects of the interaction.
This provides a holistic assessment of the conversation rather than giving equal mathematical weight to each evaluation dimension.
Score Ranges
Scores are color-coded to make it easier to identify successful conversations and those that may require attention:
- 0–40 | Red – Poor: The conversation ended without achieving the customer's goal, and the agent's performance contributed to that outcome. For example, the agent may have failed to identify the customer, misunderstood their request, or responded poorly enough to prevent the conversation from progressing. Red conversations are the highest priority for review.
- 41–70 | Orange – Fair: The customer's goal may have been reached, but the interaction did not go smoothly. Extra effort may have been required, the experience may have involved friction, or part of the interaction may have fallen short. Orange conversations often highlight opportunities for improvement.
- 71–100 | Green – Good: The customer's need was correctly identified and addressed smoothly, resulting in a positive service experience. Green conversations represent the expected standard of performance.
Summary
Eva provides a written explanation of each evaluation, helping teams understand the reasoning behind the scores and identify what worked well and where improvements may be needed. Each evaluation includes a clear explanation of what happened, why the conversation received its score, and where teams should focus their improvement efforts.