Create and run evals
This article explains how to build an eval suite for a prompt agent, author cases that grade its behavior, and run them against a deployed agent.
Evals run against the agent deployed in an environment, so publish and deploy the agent before you start a run.
Create a suite
A suite is a tag you apply to cases. Create one first so you have something to attach cases to.
-
Go to the Agents section in the navigation panel.
-
Open the prompt agent you want to evaluate.
-
Open the Evals tab on the Observe tab group.
-
Click Manage suites → New suite. This opens the "Create eval suite" dialog.
-
Fill out the suite fields.
Field
Field value
Name
Name of the suite, for example, "regression-v1."
Description
Optional summary of what the suite covers.
Judge budget (credits)
Optional cap on judge spend per run, in your organization's usage unit, which is AI credits by default.
-
Click Create.
As a result, Creatio AI Studio creates the suite. Attach cases to it in the next step.
Add a case
-
Open the Evals tab of the agent and switch to the Cases view.
-
Click Add case. This opens the "New case" editor.
-
Enter a name for the case in the unlabelled field at the top, which prompts you with "e.g. responds-with-pong."
-
Select the suite in the Attach to suite list.
-
Keep the Conversation mode selected.
-
Enter the user message the agent must handle in the final exchange. To set up context first, click Add exchange and pre-seed the earlier turns of the dialog.
-
Select the Score this turn checkbox for the final exchange and select the match strategy in the Match field.
Match strategy
What to enter
Exact
The literal text the agent must reply with.
Regex
A
/pattern/flagsexpression the reply must match.LLM judge
A plain-language rubric the LLM judge scores the reply against. Set the judge model in the Judge model field, or leave it on "Default (server-side)."
-
Add the case-level graders, if needed. Click Add case grader and select a grader from the Content checks or Operational checks group. Graders combine with AND, so the case passes only when every grader passes.
-
Select the environment to run the case in, at the bottom of the editor.
-
Click Save.
As a result, Creatio AI Studio adds the case to the agent's shared pool and tags it with the selected suites. To run the case immediately, click Save & Run instead of Save.
Do not use real data, including personally identifiable information, in evaluation cases. The Cases view shows this warning above the list because cases are stored.
A case marked with the "Vendor" tag ships with the agent. The platform replaces it on the next application update, so edits do not persist. Create a new case instead if you need a version that survives an update.
Delete a case
- Open the Evals tab of the agent and switch to the Cases view.
- Click the Delete case icon in the row of the case.
As a result, Creatio AI Studio removes the case from the agent's pool and from every suite it belongs to. The case is deleted as soon as you click the icon, without a confirmation, so check the row before you click. The runs the case took part in keep its results in the Runs view.
Run the suite
- Open the Evals tab of the agent.
- Click Start run.
- Select the suites to evaluate in the Suites (optional scope) field. Leave the field empty to evaluate every case in the agent's pool.
- Select the environment in the Environment field. Evals dispatch tasks to the agent deployed in this environment.
- Review the Preview section. It reports how many cases the run covers, for example, "1 case(s) will be evaluated."
- Click Start run.
Evals always call the agent's real tools, so a case that writes data writes it for real. Point the run at an environment where that is acceptable. A case that depends on a tool is graded on the tool's actual result, so choose an environment where the agent's integrations have credentials.
As a result, Creatio AI Studio executes the agent against every case in scope and grades each output. A row appears in the Runs view naming its scope, environment, and age, and the tab header carries the score, for example, "Evals · 1/1 · 100%."
Every user message in a case, including the earlier exchanges you pre-seed, is sent to the agent as a separate run. Each run can take up to 120 seconds, whatever the agent's Run time budget (seconds). The case as a whole gets 120 seconds per user message, so time spent waiting on one message reduces the time left for the next. For example, a case with three user messages has up to 360 seconds in total. The messages of one case never get more than 25 minutes in total.
Review the results
- Open the Evals tab of the agent and switch to the Runs view.
- Expand the run to see the per-case Case, Status, Graders, and Latency columns.
- Click the chevron on a case to inspect the result: what the agent replied, the verdict, score, confidence, and reasoning for each grader, and the latencies. A failed exact or regex match includes a diff of the expected and actual text.
As a result, you can tell which behavior broke and why. When the agent itself could not run, the detail reports "Agent execution failed" and "Graders skipped": fix the agent before you read anything into the score.
If the run reports that granted tools could not be called, cases that depend on those tools can fail on content even when the prompt is correct. Each case that expects one of those tools is marked "tool unavailable," and its status is still the graders' verdict. Credential state is checked at the organization level, so treat the warning as "not usable" rather than a verdict about one environment.
To run the cases again after a fix, deploy the fix to the environment of the run, then click Re-run on the run row. Re-run uses the environment of the original run. A run that shows "Unknown environment" was recorded before environments were tracked, so Re-run first asks you to choose where to re-run it: select the environment, then click Re-run again.
To capture a real conversation as a case, open Observe → Runs, find the run, and click Save as eval case on it. The action is on an agent run, not on an eval run, and it appears only where an administrator has turned on Allow saving eval cases in the Settings section under Offline evals. It is off out of the box.