Skip to main content

Evals in Creatio AI Studio

"It worked when I tried it" is not evidence that an agent behaves. Evals turn the behaviors you care about into repeatable, automated checks that run against a deployed agent and report a pass or fail for each one.

Evals live on each agent, on the Evals tab of the Observe tab group. The tab has two views: Runs and Cases.

The three building blocks​

Building block

Description

Case

One graded conversation. A case holds what the user says and the graders that decide whether the reply is acceptable. Cases live in a shared pool per agent.

Suite

A tag that groups cases. A case can belong to any number of suites, and a suite can optionally cap the judge spend per run.

Run

An execution of the cases in a chosen scope against the agent deployed in a chosen environment. A run reports a passed / total score.

Suites are tags rather than containers. Deleting a suite removes the tag and leaves its cases in the agent's pool. Deleting a case removes it from the agent's pool and from every suite it belongs to, and the runs it took part in keep its results.

Case modes​

A case can be authored in the following modes.

  • Conversation — a chat-style timeline of user and agent exchanges. The agent runs against the final user message and is scored on the final agent reply. Use it for most cases.
  • Raw JSON — the case exactly as the eval runner sees it. Use it when Conversation cannot express what you need. Edits made here are not translated back into the other mode.

Grading​

Grading happens in two layers, and every grader must pass for the case to pass.

Reply matching​

The final agent reply is scored with one of the following match strategies.

Match strategy

Description

Exact

The reply must equal the expected text verbatim.

Regex

The reply must match a /pattern/flags expression.

LLM judge

An LLM scores the reply against a plain-language rubric you write. The most flexible option and the one that fits most cases.

Case-level graders​

Case-level graders are transcript-wide gates and trace assertions. They are grouped into content checks and operational checks.

Grader

Description

Tool sequence

Content check. The expected tools were called. You can require a sequential order or any order, and allow extra calls.

Skill order

Content check. The expected skills were invoked in the expected order.

No PII

Operational check. Scans every agent turn for PII patterns and fails the case on detection.

Latency budget

Operational check. Fails the case when the run is slower than the cap you set.

Cost budget

Operational check. Fails the case when the run exceeds the total usage cap you set.

Forbidden skill

Operational check. Fails the case when a named skill is invoked.

Reading a run​

The Runs view lists every run with its passed / total score, scope, environment, and duration. Expand a run to see the per-case status and grader verdicts.

Each case result shows the agent output, an execution timeline with the latency of every LLM call and grader, and a grader table with the verdict, score, confidence, and the reasoning behind it. For an exact or regex mismatch, a diff shows what was expected against what came back.

Each run gives the agent the same instructions, temperature, and environment as in a conversation, and the agent calls its tools, so a case that depends on a tool is graded on the tool's actual result. A run can report that granted tools could not be called, for example, because an integration has no credentials. Each case that expects one of those tools is marked "tool unavailable," and its status is still the graders' verdict. Such cases can fail on content even when the prompt is correct, so check for the "tool unavailable" mark before you investigate the prompt.

Important

Do not use real data, including personally identifiable information, in evaluation cases. Cases are stored, so use obviously fake test data. A case that checks whether an agent leaks a card number does not need to contain one — ask whether the agent requests it, and let the No PII grader check the reply.


See also​

Create a prompt agent

Creatio AI Studio overview