Evaluations
Score an agent version against a dataset of known cases, read the results, and stop a version from being published when it scores too low.
An evaluation runs an agent version once for every case of a dataset and scores each answer against the expected one. Use it before you publish, and to compare versions. Evaluations are under Agentic ▸ Evaluations in the Orchestrator, with two tabs: Datasets and Runs. Create a dataset# A dataset is a named list of items. Each item has an Input (passed to the agent as its inputs), an Expected output, and optional Tags. The dataset's scoring settings decide how an answer is compared with the expected output. Datasets are created through the API; New dataset in the web app is not available yet. You need evaluations.create in the workspace. Create a datasetHTTPCopyPOST /api/v1/tenants/{tenantId}/workspaces/{workspaceId}/evaluations/datasets Authorization: Bearer <your access token> Content-Type: application/json { "name": "Invoice triage cases", "description": "Known invoices with the decision the AP team expects.", "scoringType": "binary", "scoringConfig": { "mode": "regex" } } Add itemsHTTPCopyPOST /api/v1/tenants/{tenantId}/workspaces/{workspaceId}/evaluations/datasets/{datasetId}/items Authorization: Bearer <your access token> Content-Type: application/json { "items": [ { "input": "{\"invoiceNumber\":\"INV-2026-0415\",\"vendor\":\"Tailspin Toys\",\"amount\":27650,\"currency\":\"EUR\"}", "expectedOutput": "\"approver\": \"CFO\"", "tags": ["threshold"] } ] } input is a JSON object written as a string. scoringType is binary, numeric, likert or custom. scoringConfig.mode chooses the comparison: Mode An answer scores 1 when exact (default) It is exactly the expected text. json_equal It is the same JSON as the expected output, ignoring key order. numeric_tolerance It is a number within scoringConfig.tolerance (default 0.01) of the expected one. contains It contains the expected text. regex The expected output, read as a regular expression, matches it. llm_judge A judge model (judge:default) rates it. Open a dataset to see its Items, Recent runs and Trends. Evaluate an agent version# Evaluations of an agent are started through the API, on a version (a candidate) of the agent. You need evaluations.create. Find the version's ID: GET /api/v1/tenants/{tenantId}/workspaces/{workspaceId}/agents/{key}/candidates lists every version with its id, version and status. Start the evaluation: Evaluate version 1 of an agentHTTPCopyPOST /api/v1/tenants/{tenantId}/workspaces/{workspaceId}/agents/invoice-triage-agent/candidates/{candidateId}/evaluations Authorization: Bearer <your access token> Content-Type: application/json { "datasetId": "<dataset id>", "minPassScore": 0.75 } The evaluation starts one run per item. They appear on the Agents Runs page with the kind Evaluation and need a Robot like any other run. Each answer is scored when its run ends. Read the results# Open Evaluations ▸ Runs. Filter by Dataset or Status if needed. Select a run. The list shows each run's Status, Avg score, Regressions, Improvements, Flaky items, Items, Duration and CI gate. The run page shows the Average score, the counts, Duration and Total cost, and whether the run passed its gate (CI gate details). The Results tab lists each item with its Score; a score below the pass mark is flagged. Job and Trace open the item's run in the Orchestrator. Other tabs show Policy, Coordination, Memory and Trace metrics results, and Compare sets this run against another run of the same dataset. Gate publishing on an evaluation# You can require that a version scores well enough before it can be published. Set a gate on the agent through the API: Set an evaluation gateHTTPCopyPUT /api/v1/tenants/{tenantId}/workspaces/{workspaceId}/agents/invoice-triage-agent Authorization: Bearer <your access token> Content-Type: application/json { "revision": 4, "gate": { "datasetId": "<dataset id>", "minScore": 0.9 } } revision is the agent's current revision (from GET …/agents/{key}). With a gate, publishing also runs the check dataset_gate: the newest finished evaluation of that exact version on the gate dataset must average at least minScore. Because the evaluation must belong to the version you publish, publish a gated agent through the API: Create a release candidate: POST …/agents/{key}/candidates with { "purpose": "release" }. The answer has its id. Evaluate it on the gate dataset, as above, and wait until the evaluation run is Completed. Publish it: POST …/agents/{key}/candidates/{candidateId}/publish. A failing check answers 409 agent_publish_blocked with the list of checks. See Publish a version for the other checks. Evaluate a process# Datasets also work for ordinary processes. POST /api/v1/tenants/{tenantId}/workspaces/{workspaceId}/evaluations/run with datasetId and automationId starts one job per item; optional baselineRunId, minPassScore and maxRegressions set what the CI gate compares and requires.
Create a dataset
Evaluate an agent version
Read the results
Gate publishing on an evaluation
Evaluate a process