Agent evaluation

Test whether an AI agent can do your work

Give an agent a job. Define what success means. Inspect what it did, where it failed and whether a change helped. Gradia connects the result to the rules and evidence behind it.

Have a workspace? Open Console. New managed workspaces require an invitation.

Who it helps
AI teams choosing an agent or testing a release, and business owners who need to know whether it can follow their workflow.
Availability and scope
Start with agent output evaluation in the Console: add cases, run exact, text or field checks, and compare supplied baseline and candidate answers. A downloadable Python runner can produce those answers from your own trusted agent code. Full workflow execution uses a prepared environment and reviewed evaluator. A universal one-click agent connector is not available.

Work with Gradia

Scope an agent evaluation

A focused, assisted engagement. No Console account needed to start the conversation.

Agree on the job, the test and the evidence needed for your next release decision. Start with supplied outputs or an existing evaluation; scope complete workflow execution separately.

Bring
One workflow, the decision you face, a person who knows the work and a description of the evidence available.
Receive
A scoped proposal with deliverables, required inputs, supported checks, review responsibilities, timing and fees. Work begins after those are agreed.
Email Gradia to scope a pilot →

Opens your email app. You can also write to rudy@gradiahq.com. Describe the workflow first; keep confidential records for the agreed transfer process.

What should an agent evaluation actually check?

The outcome, the rules it had to follow and the evidence available when it acted. A fluent answer can still be wrong. A successful tool call can still break a business rule. Gradia keeps task failures, environment problems and missing evidence distinct so you can decide what to fix.

How the engagement works

  1. Bring one job

    Start with a few inputs and expected answers. Add them in the output workbench or import a dataset, then run the local Python adapter if you need agent outputs. For a complete workflow, begin with a brief or approved Observer history. Imported records do not automatically become runnable environments.

  2. Agree on the test

    Use the full workflow path to review requirements with an expert, prepare the supported environment and define measurable outcomes. Check the evaluator with known passing and failing controls. Model judges need their own human alignment evidence.

  3. Run and inspect

    Approve the model, tools, data use and budget for the supported runtime. Review the attempt alongside its outcome and evidence. For captured Observer cases, the protected assessment has a separate reviewed setup and currently admits a task-correctness criterion.

  4. Compare a change

    For comparable runs, keep the task, evaluator and execution conditions fixed. Reviewed Observer cohorts can compare 2–40 native cases, with missing results and infrastructure outcomes kept visible. Uncertainty depends on the reviewed sampling assumptions; a single matched pair remains descriptive.

What you leave with

A result you can investigate
See the recorded actions, the applicable checks and the evidence supporting the result. Missing evidence stays visible rather than becoming a confident score.
A clearer repair decision
Separate an agent mistake from a broken test, a disagreeing judge or an infrastructure failure. Use the finding to decide which part of the system needs work.
A comparison with its limits
Compare supported candidates under matched conditions. Native benchmark reports and protected Observer comparisons have different admission and reporting rules; captured results do not automatically become published benchmarks or training data.

Illustrative example

The refund went through. Should it have?

A support agent refunds $140 when the policy requires approval above $100. The tool succeeds and the customer gets an answer, but the agent skipped a required approval. A better next action is to request approval and leave the refund pending. Evaluate both the action and the business rule.

Illustrative scenario, not a customer result or an executed model comparison. More cases, executed controls and independent review are needed before relying on a release decision.

Before you start

Do I need to build a whole Universe first?
No. The agent output workbench checks supplied answers and compares two versions without an environment build. It keeps missing attempts and errors visible. Checking real tool actions, permissions and application state needs a supported environment and approved evaluator. You can also start with Guard evidence, trace review or a benchmark audit.
Can I keep my existing observability tools?
Yes. Gradia has explicit JSON/JSONL import paths for supported trace schemas. Check format and field coverage before uploading. Imports preserve evidence provenance; they are not automatic live synchronization or a guarantee of compatibility with every vendor export.
How is the evaluator checked?
The native evaluation system supports known-answer controls, checks for ways to obtain an undeserved pass, and agreement checks against human annotations. Each evaluator must be qualified for its intended criterion. A cryptographic signature can show that a record changed; it cannot prove that a judgment was correct.
How do I start in the Console?
Choose Evaluate agent outputs when creating a project, or open Evaluation in an existing project. Try the synthetic example, add your cases and run checks. Save your dataset to resume; the workbench keeps it in the current tab. Choose Build the complete workflow for specification and environment setup. For a captured case, start in Observer. New managed workspaces currently require an invitation.

Keep the next step connected

Open Console

Start with agent output checks, or continue into a reviewed workflow assessment.

Get the local Python runner

Produce baseline or candidate outputs using your trusted agent function. Runs with your own machine and provider permissions; this is not a sandbox.

Start locally with open-source Guard

Capture and inspect a covered agent run without a managed account. Recording alone does not evaluate its quality.

Check an existing benchmark

Review the test and its evaluator before relying on its scores.

Explore workflow test environments

See how approved rules, state and changing conditions can form a Gradia Universe.