Dataset starter
What should your agent get right?
Turn a task and a reference answer into an editable evaluation contract and first case.
You leave with: Contract + reference case
Build my starter →Free tools · no account needed
Define what good looks like, see what changed, or find the next failure to investigate. Try a fictional example or use your own inputs. Download the result before you leave.
Your inputs stay in this browser. Workspace access is a separate, optional step.
Dataset starter
Turn a task and a reference answer into an editable evaluation contract and first case.
You leave with: Contract + reference case
Build my starter →Output comparison
Compare supplied baseline and candidate outputs. Find regressions, improvements and missing results.
You leave with: Case-by-case report
Compare outputs →Quality report reader
Read a ProofSeal or TestLore JSON report and get a focused next step.
You leave with: Findings + next step
Read a report →Open-source developer tool
Code Transplant extracts a TypeScript parser’s dependency boundary, adapts its interface to another application, and compares behavior against the original. Inspect the first demonstration, then run the local CLI in your own repository.
The website displays the proof; source extraction and execution happen locally.
ProofSeal and TestLore run independently in your own repository. Gradia Guard records covered execution and verifies its evidence. A Gradia workspace can connect their reports to approved evaluation criteria and independent human review.
Passing software tests and intact recordings answer specific questions. Your team still approves the requirements, case answers and business acceptance criteria.
Start with the task you just explored. Agree on success, review the contract, generate case proposals, and inspect recorded failures before choosing the next experiment.
Access is by invitation. The free tools and downloads stay available without joining.