Workflow · Advanced · 40 min
Evaluate an agent with a fixed test set workflow
Run an agent against a fixed set of inputs after every change and compare against the last run, so regressions are caught before users are.
- Claude
- ChatGPT
- Any assistant
- agents
- workflow
Inside this workflow
- When: After any change to the agent's prompt, tools or model.
- 1. Freeze the test set
- 2. Run and record
- 3. Score against expectations
- 4. Diff against the last run
- 5. Decide
- Steps
- 5
- Tools
- A test file of inputs and expected outputs, A script that runs the agent
The full workflow opens with a membership that covers AI Agents.
Sign in if you already subscribe.
Sign in
