Quality you can
point to.
An Eval is a stored quality gate: criteria, thresholds, and an LLM-as-judge that scores the artifact and returns a verdict with a full report. Low-quality work stops before it ships.
What an Eval is.
A declared EvalSuite with criteria, pass thresholds, and a judge that produces a verdict and evidence report.
- Scores artifacts against declared criteria
- Uses an LLM-as-judge, not a checklist
- Pass modes: all criteria or a minimum score
- Returns a verdict and a full reason report
- Passed eval report is required evidence
- Registered in the evals registry
Exactly how it runs here.
Evaluations apply declared criteria and return evidence-backed reports. The Content Department uses a copy quality gate before campaign work is marked ready.
How an Eval works.
- 1
A Department registers an EvalSuite with criteria and thresholds
- 2
The run produces an artifact
- 3
The judge scores it against the criteria
- 4
The verdict and report are stored with the run
- 5
The artifact advances only on a passing report
Want this discipline in your company?
The Agent Brief maps one of your workflows to the same loop, approvals, and evidence.
Apply — Free Review