Protocol

Turn subjective quality into evaluator criteria before looping on it

An evaluator cannot enforce taste that nobody has made legible.

When it fits

  • An agent is told to make a design, report or interface 'better' and keeps polishing without a stable target.

When to avoid it

  • Some quality remains irreducibly judgmental; a rubric makes it discussable but does not turn taste into objective truth.

Why it matters

Before adding an evaluator loop, write a small rubric for the dimensions that matter, with observable examples or failure anchors. Separate hard requirements from preference. Let the evaluator point to specific violations and evidence, not emit one opaque score. Periodically compare evaluator judgments with human review so the loop does not optimize a distorted proxy.

Steps

  1. A reviewer can explain why the evaluator passed or failed the artifact without appealing to its confidence.

An example

For a dashboard, grade information hierarchy, task completion and responsive behavior separately instead of asking a judge model whether it 'looks professional.'

Check your result

A reviewer can explain why the evaluator passed or failed the artifact without appealing to its confidence.

Keep this limit in mind

  • Some quality remains irreducibly judgmental; a rubric makes it discussable but does not turn taste into objective truth.

Connected ideas

Use before
Add planner-builder-evaluator roles only after the simple loop hits a ceiling

Evidence and sources

Supports

Anthropic's evaluator work operationalized subjective quality by writing concrete grading criteria before using an evaluator agent to drive iteration.

Rubrics encode judgment imperfectly and can be gamed; periodic human calibration remains important.

Harness design for long-running application development · Generator-evaluator loop and evaluator criteria

All sources (1)