AI TRAINING & EVALUATION

Good AI still needs good judgment.

I evaluate AI responses for accuracy, consistency, and how well they follow the prompt.

Creative director combining analog design judgment with an AI-assisted visual workflow

HOW I REVIEW AI WORK

A clear process, with room for human judgment.

I start with the prompt, separate requirements from preferences, and review each response against the same criteria. Then I explain the choice with specific evidence and flag anything the prompt leaves ambiguous.

JUDGMENT IN PRACTICE

The work is in the reasoning.

The decision should be clear enough that another reviewer can follow it and apply the same standard.

POLICY & ACCURACY

Catch the answer that sounds right but makes something up

What happened

Two answers are clear and polite. One promises a refund the source never authorizes.

How I judge it

Choose the answer that stays within policy. Polished wording does not excuse invented information.

VISUAL ACCURACY

Do not confuse “looks good” with “followed the prompt”

What happened

A generated design looks polished but changes the required colors, logo treatment, and hierarchy.

How I judge it

Score it against the prompt, not appearance alone. An attractive response can still miss the task.

REVIEWER CONSISTENCY

Turn a gray area into a rule everyone can use

What happened

Reviewers disagree on whether a small style problem should change which response wins.

How I judge it

Identify the affected criterion, decide whether the issue is minor or task-breaking, and document the threshold.

CORE REVIEW AREAS

What I check

01 / AI RESPONSE REVIEW

Which answer is better—and why?

Compare responses against the prompt, identify factual or reasoning problems, and explain the stronger choice.

02 / TEXT + IMAGE REVIEW

Does the visual actually follow the prompt?

Review text and images for composition, typography, brand accuracy, visual defects, realism, and instruction-following.

03 / REVIEWER CONSISTENCY

Would another reviewer reach the same conclusion?

Clarify gray areas, document recurring failures, and help reviewers apply the same standard consistently.

REPRESENTATIVE WORK SAMPLE

Human evaluation of AI responses · RLHF-style preference ranking

Choosing the better of two AI responses.

Read the prompt, compare Response A and Response B, then choose the stronger response.

PROMPT

Create a homepage hero for a premium architecture studio using the supplied identity. Keep the tone editorial and quiet, preserve the neutral palette, avoid generic sustainability claims, and use one clear call to action.

RESPONSE AMISSES KEY REQUIREMENTS

This version adds a blue gradient that was not requested, uses generic stock imagery, introduces two competing calls to action, and makes a sustainability claim that was never provided.

  • Changes the requested visual direction
  • Adds an unsupported claim
  • Uses competing calls to action
  • Weak match to the supplied brand
PREFERRED RESPONSE
RESPONSE BBEST FIT

This version keeps the neutral identity, creates one clear hierarchy, uses a single call to action, and stays close to the requested tone. It still has one smaller contrast problem that should be corrected.

  • Follows the requested palette
  • Clear information hierarchy
  • One purposeful call to action
  • Only a minor contrast issue remains

RANKING CHECK

Follows the promptWhich response follows the requested instructions most closely?
AB
Accurate & supportedWhich response stays within the information actually provided?
AB
Matches the brandWhich response preserves the supplied visual direction and tone?
AB
Clear hierarchyWhich response organizes the content more clearly?
AB
Overall preferenceWhich response is the stronger result overall?
AB
DECISION

Response B. It follows the prompt and stays within the supplied information; Response A adds unsupported claims and unrequested visual changes.