Skip to content

Reviewer influence changes only with comparable evidence

Creative Engineer reviews use a protected role-based contribution. Standard reviewer influence begins neutral and may change only after enough healthy, decisive, two-candidate comparisons can be assessed against independent peer evidence.

The comparison is not Creative versus everyone else

Creative judgment is important, but it is not stored as ground truth. The reliability method removes the target reviewer from the comparison and looks at mature Creative and Standard peer evidence separately. A decision contributes less when the item is weakly supported or the reviewer groups disagree strongly.

This protects a reviewer from being penalized because the item itself is ambiguous. It also exposes cases where Creative Engineering and the broader reviewer group may be applying different interpretations of the criterion.

What can change

  • Task-local influence

    Comparable evidence from the current task can reduce a Standard reviewer's contribution to that task.

  • General influence

    Repeated evidence across different tasks and prompts can affect the reviewer's broader contribution.

  • Access state

    Repeated, mature adverse evidence may warn or pause a Standard reviewer. A single unusual choice is not enough.

  • Creative review

    Creative contribution stays fixed, while strong disagreement can create a diagnostic review case.

Topic and use-case signals

GenMedia can calculate reviewer diagnostics for tagged areas such as motion, identity, or reference fidelity. These signals help inspect where reviewer disagreement concentrates. They do not currently replace the production event weight or automatically route tasks by expertise.

Model-level topic findings also require repeated qualified evidence before they appear. A single tagged vote does not create a public expertise label or a model claim.

Reviewer data remains restricted

Individual coefficients, warnings, calibration records, and reviewer-linked decisions are operational data, not a public leaderboard. Normal result pages expose the aggregate evidence needed to interpret the Evaluation Task without exposing another person's reliability profile.