Skip to content

Use the interface that matches the question

The review screen adapts to the task mode, media type, presentation rules, criteria, and reference requirements. Candidate identity remains hidden while the reviewer chooses a winner, records a tie or skip, or assigns a rating.

Review modes and their purpose
ModeBest forReviewer action
Side-by-side pairImages or media that can be compared at the same timeInspect both candidates and choose the better match
Sequential pairVideo or audio where focused playback mattersWatch or listen to each candidate before choosing
Reference-dependent comparisonImage-to-image, image-to-video, reference-to-video, and related tasksUse the Inputs control in the comparison or rating flow, then judge preservation and requested change
RatingIndependent quality scoring against a stable rubricAssign a score on the task-pinned range, such as 1 to 5 or 1 to 9
DeltaFrameVideo differences that depend on timing or a specific frameInspect synchronized playback and localized changes

Pairwise decisions

  • Choose a candidate

    Select the output that better satisfies the written criteria.

  • Tie

    Use tie when the candidates are meaningfully indistinguishable and the task permits it.

  • Skip

    Use skip for broken, missing, incomparable, or out-of-scope evidence instead of forcing a preference.

  • Comment

    Record a concise reason or visible failure mode without guessing the model identity.

Blind presentation

Candidates use neutral labels and their order varies across evaluators. The same evaluator receives a stable order for the same item so a refresh does not silently swap the choices. Model names become available only on result paths that the account is allowed to open.

Watch and listen requirements

For supported video protocols, MLOpt chooses 0, 25, 50, 75, or 100 percent required playback per candidate. The requirement is frozen into the Session protocol and applied separately to every presented video or audio candidate. A choice, tie, or rating remains disabled until each candidate reaches its requirement; skip stays available for broken or unusable media.

The server credits unique timeline coverage against the media duration. Replaying the same range, jumping the playhead, duplicating a sequence number, refreshing the page, or changing a client-side counter does not satisfy the requirement. Playback errors and unusable review Sessions remain quality issues rather than being treated as confident evidence.