Browse documentationReview modes
Use the interface that matches the question
The review screen adapts to the task mode, media type, presentation rules, criteria, and reference requirements. Candidate identity remains hidden while the reviewer chooses a winner, records a tie or skip, or assigns a rating.
| Mode | Best for | Reviewer action |
|---|---|---|
| Side-by-side pair | Images or media that can be compared at the same time | Inspect both candidates and choose the better match |
| Sequential pair | Video or audio where focused playback matters | Watch or listen to each candidate before choosing |
| Reference-dependent comparison | Image-to-image, image-to-video, reference-to-video, and related tasks | Use the Inputs control in the comparison or rating flow, then judge preservation and requested change |
| Rating | Independent quality scoring against a stable rubric | Assign a score on the task-pinned range, such as 1 to 5 or 1 to 9 |
| DeltaFrame | Video differences that depend on timing or a specific frame | Inspect synchronized playback and localized changes |
Pairwise decisions
- Choose a candidate
Select the output that better satisfies the written criteria.
- Tie
Use tie when the candidates are meaningfully indistinguishable and the task permits it.
- Skip
Use skip for broken, missing, incomparable, or out-of-scope evidence instead of forcing a preference.
- Comment
Record a concise reason or visible failure mode without guessing the model identity.
Blind presentation
Candidates use neutral labels and their order varies across evaluators. The same evaluator receives a stable order for the same item so a refresh does not silently swap the choices. Model names become available only on result paths that the account is allowed to open.
Watch and listen requirements
For supported video protocols, MLOpt chooses 0, 25, 50, 75, or 100 percent required playback per candidate. The requirement is frozen into the Session protocol and applied separately to every presented video or audio candidate. A choice, tie, or rating remains disabled until each candidate reaches its requirement; skip stays available for broken or unusable media.
The server credits unique timeline coverage against the media duration. Replaying the same range, jumping the playhead, duplicating a sequence number, refreshing the page, or changing a client-side counter does not satisfy the requirement. Playback errors and unusable review Sessions remain quality issues rather than being treated as confident evidence.