Browse documentationEvaluation Task design
Decide what may vary
A useful benchmark starts with a narrow question and a clear control condition. GenMedia presets describe the comparison intent, the inputs that should remain fixed, and the review mode that fits the decision.
| Design | Use it for | Keep constant |
|---|---|---|
| Model versus model | Compare two endpoint arms on the same cases | Prompt, reference media, criteria, and seed when supported |
| Version regression | Compare a release candidate with the current version | Prompt set, API settings, and presentation protocol |
| Same-seed variants | Inspect the effect of a model or parameter change with matched randomness | Seed and every input outside the tested arm |
| Reference fidelity | Judge identity, product, composition, or motion preservation | Reference assets and transformation request |
| Rating task | Score each candidate independently on a stable rubric | Rubric, scale anchors, and media presentation |
| Treatment-arm task | Compare prompt, LoRA, guidance, or another parameter treatment | Endpoint and all settings outside the treatment |
Pairwise or rating
Pairwise review asks which candidate better satisfies the written criteria. It is usually the clearest choice for model and release comparisons. Rating asks for one score per candidate on the immutable task-pinned scale and is useful when each output must be judged against a stable rubric.
When one case has more than two candidates, a reviewer sees blind two-candidate comparisons rather than a wall of outputs. Within that Evaluation Session, each candidate is used at most once for the case, so ten candidates provide five disjoint comparisons. The Evaluation Task owner sets the minimum decisions required; when it is omitted, the default is ten or the full available capacity when fewer than ten comparisons exist.
Direct fal generation currently creates endpoint-arm pairwise Evaluation Tasks. Rating tasks and treatment-arm comparisons use finished-media import so each intended arm is explicit before review.
Controls belong in the task
- Prompt
Use the same resolved instruction unless prompt wording is the treatment being tested.
- Seed
Share a seed only when both endpoint contracts support a meaningful seed field.
- References
Attach the same source images, video, or audio to every candidate arm that depends on them.
- Presentation
Keep watch time, playback behavior, criteria, and blind ordering consistent across reviewers.
- Audience
Use a cohort that can judge the criterion and can complete the media preflight honestly.
Comparisons to avoid
Do not put candidates with different prompts, missing references, unequal durations, or incompatible tasks into the same item and call the result a model comparison. Split them into separate cases or redesign the Evaluation Task so the intended difference is the only important difference.
A parameter sweep creates more cases. It does not automatically create blinded treatment arms inside one endpoint. Import the paired outputs when the parameter itself is the decision variable.