Skip to content

Decide what may vary

A useful benchmark starts with a narrow question and a clear control condition. GenMedia presets describe the comparison intent, the inputs that should remain fixed, and the review mode that fits the decision.

Common Evaluation Task designs
DesignUse it forKeep constant
Model versus modelCompare two endpoint arms on the same casesPrompt, reference media, criteria, and seed when supported
Version regressionCompare a release candidate with the current versionPrompt set, API settings, and presentation protocol
Same-seed variantsInspect the effect of a model or parameter change with matched randomnessSeed and every input outside the tested arm
Reference fidelityJudge identity, product, composition, or motion preservationReference assets and transformation request
Rating taskScore each candidate independently on a stable rubricRubric, scale anchors, and media presentation
Treatment-arm taskCompare prompt, LoRA, guidance, or another parameter treatmentEndpoint and all settings outside the treatment

Pairwise or rating

Pairwise review asks which candidate better satisfies the written criteria. It is usually the clearest choice for model and release comparisons. Rating asks for one score per candidate on the immutable task-pinned scale and is useful when each output must be judged against a stable rubric.

When one case has more than two candidates, a reviewer sees blind two-candidate comparisons rather than a wall of outputs. Within that Evaluation Session, each candidate is used at most once for the case, so ten candidates provide five disjoint comparisons. The Evaluation Task owner sets the minimum decisions required; when it is omitted, the default is ten or the full available capacity when fewer than ten comparisons exist.

Direct fal generation currently creates endpoint-arm pairwise Evaluation Tasks. Rating tasks and treatment-arm comparisons use finished-media import so each intended arm is explicit before review.

Controls belong in the task

  • Prompt

    Use the same resolved instruction unless prompt wording is the treatment being tested.

  • Seed

    Share a seed only when both endpoint contracts support a meaningful seed field.

  • References

    Attach the same source images, video, or audio to every candidate arm that depends on them.

  • Presentation

    Keep watch time, playback behavior, criteria, and blind ordering consistent across reviewers.

  • Audience

    Use a cohort that can judge the criterion and can complete the media preflight honestly.

Comparisons to avoid

Do not put candidates with different prompts, missing references, unequal durations, or incompatible tasks into the same item and call the result a model comparison. Split them into separate cases or redesign the Evaluation Task so the intended difference is the only important difference.

A parameter sweep creates more cases. It does not automatically create blinded treatment arms inside one endpoint. Import the paired outputs when the parameter itself is the decision variable.