Browse documentationQuickstart
Run a two-model, 20-item comparison
This example follows the common path: two fal endpoints receive the same 20 prompts, the fal.ai team reviews the resulting pairs without model names, and the MLOpt operator checks both the observed preference and the statistical evidence.
| Setting | Example value |
|---|---|
| Question | Which model is preferred for this prompt set? |
| Mode | Blind pairwise comparison |
| Candidates | Two endpoint arms |
| Cases | 20 prompts, one item per prompt |
| Controls | Shared prompt and reference inputs; shared seed when both endpoints support it |
| Audience | fal.ai team |
Create the task
- 1
Open New evaluation, choose the project and media modality, then select Model versus model as the comparison design.
- 2
Add two fal catalog endpoints. Private or ephemeral fal endpoints are available under the advanced endpoint control.
- 3
Select 20 compatible prompts or add a saved prompt pack. Attach every required image, video, or audio reference before continuing.
- 4
Review the resolved input fields for each endpoint. Keep shared values bound to the same source and set only the fields the Evaluation Task intends to vary.
- 5
Inspect the generation preview. Twenty prompts across two endpoints should produce 20 items and 40 generation jobs.
- 6
Confirm the fal.ai team audience, criteria, seed policy, and variant notes, then start generation.
Review generation and start evaluation
The task page updates as jobs move through queued, generating, complete, or failed states. Completed outputs appear inside their item and endpoint cell. A failed output should be resolved before the task is treated as a complete pairwise evaluation.
Reviewers receive randomized blind labels rather than endpoint names. Reference inputs remain available from the evaluation header when the task depends on an image, video, or audio source.
Read the result
- 1
Start with submitted choices and evaluator count. This is the direct record of what people selected.
- 2
Check how many decisions are eligible for statistical analysis. A submitted choice can be excluded when its session or media checks are not usable.
- 3
Inspect the pair matrix and uncertainty. Do not call a winner from Elo order alone.
- 4
Read prompt and topic slices for concentrated failures, then collect more evidence where the Evaluation Task remains unresolved.