Browse documentationOverview
A workspace for controlled media evaluation
GenMedia helps teams compare generative media systems through blinded human review. An MLOpt operator defines the comparison, generates or imports matched candidates, assigns reviewers, and reads the result with coverage and uncertainty attached.
Evaluation lifecycle
From a model question to a supported result
- 01Define
Question, arms, controls
- 02Prepare
Prompts and references
- 03Produce
Generate or import
- 04Review
Blind media judgment
- 05Interpret
Counts, evidence, limits
Common paths
Start with the task you need to complete
Select two fal endpoints, bind the same 20 cases, generate matched candidates, and open blind review.
Read guide ReviewerComplete a blind reviewInspect the source material, watch or compare both candidates, and submit the choice the criteria support.
Read guide Results readerInterpret an evaluationSeparate submitted choices from eligible evidence, then inspect uncertainty, coverage, and failure slices.
Read guide| Object | What it represents |
|---|---|
| Project | The durable workspace for one model family, optimization program, or product question. It can contain many Evaluation Tasks over time. |
| Evaluation Task | One concrete test: its Dataset version, Protocol, model or treatment arms, audience, criteria, sampling plan, and comparable items. |
| Evaluation Session | One reviewer's resumable queue inside one Evaluation Task. It stores the blind assignments and submitted decisions for that reviewer. |
| Generation Run | One production or import batch inside an Evaluation Task. A later extension creates another run rather than replacing earlier work. |
| Generation Job | One provider request for one item and endpoint arm. Jobs belong to a run and are operational lineage, not separate Evaluation Tasks. |
| Item | One comparable case, such as a shared prompt and reference image evaluated across two endpoints. |
| Candidate | One generated or imported output attached to an item. |
| Decision | A blind choice, tie, skip, comment, or rating recorded with its presented order and Evaluation Session. |
How an Evaluation Task moves through GenMedia
- 1
Write the question and decide which input or model difference you want to test.
- 2
Prepare comparable prompts, seeds, reference media, criteria, and presentation rules.
- 3
Generate candidates from fal endpoints or import outputs produced elsewhere.
- 4
Choose the fal.ai team, Creative Engineering, or a named reviewer list as the audience.
- 5
Collect blind pairwise choices or ratings on the immutable task-pinned scale, including 1 to 5 and the legacy 1 to 9 default, with media-specific controls.
- 6
Read observed choices first, then inspect coverage, uncertainty, pair evidence, and failure slices before making a release decision.
The hierarchy is deliberate
A Project is not one benchmark result. It is the long-lived place where a team can run tens or hundreds of related Evaluation Tasks. Each task keeps its own result because prompts, model revisions, protocols, reviewers, and sampling may differ.
Cross-task leaderboards appear only when MLOpt explicitly pins the same immutable Dataset, Protocol, rubric, sampling, and analysis versions. Project membership alone never makes two scores comparable, and task Elo values are not averaged together.
Questions the product can answer
- Model selection
Which model or version is preferred under the same prompts and controlled inputs?
- Regression review
Did a release candidate improve the cases that matter without introducing a visible failure mode?
- Reference fidelity
Which output better preserves the supplied person, product, composition, or motion reference?
- Prompt and parameter evaluations
Do prompt variants or generation settings change the result on a matched set of cases?
- Failure localization
Which prompts, topics, positions, or media moments account for a broader model difference?
What a result does not prove
A preference result applies to the prompts, reviewers, criteria, and presentation protocol that produced it. It is not a universal model score. A headline rank without enough pair coverage or uncertainty control is useful for orientation, but it is not a supported winner claim.
GenMedia keeps descriptive results separate from statistical conclusions so a small Evaluation Task can still be inspected without overstating what the sample can support.