Skip to content

A workspace for controlled media evaluation

GenMedia helps teams compare generative media systems through blinded human review. An MLOpt operator defines the comparison, generates or imports matched candidates, assigns reviewers, and reads the result with coverage and uncertainty attached.

Evaluation lifecycle

From a model question to a supported result

  1. 01Define

    Question, arms, controls

  2. 02Prepare

    Prompts and references

  3. 03Produce

    Generate or import

  4. 04Review

    Blind media judgment

  5. 05Interpret

    Counts, evidence, limits

Common paths

Start with the task you need to complete

The records that make an evaluation traceable
ObjectWhat it represents
ProjectThe durable workspace for one model family, optimization program, or product question. It can contain many Evaluation Tasks over time.
Evaluation TaskOne concrete test: its Dataset version, Protocol, model or treatment arms, audience, criteria, sampling plan, and comparable items.
Evaluation SessionOne reviewer's resumable queue inside one Evaluation Task. It stores the blind assignments and submitted decisions for that reviewer.
Generation RunOne production or import batch inside an Evaluation Task. A later extension creates another run rather than replacing earlier work.
Generation JobOne provider request for one item and endpoint arm. Jobs belong to a run and are operational lineage, not separate Evaluation Tasks.
ItemOne comparable case, such as a shared prompt and reference image evaluated across two endpoints.
CandidateOne generated or imported output attached to an item.
DecisionA blind choice, tie, skip, comment, or rating recorded with its presented order and Evaluation Session.

How an Evaluation Task moves through GenMedia

  1. 1

    Write the question and decide which input or model difference you want to test.

  2. 2

    Prepare comparable prompts, seeds, reference media, criteria, and presentation rules.

  3. 3

    Generate candidates from fal endpoints or import outputs produced elsewhere.

  4. 4

    Choose the fal.ai team, Creative Engineering, or a named reviewer list as the audience.

  5. 5

    Collect blind pairwise choices or ratings on the immutable task-pinned scale, including 1 to 5 and the legacy 1 to 9 default, with media-specific controls.

  6. 6

    Read observed choices first, then inspect coverage, uncertainty, pair evidence, and failure slices before making a release decision.

The hierarchy is deliberate

A Project is not one benchmark result. It is the long-lived place where a team can run tens or hundreds of related Evaluation Tasks. Each task keeps its own result because prompts, model revisions, protocols, reviewers, and sampling may differ.

Cross-task leaderboards appear only when MLOpt explicitly pins the same immutable Dataset, Protocol, rubric, sampling, and analysis versions. Project membership alone never makes two scores comparable, and task Elo values are not averaged together.

Questions the product can answer

  • Model selection

    Which model or version is preferred under the same prompts and controlled inputs?

  • Regression review

    Did a release candidate improve the cases that matter without introducing a visible failure mode?

  • Reference fidelity

    Which output better preserves the supplied person, product, composition, or motion reference?

  • Prompt and parameter evaluations

    Do prompt variants or generation settings change the result on a matched set of cases?

  • Failure localization

    Which prompts, topics, positions, or media moments account for a broader model difference?

What a result does not prove

A preference result applies to the prompts, reviewers, criteria, and presentation protocol that produced it. It is not a universal model score. A headline rank without enough pair coverage or uncertainty control is useful for orientation, but it is not a supported winner claim.

GenMedia keeps descriptive results separate from statistical conclusions so a small Evaluation Task can still be inspected without overstating what the sample can support.