Skip to content

Separate observed choices from supported conclusions

GenMedia shows what reviewers submitted even when the sample is small, then separates those counts from the subset that is usable for statistical analysis. This prevents an early preference from being presented as a conclusive model result.

Three layers of a result
LayerQuestion it answersWhat to check
ObservedWhat did the reviewers submit?Choice counts, ties, skips, ratings, items, and evaluator count
EligibleWhich decisions can enter analysis?Completed assignments, usable media sessions, decisive records, and excluded events
StatisticalHow strong is the comparative evidence?Pair coverage, interval, exact test, multiplicity correction, and unresolved pairs

Basic and Advanced results

Basic results give a completed fal.ai Standard reviewer a compact view of their own choices and the available team comparison. Advanced results add model identities, pair matrices, intervals, prompt and taxonomy slices, bias checks, reviewer agreement, and sampling guidance for authorized roles.

Elo is an operational ordering. It helps scan a result, but the pair decision and uncertainty determine whether the evidence supports a winner claim.

Guest evidence and the archive seal

Completed unclaimed guest evidence enters official analysis at up to 70 percent influence, using a 0.70 base multiplied by its private terminal behavior-quality factor. Claiming the Session into a verified account that is still pending approval preserves that 0.70 base. It does not create a reward.

If the owner becomes an active Standard reviewer before the result is sealed, the base may upgrade to 1.00 while the same behavior-quality factor remains in force. Once the official archive is sealed, later claims, approvals, role changes, or reliability changes cannot rewrite the archived weight or Elo result.

How many decisions are enough

There is no universal sample size. One clean decision is enough to populate descriptive tables, but not to support a broad model claim. The required count depends on how lopsided the choices are and how many pairs or slices are being tested.

For one two-model pair with 20 eligible decisive choices, a split of 15 to 5 or stronger clears the current pair rule; 14 to 6 does not. Ties do not enter that exact decisive count. The 97-per-pair figure shown in planning is a conservative precision target near plus or minus 10 percentage points, not a gate that hides results before vote 97.

Prompt and topic findings

Prompt-level rows show where one pair behaves differently on a particular case. Topic and use-case views group eligible tagged cases to find recurring behavior. These sections remain empty when there are no eligible decisions or the repeated evidence is too small to survive the relevant multiple-comparison check.

An empty finding section means the available evidence does not clear its rule. It does not prove that every prompt or topic behaves the same.

Read results in this order

  1. 1

    Check submitted and eligible counts, evaluator count, coverage, ties, skips, and excluded sessions.

  2. 2

    Read the head-to-head pair matrix and uncertainty before the ranking.

  3. 3

    Inspect prompt, topic, position, and reviewer-agreement sections for a concentrated failure or Evaluation Task defect.

  4. 4

    Collect the next sample where uncertainty or pair coverage remains unresolved.