Skip to content

Reuse the parts that make evaluations comparable

Projects organize work, while versioned Datasets, Protocols, Sampling Plans, model revisions, reviewer pools, and reviewer calibration modules make individual Evaluation Tasks repeatable. Reuse is explicit: changing one of these records creates a new version instead of silently changing earlier evidence.

Reusable evaluation objects
ObjectPurposeVersion rule
DatasetVersioned generated examples, candidate media, inputs, references, tags, and item orderPublished versions are immutable and can be inspected or reused
ProtocolThe comparison design, public rubric, controls, reference policy, and suggested sampling strategyA task binds one published Protocol snapshot while it is draft
Sampling PlanThe per-session minimum and maximum, task-wide judgment budget, pair exposure rule, and assignment strategyThe frozen plan is attached before the task becomes active
Model revisionA stable identity for a checkpoint or provider revision across tasksEndpoint links are explicit, audited, and append-only
Rater poolA reusable staffing cohort defined by saved demographic and capability filtersFilter changes create another aggregate-only pool version
Reviewer calibration or Golden ControlA private-answer module built from an appropriate published DatasetAssignments point to an immutable module version

Dataset and Protocol selection

Choose one canonical Project for every Evaluation Task. Then select the published Dataset version that defines the cases and the Protocol version that defines how those cases should be judged. The task keeps those exact snapshots even when authors later publish newer catalog versions.

Prompt Library remains independent, reusable authoring material. A generation run can apply the same prompt set to several models, and the resulting examples become a Dataset version suitable for evaluation, reviewer practice, golden controls, or calibration.

Sampling and staffing

MLOpt chooses the implemented assignment strategy, the decisions required from one reviewer, and an optional task-wide judgment budget. The management view shows the frozen plan and its fingerprint; unsupported controls are not offered as if they were active.

Rater pools separate a reusable, filter-defined reviewer cohort from one task's staffing plan. MLOpt can filter by consented role, country, language, age band, display preference, media preference, and hearing support. The catalog exposes aggregate match counts, never a directory of individual reviewers.

Reviewer calibration and golden controls

Reviewer calibration modules use published Reviewer Practice Datasets. Golden controls use published Golden Control Datasets. Answer keys, rationales, attempts, and item-level correctness remain private; evaluators see only their own assignment and aggregate result.

This is evaluator onboarding and quality calibration, not model training. Completing it does not alter a model result, and a private answer key never crosses into an evaluation export.

When cross-task results are allowed

Tasks inside one Project are easy to browse together, but their scores are not automatically combined. A cross-task group is published only when Dataset content, task inputs, Protocol and rubric, Sampling Plan, analysis method, and explicit model revisions match exactly.

Structured Dataset bundle Tasks still retain complete Task-level results, but they cannot currently publish an exact cross-task comparability snapshot because source-folder items do not carry canonical Prompt Library source-prompt lineage. Review these Tasks separately; Project grouping never merges their evidence.

The group view keeps the model ranking descriptive and reports pair-level uncertainty and multiplicity-adjusted support. If the contract differs, read the tasks separately.