Browse documentationReusable controls
Reuse the parts that make evaluations comparable
Projects organize work, while versioned Datasets, Protocols, Sampling Plans, model revisions, reviewer pools, and reviewer calibration modules make individual Evaluation Tasks repeatable. Reuse is explicit: changing one of these records creates a new version instead of silently changing earlier evidence.
| Object | Purpose | Version rule |
|---|---|---|
| Dataset | Versioned generated examples, candidate media, inputs, references, tags, and item order | Published versions are immutable and can be inspected or reused |
| Protocol | The comparison design, public rubric, controls, reference policy, and suggested sampling strategy | A task binds one published Protocol snapshot while it is draft |
| Sampling Plan | The per-session minimum and maximum, task-wide judgment budget, pair exposure rule, and assignment strategy | The frozen plan is attached before the task becomes active |
| Model revision | A stable identity for a checkpoint or provider revision across tasks | Endpoint links are explicit, audited, and append-only |
| Rater pool | A reusable staffing cohort defined by saved demographic and capability filters | Filter changes create another aggregate-only pool version |
| Reviewer calibration or Golden Control | A private-answer module built from an appropriate published Dataset | Assignments point to an immutable module version |
Dataset and Protocol selection
Choose one canonical Project for every Evaluation Task. Then select the published Dataset version that defines the cases and the Protocol version that defines how those cases should be judged. The task keeps those exact snapshots even when authors later publish newer catalog versions.
Prompt Library remains independent, reusable authoring material. A generation run can apply the same prompt set to several models, and the resulting examples become a Dataset version suitable for evaluation, reviewer practice, golden controls, or calibration.
Sampling and staffing
MLOpt chooses the implemented assignment strategy, the decisions required from one reviewer, and an optional task-wide judgment budget. The management view shows the frozen plan and its fingerprint; unsupported controls are not offered as if they were active.
Rater pools separate a reusable, filter-defined reviewer cohort from one task's staffing plan. MLOpt can filter by consented role, country, language, age band, display preference, media preference, and hearing support. The catalog exposes aggregate match counts, never a directory of individual reviewers.
Reviewer calibration and golden controls
Reviewer calibration modules use published Reviewer Practice Datasets. Golden controls use published Golden Control Datasets. Answer keys, rationales, attempts, and item-level correctness remain private; evaluators see only their own assignment and aggregate result.
This is evaluator onboarding and quality calibration, not model training. Completing it does not alter a model result, and a private answer key never crosses into an evaluation export.
When cross-task results are allowed
Tasks inside one Project are easy to browse together, but their scores are not automatically combined. A cross-task group is published only when Dataset content, task inputs, Protocol and rubric, Sampling Plan, analysis method, and explicit model revisions match exactly.
Structured Dataset bundle Tasks still retain complete Task-level results, but they cannot currently publish an exact cross-task comparability snapshot because source-folder items do not carry canonical Prompt Library source-prompt lineage. Review these Tasks separately; Project grouping never merges their evidence.
The group view keeps the model ranking descriptive and reports pair-level uncertainty and multiplicity-adjusted support. If the contract differs, read the tasks separately.