Start from evidence
Review the selected execution
Download the trial JSON and candidate from the report if you need portable references. The candidate is not an approved case or an oracle. Export the report’s saved trials with an explicit rubric using the existing Label Studio exchange:human, model or synthetic accurately.
Author the expectation and promote deliberately
Write the expected result toexpected.txt, and explain its authoritative source
and intended task behavior in rationale.txt. A rejected answer does not by itself
prove what the correct answer should be. Review that expectation separately.
For non-text outcomes, keep the independently verified outcome contract; a text
reference alone cannot replace database, tool or environment assertions.
--source-id based on actual provenance.
The recorded trace is cleared for a future execution, and a callable reference
cannot silently become its expected output. The manifest retains review evidence,
parent trial/content digests and the expectation rationale.
Promotion assigns the case to development. Once you investigate and tune
against a held-out failure, that source is no longer untouched validation evidence.
Keep its variants together and reserve new untouched sources for the final check.
Do not rename source IDs to make inspected data appear independent.

