Evaluation
Captured output
An answer a system already produced, stored on a test case so a person can rate it as it stands.
Also written captured outputs, captured answer, captured response.
The answer came out of your system, an earlier version of it, or somewhere else entirely. It is the thing being judged, and the question it answers is how good that response was.
That is a different question from the one an expected output answers. An expected output is the standard a fresh response gets graded against; a captured output is a response already in hand. A case can carry both, and they support different work — human review reads the captured answer, a judge run grades a new answer against the standard.
Starting with captured outputs is a reasonable way to begin, because they exist already. You can put fifty real replies in front of reviewers before anyone has written a single standard, and what the reviewers disagree about tells you what the standard needs to say.
Keeping the two in separate fields is what lets you add the second later without rebuilding the set.