mini-v4 suite uses 30 seeds per family
(510 cases); mini-v4-sample uses 10 per family (170 cases).
Reproduce and inspect
Interpreting failures
Scoring checks the procedural expected answer. A family can also listforbidden_answers that identify the particular lure the model followed; this
is diagnostic and does not replace the primary correctness score. Refusals and
provider/API errors are recorded separately from ordinary wrong answers.
For text-layer traps, run a matched pixels-only experiment to separate provider
PDF ingestion from visual model behavior:
Suite registry
The hashes above identify the registries shipped by PDF Hell 0.6.1. The JSON
run artifact is authoritative for a particular experiment.

