Create about 30 test claims, balanced across the three labels. Manually establish expected labels before running the system. Include paraphrases, changed numbers, changed dates, negation, absent evidence, and conflicting evidence.
Report label accuracy and a confusion matrix, inspect errors, and check whether explanations are supported by the cited records. Compare retrieval-only results, an API-only assessment without the collection, and the integrated approach. Do not treat the model's own judgement as the test answer.
Test invalid inputs and API failures. The system should not present an unsupported confident verdict when relevant evidence is missing.
Proposed success criterion: assessments are grounded in the supplied records, sources are traceable, and uncertainty is handled explicitly. The group should agree on a numerical target before testing and report whether it was reached.