Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Hugging Face
ProvenanceGuard was proposed as a post-generation verification layer for MCP-based LLM agents to prevent cross-source conflation, where evidence supports a claim but the answer attributes it to the wrong MCP tool output. In a medical-agent evaluation with 281 real traces and a test set where 361 claims were reviewed by experts, experts said 139 claims should not pass and ProvenanceGuard caught 138. It keeps tool-source identity through claim checking, emits per-claim source verdicts plus an allow/block decision, and runs a RARR-style repair loop when answers are blocked, instead of relying on source-blind faithfulness scoring.