AI & Computational Science

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

How the science connects

Natural language p…Information retrie…Fact-checking

AI Insight

This preprint introduces ProvenanceGuard, a verification system that checks whether LLM agents correctly attribute information to its actual source when answering questions using multiple data sources through the Model Context Protocol (MCP). The system decomposes answers into atomic claims, verifies each claim against source-specific evidence, and identifies "cross-source conflation" where correct information is incorrectly attributed to the wrong source. Testing on 281 medical-domain traces showed the system achieved 0.802 F1 score for blocking incorrect answers and 0.858 accuracy in source attribution, successfully detecting all attribution errors in controlled tests.


In healthcare and other high-stakes domains where source provenance is critical, this addresses a previously overlooked failure mode where LLM agents may provide factually correct information but attribute it to the wrong database, clinical record, or reference source. This could prevent dangerous scenarios like attributing drug interaction data to the wrong formulary or patient record.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools. Standard factuality metrics usually test whether an answer is supported by pooled evidence, missing a provenance-sensitive failure mode: a claim may be supported somewhere while being attributed to the wrong source. We call this cross-source conflation.
We introduce ProvenanceGuard, a source-aware verifier for MCP-grounded answers. It consumes captured MCP traces with stable tool IDs, source IDs, and raw outputs; decomposes answers into atomic claims; routes claims to source-specific evidence; checks support with NLI and a token-alignment proxy; compares stated attribution with the routed source; and returns per-claim verdicts plus an answer-level allow/block decision. Blocked answers can be repaired with retrieval-augmented answer revision and re-verified.
We evaluate on 281 medical-domain MCP-agent traces. A 266-trace adjudicated subset yields 2,325 LLM-assisted claim labels split by trace; 361 held-out labels are human-verified. On the 40-trace held-out split, ProvenanceGuard achieves block F1 0.802 and source accuracy 0.858 over 260 source-eligible claims, outperforming source-blind baselines that do not emit claim-to-source IDs. On a harder multi-source benchmark it reaches block F1 0.846, while source-plus-relation accuracy drops to 0.229, showing that exact source ownership remains difficult with semantically close sources. Repair-and-reverify resolves all blocked answers in the full trace set, often via conservative fallback. In 50 controlled clinical conflation probes, ProvenanceGuard detects all injected attribution swaps with no retained wrong attribution. These results show that source attribution is an independent axis for factuality verification in MCP-based agents.

Source: ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents