Building with AI

ProvenanceGuard verifies source attribution in MCP agents

A new verification layer detects when AI agents attribute facts to the wrong data source, preventing cross-source conflation in multi-tool workflows.

Researchers have introduced ProvenanceGuard, a verification method designed to ensure that AI agents attribute facts to the correct data sources. Published on September 29, 2026, by Hugging Face, this approach addresses errors where true statements are linked to the wrong tool output in Model Context Protocol (MCP) systems.

What happened

Modern AI agents often use the Model Context Protocol to access multiple tools simultaneously. An agent might query a database, search a document repository, and inspect structured records before generating a single response. While existing evaluation methods check if a claim is supported by any available evidence, they often fail to verify if the claim is supported by the specific source the agent cites. This leads to a failure mode known as cross-source conflation, where a fact is true but attributed to the wrong origin.

ProvenanceGuard acts as a post-generation verification layer that sits on top of an MCP agent without requiring retraining. It analyzes the captured trace of tool outputs and source IDs to determine if each claim in the answer matches its stated or implied source. The system decomposes answers into individual claims, routes each claim to the most relevant source, and checks for support using natural language inference models. If a mismatch is detected, the system can block the answer or trigger a repair process.

The researchers tested this method on a medical agent that accessed patient records and research articles. In a study involving 281 real traces, human experts evaluated 361 claims from a held-out set. ProvenanceGuard identified 138 out of 139 claims that should have been rejected due to incorrect attribution or lack of support. It achieved a reject/block F1 score of 0.802, outperforming source-blind baselines like MiniCheck and RAGAS Faithfulness while providing per-claim source verdicts.

How it works

ProvenanceGuard preserves source identity throughout the verification pipeline rather than pooling all evidence into a single context. When an agent produces an answer, the system breaks it down into specific claims. It then uses embedding models, such as MiniLM, to find the source most relevant to each claim. A DeBERTa natural language inference model checks whether that specific source actually supports the claim. Finally, the system compares the supporting source with the one named or implied in the answer text.

The verification process includes a calibrated decision step that combines these signals to allow or block the answer. If an answer is blocked, a repair loop similar to RARR can attempt a source-grounded revision or provide a safe fallback. In the reported local configuration, this process adds approximately half a second of overhead per answer. The system relies on local models for controlled offline processing, though the architecture can be adapted for cloud-based services.

Key details

  • ProvenanceGuard detects cross-source conflation, where a true fact is attributed to the wrong tool output.
  • The system achieved a 0.802 F1 score for blocking incorrect claims, surpassing source-blind verifiers.
  • It correctly identified the right source for claims about 86% of the time in tests with distinct sources.
  • The verification layer adds roughly 500 milliseconds of latency per answer in the tested local setup.
  • In a controlled test with 50 swapped source attributions, the method caught all 50 errors.
  • The approach works with existing MCP agents by analyzing traces without requiring model retraining.

Why it matters

For developers building reliable document AI systems, ensuring factual accuracy is not enough if the provenance is wrong. In sensitive fields like healthcare or finance, citing a general policy document instead of a specific patient record can lead to serious compliance issues or safety risks. Source-aware verification provides the granularity needed to trust automated agents in high-stakes environments. It transforms factuality from a binary check into a detailed audit trail that links every claim to its origin.

This method also improves the debugging and maintenance of agentic workflows. By exposing which tool output supports each claim, engineers can identify weak links in their retrieval pipelines. It allows teams to distinguish between an agent that lacks information and one that misattributes information. This distinction is critical for refining prompt strategies and selecting appropriate tools for specific tasks.

What you can do

  • Implement source ID tracking in your MCP tool outputs to enable downstream verification.
  • Evaluate your current RAG or agent system for cross-source conflation using held-out test sets.
  • Consider adding a post-generation verification step that checks claim-to-source alignment.
  • Use local NLI models for initial prototyping of source-aware verification before scaling to cloud services.
  • Design your agent prompts to explicitly state the source of each claim to facilitate automated checking.
  • Test your system with swapped source attributions to measure its sensitivity to provenance errors.

Tools from the Bytechap store

Keep reading

All stories