AI Briefing
Sign in

Multiverse Computing Introduces ProvenanceGuard for Source-Aware Verification of MCP Agents

·2026.09.29 22:07

Key point

The new verification layer caught 138 of 139 expert-flagged unsupported claims in medical agent tests while identifying the correct source for 86% of claims with identifiable sources.

1 / 4

Details

Multiverse Computing has introduced ProvenanceGuard, a post-generation verification layer designed to address cross-source conflation in LLM agents that use the Model Context Protocol (MCP). Unlike traditional faithfulness checkers that pool evidence, ProvenanceGuard preserves source identity to ensure claims are attributed to the correct tool output, preventing scenarios where a fact is true but cited to the wrong source.

How It Works

The system operates as a black-box wrapper around MCP agents, analyzing captured traces without retraining the agent. It executes a five-step pipeline:

  • Decomposition: Breaks the agent's answer into specific claims.
  • Routing: Identifies the most relevant source for each claim using MiniLM.
  • Support Check: Verifies if the source actually supports the claim using a DeBERTa NLI verifier.
  • Attribution Check: Compares the supporting source against the source named or implied in the answer.
  • Decision: Emits a per-claim verdict and a global allow/block decision, with optional RARR-style repair for blocked answers.

Performance Metrics

In tests involving a medical agent using patient records and research articles, ProvenanceGuard demonstrated high precision in blocking unsupported claims:

  • Blocking Accuracy: Caught 138 of 139 claims that human experts deemed unsupported, with only one false negative.
  • Source Identification: Correctly identified the supporting source for 86% of claims with identifiable sources.
  • Comparison: Achieved a 0.802 F1 score for rejecting/blocking claims, outperforming source-blind baselines like MiniCheck (0.783), RAGAS Faithfulness (0.758), AlignScore (0.662), and SummaC-ZS (0.436).
  • Attribution Swaps: Successfully detected 100% of wrong attributions in a controlled test where named sources were swapped.

Limitations and Integration

While effective, the system faces challenges in distinguishing between highly similar sources, achieving only 50.3% exact source identification in a stress test with multiple similar documents. The verification overhead is approximately 0.5 seconds per answer in local configurations. The approach has already been integrated into NVIDIA NVFlow as an optional grounding-verification stage for finance agents, demonstrating its applicability in sensitive, data-heavy environments.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.