Localizing Post-Wire Semantic Changes in MCP Agent Frameworks
Abstract
Valid Model Context Protocol (MCP) messages do not guarantee that an agent framework preserves the distinctions downstream software needs. We present a differential testing method that follows a fixed tool result through each framework's public interfaces and checks explicit consumer requirements. Applied to 18 designed fixtures in four pinned Python integrations, the method identifies 13 unique fixture-task divergences involving structured values, declared errors, and rich content. Google ADK satisfies every primary contract in its observed path; the other integrations show interface-specific changes or an execution failure. Whether a change matters depends on the consumer: treating absent optional fields as equivalent to null explains most of OpenAI's strict rich-content failures. In an exploratory replay, parsing JSON text recovers more structured values but also returns incorrect values and values for fields absent at the source. Documented settings help in conflicting-output stress cases without restoring a separate structured slot. Fault challenges expose an oracle weakness and test its repair on fresh cases. The study provides reproducible, interface-specific evidence from controlled cases, not production failure rates or model-behavior measurements. Its practical implication is that post-wire regression tests need to specify both the information required and how the consumer reads it.