The MCP provides a bounded, testable, and reproducible foundation for closed-loop agentic instrumentation research by connecting local LLMs to scientific instruments through the Model Context Protocol.
Abstract
Large language models (LLMs) can plan tool-mediated scientific work, but scientific instruments remain difficult to connect to such agents: vendor APIs may load only inside acquisition host processes, facilities may prohibit cloud-hosted agents, and natural-language interfaces can emit physically unreasonable arguments. We present a method for connecting local LLMs to scientific instruments through the Model Context Protocol (MCP). It combines: (1) a schema-bound tool surface that validates requests against physical bounds before adapter dispatch; (2) a vendor-neutral, host-process adapter pattern separating language-side reasoning from instrument-side execution; (3) a persistent lifecycle for long-running live-processing jobs; and (4) MCP-prompt-registered skills that compose typed tools into reusable multi-step protocols. Our open-source reference server exposes 30 typed tools, 5 live-job types, and 6 skills through a physics-plausible simulator implementing the same protocol surface. Validation is software-only: all 120 hardware-independent tests pass deterministically, while 15 local-LLM integration tests pass 12-15 of 15 across runs because of model nondeterminism. A preliminary single-run probe across five open tool-calling LLMs indicates that the schema-bound interface can be driven locally by small open-weight models without cloud dependency; it is not a benchmark and has no confidence intervals. The method provides a bounded, testable, and reproducible foundation for closed-loop agentic instrumentation research.
Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two key challenges: brittle multi-step and multi-turn reasoning caused by incompatible tool output types and API schemas, and performance degra...
Hai-Bo Jin, Sui-Jin Wang, Xuchen Yu et al.· 0 citations
It is argued that GraphQL constitutes a principled, testable alternative to function calling for agentic systems, combining lower cost, stronger safety, and improved cognitive robustness.
Viktor Zhakhalov· CEUR Workshop Proceedings, V...· 0 citations
This demo presents SAGE as a pipeline-native runtime system and demonstrates it through an OPC-facing control plane and includes a compact distributed comparison against a LangChain RPC baseline over a shared 15-workload RAG suite, where SAGE nearly doubles full-RAG throughput under matched 8-node settings.
Jun Liu, Shu-Hao Zhang· Workshop Proceedings of the...· 0 citations
Model-Driven Engineering (MDE) offers strong abstractions but remains hindered in practice by complex tooling and steep learning curves. Recent advances in Large Language Models (LLMs) promise intuitive, natural-language interaction with modeling environments; however, existing approaches—such as direct API interaction...
Andreas Hell, Martin Fleck, Philip Langer et al.· Proceedings of the ACM/IEEE...· 2 citations
Metis, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects is presented, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects.
We ask whether AI agents powered by locally deployed large language models can reliably automate expert-defined hardware design workflows in an industry-realistic tool-calling setting. In these environments, engineers issue repetitive, dependency-ordered operations---such as creating components, adding ports, and wirin...
Leonardo Liparulo, F. Pierri· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.