Autonomous Multi-Agent Systems Orchestrated via n8n: Infrastructure Bottlenecks and Self-Healing Architectures
Abstract
The transition from single-shot generative models to autonomous, goal-directed agents represents a structural departure from fixed-pipeline automation. Low-code orchestration platforms such as n8n increasingly supply the operational substrate for this transition, handling function invocation, state persistence, and inter-agent communication while the reasoning itself remains delegated to a large language model (LLM). This paper investigates, through two controlled case-study experiments built on a Gemini 2.0 Flash agent orchestrated in n8n, whether the sustainable operating depth of such systems is bounded by the reasoning capacity of the model or by the surrounding infrastructure. We introduce the Agentic Depth Coefficient (ADC), a formal metric for the number of consecutive successful reasoning-action cycles an agent completes before terminal failure, and the Agentic Thirst equation, which expresses the cumulative API demand of a multi-turn agentic task as a function of ADC, tool calls, and reasoning sub-steps. In Experiment 1, an openended literature-synthesis task terminated at ADC = 6 due to an HTTP 429 rate-limit response rather than any indication of reasoning failure, evidencing that free-tier infrastructure quotas, not model cognition, defined the operating ceiling. In Experiment 2, a secondary Healer agent achieved full recovery from all deterministic parsing failures induced by semantically valid but syntactically irregular natural-language input, a behavior we formalize as Semantic Override. Based on these findings we propose a Split-Stack Orchestration Architecture that allocates strategic reasoning, execution, and low-cost repair tasks to separate model tiers, reducing sanitization cost by an estimated 99.2% while decoupling execution-tier rate-limit exhaustion from strategic-tier availability. As single-instance case studies, these experiments are intended to demonstrate the existence and mechanism of infrastructure-bound and semantically self-healing behavior rather than to establish statistically generalizable failure rates; we discuss this scope explicitly in Section IX and outline the replication needed to establish the latter.