Skip to content
#edge computing Dataset Open access

Large Language Model-Enhanced Multi-Agent Reinforcement Learning for Autonomous Enterprise Workflow Orchestration via Graph Semantic Routing

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

# LeMARLA **L**arge Language Model-**E**nhanced **M**ulti-**A**gent **R**einforcement **L**earning for **A**utonomous enterprise workflow orchestration via graph semantic routing. LeMARLA couples a large language model, a heterogeneous graph attention network and a multi-agent policy under centralized training with decentralized execution. It turns a natural-language business instruction into an executable DAG of subtasks, routes each subtask to the employee best able to handle it, and adapts its plan online whenever a subtask fails. --- ## Highlights * Task decomposition: LLaMA-2-13B-chat generates the ordered subtask list and its precedence structure. A learned dependency scorer (Eq. 3) rewrites edges based on subtask semantics rather than surface heuristics.* Semantic routing: a three-relation heterogeneous graph (task, employee, skill) is embedded with type-aware attention (Eq. 4-7). The Eq. 8 routing score fuses cosine similarity of task and employee embeddings with a load-balance penalty.* Multi-agent policy: independent PPO actors observe local workload signals, share a centralized critic during training, and act with a semantic prior injected into the softmax (Eq. 9). Trained with the clipped surrogate of Eq. 12.* Online feedback: failure signals from the execution engine trigger targeted re-decomposition of the failed subgraph without recomputing completed subtasks.* Log-replay evaluation: a duration model and a failure model, fitted on the training split under log-normal likelihood, drive counterfactual outcomes on the withheld test split. Ten independent seeds, Holm-Bonferroni-corrected significance tests. --- ## Repository content Two archives sit at the root of the repository: ```lemarla-code.zip # Source code, configurations, tests, cached model outputslemarla-data.zip # Enterprise workflow dataset (11 xlsx tables + docs)``` Downloading a single ZIP is enough for most users: `lemarla-code.zip` alreadyships the full dataset under `data/lemarla_dataset/`. `lemarla-data.zip` isprovided for those who only need the tables. ### Code archive layout ```lemarla-code/├── README.md├── requirements.txt├── configs/default.yaml # Every hyperparameter (learning rates, seeds, splits)├── data/lemarla_dataset/ # 11 xlsx tables of the operational log├── src/ # 10 modules mapped to Eq. 1-13├── scripts/ # prepare_data / train / evaluate / plot / run_all├── tests/test_smoke.py # 12 unit checks over every module└── outputs/ # Tables, figures (PDF + PNG) and checkpoints``` ### Data archive layout ```lemarla-data/├── README.md├── 01_departments.xlsx├── 02_job_categories.xlsx├── 03_skill_taxonomy.xlsx├── 04_subtask_category_taxonomy.xlsx├── 05_employee_master_table.xlsx├── 06_employee_skill_certifications.xlsx├── 07_collaborates_edges.xlsx├── 08_workflow_instances_raw.xlsx├── 09_subtask_execution_log.xlsx├── 10_dependency_edges.xlsx└── 11_annotated_reference_set.xlsx``` The `data/lemarla_dataset/README.md` inside the archive documents everycolumn and lists the row counts (326 employees, 48 skills, 21,435 rawinstances, 8,742 retained, 41,537 subtasks, 29,086 dependency edges, and a300-instance / 1,521-subtask / 1,063-edge annotated reference set). --- ## Installation ```bash# 1. Clone the repository and unpack the code archivegit clone https://github.com/ /lemarla.gitcd lemarlaunzip lemarla-code.zip -d . # 2. Create a virtual environment and install the pinned requirementspython -m venv .venvsource .venv/bin/activate # Linux / macOSpip install -r lemarla-code/requirements.txt``` Requirements (from `lemarla-code/requirements.txt`): ```torch==2.1.0 torchvision==0.16.0numpy>=1.24.0 pandas>=2.0.0scikit-learn>=1.3.0 scipy>=1.11.0matplotlib>=3.7.0 seaborn>=0.12.0networkx>=3.1 openpyxl>=3.1.0tqdm>=4.65.0 PyYAML>=6.0sentence-transformers>=2.2.0ortools>=9.10.0transformers>=4.38.0torch-geometric>=2.5.0``` The pipeline runs on both GPU (recommended for full training) and CPU(sufficient for evaluation, plotting and unit tests). --- ## Quick start ```bashcd lemarla-code # Sanity: eleven module-level checkspython tests/test_smoke.py # Reproduce every experiment end to endpython scripts/run_all_experiments.py # includes MAPPO trainingpython scripts/run_all_experiments.py --skip-train # tables + figures only # Or run each stage individuallypython scripts/prepare_data.py # writes outputs/tables/table3_cascade.jsonpython scripts/train.py # writes outputs/checkpoints/{simulator.joblib, lemarla.pt}python scripts/evaluate.py # writes six xlsx tables under outputs/tables/python scripts/plot_all_figures.py # writes fourteen figures under outputs/figures/``` --- ## Reproduced results Ten independent seeds, 95% CIs computed with t-critical 2.262, Holm-Bonferronicorrected p-values. | Method | TCR (%) | E2EL (min) | LBI (1/min) | RA@1 (%) | RA@3 (%) | DL (ms) ||--------------------|----------------------|----------------------|----------------------|----------------------|----------------------|---------|| Rule-Based | 78.4 [77.8, 79.0] | 52.1 [51.0, 53.2] | 0.58 [0.57, 0.59] | 38.5 [37.8, 39.2] | 58.7 [58.3, 59.1] | 1.2 || Vanilla-MARL | 83.2 [82.6, 83.8] | 45.3 [44.3, 46.3] | 0.63 [0.62, 0.64] | 40.2 [39.6, 40.8] | 56.9 [56.1, 57.7] | 38.7 || LLM-Only | 84.6 [83.9, 85.3] | 43.5 [41.8, 45.2] | 0.48 [0.46, 0.50] | 52.4 [51.6, 53.2] | 72.3 [71.4, 73.2] | 1250.0 || GNN-MARL | 85.8 [85.0, 86.6] | 38.2 [37.2, 39.2] | 0.62 [0.60, 0.64] | 48.6 [47.9, 49.3] | 69.8 [68.8, 70.8] | 42.3 || TDAG-Assign | 87.4 [86.4, 88.4] | 35.1 [34.2, 36.0] | 0.57 [0.56, 0.58] | 55.1 [54.4, 55.8] | 77.4 [76.8, 78.0] | 852.0 || TDAG-Assign-LB | 87.1 [86.4, 87.8] | 35.4 [34.5, 36.3] | 0.66 [0.64, 0.68] | 54.8 [54.1, 55.5] | 77.0 [75.9, 78.1] | 863.0 || CP-RCPSP | 86.2 [85.7, 86.7] | 33.8 [32.7, 34.9] | 0.69 [0.67, 0.71] | 53.7 [53.0, 54.4] | 76.5 [75.7, 77.3] | 1783.0 || SOP-Orchestrator | 86.5 [85.9, 87.1] | 37.8 [36.9, 38.7] | 0.51 [0.49, 0.53] | 54.9 [54.2, 55.6] | 77.2 [76.2, 78.2] | 974.0 || **LeMARLA** | **91.7 [91.3, 92.1]**| **30.4 [29.0, 31.8]**| **0.71 [0.69, 0.73]**| **60.4 [59.7, 61.1]**| **82.6 [81.8, 83.4]**| **47.6**|| Historical | 88.5 [88.4, 88.6] | 41.2 [41.1, 41.3] | – | – | – | – | All eight baselines are separated from LeMARLA at Holm-p < 0.001 withCohen's d ≥ 3.9 on TCR. Log-replay simulator diagnostics on the 4,154-subtask validation split: ```Duration model: MAE = 1.68 min, RMSE = 2.09 min, R² = 0.813, KS = 0.031Failure model : AUC = 0.877, Brier = 0.075, ECE = 0.013``` --- ## Directory of outputs | Path | Content ||---------------------------------------------------------|--------------------------------------------|| `lemarla-code/outputs/tables/table3_cascade.json` | Preprocessing cascade || `lemarla-code/outputs/tables/Table_5_main_comparison.xlsx` | Ten-method comparison || `lemarla-code/outputs/tables/Table_6_decomposition_quality.xlsx` | Boundary / dependency P*, R*, F1* || `lemarla-code/outputs/tables/Table_7_ablation.xlsx` | Four-way ablation || `lemarla-code/outputs/tables/Table_8_scalability.xlsx` | 50-2000 employee sweep || `lemarla-code/outputs/tables/Table_9_backbones.xlsx` | LLaMA-2/3.1, Qwen2.5 backbones || `lemarla-code/outputs/tables/Table_10_relative_contribution.xlsx` | Percent contribution of each module|| `lemarla-code/outputs/figures/Figure_*.pdf` | 14 figures || `lemarla-code/outputs/figures_png/Figure_*.png` | Rasterised previews of every figure || `lemarla-code/outputs/checkpoints/simulator.joblib` | Trained duration + failure models | --- ## Correspondence between code and paper equations | Manuscript element | Code location ||-------------------------------------------|-----------------------------------------------------|| Eq. 1 Workflow orchestration objective | `src/mappo.py::joint_reward` || Eq. 2 Two-layer projection | `src/decomposition.py::TaskProjection` || Eq. 3 Dependency scoring | `src/decomposition.py::DependencyScorer` || Eq. 4 Type-specific projection | `src/het_gat.py::TypeSpecificProjection` || Eq. 5 Node-level attention | `src/het_gat.py::TypeAwareGATLayer` || Eq. 6 ELU-activated aggregation | `src/het_gat.py::TypeAwareGATLayer.forward` || Eq. 7 Semantic-level fusion | `src/het_gat.py::SemanticLevelAttention` || Eq. 8 Routing score | `src/het_gat.py::routing_score` || Eq. 9 Softmax policy with prior | `src/mappo.py::MAPPOPolicy.action_distribution` || Eq. 10 Joint reward decomposition | `src/mappo.py::MAPPOPolicy.joint_reward` || Eq. 11 Advantage from centralized critic | `src/mappo.py::MAPPOPolicy.advantage` || Eq. 12 Clipped PPO surrogate | `src/mappo.py::MAPPOPolicy.update` || Eq.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.