Skip to content
#reinforcement learning Dataset Open access

Large Language Model-Enhanced Multi-Agent Reinforcement Learning for Autonomous Enterprise Workflow Orchestration via Graph Semantic Routing

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

# LeMARLA **Large Language Model-Enhanced Multi-Agent Reinforcement Learning for Autonomous Enterprise Workflow Orchestration via Graph Semantic Routing.** LeMARLA couples an LLM-driven task decomposition module, a heterogeneous graph attention routing module over tasks, employees and skills, and a CTDE multi-agent reinforcement learning policy that consumes the routing scores as a differentiable semantic prior. The failure-driven decomposition loop, the differentiable semantic prior and the heterogeneous graph over organisational entities are optimised end to end. --- ## Repository layout ```lemarla/├── data/ Enterprise workflow dataset (see Dataset below)├── src/ Model implementation├── scripts/ Training, evaluation, plotting entry points├── configs/ Default YAML configuration├── tests/ Module-level smoke tests├── outputs/ Tables, figures, logs and checkpoints├── requirements.txt└── README.md``` ## Stack * Python 3.10+* PyTorch 2.1, PyTorch Geometric 2.5.0, Transformers 4.38, vLLM 0.4.2* LLaMA-2-13B-chat backbone in bfloat16 (weights frozen)* NVIDIA A100 80 GB GPU* OR-Tools CP-SAT 9.10 for the constraint-programming baseline ## Quickstart ```bash# 1. Install dependenciespip install -r requirements.txt # 2. Place the dataset (eleven xlsx files) under data/lemarla_dataset/# See "Dataset" section below. # 3. Run the pipeline end to endpython scripts/run_all_experiments.py # Or run each stage separatelypython scripts/prepare_data.pypython scripts/train.pypython scripts/evaluate.pypython scripts/plot_all_figures.py # 4. Verify installationpython tests/test_smoke.py``` If the LLaMA-2 checkpoint is unavailable, the decomposition module resolves the hidden state through a sentence-transformer backbone padded to the target hidden dimension so the pipeline runs unchanged. --- ## Dataset Operational logs of a manufacturing enterprise between **2024-04-01 and 2025-03-31** covering task decomposition, employee routing and workflow execution across six functional departments. All identifiers were removed by the data owner before transfer; employee-level attributes are restricted to department, job category, skill tags and workload. ### Coverage | Item | Count || ------------------------------------ | ---------- || Observation window | 12 months || Departments | 6 || Job categories | 8 || Certified skill categories | 48 || Subtask categories | 8 || Employees | 326 || Employee-skill certifications | 2,847 || Collaborates edges | 2,914 || Raw workflow instances | 21,435 || Retained workflow instances | 8,742 || Retained subtasks | 41,537 || Dependency edges | 29,086 || Annotated reference instances | 300 || Annotated reference subtasks | 1,521 || Annotated reference dependency edges | 1,063 | ### Files The dataset ships as eleven xlsx files under `data/lemarla_dataset/`: | File | Rows | Content || ---- | ----:| ------- || `01_departments.xlsx` | 6 | Department master || `02_job_categories.xlsx` | 8 | Job category master || `03_skill_taxonomy.xlsx` | 48 | Certified skill categories || `04_subtask_category_taxonomy.xlsx` | 8 | Subtask category master || `05_employee_master_table.xlsx` | 326 | Employees || `06_employee_skill_certifications.xlsx` | 2,847 | Employee-skill Possesses edges || `07_collaborates_edges.xlsx` | 2,914 | Employee-employee Collaborates edges || `08_workflow_instances_raw.xlsx` | 21,435 | Raw workflow instances with retention flag || `09_subtask_execution_log.xlsx` | 41,537 | Subtask execution log || `10_dependency_edges.xlsx` | 29,086 | Subtask precedence edges || `11_annotated_reference_set.xlsx` | 300 / 1,521 / 1,063 | Annotated reference (three sheets) | Full column dictionary and preprocessing cascade are in `data/lemarla_dataset/README.md`. ### Chronological split The 8,742 retained instances are partitioned in strict chronological order of `issue_ts` into training (70%), validation (10%) and test (20%). --- ## Modules | File | Purpose || ---- | ------- || `src/data.py` | Dataset loader, heterogeneous graph indices, chronological splitter || `src/decomposition.py` | LLaMA backbone wrapper, two-layer task projection, dependency scorer, failure-driven re-decomposition || `src/het_gat.py` | Type-specific projection, type-aware node-level attention, semantic-level fusion, routing matching score || `src/mappo.py` | Actor, centralised critic, PPO update with clipped surrogate loss || `src/simulator.py` | Gradient-boosted duration and failure models, workload state update || `src/metrics.py` | Task completion rate, end-to-end latency, load balance index, routing accuracy at rank k, decision latency || `src/baselines.py` | Rule-Based, Vanilla-MARL, LLM-Only, GNN-MARL, TDAG-Assign, TDAG-Assign-LB, CP-RCPSP, SOP-Orchestrator || `src/lemarla.py` | Full model wiring || `src/utils.py` | Configuration, logging, seeding, Holm-Bonferroni correction, Cohen's d | ## Scripts | Script | Role || ------ | ---- || `scripts/prepare_data.py` | Load the dataset and dump the preprocessing cascade || `scripts/train.py` | Fit the simulator, train the dependency scorer, run MAPPO || `scripts/evaluate.py` | Produce the aggregate outcome, ablation, scalability and backbone tables || `scripts/plot_all_figures.py` | Produce the fourteen figures || `scripts/run_all_experiments.py` | End-to-end orchestrator | ## Outputs After a successful run, `outputs/` contains: * `tables/Table_5_main_comparison.xlsx` — aggregate outcomes for LeMARLA and eight baselines over ten seeds, with 95% confidence intervals, Holm-corrected p values, and Cohen's d.* `tables/Table_6_decomposition_quality.xlsx` — boundary and dependency F1 against the annotated reference set.* `tables/Table_7_ablation.xlsx` — ablation results with paired t tests.* `tables/Table_8_scalability.xlsx` — employee-scale sweep from 50 to 2,000.* `tables/Table_9_backbones.xlsx` — LLaMA-2-13B-chat vs LLaMA-3.1-8B-Instruct vs Qwen2.5-14B-Instruct.* `tables/Table_10_relative_contribution.xlsx` — relative degradations upon module removal.* `tables/table3_cascade.json` — preprocessing cascade.* `figures/Figure_1_overall_framework.pdf` … `figures/Figure_14_radar_comparison.pdf` ## Method LeMARLA models autonomous workflow orchestration as a Dec-POMDP over virtualised employee roles. The three modules cooperate through the following signal path: 1. **Decomposition.** An instruction is parsed by LLaMA-2-13B-chat into a subtask directed acyclic graph. The dependency scoring function is trained under a binary cross-entropy loss on positive and negative subtask pairs recovered from execution timestamps and work-order routing tables. A schema-validation layer regenerates on failure and falls back to a rule template after three attempts.2. **Routing.** Type-specific projections map task, employee and skill features into a shared 128-dimensional space. Node-level attention aggregates neighbourhood information per edge type, and a semantic-level attention fuses the three edge types. The cosine similarity between task and employee embeddings, discounted by a load penalty, forms the routing matching score.3. **Policy.** MAPPO trains a centralised critic and decentralised actors. Actor logits are the sum of a local raw logit and the routing matching score, so the semantic prior enters the policy as a differentiable term. The joint reward combines execution success, the deviation between actual and estimated processing time, and the workload variance across agents. Failure signals raised by the execution engine re-decompose the local subgraph, keeping completed subtasks unchanged. ## Statistical protocol All results summarise ten independent seeds. Point estimates are means; dispersion is either sample standard deviation or a 95% confidence interval under a t distribution with nine degrees of freedom (t half-width 2.262). Pairwise comparisons use the Welch t test against the eight baselines and the paired t test against the ablation variants, controlled by the Holm-Bonferroni procedure at a family-wise level of 0.05. Effect sizes are reported as Cohen's d. ## Licence MIT.

View source

Similar papers

#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#machine learning Review Open access Jun 2014

Why Early-Stage Software Startups Fail: A Behavioral Framework

This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.

Carmine Giardino, Xiaofeng Wang, P. Abrahamsson · 175 citations · ⚡19
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#machine learning Review Open access May 2016

Key Challenges in Software Startups Across Life Cycle Stages

It is found that what perceived as biggest challenges by software startups do vary across different life cycle stages, even though its significance decreases when the learning focuses of the startups move from problem to solution and their products mature.

Xiaofeng Wang, Henry Edison, Sohaib Shahid Bajwa et al. · 62 citations · ⚡6

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.