Large Language Model-Enhanced Multi-Agent Reinforcement Learning for Autonomous Enterprise Workflow Orchestration via Graph Semantic Routing
Abstract
# LeMARLA **Large Language Model-Enhanced Multi-Agent Reinforcement Learning for Autonomous Enterprise Workflow Orchestration via Graph Semantic Routing.** LeMARLA couples an LLM-driven task decomposition module, a heterogeneous graph attention routing module over tasks, employees and skills, and a CTDE multi-agent reinforcement learning policy that consumes the routing scores as a differentiable semantic prior. The failure-driven decomposition loop, the differentiable semantic prior and the heterogeneous graph over organisational entities are optimised end to end. --- ## Repository layout ```lemarla/├── data/ Enterprise workflow dataset (see Dataset below)├── src/ Model implementation├── scripts/ Training, evaluation, plotting entry points├── configs/ Default YAML configuration├── tests/ Module-level smoke tests├── outputs/ Tables, figures, logs and checkpoints├── requirements.txt└── README.md``` ## Stack * Python 3.10+* PyTorch 2.1, PyTorch Geometric 2.5.0, Transformers 4.38, vLLM 0.4.2* LLaMA-2-13B-chat backbone in bfloat16 (weights frozen)* NVIDIA A100 80 GB GPU* OR-Tools CP-SAT 9.10 for the constraint-programming baseline ## Quickstart ```bash# 1. Install dependenciespip install -r requirements.txt # 2. Place the dataset (eleven xlsx files) under data/lemarla_dataset/# See "Dataset" section below. # 3. Run the pipeline end to endpython scripts/run_all_experiments.py # Or run each stage separatelypython scripts/prepare_data.pypython scripts/train.pypython scripts/evaluate.pypython scripts/plot_all_figures.py # 4. Verify installationpython tests/test_smoke.py``` If the LLaMA-2 checkpoint is unavailable, the decomposition module resolves the hidden state through a sentence-transformer backbone padded to the target hidden dimension so the pipeline runs unchanged. --- ## Dataset Operational logs of a manufacturing enterprise between **2024-04-01 and 2025-03-31** covering task decomposition, employee routing and workflow execution across six functional departments. All identifiers were removed by the data owner before transfer; employee-level attributes are restricted to department, job category, skill tags and workload. ### Coverage | Item | Count || ------------------------------------ | ---------- || Observation window | 12 months || Departments | 6 || Job categories | 8 || Certified skill categories | 48 || Subtask categories | 8 || Employees | 326 || Employee-skill certifications | 2,847 || Collaborates edges | 2,914 || Raw workflow instances | 21,435 || Retained workflow instances | 8,742 || Retained subtasks | 41,537 || Dependency edges | 29,086 || Annotated reference instances | 300 || Annotated reference subtasks | 1,521 || Annotated reference dependency edges | 1,063 | ### Files The dataset ships as eleven xlsx files under `data/lemarla_dataset/`: | File | Rows | Content || ---- | ----:| ------- || `01_departments.xlsx` | 6 | Department master || `02_job_categories.xlsx` | 8 | Job category master || `03_skill_taxonomy.xlsx` | 48 | Certified skill categories || `04_subtask_category_taxonomy.xlsx` | 8 | Subtask category master || `05_employee_master_table.xlsx` | 326 | Employees || `06_employee_skill_certifications.xlsx` | 2,847 | Employee-skill Possesses edges || `07_collaborates_edges.xlsx` | 2,914 | Employee-employee Collaborates edges || `08_workflow_instances_raw.xlsx` | 21,435 | Raw workflow instances with retention flag || `09_subtask_execution_log.xlsx` | 41,537 | Subtask execution log || `10_dependency_edges.xlsx` | 29,086 | Subtask precedence edges || `11_annotated_reference_set.xlsx` | 300 / 1,521 / 1,063 | Annotated reference (three sheets) | Full column dictionary and preprocessing cascade are in `data/lemarla_dataset/README.md`. ### Chronological split The 8,742 retained instances are partitioned in strict chronological order of `issue_ts` into training (70%), validation (10%) and test (20%). --- ## Modules | File | Purpose || ---- | ------- || `src/data.py` | Dataset loader, heterogeneous graph indices, chronological splitter || `src/decomposition.py` | LLaMA backbone wrapper, two-layer task projection, dependency scorer, failure-driven re-decomposition || `src/het_gat.py` | Type-specific projection, type-aware node-level attention, semantic-level fusion, routing matching score || `src/mappo.py` | Actor, centralised critic, PPO update with clipped surrogate loss || `src/simulator.py` | Gradient-boosted duration and failure models, workload state update || `src/metrics.py` | Task completion rate, end-to-end latency, load balance index, routing accuracy at rank k, decision latency || `src/baselines.py` | Rule-Based, Vanilla-MARL, LLM-Only, GNN-MARL, TDAG-Assign, TDAG-Assign-LB, CP-RCPSP, SOP-Orchestrator || `src/lemarla.py` | Full model wiring || `src/utils.py` | Configuration, logging, seeding, Holm-Bonferroni correction, Cohen's d | ## Scripts | Script | Role || ------ | ---- || `scripts/prepare_data.py` | Load the dataset and dump the preprocessing cascade || `scripts/train.py` | Fit the simulator, train the dependency scorer, run MAPPO || `scripts/evaluate.py` | Produce the aggregate outcome, ablation, scalability and backbone tables || `scripts/plot_all_figures.py` | Produce the fourteen figures || `scripts/run_all_experiments.py` | End-to-end orchestrator | ## Outputs After a successful run, `outputs/` contains: * `tables/Table_5_main_comparison.xlsx` — aggregate outcomes for LeMARLA and eight baselines over ten seeds, with 95% confidence intervals, Holm-corrected p values, and Cohen's d.* `tables/Table_6_decomposition_quality.xlsx` — boundary and dependency F1 against the annotated reference set.* `tables/Table_7_ablation.xlsx` — ablation results with paired t tests.* `tables/Table_8_scalability.xlsx` — employee-scale sweep from 50 to 2,000.* `tables/Table_9_backbones.xlsx` — LLaMA-2-13B-chat vs LLaMA-3.1-8B-Instruct vs Qwen2.5-14B-Instruct.* `tables/Table_10_relative_contribution.xlsx` — relative degradations upon module removal.* `tables/table3_cascade.json` — preprocessing cascade.* `figures/Figure_1_overall_framework.pdf` … `figures/Figure_14_radar_comparison.pdf` ## Method LeMARLA models autonomous workflow orchestration as a Dec-POMDP over virtualised employee roles. The three modules cooperate through the following signal path: 1. **Decomposition.** An instruction is parsed by LLaMA-2-13B-chat into a subtask directed acyclic graph. The dependency scoring function is trained under a binary cross-entropy loss on positive and negative subtask pairs recovered from execution timestamps and work-order routing tables. A schema-validation layer regenerates on failure and falls back to a rule template after three attempts.2. **Routing.** Type-specific projections map task, employee and skill features into a shared 128-dimensional space. Node-level attention aggregates neighbourhood information per edge type, and a semantic-level attention fuses the three edge types. The cosine similarity between task and employee embeddings, discounted by a load penalty, forms the routing matching score.3. **Policy.** MAPPO trains a centralised critic and decentralised actors. Actor logits are the sum of a local raw logit and the routing matching score, so the semantic prior enters the policy as a differentiable term. The joint reward combines execution success, the deviation between actual and estimated processing time, and the workload variance across agents. Failure signals raised by the execution engine re-decompose the local subgraph, keeping completed subtasks unchanged. ## Statistical protocol All results summarise ten independent seeds. Point estimates are means; dispersion is either sample standard deviation or a 95% confidence interval under a t distribution with nine degrees of freedom (t half-width 2.262). Pairwise comparisons use the Welch t test against the eight baselines and the paired t test against the ablation variants, controlled by the Holm-Bonferroni procedure at a family-wise level of 0.05. Effect sizes are reported as Cohen's d. ## Licence MIT.