Home
Scholarly Works
CSPPO: A multi-agent reinforcement learning method...
Journal article

CSPPO: A multi-agent reinforcement learning method for production capacity allocation in manufacturing industry chains

Abstract

Production capacity allocation in multi-tier manufacturing industry chains requires coordinated decisions under heterogeneous node capabilities, private local information, and dynamic demand. Centralized optimization scales poorly in such networks, while flat multi-agent reinforcement learning (MARL) methods lack explicit leader–follower coordination and are affected by non-stationary policy updates. This paper proposes chain-leader-guided Stackelberg proximal policy optimization (CSPPO), a hierarchical MARL method built on independent PPO. The problem is formulated as a hierarchical decentralized partially observable Markov decision process. The chain leader broadcasts privacy-preserving global guidance, followers execute decentralized capacity allocation policies, and Stackelberg sequential update mechanism (SSUM) separates follower and leader updates to reduce non-stationarity. Experiments include nine simulated industrial chain instances and an automotive glass case. The method is evaluated against ablation variants and other MARL methods. CSPPO stabilizes rapidly and achieves the highest order fulfillment rate (OFR) among the MARL baselines, with competitive Cost and production capacity load balance (PCLB) trade-offs. In the automotive glass case, it improves OFR by 12.0%, reduces Cost by 4.6%, and reduces average changeover frequency (ACF) by 91.6%. These results show that CSPPO can coordinate production capacity allocation in complex industrial chains.

Authors

Liang P; Sun X; She M; Gao Y; Liu M; Shen W

Journal

Journal of Manufacturing Systems, Vol. 89, , pp. 124–138

Publisher

Elsevier

Publication Date

December 1, 2026

DOI

10.1016/j.jmsy.2026.08.016

ISSN

0278-6125

View published work (Non-McMaster Users)