Hierarchical Reinforcement Learning–Based Optimal Control for Model-Free Linear Systems
Yong Zhang, Xiangrui Yan, Weiqing Yang, Yuyang Zhou
Source abstract
A novel model-free hierarchical reinforcement learning (HRL)–based Linear Quadratic Regulator (LQR) control framework with adaptive weight selection is proposed to address the reliance of conventional LQR methods on accurate system models and manual parameter tuning. The proposed approach adopts a two-level learning architecture in which a high-level meta-agent adaptively optimizes the LQR weighting matrices Q and R through entropy-based trajectory evaluation, while a low-level base-agent performs model-free policy iteration to update the state-feedback control law under unknown system dynamics. By decoupling weight optimization from control-law learning, the framework enables simultaneous adaptation of the cost-function parameters and the feedback gain without requiring explicit model information. To enhance learning stability and exploration during weight adaptation, Gaussian noise and an experience replay mechanism are incorporated into the learning process. Numerical simulations on second- and third-order linear systems demonstrate that the proposed HRL-based LQR method achieves effective control performance, reliable convergence, and improved adaptability in model-free environments.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.