HiLoRL: A Hierarchical Logical Model for Learning Composite Tasks
Chuan Hu, Jingyu Cao, Jinpeng Zhang, Yunze Wu, Yi Wu, Zhilei Xu, Jianzhu Ma, Yuan Zhou
OpenReview ground truth
TL;DR — We design a hierarchical reinforcement learning model to deal with composite tasks, meanwhile providing interpretability and selective domain knowledge instruction mechanism
Abstract
We propose HiLoRL, a hierarchical model to learn policies for composite tasks. Recent studies mostly focus on using human-specified logical specifications, which is laborious and produces models that perform poorly when facing tasks not entirely human-predictable. HiLoRL is composed of a high-level logical planner and low-level action policies. It initially learns a rough rule at its upper level with the help of low-level policies and then uses joint training with surrogate rewards to refine the rough rule and low-level policies. Furthermore, HiLoRL can incorporate specialized predicates derived from expert knowledge, thereby enhancing its training speed and performance. We also design a synthesis algorithm to illustrate our high-level planner's logical structure as an automaton, demonstrating our model's interpretability. HiLoRL outperforms state-of-the-art baselines in several benchmarks with continuous state and action spaces. Additionally, HiLoRL does not require human to hard-code logical structures, so it can solve logically uncertain tasks.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 34 comparisons
Ranked above opponent in 41% of matchups.
- ▲ beat Keqing: Knowledge-based Question Answering… ×8
- ▼ lost to Semi-supervised batch learning from logged… ×6
- ▼ lost to The Curse of Diversity in Ensemble-Based E… ×4
- ▼ lost to Boolformer: Symbolic Regression of Logic F… ×4
- ▼ lost to Physics-Regulated Deep Reinforcement Learn… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 34)