PapersWithELO
← ICLR 2024 leaderboard

HiLoRL: A Hierarchical Logical Model for Learning Composite Tasks

Chuan Hu, Jingyu Cao, Jinpeng Zhang, Yunze Wu, Yi Wu, Zhilei Xu, Jianzhu Ma, Yuan Zhou

neurosymbolic AIHierarchical Reinforcement LearningAdaptive Logic PlannerInterpretabilityExpert Knowledge Instruction
15.40100
Fused
band ≈ ±15 pct pts (from σ = 0.29)
16.30100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
15.20100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.42)

OpenReview ground truth

Rejected

TL;DR — We design a hierarchical reinforcement learning model to deal with composite tasks, meanwhile providing interpretability and selective domain knowledge instruction mechanism

Abstract

We propose HiLoRL, a hierarchical model to learn policies for composite tasks. Recent studies mostly focus on using human-specified logical specifications, which is laborious and produces models that perform poorly when facing tasks not entirely human-predictable. HiLoRL is composed of a high-level logical planner and low-level action policies. It initially learns a rough rule at its upper level with the help of low-level policies and then uses joint training with surrogate rewards to refine the rough rule and low-level policies. Furthermore, HiLoRL can incorporate specialized predicates derived from expert knowledge, thereby enhancing its training speed and performance. We also design a synthesis algorithm to illustrate our high-level planner's logical structure as an automaton, demonstrating our model's interpretability. HiLoRL outperforms state-of-the-art baselines in several benchmarks with continuous state and action spaces. Additionally, HiLoRL does not require human to hard-code logical structures, so it can solve logically uncertain tasks.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 41% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)