PapersWithELO
← ICLR 2024 leaderboard

HeaP: Hierarchical Policies for Web Actions using LLMs

Paloma Sodhi, S.R.K Branavan, Ryan McDonald

robotics & planningweb actionslarge language modelstask decompositionfew-shot demonstrations
66.30100
Fused
band ≈ ±16 pct pts (from σ = 0.32)
75.60100
Mimo
band ≈ ±23 pct pts (from σ = 0.46)
56.60100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.45)

OpenReview ground truth

Rejected

TL;DR — LLMs that learn to solve complex web tasks by decomposing them into low-level policy calls, achieving superior performance with significantly less data

Abstract

Large language models (LLMs) have demonstrated remarkable capabilities in performing a range of instruction-following tasks in few and zero-shot settings. However, teaching LLMs to perform tasks on the web presents fundamental challenges -- combinatorially large open-world tasks and variations across web interfaces. We tackle these challenges by leveraging LLMs to decompose web tasks into a collection of sub-tasks, each of which can be solved by a low-level, closed-loop policy. These policies constitute a shared grammar across tasks, i.e., new web tasks can be expressed as a composition of these policies. We propose a novel framework, Hierarchical Policies for Web Actions using LLMs (HeaP), that learns a set of hierarchical LLM prompts from demonstrations for planning high-level tasks and executing low-level policies. We evaluate HeaP against a range of baselines on a suite of web tasks, including MiniWoB++, WebArena, a mock airline CRM, as well as live website interactions, and show that it is able to outperform prior works using orders of magnitude less data.

Author context

Most prolific author: 1 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 28)