HeaP: Hierarchical Policies for Web Actions using LLMs
Paloma Sodhi, S.R.K Branavan, Ryan McDonald
OpenReview ground truth
TL;DR — LLMs that learn to solve complex web tasks by decomposing them into low-level policy calls, achieving superior performance with significantly less data
Abstract
Large language models (LLMs) have demonstrated remarkable capabilities in performing a range of instruction-following tasks in few and zero-shot settings. However, teaching LLMs to perform tasks on the web presents fundamental challenges -- combinatorially large open-world tasks and variations across web interfaces. We tackle these challenges by leveraging LLMs to decompose web tasks into a collection of sub-tasks, each of which can be solved by a low-level, closed-loop policy. These policies constitute a shared grammar across tasks, i.e., new web tasks can be expressed as a composition of these policies. We propose a novel framework, Hierarchical Policies for Web Actions using LLMs (HeaP), that learns a set of hierarchical LLM prompts from demonstrations for planning high-level tasks and executing low-level policies. We evaluate HeaP against a range of baselines on a suite of web tasks, including MiniWoB++, WebArena, a mock airline CRM, as well as live website interactions, and show that it is able to outperform prior works using orders of magnitude less data.
Author context
Most prolific author: 1 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 28 comparisons
Ranked above opponent in 52% of matchups.
- ▲ beat MindAgent: Emergent Gaming Interaction ×4
- ▼ lost to Zero-Shot Robotic Manipulation with Pre-Tr… ×4
- ▲ beat Pick-or-Mix: Dynamic Channel Sampling for … ×4
- ▲ beat V-DETR: DETR with Vertex Relative Position… ×4
- ▲ beat THOUGHT PROPAGATION: AN ANALOGICAL APPROAC… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 28)