PapersWithELO
← ICLR 2024 leaderboard

FireAct: Toward Language Agent Finetuning

Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik R Narasimhan, Shunyu Yao

representation learninglanguage agentlanguage modellarge language modelfinetuningagenttool use
42.50100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
51.20100
Mimo
band ≈ ±21 pct pts (from σ = 0.43)
30.00100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.43)

OpenReview ground truth

Rejected

TL;DR — Fine-tuning language models for agents is understudied, so we present a systematic study.

Abstract

Recent efforts have augmented language models (LMs) with external tools or environments, leading to the development of language agents that can reason and act. However, most of these agents rely on few-shot prompting techniques, which can result in a lack of robustness in agent performance due to the limited learning support. In this paper, we investigate the less explored direction of fine-tuning LMs to obtain language agents. With a simple, controlled setup that uses a Google search API for question answering (QA), we systematically explore a variety of base LMs, agent methods, fine-tuning data, and QA tasks. Our experiments reveal novel insights around the scaling effects of the base LM and fine-tuning data, combining trajectory data collected from different tasks and different agent methods, as well as robustness to different types of data perturbations. Overall, these findings illustrate overlooked advantages of fine-tuned language agents over existing prompting-based ones, provide empirical guidelines for fine-tuning, and indicate future directions in creating better tasks and methods for language agents.

Author context

Most prolific author: 6 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 32 comparisons

Ranked above opponent in 49% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)