FireAct: Toward Language Agent Finetuning
Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik R Narasimhan, Shunyu Yao
OpenReview ground truth
TL;DR — Fine-tuning language models for agents is understudied, so we present a systematic study.
Abstract
Recent efforts have augmented language models (LMs) with external tools or environments, leading to the development of language agents that can reason and act. However, most of these agents rely on few-shot prompting techniques, which can result in a lack of robustness in agent performance due to the limited learning support. In this paper, we investigate the less explored direction of fine-tuning LMs to obtain language agents. With a simple, controlled setup that uses a Google search API for question answering (QA), we systematically explore a variety of base LMs, agent methods, fine-tuning data, and QA tasks. Our experiments reveal novel insights around the scaling effects of the base LM and fine-tuning data, combining trajectory data collected from different tasks and different agent methods, as well as robustness to different types of data perturbations. Overall, these findings illustrate overlooked advantages of fine-tuned language agents over existing prompting-based ones, provide empirical guidelines for fine-tuning, and indicate future directions in creating better tasks and methods for language agents.
Author context
Most prolific author: 6 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 49% of matchups.
- ▲ beat Pick-or-Mix: Dynamic Channel Sampling for … ×6
- ▼ lost to Tool-Augmented Reward Modeling ×4
- ▼ lost to Listen to Motion: Robustly Learning Correl… ×4
- ▼ lost to Unified Language-Vision Pretraining in LLM… ×4
- ▼ lost to LETI: Learning to Generate from Textual In… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)