Instance Needs More Care: Rewriting Prompts for Instances Yields Better Zero-Shot Performance
Saurabh Srivastava, Chengyue Huang, Weiguo Fan, Ziyu Yao
OpenReview ground truth
TL;DR — We propose PRoMTd, an approach that rewrites the task prompt for each individual test input to be more specific, unambiguous, and complete, so as to provide better guidance to the task LLM for zero-shot tasks.
Abstract
Enabling large language models (LLMs) to perform tasks in zero-shot has been an appealing goal owing to its labor-saving (i.e., requiring no task-specific annotations); as such, zero-shot prompting approaches also enjoy better task generalizability. To improve LLMs' zero-shot performance, prior work has focused on devising more effective task instructions (e.g., ``let's think step by step'' \cite{kojima2022large}). However, we argue that, in order for an LLM to solve them correctly in zero-shot, individual test instances need more carefully designed and customized instructions. To this end, we propose PRoMTd, an approach that rewrites the task prompt for each individual test input to be more specific, unambiguous, and complete, so as to provide better guidance to the task LLM. We evaluated PRoMTd on eight datasets covering tasks including arithmetics, logical reasoning, and code generation, using GPT-4 as the task LLM. PRoMTd consistently outperforms traditional zero-shot approaches on all the datasets. Notably, we observe an absolute improvement of 10\% on the complex MATHS dataset and 5\% on the code generation task on HumanEval. In addition, we also showed that the rewritten prompt can provide better interpretability of how the LLM resolves each test instance, which can potentially be leveraged as a defense mechanism against adversarial prompting.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 42 comparisons
Ranked above opponent in 48% of matchups.
- ▲ beat FLASK: Fine-grained Language Model Evaluat… ×6
- ▼ lost to SWE-bench: Can Language Models Resolve Rea… ×4
- ▲ beat Long-Term Typhoon Trajectory Prediction: A… ×4
- ▼ lost to Beyond Size: How Gradients Shape Pruning D… ×4
- ▼ lost to Unified Language-Vision Pretraining in LLM… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 42)