Musketeer: Joint Training/Inference for Multi-task Vision-Language Model with Task Explanation Prompts
Zhaoyang Zhang, Yantao Shen, Kunyu Shi, Zhaowei Cai, Jun Fang, Siqi Deng, Hao Yang, Davide Modolo, Zhuowen Tu, Stefano Soatto
OpenReview ground truth
Abstract
We present a sequence-to-sequence vision-language model whose parameters are jointly trained on all tasks and fully shared among multiple tasks, resulting in a single model which we named Musketeer. The integration of knowledge across heterogeneous tasks is enabled by a novel feature called Task Explanation Prompt (TEP). TEP reduces interference among tasks, allowing the model to focus on their shared structure. With a single model, Musketeer achieves results comparable to or better than strong baselines trained on single tasks, almost uniformly across multiple tasks.
Author context
Most prolific author: 4 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 46 comparisons
Ranked above opponent in 45% of matchups.
- ▼ lost to Improving LoRA in Privacy-preserving Feder… ×6
- ▼ lost to Detect Every Thing with Few Examples ×4
- ▼ lost to V-DETR: DETR with Vertex Relative Position… ×4
- ▼ lost to TransCues: Boundary and Reflection-empower… ×4
- ▼ lost to AlignDiff: Aligning Diffusion Models for G… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 46)