PapersWithELO
← ICLR 2024 leaderboard

Musketeer: Joint Training/Inference for Multi-task Vision-Language Model with Task Explanation Prompts

Zhaoyang Zhang, Yantao Shen, Kunyu Shi, Zhaowei Cai, Jun Fang, Siqi Deng, Hao Yang, Davide Modolo, Zhuowen Tu, Stefano Soatto

representation learningMulti-Task LearningMulti-modality
25.50100
Fused
band ≈ ±13 pct pts (from σ = 0.26)
29.60100
Mimo
band ≈ ±18 pct pts (from σ = 0.36)
18.50100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.39)

OpenReview ground truth

Rejected

Abstract

We present a sequence-to-sequence vision-language model whose parameters are jointly trained on all tasks and fully shared among multiple tasks, resulting in a single model which we named Musketeer. The integration of knowledge across heterogeneous tasks is enabled by a novel feature called Task Explanation Prompt (TEP). TEP reduces interference among tasks, allowing the model to focus on their shared structure. With a single model, Musketeer achieves results comparable to or better than strong baselines trained on single tasks, almost uniformly across multiple tasks.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 46 comparisons

Ranked above opponent in 45% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 46)