PapersWithELO
← ICLR 2024 leaderboard

P4Q: Learning to Prompt for Quantization in Visual-language Models

Huixin Sun, Runqi Wang, Xianbin Cao, Yanjing Li, Xiaolong Jiang, Yao Hu, Baochang Zhang

self/semi-supervised learningQuantizationVision-Language Models (VLMs)
48.80100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
43.40100
Mimo
band ≈ ±20 pct pts (from σ = 0.39)
59.80100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.38)

OpenReview ground truth

Rejected

Abstract

Large-scale pre-trained Vision-Language Models (VLMs) have gained prominence in various visual and multimodal tasks, yet the deployment of VLMs on resource-constrained platforms remains challenging due to their prohibitive computational and memory overhead. Quantization of VLMs can substantially reduce the computational and memory costs, which are in urgent need. There are two prevailing paradigms, Quantization-Aware Training (QAT) can effectively quantize large-scale VLMs but incur a huge training cost, while low-bit Post-Training Quantization (PTQ) suffers from a notable performance drop. We propose a `Prompt for Quantization'' (P4Q) method, in which we design a lightweight architecture to leverage contrastive loss supervision to enhance the recognition performance of a PTQ model. Our method can effectively reduce the gap between image features and text features caused by low-bit quantization, based on learnable prompts to reorganize textual representations and a low-bit adapter to realign the distributions of image and text features. We also introduce a distillation loss based on cosine similarity predictions to distill the quantized model using a full-precision teacher. Extensive experimental results demonstrate that our P4Q method outperforms prior arts, even achieving comparable results to its full-precision counterparts. For instance, our 8-bit P4Q can theoretically compress the CLIP-ViT/B-32 by 4 $\times$ while achieving 79.42\% Top-1 accuracy, outperforming the learnable prompt fine-tuned full-precision model by 2.91\% with negligible additional parameters on the CIFAR100 dataset. Test code and checkpoints are available at \url{https://anonymous.4open.science/r/ICLR2024-P4Q-1255}

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)