P4Q: Learning to Prompt for Quantization in Visual-language Models
Huixin Sun, Runqi Wang, Xianbin Cao, Yanjing Li, Xiaolong Jiang, Yao Hu, Baochang Zhang
OpenReview ground truth
Abstract
Large-scale pre-trained Vision-Language Models (VLMs) have gained prominence in various visual and multimodal tasks, yet the deployment of VLMs on resource-constrained platforms remains challenging due to their prohibitive computational and memory overhead. Quantization of VLMs can substantially reduce the computational and memory costs, which are in urgent need. There are two prevailing paradigms, Quantization-Aware Training (QAT) can effectively quantize large-scale VLMs but incur a huge training cost, while low-bit Post-Training Quantization (PTQ) suffers from a notable performance drop. We propose a `Prompt for Quantization'' (P4Q) method, in which we design a lightweight architecture to leverage contrastive loss supervision to enhance the recognition performance of a PTQ model. Our method can effectively reduce the gap between image features and text features caused by low-bit quantization, based on learnable prompts to reorganize textual representations and a low-bit adapter to realign the distributions of image and text features. We also introduce a distillation loss based on cosine similarity predictions to distill the quantized model using a full-precision teacher. Extensive experimental results demonstrate that our P4Q method outperforms prior arts, even achieving comparable results to its full-precision counterparts. For instance, our 8-bit P4Q can theoretically compress the CLIP-ViT/B-32 by 4 $\times$ while achieving 79.42\% Top-1 accuracy, outperforming the learnable prompt fine-tuned full-precision model by 2.91\% with negligible additional parameters on the CIFAR100 dataset. Test code and checkpoints are available at \url{https://anonymous.4open.science/r/ICLR2024-P4Q-1255}
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 50% of matchups.
- ▼ lost to Differential Model Scaling using Different… ×6
- ▲ beat Out of the Ordinary: Spectrally Adapting R… ×6
- ▼ lost to A Recipe for Improved Certifiable Robustne… ×4
- ▲ beat Gaussian Process-Based Corruption-resilien… ×4
- ▲ beat Black-box Targeted Adversarial Attack on S… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)