CellPLM: Pre-training of Cell Language Model Beyond Single Cells
Hongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding, Wei Jin, Yuying Xie, Jiliang Tang
OpenReview ground truth
Abstract
The current state-of-the-art single-cell pre-trained models are greatly inspired by the success of large language models. They trained transformers by treating genes as tokens and cells as sentences. However, three fundamental differences between single-cell data and natural language data are overlooked: (1) scRNA-seq data are presented as bag-of-genes instead of sequences of RNAs; (2) Cell-cell relations are more intricate and important than inter-sentence relations; and (3) The quantity of single-cell data is considerably inferior to text data, and they are very noisy. In light of these characteristics, we propose a new pre-trained model, $\textit{CellPLM}$, which takes cells as tokens and tissues as sentences. In addition, we leverage spatially-resolved transcriptomic data in pre-training to facilitate learning cell-cell relationships and introduce a Gaussian prior distribution as an additional inductive bias to overcome data limitations. $\textit{CellPLM}$ is the first single-cell pre-trained transformer that encodes cell-cell relations and it consistently outperforms existing pre-trained and non-pre-trained models in diverse downstream tasks, with 100 times higher inference speed on generating cell embeddings than previous pre-trained models.
Author context
Most prolific author: 13 submissions (credibility 0.64).
Delta if applied: -0.1 percentile
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 38 comparisons
Ranked above opponent in 48% of matchups.
- ▲ beat A General Single-Cell Analysis Framework v… ×6
- ▼ lost to KW-Design: Pushing the Limit of Protein De… ×6
- ▲ beat OPTIMIZING STABILIZATION IN SINGULARLY PER… ×4
- ▼ lost to De novo Protein Design Using Geometric Vec… ×4
- ▲ beat NL2ProGPT: Taming Large Language Model for… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 38)