NL2ProGPT: Taming Large Language Model for Conversational Protein Design
Zekun Guo, Jianxin Lin, hnuzsf@hnu.edu.cn, Yijun Wang, Lijun Wu, xiangxiang Zeng
OpenReview ground truth
Abstract
Large Language Models (LLMs), like ChatGPT, excel in cross-modal tasks thanks to their powerful abilities in natural language comprehension, generalization, and reasoning. Meanwhile, the wealth of human-curated protein knowledge in text form presents a unique opportunity for LLMs to contribute to advanced protein design. In this work, we propose a new LLMs-based framework, namely NL2ProGPT, for macromolecular protein sequence generation that bridges the domain gap between natural and protein languages. Specifically, we first combine the protein functions and properties to create specific text guidelines for designing the protein, ensuring it follows precise controls. Second, to form a more informative and generalizable protein description, we explicitly inject protein structural information by clustering the embeddings from pre-trained protein language models. Third, we train a reward model to align the protein language model with the Rosetta energy function, following an RLAIF (reinforced learning from AI feedback) fashion. We empirically verify the effectiveness of NL2ProGPT from three aspects: (1) outperforms existing protein sequence design methods in different evaluations; (2) exhibits more than 90\% consistency in text-to-protein generation; (3) has effective exploration potential in disordered regions.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 42 comparisons
Ranked above opponent in 44% of matchups.
- ▲ beat Dispatching Ambulances using Deep Reinforc… ×8
- ▲ beat Learning Team-Level Information Integratio… ×8
- ▲ beat CLIP-Guided Reinforcement Learning for Ope… ×6
- ▼ lost to De novo Protein Design Using Geometric Vec… ×4
- ▼ lost to Dynamics-Informed Protein Design with Stru… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 42)