PapersWithELO
← ICLR 2024 leaderboard

NL2ProGPT: Taming Large Language Model for Conversational Protein Design

Zekun Guo, Jianxin Lin, hnuzsf@hnu.edu.cn, Yijun Wang, Lijun Wu, xiangxiang Zeng

physical sciencesProtein designLarge language model
13.40100
Fused
band ≈ ±14 pct pts (from σ = 0.27)
12.80100
Mimo
band ≈ ±19 pct pts (from σ = 0.39)
12.90100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.39)

OpenReview ground truth

Rejected

Abstract

Large Language Models (LLMs), like ChatGPT, excel in cross-modal tasks thanks to their powerful abilities in natural language comprehension, generalization, and reasoning. Meanwhile, the wealth of human-curated protein knowledge in text form presents a unique opportunity for LLMs to contribute to advanced protein design. In this work, we propose a new LLMs-based framework, namely NL2ProGPT, for macromolecular protein sequence generation that bridges the domain gap between natural and protein languages. Specifically, we first combine the protein functions and properties to create specific text guidelines for designing the protein, ensuring it follows precise controls. Second, to form a more informative and generalizable protein description, we explicitly inject protein structural information by clustering the embeddings from pre-trained protein language models. Third, we train a reward model to align the protein language model with the Rosetta energy function, following an RLAIF (reinforced learning from AI feedback) fashion. We empirically verify the effectiveness of NL2ProGPT from three aspects: (1) outperforms existing protein sequence design methods in different evaluations; (2) exhibits more than 90\% consistency in text-to-protein generation; (3) has effective exploration potential in disordered regions.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 42)