Generative Pre-Trained Speech Language Model with Efficient Hierarchical Transformer
Yongxin Zhu, Dan Su, andylqhe@tencent.com, Linli Xu, Dong Yu
OpenReview ground truth
Abstract
While recent advancements in speech language modeling have achieved significant progress, they face remarkable challenges in modelling the long acoustic sequence of neural audio codecs. Previous speech language models are compelled to learn acoustic tokens through a multi-stage generation process, which hinders their performance due to error propagation and information loss. In this paper, we introduce \textbf{G}enerative \textbf{P}re-Trained \textbf{S}peech Language Model (GPST), a hierarchical transformer designed for efficient speech language modeling. GPST quantizes audio waveforms into two distinct types of discrete speech representations and integrates them within a hierarchical transformer architecture, allowing for a unified one-stage generation process and enhancing Hi-Res audio generation capabilities. By training on large corpora of raw audio waveforms in an end-to-end unsupervised manner, GPST can generate syntactically consistent speech with diverse speaker identity unconditionally. When provided a brief 3-second prompt, GPST is able to produce natural and coherent personalized speech, demonstrating in-context learning abilities. Moreover, our approach can be easily extended to spoken cross-lingual speech generation by incorporating multi-lingual semantic tokens and universal acoustic tokens. Experimental results indicate that GPST significantly outperforms the existing speech language models in terms of word error rate, speech quality and speaker similarity.
Author context
Most prolific author: 6 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 36 comparisons
Ranked above opponent in 55% of matchups.
- ▼ lost to Detecting, Explaining, and Mitigating Memo… ×6
- ▲ beat MM-LDM: Multi-Modal Latent Diffusion Model… ×6
- ▼ lost to Rigid Protein-Protein Docking via Equivari… ×6
- ▲ beat Improving Compositional Text-to-image Gene… ×4
- ▼ lost to Language Model Beats Diffusion - Tokenizer… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 36)