← ICLR 2024 leaderboard

PPTSER: A Plug-and-Play Tag-guided Method for Few-shot Semantic Entity Recognition on Visually-rich Documents

Wenhui Liao, Jiapeng Wang, Longfei Xiong, Lianwen Jin

representation learningFew-shot LearningSemantic Entity RecognitionMulti-modal Pre-trained ModelsPrompt Learning
34.00100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
45.10100
Mimo
band ≈ ±21 pct pts (from σ = 0.42)
28.90100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.38)

OpenReview ground truth

Rejected

TL;DR — A simple yet effective Plug-and-Play Tag-guided method for few-shot Semantic Entity Recognition on visually-rich documents

Abstract

Visually-rich document information extraction (VIE) is a vital aspect of document understanding, wherein Semantic Entity Recognition (SER) plays a significant role. However, the study of few-shot SER on visually-rich documents remains largely unexplored despite its considerable potential for practical applications. To address this issue, we propose a simple yet effective Plug-and-Play Tag-guided method for few-shot Semantic Entity Recognition (PPTSER) on visually-rich documents. PPTSER is a pluggable method building upon off-the-shelf multi-modal pre-trained models. It leverages the semantics of the tags to guide the SER task. In essence, PPTSER reformulates SER into entity typing and span detection, handling both tasks simultaneously via cross-attention. Experimental results illustrate that PPTSER outperforms fine-tuning baseline and existing few-shot methods, especially in low-data regimes. With full training data, PPTSER achieves comparable or superior performance to fine-tuning baseline. Specifically, on the FUNSD benchmark, our method improves the performance of LayoutLMv3 in 1-shot, 3-shot and 5-shot scenarios by 15.61%, 2.13%, and 2.01%, respectively. On the XFUND-zh benchmark, it improves the performance of LayoutLMv3 by 3.73%, 6.16%, and 4.01%, respectively. Overall, PPTSER demonstrates promising generalizability, effectiveness, and plug-and-play nature for few-shot SER on visually-rich documents. The codes will be available.

Author context

Most prolific author: 2 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 36)