PapersWithELO
← ICLR 2024 leaderboard

Enhancing Clinical Note Summarization: Iterative Reflexions with Small-model Supervision and Error2Correct Demonstrations

Jingping Liu, Boyang Zhong, Weiyan Zhang, Haiyun Jiang, Weichao Ding, Tong Ruan, sj11788@rjh.com.cn, Lifeng Zhu

physical sciencesClinical Note SummarizationLarge Language ModelError2Correct demonstration
21.80100
Fused
band ≈ ±14 pct pts (from σ = 0.29)
16.90100
Mimo
band ≈ ±22 pct pts (from σ = 0.43)
20.70100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.38)

OpenReview ground truth

Rejected

TL;DR — We are the first to propose a novel iterative reflexion framework for large and small model collaboration with Error2Correct demonstrations in clinical note summarization, and achieve SoTA performance on Chinese and English datasets.

Abstract

Generating clinical notes from doctor-patient dialogues is an important task in medical artificial intelligence. Mainstream methods currently employ large language models with few-shot demonstrations to tackle this challenge. However, the absence of domain knowledge supervision in these models often results in issues like missing key information, irregular writing standards, and non-compliant language styles. To this end, in this paper, we propose a novel iterative reflexion framework with small-model supervision and Error2Correct demonstrations for clinical note summarization. In this framework, we leverage a large model to produce clinical notes and design a small model trained on domain-specific data to evaluate the generated content. To enhance the quality of the generated clinical notes, we further propose Error2Correct demonstrations, which consist of error examples, error analysis, and corresponding correct examples, to help the large model identify and rectify errors effectively. To evaluate the effectiveness of our proposed method, we conduct extensive experiments on both Chinese and English datasets. The results demonstrate that our method achieves state-of-the-art performance on both datasets for the clinical note summarization task.

Author context

Most prolific author: 3 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 32 comparisons

Ranked above opponent in 39% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 32)