Enhancing Clinical Note Summarization: Iterative Reflexions with Small-model Supervision and Error2Correct Demonstrations
Jingping Liu, Boyang Zhong, Weiyan Zhang, Haiyun Jiang, Weichao Ding, Tong Ruan, sj11788@rjh.com.cn, Lifeng Zhu
OpenReview ground truth
TL;DR — We are the first to propose a novel iterative reflexion framework for large and small model collaboration with Error2Correct demonstrations in clinical note summarization, and achieve SoTA performance on Chinese and English datasets.
Abstract
Generating clinical notes from doctor-patient dialogues is an important task in medical artificial intelligence. Mainstream methods currently employ large language models with few-shot demonstrations to tackle this challenge. However, the absence of domain knowledge supervision in these models often results in issues like missing key information, irregular writing standards, and non-compliant language styles. To this end, in this paper, we propose a novel iterative reflexion framework with small-model supervision and Error2Correct demonstrations for clinical note summarization. In this framework, we leverage a large model to produce clinical notes and design a small model trained on domain-specific data to evaluate the generated content. To enhance the quality of the generated clinical notes, we further propose Error2Correct demonstrations, which consist of error examples, error analysis, and corresponding correct examples, to help the large model identify and rectify errors effectively. To evaluate the effectiveness of our proposed method, we conduct extensive experiments on both Chinese and English datasets. The results demonstrate that our method achieves state-of-the-art performance on both datasets for the clinical note summarization task.
Author context
Most prolific author: 3 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 32 comparisons
Ranked above opponent in 39% of matchups.
- ▼ lost to Better Neural PDE Solvers Through Data-Fre… ×6
- ▲ beat Enhancing Small Medical Learners with Priv… ×6
- ▼ lost to MediTab: Scaling Medical Tabular Data Pred… ×4
- ▼ lost to Knowledge-Infused Prompting: Assessing and… ×4
- ▼ lost to Safe and Robust Watermark Injection with a… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 32)