PapersWithELO
← ICLR 2024 leaderboard

Keqing: Knowledge-based Question Answering is A Nature Chain-of-Thought mentor of LLMs

Yishi Xu, Chaojie Wang, Zhong Peng, Chenxi Zhang, Xinrun Wang, Xinyang Liu, Bo Chen, Bo An, Lei Feng

neurosymbolic AIknowledge-based question answeringlarge language modelknwledge graph
8.00100
Fused
band ≈ ±15 pct pts (from σ = 0.30)
6.60100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
7.10100
DeepSeek
band ≈ ±22 pct pts (from σ = 0.45)

OpenReview ground truth

Rejected

Abstract

Large language models (LLMs) have exhibited remarkable performance on various natural language processing (NLP) tasks, especially for question answering. However, in the face of problems beyond the scope of knowledge, these LLMs tend to talk nonsense with a straight face, where the potential solution could be incorporating an Information Retrieval (IR) module and generating response based on these retrieved knowledge. In this paper, we present a novel framework to assist LLMs, such as ChatGPT, to retrieve question-related structured information on the knowledge graph, and demonstrate that Knowledge-based question answering (Keqing) could be a nature Chain-of-Thought (CoT) mentor to guide the LLM to sequentially find the answer entities of a complex question through interpretable logical chains. Specifically, the workflow of Keqing will execute decomposing a complex question according to predefined templates, retrieving candidate entities on knowledge graph, reasoning answers of sub-questions, and finally generating response with reasoning paths, which greatly improves the reliability of LLM’s response. The experimental results on KBQA datasets show that Keqing can achieve competitive performance and illustrate the logic of answering each question.

Author context

Most prolific author: 14 submissions (credibility 0.79).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 30 comparisons

Ranked above opponent in 35% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 30)