Logo image
Exploring and Improving Consistency in Large Language Models for Multiple-Choice Question Assessment
Conference proceeding

Exploring and Improving Consistency in Large Language Models for Multiple-Choice Question Assessment

Wenjie Zhou, Xiangyu Duan and IEEE
Proceedings of ... International Joint Conference on Neural Networks, pp.1-9
30/06/2024

Abstract

Adaptation models Consistency Metrics Data models Fine-Tuning Techniques In-context learning Large language models Measurement Model Inconsistency Multi-Choice Questions Neural networks Symbols Training data
With the evolution of Large Language Models (LLMs), accurately evaluating their capabilities has become a critical focus. The Multi-choice Questions (MCQ) benchmark is widely adopted for its definitive answers and straightforward assessment approach. However, recent studies have found that there is inconsistency in the model, raising concerns about potential biases and their genuine comprehension abilities in MCQ contexts. To delve into and improve the consistency performance of LLMs in MCQ answering domain, this study first proposes new consistency metrics, including option position consistency and option symbol consistency. These metrics are designed to quantify and reveal the level of consistency in the model, thereby assessing the authenticity of its knowledge comprehension. Secondly, utilizing these metrics, we propose two novel improvement strategies: 1) An enhanced In-context Learning (ICL) prompt customization technique, which adaptively modifies prompts based on the model's demonstrated capabilities, aligning the prompts with the model's inherent abilities to filter out questions it deems consistently answerable; and 2) Consistency Supervised Fine-Tuning (CSFT), which enriches the training data set focused on consistency, followed by specialized fine-tuning to augment the model's inherent capabilities. Our research is committed to exploring and improving consistency levels in models, with the goal of bolstering their integrity and reliability in the realm of MCQ answering.

Metrics

1 Record Views

Details

Logo image