Abstract
Focus group discussions (FGD) are an important research method to obtain in-depth insights about social phenomena. Despite its prevalence and usefulness, qualitative research methods such as FGDs are incredibly laborious as such method involves inductive hand coding of text corpuses such as transcripts. To overcome this labour-intensive process, Topic Modelling techniques can be used to automate and expedite the process of analysing focus group transcripts. Therefore, this paper aims to use Topic Modelling techniques to create a more efficient system for the analysis of focus group transcripts. Concurrently, this paper aims to offer three main contributions. Firstly, this paper aims to solve the challenge of contextual Topic Modelling in dialogue system by proposing two segmentation methods to aggregate focus group transcripts into two different corpuses that share the same context. Secondly, this paper seeks to create a Topic Modelling pipeline with Transformer-based language model and demonstrate its effectiveness in Topic Modelling focus group transcript. Lastly, this paper aims to provide useful insights about the application of Topic Modelling techniques in qualitative analysis to the research community by comparing and evaluating the performance of different Topic Modelling Techniques in analyzing focus group transcripts. This paper applies four popular and state-of-the-art Topic Modelling techniques: LDA, Word2Vec, LFLDA and BERT, to analyze focus group transcripts which are aggregated into three hierarchical information level: Top level, Topic level and Stakeholder level. The resulting Topic Models are then evaluated quantitatively and qualitatively. Quantitatively, the Topic Models are evaluated with CV and C_UCI Coherence Scores. Qualitatively, the principal investigator who conducted the focus groups assessed the Topic Quality of the Topic Models. The results show that LFLDA is the best performing model out of the four Topic Models compared. On the other hand, the results also show that Topic Modelling techniques perform the best when the dataset is first aggregated at the Topic Level. However, the performance of the Topic Modelling techniques in analyzing focus group transcripts leave much space to be improved. This can be attributed to the fact that current Topic Modelling techniques do not account for the local slang used in conversational dialogue. On the flip side, the principal investigator who evaluated the Topic Models qualitatively noted great potential in such a tool in qualitative analysis. In his opinion, such a tool would greatly help in explorative analysis where inductive reasoning is used. In conclusion, the current Topic Modelling techniques require more refinement and fine-tuning to greatly improve their affinity to the local conversational slang and enhance their performance in analysing focus group transcripts.