Logo image
Modelling Text Similarity: A Survey
Conference proceeding

Modelling Text Similarity: A Survey

Wenchuan Mu and Kwan Hui Lim
Proceedings of the International Conference on Advances in Social Networks Analysis and Mining, pp.698-705
ACM Conferences
ASONAM '23: International Conference on Advances in Social Networks Analysis and Mining
06/11/2023

Abstract

Computing methodologies Computing methodologies -- Artificial intelligence Computing methodologies -- Artificial intelligence -- Natural language processing Computing methodologies -- Machine learning Computing methodologies -- Machine learning -- Learning paradigms Computing methodologies -- Machine learning -- Learning paradigms -- Unsupervised learning Computing methodologies -- Machine learning -- Learning paradigms -- Unsupervised learning -- Cluster analysis Information systems Information systems -- Information systems applications Information systems -- Information systems applications -- Data mining Mathematics of computing
Online social networking services such as Twitter and Instagram have become pervasive platforms for engaging in discussions on a wide array of topics. These platforms cater to both mainstream subjects, like music and movies, as well as more specialized areas, such as politics. With the growing volume of textual data generated on these platforms, the ability to define and identify similar texts becomes crucial for effective investigation and clustering. In this paper, we explore the challenges and significance of text similarity regression models in the context of online social networking services. We delve into the methods and techniques employed to define and find similarities among texts, enabling the extraction of meaningful patterns and insights. Specifically, we categorize text similarity regression models into four distinct types: set-theoretic, sequence-theoretic, real-vector, and end-to-end methods. This categorization is based on the mathematical formalisation of similarity used by each model. Ultimately, our survey aims to provide a comprehensive overview of the interlinkages between independently proposed methods for text similarity. By understanding the strengths and weaknesses of these methods, researchers can make informed decisions when designing novel approaches and algorithms. We hope this survey serves as a valuable resource for advancing the state-of-the-art in addressing the complex problem of text similarity.
url
https://doi.org/10.1145/3625007.3627305View
Published (Version of record) Open

Metrics

1 Record Views

Details

Logo image