Abstract
Knowledge about human mobility has numerous applications including urban planning, crowd monitoring, and targeted advertising. This knowledge can be drawn from the ever-growing massive amounts of human data generated daily with the improvements in smart technologies. In the process of extracting useful insights from raw human mobility data, machine learning can be a useful tool. However, the multiple machine learning methods and ways of application available can be varied and complex. Each specific case requires thought and consideration in terms of desired outcomes and available data. In this thesis, three systematic frameworks dealing with different mobility data situations are proposed as a starting point for such data analysis. These frameworks are applied to real-world data in each case to demonstrate their applicability, and their interesting insights are highlighted. Firstly, a multiple-perspective framework addressing the human mobility patterns surrounding a residential estate and its neighboring facility area is proposed. This framework puts forward the use of unsupervised machine learning techniques, namely clustering, to extract insights from the data when approached in three different ways - by time, person, and location - depending on the needs of the analysis. Next, a smoothing, peak detection, and clustering framework is proposed to compare the daily human mobility patterns within two residential estates. This proposed framework allows the machine learning algorithm to group daily temporal patterns based on main peak features and the resulting groups of patterns are largely distinct in node type and day type without such information being provided to the algorithm beforehand. Finally, a third framework is proposed to cluster groups of users using extracted distance-based features on both work days and off days. This framework is supplemented by two proposed cluster analysis metrics to highlight specific distance thresholds favored by users within each cluster, as well as illustrate the mobility pattern of an average user within each cluster.