Abstract
With recent technological advancements, especially the rapid development of the Internet and the popularity of consumer electronic devices (such as computers, digital cameras, and mobile phones), video production and watching have become much easier for individuals. The video content, the number of video viewers, as well as the viewing time, have significantly increased in recent years. Effectively organizing and managing videos are not only beneficial for video viewers but also necessary for video makers and distributors. In addition to the normal video content, the affective video content is also a considerable aspect when organizing and managing videos. With the rapid increase in online video sharing, manually tagging emotion labels to videos has become impossible. Therefore, it is necessary to develop methods for automatically labelling videos according to their affective content. In this thesis, various tasks ranging from direct affective video content analysis and its applications in automatic video emotion tagging to affective audio-visual correspondence learning are addressed. First, we tackle the task of predicting the affective responses of video viewers in two different contexts: dynamic and static. Various multimodal deep neural networks are designed to achieve a state-of-the-art performance on different datasets. In addition, we also analyze the importance of the use of multiple modalities (video, audio, and text in the form of video subtitles) on the performance of our proposed models for the affective response prediction. Second, our static affective response prediction models are applied to automatically tag global emotion labels to music video segments collected from various sources. From the music video segments with automatically tagged emotion labels, we construct a collection of three datasets dedicated to affective audio-visual correspondence learning. In addition to the dataset creation, two tasks including binary affective music-video correspondence classification and affective music-video retrieval are also addressed. A benchmark deep neural network model is first proposed for binary affective music-video correspondence classification. This proposed benchmark model is then modified to adapt to affective music-video retrieval. Our proposed model outperforms state-of-the-art approaches on both the binary classification and retrieval tasks on the newly created dataset collection.