Logo image
A foundation model-based framework for unsupervised gaze anomaly detection
Journal article   Peer reviewed

A foundation model-based framework for unsupervised gaze anomaly detection

Kritika Johari, Jung-Jae Kim, Wei Quin Yow and U-Xuan Tan
Knowledge-based systems, Vol.324, p.113774
03/08/2025

Abstract

And timeseries foundation model Anomaly detection Gaze behavior analysis
Traditional gaze analysis methods for online lecture largely depend on predefined average gaze features and self-reported ground-truths, limiting their ability to obtain real-time status in unsupervised settings. To address this, we propose Gaze-READ (Gaze Representative Embedding and Anomaly Detection), a framework that integrates gaze behavior analysis with unsupervised anomaly detection to systematically identify attention shifts caused by external stimuli. Our approach leverages GazeMTM (Masked Time-Series Modeling of Gaze), which employs MOMENT, a time-series foundation model, to extract gaze embeddings that capture temporal dependencies in eye movements. We establish a baseline for normal gaze behavior using a control group (students without distractions) and apply unsupervised clustering to define representative gaze patterns. By comparing this baseline with gaze data from students exposed to distractors, Gaze-READ detects deviations, flagging them as potential indicators of distraction. Experimental results show that MOMENTGET (further pre-trained MOMENT) improves gaze reconstruction, reducing mean squared error (MSE) by at least 28% compared to its pre-trained version and outperforming baseline linear interpolation for oculomotor event representation. Additionally, Gaze-READ achieves a higher clustering silhouette score (0.38) and a lower Davies–Bouldin Index (0.88) than traditional time-series clustering methods, demonstrating its effectiveness in distinguishing gaze patterns. These findings highlight the potential of our framework to enable automated, real-time engagement tracking in online learning environments, offering a scalable solution for identifying attentional shifts without requiring labeled data. •First framework leveraging foundation model for unsupervised gaze anomaly detection.•Proposed GazeMTM to adapt MOMENT via masked modeling for gaze representations.•Introduced MOMENTGET, a pre-trained variant for extracting meaningful gaze embeddings.•Developed Gaze-READ for unsupervised detection of distraction-driven gaze anomalies.•Showed gaze clusters align with lecture structure and distractors impact anomalies.

Metrics

1 Record Views

Details

Logo image