Abstract
Walkable urban environments are increasingly recognized as essential for sustainable mobility, public health, and social well-being. While macro-scale indicators of walkability are widely used, growing evidence highlights the importance of street-level physical conditions experienced at eye level. Advances in computer vision and street view imagery (SVI) offer new opportunities to quantify such streetscape characteristics, yet the applicability of existing semantic segmentation models in developing urban contexts remains underexplored. This study evaluates the suitability of five state-of-the-art semantic segmentation models for streetscape analysis using crowdsourced SVI from Phnom Penh, Cambodia. Through a comparative analysis, Oneformer was identified as the most suitable semantic segmentation model, uniquely successful in identifying street vendors through surrogate semantic class (base) and street furniture. A rigorous quantitative validation using manually annotated images confirmed the model's reliability, achieving an mIoU of 65.7% within the complex urban fabric of Phnom Penh. This performance stems from OneFormer's unified task-conditioned framework, which integrates semantic, instance, and panoptic information within a single query. Such an architecture ensures enhanced boundary stability and semantic coherence by consolidating visual noise into meaningful units, making it particularly robust for processing the irregular street elements typical of Southeast Asian cities. Applying the selected model revealed pronounced spatial variation in streetscape composition across three neighborhoods, reflecting distinct development stages and levels of informality. These findings suggest that carefully selected pretrained models can yield analytically useful representations of streetscape conditions in data-constrained settings, supporting more context-sensitive and inclusive urban analysis in rapidly developing cities.