Comparative Analysis of Instance Segmentation Models for House-Level Visual Socioeconomic Classification using Satellite Imagery

David Cahyapratama, Chairani Fauzi

Abstract


House-level visual socioeconomic information can provide a more detailed spatial understanding of residential areas and can support applied urban and commercial analysis. However, official socioeconomic data are often available only at broader administrative levels and may not capture variation among small residential clusters or individual houses. This study compares five instance segmentation models for detecting, segmenting, and classifying individual houses into lower, middle, and upper visual socioeconomic classes using high-resolution satellite imagery. The evaluated models are Mask R-CNN, Cascade Mask R-CNN, YOLO11n-seg, SOLOv2, and Mask2Former. The dataset consists of 1,209 training images with 7,052 annotations and 213 validation images with 1,353 annotations from residential areas in Banten and DKI Jakarta. Annotation reliability was assessed using 100 annotation pairs, producing a quadratic-weighted Cohen’s kappa of 0.8077. Model performance was evaluated using COCO metrics for both segmentation masks and bounding boxes. The results show that Cascade Mask R-CNN achieved the highest observed overall validation performance among the tested configurations. Under the current experimental setting, it produced the strongest combination of object-localization and mask-segmentation metrics. These findings show that comparing multiple instance segmentation models can help identify a more suitable method for house-level visual socioeconomic classification. Unlike previous studies that generally perform area-level socioeconomic estimation or building extraction alone, this study compares multiple instance-segmentation approaches for expert-defined socioeconomic classification at the individual-house level. The resulting output can serve as a supplementary visual socioeconomic layer that complements demographic, accessibility, and commercial data in applied spatial and market-development analyses.

Keywords


Instance segmentation; Socioeconomic classification; Residential area; Satellite imagery; Model comparison

Full Text:

PDF

References


A. Arshad, J. Zulfiqar, M. H. Zaib, A. Khan, and M. J. Khan, “Mapping Socioeconomic Conditions using Satellite Imagery: A Computer Vision Approach for Developing Countries,” Journal of Economy and Technology, Vol. 1, pp. 144–163, Nov. 2023, DOI: 10.1016/j.ject.2023.11.001.

F. S. K. Agyemang, R. Memon, L. J. Wolf, and S. Fox, “High-Resolution Rural Poverty Mapping in Pakistan with Ensemble Deep Learning,” PLoS One, Vol. 18, No. 4 April, Apr. 2023, DOI: 10.1371/journal.pone.0283938.

O. Hall, F. Dompae, I. Wahab, and F. M. Dzanku, “A Review of Machine Learning and Satellite Imagery for Poverty Prediction: Implications for Development Research and Applications,” Oct. 01, 2023, John Wiley and Sons Ltd. DOI: 10.1002/jid.3751.

F. Biljecki and K. Ito, “Street view Imagery in Urban Analytics and GIS: A Review,” Nov. 01, 2021, Elsevier B.V. DOI: 10.1016/j.landurbplan.2021.104217.

A. Singleton, D. Arribas-Bel, J. Murray, and M. Fleischmann, “Estimating Generalized Measures of Local Neighbourhood Context from Multispectral Satellite Images using a Convolutional Neural Network,” Comput. Environ. Urban Syst., Vol. 95, Jul. 2022, DOI: 10.1016/j.compenvurbsys.2022.101802.

E. Suel, S. Bhatt, M. Brauer, S. Flaxman, and M. Ezzati, “Multimodal Deep Learning from Satellite and Street-Level Imagery for Measuring Income, Overcrowding, and Environmental Deprivation in Urban Areas,” Remote Sens. Environ., Vol. 257, May 2021, DOI: 10.1016/j.rse.2021.112339.

H. Sarmadi, I. Wahab, O. Hall, T. Rögnvaldsson, and M. Ohlsson, “Human Bias and CNNs’ Superior Insights in Satellite based Poverty Mapping,” SCI. Rep., Vol. 14, No. 1, Dec. 2024, DOI: 10.1038/s41598-024-74150-9.

O. L. F. de Carvalho et al., “Instance Segmentation for Large, Multi-Channel Remote Sensing Imagery using Mask-RCNN and a Mosaicking Approach,” Remote Sens. (Basel)., Vol. 13, No. 1, pp. 1–24, Jan. 2021, DOI: 10.3390/rs13010039.

T. Hou and J. Li, “Application of Mask R-CNN for Building Detection in UAV Remote Sensing Images,” Heliyon, Vol. 10, No. 19, Oct. 2024, DOI: 10.1016/j.heliyon.2024.e38141.

M. Amo-Boateng, N. Ekow Nkwa Sey, A. Ampah Amproche, and M. Kyereh Domfeh, “Instance Segmentation Scheme for Roofs in Rural Areas based on Mask R-CNN,” Egyptian Journal of Remote Sensing and Space Science, Vol. 25, No. 2, pp. 569–577, Aug. 2022, DOI: 10.1016/j.ejrs.2022.03.017.

S. Gui, S. Song, R. Qin, and Y. Tang, “Remote Sensing Object Detection in the Deep Learning Era—A Review,” Jan. 01, 2024, Multidisciplinary Digital Publishing Institute (MDPI). DOI: 10.3390/rs16020327.

J. Butler and H. Leung, “A Heatmap-Supplemented R-CNN Trained using an Inflated IoU for Small Object Detection,” Remote Sens. (Basel)., Vol. 16, No. 21, Nov. 2024, DOI: 10.3390/rs16214065.

S. Chen, Y. Ogawa, C. Zhao, and Y. Sekimoto, “Large-Scale Individual Building Extraction from Open-Source Satellite Imagery via Super-Resolution-based Instance Segmentation Approach,” ISPRS Journal of Photogrammetry and Remote Sensing, Vol. 195, pp. 129–152, Jan. 2023, DOI: 10.1016/j.isprsjprs.2022.11.006.

K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask R-CNN,” in Proceedings of the IEEE International Conference on Computer Vision, Institute of Electrical and Electronics Engineers Inc., Dec. 2017, pp. 2980–2988. DOI: 10.1109/ICCV.2017.322.

Z. Cai and N. Vasconcelos, “Cascade R-CNN: Delving into High Quality Object Detection,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2018. DOI: 10.1109/CVPR.2018.00644.

H. Y. Oh, M. S. Khan, S. B. Jeon, and M. H. Jeong, “Automated Detection of Greenhouse Structures using Cascade Mask R-CNN,” Applied Sciences (Switzerland), Vol. 12, No. 11, Jun. 2022, DOI: 10.3390/app12115553.

X. Wang, R. Zhang, T. Kong, L. Li, and C. Shen, “SOLOv2: Dynamic and Fast Instance Segmentation,” in Advances in Neural Information Processing Systems, 2020.

Z. Qiu, X. Huang, Z. Sun, S. Li, and J. Wang, “GS-YOLO-Seg: A Lightweight Instance Segmentation Method for Low-Grade Graphite Ore Sorting based on Improved YOLO11-Seg,” Sustainability (Switzerland), Vol. 17, No. 12, Jun. 2025, DOI: 10.3390/su17125663.

B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-Attention Mask Transformer for Universal Image Segmentation,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2022. DOI: 10.1109/CVPR52688.2022.00135.

S. Mukodimah and C. Fauzi, “Comparison of Tree Implementation, Regression Logistics, and Random Forest to Detect Iris Types,” Technology Acceptance Model, Vol. 12, No. 2, Oct. 2021.

R. Bakeman, “KappaAcc: A Program for Assessing the Adequacy of Kappa,” Behav. Res. Methods, Vol. 55, No. 2, 2023, DOI: 10.3758/s13428-022-01836-1.

T.-Y. Lin et al., “Microsoft COCO: Common Objects in Context,” in ECCV, European Conference on Computer Vision, Sep. 2014. [Online]. Available: https://www.microsoft.com/en-us/research/publication/microsoft-coco-common-objects-in-context/

J. Cohen, “Weighted Kappa: Nominal Scale Agreement with Provision for Scaled Disagreement or Partial Credit,” Psychol. Bull., Vol. 70, No. 4, pp. 213–220, 1968, DOI: 10.1037/h0026256.




DOI: https://doi.org/10.32520/stmsi.v15i8.6743

Article Metrics

Abstract view : 1 times
PDF - 0 times

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.