Analysis and Clustering of Plantation Crop Production in Indonesia using K-Means Clustering and Principal Component Analysis Methods

Alfredo Grasia Valdano, Styawati Styawati

Abstract


Indonesia has substantial potential in the plantation sector, with production levels varying considerably across provinces. These differences are influenced by geographical conditions, land availability, and the leading commodities of each region. This study aims to identify production patterns of plantation crops in Indonesia using K-Means Clustering and to visualize the resulting clusters using Principal Component Analysis (PCA). Official 2024 BPS data covering the production of oil palm, coconut, rubber, coffee, cocoa, tea, and sugarcane across 38 provinces were analyzed. The preprocessing procedure included outlier treatment using Interquartile Range (IQR)-based capping, logarithmic transformation, normalization using StandardScaler, dimensionality reduction using PCA, and clustering using the K-Means algorithm. The optimal number of clusters was determined using the Elbow Method, Silhouette Score, and Davies-Bouldin Index (DBI). The results indicate that K = 5 was the optimal clustering solution, achieving a Silhouette Score of 0.5883 and a DBI of 0.4744. The clustering results classified Indonesian provinces into five production categories: very high, high, moderate, low, and very low. Provinces in Sumatra and Kalimantan dominated the very-high and high production categories, while several provinces fell into the moderate category. In contrast, most provinces in eastern Indonesia were classified into the low and very-low production categories. Furthermore, the combination of K-Means Clustering and PCA proved effective for categorizing provincial production data and identifying spatial patterns in plantation crop production across Indonesia. The findings can support government agencies and other stakeholders in formulating plantation-sector development policies and establishing regional priorities based on the production characteristics of individual provinces.

Keywords


data mining; K-Means clustering; principal component analysis (PCA); plantation crop production; provincial clustering

Full Text:

PDF

References


S. Fevriera and F. Safara Devi, “Analisis Produksi Kelapa Sawit Indonesia: Pendekatan Mikro dan Makro Ekonomi,” Transformatif, Vol. 12, No. 1, pp. 1–16, May 2023.

R. Situmeang, S. Marta Lina, N. Rodiyah Hisan, and S. Vaulina, “Peran Subsektor Perkebunan terhadap Pertumbuhan Ekonomi di Provinsi Riau: The Role of Plantation Subsector to Economic Growth in Riau Province,” Jurnal Agroteknologi Agribisnis dan Akuakultur, Vol. 4, No. 1, pp. 51–60, Jan. 2024.

A. Said, A. Akhmad, I. Sribianti, M. Natsir, and M. Maulina, “Analisis Pengaruh Produksi dan Luas Lahan Kelapa Sawit terhadap PDRB Sektor Pertanian: Pendekatan Regresi Linier Berganda menggunakan Data Sekunder 2013-2022,” Jurnal Ilmu Manajemen Sosial Humaniora (JIMSH), Vol. 6, No. 1, pp. 46–56, Jul. 2024, DOI: 10.51454/jimsh.v6i1.632.

N. P. Daulay, A. Alyani Putri, O. N. Ramadani, E. V. Agata, and M. Hamid, “Analisis Perkembangan Produksi Kelapa Sawit berdasarkan Luas Lahan dan Jumlah Pabrik Kelapa Sawit (Studi Pada Kabupaten Rokan Hulu),” Jurnal Pengabdian Masyarakat dan Riset Pendidikan, Vol. 4, No. 2, pp. 11335–11342, 2025, DOI: 10.31004/jerkin.v4i2.3935.

I. Akbar, I. S. Samad, R. Rahmat, and S. Rosmiana, “Data Mining Analysis of K-Means Algorithm and Decision Tree for Early Detection of Students at Risk of Dropping Out,” Journal of Informatics Information System Software Engineering and Applications (INISTA), Vol. 7, No. 2, pp. 148–162, May 2025, DOI: 10.20895/inista.v7i2.1630.

M. Rochmawati, G. W. C. Bagaskara, I. A. Adha, Y. Umaidah, and A. Voutama, “Implementasi Algoritma K-Means dalam Klasterisasi Penjualan pada Sebuah Perusahaan menggunakan Metodologi KDD,” SISTEMASI: Jurnal Sistem Informasi, Vol. 13, No. 1, pp. 54–62, Jan. 2024, [Online]. Available: http://sistemasi.ftik.unisi.ac.id

A. Mafera, H. Susilo, and I. N. Sulrieni, “Application of Data Mining With K-Means Clustering Algorithm for Hypertension Disease Classification at Puskesmas Lubuk Buaya Padang in 2024,” Journal of Medical Records and Information Technology (JOMRIT), Vol. 2, No. 2, pp. 1–5, Dec. 2024, [Online]. Available: https://jurnal.syedzasaintika.ac.id/index.php/jomrit

I. Firman Ashari, R. Banjarnahor, D. R. Farida, S. P. Aisyah, A. P. Dewi, and N. Humaya, “Application of Data Mining with the K-Means Clustering Method and Davies Bouldin Index for Grouping IMDB Movies,” Journal of Applied Informatics and Computing (JAIC), Vol. 6, No. 1, pp. 7–15, 2022, [Online]. Available: http://jurnal.polibatam.ac.id/index.php/JAIC

A. Alif et al., “Penerapan Algoritma K-means Clustering dan Hierarchical Clustering dalam Mengelompokkan Data Pengangguran di Karawang,” Algoritma, Vol. 21, No. 2, pp. 209–220, Nov. 2024, DOI: 10.33364/algoritma/v.21-2.2155.

R. Maulidina and S. Y. Riska, “Application of the K-Means Algorithm for Clustering Plantation Crop Production in Indonesia,” SMATIKA JURNAL, Vol. 13, No. 02, pp. 339–349, Dec. 2023, DOI: 10.32664/smatika.v13i02.991.

A. A. Simangunsong, I. Gunawan, Z. M. Nasution, and G. Artikel, “Pengelompokkan Hasil Produksi Tanaman Perkebunan berdasarkan Provinsi menggunakan Metode K-Means Clustering Production of Plantation Crops by Province using the K-Means Method Article Info ABSTRAK,” JOMLAI: Journal of Machine Learning and Artificial Intelligence, Vol. 1, No. 4, pp. 2828–9099, 2022, DOI: 10.55123/jomlai.v1i4.1661.

I. A. Rosyada and D. T. Utari, “Penerapan Principal Component Analysis untuk Reduksi Variabel pada Algoritma K-Means Clustering,” Jambura J. Probab. Stat, Vol. 5, No. 1, pp. 6–13, 2024, DOI: 10.34312/jjps.v5i1.18733.

R. Ishak, “Optimalisasi Seleksi Atribut K-Means Menggunakan Correlation Matrix pada Clustering Penyakit Pasien Optimization of K-Means Attribute Selection using Correlation Matrix in Patient Disease Clustering,” Jambura Journal of Electrical and Electronics Engineering, Vol. 7, No. 2, pp. 141–148, Jul. 2025.

S. Dewi and M. A. I. Pakereng, “Implementasi Principal Component Analysis pada K-Means untuk Klasterisasi Tingkat Pendidikan Penduduk Kabupaten Semarang,” JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika), Vol. 8, No. 4, pp. 1186–1195, Dec. 2023, DOI: 10.29100/jipi.v8i4.4101.

D. A. Awaliyah, Budi Prasetiyo, R. Muzayanah, and A. D. Lestari, “Optimizing Customer Segmentation in Online Retail Transactions through the Implementation of the K-Means Clustering Algorithm,” Scientific Journal of Informatics, Vol. 11, No. 2, pp. 539–548, Jun. 2024, DOI: 10.15294/sji.v11i2.6137.

R. Situmeang, S. M. Lina, N. R. Hisan, and S. Vaulina, “Peran Subsektor Perkebunan terhadap Pertumbuhan Ekonomi di Provinsi Riau: The Role of Plantation Subsector to Economic Growth in Riau Province,” Jurnal Agroteknologi Agribisnis dan Akuakultur, Vol. 4, No. 1, pp. 51–60, Jan. 2024.

P. Villalobos Perna, M. Di Febbraro, M. L. Carranza, F. Marzialetti, and M. Innangi, “Remote Sensing and Invasive Plants in Coastal Ecosystems: What We Know so Far and Future Prospects,” Land (Basel)., Vol. 12, No. 2, p. 341, Feb. 2023, DOI: 10.3390/land12020341.

C. C. Mireștean, R. I. Iancu, and D. T. Iancu, “Radiotherapy and Immunotherapy—A Future Partnership towards a New Standard,” Applied Sciences, Vol. 13, No. 9, p. 5643, May 2023, DOI: 10.3390/app13095643.

S. P. Sipayung and P. M. Hasugian, “Integrating PCA and K-Means for Evidence-based Staple Food Segmentation: An Indonesian Food Policy Approach,” Sinkron: Jurnal dan Penelitian Teknik Informatika, Vol. 9, No. 4, pp. 3175–3189, Oct. 2025, DOI: 10.33395/sinkron.v9i4.15343.

A. C. Ehresmann, M. Béjean, and J. P. Vanbremeersch, “A Mathematical Framework for Enriching Human–Machine Interactions,” Mach. Learn. Knowl. Extr., Vol. 5, No. 2, pp. 597–610, Jun. 2023, DOI: 10.3390/make5020034.

K. H. Izzuddin and A. W. Wijayanto, “Pemodelan Clustering Ward, K-Means, Diana, dan PAM dengan PCA untuk Karakterisasi Kemiskinan Indonesia Tahun 2021,” Komputika : Jurnal Sistem Komputer, Vol. 13, No. 1, pp. 41–53, Apr. 2024, DOI: 10.34010/komputika.v13i1.10803.

H. Lestari Siregar, R. Hidayanthi, and A. Langga Dewa Sakti, “Implementation of K-Means Clustering on Student Learning Achievements based on Social Economic and Social Related,” Research and Development in Education, Vol. 4, No. 2, pp. 1447–1459, Dec. 2024, DOI: 10.22219/raden.v4i2.36742.

E. Arry Kusuma and A. Dharmawati, “Analisis Pemerataan Pendidikan di Indonesia menggunakan Reduksi Dimensi PCA dan Klasterisasi K-Means,” JURNAL FASILKOM, Vol. 16, No. 1, pp. 105–112, Apr. 2026.

R. Rianti, R. Andarsyah, and R. M. Awangga, “Penerapan PCA dan Algoritma Clustering untuk Analisis Mutu Perguruan Tinggi di LLDIKTI Wilayah IV,” NUANSA INFORMATIKA, Vol. 18, No. 2, pp. 67–77, Jul. 2024, [Online]. Available: https://journal.fkom.uniku.ac.id/ilkom

I. K. O. Jabari, Shofiyah, P. K. S, N. N. Putriwijaya, and N. Yudistira, “Learning-Augmented K-Means Clustering using Dimensional Reduction,” ArXiv, Jan. 2024, DOI: 10.1145/3626641.3627239.

N. Migenda, R. Moller, and W. Schenck, “Adaptive Dimensionality Reduction for Neural Network-based Online Principal Component Analysis,” PLoS One, Vol. 16, No. 3 March, p. e0248896, Mar. 2021, DOI: 10.1371/journal.pone.0248896.

O. Assani-Amate, M. Bakhtyari, É. Roy, and V. Makarenkov, “Assessing the Impact of Dimensionality Reduction on Clustering Performance -- A Systematic Study,” May 2026, [Online]. Available: http://arxiv.org/abs/2604.22099

P. Chen, L. Wu, and L. Wang, “AI Fairness in Data Management and Analytics: A Review on Challenges, Methodologies and Applications,” Applied Sciences, Vol. 13, No. 18, p. 10258, Sep. 2023, DOI: 10.3390/app131810258.

Y. Fernando, R. Lukman, A. Jayadi, and R. Nuraini, “Sistem Otomatisasi Pemberi Pakan pada Peternakan Bebek dengan Arduino UNO dan Bluetooth menggunakan SmartPhone,” TIN: Terapan Informatika Nusantara, Vol. 3, No. 2, pp. 42–50, Jul. 2022, DOI: 10.47065/tin.v3i2.1789.

J. Arellano-Uson, E. Magaña, D. Morato, and M. Izal, “Survey on Quality of Experience Evaluation for Cloud-based Interactive Applications,” Applied Sciences, Vol. 14, No. 5, p. 1987, Feb. 2024, DOI: 10.3390/app14051987.

S. V. Ludkowski, “Inverse Spectrum and Structure of Topological Metagroups,” Mathematics, Vol. 12, No. 4, p. 511, Feb. 2024, DOI: 10.3390/math12040511.

A. Zaki, Irwan, and I. A. Sembe, “Penerapan K-Means Clustering dalam Pengelompokan Data (Studi Kasus Profil Mahasiswa Matematika FMIPA UNM),” Journal of Mathematics, Computations, and Statistics, Vol. 5, No. 2, pp. 163–176, Oct. 2022, Accessed: May 16, 2026. [Online]. Available: http://http://www.ojs.unm.ac.id/jmathcos.ojs.unm.ac.id/jmathcos

A. A. A. Daniswara and I. K. D. Nuryana, “Data Preprocessing Pola pada Penilaian Mahasiswa Program Profesi Guru,” Journal of Informatics and Computer Science, Vol. 5, No. 1, pp. 97–100, Mar. 2023.

M. S. Arif, A. Mukheimer, and D. Asif, “Enhancing the Early Detection of Chronic Kidney Disease: A Robust Machine Learning Model,” Big Data and Cognitive Computing, Vol. 7, No. 3, p. 144, Sep. 2023, DOI: 10.3390/bdcc7030144.

L. G. Zamarrenho et al., “Effects of Three Different Brazilian Green Propolis Extract Formulations on Pro- and Anti-Inflammatory Cytokine Secretion by Macrophages,” Applied Sciences, Vol. 13, No. 10, p. 6247, May 2023, DOI: 10.3390/app13106247.

B. Yang, M. H. Arshad, and Q. Zhao, “Packet-Level and Flow-Level Network Intrusion Detection based on Reinforcement Learning and Adversarial Training,” Algorithms, Vol. 15, No. 12, p. 453, Dec. 2022, DOI: 10.3390/a15120453.

M. Almulhem, S. Z. Hassan, A. Al-buainain, M. A. Sohaly, and M. A. E. Abdelrahman, “Characteristics of Solitary Stochastic Structures for Heisenberg Ferromagnetic Spin Chain Equation,” Symmetry (Basel)., Vol. 15, No. 4, p. 927, Apr. 2023, DOI: 10.3390/sym15040927.

Z. Chen, X. Xiong, F. Meng, X. Xiao, and J. Liu, “Scaling-Invariant Max-Filtering Enhancement Transformers for Efficient Visual Tracking,” Electronics (Switzerland), Vol. 12, No. 18, p. 3905, Sep. 2023, DOI: 10.3390/electronics12183905.

Y. Wang, C. Hu, Z. Li, D. Zheng, F. Cui, and X. Yang, “Theoretical and Simulation Analysis of Static and Dynamic Properties of MXene-Based Humidity Sensors,” Applied Sciences, Vol. 12, No. 16, p. 8254, Aug. 2022, DOI: 10.3390/app12168254.

N. Andriyanov, “The use of Correlation Features in the Problem of Speech Recognition,” Algorithms, Vol. 16, No. 2, p. 90, Feb. 2023, DOI: 10.3390/a16020090.




DOI: https://doi.org/10.32520/stmsi.v15i8.6534

Article Metrics

Abstract view : 1 times
PDF - 0 times

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.