Skip to content
Conference

Optimal Cluster Selection in Unsupervised Machine Learning Using K-Means Clustering

Aug 2026 · International Conference on Computing Communication Control and automation · pp. 1-5 · 0 citations · 16 references

Abstract

In unsupervised learning, Clustering is a core method used to determine unseen arrangements and structures within datasets by grouping similar instances together. Among the many clustering algorithms, Among clustering techniques, K-Means continues to be one of the most popular owing to its ease of implementation, fast execution, and broad applicability. However, a key limitation of K-means lies in the essential to predefine the optimal number of clusters (ONC), where an improper selection can significantly degrade clustering performance. This paper investigates three established methods for determining the ONC. The Elbow Method, the Silhouette Score, and the Gap Statistic. The mentioned technique is useful to the Iris dataset a classical benchmark in data analysis and evaluated using four metrics: the matrixes are used are Adjusted Rand Index (ARI), Root Mean Squared Error (RMSE), Mean Squared Error (MSE), and execution time. Among the methods, the Silhouette Score emerged as the most effective, achieving the highest ARI (0.54) and lowest RMSE (0.58), indicating superior clustering quality and robustness. The Gap Statistic performed moderately well with ARI of 0.39 and the fastest runtime (0.02 s), while the Elbow Method, though the quickest (0.01 s), produced the lowest ARI (0.33). These findings highlight the Silhouette Score as the most reliable method for optimal cluster selection in practical unsupervised learning tasks.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.