verticapy.machine_learning.model_selection.elbow¶
- verticapy.machine_learning.model_selection.elbow(input_relation: Annotated[str | vDataFrame, ''], X: Annotated[str | list[str], 'STRING representing one column or a list of columns'] | None = None, n_cluster: tuple | list = (1, 15), init: Literal['kmeanspp', 'random', None] = None, max_iter: int = 50, tol: float = 0.0001, use_kprototype: bool = False, gamma: float = 1.0, show: bool = True, chart: PlottingBase | TableSample | Axes | mFigure | Highchart | Highstock | Figure | None = None, **style_kwargs) TableSample¶
Draws an Elbow curve.
Parameters¶
- input_relation: SQLRelation
Relation used to train the model.
- X: SQLColumns, optional
listof the predictor columns. If empty, all numerical columns are used.- n_cluster: tuple | list, optional
Tuple representing the number of clusters to start and end with. This can also be a customized list with various
kvalues to test.- init: str | list, optional
The method used to find the initial cluster centers.
- kmeanspp:
Only available when
use_kprototype = FalseUse thek-means++method to initialize the centers.
- random:
Randomly subsamples the data to find initial centers.
Default value is
kmeansppifuse_kprototype = False; otherwise,random.- max_iter: int, optional
The maximum number of iterations for the algorithm.
- tol: float, optional
Determines whether the algorithm has converged. The algorithm is considered converged after no center has moved more than a distance of
tolfrom the previous iteration.- use_kprototype: bool, optional
If set to
True, the function uses theKPrototypesalgorithm instead ofKMeans.KPrototypescan handle categorical features.- gamma: float, optional
Only if
use_kprototype = True. Weighting factor for categorical columns. It determines the relative importance of numerical and categorical attributes.- show: bool, optional
If set to
True, the Plotting object is returned.- chart: PlottingObject, optional
The chart object to plot on.
- **style_kwargs
Any optional parameter to pass to the Plotting functions.
Returns¶
- TableSample
nb_clusters,total_within_cluster_ss,between_cluster_ss,total_ss, elbow_score
Examples¶
The following examples provide a basic understanding of usage. For more detailed examples, please refer to the: Elbow Curve page.
Load data for machine learning¶
We import
verticapy:import verticapy as vp
Hint
By assigning an alias to
verticapy, we mitigate the risk of code collisions with other libraries. This precaution is necessary because verticapy uses commonly known function names like “average” and “median”, which can potentially lead to naming conflicts. The use of an alias ensures that the functions fromverticapyare used as intended without interfering with functions from other libraries.For this example, we will use the iris dataset.
import verticapy.datasets as vpd data = vpd.load_iris()
123SepalLengthCm123SepalWidthCm123PetalLengthCm123PetalWidthCmAbcSpecies1 4.6 3.6 1.0 0.2 Iris-setosa 2 4.7 3.2 1.3 0.2 Iris-setosa 3 4.7 3.2 1.6 0.2 Iris-setosa 4 4.8 3.0 1.4 0.1 Iris-setosa 5 4.8 3.1 1.6 0.2 Iris-setosa 6 4.8 3.4 1.9 0.2 Iris-setosa 7 4.9 3.0 1.4 0.2 Iris-setosa 8 4.9 3.1 1.5 0.1 Iris-setosa 9 4.9 3.1 1.5 0.1 Iris-setosa 10 4.9 3.1 1.5 0.1 Iris-setosa 11 5.0 2.3 3.3 1.0 Iris-versicolor 12 5.0 3.4 1.5 0.2 Iris-setosa 13 5.1 3.5 1.4 0.2 Iris-setosa 14 5.4 3.0 4.5 1.5 Iris-versicolor 15 5.4 3.4 1.5 0.4 Iris-setosa 16 5.4 3.9 1.3 0.4 Iris-setosa 17 5.5 2.4 3.7 1.0 Iris-versicolor 18 5.5 2.4 3.8 1.1 Iris-versicolor 19 5.6 2.7 4.2 1.3 Iris-versicolor 20 5.7 3.0 4.2 1.2 Iris-versicolor 21 5.7 4.4 1.5 0.4 Iris-setosa 22 5.8 2.8 5.1 2.4 Iris-virginica 23 5.9 3.2 4.8 1.8 Iris-versicolor 24 6.1 3.0 4.6 1.4 Iris-versicolor 25 6.1 3.0 4.9 1.8 Iris-virginica 26 6.3 2.5 4.9 1.5 Iris-versicolor 27 6.3 3.3 4.7 1.6 Iris-versicolor 28 6.3 3.3 6.0 2.5 Iris-virginica 29 6.4 2.9 4.3 1.3 Iris-versicolor 30 6.5 3.0 5.5 1.8 Iris-virginica 31 6.5 3.0 5.8 2.2 Iris-virginica 32 6.7 3.0 5.0 1.7 Iris-versicolor 33 6.8 2.8 4.8 1.4 Iris-versicolor 34 6.8 3.2 5.9 2.3 Iris-virginica 35 7.0 3.2 4.7 1.4 Iris-versicolor 36 7.1 3.0 5.9 2.1 Iris-virginica 37 7.7 3.8 6.7 2.2 Iris-virginica 38 4.4 2.9 1.4 0.2 Iris-setosa 39 4.5 2.3 1.3 0.3 Iris-setosa 40 4.8 3.4 1.6 0.2 Iris-setosa 41 5.0 2.0 3.5 1.0 Iris-versicolor 42 5.1 3.3 1.7 0.5 Iris-setosa 43 5.1 3.4 1.5 0.2 Iris-setosa 44 5.2 2.7 3.9 1.4 Iris-versicolor 45 5.2 3.5 1.5 0.2 Iris-setosa 46 5.2 4.1 1.5 0.1 Iris-setosa 47 5.4 3.9 1.7 0.4 Iris-setosa 48 5.5 3.5 1.3 0.2 Iris-setosa 49 5.6 3.0 4.1 1.3 Iris-versicolor 50 5.8 2.7 3.9 1.2 Iris-versicolor 51 5.8 2.7 5.1 1.9 Iris-virginica 52 5.8 2.7 5.1 1.9 Iris-virginica 53 5.9 3.0 4.2 1.5 Iris-versicolor 54 5.9 3.0 5.1 1.8 Iris-virginica 55 6.0 2.7 5.1 1.6 Iris-versicolor 56 6.0 2.9 4.5 1.5 Iris-versicolor 57 6.1 2.8 4.7 1.2 Iris-versicolor 58 6.2 2.8 4.8 1.8 Iris-virginica 59 6.2 2.9 4.3 1.3 Iris-versicolor 60 6.3 2.3 4.4 1.3 Iris-versicolor 61 6.3 2.7 4.9 1.8 Iris-virginica 62 6.4 3.2 5.3 2.3 Iris-virginica 63 6.5 2.8 4.6 1.5 Iris-versicolor 64 6.5 3.0 5.2 2.0 Iris-virginica 65 6.5 3.2 5.1 2.0 Iris-virginica 66 6.6 2.9 4.6 1.3 Iris-versicolor 67 6.6 3.0 4.4 1.4 Iris-versicolor 68 6.7 3.1 4.4 1.4 Iris-versicolor 69 6.7 3.1 4.7 1.5 Iris-versicolor 70 6.9 3.1 4.9 1.5 Iris-versicolor 71 6.9 3.1 5.4 2.1 Iris-virginica 72 6.9 3.2 5.7 2.3 Iris-virginica 73 7.2 3.0 5.8 1.6 Iris-virginica 74 7.2 3.2 6.0 1.8 Iris-virginica 75 7.3 2.9 6.3 1.8 Iris-virginica 76 7.7 2.6 6.9 2.3 Iris-virginica 77 3.3 4.5 5.6 7.8 Iris-setosa 78 3.3 4.5 5.6 7.8 Iris-setosa 79 3.3 4.5 5.6 7.8 Iris-setosa 80 3.3 4.5 5.6 7.8 Iris-setosa 81 3.3 4.5 5.6 7.8 Iris-setosa 82 3.3 4.5 5.6 7.8 Iris-setosa 83 3.3 4.5 5.6 7.8 Iris-setosa 84 3.3 4.5 5.6 7.8 Iris-setosa 85 3.3 4.5 5.6 7.8 Iris-setosa 86 3.3 4.5 5.6 7.8 Iris-setosa 87 3.3 4.5 5.6 7.8 Iris-setosa 88 3.3 4.5 5.6 7.8 Iris-setosa 89 3.3 4.5 5.6 7.8 Iris-setosa 90 3.3 4.5 5.6 7.8 Iris-setosa 91 3.3 4.5 5.6 7.8 Iris-setosa 92 3.3 4.5 5.6 7.8 Iris-setosa 93 3.3 4.5 5.6 7.8 Iris-setosa 94 3.3 4.5 5.6 7.8 Iris-setosa 95 3.3 4.5 5.6 7.8 Iris-setosa 96 3.3 4.5 5.6 7.8 Iris-setosa 97 3.3 4.5 5.6 7.8 Iris-setosa 98 3.3 4.5 5.6 7.8 Iris-setosa 99 3.3 4.5 5.6 7.8 Iris-setosa 100 3.3 4.5 5.6 7.8 Iris-setosa Rows: 1-100 | Columns: 5Note
VerticaPy offers a wide range of sample datasets that are ideal for training and testing purposes. You can explore the full list of available datasets in the Datasets, which provides detailed information on each dataset and how to use them effectively. These datasets are invaluable resources for honing your data analysis and machine learning skills within the VerticaPy environment.
Data Exploration¶
Through a quick scatter plot, we can observe that the data has three main clusters.
data.scatter( columns = ["PetalLengthCm", "SepalLengthCm"], by = "Species", )
Elbow Curve¶
Let’s compute the optimal
kfor ourKMeansalgorithm and check if it aligns with the three clusters we observed earlier.To achieve this, let’s create the Elbow curve.
from verticapy.machine_learning.model_selection import elbow elbow( input_relation = data, X = data.get_columns(exclude_columns= "Species"), # All columns except Species n_cluster = (1, 100), init = "kmeanspp", )
Note
You can experiment with the Elbow score to determine the optimal number of clusters. The score is based on the ratio of Between -Cluster Sum of Squares to Total Sum of Squares, providing a way to assess the clustering accuracy. A score of 1 indicates a perfect clustering.
Note
It’s evident from the Elbow curve that
k=3is a suitable choice, indicating the optimal number of clusters for the KMeans algorithm.See also
best_k(): Finds theKMeans/KPrototypeskbased on a score.