Loading...

verticapy.machine_learning.model_selection.elbow

verticapy.machine_learning.model_selection.elbow(input_relation: Annotated[str | vDataFrame, ''], X: Annotated[str | list[str], 'STRING representing one column or a list of columns'] | None = None, n_cluster: tuple | list = (1, 15), init: Literal['kmeanspp', 'random', None] = None, max_iter: int = 50, tol: float = 0.0001, use_kprototype: bool = False, gamma: float = 1.0, show: bool = True, chart: PlottingBase | TableSample | Axes | mFigure | Highchart | Highstock | Figure | None = None, **style_kwargs) → TableSample

Draws an Elbow curve.

Parameters

input_relation: SQLRelation

Relation used to train the model.

X: SQLColumns, optional

list of the predictor columns. If empty, all numerical columns are used.

n_cluster: tuple | list, optional

Tuple representing the number of clusters to start and end with. This can also be a customized list with various k values to test.

init: str | list, optional

The method used to find the initial cluster centers.

  • kmeanspp:

    Only available when use_kprototype = False Use the k-means++ method to initialize the centers.

  • random:

    Randomly subsamples the data to find initial centers.

Default value is kmeanspp if use_kprototype = False; otherwise, random.

max_iter: int, optional

The maximum number of iterations for the algorithm.

tol: float, optional

Determines whether the algorithm has converged. The algorithm is considered converged after no center has moved more than a distance of tol from the previous iteration.

use_kprototype: bool, optional

If set to True, the function uses the KPrototypes algorithm instead of KMeans. KPrototypes can handle categorical features.

gamma: float, optional

Only if use_kprototype = True. Weighting factor for categorical columns. It determines the relative importance of numerical and categorical attributes.

show: bool, optional

If set to True, the Plotting object is returned.

chart: PlottingObject, optional

The chart object to plot on.

**style_kwargs

Any optional parameter to pass to the Plotting functions.

Returns

TableSample

nb_clusters,total_within_cluster_ss,between_cluster_ss,total_ss, elbow_score

Examples

The following examples provide a basic understanding of usage. For more detailed examples, please refer to the: Elbow Curve page.

Load data for machine learning

We import verticapy:

import verticapy as vp

Hint

By assigning an alias to verticapy, we mitigate the risk of code collisions with other libraries. This precaution is necessary because verticapy uses commonly known function names like “average” and “median”, which can potentially lead to naming conflicts. The use of an alias ensures that the functions from verticapy are used as intended without interfering with functions from other libraries.

For this example, we will use the iris dataset.

import verticapy.datasets as vpd

data = vpd.load_iris()
123
SepalLengthCm
Numeric(7)
123
SepalWidthCm
Numeric(7)
123
PetalLengthCm
Numeric(7)
123
PetalWidthCm
Numeric(7)
Abc
Species
Varchar(30)
14.63.61.00.2Iris-setosa
24.73.21.30.2Iris-setosa
34.73.21.60.2Iris-setosa
44.83.01.40.1Iris-setosa
54.83.11.60.2Iris-setosa
64.83.41.90.2Iris-setosa
74.93.01.40.2Iris-setosa
84.93.11.50.1Iris-setosa
94.93.11.50.1Iris-setosa
104.93.11.50.1Iris-setosa
115.02.33.31.0Iris-versicolor
125.03.41.50.2Iris-setosa
135.13.51.40.2Iris-setosa
145.43.04.51.5Iris-versicolor
155.43.41.50.4Iris-setosa
165.43.91.30.4Iris-setosa
175.52.43.71.0Iris-versicolor
185.52.43.81.1Iris-versicolor
195.62.74.21.3Iris-versicolor
205.73.04.21.2Iris-versicolor
215.74.41.50.4Iris-setosa
225.82.85.12.4Iris-virginica
235.93.24.81.8Iris-versicolor
246.13.04.61.4Iris-versicolor
256.13.04.91.8Iris-virginica
266.32.54.91.5Iris-versicolor
276.33.34.71.6Iris-versicolor
286.33.36.02.5Iris-virginica
296.42.94.31.3Iris-versicolor
306.53.05.51.8Iris-virginica
316.53.05.82.2Iris-virginica
326.73.05.01.7Iris-versicolor
336.82.84.81.4Iris-versicolor
346.83.25.92.3Iris-virginica
357.03.24.71.4Iris-versicolor
367.13.05.92.1Iris-virginica
377.73.86.72.2Iris-virginica
384.42.91.40.2Iris-setosa
394.52.31.30.3Iris-setosa
404.83.41.60.2Iris-setosa
415.02.03.51.0Iris-versicolor
425.13.31.70.5Iris-setosa
435.13.41.50.2Iris-setosa
445.22.73.91.4Iris-versicolor
455.23.51.50.2Iris-setosa
465.24.11.50.1Iris-setosa
475.43.91.70.4Iris-setosa
485.53.51.30.2Iris-setosa
495.63.04.11.3Iris-versicolor
505.82.73.91.2Iris-versicolor
515.82.75.11.9Iris-virginica
525.82.75.11.9Iris-virginica
535.93.04.21.5Iris-versicolor
545.93.05.11.8Iris-virginica
556.02.75.11.6Iris-versicolor
566.02.94.51.5Iris-versicolor
576.12.84.71.2Iris-versicolor
586.22.84.81.8Iris-virginica
596.22.94.31.3Iris-versicolor
606.32.34.41.3Iris-versicolor
616.32.74.91.8Iris-virginica
626.43.25.32.3Iris-virginica
636.52.84.61.5Iris-versicolor
646.53.05.22.0Iris-virginica
656.53.25.12.0Iris-virginica
666.62.94.61.3Iris-versicolor
676.63.04.41.4Iris-versicolor
686.73.14.41.4Iris-versicolor
696.73.14.71.5Iris-versicolor
706.93.14.91.5Iris-versicolor
716.93.15.42.1Iris-virginica
726.93.25.72.3Iris-virginica
737.23.05.81.6Iris-virginica
747.23.26.01.8Iris-virginica
757.32.96.31.8Iris-virginica
767.72.66.92.3Iris-virginica
773.34.55.67.8Iris-setosa
783.34.55.67.8Iris-setosa
793.34.55.67.8Iris-setosa
803.34.55.67.8Iris-setosa
813.34.55.67.8Iris-setosa
823.34.55.67.8Iris-setosa
833.34.55.67.8Iris-setosa
843.34.55.67.8Iris-setosa
853.34.55.67.8Iris-setosa
863.34.55.67.8Iris-setosa
873.34.55.67.8Iris-setosa
883.34.55.67.8Iris-setosa
893.34.55.67.8Iris-setosa
903.34.55.67.8Iris-setosa
913.34.55.67.8Iris-setosa
923.34.55.67.8Iris-setosa
933.34.55.67.8Iris-setosa
943.34.55.67.8Iris-setosa
953.34.55.67.8Iris-setosa
963.34.55.67.8Iris-setosa
973.34.55.67.8Iris-setosa
983.34.55.67.8Iris-setosa
993.34.55.67.8Iris-setosa
1003.34.55.67.8Iris-setosa
Rows: 1-100 | Columns: 5

Note

VerticaPy offers a wide range of sample datasets that are ideal for training and testing purposes. You can explore the full list of available datasets in the Datasets, which provides detailed information on each dataset and how to use them effectively. These datasets are invaluable resources for honing your data analysis and machine learning skills within the VerticaPy environment.

Data Exploration

Through a quick scatter plot, we can observe that the data has three main clusters.

data.scatter(
    columns = ["PetalLengthCm", "SepalLengthCm"],
    by = "Species",
)

Elbow Curve

Let’s compute the optimal k for our KMeans algorithm and check if it aligns with the three clusters we observed earlier.

To achieve this, let’s create the Elbow curve.

from verticapy.machine_learning.model_selection import elbow

elbow(
    input_relation = data,
    X = data.get_columns(exclude_columns= "Species"), # All columns except Species
    n_cluster = (1, 100),
    init = "kmeanspp",
)

Note

You can experiment with the Elbow score to determine the optimal number of clusters. The score is based on the ratio of Between -Cluster Sum of Squares to Total Sum of Squares, providing a way to assess the clustering accuracy. A score of 1 indicates a perfect clustering.

Note

It’s evident from the Elbow curve that k=3 is a suitable choice, indicating the optimal number of clusters for the KMeans algorithm.

See also

best_k() : Finds the KMeans / KPrototypes k based on a score.