Loading...

Introduction to Machine Learning

One of the last stages of the data science life cycle is the Data Modeling. Machine learning algorithms are a set of statistical techniques that build mathematical models from training data. These algorithms come in two types:
  • Supervised: these algorithms are used when we want to predict a response column.

  • Unsupervised: these algorithms are used when we want to detect anomalies or when we want to segment the data. No response column is needed.

Supervised Learning

Supervised Learning techniques map an input to an output based on some example dataset. This type of learning consists of two main types:
  • Regression: The Response is numerical (Linear Regression, SVM Regression, RF Regression…).

  • Classification: The Response is categorical (Gradient Boosting, Naive Bayes, Logistic Regression…).

For example, predicting the total charges of a Telco customer using their tenure would be a type of regression. The following code is drawing a linear regression using the TotalCharges as a function of the tenure in the telco churn dataset.

import verticapy as vp

churn = vp.read_csv("churn.csv")

from verticapy.machine_learning.vertica import LinearRegression

model = LinearRegression()
model.fit(churn, ["tenure"], "TotalCharges")
model.plot()

In contrast, when we have to predict a categorical column, we’re dealing with classification.

In the following example, we use a Linear Support Vector Classification (SVC) to predict the species of a flower based on its petal and sepal lengths.

from verticapy.datasets import load_iris

iris = load_iris()
iris.one_hot_encode()
123
SepalLengthCm
Numeric(5,2)
100%
...
123
Species_Iris-setosa
Bool
100%
123
Species_Iris-versicolor
Bool
100%
14.6...10
24.7...10
34.7...10
44.8...10
54.8...10
64.8...10
74.9...10
84.9...10
94.9...10
104.9...10
115.0...01
125.0...10
135.1...10
145.4...01
155.4...10
165.4...10
175.5...01
185.5...01
195.6...01
205.7...01
from verticapy.machine_learning.vertica import LinearSVC

model = LinearSVC()
model.fit(iris, ["PetalLengthCm", "SepalLengthCm"], "Species_Iris-setosa")
model.plot()

When we have more than two categories, we use the expression Multiclass Classification instead of Classification.

Unsupervised Learning

These algorithms are to used to segment the data (KMeans, DBSCAN, etc.) or to detect anomalies (LocalOutlierFactor, Z-Score Techniques…). In particular, they’re useful for finding patterns in data without labels. For example, let’s use a KMeans algorithm to create different clusters on the Iris dataset. Each cluster will represent a flower’s species.

from verticapy.machine_learning.vertica import KMeans

model = KMeans(n_cluster = 3)
model.fit(iris, ["PetalLengthCm", "SepalLengthCm"])
model.plot()

In this section, we went over a few of the many ML algorithms available in VerticaPy.

In the next lesson, we’ll go over Time Series