Encoding¶
Encoding features is a very important part of the data science life cycle. In data science, generality is important and having too many categories can compromise that and lead to incorrect results. In addition, some algorithmic optimizations are linear and prefer categorized information, and some can’t process non-numerical features.
There are many encoding techniques:
User-Defined Encoding: The most flexible encoding. The user can choose how to encode the different categories.
Label Encoding: Each category is converted to an integer using a bijection to [0;n-1] where n is the feature number of unique values.
One-hot Encoding: This technique creates dummies (values in {0,1}) of each category. The categories are then separated into n features.
Mean Encoding: This technique uses the frequencies of each category for a specific response column.
Discretization: This technique uses various mathematical technique to encode continuous features into categories.
To demonstrate encoding data in VerticaPy, we’ll use the well-known titanic dataset.
from verticapy.datasets import load_titanic
titanic = load_titanic()
titanic.head(100)
123 pclass100% | ... | 123 body9% | Abc 57% | |
| 1 | 1 | ... | 22 | |
| 2 | 1 | ... | [null] | |
| 3 | 1 | ... | [null] | |
| 4 | 1 | ... | [null] | |
| 5 | 1 | ... | [null] | |
| 6 | 1 | ... | [null] | |
| 7 | 1 | ... | [null] | |
| 8 | 1 | ... | 62 | |
| 9 | 1 | ... | [null] | |
| 10 | 1 | ... | [null] | |
| 11 | 1 | ... | [null] | |
| 12 | 1 | ... | [null] | |
| 13 | 1 | ... | 110 | |
| 14 | 1 | ... | [null] | |
| 15 | 1 | ... | [null] | |
| 16 | 1 | ... | 38 | |
| 17 | 1 | ... | [null] | |
| 18 | 1 | ... | 126 | |
| 19 | 1 | ... | 292 | |
| 20 | 1 | ... | 175 | |
| 21 | 1 | ... | [null] | |
| 22 | 1 | ... | 122 | |
| 23 | 1 | ... | 166 | |
| 24 | 1 | ... | [null] | |
| 25 | 1 | ... | 207 | |
| 26 | 1 | ... | [null] | |
| 27 | 1 | ... | 232 | |
| 28 | 1 | ... | [null] | |
| 29 | 1 | ... | [null] | |
| 30 | 1 | ... | [null] | |
| 31 | 1 | ... | [null] | |
| 32 | 1 | ... | [null] | |
| 33 | 1 | ... | 46 | |
| 34 | 1 | ... | 169 | |
| 35 | 1 | ... | [null] | |
| 36 | 1 | ... | [null] | |
| 37 | 1 | ... | [null] | |
| 38 | 1 | ... | [null] | |
| 39 | 1 | ... | [null] | |
| 40 | 1 | ... | [null] | |
| 41 | 1 | ... | [null] | |
| 42 | 1 | ... | [null] | |
| 43 | 1 | ... | [null] | |
| 44 | 1 | ... | [null] | |
| 45 | 1 | ... | [null] | |
| 46 | 1 | ... | [null] | |
| 47 | 1 | ... | [null] | |
| 48 | 1 | ... | [null] | |
| 49 | 1 | ... | [null] | |
| 50 | 1 | ... | [null] | |
| 51 | 1 | ... | [null] | |
| 52 | 1 | ... | [null] | |
| 53 | 1 | ... | [null] | |
| 54 | 1 | ... | [null] | |
| 55 | 1 | ... | [null] | |
| 56 | 1 | ... | [null] | |
| 57 | 1 | ... | [null] | |
| 58 | 1 | ... | [null] | |
| 59 | 1 | ... | [null] | |
| 60 | 1 | ... | [null] | |
| 61 | 1 | ... | [null] | |
| 62 | 1 | ... | [null] | |
| 63 | 1 | ... | [null] | |
| 64 | 1 | ... | [null] | |
| 65 | 1 | ... | [null] | |
| 66 | 1 | ... | [null] | |
| 67 | 1 | ... | [null] | |
| 68 | 1 | ... | [null] | |
| 69 | 1 | ... | [null] | |
| 70 | 1 | ... | [null] | |
| 71 | 1 | ... | [null] | |
| 72 | 1 | ... | [null] | |
| 73 | 1 | ... | [null] | |
| 74 | 2 | ... | [null] | |
| 75 | 2 | ... | [null] | |
| 76 | 2 | ... | [null] | |
| 77 | 2 | ... | [null] | |
| 78 | 2 | ... | [null] | |
| 79 | 2 | ... | [null] | |
| 80 | 2 | ... | [null] | |
| 81 | 2 | ... | [null] | |
| 82 | 2 | ... | [null] | |
| 83 | 2 | ... | 236 | |
| 84 | 2 | ... | [null] | |
| 85 | 2 | ... | [null] | |
| 86 | 2 | ... | [null] | |
| 87 | 2 | ... | 155 | |
| 88 | 2 | ... | [null] | |
| 89 | 2 | ... | 75 | |
| 90 | 2 | ... | 35 | |
| 91 | 2 | ... | [null] | |
| 92 | 2 | ... | [null] | |
| 93 | 2 | ... | [null] | |
| 94 | 2 | ... | [null] | |
| 95 | 2 | ... | [null] | |
| 96 | 2 | ... | [null] | |
| 97 | 2 | ... | [null] | |
| 98 | 2 | ... | 165 | |
| 99 | 2 | ... | [null] | |
| 100 | 2 | ... | [null] |
Let’s look at the age of the passengers.
titanic["age"].hist()
By using the discretize() method, we can discretize the data using equal-width binning.
titanic["age"].discretize(method = "same_width", h = 10)
titanic["age"].bar(max_cardinality = 10)
We can also discretize the data using frequency bins.
titanic = load_titanic()
titanic["age"].discretize(method = "same_freq", nbins = 5)
titanic["age"].bar(max_cardinality = 5)
Computing categories using a response column can also be a good solution.
titanic = load_titanic()
titanic["age"].discretize(method = "smart", response = "survived", nbins = 6)
titanic["age"].bar(method = "avg", of = "survived")
We can view the available techniques in the discretize() method with the help() method.
help(titanic["age"].discretize)
Help on function discretize in module verticapy.core.vdataframe._encoding:
discretize(method: Literal['auto', 'smart', 'same_width', 'same_freq', 'topk'] = 'auto', h: Annotated[Union[int, float, decimal.Decimal], 'Python Numbers'] = 0, nbins: int = -1, k: int = 6, new_category: str = 'Others', RFmodel_params: Optional[dict] = None, response: Optional[str] = None, return_enum_trans: bool = False) -> 'vDataFrame'
Discretizes the vDataColumn using the input method.
Parameters
----------
method: str, optional
The method used to discretize the vDataColumn:
- auto:
Uses method 'same_width' for numerical
vDataColumns, casts the other types to varchar.
- same_freq:
Computes bins with the same number of elements.
- same_width:
Computes regular width bins.
- smart:
Uses the Random Forest on a response
column to find the most relevant
interval to use for the discretization.
- topk:
Keeps the topk most frequent categories
and merge the other into one unique
category.
h: PythonNumber, optional
The interval size used to convert the vDataColumn.
If this parameter is equal to 0, an optimised interval is
computed.
nbins: int, optional
Number of bins used for the discretization (must be > 1)
k: int, optional
The integer k of the 'topk' method.
new_category: str, optional
The name of the merging category when using the 'topk'
method.
RFmodel_params: dict, optional
Dictionary of the Random Forest model parameters used to
compute the best splits when 'method' is set to 'smart'.
A RF Regressor is trained if the response is numerical
(except ints and bools), a RF Classifier otherwise.
Example: Write {"n_estimators": 20, "max_depth": 10} to train
a Random Forest with 20 trees and a maximum depth of 10.
response: str, optional
Response vDataColumn when method is set to 'smart'.
return_enum_trans: bool, optional
Returns the transformation instead of the vDataFrame parent,
and does not apply the transformation. This parameter is
useful for testing the look of the final transformation.
Returns
-------
vDataFrame
self._parent
To encode a categorical feature, we can use label encoding. For example, the column sex has two categories (male and female) that we can represent with 0 and 1, respectively.
titanic["sex"].label_encode()
titanic["sex"].head(100)
123 sex | |
| 1 | 1 |
| 2 | 1 |
| 3 | 1 |
| 4 | 1 |
| 5 | 1 |
| 6 | 1 |
| 7 | 1 |
| 8 | 1 |
| 9 | 0 |
| 10 | 1 |
| 11 | 1 |
| 12 | 1 |
| 13 | 1 |
| 14 | 1 |
| 15 | 1 |
| 16 | 1 |
| 17 | 1 |
| 18 | 1 |
| 19 | 1 |
| 20 | 1 |
| 21 | 1 |
| 22 | 1 |
| 23 | 1 |
| 24 | 1 |
| 25 | 1 |
| 26 | 1 |
| 27 | 1 |
| 28 | 1 |
| 29 | 1 |
| 30 | 1 |
| 31 | 1 |
| 32 | 0 |
| 33 | 1 |
| 34 | 1 |
| 35 | 1 |
| 36 | 0 |
| 37 | 0 |
| 38 | 0 |
| 39 | 0 |
| 40 | 0 |
| 41 | 0 |
| 42 | 1 |
| 43 | 0 |
| 44 | 1 |
| 45 | 0 |
| 46 | 0 |
| 47 | 0 |
| 48 | 0 |
| 49 | 0 |
| 50 | 0 |
| 51 | 0 |
| 52 | 0 |
| 53 | 1 |
| 54 | 0 |
| 55 | 1 |
| 56 | 0 |
| 57 | 0 |
| 58 | 0 |
| 59 | 0 |
| 60 | 1 |
| 61 | 0 |
| 62 | 0 |
| 63 | 0 |
| 64 | 1 |
| 65 | 0 |
| 66 | 0 |
| 67 | 1 |
| 68 | 0 |
| 69 | 1 |
| 70 | 0 |
| 71 | 0 |
| 72 | 0 |
| 73 | 1 |
| 74 | 1 |
| 75 | 1 |
| 76 | 1 |
| 77 | 0 |
| 78 | 1 |
| 79 | 1 |
| 80 | 1 |
| 81 | 1 |
| 82 | 1 |
| 83 | 1 |
| 84 | 1 |
| 85 | 1 |
| 86 | 1 |
| 87 | 1 |
| 88 | 1 |
| 89 | 1 |
| 90 | 1 |
| 91 | 1 |
| 92 | 1 |
| 93 | 1 |
| 94 | 0 |
| 95 | 1 |
| 96 | 1 |
| 97 | 1 |
| 98 | 1 |
| 99 | 1 |
| 100 | 1 |
When a feature has few categories, the most suitable choice is the one-hot encoding. Label encoding converts a categorical feature to numerical without retaining its mathematical relationships. Let’s use a one-hot encoding on the embarked column.
titanic["embarked"].one_hot_encode()
titanic.select(["embarked", "embarked_C", "embarked_Q"])
Abc embarked99% | ... | 123 embarked_C100% | 123 embarked_Q100% | |
| 1 | C | ... | 1 | 0 |
| 2 | S | ... | 0 | 0 |
| 3 | S | ... | 0 | 0 |
| 4 | S | ... | 0 | 0 |
| 5 | C | ... | 1 | 0 |
| 6 | C | ... | 1 | 0 |
| 7 | S | ... | 0 | 0 |
| 8 | C | ... | 1 | 0 |
| 9 | C | ... | 1 | 0 |
| 10 | S | ... | 0 | 0 |
| 11 | C | ... | 1 | 0 |
| 12 | S | ... | 0 | 0 |
| 13 | S | ... | 0 | 0 |
| 14 | S | ... | 0 | 0 |
| 15 | S | ... | 0 | 0 |
| 16 | S | ... | 0 | 0 |
| 17 | C | ... | 1 | 0 |
| 18 | S | ... | 0 | 0 |
| 19 | C | ... | 1 | 0 |
| 20 | S | ... | 0 | 0 |
One-hot encoding can be expensive if the column in question has a large number of categories. In that case, we should use mean encoding. Mean encoding replaces each category of a variable with its corresponding average over a partition by a response column. This makes it an efficient way to encode the data, but be careful of over-fitting.
Let’s use a mean encoding on the home.dest variable.
titanic["home.dest"].mean_encode("survived")
titanic.head(100)
123 pclass100% | ... | 123 embarked_C100% | 123 embarked_Q100% | |
| 1 | 1 | ... | 0 | 1 |
| 2 | 2 | ... | 0 | 0 |
| 3 | 3 | ... | 0 | 0 |
| 4 | 1 | ... | 0 | 0 |
| 5 | 2 | ... | 0 | 0 |
| 6 | 3 | ... | 0 | 0 |
| 7 | 2 | ... | 0 | 0 |
| 8 | 3 | ... | 0 | 0 |
| 9 | 1 | ... | 0 | 0 |
| 10 | 3 | ... | 0 | 0 |
| 11 | 2 | ... | 0 | 0 |
| 12 | 2 | ... | 0 | 0 |
| 13 | 3 | ... | 0 | 0 |
| 14 | 1 | ... | 0 | 0 |
| 15 | 3 | ... | 0 | 0 |
| 16 | 3 | ... | 0 | 1 |
| 17 | 2 | ... | 0 | 0 |
| 18 | 1 | ... | 1 | 0 |
| 19 | 1 | ... | 0 | 0 |
| 20 | 1 | ... | 0 | 0 |
| 21 | 1 | ... | 0 | 0 |
| 22 | 1 | ... | 0 | 0 |
| 23 | 1 | ... | 0 | 0 |
| 24 | 2 | ... | 0 | 0 |
| 25 | 1 | ... | 1 | 0 |
| 26 | 1 | ... | 1 | 0 |
| 27 | 1 | ... | 1 | 0 |
| 28 | 1 | ... | 1 | 0 |
| 29 | 2 | ... | 0 | 0 |
| 30 | 2 | ... | 0 | 0 |
| 31 | 2 | ... | 0 | 0 |
| 32 | 2 | ... | 0 | 0 |
| 33 | 3 | ... | 0 | 0 |
| 34 | 2 | ... | 1 | 0 |
| 35 | 2 | ... | 1 | 0 |
| 36 | 2 | ... | 0 | 0 |
| 37 | 1 | ... | 0 | 0 |
| 38 | 1 | ... | 0 | 0 |
| 39 | 1 | ... | 1 | 0 |
| 40 | 1 | ... | 1 | 0 |
| 41 | 1 | ... | 0 | 0 |
| 42 | 1 | ... | 0 | 0 |
| 43 | 1 | ... | 0 | 0 |
| 44 | 1 | ... | 1 | 0 |
| 45 | 1 | ... | 0 | 0 |
| 46 | 1 | ... | 0 | 0 |
| 47 | 1 | ... | 0 | 0 |
| 48 | 1 | ... | 0 | 0 |
| 49 | 1 | ... | 1 | 0 |
| 50 | 1 | ... | 1 | 0 |
| 51 | 1 | ... | 0 | 0 |
| 52 | 3 | ... | 0 | 1 |
| 53 | 2 | ... | 1 | 0 |
| 54 | 1 | ... | 0 | 0 |
| 55 | 1 | ... | 1 | 0 |
| 56 | 1 | ... | 0 | 0 |
| 57 | 1 | ... | 1 | 0 |
| 58 | 1 | ... | 0 | 0 |
| 59 | 1 | ... | 1 | 0 |
| 60 | 1 | ... | 1 | 0 |
| 61 | 1 | ... | 0 | 0 |
| 62 | 1 | ... | 0 | 0 |
| 63 | 1 | ... | 0 | 0 |
| 64 | 1 | ... | 1 | 0 |
| 65 | 1 | ... | 1 | 0 |
| 66 | 3 | ... | 0 | 1 |
| 67 | 1 | ... | 0 | 0 |
| 68 | 1 | ... | 1 | 0 |
| 69 | 1 | ... | 0 | 0 |
| 70 | 1 | ... | 0 | 0 |
| 71 | 1 | ... | 1 | 0 |
| 72 | 1 | ... | 1 | 0 |
| 73 | 1 | ... | 1 | 0 |
| 74 | 1 | ... | 1 | 0 |
| 75 | 1 | ... | 0 | 0 |
| 76 | 1 | ... | 1 | 0 |
| 77 | 1 | ... | 0 | 0 |
| 78 | 1 | ... | 1 | 0 |
| 79 | 1 | ... | 1 | 0 |
| 80 | 1 | ... | 0 | 0 |
| 81 | 1 | ... | 1 | 0 |
| 82 | 2 | ... | 1 | 0 |
| 83 | 2 | ... | 1 | 0 |
| 84 | 1 | ... | 1 | 0 |
| 85 | 1 | ... | 1 | 0 |
| 86 | 1 | ... | 1 | 0 |
| 87 | 1 | ... | 0 | 0 |
| 88 | 1 | ... | 1 | 0 |
| 89 | 1 | ... | 1 | 0 |
| 90 | 1 | ... | 0 | 0 |
| 91 | 1 | ... | 1 | 0 |
| 92 | 1 | ... | 0 | 0 |
| 93 | 1 | ... | 0 | 0 |
| 94 | 1 | ... | 1 | 0 |
| 95 | 3 | ... | 0 | 0 |
| 96 | 3 | ... | 0 | 0 |
| 97 | 3 | ... | 0 | 0 |
| 98 | 2 | ... | 0 | 0 |
| 99 | 2 | ... | 0 | 0 |
| 100 | 1 | ... | 0 | 0 |
VerticaPy offers many encoding techniques. For example, the case_when() and decode() methods allow the user to use a customized encoding on a column. The discretize() method allows you to reduce the number of categories in a column. It’s important to get familiar with all the techniques available so you can make informed decisions about which to use for a given dataset.