Loading...

Encoding

Encoding features is a very important part of the data science life cycle. In data science, generality is important and having too many categories can compromise that and lead to incorrect results. In addition, some algorithmic optimizations are linear and prefer categorized information, and some can’t process non-numerical features.

There are many encoding techniques:

  • User-Defined Encoding: The most flexible encoding. The user can choose how to encode the different categories.

  • Label Encoding: Each category is converted to an integer using a bijection to [0;n-1] where n is the feature number of unique values.

  • One-hot Encoding: This technique creates dummies (values in {0,1}) of each category. The categories are then separated into n features.

  • Mean Encoding: This technique uses the frequencies of each category for a specific response column.

  • Discretization: This technique uses various mathematical technique to encode continuous features into categories.

To demonstrate encoding data in VerticaPy, we’ll use the well-known titanic dataset.

from verticapy.datasets import load_titanic

titanic = load_titanic()
titanic.head(100)
123
pclass
Int
100%
...
123
body
Int
9%
Abc
Varchar(100)
57%
11...22
21...[null]
31...[null]
41...[null]
51...[null]
61...[null]
71...[null]
81...62
91...[null]
101...[null]
111...[null]
121...[null]
131...110
141...[null]
151...[null]
161...38
171...[null]
181...126
191...292
201...175
211...[null]
221...122
231...166
241...[null]
251...207
261...[null]
271...232
281...[null]
291...[null]
301...[null]
311...[null]
321...[null]
331...46
341...169
351...[null]
361...[null]
371...[null]
381...[null]
391...[null]
401...[null]
411...[null]
421...[null]
431...[null]
441...[null]
451...[null]
461...[null]
471...[null]
481...[null]
491...[null]
501...[null]
511...[null]
521...[null]
531...[null]
541...[null]
551...[null]
561...[null]
571...[null]
581...[null]
591...[null]
601...[null]
611...[null]
621...[null]
631...[null]
641...[null]
651...[null]
661...[null]
671...[null]
681...[null]
691...[null]
701...[null]
711...[null]
721...[null]
731...[null]
742...[null]
752...[null]
762...[null]
772...[null]
782...[null]
792...[null]
802...[null]
812...[null]
822...[null]
832...236
842...[null]
852...[null]
862...[null]
872...155
882...[null]
892...75
902...35
912...[null]
922...[null]
932...[null]
942...[null]
952...[null]
962...[null]
972...[null]
982...165
992...[null]
1002...[null]

Let’s look at the age of the passengers.

titanic["age"].hist()

By using the discretize() method, we can discretize the data using equal-width binning.

titanic["age"].discretize(method = "same_width", h = 10)
titanic["age"].bar(max_cardinality = 10)

We can also discretize the data using frequency bins.

titanic = load_titanic()
titanic["age"].discretize(method = "same_freq", nbins = 5)
titanic["age"].bar(max_cardinality = 5)

Computing categories using a response column can also be a good solution.

titanic = load_titanic()
titanic["age"].discretize(method = "smart", response = "survived", nbins = 6)
titanic["age"].bar(method = "avg", of = "survived")

We can view the available techniques in the discretize() method with the help() method.

help(titanic["age"].discretize)
Help on function discretize in module verticapy.core.vdataframe._encoding:

discretize(method: Literal['auto', 'smart', 'same_width', 'same_freq', 'topk'] = 'auto', h: Annotated[Union[int, float, decimal.Decimal], 'Python Numbers'] = 0, nbins: int = -1, k: int = 6, new_category: str = 'Others', RFmodel_params: Optional[dict] = None, response: Optional[str] = None, return_enum_trans: bool = False) -> 'vDataFrame'

Discretizes the vDataColumn using the input method.

Parameters
----------
method: str, optional
    The method used to discretize the vDataColumn:

    - auto:
        Uses method 'same_width' for numerical
        vDataColumns, casts the other types to varchar.
    - same_freq:
        Computes bins  with the same number of elements.
    - same_width:
        Computes regular width bins.
    - smart:
        Uses  the Random  Forest on a  response
        column  to   find   the  most  relevant
        interval to use for the discretization.
    - topk:
        Keeps the topk most frequent categories
        and  merge the  other  into one  unique
        category.
h: PythonNumber, optional
    The  interval  size  used  to  convert  the vDataColumn.
    If this parameter is equal to 0, an optimised interval is
    computed.
nbins: int, optional
    Number of bins  used for the discretization  (must be > 1)
k: int, optional
    The integer k of the 'topk' method.
new_category: str, optional
    The  name of the  merging  category when using the  'topk'
    method.
RFmodel_params: dict, optional
    Dictionary  of the  Random Forest  model  parameters used  to
    compute the best splits when 'method' is set to 'smart'.
    A RF Regressor is  trained if  the response is numerical
    (except ints and bools), a RF Classifier otherwise.
    Example: Write {"n_estimators": 20, "max_depth": 10} to train
    a Random Forest with 20 trees and a maximum depth of 10.
response: str, optional
    Response vDataColumn when method is set to 'smart'.
return_enum_trans: bool, optional
    Returns  the transformation instead of the vDataFrame parent,
    and does not apply the transformation. This parameter is
    useful for testing the look of the final transformation.

Returns
-------
vDataFrame
    self._parent

To encode a categorical feature, we can use label encoding. For example, the column sex has two categories (male and female) that we can represent with 0 and 1, respectively.

titanic["sex"].label_encode()
titanic["sex"].head(100)
123
sex
Integer
11
21
31
41
51
61
71
81
90
101
111
121
131
141
151
161
171
181
191
201
211
221
231
241
251
261
271
281
291
301
311
320
331
341
351
360
370
380
390
400
410
421
430
441
450
460
470
480
490
500
510
520
531
540
551
560
570
580
590
601
610
620
630
641
650
660
671
680
691
700
710
720
731
741
751
761
770
781
791
801
811
821
831
841
851
861
871
881
891
901
911
921
931
940
951
961
971
981
991
1001

When a feature has few categories, the most suitable choice is the one-hot encoding. Label encoding converts a categorical feature to numerical without retaining its mathematical relationships. Let’s use a one-hot encoding on the embarked column.

titanic["embarked"].one_hot_encode()
titanic.select(["embarked", "embarked_C", "embarked_Q"])
Abc
embarked
Varchar(20)
99%
...
123
embarked_C
Integer
100%
123
embarked_Q
Integer
100%
1C...10
2S...00
3S...00
4S...00
5C...10
6C...10
7S...00
8C...10
9C...10
10S...00
11C...10
12S...00
13S...00
14S...00
15S...00
16S...00
17C...10
18S...00
19C...10
20S...00

One-hot encoding can be expensive if the column in question has a large number of categories. In that case, we should use mean encoding. Mean encoding replaces each category of a variable with its corresponding average over a partition by a response column. This makes it an efficient way to encode the data, but be careful of over-fitting.

Let’s use a mean encoding on the home.dest variable.

titanic["home.dest"].mean_encode("survived")
titanic.head(100)
123
pclass
Int
100%
...
123
embarked_C
Bool
100%
123
embarked_Q
Bool
100%
11...01
22...00
33...00
41...00
52...00
63...00
72...00
83...00
91...00
103...00
112...00
122...00
133...00
141...00
153...00
163...01
172...00
181...10
191...00
201...00
211...00
221...00
231...00
242...00
251...10
261...10
271...10
281...10
292...00
302...00
312...00
322...00
333...00
342...10
352...10
362...00
371...00
381...00
391...10
401...10
411...00
421...00
431...00
441...10
451...00
461...00
471...00
481...00
491...10
501...10
511...00
523...01
532...10
541...00
551...10
561...00
571...10
581...00
591...10
601...10
611...00
621...00
631...00
641...10
651...10
663...01
671...00
681...10
691...00
701...00
711...10
721...10
731...10
741...10
751...00
761...10
771...00
781...10
791...10
801...00
811...10
822...10
832...10
841...10
851...10
861...10
871...00
881...10
891...10
901...00
911...10
921...00
931...00
941...10
953...00
963...00
973...00
982...00
992...00
1001...00

VerticaPy offers many encoding techniques. For example, the case_when() and decode() methods allow the user to use a customized encoding on a column. The discretize() method allows you to reduce the number of categories in a column. It’s important to get familiar with all the techniques available so you can make informed decisions about which to use for a given dataset.