ONE_HOT_ENCODER_FIT

Generates a sorted list of each of the category levels for each feature that will be encoded and stores the model.

Syntax

ONE_HOT_ENCODER_FIT ( 'model‑name', 'input‑relation','input‑columns' 
                  [ USING PARAMETERS [exclude_columns='excluded‑columns']
                                     [, output_view='output‑view']
                                     [, extra_levels='category‑levels'] ] )

Arguments

model‑name

Identifies the model to create, where model‑name conforms to conventions described in Identifiers. It must also be unique among all names of sequences, tables, projections, views, and models within the same schema.

input‑relation

The table or view that contains the data for one hot encoding. If the input relation is defined in Hive, use SYNC_WITH_HCATALOG_SCHEMA to sync the hcatalog schema, and then run the machine learning function.

input‑columns

Comma-separated list of columns to use from the input table/view, or asterisk (*) to select all columns.

Parameter Settings

Parameter name Set to…
exclude_columns

Comma-separated list of column names from input‑columns to exclude from processing.

output_view

The name of the view that stores the input relation and the one hot encodings. Columns are returned in the order they appear in the input relation, with the one-hot encoded columns appended after the original columns.

extra_levels

Additional levels in each category that are not present in the input relation. This parameter should be passed as a JSON string with category names as keys and lists of extra levels in each category as values.

Note: Quote hyper parameter names and string values according to the JSON standard.

Privileges

Non-superusers:

Examples

=> SELECT ONE_HOT_ENCODER_FIT ('one_hot_encoder_model','mtcars','*' 
USING PARAMETERS exclude_columns='mpg,disp,drat,wt,qsec,vs,am');
ONE_HOT_ENCODER_FIT
--------------------
Success
(1 row)

See Also