Loading...

verticapy.machine_learning.vertica.tsa.ensemble.TimeSeriesByCategory

class verticapy.machine_learning.vertica.tsa.ensemble.TimeSeriesByCategory(name: str = None, overwrite_model: bool = False, base_model: TimeSeriesModelBase | None = None)

This model is built based on multiple base models. You should look at the source models to see entire examples.

Important

This is still Beta.

Parameters

name: str, optional

Name of the model. The model is stored in the database.

overwrite_model: bool, optional

If set to True, training a model with the same name as an existing model overwrites the existing model.

base_model: TimeSeriesModelBase

The user should provide a base model which will be used for each category. It could be - ARIMA - ARMA - AR - :py:class:`~verticapy.machine_learning.vertica.tsa.MA’

Attributes

Many attributes are created during the fitting phase.

distinct: list

This provides a sequential list of the categories used to build the different models.

ts: str

The column name for time stamp.

y: str

The column name used for building the model.

_is_already_stored: bool

This tells us whether a model is stored in the Vertica database.

_get_model_names: list

This returns the list of names of the models created.

Examples

The following examples provide a basic understanding of usage.

Initialization

For this example, we will use a subset of the amazon dataset.

import verticapy.datasets as vpd

amazon_full = vpd.load_amazon()
📅
date
Date
Abc
state
Varchar(32)
123
number
Integer
11998-01-01AMAPÁ0
21998-01-01AMAZONAS0
31998-01-01DISTRITO FEDERAL0
41998-01-01ESPÍRITO SANTO0
51998-01-01MARANHÃO0
61998-01-01PARANÁ0
71998-01-01PIAUÍ0
81998-01-01RORAIMA0
91998-01-01SERGIPE0
101998-01-01SÃO PAULO0
111998-02-01GOIÁS0
121998-02-01MATO GROSSO DO SUL0
131998-02-01MINAS GERAIS0
141998-02-01PARAÍBA0
151998-02-01SANTA CATARINA0
161998-02-01SÃO PAULO0
171998-03-01AMAPÁ0
181998-03-01BAHIA0
191998-03-01MATO GROSSO DO SUL0
201998-03-01PARÁ0
211998-03-01PERNAMBUCO0
221998-03-01RIO GRANDE DO SUL0
231998-04-01CEARÁ0
241998-04-01PARANÁ0
251998-04-01PARAÍBA0
261998-04-01PARÁ0
271998-05-01ALAGOAS0
281998-05-01AMAPÁ0
291998-05-01MARANHÃO0
301998-05-01MATO GROSSO DO SUL0
311998-05-01RIO GRANDE DO SUL0
321998-05-01TOCANTINS0
331998-06-01ESPÍRITO SANTO6
341998-06-01RIO DE JANEIRO3
351998-06-01RIO GRANDE DO NORTE1
361998-06-01SÃO PAULO451
371998-07-01ESPÍRITO SANTO37
381998-07-01MARANHÃO274
391998-07-01MATO GROSSO360
401998-07-01MATO GROSSO DO SUL3712
411998-07-01PARAÍBA0
421998-07-01PARÁ638
431998-07-01RONDÔNIA365
441998-07-01SÃO PAULO596
451998-08-01ALAGOAS1
461998-08-01BAHIA815
471998-08-01DISTRITO FEDERAL48
481998-08-01ESPÍRITO SANTO38
491998-08-01MARANHÃO1176
501998-08-01MATO GROSSO228
511998-08-01MINAS GERAIS875
521998-08-01PIAUÍ711
531998-08-01RIO GRANDE DO SUL9
541998-08-01RORAIMA0
551998-08-01SERGIPE0
561998-09-01AMAPÁ20
571998-09-01DISTRITO FEDERAL33
581998-09-01PIAUÍ1991
591998-09-01RORAIMA2
601998-10-01AMAZONAS83
611998-10-01GOIÁS1034
621998-10-01MATO GROSSO576
631998-10-01PARAÍBA179
641998-10-01PARÁ3665
651998-10-01PIAUÍ2586
661998-10-01SERGIPE0
671998-11-01ALAGOAS19
681998-11-01AMAPÁ131
691998-11-01CEARÁ575
701998-11-01DISTRITO FEDERAL0
711998-11-01MARANHÃO2237
721998-11-01RIO DE JANEIRO6
731998-11-01RIO GRANDE DO SUL28
741998-11-01SÃO PAULO488
751998-12-01BAHIA82
761998-12-01MARANHÃO1399
771998-12-01MATO GROSSO100
781998-12-01PARAÍBA51
791998-12-01PERNAMBUCO59
801998-12-01RIO DE JANEIRO1
811998-12-01RONDÔNIA33
821998-12-01TOCANTINS9
831999-01-01ALAGOAS58
841999-01-01GOIÁS14
851999-01-01MATO GROSSO239
861999-01-01MINAS GERAIS36
871999-01-01PARÁ87
881999-01-01PERNAMBUCO102
891999-01-01RONDÔNIA1
901999-01-01SÃO PAULO7
911999-01-01TOCANTINS36
921999-02-01ACRE0
931999-02-01CEARÁ16
941999-02-01MATO GROSSO69
951999-02-01MATO GROSSO DO SUL28
961999-02-01PERNAMBUCO13
971999-02-01RONDÔNIA1
981999-02-01SANTA CATARINA2
991999-02-01TOCANTINS1
1001999-03-01AMAPÁ2
Rows: 1-100 | Columns: 3

We can reduce the number of states for the sake of ease in this example:

amazon = amazon_full[(amazon_full["state"] == "PERNAMBUCO") | (amazon_full["state"] == "SERGIPE")]

Now we can setup a base model that will be created for each unique state inside the dataset. For this example, we use ARIMA.

from verticapy.machine_learning.vertica.tsa import ARIMA

base_model = ARIMA(order = (2, 1, 2))

Finally we can now initiate our multiple models in one go:

from verticapy.machine_learning.vertica.tsa.ensemble import TimeSeriesByCategory

model = TimeSeriesByCategory(base_model = base_model)

Model Fitting

We can now fit the model:

model.fit(amazon, ts = "date", y = "number", by = "state")
For category: PERNAMBUCO


============
coefficients
============
parameter| value  
---------+--------
  phi_1  | 0.45168
  phi_2  | 0.01325
 theta_1 |-0.66935
 theta_2 |-0.34679


==============
regularization
==============
none

===============
timeseries_name
===============
"number"

==============
timestamp_name
==============
date

==============
missing_method
==============
linear_interpolation

===========
call_string
===========
ARIMA('"public"."_verticapy_tmp_timeseriesbycategory_v_mldb_9f553e78979e11efa8720242ac120002__000_tsbc"', '"public"."_verticapy_tmp_view_v_mldb_9ff4e3f6979e11efa8720242ac120002_"', '"number"', 'date' USING PARAMETERS p=2, d=1, q=2, missing='linear_interpolation', init_method='Zero', epsilon=1e-06, max_iterations=100);

===============
Additional Info
===============
       Name       |   Value   
------------------+-----------
        p         |     2     
        d         |     1     
        q         |     2     
       mean       |  0.27731  
      lambda      |  1.00000  
mean_squared_error|13117.49869
rejected_row_count|     0     
accepted_row_count|    239    


For category: SERGIPE


============
coefficients
============
parameter| value  
---------+--------
  phi_1  |-0.03152
  phi_2  | 0.27608
 theta_1 |-0.25561
 theta_2 |-0.73844


==============
regularization
==============
none

===============
timeseries_name
===============
"number"

==============
timestamp_name
==============
date

==============
missing_method
==============
linear_interpolation

===========
call_string
===========
ARIMA('"public"."_verticapy_tmp_timeseriesbycategory_v_mldb_9f553e78979e11efa8720242ac120002__001_tsbc"', '"public"."_verticapy_tmp_view_v_mldb_a2d9f0a2979e11efa8720242ac120002_"', '"number"', 'date' USING PARAMETERS p=2, d=1, q=2, missing='linear_interpolation', init_method='Zero', epsilon=1e-06, max_iterations=100);

===============
Additional Info
===============
       Name       |  Value  
------------------+---------
        p         |    2    
        d         |    1    
        q         |    2    
       mean       | 0.00420 
      lambda      | 1.00000 
mean_squared_error|319.36890
rejected_row_count|    0    
accepted_row_count|   239   

Important

To train a model, you can directly use the vDataFrame or the name of the relation stored in the database. The test set is optional and is only used to compute the test metrics. In verticapy, we don’t work using X matrices and y vectors. Instead, we work directly with lists of predictors and the response name.

Plots

We can conveniently plot the predictions on a line plot to observe the efficacy of our model. We need to provide the idx which represents the model number.

model.plot(idx = 0, npredictions = 5)

Note

You can find out the name of the category by the distinct attribute. The sequential list of categories correspond to idx = 0, 1 .... model.distinct.

__init__(name: str = None, overwrite_model: bool = False, base_model: TimeSeriesModelBase | None = None) → None

Must be overridden in the child class

Methods

__init__([name, overwrite_model, base_model])

Must be overridden in the child class

contour([nbins, chart])

Draws the model's contour plot.

deploySQL([vdf, ts, y, start, npredictions, ...])

Returns the SQL code needed to deploy the model.

does_model_exists(name[, raise_error, ...])

Checks whether the model is stored in the Vertica database.

drop()

Drops the model from the Vertica database.

export_models(name, path[, kind])

Exports machine learning models.

features_importance([idx, show, chart])

Computes the input submodel's features importance.

fit(input_relation, ts, y, by[, ...])

Trains the model.

get_attributes([attr_name])

Returns the model attributes.

get_match_index(x, col_list[, str_check])

Returns the matching index.

get_params()

Returns the parameters of the model.

get_plotting_lib([class_name, chart, ...])

Returns the first available library (Plotly, Matplotlib, or Highcharts) to draw a specific graphic.

get_vertica_attributes([attr_name])

Returns the model Vertica attributes.

import_models(path[, schema, kind])

Imports machine learning models.

plot([idx, vdf, ts, y, start, npredictions, ...])

Draws the input submodel.

predict([vdf, ts, y, start, npredictions, ...])

Predicts using the input relation.

register(registered_name[, raise_error])

Registers the model and adds it to in-DB Model versioning environment with a status of 'under_review'.

regression_report([metrics, start, ...])

Computes a regression report using multiple metrics to evaluate the model (r2, mse, max error...).

report([metrics, start, npredictions, method])

Computes a regression report using multiple metrics to evaluate the model (r2, mse, max error...).

score([metric, start, npredictions, method])

Computes the model score.

set_params([parameters])

Sets the parameters of the model.

summarize()

Summarizes the model.

to_binary(path)

Exports the model to the Vertica Binary format.

to_pmml(path)

Exports the model to PMML.

to_python([return_proba, ...])

Returns the Python function needed for in-memory scoring without using built-in Vertica functions.

to_sql([X, return_proba, ...])

Returns the SQL code needed to deploy the model without using built-in Vertica functions.

to_tf(path)

Exports the model to the Frozen Graph format (TensorFlow).

Attributes