verticapy.machine_learning.vertica.tsa.ensemble.TimeSeriesByCategory¶
- class verticapy.machine_learning.vertica.tsa.ensemble.TimeSeriesByCategory(name: str = None, overwrite_model: bool = False, base_model: TimeSeriesModelBase | None = None)¶
This model is built based on multiple base models. You should look at the source models to see entire examples.
Important
This is still Beta.
Parameters¶
- name: str, optional
Name of the model. The model is stored in the database.
- overwrite_model: bool, optional
If set to
True, training a model with the same name as an existing model overwrites the existing model.- base_model: TimeSeriesModelBase
The user should provide a base model which will be used for each category. It could be -
ARIMA-ARMA-AR- :py:class:`~verticapy.machine_learning.vertica.tsa.MA’
Attributes¶
Many attributes are created during the fitting phase.
- distinct: list
This provides a sequential list of the categories used to build the different models.
- ts: str
The column name for time stamp.
- y: str
The column name used for building the model.
- _is_already_stored: bool
This tells us whether a model is stored in the Vertica database.
- _get_model_names: list
This returns the list of names of the models created.
Examples¶
The following examples provide a basic understanding of usage.
Initialization¶
For this example, we will use a subset of the amazon dataset.
import verticapy.datasets as vpd amazon_full = vpd.load_amazon()
📅dateAbcstate123number1 1998-01-01 AMAPÁ 0 2 1998-01-01 AMAZONAS 0 3 1998-01-01 DISTRITO FEDERAL 0 4 1998-01-01 ESPÍRITO SANTO 0 5 1998-01-01 MARANHÃO 0 6 1998-01-01 PARANÁ 0 7 1998-01-01 PIAUÍ 0 8 1998-01-01 RORAIMA 0 9 1998-01-01 SERGIPE 0 10 1998-01-01 SÃO PAULO 0 11 1998-02-01 GOIÁS 0 12 1998-02-01 MATO GROSSO DO SUL 0 13 1998-02-01 MINAS GERAIS 0 14 1998-02-01 PARAÍBA 0 15 1998-02-01 SANTA CATARINA 0 16 1998-02-01 SÃO PAULO 0 17 1998-03-01 AMAPÁ 0 18 1998-03-01 BAHIA 0 19 1998-03-01 MATO GROSSO DO SUL 0 20 1998-03-01 PARÁ 0 21 1998-03-01 PERNAMBUCO 0 22 1998-03-01 RIO GRANDE DO SUL 0 23 1998-04-01 CEARÁ 0 24 1998-04-01 PARANÁ 0 25 1998-04-01 PARAÍBA 0 26 1998-04-01 PARÁ 0 27 1998-05-01 ALAGOAS 0 28 1998-05-01 AMAPÁ 0 29 1998-05-01 MARANHÃO 0 30 1998-05-01 MATO GROSSO DO SUL 0 31 1998-05-01 RIO GRANDE DO SUL 0 32 1998-05-01 TOCANTINS 0 33 1998-06-01 ESPÍRITO SANTO 6 34 1998-06-01 RIO DE JANEIRO 3 35 1998-06-01 RIO GRANDE DO NORTE 1 36 1998-06-01 SÃO PAULO 451 37 1998-07-01 ESPÍRITO SANTO 37 38 1998-07-01 MARANHÃO 274 39 1998-07-01 MATO GROSSO 360 40 1998-07-01 MATO GROSSO DO SUL 3712 41 1998-07-01 PARAÍBA 0 42 1998-07-01 PARÁ 638 43 1998-07-01 RONDÔNIA 365 44 1998-07-01 SÃO PAULO 596 45 1998-08-01 ALAGOAS 1 46 1998-08-01 BAHIA 815 47 1998-08-01 DISTRITO FEDERAL 48 48 1998-08-01 ESPÍRITO SANTO 38 49 1998-08-01 MARANHÃO 1176 50 1998-08-01 MATO GROSSO 228 51 1998-08-01 MINAS GERAIS 875 52 1998-08-01 PIAUÍ 711 53 1998-08-01 RIO GRANDE DO SUL 9 54 1998-08-01 RORAIMA 0 55 1998-08-01 SERGIPE 0 56 1998-09-01 AMAPÁ 20 57 1998-09-01 DISTRITO FEDERAL 33 58 1998-09-01 PIAUÍ 1991 59 1998-09-01 RORAIMA 2 60 1998-10-01 AMAZONAS 83 61 1998-10-01 GOIÁS 1034 62 1998-10-01 MATO GROSSO 576 63 1998-10-01 PARAÍBA 179 64 1998-10-01 PARÁ 3665 65 1998-10-01 PIAUÍ 2586 66 1998-10-01 SERGIPE 0 67 1998-11-01 ALAGOAS 19 68 1998-11-01 AMAPÁ 131 69 1998-11-01 CEARÁ 575 70 1998-11-01 DISTRITO FEDERAL 0 71 1998-11-01 MARANHÃO 2237 72 1998-11-01 RIO DE JANEIRO 6 73 1998-11-01 RIO GRANDE DO SUL 28 74 1998-11-01 SÃO PAULO 488 75 1998-12-01 BAHIA 82 76 1998-12-01 MARANHÃO 1399 77 1998-12-01 MATO GROSSO 100 78 1998-12-01 PARAÍBA 51 79 1998-12-01 PERNAMBUCO 59 80 1998-12-01 RIO DE JANEIRO 1 81 1998-12-01 RONDÔNIA 33 82 1998-12-01 TOCANTINS 9 83 1999-01-01 ALAGOAS 58 84 1999-01-01 GOIÁS 14 85 1999-01-01 MATO GROSSO 239 86 1999-01-01 MINAS GERAIS 36 87 1999-01-01 PARÁ 87 88 1999-01-01 PERNAMBUCO 102 89 1999-01-01 RONDÔNIA 1 90 1999-01-01 SÃO PAULO 7 91 1999-01-01 TOCANTINS 36 92 1999-02-01 ACRE 0 93 1999-02-01 CEARÁ 16 94 1999-02-01 MATO GROSSO 69 95 1999-02-01 MATO GROSSO DO SUL 28 96 1999-02-01 PERNAMBUCO 13 97 1999-02-01 RONDÔNIA 1 98 1999-02-01 SANTA CATARINA 2 99 1999-02-01 TOCANTINS 1 100 1999-03-01 AMAPÁ 2 Rows: 1-100 | Columns: 3We can reduce the number of states for the sake of ease in this example:
amazon = amazon_full[(amazon_full["state"] == "PERNAMBUCO") | (amazon_full["state"] == "SERGIPE")]
Now we can setup a base model that will be created for each unique state inside the dataset. For this example, we use ARIMA.
from verticapy.machine_learning.vertica.tsa import ARIMA base_model = ARIMA(order = (2, 1, 2))
Finally we can now initiate our multiple models in one go:
from verticapy.machine_learning.vertica.tsa.ensemble import TimeSeriesByCategory model = TimeSeriesByCategory(base_model = base_model)
Model Fitting¶
We can now fit the model:
model.fit(amazon, ts = "date", y = "number", by = "state") For category: PERNAMBUCO ============ coefficients ============ parameter| value ---------+-------- phi_1 | 0.45168 phi_2 | 0.01325 theta_1 |-0.66935 theta_2 |-0.34679 ============== regularization ============== none =============== timeseries_name =============== "number" ============== timestamp_name ============== date ============== missing_method ============== linear_interpolation =========== call_string =========== ARIMA('"public"."_verticapy_tmp_timeseriesbycategory_v_mldb_9f553e78979e11efa8720242ac120002__000_tsbc"', '"public"."_verticapy_tmp_view_v_mldb_9ff4e3f6979e11efa8720242ac120002_"', '"number"', 'date' USING PARAMETERS p=2, d=1, q=2, missing='linear_interpolation', init_method='Zero', epsilon=1e-06, max_iterations=100); =============== Additional Info =============== Name | Value ------------------+----------- p | 2 d | 1 q | 2 mean | 0.27731 lambda | 1.00000 mean_squared_error|13117.49869 rejected_row_count| 0 accepted_row_count| 239 For category: SERGIPE ============ coefficients ============ parameter| value ---------+-------- phi_1 |-0.03152 phi_2 | 0.27608 theta_1 |-0.25561 theta_2 |-0.73844 ============== regularization ============== none =============== timeseries_name =============== "number" ============== timestamp_name ============== date ============== missing_method ============== linear_interpolation =========== call_string =========== ARIMA('"public"."_verticapy_tmp_timeseriesbycategory_v_mldb_9f553e78979e11efa8720242ac120002__001_tsbc"', '"public"."_verticapy_tmp_view_v_mldb_a2d9f0a2979e11efa8720242ac120002_"', '"number"', 'date' USING PARAMETERS p=2, d=1, q=2, missing='linear_interpolation', init_method='Zero', epsilon=1e-06, max_iterations=100); =============== Additional Info =============== Name | Value ------------------+--------- p | 2 d | 1 q | 2 mean | 0.00420 lambda | 1.00000 mean_squared_error|319.36890 rejected_row_count| 0 accepted_row_count| 239
Important
To train a model, you can directly use the
vDataFrameor the name of the relation stored in the database. The test set is optional and is only used to compute the test metrics. Inverticapy, we don’t work usingXmatrices andyvectors. Instead, we work directly with lists of predictors and the response name.Plots¶
We can conveniently plot the predictions on a line plot to observe the efficacy of our model. We need to provide the
idxwhich represents the model number.model.plot(idx = 0, npredictions = 5)
Note
You can find out the name of the category by the
distinctattribute. The sequential list of categories correspond toidx = 0, 1 ....model.distinct.- __init__(name: str = None, overwrite_model: bool = False, base_model: TimeSeriesModelBase | None = None) None¶
Must be overridden in the child class
Methods
__init__([name, overwrite_model, base_model])Must be overridden in the child class
contour([nbins, chart])Draws the model's contour plot.
deploySQL([vdf, ts, y, start, npredictions, ...])Returns the SQL code needed to deploy the model.
does_model_exists(name[, raise_error, ...])Checks whether the model is stored in the Vertica database.
drop()Drops the model from the Vertica database.
export_models(name, path[, kind])Exports machine learning models.
features_importance([idx, show, chart])Computes the input submodel's features importance.
fit(input_relation, ts, y, by[, ...])Trains the model.
get_attributes([attr_name])Returns the model attributes.
get_match_index(x, col_list[, str_check])Returns the matching index.
Returns the parameters of the model.
get_plotting_lib([class_name, chart, ...])Returns the first available library (Plotly, Matplotlib, or Highcharts) to draw a specific graphic.
get_vertica_attributes([attr_name])Returns the model Vertica attributes.
import_models(path[, schema, kind])Imports machine learning models.
plot([idx, vdf, ts, y, start, npredictions, ...])Draws the input submodel.
predict([vdf, ts, y, start, npredictions, ...])Predicts using the input relation.
register(registered_name[, raise_error])Registers the model and adds it to in-DB Model versioning environment with a status of 'under_review'.
regression_report([metrics, start, ...])Computes a regression report using multiple metrics to evaluate the model (
r2,mse,max error...).report([metrics, start, npredictions, method])Computes a regression report using multiple metrics to evaluate the model (
r2,mse,max error...).score([metric, start, npredictions, method])Computes the model score.
set_params([parameters])Sets the parameters of the model.
Summarizes the model.
to_binary(path)Exports the model to the Vertica Binary format.
to_pmml(path)Exports the model to PMML.
to_python([return_proba, ...])Returns the Python function needed for in-memory scoring without using built-in Vertica functions.
to_sql([X, return_proba, ...])Returns the SQL code needed to deploy the model without using built-in Vertica functions.
to_tf(path)Exports the model to the Frozen Graph format (TensorFlow).
Attributes