Loading...

Time Series

Time series models are a type of regression on a dataset with a timestamp label.

The following example creates a time series model to predict the number of forest fires in Brazil with the amazon dataset.

from verticapy.datasets import load_amazon

amazon = load_amazon().groupby("date", "SUM(number) AS number")
amazon.head(100)
📅
date
Date
100%
123
number
Integer
100%
11998-01-010
21998-05-010
31998-10-0123495
41998-12-014448
51999-03-01667
61999-04-01717
72000-01-01778
82000-03-01848
92000-09-0123291
102000-10-0127336
112000-12-014465
122001-01-01547
132001-12-016201
142002-03-011679
152002-04-011682
162002-05-013818
172002-08-0157151
182002-10-0147722
192002-11-0128179
202002-12-0111944
212003-03-012749
222003-06-016506
232003-09-0176325
242003-10-0143295
252004-02-011255
262004-03-012040
272004-10-0140331
282005-01-014990
292005-02-012153
302005-03-011706
312005-08-0151981
322005-11-0121752
332006-05-01808
342007-01-013055
352007-02-011751
362007-03-012136
372007-05-011286
382007-07-017808
392007-09-0194526
402008-04-011253
412008-06-011287
422008-09-0139445
432008-12-014995
442009-02-011140
452009-08-0117559
462009-09-0129430
472010-04-012200
482010-09-0185408
492011-01-011416
502011-11-0112221
512012-02-011436
522012-03-012058
532012-05-013240
542012-06-015890
552012-10-0134213
562012-12-016823
572013-04-011372
582013-09-0131585
592013-12-0112006
602014-01-012633
612014-05-013189
622014-07-0110801
632014-12-0110938
642015-04-012573
652015-05-012384
662015-06-015810
672015-09-0172083
682015-11-0127529
692016-01-015979
702016-02-014147
712016-05-013568
722016-07-0119141
732016-10-0130210
742017-01-012408
752017-06-017506
762017-07-0122911
772017-11-0117585
781998-04-010
791998-08-0135549
801998-09-0141968
811998-11-016804
821999-02-011284
831999-05-011812
841999-08-0139486
851999-09-0136913
862000-02-01561
872000-04-01537
882000-06-016275
892000-11-018399
902001-04-011081
912001-05-012090
922001-07-016490
932001-10-0131038
942002-06-0110839
952002-09-0155803
962003-07-0111804
972003-08-0143736
982003-11-0123572
992003-12-0115342
1002004-01-012705

The feature date tells us that we should be working with a time series model. To do predictions on time series, we use previous values called lags.

To help visualize the seasonality of forest fires, we’ll draw some autocorrelation plots.

amazon.acf(
    ts = "date",
    column = "number",
    p = 24,
)
amazon.pacf(
    ts = "date",
    column = "number",
    p = 8,
)

Forest fires follow a predictable, seasonal pattern, so it should be easy to predict future forest fires with past data.

VerticaPy offers several models, including a multiple time series model. For this example, let’s use a ARIMA model.

from verticapy.machine_learning.vertica import ARIMA

model = ARIMA(order = (12, 0, 1))

model.fit(
    amazon,
    y = "number",
    ts = "date",
)



============
coefficients
============
parameter| value  
---------+--------
  phi_1  | 0.31880
  phi_2  |-0.10672
  phi_3  |-0.07836
  phi_4  |-0.03784
  phi_5  |-0.00728
  phi_6  | 0.05855
  phi_7  |-0.05685
  phi_8  | 0.02299
  phi_9  |-0.02943
 phi_10  |-0.03170
 phi_11  | 0.08016
 phi_12  | 0.67982
 theta_1 | 0.37215


==============
regularization
==============
none

===============
timeseries_name
===============
"number"

==============
timestamp_name
==============
date

==============
missing_method
==============
linear_interpolation

===========
call_string
===========
ARIMA('"public"."_verticapy_tmp_arima_v_mldb_01927f5497bf11efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_01e38a5c97bf11efa8720242ac120002_"', '"number"', 'date' USING PARAMETERS p=12, d=0, q=1, missing='linear_interpolation', init_method='Zero', epsilon=1e-06, max_iterations=100);

===============
Additional Info
===============
       Name       |     Value     
------------------+---------------
        p         |      12       
        d         |       0       
        q         |       1       
       mean       |  15288.86611  
      lambda      |    1.00000    
mean_squared_error|110069159.22574
rejected_row_count|       0       
accepted_row_count|      239      

Just like with other regression models, we’ll evaluate our model with the report() method.

model.report(npredictions = 50, start = 50)
value
explained_variance0.886197750335032
max_error27071.0112210089
median_absolute_error2512.26144236847
mean_absolute_error4999.39474659287
mean_squared_error61686598.2683679
root_mean_squared_error7854.08163112454
r20.881534696051036
r2_adj0.879066668885433
aic901.304394780616
bic904.702908876578

We can also draw our model using one-step ahead and dynamic forecasting.

model.plot(amazon, npredictions = 40,)

In the next lesson, we’ll go over Regression