Time Series¶
Time series models are a type of regression on a dataset with a timestamp label.
The following example creates a time series model to predict the number of forest fires in Brazil with the amazon dataset.
from verticapy.datasets import load_amazon
amazon = load_amazon().groupby("date", "SUM(number) AS number")
amazon.head(100)
📅 date100% | 123 number100% | |
| 1 | 1998-01-01 | 0 |
| 2 | 1998-05-01 | 0 |
| 3 | 1998-10-01 | 23495 |
| 4 | 1998-12-01 | 4448 |
| 5 | 1999-03-01 | 667 |
| 6 | 1999-04-01 | 717 |
| 7 | 2000-01-01 | 778 |
| 8 | 2000-03-01 | 848 |
| 9 | 2000-09-01 | 23291 |
| 10 | 2000-10-01 | 27336 |
| 11 | 2000-12-01 | 4465 |
| 12 | 2001-01-01 | 547 |
| 13 | 2001-12-01 | 6201 |
| 14 | 2002-03-01 | 1679 |
| 15 | 2002-04-01 | 1682 |
| 16 | 2002-05-01 | 3818 |
| 17 | 2002-08-01 | 57151 |
| 18 | 2002-10-01 | 47722 |
| 19 | 2002-11-01 | 28179 |
| 20 | 2002-12-01 | 11944 |
| 21 | 2003-03-01 | 2749 |
| 22 | 2003-06-01 | 6506 |
| 23 | 2003-09-01 | 76325 |
| 24 | 2003-10-01 | 43295 |
| 25 | 2004-02-01 | 1255 |
| 26 | 2004-03-01 | 2040 |
| 27 | 2004-10-01 | 40331 |
| 28 | 2005-01-01 | 4990 |
| 29 | 2005-02-01 | 2153 |
| 30 | 2005-03-01 | 1706 |
| 31 | 2005-08-01 | 51981 |
| 32 | 2005-11-01 | 21752 |
| 33 | 2006-05-01 | 808 |
| 34 | 2007-01-01 | 3055 |
| 35 | 2007-02-01 | 1751 |
| 36 | 2007-03-01 | 2136 |
| 37 | 2007-05-01 | 1286 |
| 38 | 2007-07-01 | 7808 |
| 39 | 2007-09-01 | 94526 |
| 40 | 2008-04-01 | 1253 |
| 41 | 2008-06-01 | 1287 |
| 42 | 2008-09-01 | 39445 |
| 43 | 2008-12-01 | 4995 |
| 44 | 2009-02-01 | 1140 |
| 45 | 2009-08-01 | 17559 |
| 46 | 2009-09-01 | 29430 |
| 47 | 2010-04-01 | 2200 |
| 48 | 2010-09-01 | 85408 |
| 49 | 2011-01-01 | 1416 |
| 50 | 2011-11-01 | 12221 |
| 51 | 2012-02-01 | 1436 |
| 52 | 2012-03-01 | 2058 |
| 53 | 2012-05-01 | 3240 |
| 54 | 2012-06-01 | 5890 |
| 55 | 2012-10-01 | 34213 |
| 56 | 2012-12-01 | 6823 |
| 57 | 2013-04-01 | 1372 |
| 58 | 2013-09-01 | 31585 |
| 59 | 2013-12-01 | 12006 |
| 60 | 2014-01-01 | 2633 |
| 61 | 2014-05-01 | 3189 |
| 62 | 2014-07-01 | 10801 |
| 63 | 2014-12-01 | 10938 |
| 64 | 2015-04-01 | 2573 |
| 65 | 2015-05-01 | 2384 |
| 66 | 2015-06-01 | 5810 |
| 67 | 2015-09-01 | 72083 |
| 68 | 2015-11-01 | 27529 |
| 69 | 2016-01-01 | 5979 |
| 70 | 2016-02-01 | 4147 |
| 71 | 2016-05-01 | 3568 |
| 72 | 2016-07-01 | 19141 |
| 73 | 2016-10-01 | 30210 |
| 74 | 2017-01-01 | 2408 |
| 75 | 2017-06-01 | 7506 |
| 76 | 2017-07-01 | 22911 |
| 77 | 2017-11-01 | 17585 |
| 78 | 1998-04-01 | 0 |
| 79 | 1998-08-01 | 35549 |
| 80 | 1998-09-01 | 41968 |
| 81 | 1998-11-01 | 6804 |
| 82 | 1999-02-01 | 1284 |
| 83 | 1999-05-01 | 1812 |
| 84 | 1999-08-01 | 39486 |
| 85 | 1999-09-01 | 36913 |
| 86 | 2000-02-01 | 561 |
| 87 | 2000-04-01 | 537 |
| 88 | 2000-06-01 | 6275 |
| 89 | 2000-11-01 | 8399 |
| 90 | 2001-04-01 | 1081 |
| 91 | 2001-05-01 | 2090 |
| 92 | 2001-07-01 | 6490 |
| 93 | 2001-10-01 | 31038 |
| 94 | 2002-06-01 | 10839 |
| 95 | 2002-09-01 | 55803 |
| 96 | 2003-07-01 | 11804 |
| 97 | 2003-08-01 | 43736 |
| 98 | 2003-11-01 | 23572 |
| 99 | 2003-12-01 | 15342 |
| 100 | 2004-01-01 | 2705 |
The feature date tells us that we should be working with a time series model. To do predictions on time series, we use previous values called lags.
To help visualize the seasonality of forest fires, we’ll draw some autocorrelation plots.
amazon.acf(
ts = "date",
column = "number",
p = 24,
)
amazon.pacf(
ts = "date",
column = "number",
p = 8,
)
Forest fires follow a predictable, seasonal pattern, so it should be easy to predict future forest fires with past data.
VerticaPy offers several models, including a multiple time series model. For this example, let’s use a ARIMA model.
from verticapy.machine_learning.vertica import ARIMA
model = ARIMA(order = (12, 0, 1))
model.fit(
amazon,
y = "number",
ts = "date",
)
============
coefficients
============
parameter| value
---------+--------
phi_1 | 0.31880
phi_2 |-0.10672
phi_3 |-0.07836
phi_4 |-0.03784
phi_5 |-0.00728
phi_6 | 0.05855
phi_7 |-0.05685
phi_8 | 0.02299
phi_9 |-0.02943
phi_10 |-0.03170
phi_11 | 0.08016
phi_12 | 0.67982
theta_1 | 0.37215
==============
regularization
==============
none
===============
timeseries_name
===============
"number"
==============
timestamp_name
==============
date
==============
missing_method
==============
linear_interpolation
===========
call_string
===========
ARIMA('"public"."_verticapy_tmp_arima_v_mldb_01927f5497bf11efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_01e38a5c97bf11efa8720242ac120002_"', '"number"', 'date' USING PARAMETERS p=12, d=0, q=1, missing='linear_interpolation', init_method='Zero', epsilon=1e-06, max_iterations=100);
===============
Additional Info
===============
Name | Value
------------------+---------------
p | 12
d | 0
q | 1
mean | 15288.86611
lambda | 1.00000
mean_squared_error|110069159.22574
rejected_row_count| 0
accepted_row_count| 239
Just like with other regression models, we’ll evaluate our model with the report() method.
model.report(npredictions = 50, start = 50)
| value | |
| explained_variance | 0.886197750335032 |
| max_error | 27071.0112210089 |
| median_absolute_error | 2512.26144236847 |
| mean_absolute_error | 4999.39474659287 |
| mean_squared_error | 61686598.2683679 |
| root_mean_squared_error | 7854.08163112454 |
| r2 | 0.881534696051036 |
| r2_adj | 0.879066668885433 |
| aic | 901.304394780616 |
| bic | 904.702908876578 |
We can also draw our model using one-step ahead and dynamic forecasting.
model.plot(amazon, npredictions = 40,)
In the next lesson, we’ll go over Regression