verticapy.machine_learning.vertica.tsa.ARIMA.fit¶
- ARIMA.fit(input_relation: Annotated[str | vDataFrame, ''], ts: str, y: Annotated[str | list[str], 'STRING representing one column or a list of columns'], test_relation: Annotated[str | vDataFrame, ''] = '', return_report: bool = False) str | None¶
Trains the model.
Parameters¶
- input_relation: SQLRelation
Training relation.
- ts: str
TS (Time Series) :py:class`vDataColumn` used to order the data. The :py:class`vDataColumn` type must be
date(date,datetime,timestamp…) or numerical.- y: SQLColumns
Response column.
In the case of multivariate analysis, it represents a
listof all the predictors.- test_relation: SQLRelation, optional
Relation used to test the model.
- return_report: bool, optional
[For native models] When set to
True, the model summary will be returned. Otherwise, it will be printed.
Returns¶
- str
model’s summary.
Examples¶
We import
verticapy:import verticapy as vp
For this example, we will use the airline passengers dataset.
import verticapy.datasets as vpd data = vpd.load_airline_passengers()
📅date123passengers1 1949-06-01 135 2 1950-05-01 125 3 1950-09-01 158 4 1950-11-01 114 5 1951-02-01 150 6 1951-04-01 163 7 1951-05-01 172 8 1951-07-01 199 9 1951-11-01 146 10 1952-02-01 180 11 1952-07-01 230 12 1953-02-01 196 13 1953-03-01 236 14 1953-07-01 264 15 1953-10-01 211 16 1954-10-01 229 17 1955-02-01 233 18 1955-09-01 312 19 1955-12-01 278 20 1956-01-01 284 21 1956-02-01 277 22 1956-09-01 355 23 1957-05-01 355 24 1957-09-01 404 25 1958-05-01 363 26 1958-10-01 359 27 1959-02-01 342 28 1959-04-01 396 29 1959-08-01 559 30 1959-10-01 407 31 1959-11-01 362 32 1960-05-01 472 33 1960-09-01 508 34 1960-10-01 461 35 1960-12-01 432 36 1949-03-01 132 37 1949-05-01 121 38 1949-07-01 148 39 1949-08-01 148 40 1949-10-01 119 41 1950-02-01 126 42 1950-03-01 141 43 1950-04-01 135 44 1950-08-01 170 45 1950-12-01 140 46 1951-06-01 178 47 1951-08-01 199 48 1951-10-01 162 49 1952-01-01 171 50 1952-03-01 193 51 1952-04-01 181 52 1952-08-01 242 53 1953-04-01 235 54 1953-05-01 229 55 1953-09-01 237 56 1953-11-01 180 57 1954-01-01 204 58 1954-04-01 227 59 1954-06-01 264 60 1954-07-01 302 61 1954-08-01 293 62 1954-09-01 259 63 1954-11-01 203 64 1955-03-01 267 65 1955-05-01 270 66 1955-10-01 274 67 1955-11-01 237 68 1956-05-01 318 69 1956-06-01 374 70 1956-07-01 413 71 1956-08-01 405 72 1956-11-01 271 73 1957-03-01 356 74 1957-04-01 348 75 1957-07-01 465 76 1957-11-01 305 77 1958-01-01 340 78 1958-03-01 362 79 1958-06-01 435 80 1958-07-01 491 81 1958-08-01 505 82 1959-05-01 420 83 1960-01-01 417 84 1949-02-01 118 85 1949-04-01 129 86 1949-11-01 104 87 1950-07-01 170 88 1950-10-01 133 89 1951-01-01 145 90 1951-03-01 178 91 1951-09-01 184 92 1951-12-01 166 93 1952-06-01 218 94 1952-09-01 209 95 1952-10-01 191 96 1952-12-01 194 97 1953-08-01 272 98 1953-12-01 201 99 1954-03-01 235 100 1954-05-01 234 Rows: 1-100 | Columns: 2First we import the model:
from verticapy.machine_learning.vertica.tsa import ARIMA
Then we can create the model:
model = ARIMA(order = (12, 1, 2))
We can now fit the model:
model.fit(data, "date", "passengers") ============ coefficients ============ parameter| value ---------+-------- phi_1 |-0.02408 phi_2 |-0.03398 phi_3 |-0.02702 phi_4 |-0.12197 phi_5 |-0.01651 phi_6 |-0.21558 phi_7 |-0.00477 phi_8 |-0.15146 phi_9 | 0.04249 phi_10 |-0.16296 phi_11 | 0.04043 phi_12 | 0.86090 theta_1 | 0.06580 theta_2 |-0.06794 ============== regularization ============== none =============== timeseries_name =============== "passengers" ============== timestamp_name ============== date ============== missing_method ============== linear_interpolation =========== call_string =========== ARIMA('"public"."_verticapy_tmp_arima_v_mldb_b87f5c5e979d11efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_b8a0837a979d11efa8720242ac120002_"', '"passengers"', 'date' USING PARAMETERS p=12, d=1, q=2, missing='linear_interpolation', init_method='Zero', epsilon=1e-06, max_iterations=100); =============== Additional Info =============== Name | Value ------------------+--------- p | 12 d | 1 q | 2 mean | 2.23776 lambda | 1.00000 mean_squared_error|178.86952 rejected_row_count| 0 accepted_row_count| 144