verticapy.machine_learning.vertica.linear_model.PoissonRegressor¶
- class verticapy.machine_learning.vertica.linear_model.PoissonRegressor(name: str = None, overwrite_model: bool = False, tol: float = 1e-06, penalty: Literal['none', 'l2', None] = 'none', C: Annotated[int | float | Decimal, 'Python Numbers'] = 1.0, max_iter: int = 100, solver: Literal['newton'] = 'newton', fit_intercept: bool = True)¶
Creates an
PoissonRegressorobject using the Vertica Poisson Regression algorithm.Parameters¶
- name: str, optional
Name of the model. The model is stored in the database.
- overwrite_model: bool, optional
If set to
True, training a model with the same name as an existing model overwrites the existing model.- tol: float, optional
Determines whether the algorithm has reached the specified accuracy result.
- penalty: str, optional
Determines the method of regularization.
- None:
No Regularization.
- l2:
L2Regularization.
- C: PythonNumber, optional
The regularization parameter value. The value must be zero or non-negative.
- max_iter: int, optional
Determines the maximum number of iterations the algorithm performs before achieving the specified accuracy result.
- solver: str, optional
The optimizer method used to train the model.
- newton:
Newton Method.
- fit_intercept: bool, optional
boolean, specifies whether the model includes an intercept. If set toFalse, no intercept is used in training the model. Note that settingfit_intercepttoFalsedoes not work well with the BFGS optimizer.
Attributes¶
Many attributes are created during the fitting phase.
- coef_: numpy.array
The regression coefficients. The order of coefficients is the same as the order of columns used during the fitting phase.
- intercept_: float
The expected value of the dependent variable when all independent variables are zero, serving as the baseline or constant term in the model.
- features_importance_: numpy.array
The importance of features is computed through the model coefficients, which are normalized based on their range. Subsequently, an activation function calculates the final score. It is necessary to use the
features_importance()method to compute it initially, and the computed values will be subsequently utilized for subsequent calls.
Note
All attributes can be accessed using the
get_attributes()method.Note
Several other attributes can be accessed by using the
get_vertica_attributes()method.Examples¶
The following examples provide a basic understanding of usage. For more detailed examples, please refer to the Machine Learning or the Examples section on the website.
Load data for machine learning¶
We import
verticapy:import verticapy as vp
Hint
By assigning an alias to
verticapy, we mitigate the risk of code collisions with other libraries. This precaution is necessary because verticapy uses commonly known function names like “average” and “median”, which can potentially lead to naming conflicts. The use of an alias ensures that the functions fromverticapyare used as intended without interfering with functions from other libraries.For this example, we will use the winequality dataset.
import verticapy.datasets as vpd data = vpd.load_winequality()
123fixed_acidity123volatile_acidity123citric_acid123residual_sugar123chlorides123free_sulfur_dioxide123total_sulfur_dioxide123density123pH123sulphates123alcohol123quality123goodAbccolor1 3.9 0.225 0.4 4.2 0.03 29.0 118.0 0.989 3.57 0.36 12.8 8 1 white 2 4.7 0.335 0.14 1.3 0.036 69.0 168.0 0.99212 3.47 0.46 10.5 5 0 white 3 4.7 0.455 0.18 1.9 0.036 33.0 106.0 0.98746 3.21 0.83 14.0 7 1 white 4 4.7 0.785 0.0 3.4 0.036 23.0 134.0 0.98981 3.53 0.92 13.8 6 0 white 5 4.9 0.345 0.34 1.0 0.068 32.0 143.0 0.99138 3.24 0.4 10.1 5 0 white 6 4.9 0.345 0.34 1.0 0.068 32.0 143.0 0.99138 3.24 0.4 10.1 5 0 white 7 4.9 0.42 0.0 2.1 0.048 16.0 42.0 0.99154 3.71 0.74 14.0 7 1 red 8 5.0 0.27 0.4 1.2 0.076 42.0 124.0 0.99204 3.32 0.47 10.1 6 0 white 9 5.0 0.31 0.0 6.4 0.046 43.0 166.0 0.994 3.3 0.63 9.9 6 0 white 10 5.0 0.4 0.5 4.3 0.046 29.0 80.0 0.9902 3.49 0.66 13.6 6 0 red 11 5.0 0.44 0.04 18.6 0.039 38.0 128.0 0.9985 3.37 0.57 10.2 6 0 white 12 5.1 0.11 0.32 1.6 0.028 12.0 90.0 0.99008 3.57 0.52 12.2 6 0 white 13 5.1 0.14 0.25 0.7 0.039 15.0 89.0 0.9919 3.22 0.43 9.2 6 0 white 14 5.1 0.165 0.22 5.7 0.047 42.0 146.0 0.9934 3.18 0.55 9.9 6 0 white 15 5.1 0.33 0.22 1.6 0.027 18.0 89.0 0.9893 3.51 0.38 12.5 7 1 white 16 5.1 0.33 0.22 1.6 0.027 18.0 89.0 0.9893 3.51 0.38 12.5 7 1 white 17 5.1 0.33 0.22 1.6 0.027 18.0 89.0 0.9893 3.51 0.38 12.5 7 1 white 18 5.1 0.39 0.21 1.7 0.027 15.0 72.0 0.9894 3.5 0.45 12.5 6 0 white 19 5.2 0.2 0.27 3.2 0.047 16.0 93.0 0.99235 3.44 0.53 10.1 7 1 white 20 5.2 0.21 0.31 1.7 0.048 17.0 61.0 0.98953 3.24 0.37 12.0 7 1 white 21 5.2 0.22 0.46 6.2 0.066 41.0 187.0 0.99362 3.19 0.42 9.73333333333333 5 0 white 22 5.2 0.31 0.2 2.4 0.027 27.0 117.0 0.98886 3.56 0.45 13.0 7 1 white 23 5.2 0.32 0.25 1.8 0.103 13.0 50.0 0.9957 3.38 0.55 9.2 5 0 red 24 5.2 0.34 0.37 6.2 0.031 42.0 133.0 0.99076 3.25 0.41 12.5 6 0 white 25 5.2 0.36 0.02 1.6 0.031 24.0 104.0 0.9896 3.44 0.35 12.2 6 0 white 26 5.2 0.365 0.08 13.5 0.041 37.0 142.0 0.997 3.46 0.39 9.9 6 0 white 27 5.2 0.48 0.04 1.6 0.054 19.0 106.0 0.9927 3.54 0.62 12.2 7 1 red 28 5.2 0.5 0.18 2.0 0.036 23.0 129.0 0.98949 3.36 0.77 13.4 7 1 white 29 5.3 0.16 0.39 1.0 0.028 40.0 101.0 0.99156 3.57 0.59 10.6 6 0 white 30 5.3 0.16 0.39 1.0 0.028 40.0 101.0 0.99156 3.57 0.59 10.6 6 0 white 31 5.3 0.165 0.24 1.1 0.051 25.0 105.0 0.9925 3.32 0.47 9.1 5 0 white 32 5.3 0.23 0.56 0.9 0.041 46.0 141.0 0.99119 3.16 0.62 9.7 5 0 white 33 5.3 0.3 0.3 1.2 0.029 25.0 93.0 0.98742 3.31 0.4 13.6 7 1 white 34 5.3 0.33 0.3 1.2 0.048 25.0 119.0 0.99045 3.32 0.62 11.3 6 0 white 35 5.3 0.36 0.27 6.3 0.028 40.0 132.0 0.99186 3.37 0.4 11.6 6 0 white 36 5.3 0.36 0.27 6.3 0.028 40.0 132.0 0.99186 3.37 0.4 11.6 6 0 white 37 5.3 0.4 0.25 3.9 0.031 45.0 130.0 0.99072 3.31 0.58 11.75 7 1 white 38 5.3 0.47 0.11 2.2 0.048 16.0 89.0 0.99182 3.54 0.88 13.6 7 1 red 39 5.3 0.47 0.11 2.2 0.048 16.0 89.0 0.99182 3.54 0.88 13.5666666666667 7 1 red 40 5.3 0.715 0.19 1.5 0.161 7.0 62.0 0.99395 3.62 0.61 11.0 5 0 red 41 5.4 0.22 0.29 1.2 0.045 69.0 152.0 0.99178 3.76 0.63 11.0 7 1 white 42 5.4 0.595 0.1 2.8 0.042 26.0 80.0 0.9932 3.36 0.38 9.3 5 0 white 43 5.4 0.74 0.09 1.7 0.089 16.0 26.0 0.99402 3.67 0.56 11.6 6 0 red 44 5.5 0.12 0.33 1.0 0.038 23.0 131.0 0.99164 3.25 0.45 9.8 5 0 white 45 5.5 0.12 0.33 1.0 0.038 23.0 131.0 0.99164 3.25 0.45 9.8 5 0 white 46 5.5 0.14 0.27 4.6 0.029 22.0 104.0 0.9949 3.34 0.44 9.0 5 0 white 47 5.5 0.14 0.27 4.6 0.029 22.0 104.0 0.9949 3.34 0.44 9.0 5 0 white 48 5.5 0.16 0.31 1.2 0.026 31.0 68.0 0.9898 3.33 0.44 11.65 6 0 white 49 5.5 0.16 0.31 1.2 0.026 31.0 68.0 0.9898 3.33 0.44 11.6333333333333 6 0 white 50 5.5 0.18 0.22 5.5 0.037 10.0 86.0 0.99156 3.46 0.44 12.2 5 0 white 51 5.5 0.24 0.45 1.7 0.046 22.0 113.0 0.99224 3.22 0.48 10.0 5 0 white 52 5.5 0.29 0.3 1.1 0.022 20.0 110.0 0.98869 3.34 0.38 12.8 7 1 white 53 5.5 0.31 0.29 3.0 0.027 16.0 102.0 0.99067 3.23 0.56 11.2 6 0 white 54 5.5 0.32 0.45 4.9 0.028 25.0 191.0 0.9922 3.51 0.49 11.5 7 1 white 55 5.5 0.35 0.35 1.1 0.045 14.0 167.0 0.992 3.34 0.68 9.9 6 0 white 56 5.5 0.375 0.38 1.7 0.036 17.0 98.0 0.99142 3.29 0.39 10.5 6 0 white 57 5.6 0.15 0.26 5.55 0.051 51.0 139.0 0.99336 3.47 0.5 11.0 6 0 white 58 5.6 0.15 0.31 5.3 0.038 8.0 79.0 0.9923 3.3 0.39 10.5 6 0 white 59 5.6 0.16 0.27 1.4 0.044 53.0 168.0 0.9918 3.28 0.37 10.1 6 0 white 60 5.6 0.175 0.29 0.8 0.043 20.0 67.0 0.99112 3.28 0.48 9.9 6 0 white 61 5.6 0.185 0.19 7.1 0.048 36.0 110.0 0.99438 3.26 0.41 9.5 6 0 white 62 5.6 0.185 0.19 7.1 0.048 36.0 110.0 0.99438 3.26 0.41 9.5 6 0 white 63 5.6 0.22 0.32 1.2 0.024 29.0 97.0 0.98823 3.2 0.46 13.05 7 1 white 64 5.6 0.26 0.18 1.4 0.034 18.0 135.0 0.99174 3.32 0.35 10.2 6 0 white 65 5.6 0.26 0.26 5.7 0.031 12.0 80.0 0.9923 3.25 0.38 10.8 5 0 white 66 5.6 0.26 0.5 11.4 0.029 25.0 93.0 0.99428 3.23 0.49 10.5 6 0 white 67 5.6 0.28 0.28 4.2 0.044 52.0 158.0 0.992 3.35 0.44 10.7 7 1 white 68 5.6 0.3 0.1 6.4 0.043 34.0 142.0 0.99382 3.14 0.48 9.8 5 0 white 69 5.6 0.35 0.14 5.0 0.046 48.0 198.0 0.9937 3.3 0.71 10.3 5 0 white 70 5.6 0.49 0.13 4.5 0.039 17.0 116.0 0.9907 3.42 0.9 13.7 7 1 white 71 5.6 0.49 0.13 4.5 0.039 17.0 116.0 0.9907 3.42 0.9 13.7 7 1 white 72 5.6 0.66 0.0 2.2 0.087 3.0 11.0 0.99378 3.71 0.63 12.8 7 1 red 73 5.6 0.66 0.0 2.2 0.087 3.0 11.0 0.99378 3.71 0.63 12.8 7 1 red 74 5.7 0.15 0.47 11.4 0.035 49.0 128.0 0.99456 3.03 0.34 10.5 8 1 white 75 5.7 0.18 0.26 2.2 0.023 21.0 95.0 0.9893 3.07 0.54 12.3 6 0 white 76 5.7 0.18 0.36 1.2 0.046 9.0 71.0 0.99199 3.7 0.68 10.9 7 1 white 77 5.7 0.2 0.3 2.5 0.046 38.0 125.0 0.99276 3.34 0.5 9.9 6 0 white 78 5.7 0.21 0.32 0.9 0.038 38.0 121.0 0.99074 3.24 0.46 10.6 6 0 white 79 5.7 0.21 0.37 4.5 0.04 58.0 140.0 0.99332 3.29 0.62 10.6 6 0 white 80 5.7 0.22 0.2 16.0 0.044 41.0 113.0 0.99862 3.22 0.46 8.9 6 0 white 81 5.7 0.22 0.2 16.0 0.044 41.0 113.0 0.99862 3.22 0.46 8.9 6 0 white 82 5.7 0.22 0.2 16.0 0.044 41.0 113.0 0.99862 3.22 0.46 8.9 6 0 white 83 5.7 0.22 0.2 16.0 0.044 41.0 113.0 0.99862 3.22 0.46 8.9 6 0 white 84 5.7 0.22 0.2 16.0 0.044 41.0 113.0 0.99862 3.22 0.46 8.9 6 0 white 85 5.7 0.22 0.29 3.5 0.04 27.0 146.0 0.98999 3.17 0.36 12.1 6 0 white 86 5.7 0.23 0.28 9.65 0.025 26.0 121.0 0.9925 3.28 0.38 11.3 6 0 white 87 5.7 0.25 0.26 12.5 0.049 52.5 106.0 0.99691 3.08 0.45 9.4 6 0 white 88 5.7 0.25 0.26 12.5 0.049 52.5 120.0 0.99691 3.08 0.45 9.4 6 0 white 89 5.7 0.25 0.27 11.5 0.04 24.0 120.0 0.99411 3.33 0.31 10.8 6 0 white 90 5.7 0.26 0.24 17.8 0.059 23.0 124.0 0.99773 3.3 0.5 10.1 5 0 white 91 5.7 0.26 0.24 17.8 0.059 23.0 124.0 0.99773 3.3 0.5 10.1 5 0 white 92 5.7 0.26 0.24 17.8 0.059 23.0 124.0 0.99773 3.3 0.5 10.1 5 0 white 93 5.7 0.27 0.32 1.2 0.046 20.0 155.0 0.9934 3.8 0.41 10.2 6 0 white 94 5.7 0.28 0.24 17.5 0.044 60.0 167.0 0.9989 3.31 0.44 9.4 5 0 white 95 5.7 0.32 0.18 1.4 0.029 26.0 104.0 0.9906 3.44 0.37 11.0 6 0 white 96 5.7 0.32 0.38 4.75 0.033 23.0 94.0 0.991 3.42 0.42 11.8 7 1 white 97 5.7 0.36 0.34 4.2 0.026 21.0 77.0 0.9907 3.41 0.45 11.9 6 0 white 98 5.8 0.14 0.15 6.1 0.042 27.0 123.0 0.99362 3.06 0.6 9.9 6 0 white 99 5.8 0.15 0.32 1.2 0.037 14.0 119.0 0.99137 3.19 0.5 10.2 6 0 white 100 5.8 0.17 0.34 1.8 0.045 96.0 170.0 0.99035 3.38 0.9 11.8 8 1 white Rows: 1-100 | Columns: 14Note
VerticaPy offers a wide range of sample datasets that are ideal for training and testing purposes. You can explore the full list of available datasets in the Datasets, which provides detailed information on each dataset and how to use them effectively. These datasets are invaluable resources for honing your data analysis and machine learning skills within the VerticaPy environment.
You can easily divide your dataset into training and testing subsets using the
vDataFrame.train_test_split()method. This is a crucial step when preparing your data for machine learning, as it allows you to evaluate the performance of your models accurately.data = vpd.load_winequality() train, test = data.train_test_split(test_size = 0.2)
Warning
In this case, VerticaPy utilizes seeded randomization to guarantee the reproducibility of your data split. However, please be aware that this approach may lead to reduced performance. For a more efficient data split, you can use the
vDataFrame.to_db()method to save your results intotablesortemporary tables. This will help enhance the overall performance of the process.Model Initialization¶
First we import the
PoissonRegressormodel:from verticapy.machine_learning.vertica import PoissonRegressor
Then we can create the model:
model = PoissonRegressor( tol = 1e-6, penalty = 'L2', C = 1, max_iter = 100, fit_intercept = True, )
Hint
In
verticapy1.0.x and higher, you do not need to specify the model name, as the name is automatically assigned. If you need to re-use the model, you can fetch the model name from the model’s attributes.Important
The model name is crucial for the model management system and versioning. It’s highly recommended to provide a name if you plan to reuse the model later.
Model Training¶
We can now fit the model:
model.fit( train, [ "fixed_acidity", "volatile_acidity", "citric_acid", "residual_sugar", "chlorides", "density", ], "quality", test, ) ======= details ======= predictor |coefficient|std_err |z_value |p_value ----------------+-----------+--------+--------+-------- Intercept | 3.66874 | 0.94790| 3.87040| 0.00011 fixed_acidity | 0.00054 | 0.00526| 0.10290| 0.91804 volatile_acidity| -0.20969 | 0.04493|-4.66709| 0.00000 citric_acid | 0.01685 | 0.04883| 0.34506| 0.73005 residual_sugar | -0.00252 | 0.00133|-1.89905| 0.05756 chlorides | -0.48876 | 0.19110|-2.55765| 0.01054 density | -1.81630 | 0.96359|-1.88492| 0.05944 ============== regularization ============== type| lambda ----+-------- l2 | 1.00000 =========== call_string =========== poisson_reg('"public"."_verticapy_tmp_poissonregressor_v_mldb_e7fe9f34979911efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_e82dfd06979911efa8720242ac120002_"', '"quality"', '"fixed_acidity", "volatile_acidity", "citric_acid", "residual_sugar", "chlorides", "density"' USING PARAMETERS optimizer='newton', epsilon=1e-06, max_iterations=100, regularization='l2', lambda=1, alpha=0.5, fit_intercept=true) =============== Additional Info =============== Name |Value ------------------+----- iteration_count | 8 rejected_row_count| 0 accepted_row_count|5199
Important
To train a model, you can directly use the
vDataFrameor the name of the relation stored in the database. The test set is optional and is only used to compute the test metrics. Inverticapy, we don’t work usingXmatrices andyvectors. Instead, we work directly with lists of predictors and the response name.Metrics¶
We can get the entire report using:
model.report()
value explained_variance 0.0995049092790506 max_error 3.06195928502207 median_absolute_error 0.588193903332221 mean_absolute_error 0.6368840993231 mean_squared_error 0.699515392235532 root_mean_squared_error 0.836370367860754 r2 0.0985192654578602 r2_adj 0.0943295796273004 aic -449.700198982508 bic -413.682930399334 Rows: 1-10 | Columns: 2Important
Most metrics are computed using a single SQL query, but some of them might require multiple SQL queries. Selecting only the necessary metrics in the report can help optimize performance. E.g.
model.report(metrics = ["mse", "r2"]).For
LinearModel, we can easily get the ANOVA table using:model.report(metrics = "anova")
Df SS MS F p_value Regression 6 87.9171591656714 14.652859860945235 20.834192408636817 2.0467009107248934e-23 Residual 1291 907.970979121721 0.7033082719765461 Total 1297 1007.19953775039 Rows: 1-3 | Columns: 6You can also use the
LinearModel.scorefunction to compute the R-squared value:model.score() Out[2]: 0.0985192654578603
Prediction¶
Prediction is straight-forward:
model.predict( test, [ "fixed_acidity", "volatile_acidity", "citric_acid", "residual_sugar", "chlorides", "density", ], "prediction", )
123fixed_acidity123volatile_acidity123citric_acid123residual_sugar123chlorides123free_sulfur_dioxide123total_sulfur_dioxide123density123pH123sulphates123alcohol123quality123goodAbccolor123prediction1 4.2 0.17 0.36 1.8 0.029 93.0 161.0 0.98999 3.65 0.89 12.0 7 1 white 6.20032850975741 2 4.8 0.13 0.32 1.2 0.042 40.0 98.0 0.9898 3.42 0.64 11.8 7 1 white 6.22233739662392 3 4.8 0.17 0.28 2.9 0.03 22.0 111.0 0.9902 3.38 0.34 11.3 7 1 white 6.17145214220889 4 4.8 0.225 0.38 1.2 0.074 47.0 130.0 0.99132 3.31 0.4 10.3 6 0 white 5.99445556641177 5 4.9 0.335 0.14 1.3 0.036 69.0 168.0 0.99212 3.47 0.46 10.4666666666667 5 0 white 5.93369485896087 6 5.0 0.455 0.18 1.9 0.036 33.0 106.0 0.98746 3.21 0.83 14.0 7 1 white 5.83084927996444 7 5.1 0.31 0.3 0.9 0.037 28.0 152.0 0.992 3.54 0.56 10.1 6 0 white 5.986048585754 8 5.1 0.52 0.06 2.7 0.052 30.0 79.0 0.9932 3.32 0.43 9.3 5 0 white 5.62545286257028 9 5.2 0.25 0.23 1.4 0.047 20.0 77.0 0.99001 3.32 0.62 11.4 5 0 white 6.03969361301039 10 5.2 0.335 0.2 1.7 0.033 17.0 74.0 0.99002 3.34 0.48 12.3 6 0 white 5.96609015206905 11 5.3 0.21 0.29 0.7 0.028 11.0 66.0 0.99215 3.3 0.4 9.8 5 0 white 6.14089946315679 12 5.3 0.47 0.1 1.3 0.036 11.0 74.0 0.99082 3.48 0.54 11.2 4 0 white 5.77906963413985 13 5.4 0.33 0.31 4.0 0.03 27.0 108.0 0.99031 3.3 0.43 12.2 7 1 white 5.95504793705801 14 5.4 0.42 0.27 2.0 0.092 23.0 55.0 0.99471 3.78 0.64 12.3 7 1 red 5.64880457298853 15 5.5 0.34 0.26 2.2 0.021 31.0 119.0 0.98919 3.55 0.49 13.0 8 1 white 6.00341146083158 16 5.5 0.485 0.0 1.5 0.065 8.0 103.0 0.994 3.63 0.4 9.7 4 0 white 5.63539125146962 17 5.6 0.13 0.27 4.8 0.028 22.0 104.0 0.9948 3.34 0.45 9.2 6 0 white 6.14977747784627 18 5.6 0.25 0.19 2.4 0.049 42.0 166.0 0.992 3.25 0.43 10.4 6 0 white 5.99413400083637 19 5.6 0.32 0.33 7.4 0.037 25.0 95.0 0.99268 3.25 0.49 11.1 6 0 white 5.87366327667885 20 5.6 0.32 0.33 7.4 0.037 25.0 95.0 0.99268 3.25 0.49 11.1 6 0 white 5.87366327667885 21 5.7 0.12 0.26 5.5 0.034 21.0 99.0 0.99324 3.09 0.57 9.9 6 0 white 6.15050073740849 22 5.7 0.21 0.24 2.3 0.047 60.0 189.0 0.995 3.65 0.72 10.1 6 0 white 6.0245714433699 23 5.7 0.22 0.33 1.9 0.036 37.0 110.0 0.98945 3.26 0.58 12.4 6 0 white 6.12103716839245 24 5.7 0.26 0.3 1.8 0.039 30.0 105.0 0.98995 3.48 0.52 12.5 7 1 white 6.05398272151734 25 5.7 0.27 0.16 9.0 0.053 32.0 111.0 0.99474 3.36 0.37 10.4 6 0 white 5.82730803434429 26 5.7 0.28 0.36 1.8 0.041 38.0 90.0 0.99002 3.27 0.98 11.9 7 1 white 6.02808130423772 27 5.8 0.18 0.28 1.3 0.034 9.0 94.0 0.99092 3.21 0.52 11.2 6 0 white 6.16662720559193 28 5.8 0.27 0.2 14.95 0.044 22.0 179.0 0.9962 3.37 0.37 10.2 5 0 white 5.75472077396028 29 5.8 0.3 0.09 6.3 0.042 36.0 138.0 0.99382 3.15 0.48 9.7 5 0 white 5.86497161819535 30 5.8 0.31 0.32 4.5 0.024 28.0 94.0 0.98906 3.25 0.52 13.7 7 1 white 6.00600408534323 31 5.8 0.415 0.13 1.4 0.04 11.0 64.0 0.9922 3.29 0.52 10.5 5 0 white 5.82313044692686 32 5.9 0.12 0.27 4.8 0.03 40.0 110.0 0.99226 3.55 0.68 12.1 6 0 white 6.18613859074923 33 5.9 0.14 0.2 1.6 0.04 26.0 114.0 0.99105 3.25 0.45 11.4 6 0 white 6.1861970032333 34 5.9 0.24 0.12 1.4 0.035 60.0 247.0 0.99358 3.34 0.44 9.6 6 0 white 6.03971448631876 35 5.9 0.25 0.25 11.3 0.052 30.0 165.0 0.997 3.24 0.44 9.5 6 0 white 5.80636883557108 36 5.9 0.26 0.3 1.0 0.036 38.0 114.0 0.9928 3.58 0.48 9.4 5 0 white 6.04440198763212 37 5.9 0.32 0.26 1.5 0.057 17.0 141.0 0.9917 3.24 0.36 10.7 5 0 white 5.90825168218121 38 5.9 0.36 0.41 1.3 0.047 45.0 104.0 0.9917 3.33 0.51 10.6 6 0 white 5.90548516295853 39 5.9 0.4 0.32 6.0 0.034 50.0 127.0 0.992 3.51 0.58 12.5 7 1 white 5.81203191791498 40 5.9 0.44 0.33 1.2 0.049 12.0 117.0 0.99134 3.46 0.44 11.5 5 0 white 5.79900978642892 41 6.0 0.1 0.24 1.1 0.041 15.0 65.0 0.9927 3.61 0.61 10.3 7 1 white 6.2289753672679 42 6.0 0.17 0.36 1.7 0.042 14.0 61.0 0.99144 3.22 0.54 10.8 6 0 white 6.15239517730898 43 6.0 0.23 0.34 1.3 0.025 23.0 111.0 0.98961 3.36 0.37 12.7 6 0 white 6.1506933140766 44 6.0 0.25 0.28 2.2 0.026 54.0 126.0 0.9898 3.43 0.65 12.9 8 1 white 6.09979955233949 45 6.0 0.29 0.2 12.6 0.045 45.0 187.0 0.9972 3.33 0.42 9.5 5 0 white 5.75206040548763 46 6.0 0.29 0.41 10.8 0.048 55.0 149.0 0.9937 3.09 0.59 10.9666666666667 7 1 white 5.82714794756821 47 6.0 0.33 0.27 0.8 0.185 12.0 188.0 0.9924 3.12 0.62 9.4 5 0 white 5.54229906692986 48 6.0 0.34 0.29 6.1 0.046 29.0 134.0 0.99462 3.48 0.57 10.7 6 0 white 5.81932530882925 49 6.0 0.495 0.27 5.0 0.157 17.0 129.0 0.99396 3.03 0.36 9.3 5 0 white 5.35519529166075 50 6.1 0.2 0.25 1.2 0.038 34.0 128.0 0.9921 3.24 0.44 10.1 5 0 white 6.11514947099654 51 6.1 0.22 0.5 6.6 0.045 30.0 122.0 0.99415 3.22 0.49 9.9 6 0 white 5.98957487513886 52 6.1 0.23 0.45 10.6 0.094 49.0 169.0 0.99699 3.05 0.54 8.8 5 0 white 5.74243481059561 53 6.1 0.27 0.25 1.8 0.041 9.0 109.0 0.9929 3.08 0.54 9.0 5 0 white 5.99939125662913 54 6.1 0.27 0.32 6.2 0.048 47.0 161.0 0.99281 3.22 0.6 11.0 6 0 white 5.92084951289573 55 6.1 0.27 0.44 6.7 0.041 61.0 230.0 0.99505 3.12 0.4 8.9 5 0 white 5.92151983430764 56 6.1 0.28 0.25 17.75 0.044 48.0 161.0 0.9993 3.34 0.48 9.5 5 0 white 5.67593047701386 57 6.1 0.28 0.3 7.75 0.031 33.0 139.0 0.99296 3.22 0.46 11.0 6 0 white 5.93087876468471 58 6.1 0.28 0.32 2.5 0.042 23.0 218.5 0.9935 3.27 0.6 9.8 5 0 white 5.97387666333776 59 6.1 0.32 0.33 10.7 0.036 27.0 98.0 0.99521 3.34 0.52 10.2 6 0 white 5.80266907665926 60 6.1 0.35 0.24 2.3 0.034 25.0 133.0 0.9906 3.34 0.59 12.0 7 1 white 5.93609645731085 61 6.1 0.4 0.18 9.0 0.051 28.5 259.0 0.9964 3.19 0.5 8.8 5 0 white 5.66219346658852 62 6.2 0.2 0.49 1.6 0.065 17.0 143.0 0.9937 3.22 0.52 9.2 6 0 white 6.03608303498881 63 6.2 0.235 0.34 1.9 0.036 4.0 117.0 0.99032 3.4 0.44 12.2 5 0 white 6.09485614538367 64 6.2 0.25 0.31 3.2 0.03 32.0 150.0 0.99014 3.18 0.31 12.0 6 0 white 6.07252446897251 65 6.2 0.25 0.38 7.9 0.045 54.0 208.0 0.99572 3.17 0.46 9.1 5 0 white 5.90400604220314 66 6.2 0.25 0.48 10.0 0.044 78.0 240.0 0.99655 3.25 0.47 9.5 6 0 white 5.87672473624221 67 6.2 0.28 0.45 7.5 0.045 46.0 203.0 0.99573 3.26 0.46 9.2 6 0 white 5.87972881346452 68 6.2 0.3 0.3 2.5 0.041 29.0 82.0 0.99065 3.31 0.61 11.8 7 1 white 5.98098124188257 69 6.2 0.31 0.23 3.3 0.052 34.0 113.0 0.99429 3.16 0.48 8.4 5 0 white 5.87850604688373 70 6.3 0.13 0.42 1.1 0.043 63.0 146.0 0.99066 3.13 0.72 11.2 7 1 white 6.22668370068583 71 6.3 0.19 0.29 2.0 0.022 33.0 96.0 0.98902 3.04 0.54 12.8 7 1 white 6.20307081610548 72 6.3 0.19 0.32 2.8 0.046 18.0 80.0 0.99043 2.92 0.47 11.05 6 0 white 6.10580941179001 73 6.3 0.2 0.19 12.3 0.048 54.0 145.0 0.99668 3.16 0.42 9.3 6 0 white 5.86298679000515 74 6.3 0.2 0.26 12.7 0.046 60.0 143.0 0.99526 3.26 0.35 10.8 6 0 white 5.88487858624834 75 6.3 0.2 0.4 1.5 0.037 35.0 107.0 0.9917 3.46 0.5 11.4 6 0 white 6.13409855932567 76 6.3 0.21 0.58 10.0 0.081 34.0 126.0 0.9962 2.95 0.46 8.9 5 0 white 5.83385412495296 77 6.3 0.23 0.22 3.75 0.039 37.0 116.0 0.9927 3.23 0.5 10.7 6 0 white 6.02591948441593 78 6.3 0.25 0.22 3.3 0.048 41.0 161.0 0.99256 3.16 0.5 10.5 6 0 white 5.98266990842405 79 6.3 0.25 0.22 3.3 0.048 41.0 161.0 0.99256 3.16 0.5 10.5 6 0 white 5.98266990842405 80 6.3 0.28 0.24 8.45 0.031 32.0 172.0 0.9958 3.39 0.57 9.7 7 1 white 5.88464055965873 81 6.3 0.3 0.29 2.1 0.048 33.0 142.0 0.98956 3.22 0.46 12.9 7 1 white 5.97771162735498 82 6.3 0.4 0.24 5.1 0.036 43.0 131.0 0.99186 3.24 0.44 11.3 6 0 white 5.81444932026599 83 6.3 0.47 0.0 1.4 0.055 27.0 33.0 0.9922 3.45 0.48 12.3 6 0 red 5.70335237870838 84 6.4 0.105 0.29 1.1 0.035 44.0 140.0 0.99142 3.17 0.55 10.7 7 1 white 6.26187531239827 85 6.4 0.21 0.21 5.1 0.097 21.0 105.0 0.9939 3.07 0.46 9.6 5 0 white 5.84868359298224 86 6.4 0.21 0.28 5.9 0.047 29.0 101.0 0.99278 3.15 0.4 11.0 6 0 white 6.00054244953618 87 6.4 0.22 0.49 7.5 0.054 42.0 151.0 0.9948 3.27 0.52 10.1 6 0 white 5.94270603820574 88 6.4 0.23 0.33 1.15 0.044 15.5 217.5 0.992 3.33 0.44 11.0 6 0 white 6.07003246658586 89 6.4 0.24 0.26 8.2 0.054 47.0 182.0 0.99538 3.12 0.5 9.5 5 0 white 5.87835222784195 90 6.4 0.24 0.5 11.6 0.047 60.0 211.0 0.9966 3.18 0.57 9.3 5 0 white 5.85882115247162 91 6.4 0.25 0.41 8.6 0.042 57.0 173.0 0.9965 3.0 0.44 9.1 5 0 white 5.89749993029346 92 6.4 0.26 0.21 8.2 0.05 51.0 182.0 0.99542 3.23 0.48 9.5 5 0 white 5.85984212436088 93 6.4 0.28 0.28 3.0 0.04 19.0 98.0 0.99216 3.25 0.47 11.1 6 0 white 5.98367295841877 94 6.4 0.33 0.28 1.1 0.038 30.0 110.0 0.9917 3.12 0.42 10.5 6 0 white 5.96051022485213 95 6.4 0.63 0.21 1.6 0.08 12.0 32.0 0.99689 3.58 0.66 9.8 5 0 red 5.41868965417692 96 6.5 0.13 0.37 1.0 0.036 48.0 114.0 0.9911 3.41 0.51 11.5 8 1 white 6.2400246593108 97 6.5 0.18 0.34 1.6 0.04 43.0 148.0 0.9912 3.32 0.59 11.5 8 1 white 6.14933503541038 98 6.5 0.19 0.32 1.4 0.04 31.0 132.0 0.9922 3.36 0.54 10.8 7 1 white 6.12634499562452 99 6.5 0.22 0.31 3.9 0.046 17.0 106.0 0.99098 3.15 0.31 11.5 5 0 white 6.04430130869772 100 6.5 0.22 0.32 2.2 0.028 36.0 92.0 0.99076 3.27 0.59 11.9 7 1 white 6.12739706622768 Rows: 1-100 | Columns: 15Note
Predictions can be made automatically using the test set, in which case you don’t need to specify the predictors. Alternatively, you can pass only the
vDataFrameto thepredict()function, but in this case, it’s essential that the column names of thevDataFramematch the predictors and response name in the model.Plots¶
If the model allows, you can also generate relevant plots. For example, regression plots can be found in the Machine Learning - Regression Plots.
model.plot()
Important
The plotting feature is typically suitable for models with fewer than three predictors.
Parameter Modification¶
In order to see the parameters:
model.get_params() Out[3]: {'penalty': 'l2', 'tol': 1e-06, 'C': 1, 'max_iter': 100, 'solver': 'newton', 'fit_intercept': True}
And to manually change some of the parameters:
model.set_params({'tol': 0.001})
Model Register¶
In order to register the model for tracking and versioning:
model.register("model_v1")
Please refer to /notebooks/ml/model_tracking_versioning/index.ipynb for more details on model tracking and versioning.
Model Exporting¶
To Memmodel
model.to_memmodel()
Note
MemModelobjects serve as in-memory representations of machine learning models. They can be used for both in-database and in-memory prediction tasks. These objects can be pickled in the same way that you would pickle ascikit-learnmodel.The following methods for exporting the model use
MemModel, and it is recommended to useMemModeldirectly.To SQL
You can get the SQL code by:
model.to_sql() Out[5]: '3.66874013393681 + 0.000541320019358693 * "fixed_acidity" + -0.209690301463013 * "volatile_acidity" + 0.016847826241488 * "citric_acid" + -0.00252263536545883 * "residual_sugar" + -0.488755470954402 * "chlorides" + -1.81629573612736 * "density"'
To Python
To obtain the prediction function in Python syntax, use the following code:
X = [[4.2, 0.17, 0.36, 1.8, 0.029, 0.9899]] model.to_python()(X) Out[7]: array([1.82476574])
Hint
The
to_python()method is used to retrieve predictions, probabilities, or cluster distances. For specific details on how to use this method for different model types, refer to the relevant documentation for each model.- __init__(name: str = None, overwrite_model: bool = False, tol: float = 1e-06, penalty: Literal['none', 'l2', None] = 'none', C: Annotated[int | float | Decimal, 'Python Numbers'] = 1.0, max_iter: int = 100, solver: Literal['newton'] = 'newton', fit_intercept: bool = True) None¶
Methods
__init__([name, overwrite_model, tol, ...])contour([nbins, chart])Draws the model's contour plot.
deploySQL([X])Returns the SQL code needed to deploy the model.
does_model_exists(name[, raise_error, ...])Checks whether the model is stored in the Vertica database.
drop()Drops the model from the Vertica database.
export_models(name, path[, kind])Exports machine learning models.
features_importance([show, chart])Computes the model's features importance.
fit(input_relation, X, y[, test_relation, ...])Trains the model.
get_attributes([attr_name])Returns the model attributes.
get_match_index(x, col_list[, str_check])Returns the matching index.
Returns the parameters of the model.
get_plotting_lib([class_name, chart, ...])Returns the first available library (Plotly, Matplotlib, or Highcharts) to draw a specific graphic.
get_vertica_attributes([attr_name])Returns the model Vertica attributes.
import_models(path[, schema, kind])Imports machine learning models.
plot([max_nb_points, chart])Draws the model.
predict(vdf[, X, name, inplace])Predicts using the input relation.
register(registered_name[, raise_error])Registers the model and adds it to in-DB Model versioning environment with a status of 'under_review'.
regression_report([metrics])Computes a regression report
report([metrics])Computes a regression report
score([metric])Computes the model score.
set_params([parameters])Sets the parameters of the model.
Summarizes the model.
to_binary(path)Exports the model to the Vertica Binary format.
Converts the model to an InMemory object that can be used for different types of predictions.
to_pmml(path)Exports the model to PMML.
to_python([return_proba, ...])Returns the Python function needed for in-memory scoring without using built-in Vertica functions.
to_sql([X, return_proba, ...])Returns the SQL code needed to deploy the model without using built-in Vertica functions.
to_tf(path)Exports the model to the Frozen Graph format (TensorFlow).
Attributes
object_type