Loading...

Model Tracking and Versioning

Introduction

VerticaPy is an open-source Python package on top of Vertica database that supports pandas-like virtual dataframes over database relations. VerticaPy provides scikit-type machine learning functionality on these virtual dataframes. Data is not moved out of the database while performing machine learning or statistical analysis on virtual dataframes. Instead, the computations are done at scale in a distributed fashion inside the Vertica cluster. VerticaPy also takes advantage of multiple Python libraries to create a variety of charts, providing a quick and easy method to illustrate your statistical data.

In this article, we will introduce two new MLOps tools recently added to VerticaPy: Model Tracking and Model Versioning.

Model Tracking

Data scientists usually train many ML models for a project. To help choose the best model, data scientists need a way to keep track of all candidate models and compare them using various metrics. VerticaPy provides a model tracking system to facilitate this process for a given experiment. The data scientist first creates an experiment object and then adds candidate models to that experiment. The information related to each experiment can be automatically backed up in the database, so if the Python environment is closed for any reason, like a holiday, the data scientist has peace of mind that the experiment can be easily retrieved. The experiment object also provides methods to easily compare the prediction performance of its associated models and to pick the model with the best performance on a specific test dataset.

The following example demonstrates how the model tracking feature can be used for an experiment that trains a few binary-classifier models on the Titanic dataset. First, we must load the titanic data into our database and store it as a virtual dataframe (vDF):

from verticapy.datasets import load_titanic

titanic_vDF = load_titanic()

predictors = ["age", "fare", "pclass"]

response = "survived"

We then define a vExperiment() object to track the candidate models. To define the experiment object, specify the following parameters:

  • experiment_name: The name of the experiment.

  • test_relation: Relation or vDF to use to test the model.

  • X: List of the predictors.

  • y: Response column.

Note

If experiments_type is set to clustering, test_relation, X, and Y must be set to None.

The following parameters are optional:

  • experiment_type: By default auto, meaning VerticaPy tries to detect the experiment type from the response value. However, it might be cleaner to explicitly specify the experiment type. The other valid values for this parameter are regressor (for regression models), binary (for binary classification models), multi (for multiclass classification models), and clustering (for clustering models).

  • experiment_table: The name of the table ([schema_name.]table_name) in the database to archive the experiment. The experiment information won’t be backed up in the database without specifying this parameter. If the table already exists, its previously stored experiments are loaded to the object. In this case, the user must have SELECT, INSERT, and DELETE privileges on the table. If the table doesn’t exist and the user has the necessary privileges for creating such a table, the table is created.

import verticapy.mlops.model_tracking as mt

my_experiment_1 = mt.vExperiment(
    experiment_name = "my_exp_1",
    test_relation = titanic_vDF,
    X=predictors,
    y=response,
    experiment_type="binary",
    experiment_table="my_exp_table_1",
)

After creating the experiment object, we can train different models and add them to the experiment:

# training a LogisticRegression model
from verticapy.machine_learning.vertica import LogisticRegression

model_1 = LogisticRegression("logistic_reg_m", overwrite_model = True)

model_1.fit(titanic_vDF, predictors, response)


=======
details
=======
predictor|coefficient|std_err |z_value |p_value 
---------+-----------+--------+--------+--------
Intercept|  2.54895  | 0.39019| 6.53263| 0.00000
   age   | -0.03453  | 0.00563|-6.13772| 0.00000
  fare   |  0.00408  | 0.00174| 2.34950| 0.01880
 pclass  | -0.97692  | 0.11551|-8.45772| 0.00000


==============
regularization
==============
type| lambda 
----+--------
none| 1.00000


===========
call_string
===========
logistic_reg('public.logistic_reg_m', '"public"."_verticapy_tmp_view_v_mldb_86ce9b4a97be11efa8720242ac120002_"', '"survived"', '"age", "fare", "pclass"'
USING PARAMETERS optimizer='newton', epsilon=1e-06, max_iterations=100, regularization='none', lambda=1, alpha=0.5, fit_intercept=true)

===============
Additional Info
===============
       Name       |Value
------------------+-----
 iteration_count  |  4  
rejected_row_count| 238 
accepted_row_count| 996 


my_experiment_1.add_model(model_1)

# training a LinearSVC model
from verticapy.machine_learning.vertica import LinearSVC

model_2 = LinearSVC("svc_m", overwrite_model = True)

model_2.fit(titanic_vDF, predictors, response)


=======
details
=======
predictor|coefficient
---------+-----------
Intercept|  1.07602  
   age   | -0.01396  
  fare   |  0.00164  
 pclass  | -0.42318  


===========
call_string
===========
SELECT svm_classifier('public.svc_m', '"public"."_verticapy_tmp_view_v_mldb_8bc693be97be11efa8720242ac120002_"', '"survived"', '"age", "fare", "pclass"'
USING PARAMETERS class_weights='1,1', C=1, max_iterations=100, intercept_mode='regularized', intercept_scaling=1, epsilon=0.0001);

===============
Additional Info
===============
       Name       |Value
------------------+-----
accepted_row_count| 996 
rejected_row_count| 238 
 iteration_count  |  9  


my_experiment_1.add_model(model_2)

# training a DecisionTreeClassifier model
from verticapy.machine_learning.vertica import DecisionTreeClassifier

model_3 = DecisionTreeClassifier("tree_m", overwrite_model = True, max_depth = 3)

model_3.fit(titanic_vDF, predictors, response)


===========
call_string
===========
SELECT rf_classifier('public.tree_m', '"public"."_verticapy_tmp_view_v_mldb_9b3701a897be11efa8720242ac120002_"', 'survived', '"age", "fare", "pclass"' USING PARAMETERS exclude_columns='', ntree=1, mtry=2, sampling_size=1, max_depth=3, max_breadth=1000000000, min_leaf_size=1, min_info_gain=0, nbins=32);

=======
details
=======
predictor|      type      
---------+----------------
   age   |float or numeric
  fare   |float or numeric
 pclass  |      int       


===============
Additional Info
===============
       Name       |Value
------------------+-----
    tree_count    |  1  
rejected_row_count| 238 
accepted_row_count| 996 


my_experiment_1.add_model(model_3)

So far we have only added three models to the experiment, but we could add many more in a real scenario. Using the experiment object, we can easily list the models in the experiment and pick the one with the best prediction performance based on a specified metric.

my_experiment_1.list_models()
model_name
...
csi
user_defined_metrics
1logistic_reg_m...0.3765432098765432[null]
2svc_m...0.3796680497925311[null]
3tree_m...0.35555555555555557[null]
top_model = my_experiment_1.load_best_model(metric = "auc")

The experiment object facilitates not only model tracking but also makes cleanup super easy, especially in real-world scenarios where there is often a large number of leftover models. The drop() method drops from the database the info of the experiment and all associated models other than those specified in the keeping_models list.

my_experiment_1.drop(keeping_models = [top_model.model_name])

Experiments are also helpful for performing grid search on hyper-parameters. The following example shows how they can be used to study the impact of the max_iter parameter on the prediction performance of LogisticRegression models.

# creating an experiment
my_experiment_2 = mt.vExperiment(
    experiment_name = "my_exp_2",
    test_relation = titanic_vDF,
    X = predictors,
    y = response,
    experiment_type = "binary",
)


# training LogisticRegression with different values of max_iter
for i in range(1, 5):
    model = LogisticRegression(max_iter = i)
    model.fit(titanic_vDF, predictors, response)
    my_experiment_2.add_model(model)



=======
details
=======
predictor|coefficient|std_err |z_value |p_value 
---------+-----------+--------+--------+--------
Intercept|  2.16131  | 0.37528| 5.75922| 0.00000
   age   | -0.02788  | 0.00541|-5.15077| 0.00000
  fare   |  0.00323  | 0.00163| 1.98214| 0.04746
 pclass  | -0.84961  | 0.11103|-7.65177| 0.00000


==============
regularization
==============
type| lambda 
----+--------
none| 1.00000


===========
call_string
===========
logistic_reg('"public"."_verticapy_tmp_logisticregression_v_mldb_a15fc01a97be11efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_a1e6f1d497be11efa8720242ac120002_"', '"survived"', '"age", "fare", "pclass"'
USING PARAMETERS optimizer='newton', epsilon=1e-06, max_iterations=1, regularization='none', lambda=1, alpha=0.5, fit_intercept=true)

===============
Additional Info
===============
       Name       |Value
------------------+-----
 iteration_count  |  1  
rejected_row_count| 238 
accepted_row_count| 996 



=======
details
=======
predictor|coefficient|std_err |z_value |p_value 
---------+-----------+--------+--------+--------
Intercept|  2.52830  | 0.38924| 6.49549| 0.00000
   age   | -0.03413  | 0.00561|-6.08088| 0.00000
  fare   |  0.00401  | 0.00173| 2.32080| 0.02030
 pclass  | -0.97059  | 0.11523|-8.42312| 0.00000


==============
regularization
==============
type| lambda 
----+--------
none| 1.00000


===========
call_string
===========
logistic_reg('"public"."_verticapy_tmp_logisticregression_v_mldb_a59e598e97be11efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_a624c21c97be11efa8720242ac120002_"', '"survived"', '"age", "fare", "pclass"'
USING PARAMETERS optimizer='newton', epsilon=1e-06, max_iterations=2, regularization='none', lambda=1, alpha=0.5, fit_intercept=true)

===============
Additional Info
===============
       Name       |Value
------------------+-----
 iteration_count  |  2  
rejected_row_count| 238 
accepted_row_count| 996 



=======
details
=======
predictor|coefficient|std_err |z_value |p_value 
---------+-----------+--------+--------+--------
Intercept|  2.54890  | 0.39018| 6.53255| 0.00000
   age   | -0.03453  | 0.00563|-6.13752| 0.00000
  fare   |  0.00408  | 0.00174| 2.34932| 0.01881
 pclass  | -0.97690  | 0.11550|-8.45766| 0.00000


==============
regularization
==============
type| lambda 
----+--------
none| 1.00000


===========
call_string
===========
logistic_reg('"public"."_verticapy_tmp_logisticregression_v_mldb_aa27b6da97be11efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_aa99ae3e97be11efa8720242ac120002_"', '"survived"', '"age", "fare", "pclass"'
USING PARAMETERS optimizer='newton', epsilon=1e-06, max_iterations=3, regularization='none', lambda=1, alpha=0.5, fit_intercept=true)

===============
Additional Info
===============
       Name       |Value
------------------+-----
 iteration_count  |  3  
rejected_row_count| 238 
accepted_row_count| 996 



=======
details
=======
predictor|coefficient|std_err |z_value |p_value 
---------+-----------+--------+--------+--------
Intercept|  2.54895  | 0.39019| 6.53263| 0.00000
   age   | -0.03453  | 0.00563|-6.13772| 0.00000
  fare   |  0.00408  | 0.00174| 2.34950| 0.01880
 pclass  | -0.97692  | 0.11551|-8.45772| 0.00000


==============
regularization
==============
type| lambda 
----+--------
none| 1.00000


===========
call_string
===========
logistic_reg('"public"."_verticapy_tmp_logisticregression_v_mldb_aec99eba97be11efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_af4a303e97be11efa8720242ac120002_"', '"survived"', '"age", "fare", "pclass"'
USING PARAMETERS optimizer='newton', epsilon=1e-06, max_iterations=4, regularization='none', lambda=1, alpha=0.5, fit_intercept=true)

===============
Additional Info
===============
       Name       |Value
------------------+-----
 iteration_count  |  4  
rejected_row_count| 238 
accepted_row_count| 996 


# plotting prc_auc vs max_iter
my_experiment_2.plot("max_iter", "prc_auc")
Out[23]: <Axes: xlabel='parameter=max_iter', ylabel='metric=prc_auc'>

# cleaning all the models associated to the experiment from the database
my_experiment_2.drop()
_images/my_experiment_2_plot_max_iter_prc.png

Model Versioning

In Vertica version 12.0.4, we added support for In-DB ML Model Versioning. Now, we have integrated it into VerticaPy so that users can utilize its capabilities along with the other tools in VerticaPy. In VerticaPy, model versioning is a wrapper around an SQL API already built in Vertica. For more information about the concepts of model versioning in Vertica, see the Vertica documentation.

To showcase model versioning, we will begin by registering the top_model picked from the above experiment.

top_model.register("top_model_demo")
Out[25]: True

When the model owner registers the model, its ownership changes to DBADMIN, and the previous owner receives USAGE privileges. Registered models are referred to by their registered_name and version. Only DBADMIN or a user with the MLSUPERVISOR role can change the status of a registered model. We have provided the RegisteredModel() class in VerticaPy for working with registered models.

We will now make a RegisteredModel() object for our recently registered model and change its status to production. We can then use the registered model for scoring.

import verticapy.mlops.model_versioning as mv

rm = mv.RegisteredModel("top_model_demo")

To see the list of all models registered as top_model_demo, use the list_models() method.

rm.list_models()
Abc
registered_name
Varchar(128)
...
Abc
model_type
Varchar(128)
Abc
category
Varchar(128)
1top_model_demo...SVM_CLASSIFIERVERTICA_MODELS

The model we just registered has a status of under_review. The next step is to change the status of the model to staging, which is meant for A/B testing the model. Assuming the model performs well, we will promote it to the “production” status. Please note that we should specify the right version of the registered model from the above table.

# Getting the current version
version = rm.list_models()["registered_version"][0]

# changing the status of the model to staging
rm.change_status(version = version, new_status = "staging")

# changing the status of the model to production
rm.change_status(version = version, new_status = "production")

There can only be one version of the registered model in production at any time. The following predict function applies to the model with production status by default.

If you want to run the predict function on a model with a status other than “production”, you must also specify the model version.

rm.predict(
    titanic_vDF,
    X = predictors,
    name = "predicted_value",
)
123
pclass
Int
100%
...
Abc
home.dest
Varchar(100)
57%
123
predicted_value
Integer
80%
11...Montevideo, Uruguay0
21...Trenton, NJ1
31...[null][null]
41...Montevideo, Uruguay1
51...Los Angeles, CA1
61...Lakewood, NJ1
71...Montreal, PQ1
81...Deephaven, MN / Cedar Rapids, IA1
91...New York, NY1
101...Scituate, MA1
111...[null]1
121...New York, NY1
131...[null]1
141...London / Middlesex1
151...Brighton, MA[null]
161...New York, NY1
171...New York, NY[null]
181...Springfield, MA1
191...Vancouver, BC1
201...Dorchester, MA0

DBADMIN and users who are granted SELECT privileges on the v_monitor.model_status_history table are able to monitor the status history of registered models.

rm.list_status_history()
Abc
registered_name
Varchar(128)
...
Abc
schema_name
Varchar(128)
Abc
model_name
Varchar(128)
1top_model_demo...[null][null]
2top_model_demo...[null][null]
3top_model_demo...[null][null]
4top_model_demo...[null][null]
5top_model_demo...[null][null]
6top_model_demo...[null][null]
7top_model_demo...[null][null]
8top_model_demo...[null][null]
9top_model_demo...[null][null]
10top_model_demo...[null][null]
11top_model_demo...[null][null]
12top_model_demo...[null][null]
13top_model_demo...publicsvc_m
14top_model_demo...publicsvc_m
15top_model_demo...publicsvc_m
16top_model_demo...[null][null]
17top_model_demo...[null][null]
18top_model_demo...[null][null]

Conclusion

The addition of model tracking and model versioning to the VerticaPy toolkit greatly improves VerticaPy’s MLOps capabilities. We are constantly working to improve VerticaPy and address the needs of data scientists who wish to harness the power of Vertica database to empower their data analyses. If you have any comments or questions, don’t hesitate to reach out in the VerticaPy github community.

This concludes the fundamental lessons on machine learning algorithms in VerticaPy.