Loading...

XGBoost.to_json

Connect to Vertica

For a demonstration on how to create a new connection to Vertica, see Connection. In this example, we will use an existing connection named VerticaDSN.

import verticapy as vp

vp.connect("VerticaDSN")

Create a Schema (Optional)

Schemas allow you to organize database objects in a collection, similar to a namespace. If you create a database object without specifying a schema, Vertica uses the public schema. For example, to specify the example_table in example_schema, you would use: example_schema.example_table.

To keep things organized, this example creates the xgb_to_json schema and drops it (and its associated tables, views, etc.) at the end:

vp.drop("xgb_to_json", method = "schema")
Out[1]: False

vp.create_schema("xgb_to_json")
Out[2]: True

Load Data

VerticaPy lets you load many well-known datasets like Iris, Titanic, Amazon, etc. For a full list, check out datasets.

from verticapy.datasets import load_titanic

vdf = load_titanic(
    name = "titanic",
    schema = "xgb_to_json",
)

You can also load your own data. To ingest data from a CSV file, use the read_csv() function.

Create a vDataFrame

vDataFrames allow you to prepare and explore your data without modifying its representation in your Vertica database. Any changes you make are applied to the vDataFrame as modifications to the SQL query for the table underneath.

To create a vDataFrame out of a table in your Vertica database, specify its schema and table name with the standard SQL syntax. For example, to create a vDataFrame out of the titanic table in the xgb_to_json schema:

vdf = vp.vDataFrame("xgb_to_json.titanic")

Create an XGB model

Create a XGBClassifier model.

Unlike a vDataFrame object, which simply queries the table it was created with, the VerticaPy XGBClassifier object creates and then references a model in Vertica, so it must be stored in a schema like any other database object.

This example creates the my_model XGBClassifier model in the xgb_to_json schema:

This example loads the Titanic dataset with the load_titanic function into a table called titanic in the xgb_to_json schema:

from verticapy.machine_learning.vertica.ensemble import XGBClassifier

model = XGBClassifier(
    "xgb_to_json.my_model",
    max_ntree = 4,
    max_depth = 3,
)

Prepare the Data

While Vertica XGBoost supports columns of type VARCHAR, Python XGBoost does not, so you must encode the categorical columns you want to use. You must also drop or impute missing values.

This example drops age, fare, sex, embarked and survived columns from the vDataFrame and then encodes the sex and embarked columns. These changes are applied to the vDataFrame’s query and does not affect the main xgb_to_json.titanic table stored in Vertica:

vdf = vdf[["age", "fare", "sex", "embarked", "survived"]];

vdf.dropna();

vdf["sex"].label_encode();

vdf["embarked"].label_encode();
123
age
Numeric(8)
100%
...
123
embarked
Int
100%
123
survived
Integer
100%
171.0...00
245.0...20
317.0...20
427.0...00
537.0...00
631.0...20
750.0...00
836.0...00
937.0...20
1024.0...00
1145.0...20
1240.0...20
1342.0...20
1442.0...20
1529.0...20
1646.0...00
1754.0...20
1847.0...20
1958.0...00
2045.5...20

Split your data into training and testing:

train, test = vdf.train_test_split(0.05);

Train the Model

Define the predictor and the response columns:

relation = train;

X = ["age", "fare", "sex", "embarked"]

y = "survived"

Train the model with fit():

model.fit(relation, X, y)


===========
call_string
===========
xgb_classifier('xgb_to_json.my_model', '"xgb_to_json"."_verticapy_tmp_view_v_mldb_b81e87ba97bd11efa8720242ac120002_"', '"survived"', '"age", "fare", "sex", "embarked"' USING PARAMETERS exclude_columns='', max_ntree=4, max_depth=3, learning_rate=0.1, min_split_loss=0, weight_reg=0, nbins=32, objective=crossentropy, sampling_size=1, col_sample_by_tree=1, col_sample_by_node=1, seed=2, id_column='_verticapy_tmp_id_column_v_mldb_b7f1700497bd11efa8720242ac120002_')

=======
details
=======
predictor|      type      
---------+----------------
   age   |float or numeric
  fare   |float or numeric
   sex   |      int       
embarked |      int       


==================
initial_prediction
==================
response_label| value  
--------------+--------
      0       | 0.00000
      1       | 0.00000


===============
Additional Info
===============
       Name       |Value
------------------+-----
    tree_count    |  4  
rejected_row_count|  0  
accepted_row_count| 944 

Evaluate the Model

Evaluate the model with report():

model.report()
value
auc0.8195284087945641
prc_auc0.8034688636511114
accuracy0.7817796610169492
log_loss0.253796817077216
precision0.7343283582089553
recall0.6776859504132231
f1_score0.7048710601719197
mcc0.5332824598386225
informedness0.5245017851808651
markedness0.5422101316079702
csi0.5442477876106194

Use to_json() to export the model to a JSON file. If you omit a filename, VerticaPy prints the model:

model.to_json()
Out[17]: '{"learner": {"attributes": {"scikit_learn": "{\\"use_label_encoder\\": true, \\"n_estimators\\": 4, \\"objective\\": \\"binary:logistic\\", \\"max_depth\\": 3, \\"learning_rate\\": 0.1, \\"verbosity\\": null, \\"booster\\": null, \\"tree_method\\": null, \\"gamma\\": null, \\"min_child_weight\\": null, \\"max_delta_step\\": null, \\"subsample\\": null, \\"colsample_bytree\\": 1.0, \\"colsample_bylevel\\": null, \\"colsample_bynode\\": 1.0, \\"reg_alpha\\": null, \\"reg_lambda\\": null, \\"scale_pos_weight\\": null, \\"base_score\\": null, \\"missing\\": NaN, \\"num_parallel_tree\\": null, \\"kwargs\\": {}, \\"random_state\\": null, \\"n_jobs\\": null, \\"monotone_constraints\\": null, \\"interaction_constraints\\": null, \\"importance_type\\": \\"gain\\", \\"gpu_id\\": null, \\"validate_parameters\\": null, \\"classes_\\": [0, 1], \\"n_classes_\\": 2, \\"_le\\": {\\"classes_\\": [0, 1]}, \\"_estimator_type\\": \\"classifier\\"}"}, "feature_names": [], "feature_types": [], "gradient_booster": {"model": {"trees": [{"base_weights": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "categories": [], "categories_nodes": [], "categories_segments": [], "categories_sizes": [], "default_left": [true, true, true, true, true, true, true], "id": 0, "left_children": [1, 3, 5, -1, -1, -1, -1], "loss_changes": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "parents": [855590443, 0, 0, 1, 1, 2, 2], "right_children": [2, 4, 6, -1, -1, -1, -1], "split_conditions": [0.03125, 48.030862, 10.28875, 0.052360500000000004, 0.188235, 0.00487805, -0.13239399999999998], "split_indices": [2, 1, 0, 0, 0, 0, 0], "split_type": [0, 0, 0, 0, 0, 0, 0], "sum_hessian": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "tree_param": {"num_deleted": "0", "num_feature": "4", "num_nodes": "7", "size_leaf_vector": "0"}}, {"base_weights": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "categories": [], "categories_nodes": [], "categories_segments": [], "categories_sizes": [], "default_left": [true, true, true, true, true, true, true], "id": 1, "left_children": [1, 3, 5, -1, -1, -1, -1], "loss_changes": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "parents": [220702537, 0, 0, 1, 1, 2, 2], "right_children": [2, 4, 6, -1, -1, -1, -1], "split_conditions": [0.03125, 48.030862, 10.28875, 0.047158000000000005, 0.170973, 0.004390270000000001, -0.11969700000000001], "split_indices": [2, 1, 0, 0, 0, 0, 0], "split_type": [0, 0, 0, 0, 0, 0, 0], "sum_hessian": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "tree_param": {"num_deleted": "0", "num_feature": "4", "num_nodes": "7", "size_leaf_vector": "0"}}, {"base_weights": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "categories": [], "categories_nodes": [], "categories_segments": [], "categories_sizes": [], "default_left": [true, true, true, true, true, true, true], "id": 2, "left_children": [1, 3, 5, -1, -1, -1, -1], "loss_changes": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "parents": [501392870, 0, 0, 1, 1, 2, 2], "right_children": [2, 4, 6, -1, -1, -1, -1], "split_conditions": [0.03125, 48.030862, 7.799062, 0.042522000000000004, 0.157675, 0.0172554, -0.108212], "split_indices": [2, 1, 0, 0, 0, 0, 0], "split_type": [0, 0, 0, 0, 0, 0, 0], "sum_hessian": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "tree_param": {"num_deleted": "0", "num_feature": "4", "num_nodes": "7", "size_leaf_vector": "0"}}, {"base_weights": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "categories": [], "categories_nodes": [], "categories_segments": [], "categories_sizes": [], "default_left": [true, true, true, true, true, true, true], "id": 3, "left_children": [1, 3, 5, -1, -1, -1, -1], "loss_changes": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "parents": [641943051, 0, 0, 1, 1, 2, 2], "right_children": [2, 4, 6, -1, -1, -1, -1], "split_conditions": [0.03125, 48.030862, 0.0625, 0.0383732, 0.14707, -0.0338697, -0.104911], "split_indices": [2, 1, 3, 0, 0, 0, 0], "split_type": [0, 0, 0, 0, 0, 0, 0], "sum_hessian": [0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0], "tree_param": {"num_deleted": "0", "num_feature": "4", "num_nodes": "7", "size_leaf_vector": "0"}}], "tree_info": [0, 0, 0, 0], "gbtree_model_param": {"num_trees": "4", "size_leaf_vector": "0"}}, "name": "gbtree"}, "learner_model_param": {"base_score": "3.845339E-01", "num_class": "0", "num_feature": "4"}, "objective": {"name": "binary:logistic", "reg_loss_param": {"scale_pos_weight": "1"}}}, "version": [1, 6, 2]}'

To export and save the model as a JSON file, specify a filename:

model.to_json("exported_xgb_model.json");

Unlike Python XGBoost, Vertica does not store some information like sum_hessian or loss_changes, and the exported model from to_json() replaces this information with a list of zeroes. These information are replaced by a list filled with zeros.

Make Predictions with an Exported Model

This exported model can be used with the Python XGBoost API right away, and exported models make identical predictions in Vertica and Python:

import pytest

import xgboost as xgb

model_python = xgb.XGBClassifier();

model_python.load_model("exported_xgb_model.json");

# Convert to numpy format
X_test = test["age","fare","sex","embarked"].to_numpy() ;

y_test_vertica = model.to_python(return_proba = True)(X_test);

y_test_python = model_python.predict_proba(X_test);

result = (y_test_vertica - y_test_python) ** 2;

result = result.sum() / len(result);

assert result == pytest.approx(0.0, abs = 1.0E-14)

For multiclass classifiers, the probabilities returned by the VerticaPy and the exported model may differ slightly because of normalization; while Vertica uses multinomial logistic regression, XGBoost Python uses Softmax. Again, this difference does not affect the model’s final predictions. Categorical predictors must be encoded.

Clean the Example Environment

Drop the xgb_to_json schema, using CASCADE to drop any database objects stored inside (the titanic table, the XGBClassifier model, etc.), then delete the exported_xgb_model.json file:

import os

os.remove("exported_xgb_model.json")

vp.drop("xgb_to_json", method = "schema")
Out[31]: True

Conclusion

VerticaPy lets you to create, train, evaluate, and export Vertica machine learning models. There are some notable nuances when importing a Vertica XGBoost model into Python XGBoost, but these do not affect the accuracy of the model or its predictions:

Some information computed during the training phase may not be stored (e.g. sum_hessian and loss_changes).

The exact probabilities of multiclass classifiers in a Vertica model may differ from those in Python, but bot h will make the same predictions. Python XGBoost does not support categorical predictors, so you must encode them before training the model in VerticaPy.