vDataFrame.score

In [ ]:
vDataFrame.score(y_true: str, 
                 y_score: str,
                 method: str,
                 nbins: int = 30,)

Computes the score using the input columns and the input method.

Parameters

Name Type Optional Description
y_true
str
Response column.
y_score
str
Prediction.
method
str
The method to use to compute the score.

For Classification:
  • accuracy : Accuracy
  • auc : Area Under the Curve (ROC)
  • best_cutoff : Cutoff which optimised the ROC Curve prediction.
  • bm : Informedness = tpr + tnr - 1
  • csi : Critical Success Index = tp / (tp + fn + fp)
  • f1 : F1 Score
  • logloss : Log Loss
  • mcc : Matthews Correlation Coefficient
  • mk : Markedness = ppv + npv - 1
  • npv : Negative Predictive Value = tn / (tn + fn)
  • prc_auc : Area Under the Curve (PRC)
  • precision : Precision = tp / (tp + fp)
  • recall : Recall = tp / (tp + fn)
  • specificity : Specificity = tn / (tn + fp)

For Regression:
  • max : Max Error
  • mae : Mean Absolute Error
  • median : Median Absolute Error
  • mse : Mean Squared Error
  • msle : Mean Squared Log Error
  • r2 : R-squared coefficient
  • var : Explained Variance

Plots:
  • roc : ROC Curve
  • prc : PRC Curve
  • lift : Lift Chart
nbins
int
[Only when method is set to auc|prc_auc|best_cutoff] An integer value that determines the number of decision boundaries. Decision boundaries are set at equally spaced intervals between 0 and 1, inclusive. Greater values for nbins give more precise estimations of the AUC, but can potentially decrease performance. The maximum value is 999,999. If negative, the maximum value is used.

Returns

float / tablesample : score / tablesample of the curve

Example

In [62]:
from verticapy.datasets import load_titanic
titanic = load_titanic().select(["age", "fare", "survived"])
display(titanic)
123
age
Numeric(6,3)
123
fare
Numeric(10,5)
123
survived
Int
12.000151.550000
230.000151.550000
325.000151.550000
439.0000.000000
571.00049.504200
647.000227.525000
7[null]25.925000
824.000247.520800
936.00075.241700
1025.00026.000000
1145.00035.500000
1242.00026.550000
1341.00030.500000
1448.00050.495800
15[null]39.600000
1645.00026.550000
17[null]31.000000
1833.0005.000000
1928.00047.100000
2017.00047.100000
2149.00026.000000
2236.00078.850000
2346.00061.175000
24[null]0.000000
2527.000136.779200
26[null]52.000000
2747.00025.587500
2837.00083.158300
29[null]26.550000
3070.00071.000000
3139.00071.283300
3231.00052.000000
3350.000106.425000
3439.00029.700000
3536.00031.679200
36[null]221.779200
3730.00027.750000
3819.000263.000000
3964.000263.000000
40[null]26.550000
41[null]0.000000
4237.00053.100000
4347.00038.500000
4424.00079.200000
4571.00034.654200
4638.000153.462500
4746.00079.200000
48[null]42.400000
4945.00083.475000
5040.0000.000000
5155.00093.500000
5242.00042.500000
53[null]51.862500
5455.00050.000000
5542.00052.000000
56[null]30.695800
5750.00028.712500
5846.00026.000000
5950.00026.000000
6032.500211.500000
6158.00029.700000
6241.00051.862500
63[null]26.550000
64[null]27.720800
6529.00030.000000
6630.00045.500000
6730.00026.000000
6819.00053.100000
6946.00075.241700
7054.00051.862500
7128.00082.170800
7265.00026.550000
7344.00090.000000
7455.00030.500000
7547.00042.400000
7637.00029.700000
7758.000113.275000
7864.00026.000000
7965.00061.979200
8028.50027.720800
81[null]0.000000
8245.50028.500000
8323.00093.500000
8429.00066.600000
8518.000108.900000
8647.00052.000000
8738.0000.000000
8822.000135.633300
89[null]227.525000
9031.00050.495800
91[null]50.000000
9236.00040.125000
9355.00059.400000
9433.00026.550000
9561.000262.375000
9650.00055.900000
9756.00026.550000
9856.00030.695800
9924.00060.000000
100[null]26.000000
Rows: 1-100 of 1234 | Columns: 3
In [63]:
from verticapy.learn.linear_model import LogisticRegression
model = LogisticRegression(name = "public.LR_titanic",
                           tol = 1e-4, 
                           C = 1.0, 
                           max_iter = 100, 
                           solver = 'CGD',
                           l1_ratio = 0.5)
model.fit("public.titanic", ["fare", "age"], "survived")
model.predict(titanic, name = "survived_pred")
123
age
Numeric(6,3)
123
fare
Numeric(10,5)
123
survived
Int
123
survived_pred
Float
12.000151.5500000.902287137987944
230.000151.5500000.86058029983458
325.000151.5500000.868988340490406
439.0000.0000000.342456823846109
571.00049.5042000.414029113989722
647.000227.5250000.939923121186975
7[null]25.925000[null]
824.000247.5208000.967395954658983
936.00075.2417000.635075649080458
1025.00026.0000000.487751263444919
1145.00035.5000000.452683907002755
1242.00026.5500000.429216761960316
1341.00030.5000000.447792486054408
1448.00050.4958000.49971322568723
15[null]39.600000[null]
1645.00026.5500000.418678018840286
17[null]31.000000[null]
1833.0005.0000000.380187437748189
1928.00047.1000000.558247754878829
2017.00047.1000000.596833670206284
2149.00026.0000000.402695575911559
2236.00078.8500000.647904226645164
2346.00061.1750000.548033233872829
24[null]0.000000[null]
2527.000136.7792000.836841341038647
26[null]52.000000[null]
2747.00025.5875000.408093270883223
2837.00083.1583000.659723543412694
29[null]26.550000[null]
3070.00071.0000000.499846009538692
3139.00071.2833000.610568047403653
3231.00052.0000000.566271351391256
3350.000106.4250000.697362224205815
3439.00029.7000000.451851620327145
3536.00031.6792000.470175980622248
36[null]221.779200[null]
3730.00027.7500000.476548622611
3819.000263.0000000.975906172743917
3964.000263.0000000.954958490722554
40[null]26.550000[null]
41[null]0.000000[null]
4237.00053.1000000.54917807997722
4347.00038.5000000.457050748370639
4424.00079.2000000.687374102239646
4571.00034.6542000.359641574716634
4638.000153.4625000.8500005835631
4746.00079.2000000.61571512557438
48[null]42.400000[null]
4945.00083.4750000.634571272671506
5040.0000.0000000.339224975430521
5155.00093.5000000.637150541510681
5242.00042.5000000.490387504173356
53[null]51.862500[null]
5455.00050.0000000.472650396971738
5542.00052.0000000.527078154230704
56[null]30.695800[null]
5750.00028.7125000.409339902436815
5846.00026.0000000.413117914049503
5950.00026.0000000.399240408375275
6032.500211.5000000.937672852935814
6158.00029.7000000.385443044656693
6241.00051.8625000.530132926520432
63[null]26.550000[null]
64[null]27.720800[null]
6529.00030.0000000.488825979510246
6630.00045.5000000.545014535857961
6730.00026.0000000.469804285663854
6819.00053.1000000.612131629679044
6946.00075.2417000.601136677890098
7054.00051.8625000.483424140863347
7128.00082.1708000.684873682886807
7265.00026.5500000.350713668608726
7344.00090.0000000.660862965877402
7455.00030.5000000.398676025965298
7547.00042.4000000.472047669134459
7637.00029.7000000.458986841891317
7758.000113.2750000.695421942724314
7864.00026.0000000.352054220262656
7965.00061.9792000.482967715783495
8028.50027.7208000.481820954229194
81[null]0.000000[null]
8245.50028.5000000.424275206647782
8323.00093.5000000.735622606643536
8429.00066.6000000.627415614513485
8518.000108.9000000.791394954471803
8647.00052.0000000.509122344734676
8738.0000.0000000.345703353743199
8822.000135.6333000.84410848047684
89[null]227.525000[null]
9031.00050.4958000.560551060575476
91[null]50.000000[null]
9236.00040.1250000.502784282039499
9355.00059.4000000.508953958690919
9433.00026.5500000.461182774006888
9561.000262.3750000.95637734057273
9650.00055.9000000.51340520455764
9756.00026.5500000.380732995250046
9856.00030.6958000.395956313027424
9924.00060.0000000.62034963876044
100[null]26.000000[null]
Out[63]:
Rows: 1-100 of 1234 | Columns: 4
In [64]:
# Computing AUC
titanic.score(y_true  = "survived", 
              y_score = "survived_pred",
              method  = "auc")
Out[64]:
0.6974762740166146
In [65]:
# Computing MSE
titanic.score(y_true  = "survived", 
              y_score = "survived_pred",
              method  = "mse")
Out[65]:
0.224993557369594
In [39]:
# Drawing ROC Curve
titanic.score(y_true  = "survived", 
              y_score = "survived_pred",
              method  = "roc")
In [40]:
# Drawing PRC Curve
titanic.score(y_true  = "survived", 
              y_score = "survived_pred",
              method  = "prc")

See Also

vDataFrame.aggregate Computes the vDataFrame input aggregations.