verticapy.machine_learning.vertica.ensemble.XGBClassifier.score¶
- XGBClassifier.score(metric: Literal['aic', 'bic', 'accuracy', 'acc', 'balanced_accuracy', 'ba', 'auc', 'roc_auc', 'prc_auc', 'best_cutoff', 'best_threshold', 'false_discovery_rate', 'fdr', 'false_omission_rate', 'for', 'false_negative_rate', 'fnr', 'false_positive_rate', 'fpr', 'recall', 'tpr', 'precision', 'ppv', 'specificity', 'tnr', 'negative_predictive_value', 'npv', 'negative_likelihood_ratio', 'lr-', 'positive_likelihood_ratio', 'lr+', 'diagnostic_odds_ratio', 'dor', 'log_loss', 'logloss', 'f1', 'f1_score', 'mcc', 'bm', 'informedness', 'mk', 'markedness', 'ts', 'csi', 'critical_success_index', 'fowlkes_mallows_index', 'fm', 'prevalence_threshold', 'pm', 'confusion_matrix', 'classification_report'] = 'accuracy', average: Literal[None, 'binary', 'micro', 'macro', 'scores', 'weighted'] | None = None, pos_label: Annotated[bool | float | str | timedelta | datetime, 'Python Scalar'] | None = None, cutoff: Annotated[int | float | Decimal, 'Python Numbers'] = 0.5, nbins: int = 10000) float | list[float]¶
Computes the model score.
Parameters¶
- metric: str, optional
The metric used to compute the score.
- accuracy:
Accuracy.
\[Accuracy = \frac{TP + TN}{TP + TN + FP + FN}\]
- aic:
Akaike’s Information Criterion
\[AIC = 2k - 2\ln(\hat{L})\]
- auc:
Area Under the Curve (ROC).
\[AUC = \int_{0}^{1} TPR(FPR) \, dFPR\]
- ba:
Balanced Accuracy.
\[BA = \frac{TPR + TNR}{2}\]
- best_cutoff:
Cutoff which optimised the ROC Curve prediction.
- bic:
Bayesian Information Criterion
\[BIC = -2\ln(\hat{L}) + k \ln(n)\]
- bm:
Informedness
\[BM = TPR + TNR - 1\]
- csi:
Critical Success Index
\[index = \frac{TP}{TP + FN + FP}\]
- f1:
F1 Score
\[F_1 Score = 2 \times \frac{Precision \times Recall}{Precision + Recall}\]
- fdr:
False Discovery Rate
\[FDR = 1 - PPV\]
- fm:
Fowlkes-Mallows index
\[FM = \sqrt{PPV * TPR}\]
- fnr:
False Negative Rate
\[FNR = \frac{FN}{FN + TP}\]
- for:
False Omission Rate
\[FOR = 1 - NPV\]
- fpr:
False Positive Rate
\[FPR = \frac{FP}{FP + TN}\]
- logloss:
Log Loss.
\[Loss = -\frac{1}{N} \sum_{i=1}^{N} \left( y_i \log(p_i) + (1 - y_i) \log(1 - p_i) \right)\]
- lr+:
Positive Likelihood Ratio.
\[LR+ = \frac{TPR}{FPR}\]
- lr-:
Negative Likelihood Ratio.
\[LR- = \frac{FNR}{TNR}\]
- dor:
Diagnostic Odds Ratio.
\[DOR = \frac{TP \times TN}{FP \times FN}\]
- mc:
Matthews Correlation Coefficient .. math:
MCC = \frac{TP \times TN - FP \times FN}{\sqrt{(TP + FP)(TP + FN)(TN + FP)(TN + FN)}}
- mk:
Markedness
\[MK = PPV + NPV - 1\]
- npv:
Negative Predictive Value
\[NPV = \frac{TN}{TN + FN}\]
- prc_auc:
Area Under the Curve (PRC)
\[AUC = \int_{0}^{1} Precision(Recall) \, dRecall\]
- precision:
Precision
\[Precision = TP / (TP + FP)\]
- pt:
Prevalence Threshold.
\[threshold = \frac{\sqrt{FPR}}{\sqrt{TPR} + \sqrt{FPR}}\]
- recall:
Recall.
\[Recall = \frac{TP}{TP + FN}\]
- specificity:
Specificity.
\[Specificity = \frac{TN}{TN + FP}\]
- average: str, optional
The method used to compute the final score for multiclass-classification.
- binary:
considers one of the classes as positive and use the binary confusion matrix to compute the score.
- micro:
positive and negative values globally.
- macro:
average of the score of each class.
- scores:
scores for all the classes.
- weighted:
weighted average of the score of each class.
If empty, the result will depend on the input metric. Whenever it is possible, the exact score is computed. Otherwise, the behaviour is similar to the ‘scores’ option.
- pos_label: PythonScalar, optional
Label to consider as positive. All the other classes will be merged and considered as negative for multiclass classification.
- cutoff: PythonNumber, optional
Cutoff for which the tested category is accepted as a prediction.
- nbins: int, optional
[Only when method is set to auc|prc_auc|best_cutoff] An integer value that determines the number of decision boundaries. Decision boundaries are set at equally spaced intervals between 0 and 1, inclusive. Greater values for nbins give more precise estimations of the AUC, but can potentially decrease performance. The maximum value is 999,999. If negative, the maximum value is used.
Returns¶
- float
score.
Examples¶
For this example, we will use the Iris dataset.
import verticapy.datasets as vpd data = vpd.load_iris() train, test = data.train_test_split(test_size = 0.2)
123SepalLengthCm123SepalWidthCm123PetalLengthCm123PetalWidthCmAbcSpecies1 4.6 3.6 1.0 0.2 Iris-setosa 2 4.7 3.2 1.3 0.2 Iris-setosa 3 4.7 3.2 1.6 0.2 Iris-setosa 4 4.8 3.0 1.4 0.1 Iris-setosa 5 4.8 3.1 1.6 0.2 Iris-setosa 6 4.8 3.4 1.9 0.2 Iris-setosa 7 4.9 3.0 1.4 0.2 Iris-setosa 8 4.9 3.1 1.5 0.1 Iris-setosa 9 4.9 3.1 1.5 0.1 Iris-setosa 10 4.9 3.1 1.5 0.1 Iris-setosa 11 5.0 2.3 3.3 1.0 Iris-versicolor 12 5.0 3.4 1.5 0.2 Iris-setosa 13 5.1 3.5 1.4 0.2 Iris-setosa 14 5.4 3.0 4.5 1.5 Iris-versicolor 15 5.4 3.4 1.5 0.4 Iris-setosa 16 5.4 3.9 1.3 0.4 Iris-setosa 17 5.5 2.4 3.7 1.0 Iris-versicolor 18 5.5 2.4 3.8 1.1 Iris-versicolor 19 5.6 2.7 4.2 1.3 Iris-versicolor 20 5.7 3.0 4.2 1.2 Iris-versicolor 21 5.7 4.4 1.5 0.4 Iris-setosa 22 5.8 2.8 5.1 2.4 Iris-virginica 23 5.9 3.2 4.8 1.8 Iris-versicolor 24 6.1 3.0 4.6 1.4 Iris-versicolor 25 6.1 3.0 4.9 1.8 Iris-virginica 26 6.3 2.5 4.9 1.5 Iris-versicolor 27 6.3 3.3 4.7 1.6 Iris-versicolor 28 6.3 3.3 6.0 2.5 Iris-virginica 29 6.4 2.9 4.3 1.3 Iris-versicolor 30 6.5 3.0 5.5 1.8 Iris-virginica 31 6.5 3.0 5.8 2.2 Iris-virginica 32 6.7 3.0 5.0 1.7 Iris-versicolor 33 6.8 2.8 4.8 1.4 Iris-versicolor 34 6.8 3.2 5.9 2.3 Iris-virginica 35 7.0 3.2 4.7 1.4 Iris-versicolor 36 7.1 3.0 5.9 2.1 Iris-virginica 37 7.7 3.8 6.7 2.2 Iris-virginica 38 4.4 2.9 1.4 0.2 Iris-setosa 39 4.5 2.3 1.3 0.3 Iris-setosa 40 4.8 3.4 1.6 0.2 Iris-setosa 41 5.0 2.0 3.5 1.0 Iris-versicolor 42 5.1 3.3 1.7 0.5 Iris-setosa 43 5.1 3.4 1.5 0.2 Iris-setosa 44 5.2 2.7 3.9 1.4 Iris-versicolor 45 5.2 3.5 1.5 0.2 Iris-setosa 46 5.2 4.1 1.5 0.1 Iris-setosa 47 5.4 3.9 1.7 0.4 Iris-setosa 48 5.5 3.5 1.3 0.2 Iris-setosa 49 5.6 3.0 4.1 1.3 Iris-versicolor 50 5.8 2.7 3.9 1.2 Iris-versicolor 51 5.8 2.7 5.1 1.9 Iris-virginica 52 5.8 2.7 5.1 1.9 Iris-virginica 53 5.9 3.0 4.2 1.5 Iris-versicolor 54 5.9 3.0 5.1 1.8 Iris-virginica 55 6.0 2.7 5.1 1.6 Iris-versicolor 56 6.0 2.9 4.5 1.5 Iris-versicolor 57 6.1 2.8 4.7 1.2 Iris-versicolor 58 6.2 2.8 4.8 1.8 Iris-virginica 59 6.2 2.9 4.3 1.3 Iris-versicolor 60 6.3 2.3 4.4 1.3 Iris-versicolor 61 6.3 2.7 4.9 1.8 Iris-virginica 62 6.4 3.2 5.3 2.3 Iris-virginica 63 6.5 2.8 4.6 1.5 Iris-versicolor 64 6.5 3.0 5.2 2.0 Iris-virginica 65 6.5 3.2 5.1 2.0 Iris-virginica 66 6.6 2.9 4.6 1.3 Iris-versicolor 67 6.6 3.0 4.4 1.4 Iris-versicolor 68 6.7 3.1 4.4 1.4 Iris-versicolor 69 6.7 3.1 4.7 1.5 Iris-versicolor 70 6.9 3.1 4.9 1.5 Iris-versicolor 71 6.9 3.1 5.4 2.1 Iris-virginica 72 6.9 3.2 5.7 2.3 Iris-virginica 73 7.2 3.0 5.8 1.6 Iris-virginica 74 7.2 3.2 6.0 1.8 Iris-virginica 75 7.3 2.9 6.3 1.8 Iris-virginica 76 7.7 2.6 6.9 2.3 Iris-virginica 77 3.3 4.5 5.6 7.8 Iris-setosa 78 3.3 4.5 5.6 7.8 Iris-setosa 79 3.3 4.5 5.6 7.8 Iris-setosa 80 3.3 4.5 5.6 7.8 Iris-setosa 81 3.3 4.5 5.6 7.8 Iris-setosa 82 3.3 4.5 5.6 7.8 Iris-setosa 83 3.3 4.5 5.6 7.8 Iris-setosa 84 3.3 4.5 5.6 7.8 Iris-setosa 85 3.3 4.5 5.6 7.8 Iris-setosa 86 3.3 4.5 5.6 7.8 Iris-setosa 87 3.3 4.5 5.6 7.8 Iris-setosa 88 3.3 4.5 5.6 7.8 Iris-setosa 89 3.3 4.5 5.6 7.8 Iris-setosa 90 3.3 4.5 5.6 7.8 Iris-setosa 91 3.3 4.5 5.6 7.8 Iris-setosa 92 3.3 4.5 5.6 7.8 Iris-setosa 93 3.3 4.5 5.6 7.8 Iris-setosa 94 3.3 4.5 5.6 7.8 Iris-setosa 95 3.3 4.5 5.6 7.8 Iris-setosa 96 3.3 4.5 5.6 7.8 Iris-setosa 97 3.3 4.5 5.6 7.8 Iris-setosa 98 3.3 4.5 5.6 7.8 Iris-setosa 99 3.3 4.5 5.6 7.8 Iris-setosa 100 3.3 4.5 5.6 7.8 Iris-setosa Rows: 1-100 | Columns: 5Let’s import the model:
from verticapy.machine_learning.vertica import NearestCentroid
Then we can create the model:
model = NearestCentroid(p = 2)
We can now fit the model:
model.fit( train, [ "SepalLengthCm", "SepalWidthCm", "PetalLengthCm", "PetalWidthCm", ], "Species", test, )
We can get the score:
model.score() Out[109]: 0.58
To get the score of a particular class:
model.score(pos_label= "Iris-setosa") Out[110]: 0.6
Important
For this example, a specific model is utilized, and it may not correspond exactly to the model you are working with. To see a comprehensive example specific to your class of interest, please refer to that particular class.