Loading...

Spam

This example uses the spam dataset to detect SMS spam. You can download the Jupyter Notebook of the study here.

  • v1: the SMS type (spam or ham).

  • v2: SMS content.

We will follow the data science cycle (Data Exploration - Data Preparation - Data Modeling - Model Evaluation - Model Deployment) to solve this problem.

Initialization

This example uses the following version of VerticaPy:

import verticapy as vp

vp.__version__
Out[2]: '1.1.0'

Connect to Vertica. This example uses an existing connection called VerticaDSN . For details on how to create a connection, see the Connection tutorial. You can skip the below cell if you already have an established connection.

vp.connect("VerticaDSN")

Let’s create a Virtual DataFrame of the dataset. The dataset is available here.

spam = vp.read_csv("spam.csv")

Let’s take a look at the first few entries in the dataset.

spam.head(10)
Abc
type
Varchar(20)
100%
Abc
Varchar(1820)
100%
1ham
2ham
3ham
4ham
5ham
6ham
7ham
8ham
9ham
10ham

Data Exploration and Preparation

Our dataset relies on text analysis. First, we should create some features. For example, we can use the SMS length and label encoding on the ‘type’ to get a dummy (1 if the message is a SPAM, 0 otherwise). We should also convert the message content to lowercase to simplify our analysis.

import verticapy.sql.functions as fun

spam["length"] = fun.length(spam["content"])
spam["content"].apply("LOWER({})")
spam["type"].decode('spam', 1, 0)
123
type
Integer
100%
...
Abc
Varchar(3640)
100%
123
length
Integer
100%
10...28
20...29
30...38
40...24
50...53
60...122
70...28
80...159
90...59
100...43
110...4
120...20
130...27
140...91
150...51
160...7
170...36
180...51
190...60
200...37

Let’s compute some statistics using the length of the message.

spam['type'].describe(
    method = 'cat_stats',
    numcol = 'length',
)
...
approx_90%
max
0...125.0911
1...161.0184

Note

Spam tends to be longer than a normal message. First, let’s create a view with just spam. Then, we’ll use the CountVectorizer to create a dictionary and identify keywords.

spams = spam.search(spam["type"] == 1)

from verticapy.machine_learning.vertica import CountVectorizer

dict_spams = CountVectorizer()
dict_spams.fit(spams, ["content"])
dict_spams = dict_spams.transform()
dict_spams
Abc
token
Varchar(128)
100%
...
123
df
Numeric(38)
100%
123
rnk
Integer
100%
1to...0.0288046537811972241
2call...0.0199849878025896052
3a...0.019328204165884783
4you...0.015293676111840874
5your...0.01407393507224625
6for...0.0121974103959467076
7free...0.0107900168887220877
8or...0.0106023644210921368
9now...0.0103208857196472139
10the...0.00928879714768249310
11txt...0.00919497091386751711
12have...0.00910114468005254212
13is...0.00853818727716269513
142...0.00825670857571777214
15from...0.00769375117282792315
16u...0.00759992493901294816
17ur...0.00722462000375304917
184...0.007036967536123118
19and...0.007036967536123118
20on...0.00694314130230812520

Let’s add the most occurent words in our vDataFrame and compute the correlation vector.

for word in dict_spams.head(200).values["token"]:
    if word not in ['content', 'length', 'type'] : # because there is already a column called content, length and type
        spam.regexp(
            name = word,
            pattern = word,
            method = "count",
            column = "content",
        )
spam.corr(focus = "type")

Let’s just keep the first 100-most correlated features and merge the numbers together.

words = spam.corr(focus = "type", show = False)
spam.drop(columns = words["index"][101:])

for word in words["index"][1:101]:
    if any(char.isdigit() for char in word):
        spam[word].drop()

spam.regexp(
    column = "content",
    pattern = "([0-9])+",
    method = "count",
    name = "nb_numbers",
)
123
type
Integer
100%
...
123
offer
Integer
100%
123
nb_numbers
Integer
100%
10...00
20...00
30...00
40...00
50...00
60...02
70...01
80...09
90...01
100...02
110...01
120...00
130...00
140...00
150...00
160...00
170...00
180...00
190...00
200...00

Let’s narrow down our keyword list to words of more than two characters.

columns = spam.get_columns()
for word in columns:
    if len(word.replace('"', '')) <= 2:
        spam[word].drop()

Compute the correlation vector again using the response column.

spam.corr(focus = "type")

We have enough correlated features with our response to create a fantastic model.


Machine Learning

The Naive Bayes classifier is a powerful and performant algorithm for text analytics and binary classification. Before using it on our data, let’s use a cross-validation to test the efficiency of our model.

from verticapy.machine_learning.vertica import MultinomialNB

model = MultinomialNB()

from verticapy.machine_learning.model_selection import cross_validate

cross_validate(
    model,
    spam,
    spam.get_columns(exclude_columns = ["type", "content"]),
    "type",
    cv = 5,
)
...
csi
time
1-fold...0.881.5515558719635
2-fold...0.778688524590163978.78958344459534
3-fold...0.805309734513274480.03940558433533
4-fold...0.786324786324786378.88041639328003
5-fold...0.809523809523809581.0278754234314
avg...0.795969370990406880.05776734352112
std...0.0116520970695338111.1106121438325671

We have an excellent model! Let’s learn from the data.

model.fit(
    spam,
    spam.get_columns(exclude_columns = ["type", "content"]),
    "type",
)



=======
details
=======
index|predictor |   type    
-----+----------+-----------
  0  |   type   | ResponseI 
  1  |  length  |Multinomial
  2  |   call   |Multinomial
  3  |   you    |Multinomial
  4  |   your   |Multinomial
  5  |   for    |Multinomial
  6  |   free   |Multinomial
  7  |   now    |Multinomial
  8  |   txt    |Multinomial
  9  |   from   |Multinomial
 10  |  claim   |Multinomial
 11  |  mobile  |Multinomial
 12  |   stop   |Multinomial
 13  |  reply   |Multinomial
 14  |   only   |Multinomial
 15  |   our    |Multinomial
 16  |   text   |Multinomial
 17  |   new    |Multinomial
 18  |   send   |Multinomial
 19  |  prize   |Multinomial
 20  |   won    |Multinomial
 21  |guaranteed|Multinomial
 22  | contact  |Multinomial
 23  | service  |Multinomial
 24  |   win    |Multinomial
 25  |   cash   |Multinomial
 26  |  please  |Multinomial
 27  |   draw   |Multinomial
 28  | customer |Multinomial
 29  |  urgent  |Multinomial
 30  |  phone   |Multinomial
 31  |   line   |Multinomial
 32  |  shows   |Multinomial
 33  |   week   |Multinomial
 34  |   chat   |Multinomial
 35  |   code   |Multinomial
 36  | awarded  |Multinomial
 37  |  camera  |Multinomial
 38  |  nokia   |Multinomial
 39  |   tone   |Multinomial
 40  |  video   |Multinomial
 41  |   mins   |Multinomial
 42  | receive  |Multinomial
 43  |  valid   |Multinomial
 44  |   per    |Multinomial
 45  |   msg    |Multinomial
 46  | landline |Multinomial
 47  | selected |Multinomial
 48  |  latest  |Multinomial
 49  |  weekly  |Multinomial
 50  | network  |Multinomial
 51  |  entry   |Multinomial
 52  | private  |Multinomial
 53  |   live   |Multinomial
 54  |  todays  |Multinomial
 55  | delivery |Multinomial
 56  |  award   |Multinomial
 57  | expires  |Multinomial
 58  |   mob    |Multinomial
 59  |statement |Multinomial
 60  |  price   |Multinomial
 61  |  colour  |Multinomial
 62  |identifier|Multinomial
 63  | ringtone |Multinomial
 64  |   land   |Multinomial
 65  |  points  |Multinomial
 66  |  orange  |Multinomial
 67  |   gift   |Multinomial
 68  |camcorder |Multinomial
 69  |  apply   |Multinomial
 70  |   sms    |Multinomial
 71  |  pounds  |Multinomial
 72  |   club   |Multinomial
 73  |   all    |Multinomial
 74  |  offer   |Multinomial
 75  |nb_numbers|Multinomial


=====
prior
=====
class|probability
-----+-----------
  0  |  0.88120  
  1  |  0.11880  


===========
call_string
===========
naive_bayes('"public"."_verticapy_tmp_naivebayes_v_mldb_1ec6015e97b511efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_3130faf097b611efa8720242ac120002_"', '"type"', '"length", "call", "you", "your", "for", "free", "now", "txt", "from", "claim", "mobile", "stop", "reply", "only", "our", "text", "new", "send", "prize", "won", "guaranteed", "contact", "service", "win", "cash", "please", "draw", "customer", "urgent", "phone", "line", "shows", "week", "chat", "code", "awarded", "camera", "nokia", "tone", "video", "mins", "receive", "valid", "per", "msg", "landline", "selected", "latest", "weekly", "network", "entry", "private", "live", "todays", "delivery", "award", "expires", "mob", "statement", "price", "colour", "identifier", "ringtone", "land", "points", "orange", "gift", "camcorder", "apply", "sms", "pounds", "club", "all", "offer", "nb_numbers"' USING PARAMETERS exclude_columns='', alpha=1)

=============
multinomial.0
=============
index|probability
-----+-----------
  1  |  0.97469  
  2  |  0.00089  
  3  |  0.00656  
  4  |  0.00129  
  5  |  0.00199  
  6  |  0.00018  
  7  |  0.00175  
  8  |  0.00006  
  9  |  0.00045  
 10  |  0.00000  
 11  |  0.00003  
 12  |  0.00014  
 13  |  0.00014  
 14  |  0.00048  
 15  |  0.00176  
 16  |  0.00021  
 17  |  0.00028  
 18  |  0.00044  
 19  |  0.00000  
 20  |  0.00028  
 21  |  0.00000  
 22  |  0.00003  
 23  |  0.00001  
 24  |  0.00016  
 25  |  0.00003  
 26  |  0.00024  
 27  |  0.00002  
 28  |  0.00003  
 29  |  0.00004  
 30  |  0.00027  
 31  |  0.00013  
 32  |  0.00002  
 33  |  0.00032  
 34  |  0.00006  
 35  |  0.00001  
 36  |  0.00000  
 37  |  0.00001  
 38  |  0.00001  
 39  |  0.00001  
 40  |  0.00001  
 41  |  0.00003  
 42  |  0.00002  
 43  |  0.00000  
 44  |  0.00040  
 45  |  0.00020  
 46  |  0.00001  
 47  |  0.00002  
 48  |  0.00001  
 49  |  0.00000  
 50  |  0.00003  
 51  |  0.00000  
 52  |  0.00001  
 53  |  0.00011  
 54  |  0.00002  
 55  |  0.00000  
 56  |  0.00001  
 57  |  0.00000  
 58  |  0.00003  
 59  |  0.00000  
 60  |  0.00002  
 61  |  0.00003  
 62  |  0.00000  
 63  |  0.00000  
 64  |  0.00004  
 65  |  0.00001  
 66  |  0.00001  
 67  |  0.00003  
 68  |  0.00000  
 69  |  0.00002  
 70  |  0.00006  
 71  |  0.00001  
 72  |  0.00001  
 73  |  0.00238  
 74  |  0.00003  
 75  |  0.00336  


=============
multinomial.1
=============
index|probability
-----+-----------
  1  |  0.90678  
  2  |  0.00363  
  3  |  0.00512  
  4  |  0.00237  
  5  |  0.00224  
  6  |  0.00244  
  7  |  0.00191  
  8  |  0.00212  
  9  |  0.00121  
 10  |  0.00104  
 11  |  0.00149  
 12  |  0.00140  
 13  |  0.00109  
 14  |  0.00093  
 15  |  0.00358  
 16  |  0.00123  
 17  |  0.00089  
 18  |  0.00076  
 19  |  0.00077  
 20  |  0.00064  
 21  |  0.00060  
 22  |  0.00057  
 23  |  0.00064  
 24  |  0.00107  
 25  |  0.00077  
 26  |  0.00052  
 27  |  0.00048  
 28  |  0.00051  
 29  |  0.00045  
 30  |  0.00068  
 31  |  0.00087  
 32  |  0.00041  
 33  |  0.00107  
 34  |  0.00055  
 35  |  0.00047  
 36  |  0.00037  
 37  |  0.00043  
 38  |  0.00068  
 39  |  0.00112  
 40  |  0.00043  
 41  |  0.00041  
 42  |  0.00044  
 43  |  0.00036  
 44  |  0.00073  
 45  |  0.00095  
 46  |  0.00033  
 47  |  0.00029  
 48  |  0.00032  
 49  |  0.00029  
 50  |  0.00032  
 51  |  0.00035  
 52  |  0.00025  
 53  |  0.00068  
 54  |  0.00024  
 55  |  0.00027  
 56  |  0.00075  
 57  |  0.00024  
 58  |  0.00200  
 59  |  0.00023  
 60  |  0.00023  
 61  |  0.00024  
 62  |  0.00023  
 63  |  0.00037  
 64  |  0.00083  
 65  |  0.00023  
 66  |  0.00023  
 67  |  0.00025  
 68  |  0.00020  
 69  |  0.00020  
 70  |  0.00044  
 71  |  0.00020  
 72  |  0.00029  
 73  |  0.00403  
 74  |  0.00029  
 75  |  0.02793  


===============
Additional Info
===============
       Name       | Value  
------------------+--------
      alpha       | 1.00000
accepted_row_count|  4192  
rejected_row_count|   0    


model.confusion_matrix()
Out[4]: 
array([[3641,   53],
       [  64,  434]])

Our model can reliably identify spam.

Conclusion

We’ve solved our problem in a Pandas-like way, all without ever loading data into memory!