Spam¶
This example uses the spam dataset to detect SMS spam. You can download the Jupyter Notebook of the study here.
v1: the SMS type (spam or ham).
v2: SMS content.
We will follow the data science cycle (Data Exploration - Data Preparation - Data Modeling - Model Evaluation - Model Deployment) to solve this problem.
Initialization¶
This example uses the following version of VerticaPy:
import verticapy as vp
vp.__version__
Out[2]: '1.1.0'
Connect to Vertica. This example uses an existing connection called VerticaDSN .
For details on how to create a connection, see the Connection tutorial.
You can skip the below cell if you already have an established connection.
vp.connect("VerticaDSN")
Let’s create a Virtual DataFrame of the dataset. The dataset is available here.
spam = vp.read_csv("spam.csv")
Let’s take a look at the first few entries in the dataset.
spam.head(10)
Abc type100% | Abc 100% | |
| 1 | ham | |
| 2 | ham | |
| 3 | ham | |
| 4 | ham | |
| 5 | ham | |
| 6 | ham | |
| 7 | ham | |
| 8 | ham | |
| 9 | ham | |
| 10 | ham |
Data Exploration and Preparation¶
Our dataset relies on text analysis. First, we should create some features. For example, we can use the SMS length and label encoding on the ‘type’ to get a dummy (1 if the message is a SPAM, 0 otherwise). We should also convert the message content to lowercase to simplify our analysis.
import verticapy.sql.functions as fun
spam["length"] = fun.length(spam["content"])
spam["content"].apply("LOWER({})")
spam["type"].decode('spam', 1, 0)
123 type100% | ... | Abc 100% | 123 length100% | |
| 1 | 0 | ... | 28 | |
| 2 | 0 | ... | 29 | |
| 3 | 0 | ... | 38 | |
| 4 | 0 | ... | 24 | |
| 5 | 0 | ... | 53 | |
| 6 | 0 | ... | 122 | |
| 7 | 0 | ... | 28 | |
| 8 | 0 | ... | 159 | |
| 9 | 0 | ... | 59 | |
| 10 | 0 | ... | 43 | |
| 11 | 0 | ... | 4 | |
| 12 | 0 | ... | 20 | |
| 13 | 0 | ... | 27 | |
| 14 | 0 | ... | 91 | |
| 15 | 0 | ... | 51 | |
| 16 | 0 | ... | 7 | |
| 17 | 0 | ... | 36 | |
| 18 | 0 | ... | 51 | |
| 19 | 0 | ... | 60 | |
| 20 | 0 | ... | 37 |
Let’s compute some statistics using the length of the message.
spam['type'].describe(
method = 'cat_stats',
numcol = 'length',
)
| ... | approx_90% | max | |
| 0 | ... | 125.0 | 911 |
| 1 | ... | 161.0 | 184 |
Note
Spam tends to be longer than a normal message. First, let’s create a view with just spam. Then, we’ll use the CountVectorizer to create a dictionary and identify keywords.
spams = spam.search(spam["type"] == 1)
from verticapy.machine_learning.vertica import CountVectorizer
dict_spams = CountVectorizer()
dict_spams.fit(spams, ["content"])
dict_spams = dict_spams.transform()
dict_spams
Abc token100% | ... | 123 df100% | 123 rnk100% | |
| 1 | to | ... | 0.028804653781197224 | 1 |
| 2 | call | ... | 0.019984987802589605 | 2 |
| 3 | a | ... | 0.01932820416588478 | 3 |
| 4 | you | ... | 0.01529367611184087 | 4 |
| 5 | your | ... | 0.0140739350722462 | 5 |
| 6 | for | ... | 0.012197410395946707 | 6 |
| 7 | free | ... | 0.010790016888722087 | 7 |
| 8 | or | ... | 0.010602364421092136 | 8 |
| 9 | now | ... | 0.010320885719647213 | 9 |
| 10 | the | ... | 0.009288797147682493 | 10 |
| 11 | txt | ... | 0.009194970913867517 | 11 |
| 12 | have | ... | 0.009101144680052542 | 12 |
| 13 | is | ... | 0.008538187277162695 | 13 |
| 14 | 2 | ... | 0.008256708575717772 | 14 |
| 15 | from | ... | 0.007693751172827923 | 15 |
| 16 | u | ... | 0.007599924939012948 | 16 |
| 17 | ur | ... | 0.007224620003753049 | 17 |
| 18 | 4 | ... | 0.0070369675361231 | 18 |
| 19 | and | ... | 0.0070369675361231 | 18 |
| 20 | on | ... | 0.006943141302308125 | 20 |
Let’s add the most occurent words in our vDataFrame and compute the correlation vector.
for word in dict_spams.head(200).values["token"]:
if word not in ['content', 'length', 'type'] : # because there is already a column called content, length and type
spam.regexp(
name = word,
pattern = word,
method = "count",
column = "content",
)
spam.corr(focus = "type")
Let’s just keep the first 100-most correlated features and merge the numbers together.
words = spam.corr(focus = "type", show = False)
spam.drop(columns = words["index"][101:])
for word in words["index"][1:101]:
if any(char.isdigit() for char in word):
spam[word].drop()
spam.regexp(
column = "content",
pattern = "([0-9])+",
method = "count",
name = "nb_numbers",
)
123 type100% | ... | 123 offer100% | 123 nb_numbers100% | |
| 1 | 0 | ... | 0 | 0 |
| 2 | 0 | ... | 0 | 0 |
| 3 | 0 | ... | 0 | 0 |
| 4 | 0 | ... | 0 | 0 |
| 5 | 0 | ... | 0 | 0 |
| 6 | 0 | ... | 0 | 2 |
| 7 | 0 | ... | 0 | 1 |
| 8 | 0 | ... | 0 | 9 |
| 9 | 0 | ... | 0 | 1 |
| 10 | 0 | ... | 0 | 2 |
| 11 | 0 | ... | 0 | 1 |
| 12 | 0 | ... | 0 | 0 |
| 13 | 0 | ... | 0 | 0 |
| 14 | 0 | ... | 0 | 0 |
| 15 | 0 | ... | 0 | 0 |
| 16 | 0 | ... | 0 | 0 |
| 17 | 0 | ... | 0 | 0 |
| 18 | 0 | ... | 0 | 0 |
| 19 | 0 | ... | 0 | 0 |
| 20 | 0 | ... | 0 | 0 |
Let’s narrow down our keyword list to words of more than two characters.
columns = spam.get_columns()
for word in columns:
if len(word.replace('"', '')) <= 2:
spam[word].drop()
Compute the correlation vector again using the response column.
spam.corr(focus = "type")
We have enough correlated features with our response to create a fantastic model.
Machine Learning¶
The Naive Bayes classifier is a powerful and performant algorithm for text analytics and binary classification. Before using it on our data, let’s use a cross-validation to test the efficiency of our model.
from verticapy.machine_learning.vertica import MultinomialNB
model = MultinomialNB()
from verticapy.machine_learning.model_selection import cross_validate
cross_validate(
model,
spam,
spam.get_columns(exclude_columns = ["type", "content"]),
"type",
cv = 5,
)
| ... | csi | time | |
| 1-fold | ... | 0.8 | 81.5515558719635 |
| 2-fold | ... | 0.7786885245901639 | 78.78958344459534 |
| 3-fold | ... | 0.8053097345132744 | 80.03940558433533 |
| 4-fold | ... | 0.7863247863247863 | 78.88041639328003 |
| 5-fold | ... | 0.8095238095238095 | 81.0278754234314 |
| avg | ... | 0.7959693709904068 | 80.05776734352112 |
| std | ... | 0.011652097069533811 | 1.1106121438325671 |
We have an excellent model! Let’s learn from the data.
model.fit(
spam,
spam.get_columns(exclude_columns = ["type", "content"]),
"type",
)
=======
details
=======
index|predictor | type
-----+----------+-----------
0 | type | ResponseI
1 | length |Multinomial
2 | call |Multinomial
3 | you |Multinomial
4 | your |Multinomial
5 | for |Multinomial
6 | free |Multinomial
7 | now |Multinomial
8 | txt |Multinomial
9 | from |Multinomial
10 | claim |Multinomial
11 | mobile |Multinomial
12 | stop |Multinomial
13 | reply |Multinomial
14 | only |Multinomial
15 | our |Multinomial
16 | text |Multinomial
17 | new |Multinomial
18 | send |Multinomial
19 | prize |Multinomial
20 | won |Multinomial
21 |guaranteed|Multinomial
22 | contact |Multinomial
23 | service |Multinomial
24 | win |Multinomial
25 | cash |Multinomial
26 | please |Multinomial
27 | draw |Multinomial
28 | customer |Multinomial
29 | urgent |Multinomial
30 | phone |Multinomial
31 | line |Multinomial
32 | shows |Multinomial
33 | week |Multinomial
34 | chat |Multinomial
35 | code |Multinomial
36 | awarded |Multinomial
37 | camera |Multinomial
38 | nokia |Multinomial
39 | tone |Multinomial
40 | video |Multinomial
41 | mins |Multinomial
42 | receive |Multinomial
43 | valid |Multinomial
44 | per |Multinomial
45 | msg |Multinomial
46 | landline |Multinomial
47 | selected |Multinomial
48 | latest |Multinomial
49 | weekly |Multinomial
50 | network |Multinomial
51 | entry |Multinomial
52 | private |Multinomial
53 | live |Multinomial
54 | todays |Multinomial
55 | delivery |Multinomial
56 | award |Multinomial
57 | expires |Multinomial
58 | mob |Multinomial
59 |statement |Multinomial
60 | price |Multinomial
61 | colour |Multinomial
62 |identifier|Multinomial
63 | ringtone |Multinomial
64 | land |Multinomial
65 | points |Multinomial
66 | orange |Multinomial
67 | gift |Multinomial
68 |camcorder |Multinomial
69 | apply |Multinomial
70 | sms |Multinomial
71 | pounds |Multinomial
72 | club |Multinomial
73 | all |Multinomial
74 | offer |Multinomial
75 |nb_numbers|Multinomial
=====
prior
=====
class|probability
-----+-----------
0 | 0.88120
1 | 0.11880
===========
call_string
===========
naive_bayes('"public"."_verticapy_tmp_naivebayes_v_mldb_1ec6015e97b511efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_3130faf097b611efa8720242ac120002_"', '"type"', '"length", "call", "you", "your", "for", "free", "now", "txt", "from", "claim", "mobile", "stop", "reply", "only", "our", "text", "new", "send", "prize", "won", "guaranteed", "contact", "service", "win", "cash", "please", "draw", "customer", "urgent", "phone", "line", "shows", "week", "chat", "code", "awarded", "camera", "nokia", "tone", "video", "mins", "receive", "valid", "per", "msg", "landline", "selected", "latest", "weekly", "network", "entry", "private", "live", "todays", "delivery", "award", "expires", "mob", "statement", "price", "colour", "identifier", "ringtone", "land", "points", "orange", "gift", "camcorder", "apply", "sms", "pounds", "club", "all", "offer", "nb_numbers"' USING PARAMETERS exclude_columns='', alpha=1)
=============
multinomial.0
=============
index|probability
-----+-----------
1 | 0.97469
2 | 0.00089
3 | 0.00656
4 | 0.00129
5 | 0.00199
6 | 0.00018
7 | 0.00175
8 | 0.00006
9 | 0.00045
10 | 0.00000
11 | 0.00003
12 | 0.00014
13 | 0.00014
14 | 0.00048
15 | 0.00176
16 | 0.00021
17 | 0.00028
18 | 0.00044
19 | 0.00000
20 | 0.00028
21 | 0.00000
22 | 0.00003
23 | 0.00001
24 | 0.00016
25 | 0.00003
26 | 0.00024
27 | 0.00002
28 | 0.00003
29 | 0.00004
30 | 0.00027
31 | 0.00013
32 | 0.00002
33 | 0.00032
34 | 0.00006
35 | 0.00001
36 | 0.00000
37 | 0.00001
38 | 0.00001
39 | 0.00001
40 | 0.00001
41 | 0.00003
42 | 0.00002
43 | 0.00000
44 | 0.00040
45 | 0.00020
46 | 0.00001
47 | 0.00002
48 | 0.00001
49 | 0.00000
50 | 0.00003
51 | 0.00000
52 | 0.00001
53 | 0.00011
54 | 0.00002
55 | 0.00000
56 | 0.00001
57 | 0.00000
58 | 0.00003
59 | 0.00000
60 | 0.00002
61 | 0.00003
62 | 0.00000
63 | 0.00000
64 | 0.00004
65 | 0.00001
66 | 0.00001
67 | 0.00003
68 | 0.00000
69 | 0.00002
70 | 0.00006
71 | 0.00001
72 | 0.00001
73 | 0.00238
74 | 0.00003
75 | 0.00336
=============
multinomial.1
=============
index|probability
-----+-----------
1 | 0.90678
2 | 0.00363
3 | 0.00512
4 | 0.00237
5 | 0.00224
6 | 0.00244
7 | 0.00191
8 | 0.00212
9 | 0.00121
10 | 0.00104
11 | 0.00149
12 | 0.00140
13 | 0.00109
14 | 0.00093
15 | 0.00358
16 | 0.00123
17 | 0.00089
18 | 0.00076
19 | 0.00077
20 | 0.00064
21 | 0.00060
22 | 0.00057
23 | 0.00064
24 | 0.00107
25 | 0.00077
26 | 0.00052
27 | 0.00048
28 | 0.00051
29 | 0.00045
30 | 0.00068
31 | 0.00087
32 | 0.00041
33 | 0.00107
34 | 0.00055
35 | 0.00047
36 | 0.00037
37 | 0.00043
38 | 0.00068
39 | 0.00112
40 | 0.00043
41 | 0.00041
42 | 0.00044
43 | 0.00036
44 | 0.00073
45 | 0.00095
46 | 0.00033
47 | 0.00029
48 | 0.00032
49 | 0.00029
50 | 0.00032
51 | 0.00035
52 | 0.00025
53 | 0.00068
54 | 0.00024
55 | 0.00027
56 | 0.00075
57 | 0.00024
58 | 0.00200
59 | 0.00023
60 | 0.00023
61 | 0.00024
62 | 0.00023
63 | 0.00037
64 | 0.00083
65 | 0.00023
66 | 0.00023
67 | 0.00025
68 | 0.00020
69 | 0.00020
70 | 0.00044
71 | 0.00020
72 | 0.00029
73 | 0.00403
74 | 0.00029
75 | 0.02793
===============
Additional Info
===============
Name | Value
------------------+--------
alpha | 1.00000
accepted_row_count| 4192
rejected_row_count| 0
model.confusion_matrix()
Out[4]:
array([[3641, 53],
[ 64, 434]])
Our model can reliably identify spam.
Conclusion¶
We’ve solved our problem in a Pandas-like way, all without ever loading data into memory!