Loading...

Correlation and Dependency

Finding links between variables is a very important task. The main purpose of data science is to find relationships between variables, and to understand how these relationships can help us make better decisions.

Machine learning models are also sensitive to the number of variables and how they relate and affect each other, so finding correlations and dependencies can help us make better use of our machine learning algorithms.

Let’s use the Telco Churn dataset to understand how we can find links between different variables in VerticaPy.

import verticapy as vp

churn = vp.read_csv("customers.csv")
churn.head(100)
Abc
customerID
Varchar(20)
100%
...
123
TotalCharges
Numeric(9,3)
99%
010
Churn
Boolean
100%
10002-ORFBO...593.3
❌
20003-MKNFE...542.4
❌
30004-TLHLJ...280.85
✅
40017-DINOC...2460.55
❌
50020-JDNXP...1993.2
❌
60022-TCJCI...2791.5
✅
70023-XUOPT...1215.6
✅
80040-HALCW...1090.6
❌
90042-JVWOJ...471.85
❌
100064-SUDOG...224.5
❌
110080-EMYVY...727.85
❌
120083-PIVIK...5567.55
❌
130089-IIQKO...3767.4
❌
140093-EXYQL...3673.6
❌
150093-XWZFY...4036.85
✅
160103-CSITQ...6252.7
❌
170111-KLBQG...2861.45
❌
180115-TFERT...2317.1
✅
190117-LFRMW...1448.8
✅
200133-BMFZO...181.65
✅
210141-YEAYS...2401.05
❌
220156-FVPTA...1152.7
✅
230164-XAIRP...470.2
❌
240178-CIIKR...58.0
❌
250187-WZNAB...1972.35
❌
260188-GWFLE...33.7
❌
270206-OYVOC...864.2
❌
280218-QNVAS...7113.75
❌
290219-QAERP...576.65
❌
300221-NAUXK...219.5
❌
310233-FTHAV...4765.0
❌
320237-YFUTL...5405.8
❌
330254-FNMCI...7624.2
❌
340254-KCJGT...4354.45
❌
350254-WWRKD...3431.75
❌
360259-GBZSH...181.5
✅
370265-PSUAE...1522.7
❌
380266-GMEAO...8058.55
❌
390277-ORXQS...3364.55
❌
400282-NVSJS...355.9
❌
410307-BCOPK...326.65
❌
420310-MVLET...6010.05
✅
430320-JDNQG...2331.3
✅
440322-YINQP...48.55
✅
450324-BRPCJ...6851.65
✅
460330-IVZHA...330.15
✅
470345-XMMUG...4854.3
❌
480348-SDKOL...5885.4
✅
490354-WYROK...2911.3
✅
500357-NVCRI...471.7
❌
510363-QJVFX...3432.9
✅
520365-BZUWY...1742.5
❌
530369-ZGOVK...1992.2
❌
540376-OIWME...3366.05
❌
550376-YMCJC...1943.2
✅
560378-NHQXU...1460.65
✅
570379-DJQHR...5398.6
❌
580380-NEAVX...3343.15
❌
590383-CLDDA...5897.4
❌
600386-CWRGM...967.85
❌
610388-EOPEX...139.4
✅
620396-YCHWO...3474.2
❌
630411-EZJZE...170.5
❌
640412-UCCNP...3175.85
❌
650415-MOSGF...44.4
✅
660420-HLGXF...4036.0
❌
670422-OHQHQ...295.95
❌
680423-UDIJQ...447.9
❌
690431-APWVY...2598.95
✅
700436-TWFFZ...5714.2
❌
710439-IFYUN...1294.6
❌
720461-CVKMU...1900.25
✅
730464-WJTKO...1460.85
❌
740479-HMSWA...2715.3
❌
750480-BIXDE...1743.05
❌
760481-SUMCB...4735.35
❌
770484-JPBRU...996.45
❌
780506-LVNGN...349.65
✅
790513-RBGPE...2278.75
❌
800533-UCAAU...4140.1
❌
810537-QYZZN...1857.75
❌
820547-HURJB...235.8
❌
830549-CYCQN...3974.7
❌
840565-JUPYD...6590.8
❌
850567-GGCAC...438.9
❌
860575-CUQOV...5867.0
❌
870576-WNXXC...2510.2
✅
880577-WHMEV...1374.9
❌
890581-MDMPW...2072.75
❌
900601-WZHJF...667.7
✅
910602-DDUML...3894.4
❌
920603-TPMIB...1534.05
❌
930620-XEFWH...84.2
❌
940621-HJWXJ...5029.05
❌
950628-CNQRM...1544.05
✅
960635-WKOLD...2921.75
❌
970637-YLETY...1555.65
✅
980650-BWOZN...1359.45
❌
990655-YDGFJ...1323.7
❌
1000657-DOGUM...2985.25
❌

The Pearson correlation coefficient is a very common correlation function. In this case, it helped us to find linear links between the variables. Having a strong Pearson relationship means that the two input variables are linearly correlated.

churn.corr(method = "pearson")

We can see that tenure is well-correlated to the TotalCharges, which makes sense.

churn.scatter(["tenure", "TotalCharges"])
churn.corr(["tenure", "TotalCharges"], method = "pearson")
Out[1]: 0.825880460933202

Note, however, that having a low Pearson relationship imply that the variables aren’t correlated. For example, let’s compute the Pearson correlation coefficient between tenure and TotalCharges to the power of 20.

churn["TotalCharges^20"] = churn["TotalCharges"] ** 20

churn.corr(["tenure", "TotalCharges^20"], method = "pearson")
Out[3]: 0.224994408804537

We know that the tenure and TotalCharges are strongly linearly correlated. However we can notice that the correlation between the tenure and TotalCharges to the power of 20 is not very high. Indeed, the Pearson correlation coefficient is not robust for monotonic relationships, but rank-based correlations are. Knowing this, we’ll calculate the Spearman’s rank correlation coefficient instead.

churn.corr(method = "spearman", show = False)
...
"Churn"
"TotalCharges^20"
"SeniorCitizen"...0.1508893281764730.105795342303725
"Partner"...-0.1504475449591770.343930553215626
"Dependents"...-0.1642214015797250.0866797760484616
"tenure"...-0.3696207787634350.883103368818293
"PhoneService"...0.01194198002900310.0838048547856037
"PaperlessBilling"...0.1918253316664680.151669712799097
"MonthlyCharges"...0.1848392857837580.633958405301206
"TotalCharges"...-0.2332110185851041.0
"Churn"...1.0-0.233211018585104
"TotalCharges^20"...-0.2332110185851041.0
churn.corr(method = "spearman")

The Spearman’s rank correlation coefficient determines the monotonic relationships between the variables.

churn.corr(["tenure", "TotalCharges^20"], method = "spearman")
Out[4]: 0.883103368818293

We can notice that Spearman’s rank correlation coefficient stays the same if one of the variables can be expressed using a monotonic function on the other. The same applies to Kendall rank correlation coefficient.

churn.corr(method = "kendall")

Notice that the Kendall rank correlation coefficient will also detect the monotonic relationship.

churn.corr(["tenure", "TotalCharges^20"], method = "kendall")
Out[5]: 0.731699318287362

However, the Kendall rank correlation coefficient is very computationally expensive, so we’ll generally use Pearson and Spearman when dealing with correlations between numerical variables.

Binary features are considered numerical, but this isn’t technically accurate. Since binary variables can only take two values, calculating correlations between a binary and numerical variable can lead to misleading results. To account for this, we’ll want to use the Biserial Point method to calculate the Point-Biserial correlation coefficient. This powerful method will help us understand the link between a binary variable and a numerical variable.

churn.corr(method = "biserial")

Lastly, we’ll look at the relationship between categorical columns. In this case, the Cramer's V method is very efficient. Since there is no position in the Euclidean space for those variables, the Cramer's V coefficients cannot be negative (which is a sign of an opposite relationship) and they will range in the interval [0,1].

churn.corr(method = "cramer")

Sometimes, we just need to look at the correlation between a response and other variables. The parameter focus will isolate and show us the specified correlation vector.

churn.corr(method = "cramer", focus = "Churn")

Sometimes a correlation coefficient can lead to incorrect assumptions, so we should always look at the coefficient p-value.

churn.corr_pvalue("Churn", "customerID", method = "cramer")
Out[6]: (nan, nan)

We can see that churning correlates to the type of contract (monthly, yearly, etc.) which makes sense: you would expect that different types of contracts differ in flexibility for the customer, and particularly restrictive contracts may make churning more likely.

The type of internet service also seems to correlate with churning. Let’s split the different categories to binaries to understand which services can influence the global churning rate.

churn["InternetService"].one_hot_encode()
churn.corr(
    method = "spearman",
    focus = "Churn",
    columns = [
        "InternetService_DSL",
        "InternetService_Fiber_optic",
    ],
)

We can see that the Fiber Optic option in particular seems to be directly linked to a customer’s likelihood to churn. Let’s compute some aggregations to find a causal relationship.

churn["contract"].one_hot_encode()
churn.groupby(
    [
        "InternetService_Fiber_optic",
    ],
    [
        "AVG(tenure) AS tenure",
        "AVG(totalcharges) AS totalcharges",
        'AVG("contract_month-to-month") AS "contract_month-to-month"',
        'AVG("monthlycharges") AS "monthlycharges"',
    ],
)
123
InternetService_Fiber_optic
Integer
100%
...
123
contract_month-to-month
Float(22)
100%
123
monthlycharges
Float(22)
100%
10...0.44261464403344343.7882442361287
21...0.6873385012919991.5001291989664

It seems that users with the Fiber Optic option tend more to churn not because of the option itself, but probably because of the type of contracts and the monthly charges the users are paying to get it. Be careful when dealing with identifying correlations! Remember: correlation doesn’t imply causation!

Another important type of correlation is the autocorrelation. Let’s use the amazon dataset to understand it.

from verticapy.datasets import load_amazon

amazon = load_amazon()
amazon.head(100)
📅
date
Date
100%
...
Abc
state
Varchar(32)
100%
123
number
Int
100%
11998-01-01...AMAPÁ0
21998-01-01...AMAZONAS0
31998-01-01...DISTRITO FEDERAL0
41998-01-01...ESPÍRITO SANTO0
51998-01-01...MARANHÃO0
61998-01-01...PARANÁ0
71998-01-01...PIAUÍ0
81998-01-01...RORAIMA0
91998-01-01...SERGIPE0
101998-01-01...SÃO PAULO0
111998-02-01...GOIÁS0
121998-02-01...MATO GROSSO DO SUL0
131998-02-01...MINAS GERAIS0
141998-02-01...PARAÍBA0
151998-02-01...SANTA CATARINA0
161998-02-01...SÃO PAULO0
171998-03-01...AMAPÁ0
181998-03-01...BAHIA0
191998-03-01...MATO GROSSO DO SUL0
201998-03-01...PARÁ0
211998-03-01...PERNAMBUCO0
221998-03-01...RIO GRANDE DO SUL0
231998-04-01...CEARÁ0
241998-04-01...PARANÁ0
251998-04-01...PARAÍBA0
261998-04-01...PARÁ0
271998-05-01...ALAGOAS0
281998-05-01...AMAPÁ0
291998-05-01...MARANHÃO0
301998-05-01...MATO GROSSO DO SUL0
311998-05-01...RIO GRANDE DO SUL0
321998-05-01...TOCANTINS0
331998-06-01...ESPÍRITO SANTO6
341998-06-01...RIO DE JANEIRO3
351998-06-01...RIO GRANDE DO NORTE1
361998-06-01...SÃO PAULO451
371998-07-01...ESPÍRITO SANTO37
381998-07-01...MARANHÃO274
391998-07-01...MATO GROSSO360
401998-07-01...MATO GROSSO DO SUL3712
411998-07-01...PARAÍBA0
421998-07-01...PARÁ638
431998-07-01...RONDÔNIA365
441998-07-01...SÃO PAULO596
451998-08-01...ALAGOAS1
461998-08-01...BAHIA815
471998-08-01...DISTRITO FEDERAL48
481998-08-01...ESPÍRITO SANTO38
491998-08-01...MARANHÃO1176
501998-08-01...MATO GROSSO228
511998-08-01...MINAS GERAIS875
521998-08-01...PIAUÍ711
531998-08-01...RIO GRANDE DO SUL9
541998-08-01...RORAIMA0
551998-08-01...SERGIPE0
561998-09-01...AMAPÁ20
571998-09-01...DISTRITO FEDERAL33
581998-09-01...PIAUÍ1991
591998-09-01...RORAIMA2
601998-10-01...AMAZONAS83
611998-10-01...GOIÁS1034
621998-10-01...MATO GROSSO576
631998-10-01...PARAÍBA179
641998-10-01...PARÁ3665
651998-10-01...PIAUÍ2586
661998-10-01...SERGIPE0
671998-11-01...ALAGOAS19
681998-11-01...AMAPÁ131
691998-11-01...CEARÁ575
701998-11-01...DISTRITO FEDERAL0
711998-11-01...MARANHÃO2237
721998-11-01...RIO DE JANEIRO6
731998-11-01...RIO GRANDE DO SUL28
741998-11-01...SÃO PAULO488
751998-12-01...BAHIA82
761998-12-01...MARANHÃO1399
771998-12-01...MATO GROSSO100
781998-12-01...PARAÍBA51
791998-12-01...PERNAMBUCO59
801998-12-01...RIO DE JANEIRO1
811998-12-01...RONDÔNIA33
821998-12-01...TOCANTINS9
831999-01-01...ALAGOAS58
841999-01-01...GOIÁS14
851999-01-01...MATO GROSSO239
861999-01-01...MINAS GERAIS36
871999-01-01...PARÁ87
881999-01-01...PERNAMBUCO102
891999-01-01...RONDÔNIA1
901999-01-01...SÃO PAULO7
911999-01-01...TOCANTINS36
921999-02-01...ACRE0
931999-02-01...CEARÁ16
941999-02-01...MATO GROSSO69
951999-02-01...MATO GROSSO DO SUL28
961999-02-01...PERNAMBUCO13
971999-02-01...RONDÔNIA1
981999-02-01...SANTA CATARINA2
991999-02-01...TOCANTINS1
1001999-03-01...AMAPÁ2

Our goal is to predict the number of forest fires in Brazil. To do this, we can draw an autocorrelation plot and a partial autocorrelation plot.

amazon.acf(
    column = "number",
    ts = "date",
    by = ["state"],
    p = 24,
    method = "pearson",
)
amazon.pacf(
    column = "number",
    ts = "date",
    by = ["state"],
    p = 8,
)

We can see the seasonality forest fires.

It’s mathematically impossible to build the perfect correlation function, but we still have several powerful functions at our disposal for finding relationships in all kinds of datasets.