Loading...

Linear Regression

Linear regression is one of the most popular regression algorithms and produces good predictions for well-prepared data. Its optimization function computes coefficients to express a response column as a linear relationship of its predictors.

You must verify the Gauss-Markov assumptions when using linear regression algorithms:

  • Linearity: the parameters we are estimating using the OLS method must be linear.

  • Non-Collinearity: the regressors being calculated aren’t perfectly correlated with each other.

  • Exogeneity: the regressors aren’t correlated with the error term.

  • Homoscedasticity: no matter what the values of our regressors might be, the error of the variance is constant.

To create a good linear regression model, it’s important to:

  • Impute missing values.

  • Encode categorical features (linear regression only accepts numerical variables).

  • Compute the correlation matrix to retrieve highly-correlated predictors.

  • Decompose the data (optional).

  • Normalize the data (optional, but recommended).

Example without decomposition

Let’s use the africa_education dataset to compute a linear regression model of students’ performance in school.

from verticapy.datasets import load_africa_education

africa = load_africa_education()
africa = africa.select(
    [
        "(zralocp + zmalocp) / 2 AS student_score",
        "zraloct AS teacher_score",
        "XNUMYRS AS teacher_year_teaching",
        "numstu AS number_students_school",
        "PENGLISH AS english_at_home",
        "PTRAVEL AS travel_distance",
        "PTRAVEL2 AS means_of_travel",
        "PMOTHER AS m_education",
        "PFATHER AS f_education",
        "PLIGHT AS source_of_lighting",
        "PABSENT AS days_absent",
        "PREPEAT AS repeated_grades",
        "zpsit AS sitting_place",
        "PAGE AS age",
        "zpses AS socio_eco_statut",
        "country_long AS country",
    ],
)
africa.head(100)
123
student_score
Float(22)
99%
...
123
socio_eco_statut
Numeric(9)
99%
Abc
country
Varchar(24)
100%
1357.1760554...[null]Mozambique
2548.0440936...1.0Mozambique
3427.97953655...1.0Mozambique
4431.55012475...1.0Malawi
5422.5179036...1.0Zambia
6367.82150545...1.0Mozambique
7390.59366675...1.0Malawi
8421.3447214...1.0Zambia
9430.6369854...1.0Tanzania
10536.5504598...2.0Tanzania
11442.96103005...2.0Malawi
12532.25568495...2.0Uganda
13439.3294711...2.0Mozambique
14415.27831365...2.0Tanzania
15423.0274273...2.0Namibia
16610.003642...2.0Tanzania
17443.1010291...2.0Tanzania
18501.08052265...2.0Namibia
19448.05639505...2.0Mozambique
20589.0925764...2.0Tanzania
21525.2758219...2.0Mozambique
22392.00570405...2.0Mozambique
23321.1299002...2.0Mozambique
24466.7373761...2.0Malawi
25494.23537775...2.0Uganda
26604.09734565...2.0Tanzania
27519.2366424...2.0Namibia
28502.06867415...2.0Uganda
29378.4860875...2.0Zambia
30384.113366...2.0Zambia
31531.6862021...2.0Tanzania
32387.62298345...2.0Tanzania
33465.58674165...2.0Malawi
34428.54585685...2.0Zimbabwe
35459.917909...3.0Tanzania
36448.1354862...3.0Mozambique
37543.3015243...3.0Tanzania
38468.2764465...3.0Mozambique
39503.88363135...3.0Tanzania
40561.2185357...3.0Mozambique
41485.2000255...3.0Kenya
42552.0016312...3.0Tanzania
43519.69556795...3.0Tanzania
44364.77445005...3.0Mozambique
45444.7735204...3.0Malawi
46417.89208995...3.0Namibia
47422.38353385...3.0Mozambique
48464.89787835...3.0Uganda
49418.54468005...3.0Uganda
50528.9471953...3.0Tanzania
51610.5671158...3.0Tanzania
52445.80907615...3.0Malawi
53535.07574435...3.0Tanzania
54501.44606355...3.0Mozambique
55448.38439795...3.0Namibia
56518.9185695...3.0Mozambique
57389.59055735...3.0Tanzania
58699.58499885...3.0Tanzania
59433.75709155...3.0Malawi
60540.68676775...3.0Uganda
61473.4822422...3.0Uganda
62641.08794825...3.0Tanzania
63592.3426887...3.0Tanzania
64417.34281455...3.0Malawi
65564.5413827...3.0Tanzania
66576.01604155...3.0Tanzania
67424.1239538...3.0Zambia
68595.9993573...3.0Kenya
69530.22097325...3.0Uganda
70517.0661695...3.0Zimbabwe
71416.3602296...4.0Namibia
72573.10215485...4.0Tanzania
73436.4143824...4.0Malawi
74499.7551349...4.0Malawi
75473.6045948...4.0Tanzania
76383.93699955...4.0Tanzania
77595.70664635...4.0Tanzania
78441.1832625...4.0Uganda
79565.39124245...4.0Tanzania
80536.68808725...4.0Mozambique
81439.3294711...4.0Mozambique
82496.7408418...4.0Namibia
83447.4437773...4.0Namibia
84667.8671708...4.0Tanzania
85697.6814944...4.0Tanzania
86592.3426887...4.0Tanzania
87442.48809535...4.0Mozambique
88451.69440555...4.0Mozambique
89429.86792415...4.0Mozambique
90577.2088944...4.0Kenya
91431.55012475...4.0Namibia
92471.31550095...4.0Uganda
93440.6723473...4.0Uganda
94399.7076038...4.0Namibia
95713.3214605...4.0Tanzania
96471.31550095...4.0Kenya
97431.55012475...4.0Kenya
98[null]...4.0Mozambique
99423.00807355...4.0Tanzania
100584.8568118...4.0Tanzania

First, let’s look for missing values.

africa.count_percent()
...
count
percent
"number_students_school"...19290.0100.0
"english_at_home"...19290.0100.0
"travel_distance"...19290.0100.0
"means_of_travel"...19290.0100.0
"m_education"...19290.0100.0
"source_of_lighting"...19290.0100.0
"days_absent"...19290.0100.0
"repeated_grades"...19290.0100.0
"sitting_place"...19290.0100.0
"age"...19290.0100.0
"country"...19290.0100.0
"socio_eco_statut"...19271.099.902
"student_score"...19263.099.86
"teacher_year_teaching"...19233.099.705
"f_education"...19188.099.471
"teacher_score"...17214.089.238

We’ll simply drop the missing values to avoid adding bias to the data.

africa.dropna()
123
student_score
Float(22)
100%
...
123
teacher_score
Float(22)
100%
Abc
country
Varchar(24)
100%
1657.83952525...753.8647201Kenya
2387.88641055...763.3071026Mozambique
3392.73400295...763.3071026Mozambique
4422.06612525...777.9952532Mozambique
5436.79013695...696.5427817Tanzania
6531.6862021...734.979955Tanzania
7530.3308348...735.456843Mozambique
8607.95710255...753.8647201Uganda
9492.4671919...590.2921338Mozambique
10417.34281455...613.9457789Namibia
11458.2617984...648.0909602Mozambique
12515.30886255...709.1326251Mozambique
13469.0471837...727.8266349Malawi
14469.0409539...735.456843Mozambique
15535.42243675...740.6072335Kenya
16358.63862805...740.6072335Kenya
17421.37283465...763.3071026Mozambique
18400.2268045...846.0948606Zambia
19749.46855975...604.3126412Tanzania
20730.67809685...721.8178461Tanzania

We need to encode the categorical columns to dummies to retain linearity.

africa.one_hot_encode(max_cardinality = 20)
123
student_score
Float(22)
100%
...
123
country_Zambia
Bool
100%
123
country_Zanzibar
Bool
100%
1657.83952525...00
2387.88641055...00
3392.73400295...00
4422.06612525...00
5436.79013695...00
6531.6862021...00
7530.3308348...00
8607.95710255...00
9492.4671919...00
10417.34281455...00
11458.2617984...00
12515.30886255...00
13469.0471837...00
14469.0409539...00
15535.42243675...00
16358.63862805...00
17421.37283465...00
18400.2268045...10
19749.46855975...00
20730.67809685...00

Linear regression can only handle numerical columns, so we’ll drop the categorical columns.

africa.drop(
    columns = [
        "english_at_home",
        "travel_distance",
        "means_of_travel",
        "m_education",
        "f_education",
        "source_of_lighting",
        "repeated_grades",
        "sitting_place",
        "country",
    ],
)
123
student_score
Float(22)
100%
...
123
teacher_score
Float(22)
100%
123
country_Zanzibar
Bool
100%
1657.83952525...753.86472010
2387.88641055...763.30710260
3392.73400295...763.30710260
4422.06612525...777.99525320
5436.79013695...696.54278170
6531.6862021...734.9799550
7530.3308348...735.4568430
8607.95710255...753.86472010
9492.4671919...590.29213380
10417.34281455...613.94577890
11458.2617984...648.09096020
12515.30886255...709.13262510
13469.0471837...727.82663490
14469.0409539...735.4568430
15535.42243675...740.60723350
16358.63862805...740.60723350
17421.37283465...763.30710260
18400.2268045...846.09486060
19749.46855975...604.31264120
20730.67809685...721.81784610

Let’s look at the correlation between the response column and the predictors. We’ll look to keep columns with correlations coefficients greater than 20% (the top 10 features).

x = africa.corr(focus = "student_score", show = False)
africa = africa.select(columns = x["index"][0:12])
africa.head(100)
123
student_score
Float(22)
100%
...
123
source_of_lighting_CANDLE
Integer
100%
123
country_Zambia
Integer
100%
1548.0440936...10
2427.97953655...00
3431.55012475...00
4422.5179036...11
5367.82150545...00
6390.59366675...00
7421.3447214...11
8536.5504598...00
9442.96103005...10
10532.25568495...10
11439.3294711...00
12415.27831365...00
13423.0274273...10
14610.003642...00
15443.1010291...00
16501.08052265...00
17448.05639505...00
18589.0925764...00
19525.2758219...10
20392.00570405...00
21321.1299002...00
22466.7373761...00
23494.23537775...00
24604.09734565...00
25519.2366424...00
26502.06867415...00
27378.4860875...01
28384.113366...11
29531.6862021...00
30387.62298345...00
31465.58674165...10
32428.54585685...00
33459.917909...00
34448.1354862...10
35543.3015243...00
36468.2764465...10
37503.88363135...00
38561.2185357...00
39485.2000255...00
40552.0016312...00
41519.69556795...00
42364.77445005...10
43444.7735204...00
44417.89208995...10
45422.38353385...00
46464.89787835...10
47418.54468005...00
48528.9471953...00
49610.5671158...00
50445.80907615...00
51535.07574435...00
52501.44606355...00
53448.38439795...10
54518.9185695...10
55389.59055735...00
56699.58499885...00
57433.75709155...00
58540.68676775...00
59473.4822422...00
60641.08794825...00
61592.3426887...00
62417.34281455...00
63564.5413827...00
64576.01604155...00
65424.1239538...01
66595.9993573...00
67530.22097325...00
68517.0661695...00
69416.3602296...10
70573.10215485...00
71436.4143824...00
72499.7551349...00
73473.6045948...00
74383.93699955...00
75595.70664635...00
76441.1832625...00
77565.39124245...00
78536.68808725...00
79439.3294711...10
80496.7408418...10
81447.4437773...10
82667.8671708...00
83697.6814944...00
84592.3426887...00
85442.48809535...00
86451.69440555...00
87429.86792415...10
88577.2088944...00
89431.55012475...00
90471.31550095...10
91440.6723473...00
92399.7076038...10
93713.3214605...00
94471.31550095...00
95431.55012475...00
96423.00807355...00
97584.8568118...00
98456.26620545...00
99477.77385455...10
100681.9499137...00

Let’s examine the correlation matrix to see if we have any independent predictors.

africa.corr()

Some of these features are highly-correlated, like socioeconomic status and having an electric lighting. We’ll drop the lighting column to avoid unexpected results while computing the linear regression.

africa["source_of_lighting_ELECTRIC"].drop()
123
student_score
Float(22)
100%
...
123
source_of_lighting_CANDLE
Integer
100%
123
country_Zambia
Integer
100%
1657.83952525...00
2387.88641055...00
3392.73400295...00
4422.06612525...00
5436.79013695...00
6531.6862021...00
7530.3308348...00
8607.95710255...10
9492.4671919...10
10417.34281455...10
11458.2617984...00
12515.30886255...00
13469.0471837...00
14469.0409539...10
15535.42243675...00
16358.63862805...00
17421.37283465...00
18400.2268045...01
19749.46855975...00
20730.67809685...00

Let’s normalize the dataset to follow the Gaussian-Markov assumptions.

africa.normalize(columns = africa.get_columns(exclude_columns = ["student_score"]))
123
student_score
Float(22)
100%
...
123
source_of_lighting_CANDLE
Float
100%
123
country_Zambia
Float
100%
1657.83952525...-0.5181275027104616-0.23208691930827755
2387.88641055...-0.5181275027104616-0.23208691930827755
3392.73400295...-0.5181275027104616-0.23208691930827755
4422.06612525...-0.5181275027104616-0.23208691930827755
5436.79013695...-0.5181275027104616-0.23208691930827755
6531.6862021...-0.5181275027104616-0.23208691930827755
7530.3308348...-0.5181275027104616-0.23208691930827755
8607.95710255...1.9299139918588077-0.23208691930827755
9492.4671919...1.9299139918588077-0.23208691930827755
10417.34281455...1.9299139918588077-0.23208691930827755
11458.2617984...-0.5181275027104616-0.23208691930827755
12515.30886255...-0.5181275027104616-0.23208691930827755
13469.0471837...-0.5181275027104616-0.23208691930827755
14469.0409539...1.9299139918588077-0.23208691930827755
15535.42243675...-0.5181275027104616-0.23208691930827755
16358.63862805...-0.5181275027104616-0.23208691930827755
17421.37283465...-0.5181275027104616-0.23208691930827755
18400.2268045...-0.51812750271046164.308478564962018
19749.46855975...-0.5181275027104616-0.23208691930827755
20730.67809685...-0.5181275027104616-0.23208691930827755

We can use a cross-validation to test our model.

from verticapy.machine_learning.vertica import LinearRegression
from verticapy.machine_learning.model_selection import cross_validate

cross_validate(
    LinearRegression(solver = "BFGS"),
    input_relation = africa,
    X = africa.get_columns(exclude_columns = ["student_score"]),
    y = "student_score",
)
...
bic
time
1-fold...48632.02821413656.193042278289795
2-fold...48707.66944817535.871992588043213
3-fold...48986.36892533375.665011644363403
avg...48775.3555292151655.910015503565471
std...152.37101447417860.21723780237301593

The model isn’t bad. We’re just using a few variables to get a median absolute error of 47; that is, our score has a distance of 47 from the true value. This seems high, but if we keep in mind that the final score is over 1000, our predictions are quite good.

Let’s compare the importance of our features.

model = LinearRegression(solver = "BFGS")

model.fit(
    input_relation = africa,
    X = africa.get_columns(exclude_columns = ["student_score"]),
    y = "student_score",
)



=======
details
=======
        predictor        |coefficient|std_err | t_value |p_value 
-------------------------+-----------+--------+---------+--------
        Intercept        | 504.47464 | 0.54939|918.25139| 0.00000
    socio_eco_statut     | 19.93079  | 0.69586|28.64183 | 0.00000
   means_of_travel_car   |  8.02692  | 0.52140|15.39500 | 0.00000
   socio_eco_statut_14   |  7.16913  | 0.54330|13.19556 | 0.00000
  english_at_home_never  | -12.60354 | 0.53635|-23.49890| 0.00000
  repeated_grades_never  |  8.79099  | 0.58041|15.14606 | 0.00000
           age           | -5.33563  | 0.61692|-8.64887 | 0.00000
    country_tanzania     | 20.79941  | 0.57771|36.00338 | 0.00000
      teacher_score      | 11.65665  | 0.56041|20.80013 | 0.00000
source_of_lighting_candle| -3.87087  | 0.56637|-6.83449 | 0.00000
     country_zambia      | -12.10108 | 0.54174|-22.33732| 0.00000


==============
regularization
==============
type| lambda 
----+--------
none| 1.00000


===========
call_string
===========
linear_reg('"public"."_verticapy_tmp_linearregression_v_mldb_3e89952097bd11efa8720242ac120002_"', '"public"."_verticapy_tmp_view_v_mldb_3ef1d13097bd11efa8720242ac120002_"', '"student_score"', '"socio_eco_statut", "means_of_travel_CAR", "socio_eco_statut_14", "english_at_home_NEVER", "repeated_grades_NEVER", "age", "country_Tanzania", "teacher_score", "source_of_lighting_CANDLE", "country_Zambia"'
USING PARAMETERS optimizer='bfgs', epsilon=1e-06, max_iterations=100, regularization='none', lambda=1, alpha=0.5, fit_intercept=true)

===============
Additional Info
===============
       Name       |Value
------------------+-----
 iteration_count  |  6  
rejected_row_count|  0  
accepted_row_count|17099
model.features_importance()
The following factors seem to have the greatest influence on a student’s performance:
  • Having a good teacher.

  • Being of good socio-economic status.

  • Tanzanian teachers tend to overrate their students.

  • Age (younger students tend to perform better).

  • Being able to get to school by car.

Let’s add the prediction to the vDataFrame to see how our model performs its estimations.

model.predict(africa, name = "estimated_student_score")
africa.boxplot(["estimated_student_score", "student_score"])
africa.describe(columns = ["student_score", "estimated_student_score"])
...
approx_75%
max
"student_score"...558.464704565385918.5515125
"estimated_student_score"...534.304974060098675.48182789268

Our model has trouble catching outliers: exceptionally well-performing and struggling students.

Let’s draw a residual plot.

africa["residual"] = africa["student_score"] - africa["estimated_student_score"]
africa.scatter(["residual", "student_score"])

We see a high heteroscedasticity, indicating that we can’t trust the p-value of the coefficients.

model.coef_
Out[3]: 
array([ 19.93079185,   8.02692135,   7.16912611, -12.60354222,
         8.7909853 ,  -5.33563056,  20.7994124 ,  11.65665235,
        -3.87086824, -12.10107544])

Let’s look at the model’s analysis of variance (ANOVA) table.

model.report("anova")
...
F
p_value
Regression...761.42826739681380.0
Residual...
Total...

According to the ANOVA table, at least one of our variables is influencing the prediction.

We can also see that a student’s estimated score and true score skew heavily from a normal distribution.

africa["estimated_student_score"].hist()
from verticapy.machine_learning.model_selection.statistical_tests import jarque_bera

jarque_bera(africa, "estimated_student_score")
Out[5]: (566.447140204892, 9.944120070619405e-124)

Our model doesn’t verify the basic hypothesis and therefore isn’t stable enough to be put into production. Let’s look at a second technique.

Example with decomposition

Let’s look at the same dataset, but use decomposition techniques to filter out unimportant information. We don’t have to normalize our data or look at correlations with these types of methods.

We’ll begin by repeating the data preparation process of the previous section and export the resulting vDataFrame to Vertica.

africa = load_africa_education()
africa = africa.select(
    [
        "(zralocp + zmalocp) / 2 AS student_score",
        "zraloct AS teacher_score",
        "XNUMYRS AS teacher_year_teaching",
        "numstu AS number_students_school",
        "PENGLISH AS english_at_home",
        "PTRAVEL AS travel_distance",
        "PTRAVEL2 AS means_of_travel",
        "PMOTHER AS m_education",
        "PFATHER AS f_education",
        "PLIGHT AS source_of_lighting",
        "PABSENT AS days_absent",
        "PREPEAT AS repeated_grades",
        "zpsit AS sitting_place",
        "PAGE AS age",
        "zpses AS socio_eco_statut",
        "country_long AS country",
    ],
)
africa.dropna()
africa.one_hot_encode(max_cardinality = 20)
africa.drop(
    columns = [
        "english_at_home",
        "travel_distance",
        "means_of_travel",
        "m_education",
        "f_education",
        "source_of_lighting",
        "repeated_grades",
        "sitting_place",
        "country",
    ],
)
123
student_score
Float(22)
100%
...
123
teacher_score
Float(22)
100%
123
country_Zanzibar
Bool
100%
1657.83952525...753.86472010
2387.88641055...763.30710260
3392.73400295...763.30710260
4422.06612525...777.99525320
5436.79013695...696.54278170
6531.6862021...734.9799550
7530.3308348...735.4568430
8607.95710255...753.86472010
9492.4671919...590.29213380
10417.34281455...613.94577890
11458.2617984...648.09096020
12515.30886255...709.13262510
13469.0471837...727.82663490
14469.0409539...735.4568430
15535.42243675...740.60723350
16358.63862805...740.60723350
17421.37283465...763.30710260
18400.2268045...846.09486060
19749.46855975...604.31264120
20730.67809685...721.81784610

Let’s create our principal component analysis (PCA) model.

from verticapy.machine_learning.vertica import PCA

model = PCA()
model.fit(
    africa,
    africa.get_columns(exclude_columns = ["student_score"]),
)
africa_pca = model.transform()
africa_pca.head(100)
123
student_score
Float(22)
100%
...
123
col85
Float(22)
100%
123
col86
Float(22)
100%
1657.83952525...0.000359577699693903-0.00027997200419123
2387.88641055...0.001435869869298190.00023401543263306
3392.73400295...-0.000268934860135830.000597456574135668
4422.06612525...-0.007596995199243490.00321536138185285
5436.79013695...0.0004778782548877220.000586078032193242
6531.6862021...0.000325373054157830.000223271620745539
7530.3308348...-0.007909089158002770.00276579178468772
8607.95710255...-0.0004482673326317590.00016682149011722
9492.4671919...-0.001544248728160670.000298794345840233
10417.34281455...-0.0003540762506030670.000278375154693887
11458.2617984...0.0008983683776041580.000369241593603549
12515.30886255...-0.0004666625769174950.000116772027345443
13469.0471837...-0.0008510277527155770.000578910440615927
14469.0409539...0.0002240372020072390.000266757730610997
15535.42243675...-0.0005511075493234520.000535499961478971
16358.63862805...-0.001647386541010320.00131674529425853
17421.37283465...-0.000180061726714492-0.000186331793102462
18400.2268045...0.0004115533304551220.000171698360315153
19749.46855975...-0.0002129484456707310.000414280898441772
20730.67809685...0.000591773188009201-8.19261160070404e-05
21539.2019957...-0.0008952425433214540.00089279759559188
22425.99336735...-0.0001356579578480940.000242428842493865
23397.4796701...-0.0005627475555285142.55345437112765e-05
24453.0202996...0.0003429099263032060.000779725960077904
25449.99250385...0.00013399259194440.000196258018422879
26482.2092927...-4.48518535405074e-050.000496011381165726
27510.49975695...-1.20109828432714e-050.000319115309786115
28641.9622534...-0.0008621098283968510.000851658003102353
29510.65265825...0.000415370295012070.000514312256077254
30535.58574205...-0.000571936808605480.000724992415398366
31399.7076038...-0.0003220665736634710.00045881152904497
32602.52488005...0.0002933461475505790.000866651942232944
33565.60944755...-0.001387296682338330.000944194840685255
34451.69440555...-0.007306033549503470.00264216392643703
35495.5545671...0.00139738582608655-0.00058474861860522
36456.9573452...0.00160913149549304-0.000187013743951973
37476.34012315...-0.001200091898764030.00018986107134941
38585.9983071...0.000394209879192926-0.000384525192063908
39596.11136875...0.000358976415015503-0.00037771761997554
40452.3792203...0.000681432549534532-0.000298043681504545
41614.97165705...0.00051821569540126-0.000392878537166623
42473.4818625...9.10721168901016e-05-0.000263104033196432
43508.95161015...-0.000481136699261928-0.000655703155920966
44588.9083679...1.72963967945152e-05-9.91417439334259e-05
45551.2711833...-8.35803913525755e-05-0.000353382184719902
46365.06159545...0.000353926892988085-0.000111189897703668
47386.94044515...0.00018103598217592-2.22823827304551e-05
48427.07559925...0.0001189003545214050.000234190891084192
49540.20488335...-0.0004881780361559220.000190949515654138
50470.4509993...0.000750881080153201-0.000114125196752734
51489.53970635...0.0004405334425367098.85555946701309e-05
52556.7909405...0.000321757095271491-0.00012833989827228
53528.38596655...0.0004658265536777361.45062973675688e-05
54555.0513114...0.00042550461529055-0.000282153342778823
55519.13142985...0.000724116158647454-0.000290774440477193
56482.43679495...0.000389245226583548-0.000177547632580087
57459.9863116...0.000636500826741126.36564302649487e-05
58599.1407463...0.0002096105108317885.46683853875884e-05
59476.61604135...0.000314993507149266-0.00017378746156667
60404.85261815...0.000459262665982364-8.61116170375874e-05
61413.12320985...0.00070301860311775-0.000201018517687249
62630.4264972...0.000443377852281037-0.000134958310202261
63743.28447985...0.000914103785246775-0.000513648968669834
64238.94653645...0.00153100952542612-0.00048676700709966
65498.99230345...-6.7976154939581e-05-0.000178571290525526
66497.0200491...0.0005666399220967870.000710924129080911
67413.91855045...0.000706203264469199-8.6191784619231e-05
68457.58833645...-8.55283352749521e-064.28884424216678e-05
69514.33800925...0.0002327659540981877.82518454783603e-06
70419.0689409...0.0001210743113879183.14852520442872e-05
71440.6723473...0.0001210743113879183.14852520442872e-05
72433.6962473...0.0001210743113879183.14852520442872e-05
73519.13142985...0.0001210743113879183.14852520442872e-05
74758.1966219...0.00122982804388162-0.000519809682233395
75525.65015365...0.0001032745387880542.68969994162426e-05
76640.024659...-0.000177887631074497-0.000126896049547868
77433.75709155...0.000714080760975406-0.00010092901016448
78602.12015765...0.0006727418655940578.63662531016798e-05
79500.19151275...0.000749725492647564-0.000408676209367053
80579.51566615...-0.000343642707958085-0.000335689878077163
81495.5545671...0.000711419986526615-0.000606594531118205
82428.54585685...0.0005879110000921060.000134760309745109
83519.20701115...0.000617238788763242-0.000497245873567768
84406.15482625...-0.0003953878251024197.98429174994424e-05
85590.5249487...0.000333769474608915-0.000366131716646391
86541.55778385...0.000633900800826524-0.000356379111243875
87544.01198645...0.000497228805909419-0.000315148555179469
88490.89222805...-0.000457282358303608-0.000112962872181551
89548.6765391...-7.20068130863998e-05-0.000416001493545991
90455.36049795...-0.000709146622833584-0.000277594076985878
91519.147368...-0.000207478220038972-1.14918718115311e-06
92574.3841869...0.00050064296009617-0.000113392631338879
93490.7928971...0.000184485476359688-8.25380891488927e-05
94497.5443413...-0.000609485113200837-0.000346131463047032
95611.75199915...0.0006335392848871239.95676846338726e-05
96532.5956098...0.000522265108608628-0.000322639249720833
97525.8445453...-0.0001016062602835910.000763504677147397
98501.42749965...0.000412844275820035-0.000465228143146301
99512.3963035...0.00050218862950323-0.000329715367076172
100482.2092927...-0.000464696833843353-0.000408899606932095

We can verify the Gauss-Markov assumptions with our PCA model.

africa_pca.corr(
    columns =
        [
            "student_score",
            "col1",
            "col2",
            "col3",
            "col4",
            "col5",
            "col6",
            "col7",
        ],
)

Let’s use a cross-validation to test our linear regression model.

cross_validate(
    LinearRegression(solver = "BFGS"),
    input_relation = africa_pca,
    X = africa_pca.get_columns(exclude_columns = ["student_score"]),
    y = "student_score",
)
...
bic
time
1-fold...48515.48869450625.473784923553467
2-fold...48586.028494309924.47300410270691
3-fold...48744.75968147918.33084535598755
avg...48615.4256234316322.75921146074931
std...95.879924097256393.157869569946982

As you can see, we’ve created a much more accurate model here than in our first attempt. This example emphasizes the importance of filtering noise from the data.

Conclusion

We’ve seen two techniques that can help us create powerful linear regression models. While the first method normalized the data and looked for correlations, the second method applied a PCA model. The second one allows us to confirm the Gauss-Markov assumptions - an essential part of using linear models.