vDataFrame.drop_duplicates

In [ ]:
vDataFrame.drop_duplicates(columns: list = [])

Filters the duplicated using a partition by the input vcolumns.

⚠ Warning: Dropping duplicates will make the vDataFrame structure heavier. It is recommended to always check the current structure using the 'current_relation' method and to save it using the 'to_db' method with the parameters 'inplace = True' and 'relation_type = table'

Parameters

Name Type Optional Description
columns
list
✓
List of the vcolumns names. If empty, all the vcolumns will be selected.

Returns

vDataFrame : self

Example

In [154]:
from verticapy.datasets import load_titanic
titanic = load_titanic().select(["pclass", "survived"])
display(titanic)
123
pclass
Int
123
survived
Int
110
210
310
410
510
610
710
810
910
1010
1110
1210
1310
1410
1510
1610
1710
1810
1910
2010
2110
2210
2310
2410
2510
2610
2710
2810
2910
3010
3110
3210
3310
3410
3510
3610
3710
3810
3910
4010
4110
4210
4310
4410
4510
4610
4710
4810
4910
5010
5110
5210
5310
5410
5510
5610
5710
5810
5910
6010
6110
6210
6310
6410
6510
6610
6710
6810
6910
7010
7110
7210
7310
7410
7510
7610
7710
7810
7910
8010
8110
8210
8310
8410
8510
8610
8710
8810
8910
9010
9110
9210
9310
9410
9510
9610
9710
9810
9910
10010
Rows: 1-100 of 1234 | Columns: 2
In [155]:
titanic.drop_duplicates()
1228 element(s) was/were filtered
123
pclass
Int
123
survived
Int
110
211
320
421
530
631
Out[155]:
Rows: 6 | Columns: 2