scikit learn to identify highly correlated features 1

scikit learn to identify highly correlated features

# Create correlation matrix
corr_matrix = df.corr().abs()

# Select upper triangle of correlation matrix
upper = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k=1).astype(np.bool))

# Find index of feature columns with correlation greater than 0.95
to_drop = [column for column in upper.columns if any(upper[column] > 0.95)]

Here is what the above code is Doing:
1. Create a correlation matrix
2. Select the upper triangle of the correlation matrix
3. Find the index of feature columns with a correlation greater than 0.95

Similar Posts