{"cells":[{"metadata":{},"cell_type":"markdown","source":"The aim of this notebook is to be a comprehensive starting point for anyone interested and/or participating the OSIC Pulmonary Fibrosis Progression competition.\nIt will focus on analyzing the OSIC training set provided for the competition, including:\n* Tabular clinical data\n* CT Lung Scans and Features extracted from the scans\n\n(As this is the first time I'm posting an EDA here, any advice/suggestions regarding the analysis itself, the charts used or the coding, is welcome).","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt # Plotting\nimport seaborn as sns # statistical data visualization\nimport plotly.express as px\nimport plotly.graph_objects as go # interactive plots\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Dataframe uniqueness/quality check","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Let's start by checking how many data points there are, and if anything is missing.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Number of data points: ' + str(len(train_df)))\nprint('----------------------')\n\nfor col in train_df.columns:\n    print('{} : {} unique values, {} missing.'.format(col, \n                                                          str(len(train_df[col].unique())), \n                                                          str(train_df[col].isna().sum())))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This dataset will contain a single entry for each patient, leaving only their personal information, to have a clearer and more accurate representation of the distribution of these categorical variables.\nIt will also include how many times a patient appears in the dataset, to observe how many visits each one has done.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"unique_patient_df = train_df.drop(['Weeks', 'FVC', 'Percent'], axis=1).drop_duplicates().reset_index(drop=True)\nunique_patient_df['# visits'] = [train_df['Patient'].value_counts().loc[pid] for pid in unique_patient_df['Patient']]\n\nprint('Number of data points: ' + str(len(unique_patient_df)))\nprint('----------------------')\n\nfor col in unique_patient_df.columns:\n    print('{} : {} unique values, {} missing.'.format(col, \n                                                          str(len(unique_patient_df[col].unique())), \n                                                          str(unique_patient_df[col].isna().sum())))\nunique_patient_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 1 - Single Variable Analysis","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"The start of this analysis will look at each variable by itself, to see its distribution and the range it occupies.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# 1.1 Discrete features","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Age**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,5))\nsns.countplot(x='Age', data=unique_patient_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The age variable doesn't seem to be particularly informative, it follows a Normal/Gaussian distribution.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Sex**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(5,3))\nsns.countplot(x='Sex', data=unique_patient_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is a clear unbalance in the patients sex; assuming this imbalance is also present in the test set and isn't a result of random sampling (Do males have a higher probability of getting sick ?), it could be important to check the correlation with the FVC Progression values.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Smoking Status**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10, 3))\nsns.countplot(x='SmokingStatus', data=unique_patient_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Again, there is some imbalance between the different \"Smoking Status\" categories, it may be worthwhile to check the presence of any correlation with the target variable.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Weeks (# of visits per patient)**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(20, 5))\nax = sns.countplot(x='Weeks', data=train_df)\nax.set_xticklabels(ax.get_xticklabels(), rotation='vertical')\n\nplt.figure(figsize=(5, 5))\nsns.countplot(x='# visits', data=unique_patient_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From the \"Weeks\" attribute one can see: \n* The necessary values for the prediction stage, from week -12 to 133\n* The visits become increasingly rare the more time passes from the CT scan\n* The majority of patients has been subjected to an average of 9 medical visits.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# 1.2 Continuous features","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**FVC**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10, 5))\nsns.distplot(train_df['FVC'], rug=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Percent**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10, 5))\nsns.distplot(train_df['Percent'], rug=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Both \"FVC\" and \"Percent\" seem to roughly follow a Gaussian distribution with a light skew to the left.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Expected FVC**","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"The following two variables represents the actual FVC values represented in the dataset by the \"Percent\" variable and their difference.\nInterestingly, there are a couple patients where Percent is over 100%, meaning the measured FVC is higher than what it's supposed to be.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Expected FVC'] = train_df['FVC'] * (100/train_df['Percent'])\n\nplt.figure(figsize=(10, 5))\nsns.distplot(train_df['Expected FVC'], rug=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**FVC Difference**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['FVC Difference'] = train_df['Expected FVC'] - train_df['FVC']\n\nplt.figure(figsize=(10, 5))\nsns.distplot(train_df['FVC Difference'], rug=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 2 - Feature Interaction Analysis","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Here we'll check the possible relationship between variables.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Clinical Data contingency table**","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Let's start by checking the composition of our dataset, by Sex and Smoking Status.\nThe majority (60%) is composed by males (has we have seen before) ex-smokers (the most represented category.)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.crosstab(train_df.Sex, train_df.SmokingStatus, margins=True, normalize=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There doesn't seem to be any particular correlation between the categorical variables.\nThe continous variables exhibit a clear positive correlation, between \"FVC\" and \"Percent\", while the new engineered feature FVC Difference has an almost perfect negative correlation with Percent.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.heatmap(train_df.apply(lambda x : pd.factorize(x)[0] if x.dtype=='object' else x).corr(method='pearson'),\n            annot=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The following interactive charts show the Pulmonary Capacity decline  by Patient, Sex, and Smoking Status.\nIn each of them clicking on the legend will show only the corresponding plot, double-clicking will hide/show all the lines.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Actual/Expected FVC Progression by Patient**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.line(train_df, 'Weeks', 'FVC', color='Patient',\n             title='Pulmonary Condition Progression by Patient',\n             labels={'Week':'Week #',\n                     'FVC' : 'Actual\\Expected FVC'})\nfig.update_traces(mode='lines+markers')\n\nfor i in fig.data:\n    new_element = go.Scattergl(i)\n    new_element['mode'] = 'markers'\n    new_element['name'] = i['name'] + '_EX'\n    new_element['y'] = train_df.loc[train_df['Patient'] == new_element['legendgroup']]['Expected FVC']\n    fig.add_trace(new_element)\n\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**FVC Progression by Sex**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.line(train_df, 'Weeks', 'FVC', line_group='Patient', color='Sex',\n             title='Pulmonary Condition Progression by Sex')\nfig.update_traces(mode='lines+markers')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**FVC Progression by SmokingStatus**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.line(train_df, 'Weeks', 'FVC', line_group='Patient', color='SmokingStatus',\n             title='Pulmonary Condition Progression by Sex')\nfig.update_traces(mode='lines+markers')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Partitioning the patients by Sex seems to show that female patients occupy a smaller range of FVC values, as does partitioning by SmokingStatus for people currently smoking, but we have to remember that those are the two most under-represented categories in the dataset, so these may just be faulty assumptions.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# 3 - CT Scan Features ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"In this section I'll explore the DICOM files representing the CT Lung Scans for the patients, how to load and process them, and how to extract some additional features.\n\nFor an even more informative explanation, you can visit the Radiology Data Quest blog (https://www.raddq.com/dicom-processing-segmentation-visualization-in-python/), from where I learned the tools used in the rest of this notebook and which I thank for their public explanation.\n\nIt's also necessary to thank Kaggle user Dr.Sàndor Kónya (@sandorkonya) for his wonderful Domain Expert Insights, which can be found here (Part 1: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727, Part 2: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123).","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import pydicom\nfrom glob import glob\nimport scipy.ndimage\nfrom skimage import morphology\nfrom skimage import measure\nfrom skimage.filters import threshold_otsu, median\nfrom scipy.ndimage import binary_fill_holes\nfrom skimage.segmentation import clear_border\nfrom scipy.stats import describe","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's check the scans for the first patient.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_dir = '../input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430/'\nscans = glob(patient_dir + '/*.dcm')\n\nscans[:5]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's define some helper functions to load, sort and process the images.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def load_scan(path):\n    scans = os.listdir(path)\n    slices = []\n    \n    for scan in scans:\n        with pydicom.dcmread(path + '/' + scan) as s:\n            slices.append(s)\n    \n    slices.sort(key = lambda x: int(x.InstanceNumber))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        try:\n            slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        except:\n            slice_thickness = slices[0].SliceThickness\n    \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices\n\ndef get_pixels_hu(scans):\n    image = np.stack([s.pixel_array for s in scans])\n    # Convert to int16 (from sometimes int16), \n    # should be possible as values should always be low enough (<32k)\n    image = image.astype(np.int16)\n\n    # Set outside-of-scan pixels to 1\n    # The intercept is usually -1024, so air is approximately 0\n    image[image <= -1000] = 0\n    \n    # Convert to Hounsfield units (HU)\n    intercept = scans[0].RescaleIntercept\n    slope = scans[0].RescaleSlope\n    \n    if slope != 1:\n        image = slope * image.astype(np.float64)\n        image = image.astype(np.int16)\n        \n    image += np.int16(intercept)\n    \n    return np.array(image, dtype=np.int16)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"patient = load_scan(patient_dir)\nimgs = get_pixels_hu(patient)\n\nfig, ax = plt.subplots(1, 2, figsize=(10, 10))\n\nax[0].set_title('Original Image')\nax[0].imshow(patient[15].pixel_array, cmap='gray')\nax[0].axis('off')\n\nax[1].set_title('HU Image')\nax[1].imshow(imgs[15], cmap='gray')\nax[1].axis('off')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's have a look at the distribution of pixels in the images converted to Hounsfield Units.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.hist(imgs.flatten(), bins=50, color='c')\nplt.xlabel(\"Hounsfield Units (HU)\")\nplt.ylabel(\"Frequency\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Checking with the HU table shows: \n* There is a lot of air, the peak around -1000\n* At around -500 there are some pixels classified as Lung;\n* The smaller gaussian-shaped peak at 0 indicates soft tissues;\n* The distribution tail from 700 onward shows the presence of bones.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def sample_stack(stack, rows=2, cols=3, start_with=0, show_every=5):\n    fig,ax = plt.subplots(rows,cols,figsize=[10,10])\n    for i in range(rows*cols):\n        ind = start_with + i*show_every\n        ax[i // (rows+1), int(i % cols)].set_title('slice %d' % ind)\n        ax[i // (rows+1), int(i % cols)].imshow(stack[ind],cmap='gray')\n        ax[i // (rows+1), int(i % cols)].axis('off')\n    plt.tight_layout()\n    plt.show()\n\nsample_stack(imgs)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let's try to segment the lungs from the scan.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def lung_segment(img, display=False):\n    thresh = threshold_otsu(img)\n    binary = img <= thresh\n\n    lungs = median(clear_border(binary))\n    lungs = morphology.binary_closing(lungs, selem=morphology.disk(7))\n    lungs = binary_fill_holes(lungs)\n\n    final = lungs*img\n    final[final == 0] = np.min(img)\n    \n    if display:\n        fig, ax = plt.subplots(1, 4, figsize=(15, 15))\n\n        ax[0].set_title('HU Image')\n        ax[0].imshow(img, cmap='gray')\n        ax[0].axis('off')\n\n        ax[1].set_title('Thresholded Image')\n        ax[1].imshow(binary, cmap='gray')\n        ax[1].axis('off')\n\n        ax[2].set_title('Lungs Mask')\n        ax[2].imshow(lungs, cmap='gray')\n        ax[2].axis('off')\n\n        ax[3].set_title('Final Image')\n        ax[3].imshow(final, cmap='gray')\n        ax[3].axis('off')\n    \n    return final, lungs\n\ndef lung_segment_stack(imgs):\n    \n    masks = np.empty_like(imgs)\n    segments = np.empty_like(imgs)\n    \n    for i, img in enumerate(imgs):\n        seg, mask = lung_segment(img)\n        segments[i,:,:] = seg\n        masks[i,:,:] = mask\n        \n    return segments, masks","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lung_segment(imgs[15], display=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"segmented, masks = lung_segment_stack(imgs)\nsample_stack(segmented)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Lung Volume/Total Lung Capacity Estimation**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def lung_volume(patient, masks):\n    slice_thickness = patient[0].SliceThickness\n    pixel_spacing = patient[0].PixelSpacing\n    \n    return np.round(np.sum(masks) * slice_thickness * pixel_spacing[0] * pixel_spacing[1], 3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lung_vol = lung_volume(patient, masks)\nprint('Lung Volume is: ' + str(lung_vol) + ' mm^3 (' + str(lung_vol/1e6) + ' liters)')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"According to Wikipedia the average lung capacity for an adult is about 6 litres of air (6 millions mm^3); for this particular patient we got almost 4 liters, which is probably expected due to his pulmonary condition and the imperfect segmentation.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Histogram Analysis**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def hist_analysis(segmented, display=False):\n    values = segmented.flatten()\n    values = values[values >= -1000]\n    \n    if display:\n        plt.hist(values, bins=50)\n    \n    summary_statistics = describe(values)\n    \n    return summary_statistics","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"h_stat = hist_analysis(segmented, display=True)\n\nprint('Mean is: ' + str(h_stat.mean))\nprint('Variance is: ' + str(h_stat.variance))\nprint('Skewness is: ' + str(h_stat.skewness))\nprint('Kurtosis is: ' + str(h_stat.kurtosis))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Lung Area/Diameter/Circumference**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def chest_measurements(patient, masks):\n    middle_slice = masks[len(masks)//2]\n    pixel_spacing = patient[0].PixelSpacing\n    \n    lung_area = np.round(np.sum(middle_slice.flatten()) * pixel_spacing[0] * pixel_spacing[1], 3)\n    \n    conv_h = morphology.convex_hull_image(middle_slice)\n    props = measure.regionprops(measure.label(conv_h))\n    \n    chest_diameter = np.round(props[0].major_axis_length, 3)\n    chest_circ = np.round(props[0].perimeter, 3)\n    \n    return lung_area, chest_diameter, chest_circ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ch_measure = chest_measurements(patient, masks)\nprint('Lung Area is: ' + str(ch_measure[0]) + ' mm^2')\nprint('Chest Diameter estimate is: ' + str(ch_measure[1]) + ' mm')\nprint('Chest Circumference estimate is: ' + str(ch_measure[2]) + ' mm')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I'll remove the patients with codes **ID00011637202177653955184** and **ID00052637202186188008618**, because they cause errors when trying to read the DICOM files.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#augmented_df = train_df.drop(train_df.loc[train_df['Patient'].isin(['ID00011637202177653955184', #\n#                                                                    'ID00052637202186188008618'])]\n#                             .index).reset_index(drop=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let's compute these features for each patient and check the new correlation matrix.\n\n**The kernel crashes with an out-of-memory error here because the size of the scans size and a probable memory leak in DICOM scans loading; the engineered features may be useful, but impossible to compute on a kernel for all data points.\nIf someone knows how to solve this please write it in the comments.**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#augmented_df['LungVolume'] = None\n#augmented_df['Mean'] = None\n#augmented_df['Variance'] = None\n#augmented_df['Skewness'] = None\n#augmented_df['Kurtosis'] = None\n#augmented_df['LungArea'] = None\n#augmented_df['ChestDiameter'] = None\n#augmented_df['ChestCircumference'] = None\n\n#for pid in augmented_df['Patient'].unique():\n#    patient_dir = '../input/osic-pulmonary-fibrosis-progression/train/' + pid\n#    patient = load_scan(patient_dir)\n#    scans = get_pixels_hu(patient)\n    \n#    segmented, masks = lung_segment_stack(scans)\n    \n#    augmented_df.loc[augmented_df['Patient'] == pid, 'LungVolume'] = lung_volume(patient, masks)\n    \n#    hist_stats = hist_analysis(segmented)\n#    augmented_df.loc[augmented_df['Patient'] == pid, 'Mean'] = np.round(hist_stats.mean, 3)\n#    augmented_df.loc[augmented_df['Patient'] == pid, 'Variance'] = np.round(hist_stats.variance, 3)\n#    augmented_df.loc[augmented_df['Patient'] == pid, 'Skewness'] = np.round(hist_stats.skewness, 3)\n#    augmented_df.loc[augmented_df['Patient'] == pid, 'Kurtosis'] = np.round(hist_stats.kurtosis, 3)\n    \n#    chest_stat = chest_measurements(patient, masks)\n#    augmented_df.loc[augmented_df['Patient'] == pid, 'LungArea'] = chest_stat[0]\n#    augmented_df.loc[augmented_df['Patient'] == pid, 'ChestDiameter'] = chest_stat[1]\n#    augmented_df.loc[augmented_df['Patient'] == pid, 'ChestCircumference'] = chest_stat[2]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#plt.figure(figsize=(13,13))\n#sns.heatmap(augmented_df.apply(lambda x : pd.factorize(x)[0] if x.dtype=='object' else x).corr(method='pearson'),\n#            annot=True)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}