{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Performing some basic data exploration trying get a few insights on the data we are dealing with\n- Observation are mentioned after each of the experiments\n- First time publishing a notebook, any feedback appreciated!","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\n\nPATH = '/kaggle/input/siim-isic-melanoma-classification'\nprint(os.listdir(PATH))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%matplotlib inline\n\nimport matplotlib\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/train.csv')\ntest_df = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/test.csv')\nbenign_df = train_df[train_df['benign_malignant']=='benign']\nmalignant_df = train_df[train_df['benign_malignant']=='malignant']\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 3 columns(`sex`, `age_approx`, `anatom_site_general_challenge`) with some missing values. We will look into each.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Identifying the Target Data Distribution","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"num_training = len(train_df)\nnum_benign = len(train_df[train_df['benign_malignant']=='benign'])\nnum_malignant = len(train_df[train_df['benign_malignant']=='malignant'])\n\nprint(\"Total number of records :\", len(train_df))\nprint(\"Number of Benign records :\", len(train_df[train_df['benign_malignant']=='benign']))\nprint(\"Number of Malignant records :\", len(train_df[train_df['benign_malignant']=='malignant']))\nprint(f\"Percentage of Benign records : {num_benign/num_training:.2f}\")\nprint(f\"Percentage of Malignant records : {num_malignant/num_training:.2f}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**OBSERVATION**:\n- _HUGE_ Data imbalance. Need to think of ways to make this even out or consider it during classification\n- Some methods maybe :\n    - Data Augmentation\n    - Using GANs to produce synthetic data\n    - Weighted Loss\n    - Oversampling\n    - SMOTE","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Visualizing the Distribution of sex","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plots of distrubution of sex\nfig, ax = plt.subplots(1,3, figsize=(15,5))\nax[0]=train_df['sex'].dropna().value_counts().plot(kind='bar', ax=ax[0])\nax[1]=train_df[(train_df['benign_malignant']=='benign')]['sex'].dropna().value_counts().plot(kind='bar', ax=ax[1])\nax[2]=train_df[(train_df['benign_malignant']=='malignant')]['sex'].dropna().value_counts().plot(kind='bar', ax=ax[2])\n\nfor ax_ in ax:\n    for p in ax_.patches:\n        ax_.annotate(str(p.get_height()), (p.get_x() * 1.005, p.get_height() * 1.005))\n        \nax[0].title.set_text('Train Set Sex Distribuiton')\nax[1].title.set_text('Benign Subset Sex Distribuiton')\nax[2].title.set_text('Malignant Subset Sex Distribuiton')\n\nprint(\"Number of missing rows : \", len(train_df['sex'])-train_df['sex'].count())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### A look into the Age column","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Distribution of age for each individually\nfig, ax = plt.subplots(1,3, figsize=(15,5))\n\nax[0]=train_df['age_approx'].dropna().plot(kind='density', ax=ax[0])\nax[1]=train_df['age_approx'].dropna().plot(kind='hist', ax=ax[1])\nax[2]=train_df['age_approx'].dropna().plot(kind='box', ax=ax[2])\n\nfor p in ax[1].patches:\n        ax[1].annotate(str(int(p.get_height())), (p.get_x() * 1.005, p.get_height() * 1.005))\n        \nfig.suptitle('Age Distribution')\nprint(\"Number of missing rows : \", len(train_df['age_approx'])-train_df['age_approx'].count())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(15,5))\n_ = train_df[train_df[\"target\"]==0].age_approx.hist(bins=20, ax=ax[0])\n_ = train_df[train_df[\"target\"]==1].age_approx.hist(bins=20, ax=ax[1])\n_ = train_df.groupby(\"target\").age_approx.hist(bins=20, alpha=1, ax=ax[2])\n_ = ax[0].set_title('Age Distribution for Benign Cases')\n_ = ax[1].set_title('Age Distribution for Malignant Cases')\n_ = ax[2].set_title('Age Distribution Comparision')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Possible values of age :\",np.sort(train_df.age_approx.unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**OBSERVATIONS**\n- Has somewhat of a normal distribution\n- Could be a useful column, Malignant Melanoma more prominant for ages between 45-75.\n    - However, with the true data distribution, these are also the most frequently occuring data points, so not sure how much to expect\n- Age is provided almost as a catergorical variable","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Exploring the Diagnosis Column","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"ax = train_df['diagnosis'].value_counts().plot(kind='bar')\nfor p in ax.patches:\n        ax.annotate(str(int(p.get_height())), (p.get_x() * 1.005, p.get_height() * 1.005))\n        \nax.set_title('Diagnosis Distribution')\nprint(\"Number of unknown diagnosis : \", len(train_df[train_df['diagnosis']=='unknown']))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**OBSERVATIONS** :\n- Looks the majority of the diagnosis are unknown. In this case, this might not be a useful parameter for making predictions","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(15,5))\n\nax[0]=train_df['anatom_site_general_challenge'].value_counts().dropna().plot(kind='bar', ax=ax[0])\nax[1]=benign_df['anatom_site_general_challenge'].value_counts().dropna().plot(kind='bar', ax=ax[1], colormap='plasma')\nax[2]=malignant_df['anatom_site_general_challenge'].value_counts().dropna().plot(kind='bar', ax=ax[2])\n\nfor ax_ in ax:\n    for p in ax_.patches:\n            ax_.annotate(str(int(p.get_height())), (p.get_x() * 1.005, p.get_height() * 1.005))\n\n_ = ax[0].set_title('Train Dataset Distribution')\n_ = ax[1].set_title('Benign Set Distribution')\n_ = ax[2].set_title('Malignant Set Distribution')\n_ = fig.suptitle('Anatom Sight Distribution', fontweight='black', fontsize=15)\nprint(\"Number of missing rows : \", len(train_df['anatom_site_general_challenge'])-train_df['anatom_site_general_challenge'].count())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**OBSERVATIONS**\n- Most Common by far occurs in the _Torso_ Region.\n- Not very sure what we can expect from this variable, especially since the distributions on the Benign and Malignant set are almost identical as seen above.\n- Also has a good number of missing values","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Data Purity Check","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"Total number of records : {len(train_df)}\\n\\\nNumber of patients : {len(train_df.patient_id.unique())}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, there are multiple records per patient. We should now confirm that there is no information leak between the test and train set in the form of patients","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"overlap = len(set(train_df.patient_id.unique()).intersection(set(test_df.patient_id.unique())))\nprint(f\"Number of patients common in train and test set = {overlap}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As expected. No information leak, however, when we split the data into train and validation set, we need to make sure this holds then also.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Visualizing the images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"TRAIN_IMAGE_PATH = os.path.join(PATH,'jpeg','train') \nTRAIN_IMAGES_LIST = os.listdir(TRAIN_IMAGE_PATH)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Images at Random from the entire dataset","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(4,4,figsize=(15,15))\n \nfor i, row in enumerate(ax):\n    for j, cell in enumerate(row):\n        idx = np.random.randint(len(TRAIN_IMAGES_LIST))\n        ax[i,j].imshow(plt.imread(os.path.join(TRAIN_IMAGE_PATH, TRAIN_IMAGES_LIST[idx])))\n        ax[i,j].axis('off')\n#         print(f\"Reading image {TRAIN_IMAGES_LIST[idx]}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Some BENIGN Images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(4,4,figsize=(15,15))\nBENIGN_IMAGES_LIST = list(benign_df['image_name'])\nfor i, row in enumerate(ax):\n    for j, cell in enumerate(row):\n        idx = np.random.randint(len(BENIGN_IMAGES_LIST))\n        ax[i,j].imshow(plt.imread(os.path.join(TRAIN_IMAGE_PATH, BENIGN_IMAGES_LIST[idx]+'.jpg')))\n        ax[i,j].axis('off')\n#         print(f\"Reading image {TRAIN_IMAGES_LIST[idx]}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Some MALIGNANT Images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(4,4,figsize=(15,15))\nMALIGNANT_IMAGES_LIST = list(malignant_df['image_name'])\nfor i, row in enumerate(ax):\n    for j, cell in enumerate(row):\n        idx = np.random.randint(len(MALIGNANT_IMAGES_LIST))\n        ax[i,j].imshow(plt.imread(os.path.join(TRAIN_IMAGE_PATH, MALIGNANT_IMAGES_LIST[idx]+'.jpg')))\n        ax[i,j].axis('off')\n#         print(f\"Reading image {TRAIN_IMAGES_LIST[idx]}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**OBSERVATIONS** : \n- Images are of different sizes. \n- The melanoma spread is of very different sizes. If we look to use images in neural nets, we need to figure out a way to make all images the same size.\n- There are some circular images also! Need to deal with that too!","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Quick Look at the pixel densities","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"benign_image_sample = plt.imread(os.path.join(TRAIN_IMAGE_PATH, BENIGN_IMAGES_LIST[np.random.randint(len(BENIGN_IMAGES_LIST))]+'.jpg'))\n\nfig, ax = plt.subplots(1,2,figsize=(15,5))\nax[0].imshow(benign_image_sample)\nax[0].set_xlabel(f\"Image dimensions : {benign_image_sample.shape[:2]}\")\nax[0].set_xticks([])\nax[0].set_yticks([])\n\n_ = ax[1].hist(benign_image_sample[:,:, 0].ravel(), bins=256, color='red', alpha=0.5)\n_ = ax[1].hist(benign_image_sample[:,:, 1].ravel(), bins=256, color='green', alpha=0.7)\n_ = ax[1].hist(benign_image_sample[:,:, 2].ravel(), bins=256, color='blue', alpha=0.5)\n_ = ax[1].set_xlabel('Pixel Intensities')\n_ = ax[1].set_ylabel('Pixel Counts')\n_ = ax[1].legend(['Red Channel', 'Green Channel', 'Blue Channel'])\n_ = plt.suptitle(\"Pixel Intensities for Benign Image\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"malignant_image_sample = plt.imread(os.path.join(TRAIN_IMAGE_PATH, MALIGNANT_IMAGES_LIST[np.random.randint(len(MALIGNANT_IMAGES_LIST))]+'.jpg'))\nfig, ax = plt.subplots(1,2,figsize=(15,5))\nax[0].imshow(malignant_image_sample)\nax[0].set_xlabel(f\"Image dimensions : {malignant_image_sample.shape[:2]}\")\nax[0].set_xticks([])\nax[0].set_yticks([])\n\n_ = ax[1].hist(malignant_image_sample[:,:, 0].ravel(), bins=256, color='red', alpha=0.5)\n_ = ax[1].hist(malignant_image_sample[:,:, 1].ravel(), bins=256, color='green', alpha=0.7)\n_ = ax[1].hist(malignant_image_sample[:,:, 2].ravel(), bins=256, color='blue', alpha=0.5)\n_ = ax[1].set_xlabel('Pixel Intensities')\n_ = ax[1].set_ylabel('Pixel Counts')\n_ = ax[1].legend(['Red Channel', 'Green Channel', 'Blue Channel'])\n_ = plt.suptitle(\"Pixel Intensities for Malignant Image\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**OBSERVATIONS**\n- _RED_ is the color with the highest intensity, almost always, in both cases.\n- For most examples, Green and Blue seems more or less overlapping.\n- Not sure if there is any insight that can be gained from this information but the Red channel could be key.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"---\n**Exporing the DICOM data soon!**\n\nThis is my first Kaggle Notebook and EDA! Any feedback appreciated :) \n\n**Thanks!**","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}