{"cells":[{"metadata":{},"cell_type":"markdown","source":"<h1 align = 'center'>Melanoma Detection using CNN</h1>\n\n---\n\nIn this project, I'll be working on a deep learning pipeline that processes the input images of possible melanomic moles/tumors and predicts the probability of the mole(s) being melanomous.\n\n***\nMelanoma, also known as malignant melanoma, is a type of skin cancer that develops from the pigment-producing cells known as melanocytes. Melanomas typically occur in the skin but may rarely occur in the mouth, intestines or eyes.\nAbout 25% of melanomas develop from moles. Changes in a mole that can indicate melanoma include an increase in size, irregular edges, change in color, itchiness or skin breakdown.\n\nThe primary cause of melanoma is ultraviolet light (UV) exposure in those with low levels of the skin pigment melanin. The UV light may be from the sun or other sources, such as tanning devices. Those with many moles, a history of affected family members and poor immune function are at greater risk. A number of rare genetic conditions such as xeroderma pigmentosum also increase the risk.Diagnosis is by biopsy and analysis of any skin lesion that has signs of being potentially cancerous.\n— Source: [Wikipedia](https://en.wikipedia.org/wiki/Melanoma#:~:text=Melanoma%2C%20also%20known%20as%20malignant,or%20eye%20(uveal%20melanoma).)\n\nAlso check out [this page on SkinCancer.org](https://www.skincancer.org/skin-cancer-information/melanoma/) to know more about melanoma.\n***\n\nMelanoma is one of the most aggressive forms of skin cancer. Hence, a highly efficient model capable of early detection of melanoma can prove to be a life saviour.\n\n***","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Importing Project Dependencies\n---\n\nThis notebook will be dedicated to the EDA of the training data. So we will import the dependencies accordingly. ","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport torch\nimport numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom PIL import Image\n\nfrom torch.utils.data import DataLoader, Dataset, random_split\nimport torchvision.models as models\nimport torchvision.transforms as transforms\nimport torch.nn.functional as F\nimport torch.nn as nn\nfrom torchvision.utils import make_grid","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## EDA\n\n---\n\nNow, let us perform some exploratory data analysis on the training set. The first step for this project section is cleaning the data. Let us load the train.csv file and see if the data needs any cleaning.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"Now, we don't need the column 'patient_id', so we will drop it.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.drop(columns = ['patient_id'], axis = 1, inplace = True)\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, let us check for the missing values in our dataframe. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"null_df = train_df.isnull().sum().to_frame()\nnull_df.columns = ['null_vals']\nnull_df['percent_null'] = null_df['null_vals']/len(train_df)\nnull_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we can see, around 0.002% values are missing in the 'sex' column, and around the same number is missing from the 'age_approx' column. In case of the 'anatom_site_general_challenge' column, the number is slightly higher, around 0.01%. However, since these numbers are very small as compared to the actual size of the dataset, we will simply drop these null values from our dataframe. \n\n*One thing to be noted here is that we are only dropping these values for the EDA part. While training the model, these null values won't make any difference since all the images and the target values are all present.*","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"EDA_df = train_df.dropna()\nlen(EDA_df), len(train_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now that we have cleaned our dataset, let us move on to performing the EDA on our data.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"sns.set_style('darkgrid')\n\nf, ax = plt.subplots(1, 2, figsize = (16, 8))\nsns.countplot(y = 'diagnosis', data = EDA_df, ax = ax[0])\nsns.countplot(x = 'target', data = EDA_df, ax = ax[1])\n\nf.show()\n\nprint('Percentage of different diagnosed mole types:')\nEDA_df['diagnosis'].value_counts() / len(EDA_df) * 100","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The primary observation from the above graphs is that most of the cases (more than 80%) that we have in our training dataset don't have a certain diagnosis, i.e., their diagnosis is unknown.\n\nAlso, the percentage of confirmed diagnosed cases of melanoma is only around 1.75% of the total data. This shows that our dataset has a very high sampling bias in terms of the target variable.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Now, let us perform some EDA on the data with confirmed melanoma cases.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"target_grouped = EDA_df.groupby('target')\nconfirmed_df = target_grouped.get_group(1)\nconfirmed_df","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"f, ax = plt.subplots(1,2,figsize = (12, 6))\nsns.countplot( x = 'sex', data = confirmed_df, ax = ax[0])\nsns.violinplot( x = 'sex', y = 'age_approx', data = confirmed_df, inner = 'quartile', ax = ax[1])\nf.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From the above graph, we can observe that- \n* Men tend to show a higher chance of getting melanoma, however this can again be a case of sampling bias. \n* In case of both the males and females, people aged between 40-70 fall in the high-risk bracket in terms of developing melanoma.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"f, ax = plt.subplots(1, 2, figsize = (20, 8))\nsns.countplot( x = 'anatom_site_general_challenge', data = confirmed_df, ax = ax[0])\nsns.countplot( x = 'anatom_site_general_challenge', hue = 'sex', data = confirmed_df, ax = ax[1])\nf.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Observations from the graphs above:\n* The torso has the most number of diagnosed cases of melanoma, followed by the lower extreme and upper extreme regions of the body. However, considering the high sampling bias in our data, it can't be said with certainty that torso is the most common anatomical site for melanomous tumors/moles.\n* In case of number of confirmed cases, the above graphs show that men tend to have a higher chance of developing melanoma as compared to women. \n* This trend is visible for the number of cases in almost all anatomical regions except the oral/genital region, for which women show a significantly higher number of cases as compared to men.\n\n**NOTE**- *These are just observations made on the basis of the visualization given above. These trends may or may not represent actual statistics.*","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Now that we are done with the basic EDA part, let us view a few images from our dataset.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"sns.set_style('white')\n\ndef train_img_viewer(index):\n    \"\"\"Shows image in the training dataset at the random index that the user provides.\n    \n    Args-\n        index- index of the image\n    Returns-\n        None\n    \"\"\"\n    \n    img_info = train_df.loc[index,['image_name', 'benign_malignant']]\n    img_name = img_info.image_name + '.jpg'\n    img_class = img_info.benign_malignant\n    path = '../input/siim-isic-melanoma-classification/jpeg/train'\n    img_path = os.path.join(path, img_name)\n    img = Image.open(img_path)\n    transform = transforms.ToTensor()\n    img = transform(img)\n    print(img_class)\n    plt.imshow(img.permute(1,2,0))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_img_viewer(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_img_viewer(91)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"With this, we come to the end of the first part of the project. In the next part, we will define and train our binary classification model.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}