{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Contents","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Loading train and test data","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom os import listdir\nimport time\nimport sys\n\nfrom pydicom import dcmread\nimport matplotlib.pyplot as plt\nimport matplotlib.colors\nimport seaborn as sns\n\nfrom scipy.stats import kurtosis, skew, mode\nfrom scipy import stats\n\n\nfrom PIL import Image\nimport cv2\nfrom skimage import color\n\npath = \"/kaggle/input/siim-isic-melanoma-classification/\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(path + 'train.csv')\ntrain_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df = pd.read_csv(path + 'test.csv')\ntest_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we see above, we have 33126 images in the train set, and 10982 images in the test set. There are 8 columns in the train dataframe, wheras there are 5 columns in the test dataframe. The three columns, \"diagnosis\",\t\"benign_malignant\",\t\"target\" are the target columns.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[\"target\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[\"benign_malignant\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[\"diagnosis\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[(train_df[\"diagnosis\"] == \"melanoma\") &(train_df[\"benign_malignant\"] == \"malignant\") & (train_df[\"target\"] == 1)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we see above, the \"target\" values of 1 only appears when the \"benign_malignant\" column is \"malignant\" and the \"diagnosis\" column is \"melanoma\". There are 584 such cases, whereas the resting 32542 cases are marked as 0 in the \"target\" column. There is a big imbalance between the target values, targets with value of 1 constitutes only 584/32542, 1.8% of all the train set.\n\nWe also notice that the majority of the \"diagnosis\" falls into \"unknown\" class with the ratio of 27124/32542, 83%.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[\"image_name\"].nunique()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df[\"image_name\"].nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"\"image_name\"s are unique.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[\"patient_id\"].nunique()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df[\"patient_id\"].nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are only 2056 patients in the train set and 690 images in the test set, which means patients have several images in the set.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[\"patient_id\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df[\"patient_id\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The maximum number of images per patient is 115 in the train set and 240 images in the test set, wheras, the minimum is 2 in the train set and 3 in the test set.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"a = train_df[\"age_approx\"].unique()\na.sort()\na","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[train_df[\"age_approx\"].isna()].shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"a = test_df[\"age_approx\"].unique()\na.sort()\na","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df[test_df[\"age_approx\"].isna()].shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The ages are in 5-years bin and between 0-90 in the train set and 10-90 in the test set.\nThere are 68 null value in the \"age_approx\" whereas there isn't any in the test set.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"The distribution graphs of ages in the datasets are as follows:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1,2)\n\nsns.distplot(train_df[train_df[\"age_approx\"].notna()][\"age_approx\"], ax=ax[0], color=\"#992299\")\nax[0].set_title(\"distribution of age in train\")\n    \nsns.distplot(test_df[test_df[\"age_approx\"].notna()][\"age_approx\"], ax=ax[1], color=\"#ee2200\")\nax[1].set_title(\"distribution of age in test\");\n    \nfig.set_size_inches(10, 3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[\"sex\"].value_counts(normalize=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df[\"sex\"].value_counts(normalize=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Test dataset has more females than train dataset.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"\"anatom_site_general_challenge\" column has the values distributed like:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[\"anatom_site_general_challenge\"].value_counts(normalize=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df[\"anatom_site_general_challenge\"].value_counts(normalize=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1,2, figsize=(20,8))\n\na = train_df[\"anatom_site_general_challenge\"].value_counts()\n\nf = sns.barplot(x=a.keys(),y=a.values, ax=ax[0])\nax[0].set_title(\"anatomical sites in train\")\n\nn=0\nfor key in a.keys():\n    f.text(n,a[n]+200 , a[key], color='black', ha=\"center\")\n    n+=1\n\n    \na = test_df[\"anatom_site_general_challenge\"].value_counts()\n\nf = sns.barplot(x=a.keys(),y=a.values, ax=ax[1])\nax[1].set_title(\"anatomical sites in test\")\n\nn=0\nfor key in a.keys():\n    f.text(n,a[n]+100 , a[key], color='black', ha=\"center\")\n    n+=1 ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we see above, the distribution of anatomical sites in the train and test datasets are similar. The majority of the images are form torso, and the lowest number of images are from oral/genital sites.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Now, let's how many null values present in the columns:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"for col in train_df.columns:\n    print(col,\":\",len(train_df[train_df[col].isna()]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for col in test_df.columns:\n    print(col,\":\",len(test_df[test_df[col].isna()]))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Images","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Let's have a look at some images:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"image_classes = train_df[\"diagnosis\"].unique()\n\nfig, ax = plt.subplots(len(image_classes),5,figsize=(50,50))\n\nm=0\nfor imclass in image_classes:\n    image_names = train_df.loc[train_df[\"diagnosis\"]==imclass,\"image_name\"].values[:5]\n    n=0\n    for image_name in image_names:\n        image = cv2.imread(path + \"jpeg/train/\" + image_name + \".jpg\")\n        ax[m,n].imshow(cv2.cvtColor(image, cv2.COLOR_BGR2RGB));\n        if (n == 2) | (m>6):\n            ax[m,n].set_title(label=imclass, fontdict={'fontsize':50})\n        n+=1\n    m+=1","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}