{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Review of Records and Images","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"This notebook is for Kaggkle \"SIIM-ISIC Melanoma Classification\" challenge at https://www.kaggle.com/c/siim-isic-melanoma-classification.\n\nMuch of the code here is based on https://www.kaggle.com/nxrprime/siim-d3-eda-augmentations-and-resnext, https://www.kaggle.com/parulpandey/melanoma-classification-eda-starter, and some other sources. I am grateful to their authors.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Import packages.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nfrom os import listdir\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nimport pydicom","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Provide parameters.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dir_top = '../input/siim-isic-melanoma-classification/'\n\ndir_ts_dcm = dir_top + 'test/'\ndir_tr_dcm = dir_top + 'train/'\n\ndir_ts_jpg = dir_top + 'jpeg/test/'\ndir_tr_jpg = dir_top + 'jpeg/train/'\n\nfile_skin = \"https://lipy.us/docs/SkinAreas.png\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Load CSV datasets.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_ts = pd.read_csv(dir_top + 'test.csv')\ndf_tr = pd.read_csv(dir_top + 'train.csv')\ndf_ts = df_ts.rename(columns={'anatom_site_general_challenge':'anatom_site','age_approx':'age_apx','benign_malignant':'b_m'})\ndf_tr = df_tr.rename(columns={'anatom_site_general_challenge':'anatom_site','age_approx':'age_apx','benign_malignant':'b_m'})\ndf_tr.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_tr_b = df_tr[ df_tr['b_m']=='benign' ]\ndf_tr_m = df_tr[ df_tr['b_m']=='malignant' ]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Handle missing values.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def miss_one(df, s):\n    \n    ros = len(df)\n    mis = df.isnull().sum()\n    new = df.fillna(0).astype('str')\n    \n    tab = pd.concat([mis, round(100*mis/ros,1)], axis=1)\n    tab = tab.rename(columns = {0:s+' Mis', 1:s+' Mis%'})\n    tab = tab[ tab.iloc[:,1] != 0 ]\n    \n    msg = s + \" has \" + str(ros) + \" rows and \" + str(df.shape[1]) + \" columns. \"\n    msg = msg + str(tab[tab.iloc[:,1]!=0].shape[0]) + \" columns have missing values.\"    \n    return new, tab, msg","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_ts, tab_ts, msg_ts = miss_one(df_ts, \"Test\")\ndf_tr, tab_tr, msg_tr = miss_one(df_tr, \"Train\")\ndf_tr_b, tab_tr_b, msg_tr_b = miss_one(df_tr_b, \"Train Ben\")\ndf_tr_m, tab_tr_m, msg_tr_m = miss_one(df_tr_m, \"Train Mal\")\n\ntab_all = pd.concat([tab_ts, tab_tr, tab_tr_b, tab_tr_m], axis=1)\ntab_all = tab_all.sort_values('Train Mis%', ascending=False)\n\nprint(msg_ts + \"\\n\" + msg_tr + \"\\n\" + msg_tr_b + \"\\n\" + msg_tr_m)\ntab_all","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Define grouping by a column name.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def divby(a, b):\n    if b==0:\n        return 0\n    else:\n        return round(100 * a / b, 0)\n    \ndef gdfplot(gpc, ref, val):\n    \n    if gpc not in df_ts.columns:\n        df_ts[gpc] = '?'\n    \n    dfTs = df_ts.groupby([gpc]).size()\n    dfTr = df_tr.groupby([gpc]).size()\n    dfTrM = df_tr_m.groupby([gpc]).size()    \n    \n    dfC = pd.concat([dfTs, dfTr, dfTrM], axis=1).reset_index()\n    dfC = dfC.rename(columns={0:'Test', 1:'Train', 2:'MalTrain', 'index':gpc}).fillna(0)\n    dfC['MalPerc'] = dfC.apply(lambda x: divby(x.MalTrain, x.Train), axis = 1)\n    if ref != '':\n        dfC[ref] = val\n        \n    print(dfC)\n    \n    dfD = dfC.drop([gpc], axis=1)\n    dfD = round(100 * (dfD - dfD.min())/(dfD.max()-dfD.min()), 0)\n    dfD[gpc] = dfC[gpc].apply(lambda x: x[:10])\n    print(\"\\nNormalized values\")\n    print(dfD)\n    \n    plt.figure(figsize=(13, 2), dpi= 80, facecolor='w', edgecolor='k')    \n    plt.plot(gpc, 'Test', data=dfD, marker='', color='red', linewidth=1, label='Test')\n    plt.plot(gpc, 'Train', data=dfD, marker='', color='yellow', linewidth=1, label='Train')\n    plt.plot(gpc, 'MalTrain', data=dfD, marker='', color='green', linewidth=1, label='MalTrain')\n    plt.plot(gpc, 'MalPerc', data=dfD, marker='', color='blue', linewidth=1, label='MalTrain')\n    if ref != '':\n        plt.plot(gpc, ref, data=dfD, marker='', color='olive', linewidth=1, label=ref)\n    plt.legend(bbox_to_anchor=(1.15, 1.0))\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Group by anatomical site.\n\nLund-Browder diagram is based on https://en.wikipedia.org/wiki/Lund_and_Browder_chart and https://www.ncbi.nlm.nih.gov/pmc/articles/PMC449823/.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(8, 8))\nplt.imshow(mpimg.imread(file_skin))\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"scrolled":false,"trusted":true},"cell_type":"code","source":"gdfplot('anatom_site', 'SitePerc', [0, 8, 32, 2, 11, 32, 15])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It seems that test count, training count, and malign training count for a site generally increase with skin area of the site. Ratio of each count to skin area is higher for head/neck, probably because of the visibility of head/neck.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Group by age.\n\nPopulation distribution by age is based on https://www.census.gov/data/.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"gdfplot('age_apx', 'AgePerc', [0,12.2,6.4,6.4,6.6,7.2,6.8,6.6,6.0,6.3,6.3,6.5,6.3,5.4,4.4,2.9,1.9,1.8])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It seems that test count, training count, and malign training count per population are higher for adults, probably because of higher outdoor activities.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Group by gender.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"gdfplot('sex', 'SexPerc', [0, 50, 50])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It seems that test count, training count, mand align training count per population are higher for males, probably because of higher outdoor activities.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Group by diagnosis.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"gdfplot('diagnosis', '', [])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Review samples of images.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"jpg_ts = dir_ts_jpg + df_ts['image_name'] + '.jpg'\njpg_tr_b = dir_tr_jpg + df_tr_m['image_name'] + '.jpg'\njpg_tr_m = dir_tr_jpg + df_tr_m['image_name'] + '.jpg'\n\ndcm_ts = dir_ts_dcm + df_ts['image_name'] + '.dcm'\ndcm_tr_b = dir_tr_dcm + df_tr_m['image_name'] + '.dcm'\ndcm_tr_m = dir_tr_dcm + df_tr_m['image_name'] + '.dcm'\n\nsam_jpg_tr_b = jpg_tr_b.sample(n=6, replace=True, axis=0).reset_index()\nsam_jpg_tr_m = jpg_tr_m.sample(n=6, replace=True, axis=0).reset_index()\n\nsam_dcm_tr_b = dcm_tr_b.sample(n=6, replace=True, axis=0).reset_index()\nsam_dcm_tr_m = dcm_tr_m.sample(n=6, replace=True, axis=0).reset_index()\n\nprint(sam_jpg_tr_b.iloc[0,1])","execution_count":null,"outputs":[]},{"metadata":{"scrolled":false,"trusted":true},"cell_type":"code","source":"fig, axs = plt.subplots(nrows=6, ncols=5, figsize=(15, 12))\n\ndef jpgplot(H, r, c):\n    \n    img = mpimg.imread(H.iloc[c, 1])    \n    axs[r, c].imshow(img, cmap='gray')\n    axs[r, c].set_xticklabels([])\n    axs[r, c].set_yticklabels([])\n    \n    axs[r+2, c].hist(img[:, :, 0].ravel(), bins = 256, color = 'Red', alpha = 0.5)\n    axs[r+2, c].hist(img[:, :, 1].ravel(), bins = 256, color = 'Green', alpha = 0.5)\n    axs[r+2, c].hist(img[:, :, 2].ravel(), bins = 256, color = 'Blue', alpha = 0.5)\n    axs[r+2, c].set_xticklabels([])\n    axs[r+2, c].set_yticklabels([])\n    \ndef dcmplot(H, r, c):\n    \n    img = pydicom.dcmread(H.iloc[c, 1])\n    axs[r, c].imshow(-img.pixel_array, cmap=plt.cm.bone)\n    axs[r, c].set_xticklabels([])\n    axs[r, c].set_yticklabels([])\n    \nfor c in range(5):\n    \n    jpgplot(sam_jpg_tr_b, 0, c)\n    jpgplot(sam_jpg_tr_m, 1, c)\n    dcmplot(sam_dcm_tr_b, 4, c)\n    dcmplot(sam_dcm_tr_m, 5, c)\n    \nfor axe, col in zip(axs[0], ['Sample 1','Sample 2','Sample 3','Sample 4','Sample 5']):\n    axe.set_title(col, rotation=0, size='small')\n    \nfor axe, row in zip(axs[:,0], ['Benign Jpg','Malign Jpg','Benign Histo','Malign Histo','Benign Dcm','Malign Histo']):\n    axe.set_ylabel(row, rotation=90, size='small')\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It seems that color channel histograms cannot be used to classify skin images into benign and malign.","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}