{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Project Topic\n\nThis is a notebook for [Histopathologic Cancer Detection](https://www.kaggle.com/competitions/histopathologic-cancer-detection/overview). The task of this project is to detect cancer in photos taken by microscope. The size of image is 96x96, and when at least one of cancer cell is included in the center 32x32, it is positive. Cancer cells out of this zone is not counted, but accuracy gets much better when whole 96x96 is used. \n\n\n**target**\n\nTarget of this project is to achieve **auc > 0.9** in kaggle private score by CNN. Because of bias in data, it is even possible to achieve auc 0.8 only according to color in photo (refer to **chapter 2.5 prediction by mean RGB value**). Therefore, CNN must achieve more than auc 0.9. \n\n","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport os\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport math\nimport time\nimport cv2\nfrom sklearn import metrics\nimport gc\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-01T03:37:54.191753Z","iopub.execute_input":"2023-05-01T03:37:54.192514Z","iopub.status.idle":"2023-05-01T03:37:54.198123Z","shell.execute_reply.started":"2023-05-01T03:37:54.192473Z","shell.execute_reply":"2023-05-01T03:37:54.196967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers\nfrom tensorflow.keras import models","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:37:57.026527Z","iopub.execute_input":"2023-05-01T03:37:57.026922Z","iopub.status.idle":"2023-05-01T03:38:04.660689Z","shell.execute_reply.started":"2023-05-01T03:37:57.026864Z","shell.execute_reply":"2023-05-01T03:38:04.659526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Data\n\nThere are 220,0025 train data, and it takes about 30min to read all of them. To reduce data preprocessing time, image file (.npy) is created by [another notebook](https://www.kaggle.com/hidetaketakahashi/read-file-for-cancer-detection). Those files can be loaded as numpy array. This file is re-sampled data to balance positive  and negative. \n\n\n**data source**\n\nThe original data is available in kaggle competition website. \n\nhttps://www.kaggle.com/competitions/histopathologic-cancer-detection/data\n\n\nnumpy file is available here:\n\nhttps://www.kaggle.com/datasets/hidetaketakahashi/cancerdetection-npy","metadata":{}},{"cell_type":"code","source":"sample_df = pd.read_csv(\"/kaggle/input/histopathologic-cancer-detection/sample_submission.csv\")\nlabel_df = pd.read_csv(\"/kaggle/input/histopathologic-cancer-detection/train_labels.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:38:04.662871Z","iopub.execute_input":"2023-05-01T03:38:04.663800Z","iopub.status.idle":"2023-05-01T03:38:05.316515Z","shell.execute_reply.started":"2023-05-01T03:38:04.663719Z","shell.execute_reply":"2023-05-01T03:38:05.315220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"n of train data = \", label_df.shape[0])\nprint(\"n of test data  = \", sample_df.shape[0])","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:38:05.319488Z","iopub.execute_input":"2023-05-01T03:38:05.319844Z","iopub.status.idle":"2023-05-01T03:38:05.332067Z","shell.execute_reply.started":"2023-05-01T03:38:05.319802Z","shell.execute_reply":"2023-05-01T03:38:05.330589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fpath = \"/kaggle/input/cancerdetection-npy/X_test.npy\"\nX_test = np.load(fpath)","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:38:05.335321Z","iopub.execute_input":"2023-05-01T03:38:05.336075Z","iopub.status.idle":"2023-05-01T03:38:24.167663Z","shell.execute_reply.started":"2023-05-01T03:38:05.336031Z","shell.execute_reply":"2023-05-01T03:38:24.166583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fpath = \"/kaggle/input/cancerdetection-npy/X_val.npy\"\nX_val = np.load(fpath)\n\nfpath = \"/kaggle/input/cancerdetection-npy/y_val.npy\"\ny_val = np.load(fpath)","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:38:24.169975Z","iopub.execute_input":"2023-05-01T03:38:24.170729Z","iopub.status.idle":"2023-05-01T03:38:27.243255Z","shell.execute_reply.started":"2023-05-01T03:38:24.170687Z","shell.execute_reply":"2023-05-01T03:38:27.242083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Due to memory size restriction, train data size is limited to 50000 for this report.","metadata":{}},{"cell_type":"code","source":"fpath = \"/kaggle/input/cancerdetection-npy/X_train.npy\"\nX_train = np.load(fpath)[0:50000]\n\nfpath = \"/kaggle/input/cancerdetection-npy/y_train.npy\"\ny_train = np.load(fpath)[0:50000]","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:38:27.245821Z","iopub.execute_input":"2023-05-01T03:38:27.246615Z","iopub.status.idle":"2023-05-01T03:39:14.121786Z","shell.execute_reply.started":"2023-05-01T03:38:27.246568Z","shell.execute_reply":"2023-05-01T03:39:14.120754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.shape, X_val.shape, y_train.shape, y_val.shape, X_test.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:39:14.123156Z","iopub.execute_input":"2023-05-01T03:39:14.123671Z","iopub.status.idle":"2023-05-01T03:39:14.132809Z","shell.execute_reply.started":"2023-05-01T03:39:14.123631Z","shell.execute_reply":"2023-05-01T03:39:14.131628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_train = X_train.shape[0]\nn_test = X_test.shape[0]\nn_val = X_val.shape[0]","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:39:14.134546Z","iopub.execute_input":"2023-05-01T03:39:14.135247Z","iopub.status.idle":"2023-05-01T03:39:14.141854Z","shell.execute_reply.started":"2023-05-01T03:39:14.135209Z","shell.execute_reply":"2023-05-01T03:39:14.140692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. EDA","metadata":{}},{"cell_type":"markdown","source":"## 2.1 Labels\n\nIn the original train data, 40% of labels are positive (cancer). To make training efficient, the train data is re-sampled to be 50% positive.","metadata":{}},{"cell_type":"code","source":"label_df[\"label\"].mean(), y_train.mean()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:38:33.896960Z","iopub.execute_input":"2023-04-30T10:38:33.897420Z","iopub.status.idle":"2023-04-30T10:38:33.915637Z","shell.execute_reply.started":"2023-04-30T10:38:33.897383Z","shell.execute_reply":"2023-04-30T10:38:33.914727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize = (9, 3.5))\nax[0].hist(label_df[\"label\"])\nax[1].hist(y_train.flatten(), color = \"darkorange\")\nax[0].set_title(\"labels of original train data\")\nax[1].set_title(\"labels of train data (balanced sampling)\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:38:33.921566Z","iopub.execute_input":"2023-04-30T10:38:33.921836Z","iopub.status.idle":"2023-04-30T10:38:34.272340Z","shell.execute_reply.started":"2023-04-30T10:38:33.921811Z","shell.execute_reply":"2023-04-30T10:38:34.271329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2 Train and Test Distribution in RGB\n\nWhen test data is available, it is better to compare distribution of train and test data. As shown in the histrogram of RGB respectively, train and test has similar distribution. ","metadata":{}},{"cell_type":"code","source":"np.random.seed(1)\nsample_train = np.random.choice(n_train, 5000)\nsample_test = np.random.choice(n_test, 5000)\n\nfig, ax = plt.subplots(1, 3, figsize = (12, 3))\n\ntitles = [\"R\", \"G\", \"B\"]\nfor i in range(3):\n    ax[i].hist(X_train[sample_train, :,:,i].flatten(), density = True, alpha = 0.7, bins = 20, label = \"train\")\n    ax[i].hist(X_test[sample_test, :,:,i].flatten(), density = True, alpha = 0.7, bins = 20, label = \"test\")\n    ax[i].set_title(titles[i])\n    ax[i].legend()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:38:34.273751Z","iopub.execute_input":"2023-04-30T10:38:34.275783Z","iopub.status.idle":"2023-04-30T10:38:40.589987Z","shell.execute_reply.started":"2023-04-30T10:38:34.275743Z","shell.execute_reply":"2023-04-30T10:38:40.588845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.3 Distribution of cancer and normal in RGB\n\n","metadata":{"execution":{"iopub.status.busy":"2023-04-30T01:16:44.996977Z","iopub.execute_input":"2023-04-30T01:16:44.997461Z","iopub.status.idle":"2023-04-30T01:16:45.005889Z","shell.execute_reply.started":"2023-04-30T01:16:44.997421Z","shell.execute_reply":"2023-04-30T01:16:45.004526Z"}}},{"cell_type":"code","source":"sample_pos = y_val.flatten() == 1\nsample_neg = y_val.flatten() == 0\n\nfig, ax = plt.subplots(1, 3, figsize = (12, 3))\n\ntitles = [\"R\", \"G\", \"B\"]\nfor i in range(3):\n    ax[i].hist(X_val[sample_neg, :,:,i].flatten(), density = True, alpha = 0.7, bins = 20, label = \"normal\")\n    ax[i].hist(X_val[sample_pos, :,:,i].flatten(), density = True, alpha = 0.7, bins = 20, label = \"cancer\")\n    ax[i].set_title(titles[i])\n    ax[i].legend()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:38:40.591443Z","iopub.execute_input":"2023-04-30T10:38:40.591907Z","iopub.status.idle":"2023-04-30T10:38:46.272539Z","shell.execute_reply.started":"2023-04-30T10:38:40.591859Z","shell.execute_reply":"2023-04-30T10:38:46.271442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.4 Plot of photo\n\nSince the distribution of cancer and normal are different in RGB, there should be visible difference in photos. According to images, there are lots of pink colored photo in cancer group, this color bias can be information for prediction. ","metadata":{}},{"cell_type":"code","source":"def plot_photo(X):\n    \n    N = X.shape[0]\n    nc = 10\n    nr = math.ceil(N/nc)\n    \n    fig, ax = plt.subplots(nr, nc, figsize = (14, nr*1.4))\n    \n    for k in range(N):\n        i =  int(k/nc)\n        j = k % nc\n        ax[i,j].imshow(X[k])\n        ax[i,j].tick_params(left = False, right = False , labelleft = False, labelbottom = False, bottom = False)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:38:46.273929Z","iopub.execute_input":"2023-04-30T10:38:46.274820Z","iopub.status.idle":"2023-04-30T10:38:46.282451Z","shell.execute_reply.started":"2023-04-30T10:38:46.274782Z","shell.execute_reply":"2023-04-30T10:38:46.281445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Normal (non-cancer)\n\nIn the normal (non-cancer) group, most of photo are **purple** colored. ","metadata":{}},{"cell_type":"code","source":"plot_photo(X_val[sample_neg][0:80])","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:38:46.284180Z","iopub.execute_input":"2023-04-30T10:38:46.284576Z","iopub.status.idle":"2023-04-30T10:38:52.231083Z","shell.execute_reply.started":"2023-04-30T10:38:46.284538Z","shell.execute_reply":"2023-04-30T10:38:52.229754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Cancers\n\nThere are more **pink** colored photo in cancer group compared with normal group.","metadata":{}},{"cell_type":"code","source":"plot_photo(X_val[sample_pos][0:80])","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:38:52.232534Z","iopub.execute_input":"2023-04-30T10:38:52.233782Z","iopub.status.idle":"2023-04-30T10:38:58.034829Z","shell.execute_reply.started":"2023-04-30T10:38:52.233731Z","shell.execute_reply":"2023-04-30T10:38:58.033701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.5 Prediction from mean RGB value\n\nBased on the above EDA, it seems to be possible to predict labels according to color of image. In this section, mean value of RGB of photo is calculated, and this feature is used for RandomForest Classifier. ","metadata":{}},{"cell_type":"code","source":"# Calculation of mean RGB\nZ_train = np.zeros((n_train, 3))\nZ_val = np.zeros((n_val, 3))\nZ_test = np.zeros((n_test, 3))\n\nfor j in range(3):\n\n    for i in range(n_train):\n        Z_train[i,j] = np.mean(X_train[i,:,:,j])\n\n    for i in range(n_val):\n        Z_val[i,j] = np.mean(X_val[i,:,:,j])\n\n    for i in range(n_test):\n        Z_test[i,j] = np.mean(X_test[i,:,:,j])","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:38:58.036015Z","iopub.execute_input":"2023-04-30T10:38:58.036979Z","iopub.status.idle":"2023-04-30T10:39:04.519627Z","shell.execute_reply.started":"2023-04-30T10:38:58.036941Z","shell.execute_reply":"2023-04-30T10:39:04.518565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:39:04.521472Z","iopub.execute_input":"2023-04-30T10:39:04.521902Z","iopub.status.idle":"2023-04-30T10:39:04.810571Z","shell.execute_reply.started":"2023-04-30T10:39:04.521859Z","shell.execute_reply":"2023-04-30T10:39:04.809575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf = RandomForestClassifier(max_depth=10, random_state=0)\nclf.fit(Z_train, y_train.flatten())","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:39:04.812205Z","iopub.execute_input":"2023-04-30T10:39:04.812595Z","iopub.status.idle":"2023-04-30T10:39:09.988879Z","shell.execute_reply.started":"2023-04-30T10:39:04.812554Z","shell.execute_reply":"2023-04-30T10:39:09.987912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As a result, validation auc is **0.844**.","metadata":{}},{"cell_type":"code","source":"yhat_train = clf.predict_proba(Z_train)\nfpr, tpr, thresholds = metrics.roc_curve(y_train, yhat_train[:,1], pos_label=1)\ntrain_auc = metrics.auc(fpr, tpr)\n\nyhat_val = clf.predict_proba(Z_val)\nfpr, tpr, thresholds = metrics.roc_curve(y_val, yhat_val[:,1], pos_label=1)\nval_auc = metrics.auc(fpr, tpr)\n\nprint(\"train auc\", np.round(train_auc, 3))\nprint(\"validation auc\", np.round(val_auc, 3))","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:39:09.990168Z","iopub.execute_input":"2023-04-30T10:39:09.991013Z","iopub.status.idle":"2023-04-30T10:39:10.605023Z","shell.execute_reply.started":"2023-04-30T10:39:09.990975Z","shell.execute_reply":"2023-04-30T10:39:10.603860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This prediction got **0.8151 in kaggle private score**. It means that CNN shall be much more accurate to be worth to use GPU and longer training time. ","metadata":{}},{"cell_type":"code","source":"#yhat_test = clf.predict_proba(Z_test)[:,1]","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:39:10.606527Z","iopub.execute_input":"2023-04-30T10:39:10.607084Z","iopub.status.idle":"2023-04-30T10:39:10.611627Z","shell.execute_reply.started":"2023-04-30T10:39:10.607043Z","shell.execute_reply":"2023-04-30T10:39:10.610572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test_submit = sample_df.copy()\n#test_submit[\"label\"] = yhat_test\n#test_submit.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:39:10.613145Z","iopub.execute_input":"2023-04-30T10:39:10.613503Z","iopub.status.idle":"2023-04-30T10:39:10.626802Z","shell.execute_reply.started":"2023-04-30T10:39:10.613468Z","shell.execute_reply":"2023-04-30T10:39:10.625775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del X_test\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:39:14.146630Z","iopub.execute_input":"2023-05-01T03:39:14.146991Z","iopub.status.idle":"2023-05-01T03:39:14.330991Z","shell.execute_reply.started":"2023-05-01T03:39:14.146963Z","shell.execute_reply":"2023-05-01T03:39:14.329731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.6 Conclusion of EDA\n\n**train and test data have similar distribution**\n\nSince train and test images have similar distributions, it is possible to randomly select validation data from original train data. \n\n**There are bias in data** \n\nCancer and normal have different distribution in RGB. It is even possible to predict cancer at auc 0.8 **only according to the mean color of a image**. Thus, CNN must achieve much higher result.\n\n**Data cleaning plan**\n\nThe type of data is `uint8` that is small but still occupies large size of RAM. Preprocessing data will convert it to `float` which is too big. So, in this project, preprocess functions (`Rescale` and `Randomflip`) will be built in CNN model. ","metadata":{}},{"cell_type":"markdown","source":"# 3. CNN Tuning and Model Selection\n\nIn this chapter, 3 models of CNN are compared. Those models have 5 layers of Conv & MaxPool units.  \n\n**models**\n\n* model1: 5 x (Conv MaxPool)\n* model2: 5 x (Conv Conv MaxPool)\n* model3: 5 x (Conv Conv MaxPool) without Dropout\n\nFollowing elements are common in these three models. Those settings are determined by many trial of training. Reasons are briefly described. \n\n**Common elements**\n\nModels contain rescale and randomflip layer as preprocessing image.  \n\n* Rescaling layer\n* RandomFlip layer\n\nAs I tried some filter sizes, size = 3 was most stable and performance was better (faster learning and good accuracy) than size 5 and 7. MaxPooling layer should be after one or two Conv layers. I have tried Conv with stride = 2 instead of MaxPooling, but MaxPool was more stable. \n\n* Convolusion filter size = 3\n* MaxPooling size = 2 stride = 2\n\nBatch size and learning rate are closely related. I tried batch size 32, 64, and 128. 128 with default learning rate was the best. Furthermore, 50% of dropout is widely used in kaggle, and I can agree. When drop out ratio is small ( < 0.4), effect is not enough. For these model with 50,000 train images, the best epoch seems to be about 8, but I set epoch = 12, to clearly show overfitting symptoms. \n\n* batch size = 128 with adam lr = 0.0001 (default)\n* 50% Dropout in Dense layer (except for model3)\n* The number of epoch is 12 for all three models, which is enough to check overfitting.\n\n\n\n## 5 x (Conv, MaxPool)","metadata":{}},{"cell_type":"code","source":"X_shape = X_train[0].shape","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:39:14.332784Z","iopub.execute_input":"2023-05-01T03:39:14.333939Z","iopub.status.idle":"2023-05-01T03:39:14.338704Z","shell.execute_reply.started":"2023-05-01T03:39:14.333875Z","shell.execute_reply":"2023-05-01T03:39:14.337484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.utils.set_random_seed(\n    seed = 2\n)\n\nmodel1 = models.Sequential()\n\nNF =64\nFS = 3\n\nmodel1.add(layers.Rescaling(scale=1./127.5, offset=-1., input_shape = X_shape))\nmodel1.add(layers.RandomFlip(mode=\"horizontal_and_vertical\", seed=1, input_shape = X_shape))\n\nmodel1.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\", input_shape = X_shape))\nmodel1.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel1.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel1.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel1.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel1.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel1.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel1.add(layers.MaxPooling2D((2,2)))\n\nFS = 3\nNF = NF\nmodel1.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel1.add(layers.MaxPooling2D((2,2)))\n\nmodel1.add(layers.Flatten())\nmodel1.add(layers.Dropout(0.5, seed = 1))\nmodel1.add(layers.Dense(512*9, activation = \"relu\"))\n\nmodel1.add(layers.Dense(2))\n\nmodel1.summary()","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:39:14.340746Z","iopub.execute_input":"2023-05-01T03:39:14.341176Z","iopub.status.idle":"2023-05-01T03:39:16.987503Z","shell.execute_reply.started":"2023-05-01T03:39:14.341139Z","shell.execute_reply":"2023-05-01T03:39:16.986666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)\nmodel1.compile(\n    optimizer = \"adam\",\n    loss=loss_fn, \n    metrics=['accuracy'] \n    )","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:39:16.988555Z","iopub.execute_input":"2023-05-01T03:39:16.988929Z","iopub.status.idle":"2023-05-01T03:39:17.022584Z","shell.execute_reply.started":"2023-05-01T03:39:16.988873Z","shell.execute_reply":"2023-05-01T03:39:17.021500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history1 = model1.fit(X_train, y_train, batch_size = 128, epochs = 12, validation_data = (X_val, y_val))\ngc.collect()\ntf.keras.backend.clear_session()","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:39:17.024039Z","iopub.execute_input":"2023-05-01T03:39:17.025058Z","iopub.status.idle":"2023-05-01T03:46:43.213991Z","shell.execute_reply.started":"2023-05-01T03:39:17.025019Z","shell.execute_reply":"2023-05-01T03:46:43.212946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()\ntf.keras.backend.clear_session()","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:46:43.216122Z","iopub.execute_input":"2023-05-01T03:46:43.216496Z","iopub.status.idle":"2023-05-01T03:46:43.400559Z","shell.execute_reply.started":"2023-05-01T03:46:43.216459Z","shell.execute_reply":"2023-05-01T03:46:43.399340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_scores(history):\n    \n    fig, ax = plt.subplots(1,2, figsize = (12, 5))\n    ax[0].plot(history.history[\"accuracy\"], label = \"train\")\n    ax[0].plot(history.history[\"val_accuracy\"], label = \"val\")\n    ax[0].set_xlabel(\"epochs\")\n    ax[0].set_ylabel(\"accuracy\")\n    ax[0].set_ylim([0.7,1])\n    ax[0].set_title(\"Accuracy\")\n    ax[0].legend()\n\n    ax[1].plot(history.history[\"loss\"], label = \"train\")\n    ax[1].plot(history.history[\"val_loss\"], label = \"val\")\n    ax[1].set_xlabel(\"epochs\")\n    ax[1].set_ylabel(\"loss\")\n    ax[1].set_ylim([0,1])\n    ax[1].set_title(\"Loss\")\n    ax[1].legend()\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:46:43.402102Z","iopub.execute_input":"2023-05-01T03:46:43.402952Z","iopub.status.idle":"2023-05-01T03:46:43.412342Z","shell.execute_reply.started":"2023-05-01T03:46:43.402909Z","shell.execute_reply":"2023-05-01T03:46:43.411332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Teh model started overfitting after epoch 5. Train loss decreased whereas validation loss increased. ","metadata":{}},{"cell_type":"code","source":"plot_scores(history1)","metadata":{"execution":{"iopub.status.busy":"2023-05-01T03:46:43.413682Z","iopub.execute_input":"2023-05-01T03:46:43.414145Z","iopub.status.idle":"2023-05-01T03:46:43.771553Z","shell.execute_reply.started":"2023-05-01T03:46:43.414107Z","shell.execute_reply":"2023-05-01T03:46:43.769978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**validation auc**\n\nTwo predictions are compared. First, prediction from original image. It got **auc 0.957**. Another one is ensemble of predictions from flipped images. It got **auc 0.975**. The result is improved because CNN is sensitive to image's angle.","metadata":{}},{"cell_type":"code","source":"yhat_val = model1.predict(X_val)\nyhat_val = tf.nn.softmax(yhat_val).numpy()[:,1]\n\nfpr, tpr, thresholds = metrics.roc_curve(y_val, yhat_val, pos_label=1)\nval_auc = metrics.auc(fpr, tpr)\nprint(\"val auc\", val_auc)","metadata":{"execution":{"iopub.status.busy":"2023-05-01T04:21:51.501828Z","iopub.execute_input":"2023-05-01T04:21:51.502997Z","iopub.status.idle":"2023-05-01T04:21:53.825139Z","shell.execute_reply.started":"2023-05-01T04:21:51.502943Z","shell.execute_reply":"2023-05-01T04:21:53.824152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def pred_val(model):\n    \n    #original\n    yhat_test1 = model.predict(X_val)\n    yhat_test1 = tf.nn.softmax(yhat_test1).numpy()[:,1]\n    \n    gc.collect()\n    \n    #flipped\n    yhat_test2 = model.predict(X_val[:,::-1,::-1,:])\n    yhat_test2 = tf.nn.softmax(yhat_test2).numpy()[:,1]\n    \n    gc.collect()\n    \n    #flipped\n    yhat_test3 = model.predict(X_val[:,::-1,:,:])\n    yhat_test3 = tf.nn.softmax(yhat_test3).numpy()[:,1]\n    \n    gc.collect()\n    \n    #flipped\n    yhat_test4 = model.predict(X_val[:,:,::-1,:])\n    yhat_test4 = tf.nn.softmax(yhat_test4).numpy()[:,1]\n    \n    gc.collect()\n    \n    #average\n    yhat_test =  (yhat_test1 + yhat_test2 + yhat_test3 + yhat_test4)/4\n    \n    return yhat_test","metadata":{"execution":{"iopub.status.busy":"2023-05-01T04:21:54.731805Z","iopub.execute_input":"2023-05-01T04:21:54.732883Z","iopub.status.idle":"2023-05-01T04:21:54.741357Z","shell.execute_reply.started":"2023-05-01T04:21:54.732831Z","shell.execute_reply":"2023-05-01T04:21:54.740089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"yhat_val = pred_val(model1)\nfpr, tpr, thresholds = metrics.roc_curve(y_val, yhat_val, pos_label=1)\nval_auc = metrics.auc(fpr, tpr)\nprint(\"val auc (ensembled)\", val_auc)","metadata":{"execution":{"iopub.status.busy":"2023-05-01T04:21:55.107300Z","iopub.execute_input":"2023-05-01T04:21:55.108084Z","iopub.status.idle":"2023-05-01T04:22:09.912742Z","shell.execute_reply.started":"2023-05-01T04:21:55.108042Z","shell.execute_reply":"2023-05-01T04:22:09.911548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5 x (Conv Conv MaxPool)","metadata":{}},{"cell_type":"code","source":"tf.keras.utils.set_random_seed(\n    seed = 2\n)\n\nmodel = models.Sequential()\n\nNF =64\nFS = 3\n\nmodel.add(layers.Rescaling(scale=1./127.5, offset=-1., input_shape = X_shape))\nmodel.add(layers.RandomFlip(mode=\"horizontal_and_vertical\", seed=1, input_shape = X_shape))\n\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\", input_shape = X_shape))\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.MaxPooling2D((2,2)))\n\nFS = 3\nNF = NF\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel.add(layers.MaxPooling2D((2,2)))\n\nmodel.add(layers.Flatten())\nmodel.add(layers.Dropout(0.5, seed = 1))\nmodel.add(layers.Dense(512*9, activation = \"relu\"))\n\nmodel.add(layers.Dense(2))\n\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:49:03.398645Z","iopub.execute_input":"2023-04-30T10:49:03.399754Z","iopub.status.idle":"2023-04-30T10:49:03.686048Z","shell.execute_reply.started":"2023-04-30T10:49:03.399702Z","shell.execute_reply":"2023-04-30T10:49:03.685295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)\nmodel.compile(\n    optimizer = \"adam\",\n    loss=loss_fn, \n    metrics=['accuracy'] \n    )","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:49:04.861496Z","iopub.execute_input":"2023-04-30T10:49:04.862536Z","iopub.status.idle":"2023-04-30T10:49:04.878241Z","shell.execute_reply.started":"2023-04-30T10:49:04.862476Z","shell.execute_reply":"2023-04-30T10:49:04.877274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(X_train, y_train, batch_size = 128, epochs = 12, validation_data = (X_val, y_val))","metadata":{"execution":{"iopub.status.busy":"2023-04-30T10:49:05.385239Z","iopub.execute_input":"2023-04-30T10:49:05.385628Z","iopub.status.idle":"2023-04-30T11:02:46.934353Z","shell.execute_reply.started":"2023-04-30T10:49:05.385593Z","shell.execute_reply":"2023-04-30T11:02:46.933316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()\ntf.keras.backend.clear_session()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:02:46.936065Z","iopub.execute_input":"2023-04-30T11:02:46.936388Z","iopub.status.idle":"2023-04-30T11:02:47.279466Z","shell.execute_reply.started":"2023-04-30T11:02:46.936357Z","shell.execute_reply":"2023-04-30T11:02:47.278425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Compared with the previous model, training accuracy is slightly lower but validation accuracy was improved. Overfitting is mitigated by additional conv layer. ","metadata":{}},{"cell_type":"code","source":"plot_scores(history)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:02:47.284133Z","iopub.execute_input":"2023-04-30T11:02:47.284440Z","iopub.status.idle":"2023-04-30T11:02:47.629585Z","shell.execute_reply.started":"2023-04-30T11:02:47.284412Z","shell.execute_reply":"2023-04-30T11:02:47.628549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**validation auc**\n\nPrediction of single image and flipped images are compared. ","metadata":{}},{"cell_type":"code","source":"#single image\nyhat_val = model.predict(X_val)\nyhat_val = tf.nn.softmax(yhat_val).numpy()[:,1]\n\nfpr, tpr, thresholds = metrics.roc_curve(y_val, yhat_val, pos_label=1)\nval_auc = metrics.auc(fpr, tpr)\nprint(\"val auc\", val_auc)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:02:47.632480Z","iopub.execute_input":"2023-04-30T11:02:47.632862Z","iopub.status.idle":"2023-04-30T11:02:53.543495Z","shell.execute_reply.started":"2023-04-30T11:02:47.632823Z","shell.execute_reply":"2023-04-30T11:02:53.542174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#flipped and ensembled\nyhat_val = pred_val(model)\nfpr, tpr, thresholds = metrics.roc_curve(y_val, yhat_val, pos_label=1)\nval_auc = metrics.auc(fpr, tpr)\nprint(\"val auc (ensembled)\", val_auc)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:02:53.545145Z","iopub.execute_input":"2023-04-30T11:02:53.545641Z","iopub.status.idle":"2023-04-30T11:03:16.736185Z","shell.execute_reply.started":"2023-04-30T11:02:53.545599Z","shell.execute_reply":"2023-04-30T11:03:16.733976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"## 5 x (Conv Conv MaxPool) without Dropout\n\nNext, Dropout layer is removed from the previous architecture. Others are the same. ","metadata":{}},{"cell_type":"code","source":"tf.keras.utils.set_random_seed(\n    seed = 2\n)\n\nmodel3 = models.Sequential()\n\nNF =64\nFS = 3\n\nmodel3.add(layers.Rescaling(scale=1./127.5, offset=-1., input_shape = X_shape))\nmodel3.add(layers.RandomFlip(mode=\"horizontal_and_vertical\", seed=1, input_shape = X_shape))\n\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\", input_shape = X_shape))\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.MaxPooling2D((2,2)))\n\n\nFS = 3\nNF = NF*2\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.MaxPooling2D((2,2)))\n\nFS = 3\nNF = NF\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.Conv2D(NF, (FS,FS), activation = \"relu\", padding = \"same\"))\nmodel3.add(layers.MaxPooling2D((2,2)))\n\nmodel3.add(layers.Flatten())\n#model3.add(layers.Dropout(0.5, seed = 1))\nmodel3.add(layers.Dense(512*9, activation = \"relu\"))\n\nmodel3.add(layers.Dense(2))\n\nmodel3.summary()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:03:16.737927Z","iopub.execute_input":"2023-04-30T11:03:16.738267Z","iopub.status.idle":"2023-04-30T11:03:17.058394Z","shell.execute_reply.started":"2023-04-30T11:03:16.738238Z","shell.execute_reply":"2023-04-30T11:03:17.057481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)\nmodel3.compile(\n    optimizer = \"adam\",\n    loss=loss_fn, \n    metrics=['accuracy'] \n    )","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:03:17.059822Z","iopub.execute_input":"2023-04-30T11:03:17.060461Z","iopub.status.idle":"2023-04-30T11:03:17.096739Z","shell.execute_reply.started":"2023-04-30T11:03:17.060427Z","shell.execute_reply":"2023-04-30T11:03:17.095837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history3 = model3.fit(X_train, y_train, batch_size = 128, epochs = 12, validation_data = (X_val, y_val))","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:03:17.099887Z","iopub.execute_input":"2023-04-30T11:03:17.100453Z","iopub.status.idle":"2023-04-30T11:17:43.486723Z","shell.execute_reply.started":"2023-04-30T11:03:17.100422Z","shell.execute_reply":"2023-04-30T11:17:43.485529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()\ntf.keras.backend.clear_session()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:17:43.488423Z","iopub.execute_input":"2023-04-30T11:17:43.489349Z","iopub.status.idle":"2023-04-30T11:17:43.849664Z","shell.execute_reply.started":"2023-04-30T11:17:43.489306Z","shell.execute_reply":"2023-04-30T11:17:43.848567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"train accuracy kept increasing while validation accuracy stayed around 0.9, it is a sort of overfitting.","metadata":{}},{"cell_type":"code","source":"plot_scores(history3)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:17:43.856604Z","iopub.execute_input":"2023-04-30T11:17:43.857347Z","iopub.status.idle":"2023-04-30T11:17:44.216021Z","shell.execute_reply.started":"2023-04-30T11:17:43.857301Z","shell.execute_reply":"2023-04-30T11:17:44.215030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"yhat_val = model3.predict(X_val)\nyhat_val = tf.nn.softmax(yhat_val).numpy()[:,1]\n\nfpr, tpr, thresholds = metrics.roc_curve(y_val, yhat_val, pos_label=1)\nval_auc = metrics.auc(fpr, tpr)\nprint(\"val auc\", val_auc)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:17:44.217429Z","iopub.execute_input":"2023-04-30T11:17:44.218781Z","iopub.status.idle":"2023-04-30T11:17:48.759963Z","shell.execute_reply.started":"2023-04-30T11:17:44.218740Z","shell.execute_reply":"2023-04-30T11:17:48.757870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Results \n\nValidation accuracy and loss of three models are plotted in the below figure. Model1 converged fastest, but after 4 epochs, model 2 exceeded it. The best epoch was 9 (at x = 8 in figure) in model2 that had highest accuracy of all. Thus, the **model 2 is the best architecture**. A similar model in [other notebook](https://www.kaggle.com/code/hidetaketakahashi/cancerdetection-cnn2) achieved auc 0.93 in kaggle private score. `LeakyRelu` activation is selected there to avoid dead neuron problem happened in [the model with relu](https://www.kaggle.com/code/hidetaketakahashi/cancerdetection-cnn2-relu). ","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2, figsize = (12, 5))\nax[0].plot(history1.history[\"val_accuracy\"], label = \"model1\")\nax[0].plot(history.history[\"val_accuracy\"], label = \"model2\", lw = 2)\nax[0].plot(history3.history[\"val_accuracy\"], label = \"model3\")\nax[0].set_xlabel(\"epochs\")\nax[0].set_ylabel(\"accuracy\")\nax[0].set_ylim([0.8,1])\nax[0].set_title(\"Validation Accuracy\")\nax[0].legend()\n\nax[1].plot(history1.history[\"val_loss\"], label = \"model1\")\nax[1].plot(history.history[\"val_loss\"], label = \"model2\", lw = 2)\nax[1].plot(history3.history[\"val_loss\"], label = \"model3\")\nax[1].set_xlabel(\"epochs\")\nax[1].set_ylabel(\"loss\")\nax[1].set_ylim([0,0.7])\nax[1].set_title(\"Validation Loss\")\nax[1].legend()\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:52:03.195922Z","iopub.execute_input":"2023-04-30T11:52:03.196314Z","iopub.status.idle":"2023-04-30T11:52:03.685694Z","shell.execute_reply.started":"2023-04-30T11:52:03.196279Z","shell.execute_reply":"2023-04-30T11:52:03.684574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Analysis: Where is cancer?\n\nEven though CNN is kind of black box in calculation, we can check which part of image makes prediction confident. It is possible by selecting samples of high probability of positive. Then cropping a part of image and check prediction. When probability drops significantly, this area may contain cancer cells.","metadata":{}},{"cell_type":"code","source":"yhat_val = model.predict(X_val)\nyhat_val = tf.nn.softmax(yhat_val).numpy()[:,1]\npos_samples = np.where(yhat_val > 0.98)[0]\npos_samples[0:10]","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:17:48.762269Z","iopub.execute_input":"2023-04-30T11:17:48.762706Z","iopub.status.idle":"2023-04-30T11:17:52.990574Z","shell.execute_reply.started":"2023-04-30T11:17:48.762664Z","shell.execute_reply":"2023-04-30T11:17:52.989557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_analysis(ID, size = 32, stride = 32):\n    print(\"true label = \", y_val[ID][0], \" index\", ID)\n    \n    fig, ax = plt.subplots(figsize = (3, 3))\n    ax.imshow(X_val[ID])\n    ax.set_title(\"original image\")\n    \n    fig, ax = plt.subplots(3, 3, figsize = (8, 8))\n\n    for i in range(3):\n        for j in range(3):\n            X_sample = X_val[ID].copy()\n            start_x = i*stride\n            end_x = i*stride + size\n            start_y = j*stride\n            end_y = j*stride + size\n            \n            #cropping image\n            X_sample[start_x:end_x, start_y:end_y] = 255\n            \n            #predict probability\n            prob = model.predict(X_sample.reshape(1,96,96,3), verbose = 0)\n            prob = tf.nn.softmax(prob).numpy()[:,1]\n            \n            #show cropped image\n            ax[i,j].imshow(X_sample)\n            \n            #write prob as title\n            ax[i,j].set_title(\"p = \" + str(np.round(prob[0], 3)))\n            ax[i,j].tick_params(left = False, right = False , labelleft = False ,\n                labelbottom = False, bottom = False)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:29:00.732192Z","iopub.execute_input":"2023-04-30T11:29:00.733142Z","iopub.status.idle":"2023-04-30T11:29:00.743215Z","shell.execute_reply.started":"2023-04-30T11:29:00.733102Z","shell.execute_reply":"2023-04-30T11:29:00.742114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the following case, probability dropped when **bottom left** zone is copped. ","metadata":{}},{"cell_type":"code","source":"#index 10\nplot_analysis(10, 50)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:29:01.790271Z","iopub.execute_input":"2023-04-30T11:29:01.790675Z","iopub.status.idle":"2023-04-30T11:29:03.504939Z","shell.execute_reply.started":"2023-04-30T11:29:01.790640Z","shell.execute_reply":"2023-04-30T11:29:03.503759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the next case, probability dropped when **upper middle** zone is cropped. ","metadata":{}},{"cell_type":"code","source":"#index 1512\nplot_analysis(1512, 45)","metadata":{"execution":{"iopub.status.busy":"2023-04-30T11:29:13.515745Z","iopub.execute_input":"2023-04-30T11:29:13.516877Z","iopub.status.idle":"2023-04-30T11:29:15.409435Z","shell.execute_reply.started":"2023-04-30T11:29:13.516825Z","shell.execute_reply":"2023-04-30T11:29:15.408386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Conclusion\n\nThe model2 [trained with 100,000 images](https://www.kaggle.com/code/hidetaketakahashi/cancerdetection-cnn2) achieved target auc in kaggle private score. It consists of 5 layers of `Conv` `Conv` `MaxPooling` units with 50% `Dropout` at `FC layer`. Here are some findings from this project.\n\n**Convolution layer migated overfitting**\n\nmodel 1 (**Conv&MaxPool**) converged faster but it had obvious overfitting, and model 2 (**Conv&Conv&MaxPool**) mitigated overfitting. Validation accuracy was also improved by additional conv layers.  \n\n**Dropout layer improved validation accuracy**\n\nModel 2 (with dropout) has better accuracy than model 3 (without dropout) that demonstrated effect of dropout. Dropout is well known technique to mitigate overfit, but it also improves accuracy according to this result.\n\n**Leaky relu solved dead neuron**\n\nIt is not described in this notebook, but by comparing with other notebooks ([leaky relu](https://www.kaggle.com/code/hidetaketakahashi/cancerdetection-cnn2) and [relu](https://www.kaggle.com/hidetaketakahashi/cancerdetection-cnn2-relu)),`Leaky relu` avoided dead neuron caused by `relu` when the model has many layers. When many of neurons are dead, prediction accuracy does not increase from 0.5. \n\n**Flipping image and ensemble predictions is effective**\n\nTo get higher auc, input image for testing is flipped horizontally and vertically. Thus, one image has four prediction and they are ensembled. `yhat` = `mean`(`yhat1`(original), `yhat2`(vertical flip), `yhat3` (horizontal flip), `yhat4` (horizontal and vertical flip)). Almost everytime, it improved auc. \n\n**CNN can be visualized somehow**\n\nIt is difficult to figure out which cells are cancers for non-professional persons. However, some technique of CNN can assist diagnosis. By cropping a part of image, and check probability of cancer by CNN, we can guess what are cancer cells look like.  ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}