{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection with CNN\n\n### The problem and data\n\nThis project focuses on creating an algorithm to identify metastatic cancer in small image patches taken from larger digital pathology scans. In this dataset, we are provided with a large number of small pathology images to classify. Files are named with an image id. The *train_labels.csv* file provides the ground truth for the images in the train folder. We need to predicting the labels for the images in the test folder. A positive label indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue. Tumor tissue in the outer region of the patch does not influence the label. This outer region is provided to enable fully-convolutional models that do not use zero-padding, to ensure consistent behavior when applied to a whole-slide image.\n\nIn this project, we will explore the data, create two CNN models, fit the models into data, and finally compare their performance and select the model. ","metadata":{"papermill":{"duration":0.008585,"end_time":"2022-11-23T23:29:11.948503","exception":false,"start_time":"2022-11-23T23:29:11.939918","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# General libraries\nimport os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n%matplotlib inline\nfrom sklearn.utils import shuffle\nimport fnmatch\nimport random\nos.system('pip install visualkeras')\nimport visualkeras\nimport cv2 as cv\n\n#Model\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3' \nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.layers import RandomFlip, RandomZoom, RandomRotation\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, AveragePooling2D\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten, Activation, ZeroPadding2D\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import BatchNormalization\nfrom tensorflow.keras.optimizers import Adam","metadata":{"papermill":{"duration":13.399212,"end_time":"2022-11-23T23:29:25.354872","exception":false,"start_time":"2022-11-23T23:29:11.955660","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:53:54.697889Z","iopub.execute_input":"2023-02-17T13:53:54.698518Z","iopub.status.idle":"2023-02-17T13:54:08.333040Z","shell.execute_reply.started":"2023-02-17T13:53:54.698410Z","shell.execute_reply":"2023-02-17T13:54:08.331736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explore the data folder structure\nHere we are going to explore the data folder structure and load the names of the image files into **train_data** and **test_data**","metadata":{}},{"cell_type":"code","source":"#list the data folder structure\nos.listdir('../input/histopathologic-cancer-detection/')","metadata":{"papermill":{"duration":0.021008,"end_time":"2022-11-23T23:29:25.383358","exception":false,"start_time":"2022-11-23T23:29:25.362350","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:08.334965Z","iopub.execute_input":"2023-02-17T13:54:08.335717Z","iopub.status.idle":"2023-02-17T13:54:08.346935Z","shell.execute_reply.started":"2023-02-17T13:54:08.335668Z","shell.execute_reply":"2023-02-17T13:54:08.345790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count number of tif images in train and test folders \ntrain_path = '../input/histopathologic-cancer-detection/train/'\ntest_path = '../input/histopathologic-cancer-detection/test/'\nprint(\"# of TIFF images in train folder = \",len(fnmatch.filter(os.listdir(train_path), '*.*')))\nprint(\"# of TIFF images in test folder = \",len(fnmatch.filter(os.listdir(test_path), '*.*')))\n","metadata":{"papermill":{"duration":3.691961,"end_time":"2022-11-23T23:29:29.082053","exception":false,"start_time":"2022-11-23T23:29:25.390092","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:08.348612Z","iopub.execute_input":"2023-02-17T13:54:08.348986Z","iopub.status.idle":"2023-02-17T13:54:08.677318Z","shell.execute_reply.started":"2023-02-17T13:54:08.348954Z","shell.execute_reply":"2023-02-17T13:54:08.675900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')\nsample_submission = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\n\ntest_data = pd.DataFrame({'id':os.listdir(test_path)})\n\n","metadata":{"papermill":{"duration":0.763717,"end_time":"2022-11-23T23:29:29.853067","exception":false,"start_time":"2022-11-23T23:29:29.089350","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:08.680410Z","iopub.execute_input":"2023-02-17T13:54:08.681389Z","iopub.status.idle":"2023-02-17T13:54:09.041228Z","shell.execute_reply.started":"2023-02-17T13:54:08.681337Z","shell.execute_reply":"2023-02-17T13:54:09.039835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Exploratory Data Analysis","metadata":{"papermill":{"duration":0.006733,"end_time":"2022-11-23T23:29:29.891564","exception":false,"start_time":"2022-11-23T23:29:29.884831","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(\"---------- train data ---------------\")\nprint(train_data.head())\ntrain_data['id'] = train_data['id'] + '.tif'\ntrain_data['label'] = train_data['label'].astype(str)\nprint(\"---------- train data (after appending file extension)---------------\")\nprint(train_data.head())\nprint(\"---------- test data ---------------\")\nprint(test_data.head())\n\nprint(train_data.shape)\n\nprint(train_data.info())\n\nprint(train_data.label.unique())","metadata":{"papermill":{"duration":0.030491,"end_time":"2022-11-23T23:29:29.929097","exception":false,"start_time":"2022-11-23T23:29:29.898606","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:09.042945Z","iopub.execute_input":"2023-02-17T13:54:09.043483Z","iopub.status.idle":"2023-02-17T13:54:09.287160Z","shell.execute_reply.started":"2023-02-17T13:54:09.043409Z","shell.execute_reply":"2023-02-17T13:54:09.285652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Visualizations\n\nFrom above EDA we noticed there are only two distinct values avaible in label column 0 and 1, which represents **non-cancerous** and **canerous** respectively.\nHere we would like to render a pie chart to get a sense on the percentage for these two categories.","metadata":{"papermill":{"duration":0.007166,"end_time":"2022-11-23T23:29:29.972730","exception":false,"start_time":"2022-11-23T23:29:29.965564","status":"completed"},"tags":[]}},{"cell_type":"code","source":"unique_counts = train_data['label'].value_counts()\nprint(unique_counts)\n\nplt.pie(unique_counts.tolist(), \n        labels = ['0-Non-cancerous','1-Cancerous'],\n        autopct='%1.2f%%'\n       )\n","metadata":{"papermill":{"duration":0.233584,"end_time":"2022-11-23T23:29:30.492275","exception":false,"start_time":"2022-11-23T23:29:30.258691","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:09.289357Z","iopub.execute_input":"2023-02-17T13:54:09.290315Z","iopub.status.idle":"2023-02-17T13:54:09.427276Z","shell.execute_reply.started":"2023-02-17T13:54:09.290242Z","shell.execute_reply":"2023-02-17T13:54:09.425865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pie chart reveals the percentage of two categories 0-Non-cancerous and 1-Cancerous are 59.50% and 40.50% respectively.\n\nSome sample images from the train data are also randomly presented below:","metadata":{"papermill":{"duration":0.023054,"end_time":"2022-11-23T23:29:30.539036","exception":false,"start_time":"2022-11-23T23:29:30.515982","status":"completed"},"tags":[]}},{"cell_type":"code","source":"plt.figure(1, figsize=(16, 16))\nn = 0\nfor i in range(16):\n  n += 1\n  random_img = random.choice(train_data.id)\n  imgs = cv.imread(train_path+random_img)\n  plt.subplot(4, 4, n)\n  plt.imshow(imgs)","metadata":{"papermill":{"duration":1.352968,"end_time":"2022-11-23T23:29:31.909665","exception":false,"start_time":"2022-11-23T23:29:30.556697","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:09.429385Z","iopub.execute_input":"2023-02-17T13:54:09.429874Z","iopub.status.idle":"2023-02-17T13:54:12.610587Z","shell.execute_reply.started":"2023-02-17T13:54:09.429825Z","shell.execute_reply":"2023-02-17T13:54:12.608866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Architecture\n\nWe will be using convolutional neural network (CNN) to solve this binary classification problem:\n\n\nThe first model will start with a typical structure like: input --> Conv2D --> Conv2D --> MaxPool --> Conv2D --> Conv2D --> MaxPool --> Flatten --> Output.  Relu activation on hiden layers and sigmoid activation on the output layer.\n\nImages will be normalized by taking the pixels and dividing each pixel by 255 to have values between 0-1.\n\nThe second model will repeat Conv2D --> Conv2D --> MaxPool one more round and add dropout and batch nomalization to avoid overfit.\n\nA text summary and diagram of the model architecture will be printed after the model construct for better presentation.","metadata":{"papermill":{"duration":0.021432,"end_time":"2022-11-23T23:29:31.997266","exception":false,"start_time":"2022-11-23T23:29:31.975834","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# set model constants\nbatchSize = 256","metadata":{"papermill":{"duration":0.030522,"end_time":"2022-11-23T23:29:32.049766","exception":false,"start_time":"2022-11-23T23:29:32.019244","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:12.613000Z","iopub.execute_input":"2023-02-17T13:54:12.613648Z","iopub.status.idle":"2023-02-17T13:54:12.618069Z","shell.execute_reply.started":"2023-02-17T13:54:12.613607Z","shell.execute_reply":"2023-02-17T13:54:12.616939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# modify training data by normalizing it \n# and split data into training and validation sets\ndatagen = ImageDataGenerator(rescale=1./255.,\n                            validation_split=0.15)","metadata":{"papermill":{"duration":0.031821,"end_time":"2022-11-23T23:29:32.386600","exception":false,"start_time":"2022-11-23T23:29:32.354779","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:12.619818Z","iopub.execute_input":"2023-02-17T13:54:12.620172Z","iopub.status.idle":"2023-02-17T13:54:12.631034Z","shell.execute_reply.started":"2023-02-17T13:54:12.620138Z","shell.execute_reply":"2023-02-17T13:54:12.629964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate training data\ntrain_generator = datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"training\",\n    batch_size=batchSize,\n    seed=985723,\n    class_mode=\"binary\",\n    target_size=(64,64))  ","metadata":{"papermill":{"duration":542.771753,"end_time":"2022-11-23T23:38:35.180264","exception":false,"start_time":"2022-11-23T23:29:32.408511","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:12.635490Z","iopub.execute_input":"2023-02-17T13:54:12.636192Z","iopub.status.idle":"2023-02-17T13:54:29.234178Z","shell.execute_reply.started":"2023-02-17T13:54:12.636150Z","shell.execute_reply":"2023-02-17T13:54:29.231897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate validation data\nvalid_generator = datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"validation\",\n    batch_size=batchSize,\n    seed=985723,\n    class_mode=\"binary\",\n    target_size=(64,64))  ","metadata":{"papermill":{"duration":215.394704,"end_time":"2022-11-23T23:42:10.597302","exception":false,"start_time":"2022-11-23T23:38:35.202598","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.235486Z","iopub.status.idle":"2023-02-17T13:54:29.236238Z","shell.execute_reply.started":"2023-02-17T13:54:29.235927Z","shell.execute_reply":"2023-02-17T13:54:29.235963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Model 1 construct**","metadata":{"papermill":{"duration":0.021282,"end_time":"2022-11-23T23:42:10.706496","exception":false,"start_time":"2022-11-23T23:42:10.685214","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# tpu = None\n# try:\n#     tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n#     tf.config.experimental_connect_to_cluster(tpu)\n#     tf.tpu.experimental.initialize_tpu_system(tpu)\n#     strategy = tf.distribute.TPUStrategy(tpu)\n# except ValueError:\n#     strategy = tf.distribute.get_strategy()","metadata":{"execution":{"iopub.status.busy":"2023-02-17T13:54:29.239068Z","iopub.status.idle":"2023-02-17T13:54:29.240029Z","shell.execute_reply.started":"2023-02-17T13:54:29.239704Z","shell.execute_reply":"2023-02-17T13:54:29.239735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model1_auc = tf.keras.metrics.AUC()\n\n\n# with tpu_strategy.scope():    \n#create model\nmodel1 = Sequential()\n\nmodel1.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\nmodel1.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\nmodel1.add(MaxPooling2D(pool_size=(2,2)))\n\nmodel1.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\nmodel1.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\nmodel1.add(MaxPooling2D(pool_size=(2,2)))\n\nmodel1.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\nmodel1.add(Flatten())\nmodel1.add(Dense(1, activation='sigmoid'))\n\n#build model by input size\nmodel1.build(input_shape=(batchSize, 64, 64, 3))\n\n#compile\nadam_optimizer = Adam(learning_rate=0.0001)\nmodel1.compile(loss='binary_crossentropy', metrics=['accuracy', model1_auc], optimizer=adam_optimizer)\n\n#Summary of model\nmodel1.summary()\n\nvisualkeras.layered_view(model1, type_ignore=[ZeroPadding2D, Flatten], legend=True)","metadata":{"papermill":{"duration":5.631386,"end_time":"2022-11-23T23:42:16.359353","exception":false,"start_time":"2022-11-23T23:42:10.727967","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.241842Z","iopub.status.idle":"2023-02-17T13:54:29.242820Z","shell.execute_reply.started":"2023-02-17T13:54:29.242493Z","shell.execute_reply":"2023-02-17T13:54:29.242525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EPOCHS = 10\n# train the model\nhistory_model1 = model1.fit(\n                        train_generator,\n                        epochs = EPOCHS,\n                        validation_data = valid_generator)","metadata":{"papermill":{"duration":5558.481333,"end_time":"2022-11-24T01:14:54.864336","exception":false,"start_time":"2022-11-23T23:42:16.383003","status":"completed"},"scrolled":true,"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.244719Z","iopub.status.idle":"2023-02-17T13:54:29.245672Z","shell.execute_reply.started":"2023-02-17T13:54:29.245368Z","shell.execute_reply":"2023-02-17T13:54:29.245401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(history_model1.history.keys())\n\n# plot model accuracy per epoch \nplt.plot(history_model1.history['accuracy'])\nplt.plot(history_model1.history['val_accuracy'])\nplt.title('Model One Accuracy vs epoch')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\n# plot model loss per epoch\nplt.plot(history_model1.history['loss'])\nplt.plot(history_model1.history['val_loss'])\nplt.title('Model 1 Loss vs Epoch')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\n# plot model ROC per epoch\nplt.plot(history_model1.history['auc'])\nplt.plot(history_model1.history['val_auc'])\nplt.title('Model 1 AUC ROC vs epoch')\nplt.ylabel('ROC')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();","metadata":{"papermill":{"duration":1.362487,"end_time":"2022-11-24T01:14:56.641285","exception":false,"start_time":"2022-11-24T01:14:55.278798","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.247469Z","iopub.status.idle":"2023-02-17T13:54:29.248436Z","shell.execute_reply.started":"2023-02-17T13:54:29.248033Z","shell.execute_reply":"2023-02-17T13:54:29.248063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Model build 2**","metadata":{"papermill":{"duration":0.406122,"end_time":"2022-11-24T01:14:57.514389","exception":false,"start_time":"2022-11-24T01:14:57.108267","status":"completed"},"tags":[]}},{"cell_type":"code","source":"model2_auc = tf.keras.metrics.AUC()\n    \n# with tpu_strategy.scope():       \n#create model\nmodel2 = Sequential()\n\nmodel2.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\nmodel2.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\nmodel2.add(MaxPooling2D(pool_size=(2,2)))\nmodel2.add(Dropout(0.1))\n\nmodel2.add(BatchNormalization())\nmodel2.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\nmodel2.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\nmodel2.add(MaxPooling2D(pool_size=(2,2)))\nmodel2.add(Dropout(0.1))\n\nmodel2.add(BatchNormalization())\nmodel2.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\nmodel2.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\nmodel2.add(MaxPooling2D(pool_size=(2,2)))\n\nmodel2.add(BatchNormalization())\nmodel2.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\nmodel2.add(Flatten())\nmodel2.add(Dense(1, activation='sigmoid'))\n\n#build model by input size\nmodel2.build(input_shape=(batchSize, 64, 64, 3))\n\n#compile\nadam_optimizer = Adam(learning_rate=0.0001)\nmodel2.compile(loss='binary_crossentropy', metrics=['accuracy', model2_auc], optimizer=adam_optimizer)\n\n# Summary of the 2nd model\nmodel2.summary()\n\nvisualkeras.layered_view(model2, type_ignore=[ZeroPadding2D, Flatten], legend=True)","metadata":{"papermill":{"duration":0.509497,"end_time":"2022-11-24T01:14:58.433827","exception":false,"start_time":"2022-11-24T01:14:57.924330","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.249910Z","iopub.status.idle":"2023-02-17T13:54:29.250362Z","shell.execute_reply.started":"2023-02-17T13:54:29.250143Z","shell.execute_reply":"2023-02-17T13:54:29.250162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EPOCHS = 10\n\n# train model\nhistory_model2 = model2.fit(\n                        train_generator,\n                        epochs = EPOCHS,\n                        validation_data = valid_generator)","metadata":{"papermill":{"duration":5150.514514,"end_time":"2022-11-24T02:40:49.366365","exception":false,"start_time":"2022-11-24T01:14:58.851851","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.251827Z","iopub.status.idle":"2023-02-17T13:54:29.252291Z","shell.execute_reply.started":"2023-02-17T13:54:29.252069Z","shell.execute_reply":"2023-02-17T13:54:29.252089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history_model2.history['accuracy'])\nplt.plot(history_model2.history['val_accuracy'])\nplt.title('Model 2 Accuracy')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\nplt.plot(history_model2.history['loss'])\nplt.plot(history_model2.history['val_loss'])\nplt.title('Model 2 Loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\nplt.plot(history_model2.history['auc_1'])\nplt.plot(history_model2.history['val_auc_1'])\nplt.title('Model Two AUC ROC vs Epoch')\nplt.ylabel('ROC')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n","metadata":{"papermill":{"duration":1.552981,"end_time":"2022-11-24T02:40:51.793701","exception":false,"start_time":"2022-11-24T02:40:50.240720","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.254978Z","iopub.status.idle":"2023-02-17T13:54:29.255960Z","shell.execute_reply.started":"2023-02-17T13:54:29.255619Z","shell.execute_reply":"2023-02-17T13:54:29.255652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Results and Analysis\n\nFrom above plots for two models, we learned that both of them performed well. Since the second one has better accuracy and lower loss, we decide to choose the second one for prediction.","metadata":{}},{"cell_type":"code","source":"# test data\ndatagen_test = ImageDataGenerator(rescale=1./255.)\n\ntest_generator = datagen_test.flow_from_dataframe(\n    dataframe=test_data,\n    directory=test_path,\n    x_col='id', \n    y_col=None,\n    target_size=(64,64),         \n    batch_size=1,\n    shuffle=False,\n    class_mode=None)","metadata":{"papermill":{"duration":167.647781,"end_time":"2022-11-24T02:43:40.338338","exception":false,"start_time":"2022-11-24T02:40:52.690557","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.257446Z","iopub.status.idle":"2023-02-17T13:54:29.258670Z","shell.execute_reply.started":"2023-02-17T13:54:29.258336Z","shell.execute_reply":"2023-02-17T13:54:29.258379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#prediction with model 2\n\npredictions = model2.predict(test_generator, verbose=1)","metadata":{"papermill":{"duration":515.048541,"end_time":"2022-11-24T02:52:16.204514","exception":false,"start_time":"2022-11-24T02:43:41.155973","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.259978Z","iopub.status.idle":"2023-02-17T13:54:29.260611Z","shell.execute_reply.started":"2023-02-17T13:54:29.260293Z","shell.execute_reply":"2023-02-17T13:54:29.260325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create submission\npredictions = np.transpose(predictions)[0]\nsubmission_df = pd.DataFrame()\nsubmission_df['id'] = test_data['id'].apply(lambda x: x.split('.')[0])\nsubmission_df['label'] = list(map(lambda x: 0 if x < 0.5 else 1, predictions))\nsubmission_df.head()","metadata":{"papermill":{"duration":1.623376,"end_time":"2022-11-24T02:52:19.494337","exception":false,"start_time":"2022-11-24T02:52:17.870961","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.262217Z","iopub.status.idle":"2023-02-17T13:54:29.262875Z","shell.execute_reply.started":"2023-02-17T13:54:29.262559Z","shell.execute_reply":"2023-02-17T13:54:29.262589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Submission label pie chart\nunique_counts = submission_df['label'].value_counts()\nprint(unique_counts)\n\nplt.pie(unique_counts.tolist(), \n        labels = ['0-Non-cancerous','1-Cancerous'],\n        autopct='%1.2f%%'\n       )","metadata":{"papermill":{"duration":1.67007,"end_time":"2022-11-24T02:52:25.316537","exception":false,"start_time":"2022-11-24T02:52:23.646467","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.265003Z","iopub.status.idle":"2023-02-17T13:54:29.266538Z","shell.execute_reply.started":"2023-02-17T13:54:29.266177Z","shell.execute_reply":"2023-02-17T13:54:29.266208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Competition Submission\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"papermill":{"duration":2.154788,"end_time":"2022-11-24T02:52:28.829216","exception":false,"start_time":"2022-11-24T02:52:26.674428","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-17T13:54:29.267887Z","iopub.status.idle":"2023-02-17T13:54:29.268513Z","shell.execute_reply.started":"2023-02-17T13:54:29.268164Z","shell.execute_reply":"2023-02-17T13:54:29.268191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion\nCNN network model works very well in getting smaller training data without losing the critial features. With image normalization, less compution will be required. Regulation with dropout and batch normalization resolved overfitting problems.\nFuther augmentation techniques like resizing, flipping, rotating, cropping, padding, can help to address issues like overfitting and data scarcity, and they make the model robust with better performance.","metadata":{"papermill":{"duration":1.296736,"end_time":"2022-11-24T02:52:31.665903","exception":false,"start_time":"2022-11-24T02:52:30.369167","status":"completed"},"tags":[]}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}