{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Overview\n\nThis aim of this project is to improve cancer detection in lymph nodes by using computer vision machine learning technqiues. \nWe will examine the data given in this competition to get a better understanding of it. Once we are done examinig the pictures  we will run multiple convolutional neural network models which will help us to classify cancerous (1) from non-cancerous ones(0). \nThe issue with Cancer is that when it becomes detectable by then it is too late, hence recognizing and receving treatment faster is the key here.\nWith improved and faster cancer detection, there will be  faster treatments and less mortality. \n\nLayout for this notebook was given in assignment brief and is as follows:\n\n1.     Brief Description of the Problem and Data\n2.     Exploratory Data Analysis (EDA) - Inspect, Visualize and Clean the Data\n3.     Describe Model Architecture\n4.     Results and Analysis\n5.     Conclusion\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"https://github.com/mrjaiswa/Deeplearning-CNN-Cancer-Detect","metadata":{}},{"cell_type":"markdown","source":"# Layout\n\n    Brief Description of the Problem and Data\n    \n    Exploratory Data Analysis (EDA) - Inspect, Visualize and Clean the Data\n    \n    Describe Model Architecture\n    \n    Results and Analysis\n    \n    Conclusion\n","metadata":{}},{"cell_type":"markdown","source":"# CNN Breakout :","metadata":{}},{"cell_type":"markdown","source":"\nA Convolutional Neural Network (CNN, or ConvNet) are a special kind of multi-layer neural networks, designed to recognize visual patterns directly from pixel images with minimal preprocessing.\nThere are two main parts to a CNN architecture\n\n A convolution tool that separates and identifies the various features of the image for analysis in a process called as Feature Extraction. \n The network of feature extraction consists of many pairs of convolutional or pooling layers. \n A fully connected layer that utilizes the output from the convolution process and predicts the class of the image based on the features extracted in previous stages.\n This CNN model of feature extraction aims to reduce the number of features present in a dataset. It creates new features which summarises the existing features contained in an original   set of features. \n    \n  \nThere are three types of layers that make up the CNN which are the convolutional layers, pooling layers, and fully-connected (FC) layers. When these layers are stacked, a CNN architecture will be formed. \n1. Convolutional Layer\n\nThis layer is the first layer that is used to extract the various features from the input images. \nIn this layer, the mathematical operation of convolution is performed between the input image and a filter of a particular size MxM. \nBy sliding the filter over the input image, the dot product is taken between the filter and the parts of the input image with respect to the size of the filter (MxM).\n\nThe output is termed as the Feature map which gives us information about the image such as the corners and edges. Later, this feature map is fed to other layers to learn several other features of the input image.\n\nThe convolution layer in CNN passes the result to the next layer once applying the convolution operation in the input. Convolutional layers in CNN benefit a lot as they ensure the spatial relationship between the pixels is intact.\n2. Pooling Layer\n\nIn most cases, a Convolutional Layer is followed by a Pooling Layer. The primary aim of this layer is to decrease the size of the convolved feature map to reduce the computational costs. This is performed by decreasing the connections between layers and independently operates on each feature map. Depending upon method used, there are several types of Pooling operations. It basically summarises the features generated by a convolution layer.\n\nIn Max Pooling, the largest element is taken from feature map. Average Pooling calculates the average of the elements in a predefined sized Image section. The total sum of the elements in the predefined section is computed in Sum Pooling. The Pooling Layer usually serves as a bridge between the Convolutional Layer and the FC Layer.\n\nThis CNN model generalises the features extracted by the convolution layer, and helps the networks to recognise the features independently. With the help of this, the computations are also reduced in a network.\n\n\n3. Fully Connected Layer\n\nThe Fully Connected (FC) layer consists of the weights and biases along with the neurons and is used to connect the neurons between two different layers. These layers are usually placed before the output layer and form the last few layers of a CNN Architecture.\n\nIn this, the input image from the previous layers are flattened and fed to the FC layer. The flattened vector then undergoes few more FC layers where the mathematical functions operations usually take place. In this stage, the classification process begins to take place. The reason two layers are connected is that two fully connected layers will perform better than a single connected layer. These layers in CNN reduce the human supervision\n4. Dropout\n\nUsually, when all the features are connected to the FC layer, it can cause overfitting in the training dataset. Overfitting occurs when a particular model works so well on the training data causing a negative impact in the model’s performance when used on a new data.\n\nTo overcome this problem, a dropout layer is utilised wherein a few neurons are dropped from the neural network during training process resulting in reduced size of the model. On passing a dropout of 0.3, 30% of the nodes are dropped out randomly from the neural network.\n\nDropout results in improving the performance of a machine learning model as it prevents overfitting by making the network simpler. It drops neurons from the neural networks during training.\n\n\n5. Activation Functions\n\nFinally, one of the most important parameters of the CNN model is the activation function. They are used to learn and approximate any kind of continuous and complex relationship between variables of the network. In simple words, it decides which information of the model should fire in the forward direction and which ones should not at the end of the network.\n\nIt adds non-linearity to the network. There are several commonly used activation functions such as the ReLU, Softmax, tanH and the Sigmoid functions. Each of these functions have a specific usage. For a binary classification CNN model, sigmoid and softmax functions are preferred an for a multi-class classification, generally softmax us used. In simple terms, activation functions in a CNN model determine whether a neuron should be activated or not. It decides whether the input to the work is important or not to predict using mathematical operations.\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"!pip install --upgrade pip\n!pip install seaborn\n!pip install plotly\n!pip install scikit-image","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:30:16.737067Z","iopub.execute_input":"2023-06-20T13:30:16.737801Z","iopub.status.idle":"2023-06-20T13:30:55.745372Z","shell.execute_reply.started":"2023-06-20T13:30:16.737768Z","shell.execute_reply":"2023-06-20T13:30:55.744184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\nimport string\nwarnings.filterwarnings('ignore')\nimport numpy as np \nimport pandas as pd \nimport os\nimport random\nfrom sklearn.utils import shuffle\nimport shutil\n\n# Visualizations\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport matplotlib.patches as patches\n\n# Work with images\nfrom skimage.transform import rotate\nfrom skimage import io\nimport cv2 as cv\nfrom keras.preprocessing.image import ImageDataGenerator\n\n# Model Development\nfrom sklearn.model_selection import train_test_split\nimport tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.layers import RandomFlip, RandomZoom, RandomRotation\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, AveragePooling2D\nfrom tensorflow.keras.layers import Dense, Flatten, Dropout\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import BatchNormalization\nfrom tensorflow.keras.optimizers import Adam\nimport pandas_profiling as pp\nfrom tifffile import imread\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPool2D, Flatten, Dense\nfrom tensorflow.keras.wrappers.scikit_learn import KerasClassifier\nfrom sklearn.model_selection import GridSearchCV\nfrom keras.layers import Dense, Conv2D, MaxPool2D , Flatten, BatchNormalization, Activation","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:31:21.831199Z","iopub.execute_input":"2023-06-20T13:31:21.831636Z","iopub.status.idle":"2023-06-20T13:31:22.711096Z","shell.execute_reply.started":"2023-06-20T13:31:21.831599Z","shell.execute_reply":"2023-06-20T13:31:22.710096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Brief Run Down of the Problem and how this Data should be interpreted: \n\n\n The dataset contains the histopathological Images, each image is 96px * 96px with 3 channels.\n The training set contains 220,025 unique images and the test set contains about 57,500.\n Here we have two columns: 'id' which is the unique image ID correpsonding to the training directory, and 'label' which tells us the classification category.\n Each label is either a 0 or 1, depending whether the image is non-cancerous (0) or cancerous (1).\n In the competition description, we find that if at least one pixel of an image is identified as cancerous then the whole image is therefore marked with a 1, otherwise it is 0. \n No missing values in this data which will make preprocessing more efficient.","metadata":{}},{"cell_type":"markdown","source":"\n# **EDA**\nLoad Data","metadata":{}},{"cell_type":"code","source":"test_path = '../input/histopathologic-cancer-detection/test/'\ntrain_path = '../input/histopathologic-cancer-detection/train/'\nsample_submission = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\ntrain_data = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:31:50.315354Z","iopub.execute_input":"2023-06-20T13:31:50.315815Z","iopub.status.idle":"2023-06-20T13:31:50.803478Z","shell.execute_reply.started":"2023-06-20T13:31:50.315780Z","shell.execute_reply":"2023-06-20T13:31:50.802478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are some of the helper functions ","metadata":{}},{"cell_type":"code","source":"def basic_task(train,set_first):\n    if set_first == '1':\n        print(\"The number of Null Values in our dataset\")\n#         print(train.isnull().sum())\n#         print(train.isna())\n        if train.isnull().values.sum() > 1:\n            train.dropna(inplace=True)\n        print(\"This is how our data looks like\")\n#         print(train['Text'].head())\n\n\n        train.describe()\n        print(\"Number of articles that are unique\")\n\n#         print(train['ArticleId'].nunique())\n        profiler(train)\n        train.info()\n\ndef profiler(train):\n    profile = pp.ProfileReport(train, title=\"Pandas Profiling Report\", explorative=True)\n    profile.to_notebook_iframe()\n    profile.to_file(\"first_profile.html\")\n\n    \n\n    print(\"Generating Data Frame Profile\")","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:31:53.818117Z","iopub.execute_input":"2023-06-20T13:31:53.818512Z","iopub.status.idle":"2023-06-20T13:31:53.825838Z","shell.execute_reply.started":"2023-06-20T13:31:53.818481Z","shell.execute_reply":"2023-06-20T13:31:53.824847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"basic_task(train_data,set_first='1')","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:32:40.929186Z","iopub.execute_input":"2023-06-20T13:32:40.929588Z","iopub.status.idle":"2023-06-20T13:32:50.882569Z","shell.execute_reply.started":"2023-06-20T13:32:40.929556Z","shell.execute_reply":"2023-06-20T13:32:50.881460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plotter(train,hist):\n    if hist=='1':\n        \n        sns.countplot(x=train['label'], palette='colorblind').set(title='Label and Counts');\n        fig = px.pie(train_data, \n             values = train_data['label'].value_counts().values, \n             names = train_data['label'].unique())\n        fig.update_layout(\n            title={\n                'text': \"Label Percentage Pie Chart\",\n                \n                'xanchor': 'center',\n                'yanchor': 'top'})\n        fig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:33:03.197025Z","iopub.execute_input":"2023-06-20T13:33:03.197420Z","iopub.status.idle":"2023-06-20T13:33:03.204293Z","shell.execute_reply.started":"2023-06-20T13:33:03.197389Z","shell.execute_reply":"2023-06-20T13:33:03.203294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(pd.DataFrame(data={'Label Counts': train_data['label'].value_counts()}))\nplotter(train_data,hist='1')","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:33:04.298677Z","iopub.execute_input":"2023-06-20T13:33:04.299295Z","iopub.status.idle":"2023-06-20T13:33:04.632281Z","shell.execute_reply.started":"2023-06-20T13:33:04.299259Z","shell.execute_reply":"2023-06-20T13:33:04.631319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images_train = train_data\nimages_train['label'] = images_train['label'].astype(str)\nimages_train['id'] = train_data['id'] + '.tif'\nrows, cols = 6, 6\n\nfig, axes = plt.subplots(rows, cols, figsize=(6, 6))\n\nfor i in range(6 * 6):\n    image = imread(train_path + images_train['id'][i])\n\n    row, col = i // cols, i % cols\n\n    axes[row, col].imshow(image)\n    axes[row, col].axis('off')\n\nplt.subplots_adjust(wspace = 0.1, hspace = 0.3)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:33:23.797475Z","iopub.execute_input":"2023-06-20T13:33:23.797852Z","iopub.status.idle":"2023-06-20T13:33:25.677209Z","shell.execute_reply.started":"2023-06-20T13:33:23.797823Z","shell.execute_reply":"2023-06-20T13:33:25.676352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(5, 5, figsize=(15, 15))\nfor i, axis in enumerate(ax.flat):\n    file = str(train_path + images_train.id[i] )\n    image = io.imread(file)\n    axis.imshow(image)\n    box = patches.Rectangle((32,32),32,32, linewidth=2, edgecolor='r',facecolor='none', linestyle='-')\n    axis.add_patch(box)\n    axis.set(xticks=[], yticks=[], xlabel = images_train.label[i]);","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:33:39.717936Z","iopub.execute_input":"2023-06-20T13:33:39.718292Z","iopub.status.idle":"2023-06-20T13:33:41.812546Z","shell.execute_reply.started":"2023-06-20T13:33:39.718263Z","shell.execute_reply":"2023-06-20T13:33:41.811493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tpu = None\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError:\n    strategy = tf.distribute.get_strategy()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:33:53.562478Z","iopub.execute_input":"2023-06-20T13:33:53.562872Z","iopub.status.idle":"2023-06-20T13:33:53.586486Z","shell.execute_reply.started":"2023-06-20T13:33:53.562842Z","shell.execute_reply":"2023-06-20T13:33:53.585576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nRANDOM_STATE = 49\nBATCH_SIZE = 256\n# Set up GPU configuration\ngpus = tf.config.experimental.list_physical_devices('GPU')\nif gpus:\n    try:\n        tf.config.experimental.set_visible_devices(gpus[0], 'GPU')\n        tf.config.experimental.set_memory_growth(gpus[0], True)\n        print(\"GPU is set up successfully.\")\n    except RuntimeError as e:\n        print(e)\n\n# Define the train and validation data split\ntrain, valid = train_test_split(train_data, test_size=0.2)\n\n# Define the data generators\ntrain_datagen =  ImageDataGenerator(rescale=1./255.,\n                            validation_split=0.15)\n\ntest_datagen =  ImageDataGenerator(rescale=1./255.,\n                            validation_split=0.15)\n\n# Set up the data generators with GPU support\ntrain_generator = train_datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"training\",\n    batch_size=BATCH_SIZE,\n    seed=RANDOM_STATE,\n    class_mode=\"binary\",\n    target_size=(64,64))  \n\nvalid_generator = test_datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"validation\",\n    batch_size=BATCH_SIZE,\n    seed=RANDOM_STATE,\n    class_mode=\"binary\",\n    target_size=(64,64)) \n\n# Train your model using the data generators\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T14:26:01.158750Z","iopub.execute_input":"2023-06-20T14:26:01.159153Z","iopub.status.idle":"2023-06-20T14:41:24.395884Z","shell.execute_reply.started":"2023-06-20T14:26:01.159122Z","shell.execute_reply":"2023-06-20T14:41:24.394751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Points that influence CNN architectural design choices:\n\n--     CNN use filters to transform or extract features from the image/pixels by sliding through the image and applying some matrix operations at each level. \n       As this is mostly abstraction design choices are highly influenced by choice of filters(kernels) .\n\n--     Kernel size over the years has mainly converged to a 3x3 by the community. This seems to be very efficient as well as effective, and that is what we will be using\n\n--     Initial layer feature extractions are more broad and deeper layers become increasingly granular. This is so the model can learn different degrees of patterns in the image. This          is done by using the filter size and number of filters. \n    \n    \n--     Padding is used so as to preserve as much information as possible with each kernel. Without padding a 3x3 filter(kernel) would mean the outer edge of pixels would be lost. A 5x5        filter(kernel) would lose the outside 2 pixels around the entire image and so on.\n\n--     Striding is used to reduce complexity and simplify the features the model uses. \n       Striding is exactly what it sounds like, when the filter(kernel) traverses across and down the image it can do so one pixel at a time(stride=1), or in larger strides skipping            some pixels and abstracting away pixels in between along the way. This will significantly decrease the size/number of features/weights and make the model simpler and computation        much less intensive.\n\n--     Pooling can be used similar to striding wherein a specified number of pixels is abstracted into the max value in it's pool.The pool slides across and down the image applying this        transformation just like the convolution filter/kernel.\n\n--     Dropout can be used to keep a random set of weights hidden(set to zero) so that the model must optimize using less weights and therefore will typically generalize better at              inference time. We can apply dropout after any set of layers but it is typically done directly before or after a subsampling/abstracting layer such as a pooling layer. We can            choose the percentage of weights to keep hidden.\n\n--     Regularization is when we transform the data to keep the values in the same range and typically limit the range to [0, 1] for example. When Dropout is being used regularization          may not be unnecessary.\n\n","metadata":{}},{"cell_type":"code","source":"tpu = None\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError:\n    strategy = tf.distribute.get_strategy()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T14:42:29.060908Z","iopub.execute_input":"2023-06-20T14:42:29.061318Z","iopub.status.idle":"2023-06-20T14:42:29.067682Z","shell.execute_reply.started":"2023-06-20T14:42:29.061286Z","shell.execute_reply":"2023-06-20T14:42:29.066495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nROC_1 = tf.keras.metrics.AUC()\n\n# use GPU\ngpus = tf.config.experimental.list_physical_devices('GPU')\nif gpus:\n    try:\n        tf.config.experimental.set_visible_devices(gpus[0], 'GPU')\n        tf.config.experimental.set_memory_growth(gpus[0], True)\n        print(\"GPU is set up successfully.\")\n    except RuntimeError as e:\n        print(e)\n\n# Define the model\nmodel = Sequential()\nmodel.add(Conv2D(filters=16, kernel_size=(3, 3), activation='relu'))\nmodel.add(Conv2D(filters=16, kernel_size=(3, 3), activation='relu'))\nmodel.add(MaxPool2D(pool_size=(2, 2)))\nmodel.add(Conv2D(filters=32, kernel_size=(3, 3), activation='relu'))\nmodel.add(Conv2D(filters=32, kernel_size=(3, 3), activation='relu'))\nmodel.add(MaxPool2D(pool_size=(2, 2)))\nmodel.add(Flatten())\nmodel.add(Dense(units=256, activation='relu'))\nmodel.add(Dense(units=1, activation='sigmoid'))\n\n# Compile the model\nopt = Adam(learning_rate=0.0001)\nmodel.compile(optimizer=opt, loss='binary_crossentropy', metrics=['accuracy', ROC_1])\n\n# Set the batch size\nbatch_size = 256\n\n# Build the model\nmodel.build(input_shape=(batch_size, 64, 64, 3))\n\n# Display the model summary\nmodel.summary()\n\n# Train your model\n\n\n\nhist = model.fit(train_generator, validation_data=valid_generator, epochs=10)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T14:44:19.454145Z","iopub.execute_input":"2023-06-20T14:44:19.454542Z","iopub.status.idle":"2023-06-20T16:38:22.287648Z","shell.execute_reply.started":"2023-06-20T14:44:19.454509Z","shell.execute_reply":"2023-06-20T16:38:22.286678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predictions = model.predict(test_generator, verbose=1)\nmodel.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(hist.history['accuracy'])\nplt.plot(hist.history['val_accuracy'])\nplt.title('Model One Accuracy per Epoch')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2023-06-20T16:38:22.291196Z","iopub.execute_input":"2023-06-20T16:38:22.291503Z","iopub.status.idle":"2023-06-20T16:38:22.568712Z","shell.execute_reply.started":"2023-06-20T16:38:22.291476Z","shell.execute_reply":"2023-06-20T16:38:22.567770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Trying Different Model Here:","metadata":{}},{"cell_type":"code","source":"plt.plot(hist.history['loss'])\nplt.plot(hist.history['val_loss'])\nplt.title('Model Two Loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2023-06-20T16:38:22.571391Z","iopub.execute_input":"2023-06-20T16:38:22.571968Z","iopub.status.idle":"2023-06-20T16:38:22.873782Z","shell.execute_reply.started":"2023-06-20T16:38:22.571932Z","shell.execute_reply":"2023-06-20T16:38:22.872807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ROC_model2 = tf.keras.metrics.AUC()\n\nwith strategy.scope():\n    model_two = Sequential()\n\n    model_two.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu', ))\n    model_two.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\n    model_two.add(MaxPooling2D(pool_size=(2,2)))\n    model_two.add(Dropout(0.1))\n\n    model_two.add(BatchNormalization())\n    model_two.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\n    model_two.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\n    model_two.add(AveragePooling2D(pool_size=(2,2)))\n    model_two.add(Dropout(0.1))\n\n    model_two.add(BatchNormalization())\n    model_two.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\n    model_two.add(Flatten())\n    model_two.add(Dense(1, activation='sigmoid'))\n\n    #build model by input size\n    model_two.build(input_shape=(batch_size, 64, 64, 3))       # original image = (96, 96, 3) \n\n    #compile\n    adam_optimizer = Adam(learning_rate=0.0001)\n    model_two.compile(loss='binary_crossentropy', metrics=['accuracy', ROC_model2], optimizer=adam_optimizer)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T16:38:22.876125Z","iopub.execute_input":"2023-06-20T16:38:22.876513Z","iopub.status.idle":"2023-06-20T16:38:23.002086Z","shell.execute_reply.started":"2023-06-20T16:38:22.876471Z","shell.execute_reply":"2023-06-20T16:38:23.001195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hist_model2 = model_two.fit(train_generator, validation_data=valid_generator, epochs=10)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T16:38:23.003289Z","iopub.execute_input":"2023-06-20T16:38:23.003643Z","iopub.status.idle":"2023-06-20T18:26:44.394828Z","shell.execute_reply.started":"2023-06-20T16:38:23.003609Z","shell.execute_reply":"2023-06-20T18:26:44.393698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(hist_model2.history['accuracy'])\nplt.plot(hist_model2.history['val_accuracy'])\nplt.title('Model One Accuracy per Epoch')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:26:44.396495Z","iopub.execute_input":"2023-06-20T18:26:44.396880Z","iopub.status.idle":"2023-06-20T18:26:44.701400Z","shell.execute_reply.started":"2023-06-20T18:26:44.396850Z","shell.execute_reply":"2023-06-20T18:26:44.700365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(hist_model2.history['loss'])\nplt.plot(hist_model2.history['val_loss'])\nplt.title('Model Two Loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\n# plot model ROC per epoch\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(hist_model2.history['auc_1'])\nplt.plot(hist_model2.history['val_auc_1'])\nplt.title('Model Two AUC ROC per Epoch')\nplt.ylabel('ROC')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:26:44.703801Z","iopub.execute_input":"2023-06-20T18:26:44.704458Z","iopub.status.idle":"2023-06-20T18:26:44.755250Z","shell.execute_reply.started":"2023-06-20T18:26:44.704399Z","shell.execute_reply":"2023-06-20T18:26:44.754066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def create_model(learning_rate):\n    model = Sequential()\n    model.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\n    model.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\n    model.add(MaxPool2D(pool_size=(2,2)))\n    model.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\n    model.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\n    model.add(MaxPool2D(pool_size=(2,2)))\n    model.add(Flatten())\n    model.add(Dense(units=256, activation='relu'))\n    model.add(Dense(units=1, activation='sigmoid'))\n    model.compile(optimizer=Adam(learning_rate=learning_rate), loss='binary_crossentropy', metrics=['accuracy'])\n    return model\n\n# Create the KerasClassifier wrapper\nmodel = KerasClassifier(build_fn=create_model)\n\n# Define the hyperparameters grid\nparam_grid = {\n    'learning_rate': [0.001, 0.0001],\n    'batch_size': [128, 256],\n    'epochs': [5, 10]\n}\n\n# Perform grid search\ngrid_search = GridSearchCV(estimator=model, param_grid=param_grid, cv=3)\ngrid_result = grid_search.fit(train_generator, validation_data=valid_generator)\n\n# Print the best parameters and score\nprint(\"Best Parameters: \", grid_result.best_params_)\nprint(\"Best Score: \", grid_result.best_score_)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"resnet_model = Sequential()\nresnet_model.add(ResNet50(include_top=False, pooling='avg', weights='imagenet'))\nresnet_model.add(Dense(2, activation='softmax'))\n\nresnet_model.compile(optimizer=Adam(), loss='categorical_crossentropy', metrics=['accuracy'])\n\n# Train the ResNet model\nresnet_history = resnet_model.fit(train_generator, epochs=10, validation_data=valid_generator)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:27:01.275409Z","iopub.execute_input":"2023-06-20T18:27:01.276458Z","iopub.status.idle":"2023-06-20T18:27:01.288639Z","shell.execute_reply.started":"2023-06-20T18:27:01.276403Z","shell.execute_reply":"2023-06-20T18:27:01.287452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.DataFrame({'id':os.listdir(test_path)})\nsubmission_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:27:02.742297Z","iopub.execute_input":"2023-06-20T18:27:02.743238Z","iopub.status.idle":"2023-06-20T18:27:08.185618Z","shell.execute_reply.started":"2023-06-20T18:27:02.743203Z","shell.execute_reply":"2023-06-20T18:27:08.184452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.DataFrame({'id':os.listdir(test_path)})\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:32:28.106154Z","iopub.execute_input":"2023-06-20T18:32:28.107070Z","iopub.status.idle":"2023-06-20T18:32:28.154636Z","shell.execute_reply.started":"2023-06-20T18:32:28.107037Z","shell.execute_reply":"2023-06-20T18:32:28.153504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datagen_test = ImageDataGenerator(rescale=1./255.)\n\ntest_generator = datagen_test.flow_from_dataframe(\n    dataframe=test_df,\n    directory=test_path,\n    x_col='id', \n    y_col=None,\n    target_size=(64,64),         # original image = (96, 96) \n    batch_size=1,\n    shuffle=False,\n    class_mode=None)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:32:33.555431Z","iopub.execute_input":"2023-06-20T18:32:33.555809Z","iopub.status.idle":"2023-06-20T18:36:27.713261Z","shell.execute_reply.started":"2023-06-20T18:32:33.555779Z","shell.execute_reply":"2023-06-20T18:36:27.712300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model_two.predict(test_generator, verbose=1)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:36:32.658698Z","iopub.execute_input":"2023-06-20T18:36:32.659071Z","iopub.status.idle":"2023-06-20T18:48:08.224278Z","shell.execute_reply.started":"2023-06-20T18:36:32.659041Z","shell.execute_reply":"2023-06-20T18:48:08.223148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = np.transpose(predictions)[0]\ncopy_df = pd.DataFrame()\ncopy_df['id'] = submission_df['id'].apply(lambda x: x.split('.')[0])\ncopy_df['label'] = list(map(lambda x: 0 if x < 0.5 else 1, predictions))\ncopy_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:48:12.044903Z","iopub.execute_input":"2023-06-20T18:48:12.045276Z","iopub.status.idle":"2023-06-20T18:48:12.272206Z","shell.execute_reply.started":"2023-06-20T18:48:12.045247Z","shell.execute_reply":"2023-06-20T18:48:12.270928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"copy_df['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:48:18.848648Z","iopub.execute_input":"2023-06-20T18:48:18.849020Z","iopub.status.idle":"2023-06-20T18:48:18.857657Z","shell.execute_reply.started":"2023-06-20T18:48:18.848989Z","shell.execute_reply":"2023-06-20T18:48:18.856706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"copy_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T18:48:22.305897Z","iopub.execute_input":"2023-06-20T18:48:22.306306Z","iopub.status.idle":"2023-06-20T18:48:22.517814Z","shell.execute_reply.started":"2023-06-20T18:48:22.306271Z","shell.execute_reply":"2023-06-20T18:48:22.516691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n4. Results and Analysis\n\nWe can see from the above plots and diagrams for each model how well they performed with the training sets. \nWe see that model-one seemed to steady out a bit more than our more complex model (model-two) with regards to the ROC metric. \n\n","metadata":{}}]}